In this interview from theCUBE + NYSE Wired: AI Factories - Data Centers of the Future, Gilad Shainer, senior vice president of networking at NVIDIA, joins theCUBE's Dave Vellante to discuss Scale-In, NVIDIA's new networking pillar built for the demands of agentic AI. Shainer explains that agentic workloads chain together many inference calls across compute, storage and memory rather than a single query, requiring the full-stack Vera Rubin AI factory to keep pace. He details how BlueField-4 now manages 7.2 terabits of secure traffic, connecting to every ConnectX SuperNIC to secure east-west communication across GPU, CPU and storage servers rather than simply guarding perimeter access.
The conversation also explores how DOCA exposes BlueField's in-silicon security engines to developers without adding performance overhead, since telemetry collection happens entirely out of band. Shainer draws a historical parallel to how browsers, virtualization and DPUs each emerged to secure new computing models without slowing innovation, positioning OpenShell and BlueField-4 as the next step for containing autonomous agents. He touches on how BlueField-4 pairs the Grace CPU with the ConnectX SuperNIC to analyze telemetry across GPUs, memory and storage from a fully isolated device that agents cannot influence. From open-source access through DOCA to NVIDIA's DSX AI Factory reference designs, Shainer outlines how partners can adopt Scale-In infrastructure to keep pace with agentic AI without compromising security.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Gilad Shainer, NVIDIA
In this interview from theCUBE + NYSE Wired: AI Factories - Data Centers of the
Future, Gilad Shainer, senior vice president of networking at NVIDIA, joins
theCUBE's Dave Vellante to discuss Scale-In, NVIDIA's new networking pillar
built for the demands of agentic AI. Shainer explains that agentic workloads
chain together many inference calls across compute, storage and memory rather
than a single query, requiring the full-stack Vera Rubin AI factory to keep
pace. He details how BlueField-4 now manages 7.2 terabits of secure traffic,
connecting to every ConnectX SuperNIC to secure east-west communication across
GPU, CPU and storage servers rather than simply guarding perimeter access. The
conversation also explores how DOCA exposes BlueField's in-silicon security
engines to developers without adding performance overhead, since telemetry
collection happens entirely out of band. Shainer draws a historical parallel to
how browsers, virtualization and DPUs each emerged to secure new computing
models without slowing innovation, positioning OpenShell and BlueField-4 as the
next step for containing autonomous agents. He touches on how BlueField-4 pairs
the Grace CPU with the ConnectX SuperNIC to analyze telemetry across GPUs,
memory and storage from a fully isolated device that agents cannot influence.
From open-source access through DOCA to NVIDIA's DSX AI Factory reference
designs, Shainer outlines how partners can adopt Scale-In infrastructure to keep
pace with agentic AI without compromising security.
In this interview from theCUBE + NYSE Wired: AI Factories - Data Centers of the Future, Gilad Shainer, senior vice president of networking at NVIDIA, joins theCUBE's Dave Vellante to discuss Scale-In, NVIDIA's new networking pillar built for the demands of agentic AI. Shainer explains that agentic workloads chain together many inference calls across compute, storage and memory rather than a single query, requiring the full-stack Vera Rubin AI factory to keep pace. He details how BlueField-4 now manages 7.2 terabits of secure traffic, connecting to every Conne...Read more
exploreKeep Exploring
What makes agentic AI workloads different from traditional workloads, and how is infrastructure like the Vera Rubin "AI factory" (e.g., using scale-up NVLink) being designed to meet those requirements?add
How is GPU-based AI infrastructure scaled to support very large jobs, and what networking and storage technologies and approaches (Scale-Out, Scale-Across, Context Scale, and Scale-In) are used?add
What is the "Scale‑In" infrastructure and how does BlueField‑4 change the role of frontend/backend networks and security in AI data centers?add
How has AI infrastructure evolved from training-focused systems to support agentic AI, and why is a new "Scale‑In" approach required?add
How can you ensure comprehensive security when scaling AI models and deploying agents, and what role does the BlueField DPU play in providing out-of-band telemetry and protection without degrading GPU or storage performance?add
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, co-hosting theCUBE here with Dave Vellante, my co-host. Hi everybody, I'm Dave Vellante. Welcome to theCUBE, NYSE Wired's AI Factory series. AI networking has already gone through several major phases. NVLink gave NVIDIA a way to scale up GPUs into larger units of compute, and InfiniBand and Spectrum-X help scale those systems across racks and across AI factories. Now NVIDIA is introducing a new pillar, Scale-In. Agentic AI puts new stressors on the network. So moving data quickly between GPUs across racks and across AI factories, those represented major breakthroughs in networking. AI factories now have to connect users, agents, applications, enterprise data and storage while at the same time enforcing security, isolation, and policy. Of course, all this at machine speeds. NVIDIA is pushing more of that infrastructure work into BlueField-4, DOCA, and Spectrum-X Ethernet, effectively creating a new accelerated control and services layer around the AI factory. So the question is, is Scale-In simply the next evolution of the DPU, or is NVIDIA trying to define an entirely new infrastructure architecture for the agentic era. And to help us unpack that, Gilad Shainer is here, Senior Vice President of Networking at NVIDIA. Gilad, good to see you again. Thanks for coming in.
Gilad Shainer
>> Thank you. Happy to be here.
Dave Vellante
>> So what is it about agentic AI that necessitates a new type of network that you're calling Scale-In?
Gilad Shainer
>> Yeah, so first, agentic AI, it's probably the most complex workload ever created. When we look at an agent, agent is not just one inference call. There are many inference calls. It's not one query. There is a flow of operations and a flow of operations that includes compute and storage and memory and networking. And for serving Agentic AI, we're building Vera Rubin, which is a full-stack AI factory, which is a full combination or integrations of compute storage, memory, software, networking. It's a full AI factory. Now, building an AI factory, unlike traditional data centers that just require a single network to get in and out because workloads were just running on a single node, building an AI factory requires multiple infrastructures, purposely built infrastructure to create that AI factory, that single unit of computing. As you mentioned, we started with scale up NVLink and NVLink, it's actually to create a GPU unit.
Dave Vellante
>> Right.
Gilad Shainer
>> And after you have the GPU unit, you need to scale it out. And this is where we're using InfiniBand or Spectrum-X Ethernet. And you can scale it out to hundreds of thousands of GPUs. We also introduce scale across because sometimes the amount of GPUs you can put in a single location is not enough for running a large job. And you need to harvest the compute capabilities of multiple AI factories. And this is what we call what we do with Scale-Across. We also introduce the context scale infrastructure for context. As context is growing, not necessarily you have enough storage within or enough memory within the compute server for holding context and you need to actually have other means to hold context. And this is where we get an infrastructure where we're actually using BlueField as a storage controller to build a storage infrastructure for context. That's for...
Dave Vellante
>> Is that STX or?
Gilad Shainer
>> That's CMX and that's STX.
Dave Vellante
>> Yeah, okay.
Gilad Shainer
>> Exactly. And then recently we also introduced a new infrastructure which is called Scale-In. Now in the traditional data centers, there was typically only one network, which is the access network, also called north-south network. When you start building supercomputers in a sense, or the first AI systems, AI factories, we had the backend network and a frontend network, right? The backend network includes scale up, scale out, scale across. And there was a frontend network that was just providing access, right? Access of clients to servers. And it was called frontend because it was completely disconnected from the backend. That's what is the frontend and backend. In agentic AI, that's not the case anymore. There is no just frontend network. There is an infrastructure that needs to scale in modules and users and storage into that AI factory, but ensures security across everything, which means it's that element that was previously, if you will, called DPU in a sense, but previously was just handling access to the data center. Now it needs to make sure that everything is secure within the data center. So that BlueField, which is now BlueField 4, has full connectivity to the backend network. Ensures security when you access the network, but also it ensures security across the entire east-west infrastructure. So every BlueField-4 connects to every ConnectX NIC. So if in a previous generation BlueField was only covering the access network, which means there is a 400-gig pipe with secure access, now with BlueField-4, it's actually managed 7.2 terabit of secure traffic across everything. It's created connectivity to the CPU servers, it's created connectivity to the storage servers, it created connectivity to the GPU servers. Essentially, that scale in infrastructure is connected to all other infrastructures, manage all other infrastructure, secure all other infrastructures. And that's why it's no longer a front end. Now it's actually one of the core infrastructure that controls everything within the AI factory.
Dave Vellante
>> So Gilad, which part of the AI factory was doing that work before? Was it all in the CPU or partially the GPU? And is scale-in taking work off of what was previously the responsibility of that system?
Gilad Shainer
>> If you look at the infrastructure, those infrastructures evolve over time to cover things that are now needed and previously were not.
Dave Vellante
>> Okay.
Gilad Shainer
>> When you look at the first data centers that were built, AI factories that were built were focused on training. And when you focus on training, there is one job that runs across everything, right? This is where we invested in scale out and building an infrastructure that ensures no jitter, no hiccups in communication between GPUs. So everything works as a single unit. But in agentic AI, it's a completely different case because now we're dealing with many agents that need to run on the infrastructure. Now you need to increase security because you need to make sure that those agents are contained in their sandbox and are not doing wrong things. So now actually you need to build an infrastructure that needs to ensure security all around, from accessing to what happens on a scale-out, what happens in the memory, what happens on a storage. That requires a new infrastructure that was not needed when you only ran training. Okay, so this is how things evolve. And this is why now we needed to bring an infrastructure which actually enabled the security across everything. And that's what Scale-In is.
Dave Vellante
>> How does Scale-In sort of interact, integrate with NVLink, Spectrum-X, InfiniBand? And I presume the answer is you've extremely co-designed this, but maybe you could explain that.
Gilad Shainer
>> Yeah, of course. So BlueField, you can think about BlueField or BlueField-4 as a device that has in-silicon security engines that are working out of band and collecting telemetry from everything, from memory, from storage, from GPUs, from CPUs, from everything. And that's why that infrastructure connects to all other infrastructures and ensures security all around. When it connects to the CX NIC, to the ConnectX NICs, SuperNICs. It ensures that they control the encryption. They control the keys, for example. That out-of-the-band monitoring enables BlueField to actually understand exactly what agents are doing. Again, be able to enforce security and make sure that it knows what agent is allowed to do, make sure that an agent is doing what it's allowed to do, what the agent is doing in the memory, what storage the agent can access and so forth. BlueField-4 as a scaling infrastructure connects to everything, all other infrastructures. It actually covers all operations that are happening, storage, memory, compute, and making sure that everything works as it's supposed to work and agents are contained in their sandbox.
Dave Vellante
>> Is that what DOCA is? I'm not up to speed technically on DOCA. It's data center infrastructure on a chip. Can you explain that in a little bit more detail and how it affects developers?
Gilad Shainer
>> You have the BlueField devices, right? For example, you have the GPU devices and GPU devices give you the capabilities in silicon, but then you need to expose it to the workloads, the application running on top. This is what you do with CUDA, right? DOCA is the same thing for DPUs. DPU brings the in-silicon engines, in-silicon security engines, storage accelerations and so forth. And provides all those capabilities, but you need to expose that. And DOCA is the way that we expose that. So DOCA is the frameworks, the software, the APIs that enable you to access all of those engines that BlueField-4 provides.
Dave Vellante
>> So developers will access DOCA and it's optimized specifically for this purpose. Exactly. Yeah. Are there trade-offs or bottlenecks that you introduce and how do you manage those when you introduce scaling?
Gilad Shainer
>> Well, the issue that scaling solves is how do I make sure that there is security all around?
Dave Vellante
>> Yeah.
Gilad Shainer
>> Okay. This is the primary aspect of it. You need to scaling models, you're bringing agents, there is access. How do you control everything, right? And you could not control or provides security in the same place where agents are running. You need to do that from something that is completely not connected, completely isolated. What is the device that is isolated from where agents is running? BlueField, the DPU is that isolated device. And when you have enough compute in that BlueField, because there is a BlueField-4, with Grace in it. And when you have that compute capability in that, you can collect out-of-the-band information from the entire AI factory, telemetry and then be able to analyze that and ensure security all around. So actually BlueField is doing things that you cannot do in any other device within that AI factory.
Dave Vellante
>> It's a very powerful BlueField essentially acting as an offload. So that's why you don't have a performance degradation. Is that correct?
Gilad Shainer
>> There's no performance degradation because it works out of band. It does not impact the GPU work. It doesn't impact the storage work. It actually collects telemetry out of the band. So it does not impact the performance of the AI factory. But it ensures security all around.
Dave Vellante
>> I'm thinking about general-purpose computing, and if I recall, it must have been a while ago, maybe 2019, when you guys pointed out all the work that the CPU was doing for instance, storage or security. And it wasn't really offloaded. The burden was on the CPU. And that was a real memory management, real bottleneck for general-purpose computing. It sounds like you guys have thought that through and this is how this network is going to evolve, these networks. How do you see the portfolio evolving now? Because now that you've got this additional pillar. There's probably something else coming around the corner that you're thinking about, but how should we think about the future of that architecture and infrastructure?
Gilad Shainer
>> So the AI Factory continues to evolve. And it continued to evolve at a very fast pace to support the new workloads that are being created. Okay. In the early days of AI factories, it was focused on training. So at scale, you have to have good GPU units. And that GPU unit was an 8 GPU, for example, if you remember NVLink 8.
Dave Vellante
>> Yeah.
Gilad Shainer
>> And then we invested in scale out because we wanted to make sure that all the GPUs are fully synchronized. And you don't create jitters. There's no hiccups because if one GPU has a hiccup in data delivery, then all others are waiting.
Dave Vellante
>> We talked about jitter last time you came on.
Gilad Shainer
>> And NVIDIA continued to grow, right? So we went from 8 to 72. We talk about 576, we're talking about 1152. So that's continued to evolve. Scale-out is now interconnecting hundreds of thousands of GPUs, but now it's extended to scale across because you need to go beyond that, right? Storage was something that was critical for context scale for inference and now scale-in and BlueField-4 is enabling that security that was required because of agentic AI, because of agents. Things are continuing to evolve and we're on an annual pace of new AI factories, new generations. So, you should expect that there are going to be more infrastructure, more elements are going to be built in order to support what we want to do. Now, the concept, by the way, of what we're doing with BlueField-4 and scale-in, it's not a new concept. It's actually enabling the continuous development and moving forward fast, with technology cadence. And I can give you a couple of examples from the past.
Dave Vellante
>> Yeah,
Gilad Shainer
>> please.when we first introduced web browsers and web pages, right? That was a great thing because now you can access more information, more data out there. But when you start running pages on your computer, on your laptop, that introduced a security issue because now you're running code that someone else has written. And of course, you cannot count on that web developer to ensure that everything is secured. And that's why browsers were built because the browser is running those pages in a containerized environment. So now you can actually have a full secure system. And continue using the web. You didn't want to slow down the web. Does that make sense? When we built the data centers and we wanted to increase the number of jobs running on a server and not just having one user on a server, we wanted to have many users on a server. This is where we introduced virtualization. And virtualization enabled a single server to run multiple users. We didn't count on those users to make sure that everything is secured, right? Because you cannot do it. You cannot do that. And that's why we created hypervisors. And hypervisors are runtime environments running underneath the users or between the users and the hardware to make sure that there is security. But you don't want that security to be done on the same device that hosts the users, right? That's why we created DPUs. And DPUs actually offload elements of the hypervisors to a completely isolated device, separating the application domain from the infrastructure domain. And now you can secure your infrastructure and you continue running user workloads, right? The goal is not to slow down technology. You want to increase the pace of technology. On the other side, you want to bring the right security that enables that cadence. Same thing happening now, 'cause now we have agents and you cannot make sure, or you cannot enforce the agents or make sure the agents are behaving Correctly. That's why you need to bring the infrastructure that will contain it, that will bring security. That's what we're doing with NVIDIA OpenShell, because OpenShell is a runtime that runs underneath the agents, that controls where the agents are going. Now, do I want to have everything there? It's a great infrastructure. The answer is no. I want to have also another device which is completely separated, completely isolated from the agents. That can monitor everything out of band. That's what we're doing with BlueField-4 bringing all the security elements together with OpenShell. And now you have a full secure infrastructure for agents. The same concept as the past, but now doing that for agents to make sure that we continue running with the technology, continue to develop at a speed of light pace because there's other things we want to do.
Dave Vellante
>> So this is a natural evolution I'm hearing of the networking architecture, the DPUs, and it coincides with the workload of Agentic. And if I heard you correctly, you said you're using Grace?
Gilad Shainer
>> Yes, so BlueField-4, it's a combination of the CX-9 SuperNIC and Grace CPU, fully integrated in that sense. And that enables that device to access all networking infrastructure because it has a networking device. It can do all the DMA operations, it can access the memory, can access the storage, collect all the telemetry and so forth. And then it has Grace with it because you want to run everything around it, right? You want to get the understanding of what's happening. You want to make sure that if you need to go and contain things, change things, then you can do that on the device. Now, in the future, we may also connect the GPU to that because then you can actually even further follow up exactly what the agents are doing. So you have a very powerful device which has the best NIC inside and the best CPU inside. And that actually forms BlueField-4.
Dave Vellante
>> Well, and that Grace CPU that people think the lifetime of silicon is limited. But here's an example where Grace is more than powerful enough and it's more cost effective than, let's say, using Vera, right? And so people, I think, forget all the different applications that can work. You mentioned OpenShell before. Security, obviously a hot topic with open source, open weights. I wonder if you could expand on that and add some additional color as to how you guys are thinking about OpenShell and security.
Gilad Shainer
>> Yeah, of course. OpenShell is basically open source because you actually want to see what's inside. You want to know what's inside. And OpenShell is a runtime. It's a runtime that was built in order to bring security to agents. And as we talked, you cannot count on an agent to contain itself. You need some infrastructure that will contain agents, right? Will ensure that they're doing what they can do with the permissions that they have, right? They don't go and do things that they're not permitted to do. This is where we build OpenShell. And OpenShell is a runtime. It sits underneath the agents. It makes sure that the agents are doing what agents are doing and are not doing what they're not supposed to do. And OpenShell is a great infrastructure, that great platform that we built, and it can run on every platform that people are using, based on NVIDIA technology, obviously. But alongside OpenShell, we also wanted to have a device which is completely isolated and can collect all the information out of the GPU, has in-silicon security engines and actually can do the enforcement and can monitor everything from an isolated device that no one can access, no agent can influence. And this is how the combination of OpenShell and BlueField-4 brings all the security elements that we need for GenAI.
Dave Vellante
>> Doing a lot of work and it's making the AI factory more secure. In this era. Last question is, are there any prerequisites for customers or partners? Is it just a matter of sort of leaning into DOCA? What do they have to do to take advantage?
Gilad Shainer
>> So everything that we build, of course, is accessible to everyone, open source, it's open source. And NVIDIA provide a great open source infrastructure. BlueField-4 comes with all the in-silicon security engines. And with DOCA, DOCA exposes all the APIs, all the information out there. We have working with a great partner of us. There is a great ecosystem of partner from us that actually leverage the technology that we build, leveraging the reference architecture that we create, and leveraging our DSX AI Factory because we actually building a full reference design, a full digital twin of everything we build. So you can actually run, explore, and be part of the ecosystem that we build.
Dave Vellante
>> Gilad, thanks so much. You're doing some amazing work. I really appreciate you coming back into the New York Stock Exchange studio here.
Gilad Shainer
>> Yeah, happy to be here.
Dave Vellante
>> Okay. Thank you for watching. This is Dave Vellante for theCUBE's AI Factories series at the New York Stock Exchange. Keep it right there for more great content from NYSE.