This discussion examines Cisco Secure AI Factory with NVIDIA and its approach to accelerating production artificial intelligence, abbreviated as AI, at rack scale. The conversation highlights rack-scale reference architectures, integration of Spectrum-X into Cisco systems and validated AI factory deployments for data center and AI infrastructure teams.
Will Eatherton of Cisco, senior vice president of data center, internet and cloud infrastructure engineering; Gilad Shainer of NVIDIA, senior vice president of networking; and Marc Hamilton of NVIDIA, vice president of solutions architecture and engineering participate.
The discussion covers Cisco Silicon One and Nexus management and innovations in NVIDIA Spectrum-X and remote direct memory access, abbreviated as RDMA. It also examines NCP reference architectures and validated designs, liquid-cooled HGX and MGX systems and operational practices to reduce time-to-first-token.
Shainer explains that Spectrum-X and lossless fabric techniques reduce jitter and increase tokens per watt. Eatherton announces that Cisco Secure AI Factory is orderable in September, expanding liquid-cooled compute and integrated go-to-market services. Hamilton emphasizes the NCP reference architecture and exemplar benchmarks as essential for availability, performance and predictable time-to-revenue.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Cisco Secure AI Factory with NVIDIA Expands to Rack Scale. If you don’t think you received an email check your
spam folder.
Sign in to Cisco Secure AI Factory with NVIDIA Expands to Rack Scale.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for Cisco Secure AI Factory with NVIDIA Expands to Rack Scale
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for Cisco Secure AI Factory with NVIDIA Expands to Rack Scale.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Cisco Secure AI Factory with NVIDIA Expands to Rack Scale. If you don’t think you received an email check your
spam folder.
Sign in to Cisco Secure AI Factory with NVIDIA Expands to Rack Scale.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to Cisco Secure AI Factory with NVIDIA Expands to Rack Scale
Please sign in with LinkedIn to continue to Cisco Secure AI Factory with NVIDIA Expands to Rack Scale. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Exclusive News: Cisco Secure AI Factory with NVIDIA Expands to Rack Scale
This discussion examines Cisco Secure AI Factory with NVIDIA and its approach to accelerating production artificial intelligence, abbreviated as AI, at rack scale. The conversation highlights rack-scale reference architectures, integration of Spectrum-X into Cisco systems and validated AI factory deployments for data center and AI infrastructure teams.
Will Eatherton of Cisco, senior vice president of data center, internet and cloud infrastructure engineering; Gilad Shainer of NVIDIA, senior vice president of networking; and Marc Hamilton of NVIDIA, vice president of solutions architecture and engineering participate.
The discussion covers Cisco Silicon One and Nexus management and innovations in NVIDIA Spectrum-X and remote direct memory access, abbreviated as RDMA. It also examines NCP reference architectures and validated designs, liquid-cooled HGX and MGX systems and operational practices to reduce time-to-first-token.
Shainer explains that Spectrum-X and lossless fabric techniques reduce jitter and increase tokens per watt. Eatherton announces that Cisco Secure AI Factory is orderable in September, expanding liquid-cooled compute and integrated go-to-market services. Hamilton emphasizes the NCP reference architecture and exemplar benchmarks as essential for availability, performance and predictable time-to-revenue.
play_circle_outlineSpeed and time-to-first-token as critical for revenue generation
replyShare Clip
play_circle_outlineLossless networking, RDMA, adaptive routing, congestion control for performance
replyShare Clip
play_circle_outlineCisco and NVIDIA Secure AI Factory: Scalable rack-scale liquid-cooled Blackwell/Vera Rubin HGX and MGX for NeoClouds, enterprises, and sovereign clouds
replyShare Clip
play_circle_outlineCisco Validated Infrastructure Services (CVIS) and go-to-market integration
replyShare Clip
play_circle_outlineNCP Reference Architecture and NVIDIA Cloud Partner certification importance
replyShare Clip
play_circle_outlineRepeatable validated designs, automated provisioning and validation at scale
Exclusive News: Cisco Secure AI Factory with NVIDIA Expands to Rack Scale
Gilad Shainer
SVP of NetworkingNVIDIA
Will Eatherton
SVP, Data Center, Internet & Cloud Infrastructure EngineeringCisco
Marc Hamilton
VP, Solutions Architecture and Engineering at NVIDIANVIDIA
search
(INTRO)
John Furrier
>> Hello and welcome to theCUBE studios in Palo Alto, California. I'm John Furrier, host of theCUBE. We have experts here from Cisco and NVIDIA to discuss the advances in rack -scale AI factories and the role of networking. As the criticality of deploying AI workloads continues to surge as production becomes the goal, large -scale AI factories become the most important trend in computer history as there is a race to fill the worldwide demand. NVIDIA and Cisco have come together to deliver Cisco secure AI factories with NVIDIA, expanding rack scale systems to the next level. We have here Cisco's SVP and head of networking engineering, Will Eatherton from NVIDIA, SVP, and CUBE alum, Gilad Shainer. Welcome back. And VP of Solutions Architecture and Engineering, Marc Hamilton, also from NVIDIA. Gentlemen, thank you for joining us here today for this announcement. Will, we'll start with you. Everyone's racing, as I mentioned, to fulfill the demand. We're seeing the rise of the neoclouds. They're expanding super fast. They have different approaches. Sovereign programs are on the horizon being architected now. And big enterprises are trying to figure out how to bring in the AI intelligence into their workflows. What's actually broken between ordering the systems, the GPUs, the compute, and actually getting all those systems into production?
Will Eatherton
>> Sure. I think we're seeing across all three of those domains, as you mentioned, convergence right now. So when we work with NeoClouds, from the moment that they put in the PO for the GPU, they already have their end customers lined up. And one of the broader challenges is just the speed, the expectation that from the moment that that is on the project plan, that the GPUs have to be racked, that that then has to go through all of the final deployment aspects, the software, and then a bring up and hand over to their end customer. So it's a rush. It's a rush to revenue for them. enterprises, many of them, actually, as a software development leader, I'm one of them, are spending a large amount on tokens right now. And the rush and the pressure is getting these systems up so they can start offloading what has been an API interface into using local inference. And so I think it's all converging on common architectures, common systems, but speed. I think speed is what is the broader challenge.
John Furrier
>> Mark, what's your take? The global view from NVIDIA has always been, let's get these bottlenecks fixed, let's deploy, let's make sure people have everything that they need. What's happening on your end?
Gilad Shainer
>> Well, first of all, to put this in perspective, the build-out of AI infrastructure is really the largest scale infrastructure build-out that the world has seen ever. people talk about the building of the transcontinental railroads or the U.S. interstate highway system. This is on par, if not greater, than that. And so, So this is where speed comes in. Everyone is rushing to adopt AI and to start to make money with AI. Traditionally, your servers, your networking to an enterprise, it's been a cost, and you've tried to reduce your costs. But factories, factories produce revenue, and this is what an AI factory does. It takes in raw materials, takes in data or tokens, and it takes in electricity. It converts that raw material using AI infrastructure into a finished good, into tokens, and those tokens have value. So it's not just about buying GPUs and having those GPUs delivered or having the network switch delivered. It's about having that AI factory from start to finish delivered and how quickly you can start producing tokens and revenue. And this is where having a standard reference architecture that the entire ecosystem can build to is important. And to me, it's really almost impossible to think about taking on the largest infrastructure project in the world without having Cisco at the table. And so it's great to have now this Cisco rack scale solution that they're bringing to market.
John Furrier
>> Gilad, weigh in on this, because we've seen everyone wanting the silicon, they want the systems, but now we're in the execution phase, delivering, standing up, and getting revenue. The faster you can get those tokens, the faster you can get those applications enabled, this becomes the focus now, operational excellence and execution at the highest quality.
Gilad Shainer
>> Yeah, building a factory, it's not connecting components and hoping for the best. Building AI factory means that you need to build a supercomputer and a supercomputer that needs to be built quickly, needs to be built fast, and needs to provide the highest numbers of tokens per second, the highest numbers of tokens per watt, and so forth. when we start building the infrastructure we start building a Spectrum-X Ethernet for example we looked into building a purpose-built infrastructure for ai purpose-built for supercomputers that needs to run distributed computing workloads so we focused on eliminating jitter for example we brought elements of lossless into scale out infrastructure we brought abilities of NICs and switches to work as one single infrastructure to eliminate Jitter with adaptive routing and RDMA, congestion control, and so much hardware capabilities. And it's a great platform. And we work with Cisco. We made an announcement, I think in October of last year, of Cisco leveraging the technology that we build. And now we're at a point where we see that or we bring to our mutual customers the capabilities of Spectrum-X into something that they're used to already. So Spectrum-X, the technology that we build, or Spectrum-X inside the Cisco systems, means that now there is an AI-optimized fabric that comes with the operating model that the enterprise world already runs on. Okay, that's what we're able to do together. And this is why this time is so important for our mutual customers.
John Furrier
>> Will, I want to get into the announcement. He mentioned distributed computing. I mentioned it's a historic moment in the industry. We've never seen this kind of scale and execution with computing. Rack scale systems, they're being built, they're being operated and invested in continuously. These cycles look like they're not going to stop for a long time. So let's get into what Cisco is announcing today. Cause you mentioned the previous announcements. I was in DC. I think we chatted about it then. Now we have the new secure rack scale system. What is Cisco actually announcing? Gotcha.
Will Eatherton
>> So as Gilad mentioned, we've been working together on getting all the networking components for the past year. So we have the Spectrum-X silicon with the Cisco software on top. We have the Cisco Silicon One switch architecture that can have the Spectrum-X license. We've been getting all those pieces together, working together on the reference architecture prep. What we're announcing today is that across the Sovereign Enterprise and NeoClouds, that we need to go big. And so we are going broader with compute. So that is liquid cooled, starting with Blackwell, moving to Vera Rubin, liquid cooled on the HGX form factor, MGX form factor. So that gives us a breadth of compute systems that we can then wrap around from a Cisco standpoint, from a sales support and a software standpoint. And so it's all those key components that we're announcing as far as the new products. This is then a solution that we've been working on. I partnered with Gilad, Mark, and the rest of the NVIDIA team on putting together how this is going to work for an NCP RA compatible construct. So we've worked through what that whole cluster looked like, how we're going to support that, and put together, for instance, a Cisco CVIS, Cisco Validated Infrastructure Services, that will help essentially be a front end across a broad range of enterprises, sovereign, neoclouds that are around the world. and then work with Mark and the NVIDIA team behind us so that we can help NVIDIA scale so that we can help our customers get these outcomes. And this full solution is going to be orderable in September. And so the Secure AI Factory is going rack scale.
John Furrier
>> So there's a technology partnership on the silicon and the certification, but it's also a go-to-market partnership with Cisco's validated infrastructure service, which is what, an end-to-end? And that's the motion. So take us through the two parts. Take us through the partner go-to-market piece.
Will Eatherton
>> Right. So Cisco will be selling the full solution. So that is the key aspects that will be the compute solution. We then will do the pre-sales work. We'll work with both our Cisco professional services, but with partners like WWT and Computacenter and so on around the day zero, day one activities. We then sell the networking. And then we're also solving for customers that are interested in the NCP RA. We will attach with that support from a Cisco CVD so that we're able to do a certification that it's meeting the performance standards that we are looking for. But also NVIDIA has put very clear guardrails and construct around.
John Furrier
>> So Gilad mentioned Spectrum-X. So Cisco Silicon One is what, the front end and NVIDIA Spectrum-X? So, yeah.
Will Eatherton
>> Silicon on the back end? Right. So the key architecture points would be Cisco, Silicon One running, for instance, NX-OS or SONiC on the front end. And we have cases where customers want, for instance, high performance with storage. There's an option for a Spectrum-X license, which then is part of that overall ecosystem that Gilad mentioned. Then on the back end, it's Spectrum Silicon running, again, NX-OS or SONiC. And then overall, that is a common management system. We have customers that for instance, also are doing more with routing interconnect of multiple data centers, a DCI data center interconnect. So everything from the routing, the front end, the storage networking and the backend, again, working seamlessly with storage partners like VAST, and then with the compute that we're pulling in, both the box level and the rack scale. And so, yeah, those are all the key components, John.
John Furrier
>> Mark, I want to bring you into this because Gilad mentioned building a supercomputer. this is great from the AI standpoint, but it's hard, right? So talk about the reference architecture and the NCP, NVIDIA Cloud Partner certification. How important is this certification in this era? Because these rack scale systems are not for the faint of heart, they are very complex, but they produce a lot of value. They want them fast. In rack scale, this is a huge thing. Explain the importance of this program.
Gilad Shainer
>> Well, the NCP reference architecture is really so important because an AI factory or AI data center is a five -layer cake, right? It's not just about the chips and the networking, right? It starts with your land power shell, your data center, then it's, of course, the chips, the networking chips, the compute chips. It's your AI infrastructure. This is where the servers, networking, and storage come together. It's your AI models and it's your AI applications. If you don't build that whole thing end to end and have a place where you can test every application, every model, every network architecture, it's nearly impossible to optimize it and make sure that it is the right solution. So through our NCP reference architecture and our enterprise reference architectures, this is exactly what we do. I think the fact that Cisco is taking this to market as one solution that they can sell and the customers can order is also important, especially for AI factory. Traditional enterprises had server teams and had networking teams. And the server team went off and bought servers. The networking team went off and bought networking. And the two didn't come together until very late. In an AI factory, because it's a five-layer cake, everything has to work together. There are so many mistakes when customers try to go to one vendor to order networking, another vendor to order servers. So I'm really excited to see this Cisco solution. You talked about also then why is it hard? You get all this infrastructure delivered, a bunch of pallets on your loading dock. And that time to first token, when you can start generating revenue is important. Two other things that are equally important that the reference architecture helps with. If you get your first token ready, but your availability is not good, You've got a $10 million or $10 billion AI factory, and 30%, 40%, 50 % of it is unavailable because it wasn't cabled right, because there was some problem with software, firmware wasn't loaded correctly, that's going to drive up your costs tremendously. The second thing, besides availability, is, again, the performance. You need availability and performance. And because NVIDIA in our reference architecture integrates all five layers of that stack, tunes across all those five layers, we can guarantee the best performance. And in the case of NVIDIA Cloud Partners, in fact, they run our, what we call, Exemplar Cloud benchmark suite, that they can then go out to the public and demonstrate that, show the results, compared to our gold standards that we run in our labs ourselves.
John Furrier
>> Gilad, talk about the end-to-end, when people bring their own networking to the delivery model, and maintaining the level of quality it takes. It's not like the old days, it's a whole new modern era. What's your view on this?
Gilad Shainer
>> Yeah. First, remember that what we build is a single unit of computing. And that single unit of computing that generates tokens is built on a huge amount of components. You're talking about hundreds of thousands of GPUs. You're talking about many switches, many cables, storage, NICs. There are so many elements. and infrastructure from scale up to scale out to access network and so forth, there are so many elements and still it needs to work as a single unit. And when neoclouds and enterprises buy AI factories, this is an important investment for them. That's actually how they're going to make money. That's how they're going to generate money. So they're going to make an important investment They're going to make a big investment, and they expect that all of those hundreds of thousands of components are going to work as one entity and they have to. That's what they expect. So the complexity here is huge, and that's why we build everything which is purposely built for AI. Every component is purposely built to generate tokens because that's how you generate money. And that's why the NCP RA reference is so important because that makes sure that they can bring the things together, deploy it very quickly and get tokens and make money. So now bringing that technology that we created and working together with Cisco, and now Cisco can deliver everything and unify the management of that big entity, which is a single entity of computing, if you think about it, managing that through Cisco Cloud Control and leveraging all the elements that NVIDIA brought inside. That's the importance of what's happening today, by the way. And this is where we're so happy to work with Cisco, as Will said, for quite some time. In a sense, it's a few months because the world is moving very, very fast. It is a lot of time, believe me. But there is an amazing collaboration between the two companies. Now Cisco can bring it all together with the same management that enterprise knows how to run, the same management that they're already running. So they can move quickly. they can leverage what they acquired very quickly. And then Cisco brings that with all of the NVIDIA technology inside, with all the capabilities to generate tokens, with all the capabilities to leverage AI and enable our mutual customers to move to the next level of their businesses.
John Furrier
>> Yeah, it's well said. I love how it brings conventional distributed computing to the artificial intelligence modern era. I guess I'd have to ask you, so you mentioned customers. So a Neo cloud, an enterprise, or a sovereign approach deploying the Cisco architecture is getting the same standard NVIDIA holds its own cloud and partners to. Is that right?
Gilad Shainer
>> They're getting the same performance as we make sure that they're going to get from the technology side. That's the reason that we build the reference architecture to make sure that they can get the full performance of NVIDIA components. And that guarantee is now coming with the Cisco AI Factory. That's how both of us guarantee that they're getting the best performance out of their investment. And one important element is when we build the InfiniBand switches, when we build Spectrum-X Ethernet, when we build those elements, we wanted to make sure that we have the full flexibility so you can actually be able to run your own things on top of that. So in that case, for example, Cisco can take their Cisco NX-OS and run them on the open interfaces that we brought into SpectrumX. Or they can take SpectrumX licenses and be able to connect it to the access network with Silicon One switches and so forth. So we brought the capabilities, we brought the technology, it's built in a fully open interface so you can go and customize that. And now everything's now working together as a single unit coming from Cisco.
John Furrier
>> Will, weigh in on the NCP because this is a major win for customers.
Will Eatherton
>> Yeah, so I want to mention it hasn't always been easy. So I think as we are working through and as Gilad mentioned, I think from an architecture standpoint, it was pretty clear how to approach it. But for instance, in working with Mark and his team, there's been key areas. For instance, they have a large array of scripts and technology that they've developed around the stack that is the default stack for NVIDIA. So we've been working as part of the Cisco Validated Designs, the CVDs, to take those and then make the tweaks in those tools for our Nexus, for our management layer. And so I think, again, bringing the core aspects of what NVIDIA has been bringing to market and then being able to adapt that so we can support our joint end customers is a big deal. And, I appreciate that Gilad mentioned the Cisco Cloud Control and really the experience Cisco has working with, telcos and enterprises in managing software and really this day two topic. And this is something, Mark and I talk a lot about, a lot of the industry focus is up to the point that you light up the cluster and you get your first token out. That's been a big focus. That's great. But the day two aspects around monitoring and health and availability and software upgrades, and these are things that from a Cisco standpoint, we've put a lot of focus on here, over the years. And as we get into AgenticOps that we're building into Cisco Cloud Control, the outcome of that is that, again, from whether it's routing and data center interconnect, whether it's front end networking storage networks, the front end, maybe an enterprise that is working in EVPN and VXLAN and other technologies to the backend with a Spectrum ASIC, a Spectrum-X technology, NX-OS or SONiC, and then integrating that into our overall management that then can be managed software life cycle and monitoring all as part of one cohesive system. And that is the big thing that we're excited to be bringing to our customers.
John Furrier
>> Mark, weigh in on this because there's a lot of stuff going on between Cisco and NVIDIA. He mentioned cloud control. There's also a lot of other Cisco products involved. Silicon One, and you got Spectrum-X. Talk about some of these reference archi, validated designs and reference architecture. And how do you get these repeatable provisioning and validation automation stuff done? There's a lot of implementation going on here.
Gilad Shainer
>> every customer that starts to build an AI factory starts off thinking, oh, I've built data centers before. I know how to do this and I can do it better. And in fact, even in the supercomputing space in the past, if you were the national labs or if you were a large university, that was part of the experience to go through and do a custom design and try to optimize on a little bit faster CPU or a little bit less memory to save money. And again, the real cost savings in an AI factory is not about cost savings, but it is about token generation and how do you drive that revenue. And so having a repeatable way to do it. The other great thing about this for Cisco's enterprise customers is, remember, most enterprises today want some sort of multi -cloud hybrid type of solution. So, being able to have the Cisco AI Factory on -prem and then being able to go to a neocloud and being able to access the same exact solution in any neocloud around the world running the Cisco AI Factory is a huge win and a huge advantage for Cisco's enterprise customers.
John Furrier
>> Yeah, and I think it's not your old playbook of building data centers. These are systems. Deployments have to be done right and they have to be done fast. and again, time to first token is revenue. And that's unique in this era, fast revenue producing. All right guys, final thoughts for all three of you. when these systems are in production across neoclouds, sovereign environments and enterprises, what's the one thing the market should remember? Gilad, we'll start with you.
Gilad Shainer
>> It's a good question. It's a good question. So what we're announcing, First, it's a great partnership. So people should remember October 25. This is when Cisco announced the Spectrum-X switches. And people should remember August 26, when we're announcing, or Cisco is announcing bringing a full AI factory, which is NCP compliant and enabling neoclouds and enterprises, for example, to continue running the same way they're running their traditional data centers. But now, they're going to run full AI factories and generate tokens and drive their businesses. So there's a great technology. There is a great roadmap ahead of us. And now that things are moving to production and people can go in and acquire from Cisco a full AI factory and just, connect it to the power and be able to generate tokens and make money. The only thing that they need to remember when they go back home is to buy milk in a supermarket. markets.
John Furrier
>> That'd be good. Mark, this is a market focus. There's a lot of headroom, a lot of things being done now. It's moving super fast. What's the one thing Mark, we should remember?
Gilad Shainer
>> Well, the one thing I would remember is that an AI factory is a living, breathing thing. Every week, there's a new AI model that's launched. And of course, recently, there's been huge interest in open AI models, be it NVIDIA's own Nemotron model or other open models. And so really building an AI factory is not just about that time to first token. We call it continuous bring -up optimization and upgrade because you install the AI factory, you get that new AI model that is launched and you want to optimize AI factory for that and then by the time you're doing that you then need to go through and do an upgrade. So taking Cisco's enterprise knowledge, Cisco's tools, and applying it to an AI factory for this continuous bring up, optimize, and upgrade motion is going to be a real market mover, I believe.
John Furrier
>> So you're saying the software stack is pretty critical because that's going to be the pacing item to keep the capabilities coming in. Every week there's a new model released. All right, Will, close us out. Cisco, why now? Enterprise market, we expect to be booming, and the neoclouds are building out as fast as they can. Obviously, we know what the hyperscalers are doing. This is distributed computing.
Will Eatherton
>> So I'm going to end with scale. This announcement is all about scale on multiple fronts. From a Cisco standpoint, it's about scaling, scaling the size of the system, the rack scale system, scaling the liquid cooling, scaling just the magnitude and the amount of clusters that we're going to be able to go out and work with our customers. This is about scale and how we're partnering with NVIDIA. You being able to wrap around the NCP RA, wrap around the NVIS with the Cisco Validated Infrastructure Services and help scale globally across a broader range of neoclouds, telcos, enterprises. It's about how do our customers scale? So, beyond first token, how do you scale the size of your clusters, the number of clusters, how do you scale the operations so that you get the uptime and the token accumulation, not only day one, but over time. So it's all about scale.
John Furrier
>> Gentlemen, thank you very much. Secure AI Factory with NVIDIA, Expanding Rack Scale. Thanks so much for taking the time and weighing in with your expert opinion. We've got a mixture of experts here. Thanks for coming on. You're welcome. Thank you, Mark, Will. Thanks. Thank you.
Exclusive News: Cisco Secure AI Factory with NVIDIA Expands to Rack Scale
search
(INTRO)
John Furrier
>> Hello and welcome to theCUBE studios in Palo Alto, California. I'm John Furrier, host of theCUBE. We have experts here from Cisco and NVIDIA to discuss the advances in rack -scale AI factories and the role of networking. As the criticality of deploying AI workloads continues to surge as production becomes the goal, large -scale AI factories become the most important trend in computer history as there is a race to fill the worldwide demand. NVIDIA and Cisco have come together to deliver Cisco secure AI factories with NVIDIA, expanding rack scale systems to the next level. We have here Cisco's SVP and head of networking engineering, Will Eatherton from NVIDIA, SVP, and CUBE alum, Gilad Shainer. Welcome back. And VP of Solutions Architecture and Engineering, Marc Hamilton, also from NVIDIA. Gentlemen, thank you for joining us here today for this announcement. Will, we'll start with you. Everyone's racing, as I mentioned, to fulfill the demand. We're seeing the rise of the neoclouds. They're expanding super fast. They have different approaches. Sovereign programs are on the horizon being architected now. And big enterprises are trying to figure out how to bring in the AI intelligence into their workflows. What's actually broken between ordering the systems, the GPUs, the compute, and actually getting all those systems into production?
Will Eatherton
>> Sure. I think we're seeing across all three of those domains, as you mentioned, convergence right now. So when we work with NeoClouds, from the moment that they put in the PO for the GPU, they already have their end customers lined up. And one of the broader challenges is just the speed, the expectation that from the moment that that is on the project plan, that the GPUs have to be racked, that that then has to go through all of the final deployment aspects, the software, and then a bring up and hand over to their end customer. So it's a rush. It's a rush to revenue for them. enterprises, many of them, actually, as a software development leader, I'm one of them, are spending a large amount on tokens right now. And the rush and the pressure is getting these systems up so they can start offloading what has been an API interface into using local inference. And so I think it's all converging on common architectures, common systems, but speed. I think speed is what is the broader challenge.
John Furrier
>> Mark, what's your take? The global view from NVIDIA has always been, let's get these bottlenecks fixed, let's deploy, let's make sure people have everything that they need. What's happening on your end?
Gilad Shainer
>> Well, first of all, to put this in perspective, the build-out of AI infrastructure is really the largest scale infrastructure build-out that the world has seen ever. people talk about the building of the transcontinental railroads or the U.S. interstate highway system. This is on par, if not greater, than that. And so, So this is where speed comes in. Everyone is rushing to adopt AI and to start to make money with AI. Traditionally, your servers, your networking to an enterprise, it's been a cost, and you've tried to reduce your costs. But factories, factories produce revenue, and this is what an AI factory does. It takes in raw materials, takes in data or tokens, and it takes in electricity. It converts that raw material using AI infrastructure into a finished good, into tokens, and those tokens have value. So it's not just about buying GPUs and having those GPUs delivered or having the network switch delivered. It's about having that AI factory from start to finish delivered and how quickly you can start producing tokens and revenue. And this is where having a standard reference architecture that the entire ecosystem can build to is important. And to me, it's really almost impossible to think about taking on the largest infrastructure project in the world without having Cisco at the table. And so it's great to have now this Cisco rack scale solution that they're bringing to market.
John Furrier
>> Gilad, weigh in on this, because we've seen everyone wanting the silicon, they want the systems, but now we're in the execution phase, delivering, standing up, and getting revenue. The faster you can get those tokens, the faster you can get those applications enabled, this becomes the focus now, operational excellence and execution at the highest quality.
Gilad Shainer
>> Yeah, building a factory, it's not connecting components and hoping for the best. Building AI factory means that you need to build a supercomputer and a supercomputer that needs to be built quickly, needs to be built fast, and needs to provide the highest numbers of tokens per second, the highest numbers of tokens per watt, and so forth. when we start building the infrastructure we start building a Spectrum-X Ethernet for example we looked into building a purpose-built infrastructure for ai purpose-built for supercomputers that needs to run distributed computing workloads so we focused on eliminating jitter for example we brought elements of lossless into scale out infrastructure we brought abilities of NICs and switches to work as one single infrastructure to eliminate Jitter with adaptive routing and RDMA, congestion control, and so much hardware capabilities. And it's a great platform. And we work with Cisco. We made an announcement, I think in October of last year, of Cisco leveraging the technology that we build. And now we're at a point where we see that or we bring to our mutual customers the capabilities of Spectrum-X into something that they're used to already. So Spectrum-X, the technology that we build, or Spectrum-X inside the Cisco systems, means that now there is an AI-optimized fabric that comes with the operating model that the enterprise world already runs on. Okay, that's what we're able to do together. And this is why this time is so important for our mutual customers.
John Furrier
>> Will, I want to get into the announcement. He mentioned distributed computing. I mentioned it's a historic moment in the industry. We've never seen this kind of scale and execution with computing. Rack scale systems, they're being built, they're being operated and invested in continuously. These cycles look like they're not going to stop for a long time. So let's get into what Cisco is announcing today. Cause you mentioned the previous announcements. I was in DC. I think we chatted about it then. Now we have the new secure rack scale system. What is Cisco actually announcing? Gotcha.
Will Eatherton
>> So as Gilad mentioned, we've been working together on getting all the networking components for the past year. So we have the Spectrum-X silicon with the Cisco software on top. We have the Cisco Silicon One switch architecture that can have the Spectrum-X license. We've been getting all those pieces together, working together on the reference architecture prep. What we're announcing today is that across the Sovereign Enterprise and NeoClouds, that we need to go big. And so we are going broader with compute. So that is liquid cooled, starting with Blackwell, moving to Vera Rubin, liquid cooled on the HGX form factor, MGX form factor. So that gives us a breadth of compute systems that we can then wrap around from a Cisco standpoint, from a sales support and a software standpoint. And so it's all those key components that we're announcing as far as the new products. This is then a solution that we've been working on. I partnered with Gilad, Mark, and the rest of the NVIDIA team on putting together how this is going to work for an NCP RA compatible construct. So we've worked through what that whole cluster looked like, how we're going to support that, and put together, for instance, a Cisco CVIS, Cisco Validated Infrastructure Services, that will help essentially be a front end across a broad range of enterprises, sovereign, neoclouds that are around the world. and then work with Mark and the NVIDIA team behind us so that we can help NVIDIA scale so that we can help our customers get these outcomes. And this full solution is going to be orderable in September. And so the Secure AI Factory is going rack scale.
John Furrier
>> So there's a technology partnership on the silicon and the certification, but it's also a go-to-market partnership with Cisco's validated infrastructure service, which is what, an end-to-end? And that's the motion. So take us through the two parts. Take us through the partner go-to-market piece.
Will Eatherton
>> Right. So Cisco will be selling the full solution. So that is the key aspects that will be the compute solution. We then will do the pre-sales work. We'll work with both our Cisco professional services, but with partners like WWT and Computacenter and so on around the day zero, day one activities. We then sell the networking. And then we're also solving for customers that are interested in the NCP RA. We will attach with that support from a Cisco CVD so that we're able to do a certification that it's meeting the performance standards that we are looking for. But also NVIDIA has put very clear guardrails and construct around.
John Furrier
>> So Gilad mentioned Spectrum-X. So Cisco Silicon One is what, the front end and NVIDIA Spectrum-X? So, yeah.
Will Eatherton
>> Silicon on the back end? Right. So the key architecture points would be Cisco, Silicon One running, for instance, NX-OS or SONiC on the front end. And we have cases where customers want, for instance, high performance with storage. There's an option for a Spectrum-X license, which then is part of that overall ecosystem that Gilad mentioned. Then on the back end, it's Spectrum Silicon running, again, NX-OS or SONiC. And then overall, that is a common management system. We have customers that for instance, also are doing more with routing interconnect of multiple data centers, a DCI data center interconnect. So everything from the routing, the front end, the storage networking and the backend, again, working seamlessly with storage partners like VAST, and then with the compute that we're pulling in, both the box level and the rack scale. And so, yeah, those are all the key components, John.
John Furrier
>> Mark, I want to bring you into this because Gilad mentioned building a supercomputer. this is great from the AI standpoint, but it's hard, right? So talk about the reference architecture and the NCP, NVIDIA Cloud Partner certification. How important is this certification in this era? Because these rack scale systems are not for the faint of heart, they are very complex, but they produce a lot of value. They want them fast. In rack scale, this is a huge thing. Explain the importance of this program.
Gilad Shainer
>> Well, the NCP reference architecture is really so important because an AI factory or AI data center is a five -layer cake, right? It's not just about the chips and the networking, right? It starts with your land power shell, your data center, then it's, of course, the chips, the networking chips, the compute chips. It's your AI infrastructure. This is where the servers, networking, and storage come together. It's your AI models and it's your AI applications. If you don't build that whole thing end to end and have a place where you can test every application, every model, every network architecture, it's nearly impossible to optimize it and make sure that it is the right solution. So through our NCP reference architecture and our enterprise reference architectures, this is exactly what we do. I think the fact that Cisco is taking this to market as one solution that they can sell and the customers can order is also important, especially for AI factory. Traditional enterprises had server teams and had networking teams. And the server team went off and bought servers. The networking team went off and bought networking. And the two didn't come together until very late. In an AI factory, because it's a five-layer cake, everything has to work together. There are so many mistakes when customers try to go to one vendor to order networking, another vendor to order servers. So I'm really excited to see this Cisco solution. You talked about also then why is it hard? You get all this infrastructure delivered, a bunch of pallets on your loading dock. And that time to first token, when you can start generating revenue is important. Two other things that are equally important that the reference architecture helps with. If you get your first token ready, but your availability is not good, You've got a $10 million or $10 billion AI factory, and 30%, 40%, 50 % of it is unavailable because it wasn't cabled right, because there was some problem with software, firmware wasn't loaded correctly, that's going to drive up your costs tremendously. The second thing, besides availability, is, again, the performance. You need availability and performance. And because NVIDIA in our reference architecture integrates all five layers of that stack, tunes across all those five layers, we can guarantee the best performance. And in the case of NVIDIA Cloud Partners, in fact, they run our, what we call, Exemplar Cloud benchmark suite, that they can then go out to the public and demonstrate that, show the results, compared to our gold standards that we run in our labs ourselves.
John Furrier
>> Gilad, talk about the end-to-end, when people bring their own networking to the delivery model, and maintaining the level of quality it takes. It's not like the old days, it's a whole new modern era. What's your view on this?
Gilad Shainer
>> Yeah. First, remember that what we build is a single unit of computing. And that single unit of computing that generates tokens is built on a huge amount of components. You're talking about hundreds of thousands of GPUs. You're talking about many switches, many cables, storage, NICs. There are so many elements. and infrastructure from scale up to scale out to access network and so forth, there are so many elements and still it needs to work as a single unit. And when neoclouds and enterprises buy AI factories, this is an important investment for them. That's actually how they're going to make money. That's how they're going to generate money. So they're going to make an important investment They're going to make a big investment, and they expect that all of those hundreds of thousands of components are going to work as one entity and they have to. That's what they expect. So the complexity here is huge, and that's why we build everything which is purposely built for AI. Every component is purposely built to generate tokens because that's how you generate money. And that's why the NCP RA reference is so important because that makes sure that they can bring the things together, deploy it very quickly and get tokens and make money. So now bringing that technology that we created and working together with Cisco, and now Cisco can deliver everything and unify the management of that big entity, which is a single entity of computing, if you think about it, managing that through Cisco Cloud Control and leveraging all the elements that NVIDIA brought inside. That's the importance of what's happening today, by the way. And this is where we're so happy to work with Cisco, as Will said, for quite some time. In a sense, it's a few months because the world is moving very, very fast. It is a lot of time, believe me. But there is an amazing collaboration between the two companies. Now Cisco can bring it all together with the same management that enterprise knows how to run, the same management that they're already running. So they can move quickly. they can leverage what they acquired very quickly. And then Cisco brings that with all of the NVIDIA technology inside, with all the capabilities to generate tokens, with all the capabilities to leverage AI and enable our mutual customers to move to the next level of their businesses.
John Furrier
>> Yeah, it's well said. I love how it brings conventional distributed computing to the artificial intelligence modern era. I guess I'd have to ask you, so you mentioned customers. So a Neo cloud, an enterprise, or a sovereign approach deploying the Cisco architecture is getting the same standard NVIDIA holds its own cloud and partners to. Is that right?
Gilad Shainer
>> They're getting the same performance as we make sure that they're going to get from the technology side. That's the reason that we build the reference architecture to make sure that they can get the full performance of NVIDIA components. And that guarantee is now coming with the Cisco AI Factory. That's how both of us guarantee that they're getting the best performance out of their investment. And one important element is when we build the InfiniBand switches, when we build Spectrum-X Ethernet, when we build those elements, we wanted to make sure that we have the full flexibility so you can actually be able to run your own things on top of that. So in that case, for example, Cisco can take their Cisco NX-OS and run them on the open interfaces that we brought into SpectrumX. Or they can take SpectrumX licenses and be able to connect it to the access network with Silicon One switches and so forth. So we brought the capabilities, we brought the technology, it's built in a fully open interface so you can go and customize that. And now everything's now working together as a single unit coming from Cisco.
John Furrier
>> Will, weigh in on the NCP because this is a major win for customers.
Will Eatherton
>> Yeah, so I want to mention it hasn't always been easy. So I think as we are working through and as Gilad mentioned, I think from an architecture standpoint, it was pretty clear how to approach it. But for instance, in working with Mark and his team, there's been key areas. For instance, they have a large array of scripts and technology that they've developed around the stack that is the default stack for NVIDIA. So we've been working as part of the Cisco Validated Designs, the CVDs, to take those and then make the tweaks in those tools for our Nexus, for our management layer. And so I think, again, bringing the core aspects of what NVIDIA has been bringing to market and then being able to adapt that so we can support our joint end customers is a big deal. And, I appreciate that Gilad mentioned the Cisco Cloud Control and really the experience Cisco has working with, telcos and enterprises in managing software and really this day two topic. And this is something, Mark and I talk a lot about, a lot of the industry focus is up to the point that you light up the cluster and you get your first token out. That's been a big focus. That's great. But the day two aspects around monitoring and health and availability and software upgrades, and these are things that from a Cisco standpoint, we've put a lot of focus on here, over the years. And as we get into AgenticOps that we're building into Cisco Cloud Control, the outcome of that is that, again, from whether it's routing and data center interconnect, whether it's front end networking storage networks, the front end, maybe an enterprise that is working in EVPN and VXLAN and other technologies to the backend with a Spectrum ASIC, a Spectrum-X technology, NX-OS or SONiC, and then integrating that into our overall management that then can be managed software life cycle and monitoring all as part of one cohesive system. And that is the big thing that we're excited to be bringing to our customers.
John Furrier
>> Mark, weigh in on this because there's a lot of stuff going on between Cisco and NVIDIA. He mentioned cloud control. There's also a lot of other Cisco products involved. Silicon One, and you got Spectrum-X. Talk about some of these reference archi, validated designs and reference architecture. And how do you get these repeatable provisioning and validation automation stuff done? There's a lot of implementation going on here.
Gilad Shainer
>> every customer that starts to build an AI factory starts off thinking, oh, I've built data centers before. I know how to do this and I can do it better. And in fact, even in the supercomputing space in the past, if you were the national labs or if you were a large university, that was part of the experience to go through and do a custom design and try to optimize on a little bit faster CPU or a little bit less memory to save money. And again, the real cost savings in an AI factory is not about cost savings, but it is about token generation and how do you drive that revenue. And so having a repeatable way to do it. The other great thing about this for Cisco's enterprise customers is, remember, most enterprises today want some sort of multi -cloud hybrid type of solution. So, being able to have the Cisco AI Factory on -prem and then being able to go to a neocloud and being able to access the same exact solution in any neocloud around the world running the Cisco AI Factory is a huge win and a huge advantage for Cisco's enterprise customers.
John Furrier
>> Yeah, and I think it's not your old playbook of building data centers. These are systems. Deployments have to be done right and they have to be done fast. and again, time to first token is revenue. And that's unique in this era, fast revenue producing. All right guys, final thoughts for all three of you. when these systems are in production across neoclouds, sovereign environments and enterprises, what's the one thing the market should remember? Gilad, we'll start with you.
Gilad Shainer
>> It's a good question. It's a good question. So what we're announcing, First, it's a great partnership. So people should remember October 25. This is when Cisco announced the Spectrum-X switches. And people should remember August 26, when we're announcing, or Cisco is announcing bringing a full AI factory, which is NCP compliant and enabling neoclouds and enterprises, for example, to continue running the same way they're running their traditional data centers. But now, they're going to run full AI factories and generate tokens and drive their businesses. So there's a great technology. There is a great roadmap ahead of us. And now that things are moving to production and people can go in and acquire from Cisco a full AI factory and just, connect it to the power and be able to generate tokens and make money. The only thing that they need to remember when they go back home is to buy milk in a supermarket. markets.
John Furrier
>> That'd be good. Mark, this is a market focus. There's a lot of headroom, a lot of things being done now. It's moving super fast. What's the one thing Mark, we should remember?
Gilad Shainer
>> Well, the one thing I would remember is that an AI factory is a living, breathing thing. Every week, there's a new AI model that's launched. And of course, recently, there's been huge interest in open AI models, be it NVIDIA's own Nemotron model or other open models. And so really building an AI factory is not just about that time to first token. We call it continuous bring -up optimization and upgrade because you install the AI factory, you get that new AI model that is launched and you want to optimize AI factory for that and then by the time you're doing that you then need to go through and do an upgrade. So taking Cisco's enterprise knowledge, Cisco's tools, and applying it to an AI factory for this continuous bring up, optimize, and upgrade motion is going to be a real market mover, I believe.
John Furrier
>> So you're saying the software stack is pretty critical because that's going to be the pacing item to keep the capabilities coming in. Every week there's a new model released. All right, Will, close us out. Cisco, why now? Enterprise market, we expect to be booming, and the neoclouds are building out as fast as they can. Obviously, we know what the hyperscalers are doing. This is distributed computing.
Will Eatherton
>> So I'm going to end with scale. This announcement is all about scale on multiple fronts. From a Cisco standpoint, it's about scaling, scaling the size of the system, the rack scale system, scaling the liquid cooling, scaling just the magnitude and the amount of clusters that we're going to be able to go out and work with our customers. This is about scale and how we're partnering with NVIDIA. You being able to wrap around the NCP RA, wrap around the NVIS with the Cisco Validated Infrastructure Services and help scale globally across a broader range of neoclouds, telcos, enterprises. It's about how do our customers scale? So, beyond first token, how do you scale the size of your clusters, the number of clusters, how do you scale the operations so that you get the uptime and the token accumulation, not only day one, but over time. So it's all about scale.
John Furrier
>> Gentlemen, thank you very much. Secure AI Factory with NVIDIA, Expanding Rack Scale. Thanks so much for taking the time and weighing in with your expert opinion. We've got a mixture of experts here. Thanks for coming on. You're welcome. Thank you, Mark, Will. Thanks. Thank you.