Vik Malyala of Supermicro, chief business officer, outlines Supermicro's approach to artificial intelligence, AI infrastructure and rack-scale system design at AMD Advancing AI 2026. John Furrier of theCUBE Research and Dave Vellante of theCUBE Research moderate the discussion and frame topics including agentic AI adoption, PCIe and HGX architectures, storage integration and strategies for scaling GPU clusters and Helios-class deployments.
Malyala presents rack-scale designs, supply-chain coordination and configurable platforms for mixed central processing unit CPU and graphics processing unit GPU workloads. They highlight that persistent supply constraints and memory allocation dynamics lengthen delivery timelines and that close co-design with vendors such as AMD is critical for high-power high-density systems. Malyala also emphasizes that agentic AI increases CPU orchestration demands, making CPU architecture, hybrid configurability and proper storage and cooling planning essential.
The hosts emphasize data center readiness and balanced system design while discussing peripheral component interconnect express, PCIe architectures, storage integration and approaches to scale GPU clusters for Helios-class deployments. The conversation provides practical insights for architects and procurement teams focused on AI infrastructure, GPU cluster scaling and high-density data center deployments.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Advancing AI 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for Advancing AI 2026
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for Advancing AI 2026.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Advancing AI 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to Advancing AI 2026
Please sign in with LinkedIn to continue to Advancing AI 2026. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Vik Malyala, Supermicro
Vik Malyala of Supermicro, chief business officer, outlines Supermicro's approach to artificial intelligence, AI infrastructure and rack-scale system design at AMD Advancing AI 2026. John Furrier of theCUBE Research and Dave Vellante of theCUBE Research moderate the discussion and frame topics including agentic AI adoption, PCIe and HGX architectures, storage integration and strategies for scaling GPU clusters and Helios-class deployments.
Malyala presents rack-scale designs, supply-chain coordination and configurable platforms for mixed central processing unit CPU and graphics processing unit GPU workloads. They highlight that persistent supply constraints and memory allocation dynamics lengthen delivery timelines and that close co-design with vendors such as AMD is critical for high-power high-density systems. Malyala also emphasizes that agentic AI increases CPU orchestration demands, making CPU architecture, hybrid configurability and proper storage and cooling planning essential.
The hosts emphasize data center readiness and balanced system design while discussing peripheral component interconnect express, PCIe architectures, storage integration and approaches to scale GPU clusters for Helios-class deployments. The conversation provides practical insights for architects and procurement teams focused on AI infrastructure, GPU cluster scaling and high-density data center deployments.
>> Welcome back, we're at theCUBE's live coverage here in San Francisco, California, AMD's Advancing AI event, all the leaders are here, industry participants, it's a free event, so a lot of practitioners and technologists here, checking out the rack scale systems, checking out all the new gear and networking servers, everything's here, power the next generation of AI. I'm John Furrier, host of theCUBE with Dave Vellante, my co -host, Vik Malyala here, he's the chief business officer of Supermicro, theCUBE alumni, recently I interviewed him. Long time no see. Supermicro, Dave just interviewed him. You guys had a big storage summit with theCUBE. Vik, great see you.
Vik Malyala
>> Thank you for having me as always. It's a pleasure talk you.
John Furrier
>> A lot's happened since supercomputing when we were riffing on the Neo -Cloud growth, AI infrastructure, a lot of content we put out together. But the big thing is just the massive build out, the demand curve. Dave was mentioning it before we came on camera. and just the nature of these AI factories. They're kind of taking shape as a new architecture. It's not the rack and stack days anymore, though there's still a lot of that going on, but it's a whole different computing paradigm. Explain the current situation. Where are we today? November we just chatted, lots happened since November.
Vik Malyala
>> I'm telling you, whatever we think is the pace at which the industry is moving, three months later it is proving us wrong, right? What I have seen is that, let's say, last time in November when we were talking about, there was hardly any discussion about agentic AI, was there? It was all about RAG models, and people are trying develop these different applications, but then in a snap it all changed. Now we are talking about heavy use of agentic AI across the board, and we have a strong demand that is coming up, not just on a GPU, but also on the CPU because of that. As far as the GPUs are concerned now, enterprises are starting pick up. It's not just all about the training clusters, but also the application stack that is being used on them, which is basically making the enterprises adopt that. So, what I see is that it's a very, very broad scale adoption of AI infrastructure, which is propelling this massive, massive demand.
John Furrier
>> And you guys have such a great history. I always kind of flex Supermicro success over the many generations. You know supply chain. And right now that is the number one thing people care about. Prices go up in memory, but people want solutions now. Talk about the impact of what that's done and also how that's changed how you guys think about building these large AI factories and these systems.
Vik Malyala
>> It's a complex equation. there's no ifs and buts about it. And if you take a look at the likes of AMD and NVIDIA, for example, for the rack scale solutions, they are bringing the predictability in the supply as well as pricing by them working directly with the memory manufacturers on the supply side at least, and in most of the cases, some level of pricing also. But as far as the standard HGX platforms and everything else is concerned, this is going be a very tricky market for quite some time come. And luckily for Supermicro, we've been working with pretty much every one of the industry leaders, whether it is memory or whether it's flash. Think of Micron, Samsung, Hynix, Solidigm, and every one of them for the longest time. And as we build these platforms and solutions, we give them enough visibility into where the systems are going, who the customers are that we are supporting, and what problem we are trying solve, which sometimes helps actually in getting an allocation at the right time. But no matter how much planning that we have, we know for a fact the supply is a lot more constrained than ever before. and we continue work with them see how best we can address it. And most importantly, we set expectations customers that it's no longer possible have a system shipped in a week or 10 days. It takes a longer time. And by taking the orders in, by having a proper forecast, it is also helping us plan better.
John Furrier
>> This event's like a, I won't say coming out party, but I don't want say that word. But we've been following AMD. They've been in the x86. We've seen that growth. but I've kind of been following. We can see them hiding the ball a little bit. They had FPGA, GPUs coming. What is the new AMD like? Talk about the relationship you have with AMD and how the, I won't say new and improved AMD, just now the AI systems side of AMD starting show, starting see the results. Explain what it means the market and what people should know about it.
Vik Malyala
>> I'll give a parallel story here. Back in 2015, 2016, when I was working with Dr. Lisa Su on the first generation of the EPYC platform, what ended up happening was the hardware was fantastic. You have many cores and the PCI Express lanes and everything, but the software and the ecosystem wasn't quite ready. So while it's adopted in a very specific marketplace, whether it is HPC, whether it is hyperscalers, the general adoption took an extra two cycles after Milan, then the next step like Genoa, Bergamo, and all these things. So what I have seen at every generation, more people started adopting, the ecosystem started expanding, and the point where they actually have the market leadership in x86, right? The same situation I see happening with respect Instinct also, initially hardware is fantastic, but the ecosystem is the one that needs be developing. I was told that some 12 ,000 people signed up for this event which is something that no one would have imagined two years ago if you had asked me in terms of the software developers. What I am seeing is that as people start figuring out a way use the GPUs, then the adoption of the GPUs beyond the training clusters is going happen. And we have seen a good demand build up, especially around the MI350 and MI355, and the demand is kind of extending into the Helios. And there's a lot of people asking about it and people are trying understand. Mind you, it's not just about the technology. People need figure out a way host them in their data center. It's not going be easy just plug a Helios in and make it work.
John Furrier
>> It's like the GPU is the initiation. Wow, I love this value. It's like, wait a minute, I need more compute. So the discovery on the ecosystem becomes kind of its own progression.
Dave Vellante
>> It is, it is. So you guys hinted that you're going two, that your margins are up, your backlog is, I think, 60 billion was the number you published. Is it as simple as this demand is so far outstripping supply?
Vik Malyala
>> we have done something right, right? So ultimately, we need take care of our customers' demands and as customers become successful, the business starts grow. And that's precisely what happened in this case. We have growing set of customers and wider adoption of the platforms. Think of like an enterprise, think of banking and financial sector, as well as the traditional compute and GPU environments. These are the ones that are actually creating the demand for us. And as far as the margins are concerned, it's a mix of various things, right? You're talking about data center building blocks that we are bringing expedite the deployment phase of these GPU clusters as part of it. Service organization coming up make sure that the things are going be up and running and be able stay up and running for at a high percentage of time. And the relationship that we have with ISVs for the storage and what we are bringing in terms of the storage solutions customers, so a combination of all these things is the one that is driving though the adoption as well as the impact on the margin. And margin is something that's going fluctuate, especially because of how the whole supply chain dynamics are working today, but we are certainly emphasizing on value and how we can actually help customers bring some predictability into their equation.
Dave Vellante
>> Yeah, Supermicro Storage Summit, I was there, it was a great event, some really good content. You were talking before about how things change so fast. I actually, earlier today, John and I came up with a list, it's like, oh yeah, RAG-based chatbots are going add huge value. You got token costs, who cares about token costs? And frontier models are going dominate everything. these other small models don't mean anything. How are you and AMD engineering for these rapid changes and seeing around corners? How do you accommodate this change?
Vik Malyala
>> So none of us have the crystal ball, but what we can do is bring the configurability into the platform design and be able size it differently. So one of the things that we have done is if you were take a look at the GPUs itself, right? You have the MI350X and MI355X, that is the HGX platforms predominantly being used, especially with AMD. And then we also have the inferencing point of view, which supports different accelerators. We have several accelerators that are actually based on AMD as a compute platform, and they have these accelerators as a PCIe, and AMD themselves are coming up with MI350P, which is also a PCIe based accelerator. Why I'm saying it is that for the guys who want like think of the Rolls-Royce, you have these platforms at the rack scale and it's fantastic, right? But for the ones who are trying use for different applications, they may not need that in order have the right ROI for what they are trying do. They can go with partially populated systems. For example, I can have a system that supports up eight of these PCIe GPUs and people can actually start with two or four or scale up eight depending on what they're trying do and if it goes beyond eight and if they're trying scale up, then we have the rack scale solution and multiple racks we can connect and create a cluster and whatnot. So there's no rocket science per se but what we are trying do is give the option the customers on what might work for them and this is one of the things. And the second part of what we are seeing is, you mentioned you started with the RAG models and whatnot. That one is like a single loop, right? you ask a question, the model is going run, it spits out the answers, and you're good go. Whether it's right or wrong, and how much of it is hallucination, it all depends on the data it's trained on, and how the models are developed and whatnot. They're getting a lot more accurate. But this agentic AI is a completely different spin, right? you're talking about it needs think, it needs reason, and it needs go into multiple loops, it needs memorize things, and it spawns off so many more sub -agents who are doing all their work. And what is happening with this is there is a significant demand for the compute, not necessarily the GPUs. From the GPU's point of view, they're doing a fantastic job. Think of what you mentioned about the RAG models, and when it gets that point, it has matrix multiplications it can do very quickly, it can spit out the answers, that's fantastic. But on the compute side of it, how do you make sure that the right type of queries are going the GPU and the rest of it is handled on the CPU. In the RAG model, all it is doing is just sequencing, managing the pipeline, orchestrating, done. But with respect agentic AI, now we are talking about the workload that's happening on the CPU, which is impacting 50 90 % depending on where you read and what you read about in terms of the latencies. So if the impact is that much that will be handled by the CPU, then it's a criminal waste not have the best computer architecture.
John Furrier
>> And the cost, financially, and the technical costs are massive. Because you got feed the GPUs, you got do routing, so there's a whole other orchestration layer here.
Dave Vellante
>> But finish that thought, so it's criminal not have the best CPU architecture orchestrate all of that, right?
Vik Malyala
>> Right, right, right.And that's where I think, now you have the CPU, you have the GPU, you're connecting, bringing the AMD into the equation, having the Pollara cards right now, and then the Pensando part of the equation. So what we are trying do is, depending on what customers are trying do, we can size it right. So have a PCIe based accelerators, or the standard HGX type of platforms. How much of the compute clusters that you need, and how much that you need it for agentic AI. What is the kind of bandwidth that is needed between these compute as well as the GPU platforms. How do you pair it with the appropriate storage, so that when things go past the size of available memory and storage available within the system, how it's going expand into the storage for the KV cache allocation, right? So the entire thought process of Supermicro as a building block is absolutely making sense right now in this kind of a scenario, because as things get more complex, ultimately what matters is how are we bringing the right value customers and sizing it right and making it run most optimally.
John Furrier
>> Talk about the AMD relationship when it comes co-designing and engineering. a lot's going into these rack scale systems. They're essentially super servers as one. And then you have distributed computing also, other areas are going have footprint. You have the NIC cards, they're going actually target workloads, so they all have talk each other. How do you guys partner now in the AI era, and where is it the same and where is it different?
Vik Malyala
>> So take an example of the PCIe based cards, right? It started with 75 watts, 150 watts, 225 watts, 300 watts. Now, 600 watts. So AMD gives us the visibility, hey, my GPUs in future are going be air cooled, but I need a platform that supports 600 watts. Or, I'm going have 600 watts, but I would like bring the maximum density, in which case I will say, well, if that's the case, let's see if that design can accommodate liquid cooling so I can actually put that many cards. And the other part of it is, let's say AMD silicon for the GPUs is based on PCI Express Gen 5 or Gen 6. Based on that I can say, well, I can actually put in the Venice platform with the PCI Express Gen 6 so you have much higher throughput that is going be supported in the platform. So what we typically look at in this scenario is what are the different technologies coming from them, what is the time frame, and when I bring a platform, how am I going make it run in the most optimal manner for that. Does it require liquid cooling? Does it need be air cooling? What is the form factor? What could be the power delivery? And we look at where that product could be adopted. Like I mentioned, when people start adopting Helios, let's say deployment starting whenever it is, we will be ready with AMD have our compute platforms go along with it. And the public information is that AMD Helios is going be running with ORv3 form factor with the power delivery using the DC bus bar. The entire server product portfolio we have transformed, I'm going say, not entire, but most of them, like the Hyper, the CloudDC, the Twin platform, the FlexTwin, all of them we have started designing with the power delivery using the DC bus bar. Why? Because as they go into the data center, they don't need figure out, okay, now I need have standard power supplies, I need have a 63 amps or 100 amp power drops and the PDUs. Helios does need a DC bus bar. So I'll just do the same DC bus bar on the compute also, That way you don't have go and rediscover. This is how we are looking at what would be needed for the customers in a given timeframe and what kind of a platform that I need develop or platforms I need develop in order for us get that rolling quickly.
Dave Vellante
>> So where do you see Helios fitting that rack scale system? Where do you see the demand for that? What type of workloads is it going support? What's the fundamental customer value prop?
Vik Malyala
>> So at this point, people if they're doing the primary inferencing workload, HGX platforms are doing a phenomenal job. But when it comes frontier models, that's where I see the immediate demand for Helios, because the models are getting trickier. I say trickier because it's no longer just about how many trillions of parameters, it's all about what are the optimizations that they're bringing into that. And when it comes inferencing, I think it was like a year ago or whenever it is, when we talk about the DeepSeek, and everyone's like, oh my god, the world is coming an end. And after that, everyone has multiplied their numbers by like three or four times. It's happening again now with all the open models.
John Furrier
>> Exactly, innovation.
Vik Malyala
>> The point is, all of it is going drive the democratization. So whenever people figure out a way run things more efficiently, it's not a negative thing. It basically makes it easy for others opt in.
John Furrier
>> It's called engineering.
Vik Malyala
>> It's called engineering. Buying opportunity.
John Furrier
>> Bingo. This is where I like the compute direction because when that gets smaller, faster, cheaper, that sounds familiar, Moore's Law. So you start see a kind of a dynamic where the engineering is going enable things run cost effectively. That should help the enterprise for sure.
Vik Malyala
>> For sure.
John Furrier
>> The Neo clouds want have versatility. They want have great price performance, get their margins, because they want have higher margins.
Vik Malyala
>> And one thing that is happening, though, this is not specific AMD, it's even NVIDIA and whatnot, is that every one of them have announced, when I say all their technology partners have announced the roadmap, like this is how the power delivery is going be, this is how the system is going be, this is how much cooling I would need, this is how heavy the system is going be. So this way the data centers are being retrofitted or built. Otherwise what happens is the traditional data centers, they have no way of taking this 240, 50 kilowatts per rack. 7,000 pounds and whatever. There you go. And you need change the elevators and you need two, oh by the way, we have seen cases even with the, forget about Helios or anything, Even with the previous generation, the current generation platforms we are shipping, some data centers have retrofitted, everything was fine. We did an audit and everything, it was fine. When you start rolling in the racks, the tiles crack.
John Furrier
>> Yeah, too heavy.
Vik Malyala
>> Because these are heavy. So that's the reason, it's not just about us working with AMD in developing the platforms, we also need work with the data centers ensure that they can actually handle this. Otherwise, I'm going have a very heavy paperweight. The data center is the computer.
John Furrier
>> It is. The data center is the computer.
Dave Vellante
>> It's like the salesperson, the Xerox salesperson back in the day, where are you going put it?
John Furrier
>> Business is good, I absolutely love this, and then the bubble conversation kind of dies down because you look at the backlog on all the top hyperscalers, neoclouds, they're not even recognizing revenue, they have backlog. That's good. So there's backlog, so this bubble discussion, come on.
Dave Vellante
>> I don't think it dies down.
John Furrier
>> Now the supply chain, pricing, okay I see that, that's supply issue, but the demand going forward is...
Vik Malyala
>> with any market of any type, there are going be some fails, for sure. But I don't see it as a bubble at all, because I see this pent up demand is so strong. that's part of the reason why you have announced.
Dave Vellante
>> You don't think we're in a bubble?
Vik Malyala
>> Not a pop bubble, not a dot com bubble.
Dave Vellante
>> I think we're in a bubble, it just hasn't popped. That's it. It won't pop. I don't think it's a bubble. It'll settle. It won't pop. It's expanding, the market's going like this, isn't it? If the application stack doesn't improve. dot-com days. It's all good until it stops. But the idea is this, right? Give me your perspective on that.
Vik Malyala
>> Ultimately, all these tools and everything need make our life more comfortable and better, right? Hardware alone is not going do that. Firmware stack is not going do that. It's ultimately the user experience and what value it's bringing into the equation. just with RAG models we have seen so much improvement in how we are able handle. nowadays, you know. Coding is great. Exactly, right? You take it for granted, you're right. Right, and then going a step further with Agentic AI, now it's a completely different experience. And look at the different verticals it can make an impact. We hardly started seeing these things in any of these verticals. So my take is that there's going be a whole bunch of software developers that are going scale up and do different types of work. not just like writing a bunch of lines of code, but actually architecting solutions that people can take advantage of, and that's what is going make this hardware better used.
John Furrier
>> I don't think it's a problem because of that. I think you're right, Vik, because one of the metrics that's coming out of all of our conversations over the past year here on theCUBE is whoever can bring the value the user, and the value them is agency. Empower me do new things. That's like the web when it came out. That's like all these new environments. Well, it's obvious. the old way's not as good as the new way. But it's natural language, it's not GUI.
Vik Malyala
>> Correct, correct.
John Furrier
>> It's a whole nother experience that's driving the stack be rethought. So if you can't provide value that's easy use and simple execute, you're out.
Vik Malyala
>> Agree, so take an example of enterprise customers.If a million tokens cost 40 cents versus a million tokens cost $40, the CFO is not going sign the check if it is going cost 40 80 dollars for a million tokens because he's going close the shop, even if they're able run the business very efficiently, the outcome is not going be as great. Versus if they're able bring the number of token the cost of tokens down a reasonable level, then it will be adopted and the business is profitable. So businesses need be profitable.
Dave Vellante
>> It's going be 50 cents at some point.
Vik Malyala
>> Yeah, it will be, but the idea again here is tokens, number one, cost coming down. Things running more efficiently. And more importantly, when people are using the models, they get smarter, so they're not going let's say, boil the ocean do simple things, but they are going do only the very specific, vertical specific applications and tune that, which is going run more efficiently. So, we have a long road ahead of us in terms of what are the things that can be done with it.
Dave Vellante
>> Yeah, I think so, the bubble's going keep getting bigger and bigger and bigger and bigger and bigger.
John Furrier
>> That's called TAM, Dave. It's called TAM. That's called market opportunity. Vik, great have you on, as usual, great conversation. We'll see you again soon, probably Supercomputing or Open Compute, one of these shows. Love what you guys do, just keep making the faster machines, systems for us. We need more tokens, cheaper, faster, more AI.
Dave Vellante
>> Driving the cost down, Vik, that's your job.
John Furrier
>> Efficiencies, we have bring it down.
Dave Vellante
>> Tech is deflationary. You've done a great job there, keep it up.
John Furrier
>> And get as much memory and give it theCUBE so we have a side business of reselling memory. because there's a lot of demand. Well, according you, if there's a bubble, then it needs memory, isn't it?
Dave Vellante
>> No, no, bubble's a good thing. It's just when it pops, that's the bad thing.
John Furrier
>> There you go. Chief Business Officer of Supermicro. Again, another example of the suppliers rethinking how they build the systems that are powering AI and advancing AI. It's a team sport up and down the stack. The ecosystem's a huge part of it. We want see them go faster. We need more products. Doing our part here in theCUBE with the ecosystem. I'm John Furrier with Dave Vellante. Thanks for watching.
>> Welcome back, we're at theCUBE's live coverage here in San Francisco, California, AMD's Advancing AI event, all the leaders are here, industry participants, it's a free event, so a lot of practitioners and technologists here, checking out the rack scale systems, checking out all the new gear and networking servers, everything's here, power the next generation of AI. I'm John Furrier, host of theCUBE with Dave Vellante, my co -host, Vik Malyala here, he's the chief business officer of Supermicro, theCUBE alumni, recently I interviewed him. Long time no see. Supermicro, Dave just interviewed him. You guys had a big storage summit with theCUBE. Vik, great see you.
Vik Malyala
>> Thank you for having me as always. It's a pleasure talk you.
John Furrier
>> A lot's happened since supercomputing when we were riffing on the Neo -Cloud growth, AI infrastructure, a lot of content we put out together. But the big thing is just the massive build out, the demand curve. Dave was mentioning it before we came on camera. and just the nature of these AI factories. They're kind of taking shape as a new architecture. It's not the rack and stack days anymore, though there's still a lot of that going on, but it's a whole different computing paradigm. Explain the current situation. Where are we today? November we just chatted, lots happened since November.
Vik Malyala
>> I'm telling you, whatever we think is the pace at which the industry is moving, three months later it is proving us wrong, right? What I have seen is that, let's say, last time in November when we were talking about, there was hardly any discussion about agentic AI, was there? It was all about RAG models, and people are trying develop these different applications, but then in a snap it all changed. Now we are talking about heavy use of agentic AI across the board, and we have a strong demand that is coming up, not just on a GPU, but also on the CPU because of that. As far as the GPUs are concerned now, enterprises are starting pick up. It's not just all about the training clusters, but also the application stack that is being used on them, which is basically making the enterprises adopt that. So, what I see is that it's a very, very broad scale adoption of AI infrastructure, which is propelling this massive, massive demand.
John Furrier
>> And you guys have such a great history. I always kind of flex Supermicro success over the many generations. You know supply chain. And right now that is the number one thing people care about. Prices go up in memory, but people want solutions now. Talk about the impact of what that's done and also how that's changed how you guys think about building these large AI factories and these systems.
Vik Malyala
>> It's a complex equation. there's no ifs and buts about it. And if you take a look at the likes of AMD and NVIDIA, for example, for the rack scale solutions, they are bringing the predictability in the supply as well as pricing by them working directly with the memory manufacturers on the supply side at least, and in most of the cases, some level of pricing also. But as far as the standard HGX platforms and everything else is concerned, this is going be a very tricky market for quite some time come. And luckily for Supermicro, we've been working with pretty much every one of the industry leaders, whether it is memory or whether it's flash. Think of Micron, Samsung, Hynix, Solidigm, and every one of them for the longest time. And as we build these platforms and solutions, we give them enough visibility into where the systems are going, who the customers are that we are supporting, and what problem we are trying solve, which sometimes helps actually in getting an allocation at the right time. But no matter how much planning that we have, we know for a fact the supply is a lot more constrained than ever before. and we continue work with them see how best we can address it. And most importantly, we set expectations customers that it's no longer possible have a system shipped in a week or 10 days. It takes a longer time. And by taking the orders in, by having a proper forecast, it is also helping us plan better.
John Furrier
>> This event's like a, I won't say coming out party, but I don't want say that word. But we've been following AMD. They've been in the x86. We've seen that growth. but I've kind of been following. We can see them hiding the ball a little bit. They had FPGA, GPUs coming. What is the new AMD like? Talk about the relationship you have with AMD and how the, I won't say new and improved AMD, just now the AI systems side of AMD starting show, starting see the results. Explain what it means the market and what people should know about it.
Vik Malyala
>> I'll give a parallel story here. Back in 2015, 2016, when I was working with Dr. Lisa Su on the first generation of the EPYC platform, what ended up happening was the hardware was fantastic. You have many cores and the PCI Express lanes and everything, but the software and the ecosystem wasn't quite ready. So while it's adopted in a very specific marketplace, whether it is HPC, whether it is hyperscalers, the general adoption took an extra two cycles after Milan, then the next step like Genoa, Bergamo, and all these things. So what I have seen at every generation, more people started adopting, the ecosystem started expanding, and the point where they actually have the market leadership in x86, right? The same situation I see happening with respect Instinct also, initially hardware is fantastic, but the ecosystem is the one that needs be developing. I was told that some 12 ,000 people signed up for this event which is something that no one would have imagined two years ago if you had asked me in terms of the software developers. What I am seeing is that as people start figuring out a way use the GPUs, then the adoption of the GPUs beyond the training clusters is going happen. And we have seen a good demand build up, especially around the MI350 and MI355, and the demand is kind of extending into the Helios. And there's a lot of people asking about it and people are trying understand. Mind you, it's not just about the technology. People need figure out a way host them in their data center. It's not going be easy just plug a Helios in and make it work.
John Furrier
>> It's like the GPU is the initiation. Wow, I love this value. It's like, wait a minute, I need more compute. So the discovery on the ecosystem becomes kind of its own progression.
Dave Vellante
>> It is, it is. So you guys hinted that you're going two, that your margins are up, your backlog is, I think, 60 billion was the number you published. Is it as simple as this demand is so far outstripping supply?
Vik Malyala
>> we have done something right, right? So ultimately, we need take care of our customers' demands and as customers become successful, the business starts grow. And that's precisely what happened in this case. We have growing set of customers and wider adoption of the platforms. Think of like an enterprise, think of banking and financial sector, as well as the traditional compute and GPU environments. These are the ones that are actually creating the demand for us. And as far as the margins are concerned, it's a mix of various things, right? You're talking about data center building blocks that we are bringing expedite the deployment phase of these GPU clusters as part of it. Service organization coming up make sure that the things are going be up and running and be able stay up and running for at a high percentage of time. And the relationship that we have with ISVs for the storage and what we are bringing in terms of the storage solutions customers, so a combination of all these things is the one that is driving though the adoption as well as the impact on the margin. And margin is something that's going fluctuate, especially because of how the whole supply chain dynamics are working today, but we are certainly emphasizing on value and how we can actually help customers bring some predictability into their equation.
Dave Vellante
>> Yeah, Supermicro Storage Summit, I was there, it was a great event, some really good content. You were talking before about how things change so fast. I actually, earlier today, John and I came up with a list, it's like, oh yeah, RAG-based chatbots are going add huge value. You got token costs, who cares about token costs? And frontier models are going dominate everything. these other small models don't mean anything. How are you and AMD engineering for these rapid changes and seeing around corners? How do you accommodate this change?
Vik Malyala
>> So none of us have the crystal ball, but what we can do is bring the configurability into the platform design and be able size it differently. So one of the things that we have done is if you were take a look at the GPUs itself, right? You have the MI350X and MI355X, that is the HGX platforms predominantly being used, especially with AMD. And then we also have the inferencing point of view, which supports different accelerators. We have several accelerators that are actually based on AMD as a compute platform, and they have these accelerators as a PCIe, and AMD themselves are coming up with MI350P, which is also a PCIe based accelerator. Why I'm saying it is that for the guys who want like think of the Rolls-Royce, you have these platforms at the rack scale and it's fantastic, right? But for the ones who are trying use for different applications, they may not need that in order have the right ROI for what they are trying do. They can go with partially populated systems. For example, I can have a system that supports up eight of these PCIe GPUs and people can actually start with two or four or scale up eight depending on what they're trying do and if it goes beyond eight and if they're trying scale up, then we have the rack scale solution and multiple racks we can connect and create a cluster and whatnot. So there's no rocket science per se but what we are trying do is give the option the customers on what might work for them and this is one of the things. And the second part of what we are seeing is, you mentioned you started with the RAG models and whatnot. That one is like a single loop, right? you ask a question, the model is going run, it spits out the answers, and you're good go. Whether it's right or wrong, and how much of it is hallucination, it all depends on the data it's trained on, and how the models are developed and whatnot. They're getting a lot more accurate. But this agentic AI is a completely different spin, right? you're talking about it needs think, it needs reason, and it needs go into multiple loops, it needs memorize things, and it spawns off so many more sub -agents who are doing all their work. And what is happening with this is there is a significant demand for the compute, not necessarily the GPUs. From the GPU's point of view, they're doing a fantastic job. Think of what you mentioned about the RAG models, and when it gets that point, it has matrix multiplications it can do very quickly, it can spit out the answers, that's fantastic. But on the compute side of it, how do you make sure that the right type of queries are going the GPU and the rest of it is handled on the CPU. In the RAG model, all it is doing is just sequencing, managing the pipeline, orchestrating, done. But with respect agentic AI, now we are talking about the workload that's happening on the CPU, which is impacting 50 90 % depending on where you read and what you read about in terms of the latencies. So if the impact is that much that will be handled by the CPU, then it's a criminal waste not have the best computer architecture.
John Furrier
>> And the cost, financially, and the technical costs are massive. Because you got feed the GPUs, you got do routing, so there's a whole other orchestration layer here.
Dave Vellante
>> But finish that thought, so it's criminal not have the best CPU architecture orchestrate all of that, right?
Vik Malyala
>> Right, right, right.And that's where I think, now you have the CPU, you have the GPU, you're connecting, bringing the AMD into the equation, having the Pollara cards right now, and then the Pensando part of the equation. So what we are trying do is, depending on what customers are trying do, we can size it right. So have a PCIe based accelerators, or the standard HGX type of platforms. How much of the compute clusters that you need, and how much that you need it for agentic AI. What is the kind of bandwidth that is needed between these compute as well as the GPU platforms. How do you pair it with the appropriate storage, so that when things go past the size of available memory and storage available within the system, how it's going expand into the storage for the KV cache allocation, right? So the entire thought process of Supermicro as a building block is absolutely making sense right now in this kind of a scenario, because as things get more complex, ultimately what matters is how are we bringing the right value customers and sizing it right and making it run most optimally.
John Furrier
>> Talk about the AMD relationship when it comes co-designing and engineering. a lot's going into these rack scale systems. They're essentially super servers as one. And then you have distributed computing also, other areas are going have footprint. You have the NIC cards, they're going actually target workloads, so they all have talk each other. How do you guys partner now in the AI era, and where is it the same and where is it different?
Vik Malyala
>> So take an example of the PCIe based cards, right? It started with 75 watts, 150 watts, 225 watts, 300 watts. Now, 600 watts. So AMD gives us the visibility, hey, my GPUs in future are going be air cooled, but I need a platform that supports 600 watts. Or, I'm going have 600 watts, but I would like bring the maximum density, in which case I will say, well, if that's the case, let's see if that design can accommodate liquid cooling so I can actually put that many cards. And the other part of it is, let's say AMD silicon for the GPUs is based on PCI Express Gen 5 or Gen 6. Based on that I can say, well, I can actually put in the Venice platform with the PCI Express Gen 6 so you have much higher throughput that is going be supported in the platform. So what we typically look at in this scenario is what are the different technologies coming from them, what is the time frame, and when I bring a platform, how am I going make it run in the most optimal manner for that. Does it require liquid cooling? Does it need be air cooling? What is the form factor? What could be the power delivery? And we look at where that product could be adopted. Like I mentioned, when people start adopting Helios, let's say deployment starting whenever it is, we will be ready with AMD have our compute platforms go along with it. And the public information is that AMD Helios is going be running with ORv3 form factor with the power delivery using the DC bus bar. The entire server product portfolio we have transformed, I'm going say, not entire, but most of them, like the Hyper, the CloudDC, the Twin platform, the FlexTwin, all of them we have started designing with the power delivery using the DC bus bar. Why? Because as they go into the data center, they don't need figure out, okay, now I need have standard power supplies, I need have a 63 amps or 100 amp power drops and the PDUs. Helios does need a DC bus bar. So I'll just do the same DC bus bar on the compute also, That way you don't have go and rediscover. This is how we are looking at what would be needed for the customers in a given timeframe and what kind of a platform that I need develop or platforms I need develop in order for us get that rolling quickly.
Dave Vellante
>> So where do you see Helios fitting that rack scale system? Where do you see the demand for that? What type of workloads is it going support? What's the fundamental customer value prop?
Vik Malyala
>> So at this point, people if they're doing the primary inferencing workload, HGX platforms are doing a phenomenal job. But when it comes frontier models, that's where I see the immediate demand for Helios, because the models are getting trickier. I say trickier because it's no longer just about how many trillions of parameters, it's all about what are the optimizations that they're bringing into that. And when it comes inferencing, I think it was like a year ago or whenever it is, when we talk about the DeepSeek, and everyone's like, oh my god, the world is coming an end. And after that, everyone has multiplied their numbers by like three or four times. It's happening again now with all the open models.
John Furrier
>> Exactly, innovation.
Vik Malyala
>> The point is, all of it is going drive the democratization. So whenever people figure out a way run things more efficiently, it's not a negative thing. It basically makes it easy for others opt in.
John Furrier
>> It's called engineering.
Vik Malyala
>> It's called engineering. Buying opportunity.
John Furrier
>> Bingo. This is where I like the compute direction because when that gets smaller, faster, cheaper, that sounds familiar, Moore's Law. So you start see a kind of a dynamic where the engineering is going enable things run cost effectively. That should help the enterprise for sure.
Vik Malyala
>> For sure.
John Furrier
>> The Neo clouds want have versatility. They want have great price performance, get their margins, because they want have higher margins.
Vik Malyala
>> And one thing that is happening, though, this is not specific AMD, it's even NVIDIA and whatnot, is that every one of them have announced, when I say all their technology partners have announced the roadmap, like this is how the power delivery is going be, this is how the system is going be, this is how much cooling I would need, this is how heavy the system is going be. So this way the data centers are being retrofitted or built. Otherwise what happens is the traditional data centers, they have no way of taking this 240, 50 kilowatts per rack. 7,000 pounds and whatever. There you go. And you need change the elevators and you need two, oh by the way, we have seen cases even with the, forget about Helios or anything, Even with the previous generation, the current generation platforms we are shipping, some data centers have retrofitted, everything was fine. We did an audit and everything, it was fine. When you start rolling in the racks, the tiles crack.
John Furrier
>> Yeah, too heavy.
Vik Malyala
>> Because these are heavy. So that's the reason, it's not just about us working with AMD in developing the platforms, we also need work with the data centers ensure that they can actually handle this. Otherwise, I'm going have a very heavy paperweight. The data center is the computer.
John Furrier
>> It is. The data center is the computer.
Dave Vellante
>> It's like the salesperson, the Xerox salesperson back in the day, where are you going put it?
John Furrier
>> Business is good, I absolutely love this, and then the bubble conversation kind of dies down because you look at the backlog on all the top hyperscalers, neoclouds, they're not even recognizing revenue, they have backlog. That's good. So there's backlog, so this bubble discussion, come on.
Dave Vellante
>> I don't think it dies down.
John Furrier
>> Now the supply chain, pricing, okay I see that, that's supply issue, but the demand going forward is...
Vik Malyala
>> with any market of any type, there are going be some fails, for sure. But I don't see it as a bubble at all, because I see this pent up demand is so strong. that's part of the reason why you have announced.
Dave Vellante
>> You don't think we're in a bubble?
Vik Malyala
>> Not a pop bubble, not a dot com bubble.
Dave Vellante
>> I think we're in a bubble, it just hasn't popped. That's it. It won't pop. I don't think it's a bubble. It'll settle. It won't pop. It's expanding, the market's going like this, isn't it? If the application stack doesn't improve. dot-com days. It's all good until it stops. But the idea is this, right? Give me your perspective on that.
Vik Malyala
>> Ultimately, all these tools and everything need make our life more comfortable and better, right? Hardware alone is not going do that. Firmware stack is not going do that. It's ultimately the user experience and what value it's bringing into the equation. just with RAG models we have seen so much improvement in how we are able handle. nowadays, you know. Coding is great. Exactly, right? You take it for granted, you're right. Right, and then going a step further with Agentic AI, now it's a completely different experience. And look at the different verticals it can make an impact. We hardly started seeing these things in any of these verticals. So my take is that there's going be a whole bunch of software developers that are going scale up and do different types of work. not just like writing a bunch of lines of code, but actually architecting solutions that people can take advantage of, and that's what is going make this hardware better used.
John Furrier
>> I don't think it's a problem because of that. I think you're right, Vik, because one of the metrics that's coming out of all of our conversations over the past year here on theCUBE is whoever can bring the value the user, and the value them is agency. Empower me do new things. That's like the web when it came out. That's like all these new environments. Well, it's obvious. the old way's not as good as the new way. But it's natural language, it's not GUI.
Vik Malyala
>> Correct, correct.
John Furrier
>> It's a whole nother experience that's driving the stack be rethought. So if you can't provide value that's easy use and simple execute, you're out.
Vik Malyala
>> Agree, so take an example of enterprise customers.If a million tokens cost 40 cents versus a million tokens cost $40, the CFO is not going sign the check if it is going cost 40 80 dollars for a million tokens because he's going close the shop, even if they're able run the business very efficiently, the outcome is not going be as great. Versus if they're able bring the number of token the cost of tokens down a reasonable level, then it will be adopted and the business is profitable. So businesses need be profitable.
Dave Vellante
>> It's going be 50 cents at some point.
Vik Malyala
>> Yeah, it will be, but the idea again here is tokens, number one, cost coming down. Things running more efficiently. And more importantly, when people are using the models, they get smarter, so they're not going let's say, boil the ocean do simple things, but they are going do only the very specific, vertical specific applications and tune that, which is going run more efficiently. So, we have a long road ahead of us in terms of what are the things that can be done with it.
Dave Vellante
>> Yeah, I think so, the bubble's going keep getting bigger and bigger and bigger and bigger and bigger.
John Furrier
>> That's called TAM, Dave. It's called TAM. That's called market opportunity. Vik, great have you on, as usual, great conversation. We'll see you again soon, probably Supercomputing or Open Compute, one of these shows. Love what you guys do, just keep making the faster machines, systems for us. We need more tokens, cheaper, faster, more AI.
Dave Vellante
>> Driving the cost down, Vik, that's your job.
John Furrier
>> Efficiencies, we have bring it down.
Dave Vellante
>> Tech is deflationary. You've done a great job there, keep it up.
John Furrier
>> And get as much memory and give it theCUBE so we have a side business of reselling memory. because there's a lot of demand. Well, according you, if there's a bubble, then it needs memory, isn't it?
Dave Vellante
>> No, no, bubble's a good thing. It's just when it pops, that's the bad thing.
John Furrier
>> There you go. Chief Business Officer of Supermicro. Again, another example of the suppliers rethinking how they build the systems that are powering AI and advancing AI. It's a team sport up and down the stack. The ecosystem's a huge part of it. We want see them go faster. We need more products. Doing our part here in theCUBE with the ecosystem. I'm John Furrier with Dave Vellante. Thanks for watching.