In this interview from AMD Advancing AI 2026 in San Francisco, Suresh Andani, corporate vice president of compute and enterprise AI at AMD, joins theCUBE's Dave Vellante and John Furrier to discuss the shift from raw hardware specs to outcome-based economics in enterprise AI deployment. Andani explains how enterprises are moving beyond frontier-only strategies toward hybrid architectures that route tasks between frontier APIs and on-premises open-weight models based on token economics, security and data sovereignty. He unveils AMD's newly launched MI350P GPU, a PCIe-form-factor, HBM-class card built for existing air-cooled data centers under 30 kilowatts. It supports models up to 260 billion parameters and more than 1,000 concurrent users, giving enterprises and edge deployments a path to run agentic workloads without a full infrastructure overhaul.
The conversation also explores why open-weight models are becoming central to enterprise control, letting CIOs avoid vendor lock-in and tune infrastructure to their own workflows rather than depend solely on proprietary frontier models. Andani frames AMD's role as a full-stack provider spanning silicon, platforms and ISV partnerships with companies including Nutanix, enabling enterprises to build, buy or fully own their agent tech stacks. He details how connecting siloed systems like ERP, CRM and ITSM through self-built agents reduces hallucination and strengthens data sovereignty across an organization. Citing early customer data showing payback periods as short as six months for enterprises processing a billion tokens daily, Andani makes the case that agentic AI is finally moving out of the pilot phase. From token routing on Ryzen AI PCs to escalating workloads into data-center-scale infrastructure, he outlines a roadmap for enterprises to become their own token generators rather than remaining token consumers.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Advancing AI 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for Advancing AI 2026
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for Advancing AI 2026.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Advancing AI 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to Advancing AI 2026
Please sign in with LinkedIn to continue to Advancing AI 2026. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Suresh Andani, AMD
In this interview from AMD Advancing AI 2026 in San Francisco, Suresh Andani, corporate vice president of compute and enterprise AI at AMD, joins theCUBE's Dave Vellante and John Furrier to discuss the shift from raw hardware specs to outcome-based economics in enterprise AI deployment. Andani explains how enterprises are moving beyond frontier-only strategies toward hybrid architectures that route tasks between frontier APIs and on-premises open-weight models based on token economics, security and data sovereignty. He unveils AMD's newly launched MI350P GPU, a PCIe-form-factor, HBM-class card built for existing air-cooled data centers under 30 kilowatts. It supports models up to 260 billion parameters and more than 1,000 concurrent users, giving enterprises and edge deployments a path to run agentic workloads without a full infrastructure overhaul.
The conversation also explores why open-weight models are becoming central to enterprise control, letting CIOs avoid vendor lock-in and tune infrastructure to their own workflows rather than depend solely on proprietary frontier models. Andani frames AMD's role as a full-stack provider spanning silicon, platforms and ISV partnerships with companies including Nutanix, enabling enterprises to build, buy or fully own their agent tech stacks. He details how connecting siloed systems like ERP, CRM and ITSM through self-built agents reduces hallucination and strengthens data sovereignty across an organization. Citing early customer data showing payback periods as short as six months for enterprises processing a billion tokens daily, Andani makes the case that agentic AI is finally moving out of the pilot phase. From token routing on Ryzen AI PCs to escalating workloads into data-center-scale infrastructure, he outlines a roadmap for enterprises to become their own token generators rather than remaining token consumers.
>> Welcome back to theCUBE live stream here in San Francisco. I'm John Furrier, host of theCUBE with Dave Vellante, my co -host. We're here for AMD's Advancing AI event, global event. We have all the industry players here, as well as some free agents. We have a lot of practitioners, a lot of engineers here, people building out the next generation of AI infrastructure and all the applications on top. Suresh Andani's here, Corporate Vice President of Compute and Enterprise AI for AMD, he's the man who's going to bring all that compute magic to the marketplace and enable the enterprises to go to the next level. It's great to see you.
Suresh Andani
>> Likewise.
Dave Vellante
>> Thanks for coming back on theCUBE. What's going on?
John Furrier
>> Alright, so I got to get into this because the enterprise has been waiting, we're seeing inference all over the place, starting to see people resettling in. This has been a shift, we've been documenting this shift. This is a generational shift where now the computing industry has moved over to a net new architecture. It's cloud, it's on premise, it's edge, all of it is distributed computing, hybrid. This is where compute, GPUs, FPGAs, all have to work together in concert. It's a whole new system.
Suresh Andani
>> Yes, yeah, no, you're absolutely right, John. So in traditional sense, we talked about hybrid only two ways, right? Either we're running on-prem or in the edge or in the cloud, right? Now when we talk about more distributed AI, it's not just where we're running the workload, but what models you are running are also hybrid. A lot of enterprises are now figuring out that, not all tasks need frontier models. frontier models are great, by the way. If you need the deepest context, if you really need the high concurrency, frontier models have really done well for the enterprises, but not every enterprise AI task needs to run on a frontier model. A lot of tasks because of token economics, because of security reasons, data sovereignty reasons and control reasons are better suited to run on-prem on open-weight models. So as a provider of AI compute, we want to make sure that we are enabling those enterprises to not only run their AI tasks through the frontier APIs in the cloud, but we are providing them infrastructure that they can host open-weight models very efficiently on-prem. And that's a key, every enterprise conversation I'm in, that's a key topic of discussion.
John Furrier
>> Two weeks ago I interviewed Mark Papermaster at the RAISE Summit. He's 40 years veteran CTO at AMD, you know him very well. He's an efficiency guy, self-described. When you look at the complexity of these new systems, even the way compute, GPU, memory, how they're all designed around each other, that's where the efficiency comes in. But when you want to move from pilot to production, you got to deal with things, I wrote this on my LinkedIn post, token economics, governance, data sovereignty, operational simplicity rather than raw hardware specs. So that means the speeds and feeds don't matter, but they do matter when you have to take a holistic view of these systems. Talk about the importance of this, because this is an efficiency thing. We've heard model routing that helps on token economics. Sovereignty, Dave and I did a special report with Amit Govrin, our new analyst, on sovereignty. It's a whole other ball game. So all these things are factoring into the output, the work product, the value created. What's different than just, I got speeds and feeds, here's a new processor, and here's a new chip. What's the difference? What should people know about this new wave?
Suresh Andani
>> If you think about enterprises, they are more about, I'm paying for an outcome, I'm not paying for the chips, I'm not paying for a server, I'm really paying for an outcome. If I'm a big oil and gas company, and I have an issue with contract leakage, my outcome that I'm driving is can I use AI to reduce the dollars I am leaking on contract, that's what they're trying to drive. But it's also true that the TCO is also driven significantly by what's at the foundational layer. So speeds and feeds matter to bring the end customer value in terms of outcomes.
John Furrier
>> For the economics.
Suresh Andani
>> For the economics, yeah.
John Furrier
>> Yeah, so benchmarks, the old school, speeds and feeds, benchmarks, are moving to economic benchmarks.
Dave Vellante
>> Yeah, and so the point about outcomes and infrastructure is, large organizations and even mid -sized organizations, they're thinking about their infrastructure as a platform to build outcomes on top of. So they have to put in that core capex. that's almost just compulsory table stakes, and then the outcomes get built on top of that, what the agents are going to do, what projects they're going to run, and they're measuring that, they actually want to pay for that via outcomes and points and milestones.
Suresh Andani
>> Absolutely, and if you look at enterprises, exactly right, Dave, right? If you look at enterprises, some of them are DIYers. They can take a platform that we provide through one of our OEM partners or cloud partners, and then they can build the whole agent tech stack on top of it and they can build their applications. There are other enterprises who need platforms for AI, which we partner with Nutanix, our partners, we're going to talk with them next. And there are certain enterprises that need pre -built agents, like right out of the box, I need a chatbot agent. So for us, really, it's to deliver that whole stack from not just the GPU, CPU, and networking chips, the platforms, the control plane for AI, and then the actual ISV application layers. Right, so that's what we're trying to do.
Dave Vellante
>> You were talking about open weight models, your first comment. I want to run something by you. the reasoning behind open weight models we heard from Alex Karp's rant. You don't want to give up your alpha. If you turn it over to the LLMs, it's interesting. I heard, John, somebody on TV yesterday talking about these Chinese open weight models and they called it AI communism. It's actually the opposite. It's not, notwithstanding the possibility of malware. But the concept of open weight models is not AI communism, it's actually AI capitalism, and let me explain, and I want to get your feedback on that. AI communism was everybody gets the same general intelligence and applies it to their enterprise, and the large language models suck that knowledge out and then everybody gets it. I'm not saying that's what's going to happen, and I think they'll figure that out. But that's AI communism, where everybody gets the same.
John Furrier
>> Well, that's a -
Dave Vellante
>> Let me finish, and then AI capitalism is you can tune your own models via open weights. So I just found that ironic that the speaker sort of misrepresented, just because China is a communist country, that they actually, their open weight model philosophy is more along the lines of AI capitalism than it is on AI communism, if that makes sense.
John Furrier
>> Well, I'll comment, then I'll get Suresh's point, because I have an opinion on this. So the communism piece is really thinking about the general intelligence. But as I mentioned in the last one, Sagi from the venture group, Fireworks and AMD investment, they're going the other direction. They're specialized intelligence. So there's a power law of model functionality that's been decoupled from the infrastructure. So to me, the way we see this is that I can use general intelligence for general things, but if I have domain skills like robotics in a factory, I might want to use my own model there. Or if someone in open source creates an open weight model that does something really, really good, they can run all of them. So communism's bad in the sense of everyone gets bread and milk, but they can't have the fine foods or whatever, however you want to look at it. But when you get into real domain specific, the specialism shines. And we're seeing that with robotics, because safety's the number one concern. Sovereignty, they got to have regional boundaries, which is why it's a cloud problem.
Dave Vellante
>> So let's get to sovereignty, but so the question then Suresh is, what does that mean to you? If there is at least a near term and probably mid term and long term demand for open weight models, what does that mean for AMD? What requirements does that put on you?
Suresh Andani
>> Yeah, it's huge. you talked about AI capitalism and the way I look at it is open weights models are going to be critical for democratizing AI. I'm bringing another democracy in here, another political term. but a lot of CIOs I talk to and line of business GMs, they want a lot of control over end to end stack, not just what chips and servers I'm buying, but what models I want to use, because if they get locked into any one given model, they really are locked in. What by that is if a frontier model changes from version A to version B, you'd lose access to version A, which was running probably better for my workloads at the right economics. With their ability to access open-weight models and now the open-weight models are catching up to the performance of proprietary models, it gives these enterprises choice. AMD is all about openness and choice. So from my standpoint, when I talk about open-weight models, I think about how do we democratize AIfor everybody more than anything else.
Dave Vellante
>> So they can keep their alpha, and that brings us to the sovereignty comment. John and Amit wrote an awesome piece. I actually edited it and made some substantial changes, so I put my name on it too. But they laid out five dimensions of sovereignty. Territorial, operational, technological, legal, and financial. And the real point for the audience is that sovereignty is an ownership model. It's not like physically where the data residency is. It's much, much more than that. It's taking control of your stack.
Suresh Andani
>> That's right.
Dave Vellante
>> And then they came up with a system that allows us to actually evaluate the degree of sovereignty along those five lines. And you can pass some, you can maybe not so much on others, on a scale of one to 10. And so that is something that virtually every organization, not just international companies, want to take control of.
Suresh Andani
>> Yeah, no, absolutely. And sovereignty has again many, it's a word that's used in many different contexts. Control is one of them. We talk more about privacy, security, but a lot of enterprises want to govern what's fed to AI, what comes out of AI. So, it is as much about controlling your infrastructure, the models you want to use, as much as it is about privacy and security.
Dave Vellante
>> And then there's trade -offs too. If we're going to go completely air -gapped, we're probably going to give up some phone home benefits or some patching benefits. So, there's a spectrum there. I'm not arguing for a purist.
John Furrier
>> No, no, no, but I think the open -weight model conversation's important because it's not mutually exclusive, democratization of access is one. Cost, we talk to CFOs and CEOs all the time with theCUBE and theCUBE Research on our team and they don't ask how many teraflops am I getting, they're asking how much is inference going to cost? How can I predict my operating expenses? How do I keep my data sovereign? In fact, outcomes. So the open weights is a really good cost structure. Now with AMD, I don't have to buy big honking GPUs to run an open weight thin model that does something really well. And then I could use all the big resources for the bigger thing. So I think it's going to be an architectural ingredient balance. I need cost, price performance. so you go, oh, with open weights, we're not really hurting anyone doing customer support? It's really, really good at that, or as an example, versus reasoning and solving complex biology problems. we're talking about a whole nother level of infrastructure.
Dave Vellante
>> Are you saying that doesn't necessarily mean those specialized models don't necessarily have to be open weight, or are you advocating for open weight?
John Furrier
>> In robotics, for instance, the data we have on the robotics market is that in certain use cases where AI safety's important, I'll let you get your reaction on this, what you think, but they use their proprietary small language models for their domain. Because they've created a deterministic workflow that says this is a safety thing. Now they have other AI in there, but if we're using, say, a call center service, I might say, Here's all the unstructured data, use an open weight model. I don't want to break the bank on that, but I want to make sure I have this over here. So open weights over here, proprietary here. And for the rest of the stuff, the general models will work for general work. So you start to see this office suite-like approach, which is what users want. CIOs don't want to pay an arm and a leg to do, what's the weather like in San Francisco?
Dave Vellante
>> So Suresh, a long-winded setup for, you got news. Yeah. So what's the news and how does it fit into this whole discussion that we've been having?
Suresh Andani
>> Yeah, so we talked a lot about why tokenomics is becoming not just a CIO, CFO level discussion, it's becoming a board level discussion. Part of that is, if I'm able to go host open-weight models in my existing data center, which has been air cooled for the last 10 years, what is the GPU form factor that can fit, how can I bring AI to my data, not trying to take data to AI, and how can I bring it in my existing infrastructure which fits in a regular server, which is air cooled through a PCIe slot. So we announced our MI350P series GPUs in May, but today we are officially launching those at this event. What that does is it's a PCIe form factor card. It is an HBM class GPU that fits in your existing servers. It can run models all the way up to 260 billion parameters, really supporting more than 1,000 concurrent users at the same time, so it has enough KV cache to do deep context, high concurrency, and very low latency workloads. So a lot of enterprises whose rack power density is less than 30 kilowatts air cooled, this fits perfectly to get them the outcomes. That's most enterprises here. That's most. In terms of numbers. And edge. 85 % And edge.
John Furrier
>> And edge. And edge. Edge is huge. Talk about the, actually I did a survey on what people think a beefy edge is. 30 kilowatts is about the number. Seems to be the number. Talk about the scope of that value proposition. Because you talk about a PCIe card, fits into existing infrastructure. You mentioned some stats. Scope the equivalent value that would be, go back two years, a year ago. What would I be paying? Just scope and then map it to the capability.
Suresh Andani
>> If we're doing something like deep code generation where you have 20 ,000 engineers, you really need something bigger than a PCIe card. But if we're an enterprise running your business process your process automations, your summarization, your customer support, your workflow automation, all these workloads, this fits easily now. You can get some of those agents through the enterprise software that you have been buying, your CRMs, your ERPs, but they don't talk to each other, those systems. What this card allows you to do for an IT development team is to really build their own agents on their existing infrastructure, and now they can have their ERP system talk to their CRM system, talk to their IT, because these are all talking, you are controlling the whole data flow versus buying agents, license, seat-based license agents from your software vendor, right? So, it provides a huge value over there, productivity -wise. I don't want to say how it is productive and how many people's jobs it can do, because that's not the narrative here. It's about improving the productivity through the self-reliance.
John Furrier
>> Where they have workloads already scoped with resources. So I got my ERP running there, I got this system over there. They know their estate, they plug this in, they get instant AI.
Suresh Andani
>> If I'm an enterprise, my workflow is never like, if I have to go process a customer order, I have to go into my SAP system, I have to go into my Salesforce system, I have to go into my ITSM system, I have to go into my Jira, and really to put all these things together, you want control of the whole workflow and that's where you have full control over these agents you are self-building.
Dave Vellante
>> Well, and this sets up the system of intelligence conversation, what people call the ontology, because we're talking about infrastructure that allows you to connect to all those different systems and the way that organizations work today, let's face it, this whole idea of a single version of the truth and even determinism, it's a myth. It's determinism within each individual department
Suresh Andani
>> Yes.
Dave Vellante
>> And then when you get all the departments together, it's a big debate about where the data come from, and we have different data, and then you have to harmonize that data. You've got to make sure it's current, we're using the same definitions, the same taxonomy. But this, we're saying that at least that's the physical, the fundamental connection points are the first step in actually harmonizing all that data. Now you can write software that does that hard work, and some software out there already, Palantir obviously has some and others. And then that becomes a whole new operating model. Your point about jobs, okay fine, but to me the point is more, I can double, triple, maybe even 10x the output of my company with comparable labor. That's where it gets interesting.
Suresh Andani
>> That's exactly, so it's more about how do you make that workforce more productive, and your complex workflows in an enterprise talk to each other with ontology and the semantic layer that you talked about, and having full control of how all of these systems talk to each other. Because a single system agent, where yourworkflows are across multiple systems, hallucinates a lot. So this gives that.
John Furrier
>> And by the way, the number of agents that can be deployed actually increases the work.
Suresh Andani
>> Well, exactly.
John Furrier
>> And the cost, the theme in theCUBE the past six months has been clear on agents. If the cost of the AI is greater than the human capital, don't do it. So here, you got a solution that says plug in the card in your existing system.
Suresh Andani
>> There's no rate limiting you need to do once you pay for the capex. We have studies which say the payback period for an enterprise who is spending a billion tokens a day is now in the six month range.
John Furrier
>> By the way, we are from the school of thought that jobs will increase because in cyber security, other areas, all the quote job displacement actually just shifts the jobs. So we're seeing a lot more organizing AI and managing AI than actually being a domain worker.
Dave Vellante
>> Well, it enables the organization to capture the tacit knowledge and ingest that tacit knowledge of the enterprise so it's preserved. There's now a corporate memory, if you will. The vision we have is a digital representation of the enterprise, a digital twin, a term that we use a lot. That digital, real-time digital twin, that's going to drive a lot of compute.
Suresh Andani
>> No, for us exactly, it's that, Dave, like you mentioned, right? We want to start with if you can run your workload on a Ryzen AI PRO laptop, go do that. Why not? But there are tasks that will need to go out. Then you go into your centralized token router, which can help you. And we are providing those solutions to our enterprise customers to say, I can run this on a PCIe card in my data center. But if it's code generation, I probably might need an MI300X eight-way server in my data center. And still, if I need more concurrency, more context, I can escape to the frontier APIs to go run it. So this democratizes where you are running your workloads.
Dave Vellante
>> Right, that's a key point.
Suresh Andani
>> I can be my own token generator, and take control of my own cost, or if I want to go through an API, that's great too, and there's certain workloads that are going to require that. Yeah, you start becoming from total token consumer to a token generator and serving your enterprise.
John Furrier
>> And that's going to create creativity, takes the tokenomics problem off the table.
Suresh Andani
>> That's right.
John Furrier
>> All right, just to summarize the news. You already announced the product. It's shipping, just repeat the news again.
Suresh Andani
>> Yeah, so we are launching today. We are starting this week, we are shipping production ready silicon to a lot of our OEM and cloud partners. And you will see their solutions, the commercial servers and cloud instances coming out in Q3 and Q4 time frame.
John Furrier
>> What's the early data on the early adopters in beta? So what operationally, value -wise, can you share some anecdotes?
Suresh Andani
>> Yeah, absolutely. We have 30 plus, we just launched today, we have 30 plus customers already waiting to test the MI350P, and those enterprises vary from FSIs to oil and gas to legal to like across the board, to healthcare, retail, across the verticals, a lot of them are using it for use cases which are pretty simple. It's a token factory. You are trying to deliver.
John Furrier
>> They're injecting intelligence into their pre -existing applications.
Suresh Andani
>> Exactly, they're building their, exactly. They have their applications, they're agentifying those applications. What they're looking for is a layer that can give me best token per second, per dollar, per watt in my infrastructure that I already own through the software stack that I already run. That's what we are solving for them.
Dave Vellante
>> And they want to build that in a sovereign manner so they can take control over that.
Suresh Andani
>> Absolutely.
John Furrier
>> Well Suresh, great to see you, and congratulations on the news. You got a very busy job. You're the man about town on getting compute in the enterprise, which we think is going to be a really incredible 2027. We've seen some slowdown, mainly because the coding's come in, everyone loves that, the token maxing, they're reining that in. Just quickly to end the segment, how do you see the outlook for the enterprise? Do you feel the same buzz we feel? Certainly enthusiasm, but confidence is getting there. What's your outlook for the year, just from a macro standpoint? Floodgates opening for agents, what's the forecast?
Suresh Andani
>> Yeah, so agent tech is really driving real outcomes for the enterprises now. Up till now, as we all know, AI was more about training the models. But we are seeing now enterprises across different verticals, literally the first second I walk into an enterprise CIO discussion, they want to just talk about, how can you agentify my workload? So there's a lot of excitement around this.
John Furrier
>> And I don't want to pay out the nose for it too, that's another thing.
Suresh Andani
>> Exactly, and I think they've been stuck in the pilot phase because A, they didn't have the right platforms, they didn't have the right solution stacks, that's all coming together now, we are working very closely, we cannot do this alone. We're doing all of this with our ecosystem partners, We're going to talk to Tarkan very soon, Nutanix, and what we are doing together, a lot of other ISVs, to bring outcome -based solutions to our enterprises, and they are seeing that now.
Dave Vellante
>> And a lot of customers just don't have the skills, and so they're looking for the vendor R &D, if you will, to close those gaps, and that's exactly what's happened over the last 12 to 18 months.
Suresh Andani
>> Yeah, I can go and say, I can reduce your contract leakage issue by 70%, save you $250 million. That goes a long way than if I say, I'm going to give you MI350P.
John Furrier
>> You're actually guaranteeing that, from what I heard. At FinOps this past couple, last month, Mike was there and your team, they're like, we pretty much guarantee the savings.
Suresh Andani
>> Yeah. Yeah.
John Furrier
>> That's the confidence you have.
Suresh Andani
>> That's right. A lot of cloud savings we have talked about, Mike from my team talked a lot about that. Now we are talking about the same story, like AI for FinOps and FinOps for AI story. If we're using AI to better FinOps and you'll be able to reduce the AI costs. A lot of excitement.
John Furrier
>> Well, we're Advancing AI, that's the name of the event, so it's great to see. We'll see you at theCUBE + NYSE Wired, our new brand's pool party, third annual Infrastructure Leaders on the 28th. Thanks for participating, and again, the leaders are being featured. Again, we appreciate the work you do, and go faster. Come on, the world wants more.
Suresh Andani
>> Yeah, that's what I hear from Lisa every day.
John Furrier
>> We'll see you next week at the NYSE Wired Pool Party. I'm John Furrier with Dave Vellante, doing our part on the ground here in San Francisco, getting all the data, sharing it as fast as possible, as always on theCUBE. We've got 10 interviews today, 10 tomorrow, so stay with us for more after this short break.
>> Welcome back to theCUBE live stream here in San Francisco. I'm John Furrier, host of theCUBE with Dave Vellante, my co -host. We're here for AMD's Advancing AI event, global event. We have all the industry players here, as well as some free agents. We have a lot of practitioners, a lot of engineers here, people building out the next generation of AI infrastructure and all the applications on top. Suresh Andani's here, Corporate Vice President of Compute and Enterprise AI for AMD, he's the man who's going to bring all that compute magic to the marketplace and enable the enterprises to go to the next level. It's great to see you.
Suresh Andani
>> Likewise.
Dave Vellante
>> Thanks for coming back on theCUBE. What's going on?
John Furrier
>> Alright, so I got to get into this because the enterprise has been waiting, we're seeing inference all over the place, starting to see people resettling in. This has been a shift, we've been documenting this shift. This is a generational shift where now the computing industry has moved over to a net new architecture. It's cloud, it's on premise, it's edge, all of it is distributed computing, hybrid. This is where compute, GPUs, FPGAs, all have to work together in concert. It's a whole new system.
Suresh Andani
>> Yes, yeah, no, you're absolutely right, John. So in traditional sense, we talked about hybrid only two ways, right? Either we're running on-prem or in the edge or in the cloud, right? Now when we talk about more distributed AI, it's not just where we're running the workload, but what models you are running are also hybrid. A lot of enterprises are now figuring out that, not all tasks need frontier models. frontier models are great, by the way. If you need the deepest context, if you really need the high concurrency, frontier models have really done well for the enterprises, but not every enterprise AI task needs to run on a frontier model. A lot of tasks because of token economics, because of security reasons, data sovereignty reasons and control reasons are better suited to run on-prem on open-weight models. So as a provider of AI compute, we want to make sure that we are enabling those enterprises to not only run their AI tasks through the frontier APIs in the cloud, but we are providing them infrastructure that they can host open-weight models very efficiently on-prem. And that's a key, every enterprise conversation I'm in, that's a key topic of discussion.
John Furrier
>> Two weeks ago I interviewed Mark Papermaster at the RAISE Summit. He's 40 years veteran CTO at AMD, you know him very well. He's an efficiency guy, self-described. When you look at the complexity of these new systems, even the way compute, GPU, memory, how they're all designed around each other, that's where the efficiency comes in. But when you want to move from pilot to production, you got to deal with things, I wrote this on my LinkedIn post, token economics, governance, data sovereignty, operational simplicity rather than raw hardware specs. So that means the speeds and feeds don't matter, but they do matter when you have to take a holistic view of these systems. Talk about the importance of this, because this is an efficiency thing. We've heard model routing that helps on token economics. Sovereignty, Dave and I did a special report with Amit Govrin, our new analyst, on sovereignty. It's a whole other ball game. So all these things are factoring into the output, the work product, the value created. What's different than just, I got speeds and feeds, here's a new processor, and here's a new chip. What's the difference? What should people know about this new wave?
Suresh Andani
>> If you think about enterprises, they are more about, I'm paying for an outcome, I'm not paying for the chips, I'm not paying for a server, I'm really paying for an outcome. If I'm a big oil and gas company, and I have an issue with contract leakage, my outcome that I'm driving is can I use AI to reduce the dollars I am leaking on contract, that's what they're trying to drive. But it's also true that the TCO is also driven significantly by what's at the foundational layer. So speeds and feeds matter to bring the end customer value in terms of outcomes.
John Furrier
>> For the economics.
Suresh Andani
>> For the economics, yeah.
John Furrier
>> Yeah, so benchmarks, the old school, speeds and feeds, benchmarks, are moving to economic benchmarks.
Dave Vellante
>> Yeah, and so the point about outcomes and infrastructure is, large organizations and even mid -sized organizations, they're thinking about their infrastructure as a platform to build outcomes on top of. So they have to put in that core capex. that's almost just compulsory table stakes, and then the outcomes get built on top of that, what the agents are going to do, what projects they're going to run, and they're measuring that, they actually want to pay for that via outcomes and points and milestones.
Suresh Andani
>> Absolutely, and if you look at enterprises, exactly right, Dave, right? If you look at enterprises, some of them are DIYers. They can take a platform that we provide through one of our OEM partners or cloud partners, and then they can build the whole agent tech stack on top of it and they can build their applications. There are other enterprises who need platforms for AI, which we partner with Nutanix, our partners, we're going to talk with them next. And there are certain enterprises that need pre -built agents, like right out of the box, I need a chatbot agent. So for us, really, it's to deliver that whole stack from not just the GPU, CPU, and networking chips, the platforms, the control plane for AI, and then the actual ISV application layers. Right, so that's what we're trying to do.
Dave Vellante
>> You were talking about open weight models, your first comment. I want to run something by you. the reasoning behind open weight models we heard from Alex Karp's rant. You don't want to give up your alpha. If you turn it over to the LLMs, it's interesting. I heard, John, somebody on TV yesterday talking about these Chinese open weight models and they called it AI communism. It's actually the opposite. It's not, notwithstanding the possibility of malware. But the concept of open weight models is not AI communism, it's actually AI capitalism, and let me explain, and I want to get your feedback on that. AI communism was everybody gets the same general intelligence and applies it to their enterprise, and the large language models suck that knowledge out and then everybody gets it. I'm not saying that's what's going to happen, and I think they'll figure that out. But that's AI communism, where everybody gets the same.
John Furrier
>> Well, that's a -
Dave Vellante
>> Let me finish, and then AI capitalism is you can tune your own models via open weights. So I just found that ironic that the speaker sort of misrepresented, just because China is a communist country, that they actually, their open weight model philosophy is more along the lines of AI capitalism than it is on AI communism, if that makes sense.
John Furrier
>> Well, I'll comment, then I'll get Suresh's point, because I have an opinion on this. So the communism piece is really thinking about the general intelligence. But as I mentioned in the last one, Sagi from the venture group, Fireworks and AMD investment, they're going the other direction. They're specialized intelligence. So there's a power law of model functionality that's been decoupled from the infrastructure. So to me, the way we see this is that I can use general intelligence for general things, but if I have domain skills like robotics in a factory, I might want to use my own model there. Or if someone in open source creates an open weight model that does something really, really good, they can run all of them. So communism's bad in the sense of everyone gets bread and milk, but they can't have the fine foods or whatever, however you want to look at it. But when you get into real domain specific, the specialism shines. And we're seeing that with robotics, because safety's the number one concern. Sovereignty, they got to have regional boundaries, which is why it's a cloud problem.
Dave Vellante
>> So let's get to sovereignty, but so the question then Suresh is, what does that mean to you? If there is at least a near term and probably mid term and long term demand for open weight models, what does that mean for AMD? What requirements does that put on you?
Suresh Andani
>> Yeah, it's huge. you talked about AI capitalism and the way I look at it is open weights models are going to be critical for democratizing AI. I'm bringing another democracy in here, another political term. but a lot of CIOs I talk to and line of business GMs, they want a lot of control over end to end stack, not just what chips and servers I'm buying, but what models I want to use, because if they get locked into any one given model, they really are locked in. What by that is if a frontier model changes from version A to version B, you'd lose access to version A, which was running probably better for my workloads at the right economics. With their ability to access open-weight models and now the open-weight models are catching up to the performance of proprietary models, it gives these enterprises choice. AMD is all about openness and choice. So from my standpoint, when I talk about open-weight models, I think about how do we democratize AIfor everybody more than anything else.
Dave Vellante
>> So they can keep their alpha, and that brings us to the sovereignty comment. John and Amit wrote an awesome piece. I actually edited it and made some substantial changes, so I put my name on it too. But they laid out five dimensions of sovereignty. Territorial, operational, technological, legal, and financial. And the real point for the audience is that sovereignty is an ownership model. It's not like physically where the data residency is. It's much, much more than that. It's taking control of your stack.
Suresh Andani
>> That's right.
Dave Vellante
>> And then they came up with a system that allows us to actually evaluate the degree of sovereignty along those five lines. And you can pass some, you can maybe not so much on others, on a scale of one to 10. And so that is something that virtually every organization, not just international companies, want to take control of.
Suresh Andani
>> Yeah, no, absolutely. And sovereignty has again many, it's a word that's used in many different contexts. Control is one of them. We talk more about privacy, security, but a lot of enterprises want to govern what's fed to AI, what comes out of AI. So, it is as much about controlling your infrastructure, the models you want to use, as much as it is about privacy and security.
Dave Vellante
>> And then there's trade -offs too. If we're going to go completely air -gapped, we're probably going to give up some phone home benefits or some patching benefits. So, there's a spectrum there. I'm not arguing for a purist.
John Furrier
>> No, no, no, but I think the open -weight model conversation's important because it's not mutually exclusive, democratization of access is one. Cost, we talk to CFOs and CEOs all the time with theCUBE and theCUBE Research on our team and they don't ask how many teraflops am I getting, they're asking how much is inference going to cost? How can I predict my operating expenses? How do I keep my data sovereign? In fact, outcomes. So the open weights is a really good cost structure. Now with AMD, I don't have to buy big honking GPUs to run an open weight thin model that does something really well. And then I could use all the big resources for the bigger thing. So I think it's going to be an architectural ingredient balance. I need cost, price performance. so you go, oh, with open weights, we're not really hurting anyone doing customer support? It's really, really good at that, or as an example, versus reasoning and solving complex biology problems. we're talking about a whole nother level of infrastructure.
Dave Vellante
>> Are you saying that doesn't necessarily mean those specialized models don't necessarily have to be open weight, or are you advocating for open weight?
John Furrier
>> In robotics, for instance, the data we have on the robotics market is that in certain use cases where AI safety's important, I'll let you get your reaction on this, what you think, but they use their proprietary small language models for their domain. Because they've created a deterministic workflow that says this is a safety thing. Now they have other AI in there, but if we're using, say, a call center service, I might say, Here's all the unstructured data, use an open weight model. I don't want to break the bank on that, but I want to make sure I have this over here. So open weights over here, proprietary here. And for the rest of the stuff, the general models will work for general work. So you start to see this office suite-like approach, which is what users want. CIOs don't want to pay an arm and a leg to do, what's the weather like in San Francisco?
Dave Vellante
>> So Suresh, a long-winded setup for, you got news. Yeah. So what's the news and how does it fit into this whole discussion that we've been having?
Suresh Andani
>> Yeah, so we talked a lot about why tokenomics is becoming not just a CIO, CFO level discussion, it's becoming a board level discussion. Part of that is, if I'm able to go host open-weight models in my existing data center, which has been air cooled for the last 10 years, what is the GPU form factor that can fit, how can I bring AI to my data, not trying to take data to AI, and how can I bring it in my existing infrastructure which fits in a regular server, which is air cooled through a PCIe slot. So we announced our MI350P series GPUs in May, but today we are officially launching those at this event. What that does is it's a PCIe form factor card. It is an HBM class GPU that fits in your existing servers. It can run models all the way up to 260 billion parameters, really supporting more than 1,000 concurrent users at the same time, so it has enough KV cache to do deep context, high concurrency, and very low latency workloads. So a lot of enterprises whose rack power density is less than 30 kilowatts air cooled, this fits perfectly to get them the outcomes. That's most enterprises here. That's most. In terms of numbers. And edge. 85 % And edge.
John Furrier
>> And edge. And edge. Edge is huge. Talk about the, actually I did a survey on what people think a beefy edge is. 30 kilowatts is about the number. Seems to be the number. Talk about the scope of that value proposition. Because you talk about a PCIe card, fits into existing infrastructure. You mentioned some stats. Scope the equivalent value that would be, go back two years, a year ago. What would I be paying? Just scope and then map it to the capability.
Suresh Andani
>> If we're doing something like deep code generation where you have 20 ,000 engineers, you really need something bigger than a PCIe card. But if we're an enterprise running your business process your process automations, your summarization, your customer support, your workflow automation, all these workloads, this fits easily now. You can get some of those agents through the enterprise software that you have been buying, your CRMs, your ERPs, but they don't talk to each other, those systems. What this card allows you to do for an IT development team is to really build their own agents on their existing infrastructure, and now they can have their ERP system talk to their CRM system, talk to their IT, because these are all talking, you are controlling the whole data flow versus buying agents, license, seat-based license agents from your software vendor, right? So, it provides a huge value over there, productivity -wise. I don't want to say how it is productive and how many people's jobs it can do, because that's not the narrative here. It's about improving the productivity through the self-reliance.
John Furrier
>> Where they have workloads already scoped with resources. So I got my ERP running there, I got this system over there. They know their estate, they plug this in, they get instant AI.
Suresh Andani
>> If I'm an enterprise, my workflow is never like, if I have to go process a customer order, I have to go into my SAP system, I have to go into my Salesforce system, I have to go into my ITSM system, I have to go into my Jira, and really to put all these things together, you want control of the whole workflow and that's where you have full control over these agents you are self-building.
Dave Vellante
>> Well, and this sets up the system of intelligence conversation, what people call the ontology, because we're talking about infrastructure that allows you to connect to all those different systems and the way that organizations work today, let's face it, this whole idea of a single version of the truth and even determinism, it's a myth. It's determinism within each individual department
Suresh Andani
>> Yes.
Dave Vellante
>> And then when you get all the departments together, it's a big debate about where the data come from, and we have different data, and then you have to harmonize that data. You've got to make sure it's current, we're using the same definitions, the same taxonomy. But this, we're saying that at least that's the physical, the fundamental connection points are the first step in actually harmonizing all that data. Now you can write software that does that hard work, and some software out there already, Palantir obviously has some and others. And then that becomes a whole new operating model. Your point about jobs, okay fine, but to me the point is more, I can double, triple, maybe even 10x the output of my company with comparable labor. That's where it gets interesting.
Suresh Andani
>> That's exactly, so it's more about how do you make that workforce more productive, and your complex workflows in an enterprise talk to each other with ontology and the semantic layer that you talked about, and having full control of how all of these systems talk to each other. Because a single system agent, where yourworkflows are across multiple systems, hallucinates a lot. So this gives that.
John Furrier
>> And by the way, the number of agents that can be deployed actually increases the work.
Suresh Andani
>> Well, exactly.
John Furrier
>> And the cost, the theme in theCUBE the past six months has been clear on agents. If the cost of the AI is greater than the human capital, don't do it. So here, you got a solution that says plug in the card in your existing system.
Suresh Andani
>> There's no rate limiting you need to do once you pay for the capex. We have studies which say the payback period for an enterprise who is spending a billion tokens a day is now in the six month range.
John Furrier
>> By the way, we are from the school of thought that jobs will increase because in cyber security, other areas, all the quote job displacement actually just shifts the jobs. So we're seeing a lot more organizing AI and managing AI than actually being a domain worker.
Dave Vellante
>> Well, it enables the organization to capture the tacit knowledge and ingest that tacit knowledge of the enterprise so it's preserved. There's now a corporate memory, if you will. The vision we have is a digital representation of the enterprise, a digital twin, a term that we use a lot. That digital, real-time digital twin, that's going to drive a lot of compute.
Suresh Andani
>> No, for us exactly, it's that, Dave, like you mentioned, right? We want to start with if you can run your workload on a Ryzen AI PRO laptop, go do that. Why not? But there are tasks that will need to go out. Then you go into your centralized token router, which can help you. And we are providing those solutions to our enterprise customers to say, I can run this on a PCIe card in my data center. But if it's code generation, I probably might need an MI300X eight-way server in my data center. And still, if I need more concurrency, more context, I can escape to the frontier APIs to go run it. So this democratizes where you are running your workloads.
Dave Vellante
>> Right, that's a key point.
Suresh Andani
>> I can be my own token generator, and take control of my own cost, or if I want to go through an API, that's great too, and there's certain workloads that are going to require that. Yeah, you start becoming from total token consumer to a token generator and serving your enterprise.
John Furrier
>> And that's going to create creativity, takes the tokenomics problem off the table.
Suresh Andani
>> That's right.
John Furrier
>> All right, just to summarize the news. You already announced the product. It's shipping, just repeat the news again.
Suresh Andani
>> Yeah, so we are launching today. We are starting this week, we are shipping production ready silicon to a lot of our OEM and cloud partners. And you will see their solutions, the commercial servers and cloud instances coming out in Q3 and Q4 time frame.
John Furrier
>> What's the early data on the early adopters in beta? So what operationally, value -wise, can you share some anecdotes?
Suresh Andani
>> Yeah, absolutely. We have 30 plus, we just launched today, we have 30 plus customers already waiting to test the MI350P, and those enterprises vary from FSIs to oil and gas to legal to like across the board, to healthcare, retail, across the verticals, a lot of them are using it for use cases which are pretty simple. It's a token factory. You are trying to deliver.
John Furrier
>> They're injecting intelligence into their pre -existing applications.
Suresh Andani
>> Exactly, they're building their, exactly. They have their applications, they're agentifying those applications. What they're looking for is a layer that can give me best token per second, per dollar, per watt in my infrastructure that I already own through the software stack that I already run. That's what we are solving for them.
Dave Vellante
>> And they want to build that in a sovereign manner so they can take control over that.
Suresh Andani
>> Absolutely.
John Furrier
>> Well Suresh, great to see you, and congratulations on the news. You got a very busy job. You're the man about town on getting compute in the enterprise, which we think is going to be a really incredible 2027. We've seen some slowdown, mainly because the coding's come in, everyone loves that, the token maxing, they're reining that in. Just quickly to end the segment, how do you see the outlook for the enterprise? Do you feel the same buzz we feel? Certainly enthusiasm, but confidence is getting there. What's your outlook for the year, just from a macro standpoint? Floodgates opening for agents, what's the forecast?
Suresh Andani
>> Yeah, so agent tech is really driving real outcomes for the enterprises now. Up till now, as we all know, AI was more about training the models. But we are seeing now enterprises across different verticals, literally the first second I walk into an enterprise CIO discussion, they want to just talk about, how can you agentify my workload? So there's a lot of excitement around this.
John Furrier
>> And I don't want to pay out the nose for it too, that's another thing.
Suresh Andani
>> Exactly, and I think they've been stuck in the pilot phase because A, they didn't have the right platforms, they didn't have the right solution stacks, that's all coming together now, we are working very closely, we cannot do this alone. We're doing all of this with our ecosystem partners, We're going to talk to Tarkan very soon, Nutanix, and what we are doing together, a lot of other ISVs, to bring outcome -based solutions to our enterprises, and they are seeing that now.
Dave Vellante
>> And a lot of customers just don't have the skills, and so they're looking for the vendor R &D, if you will, to close those gaps, and that's exactly what's happened over the last 12 to 18 months.
Suresh Andani
>> Yeah, I can go and say, I can reduce your contract leakage issue by 70%, save you $250 million. That goes a long way than if I say, I'm going to give you MI350P.
John Furrier
>> You're actually guaranteeing that, from what I heard. At FinOps this past couple, last month, Mike was there and your team, they're like, we pretty much guarantee the savings.
Suresh Andani
>> Yeah. Yeah.
John Furrier
>> That's the confidence you have.
Suresh Andani
>> That's right. A lot of cloud savings we have talked about, Mike from my team talked a lot about that. Now we are talking about the same story, like AI for FinOps and FinOps for AI story. If we're using AI to better FinOps and you'll be able to reduce the AI costs. A lot of excitement.
John Furrier
>> Well, we're Advancing AI, that's the name of the event, so it's great to see. We'll see you at theCUBE + NYSE Wired, our new brand's pool party, third annual Infrastructure Leaders on the 28th. Thanks for participating, and again, the leaders are being featured. Again, we appreciate the work you do, and go faster. Come on, the world wants more.
Suresh Andani
>> Yeah, that's what I hear from Lisa every day.
John Furrier
>> We'll see you next week at the NYSE Wired Pool Party. I'm John Furrier with Dave Vellante, doing our part on the ground here in San Francisco, getting all the data, sharing it as fast as possible, as always on theCUBE. We've got 10 interviews today, 10 tomorrow, so stay with us for more after this short break.