This discussion examines benchmarking for artificial intelligence infrastructure and agentic workloads with MLPerf and MLCommons. David Kanter of MLCommons, founder and head of MLPerf, explains MLPerf’s origins and its role in establishing trusted benchmarks for speed, energy efficiency and reliability across the industry. Kanter outlines recent work such as MLPerf Endpoints v0.7 and describes the shift from hardware-centric metrics to software-driven application-aware measurements; they emphasize the importance of benchmarks that are current, comparable, comprehensive and contextualized.
theCUBE Research hosts John Furrier of theCUBE and Dave Vellante of theCUBE frame the conversation around enterprise deployment, agentic AI and evolving use cases. Key takeaways include the 4C approach to benchmarking, forthcoming agentic benchmarks focused on code development, Q&A and customer support, and the need to measure trade-offs among speed, accuracy, cost and energy. Furrier and Vellante discuss how these metrics inform CIO and CISO decisions on infrastructure selection, governance and deployment strategies for enterprise and edge computing.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Mixture of Experts Series. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Mixture of Experts Series.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For theCUBE + NYSE Wired: Mixture of Experts Series
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for theCUBE + NYSE Wired: Mixture of Experts Series.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Mixture of Experts Series. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Mixture of Experts Series.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: Mixture of Experts Series
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: Mixture of Experts Series. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
David Kanter, ML Commons
This discussion examines benchmarking for artificial intelligence infrastructure and agentic workloads with MLPerf and MLCommons. David Kanter of MLCommons, founder and head of MLPerf, explains MLPerf’s origins and its role in establishing trusted benchmarks for speed, energy efficiency and reliability across the industry. Kanter outlines recent work such as MLPerf Endpoints v0.7 and describes the shift from hardware-centric metrics to software-driven application-aware measurements; they emphasize the importance of benchmarks that are current, comparable, comprehensive and contextualized.
theCUBE Research hosts John Furrier of theCUBE and Dave Vellante of theCUBE frame the conversation around enterprise deployment, agentic AI and evolving use cases. Key takeaways include the 4C approach to benchmarking, forthcoming agentic benchmarks focused on code development, Q&A and customer support, and the need to measure trade-offs among speed, accuracy, cost and energy. Furrier and Vellante discuss how these metrics inform CIO and CISO decisions on infrastructure selection, governance and deployment strategies for enterprise and edge computing.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, co-host of theCUBE, here with Dave Vellante, my co-host. Hello, I'm John Furrier, your host of theCUBE, here at theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street, part of our NYSE Wired program and community. This is our mixture of experts series where we bring in people who are experts in their field, doing great work and innovating. David Kanter's here. He's the co-founder of MLCommons and head of MLPerf. If you know all about machine learning, you know what that organization has done. Nonprofit doing really amazing work helping people figure out what's safe, what's real, what's not. David, great to see you. Thanks for coming on theCUBE. Saw you at AMD's event in San Francisco.
David Kanter
>> Absolutely a pleasure. It's great to be
John Furrier
>> here. You get to be mixed up with all the other expertsYeah, yeah. This is kind of a— originally was kind of a goof on AI when we did the series, but it's actually a great way to bring in our community and kind of make sure experts kind of share in the data. And one of the things that everyone loves about AI is that it's got a great utility. But pre the transformer technology, machine learning has been around for a long, long time. You know, fraud detection, every bank has it, supervised, unsupervised machine learning. It really was the genesis of what got all the deep tech nerds in the labs, looking at what's coming out. And then the field just went supernova from there. So it's really a valuable organization and lesson and also a template for the future. Explain what you do at MLCommons and MLPerf, how it all came together. What is it for people who don't know what it is and what it does and where it is?
David Kanter
>> Absolutely. So we got started in 2018, back pre-ChatGPT, pre-generative AI, and everyone was looking at we knew these AI models could do incredible things. We wanted to improve performance, to improve capabilities, but there was no standard way of measuring things. And so a group of us all came together from industry and academia to build that standard set of benchmarks to measure speed and energy efficiency, and that became MLPerf. And before that—
John Furrier
>> and by the way, that then became the cited benchmark stat in every presentation at that time.
David Kanter
>> Exactly right. And part of the thing that's really wonderful about getting to be involved in this group is we bring together everyone from all across the industry and through consensus we build these trusted standards like MLPerf. And, you know, at the time it would almost be as if you were buying a car and one guy says, hey, my car can do 0 to 60 in a second. The next guy says, my car has a turn signal. And the third guy says, I've got airbags. Which one do you want to buy? You don't really know. And so you need some way to compare them all. I live in San Francisco, so I might go for the airbags. But that was the genesis of MLPerf. And then we realized sort of the impact you can have to help drive the whole industry. And we said we should put this into a nonprofit. And then look for other ways that we can deploy our expertise in measurement and data to help make AI better, right? And so after we first did performance, we then zeroed in on measuring power efficiency, building large open datasets, and then over time we've started looking at benchmarks in risk and reliability of helping to make sure that the outputs of generative models are kind of in line with what we want.
John Furrier
>> The evolution of AI now is the number one conversation is safety, right? And then you see the Anthropic versus, say, OpenAI approach, fast and loose, more conservative. There's a general consensus, a lot of consensus around no one really knows what the hell that means. So take us through kind of what you guys are focused on now because you guys have the playbook on open AI. We see the success of open source. damn, it's the most successful trend ever in the computer industry. Look at what it's done. Now you've got OpenWeights. So you got a lot of open things happening. What are you guys focused on now? How is this translating into some of the conversations today?
David Kanter
>> Yeah, so I'd say one of the most critical things is when you're looking at anything, whether it's performance or risk or responsibility, it's about measuring it in the right way. Having written down what you're doing, what you're trying to accomplish, how much precision you have. And so for us, one of the things that we released last week was MLPerf Endpoints v0.7, which is a rethinking of our inference benchmarks for the modern era.
David Kanter
>> Right.
David Kanter
>> And we see this race to deploy as everyone's discovered that there's so many valuable things you can do with AI. How do we deploy it across the enterprise for consumers? We see building more data centers, needing more power and more performant systems. So we had to evolve our benchmarks to match that pace of innovation and to help customers really make the decisions that they need. You look at a Fortune 500 company.
David Kanter
>> Yeah.
David Kanter
>> They're not just saying, hey, I want AI for the C-suite. They're saying I have dozens or hundreds of applications that I'm going to deploy. Each one's different. How do I find the infrastructure that's going to pair up in the right way? And so I look at that as being our job is how do we build the tools that help empower those folks to make the decisions to help deploy AI?
John Furrier
>> So you're really kind of taking the DNA of MLPerf, MLCommons, machine learning.
John Furrier
>> Yeah
John Furrier
>> . Applying that to the AI growth wave, which is infrastructure and trying to help people navigate that. So I love that. The question that we're seeing now is I just had— I just wrote a post, went to lunch, actually, I wrote a post, but on agents and a lot of the fear with agents is you have more nondeterministic workloads. Again, that's cool, but you have workloads. And they're different. So now you have different conditions. What's the scope of some of the things you guys are getting your arms around in the open? Because a decision for Company A will be different than Company B because I might want to have more compute, less GPU, or pre-fill decode. all these things are kind of now coming into the systems. It's not a clear general-purpose benchmarking market. And so how do you guys think about that? What's the community doing on— is there any Data you can share, thoughts, personal thoughts?
David Kanter
>> Thoughts for sure. And stay tuned. We'll have data later this year for sure. But I think one of the things you really touched on is when we see blending AI inference with standard computing workloads through agentic flows, right? The world is your oyster, right? Before it was like, oh, maybe you're doing recommendation or translation. Well, Now you might be pairing that with hey, this is the conventional workflow that I have in my bank, but now I'm going to stick AI in here to accomplish my goal. Or, of course, the thing that we've seen the most demand for is coding.
John Furrier
>> Yeah, right
David Kanter
>> . And it sort of makes sense. The folks who are developing the tools are like, hey, wait, I can do what with this? Like, let's get some acceleration.
John Furrier
>> Hugging Face. Hello. Testbed went off the rails again. There are so many use cases where it could go off. Yeah, may or may not be related to anything other than the environment.
David Kanter
>> That's right. And one of the core insights in MLPerf actually was it's not just about speed, but it's about how fast you get to the right answer. And to the point you made earlier, you could achieve the same task with a lot of accelerated inference compute and maybe less conventional compute. Or maybe a different balance. And it depends on what you have. Imagine you want to pick what's the best South Indian restaurant in Manhattan? You could look at the top 10 if you're really confident of that top 10, or maybe you've got a system that's even more accurate and you just say, I really need the top 3 from this system, so I don't need to read all those reviews. Both of those will hopefully get me to a delicious meal. But what's the right path? That's really tricky. And so we're just in the starting stages.
John Furrier
>> That's a really good point. I think what you just said was compelling because most people look at the user experience of, say, ChatGPT, which most consumers experience as getting a good answer fast. The first answer, kind of search results. Yeah. Hey, that's great. Where do I find food? Boom. Answers. Reasoning is a different— it's not a search paradigm. Are you doing discovery?
John Furrier
>> That's right.
John Furrier
>> With multi-step reasoning, which is completely different. So it's a very nuanced point, but it changes the configuration of the data, what systems I might want to use. Maybe it's a complex answer. Maybe you have certain info needs that might require a Pareto curve of the variable routing, right? Who knows? Cha-ching on the tokens. So this is a cost trade-off. This is a math equation.
David Kanter
>> That's exactly right. And so part of our goal with MLPerf Endpoints is how do we present that trade-off so that if you're a CIO or a CISO, you can say, all right, I've got all these applications, these need to be really fast. These are for frontline. I need an answer for the customer quick. These other things may be back office are going to look different. And so how can you get the right infrastructure for everything? But I think to your point, the world of agentic, before you had AI walled off and it was its own thing, you had separate AI infrastructure people. But now, one of the things I used to say is to me, inference is a lot like salt in cooking. You don't eat pure salt usually.
John Furrier
>> Yeah
David Kanter
>> , but if you know from any recipe book, you add a little bit of salt and it makes almost everything taste better. And that's what we're seeing. So now it's not AI here and regular computing there. It's all braided together and all throughout. So almost everything that we have today in the enterprise is going to be recast with an agentic side. And so we're still in the early stages. And when I think about what we want to do, we have an agentic benchmark coming out later this year focusing on some of the things that we think are most popular, code development and software engineering, as well as Q&A and customer support. But I wouldn't be surprised if we have dozens of use cases in the future as we're discovering them live.
John Furrier
>> You know, you and I, when we were chatting at the AMD event in San Francisco, we were talking about some of the historical views. We've lived through many cycles of innovation. This one's obviously the most kick-ass ever because it's got everything popping. You got infrastructure up and down the stack. But in the old days when I was breaking into the business at Hewlett-Packard and before that IBM, PCs and servers, they all had the benchmarks. And that was really twofold. One, to do an industry service to kind of level the playing field on horses on the track, apples to apples, making sure everyone knows what's what. But it also helped customers scope what they wanted to buy. Yeah, it was really an economic beacon too for the, okay, I need a mid-range system for these desktops or whatever. When you get to AI, a lot of that's kind of going on, but it's not as simple. What's the biggest change in your mind today trying to rally the industry around MLCommons while looking at the aperture of use cases. Do you guys look at that as an opportunity? Do you— are there some first principles and then playbook tactics you guys are using? Because everyone wants the same thing. What do I buy it for? What and when? I don't want to waste any money. I don't want GPU cycles wasted. I want to use the right token, expensive tokens for the right models and let people do their job.
David Kanter
>> So I would say, one of the— when I look at what's changed, right, we've got a much broader audience and the rate of evolution is just incredible. And so I think for us that resolves down to we have to shift from the speed of hardware because the truth is MLPerf was started by hardware folks. To the speed of software, right? you're used to your apps on your phone getting updated every week or so, and it—
John Furrier
>> not when my battery's low though, right?
David Kanter
>> Of courseBut, my head of marketing drew out this great chart. And if you look at sort of leading edge frontier labs and capabilities, they're adding a new thing every 2 weeks. And so we have to shift to that sort of speed and get results and benchmarks that are going to stay current. We have to have benchmarks that are comprehensive, that really map out all the options a buyer is going to look at and compare them in sensible economic terms and then contextualize them. Right. it's because it's not just the infrastructure people anymore. It's everyone.
David Kanter
>> Yeah.
David Kanter
>> That contextualization is a big challenge. So we call that sort of the 4C challenge. We want things to be current.
David Kanter
>> Yeah.
David Kanter
>> Comparable, comprehensive, and then contextualized. And of course, you're part of that contextualization as well. You help illuminate the path for AI for many people.
John Furrier
>> Well, a lot of people want to know what's on the roadmap. They want to connect the dots. And again, back to the open piece, I think that is really the most important because you guys were grounded in the early days of AI, cloud native, Linux Foundation, CNCF, again, a unique approach, offered up some nice benefits with cloud native. So the open source equation is key to success. In a way, you're AI Commons now. ML's kind of too small in my mind because you're everything. You're helping everybody. You can call it Agent Commons, Physical AI Commons. basically it's—
John Furrier
>> You should be careful because we might bring you in for renaming ours.
John Furrier
>> No, no, but it's broad and you're doing great work. And also you're member funded. So explain that piece. This is not like you guys have a particular agenda. Talk about the scope and the mission, because I think this is also an important balancing piece.
David Kanter
>> Yeah, no, that's exactly right. So we are a member-driven organization. We have over 125 members on 6 out of 7 continents. We're still waiting for some folks in Antarctica to sign up, but one day, one day. And it's drawn from all pieces of the AI industry, and it's really focused on how can we make AI better for everyone through measurements.
David Kanter
>> Yeah.
David Kanter
>> And it's those members that fund us.
David Kanter
>> Right.
David Kanter
>> And so we're a nonprofit. There's a degree of transparency, of good governance that lets us do great work and that gives everyone the trust and the knowledge that we can pool our members' expertise to help shed light on what's going on and hopefully drive us down a path where AI is just going to do tremendous good. I think the applications in medicine or you've taken a Waymo in San Francisco. Yeah, it's almost magical, right?
John Furrier
>> Yeah, it's awesome. And if you look at the applications to the human society, again, back to the original mission of MLCommons and MLPerf was to make it better for people. That is what people want today in AI. Share what the activities are like for a member, what goes on behind the curtain. Every kind of project, open group has kind of different paths. What's it like? How do you guys engage? How do you get consensus? Take us through some of the sausage making and some of the behind the curtain.
David Kanter
>> Yeah. So the first thing is membership is open to everyone. We have individual members, we have academics. If you're interested, you can come get involved and help guide what we're building. And so for a lot of the members, they'll have representatives that show up and say, here is something that we think is important, we'd like to incorporate. Or yes, you've got a plan. The plan is 90% right. But if we tweak it just this way, it'll help us include our solutions in here. And so we can make the benchmarks broader.
John Furrier
>> And so you guys are more intentional in your focus than say, let a thousand flowers bloom, may the best projects win, which is kind of like a Linux Foundation. That's— yeah, that works for them. Yeah. But is that the same? You guys have a different approach?
David Kanter
>> I would say the Linux Foundation is in a lot of ways, sort of one of the original architects of this kind of an organization. And they're absolutely huge and have a huge number of projects. I think we're a lot more focused on We want to be doing things where we're uniquely suited. And so is it relevant to AI? Does measurement expertise play in? Does data play in? There's a set of things that we're really good at and align with what we do, and we want to focus on that. People ask me all the time, do you want to host an open source project? And I say there's other folks who are far more expert in doing that. The Linux Foundation has been doing it for 15 years.
John Furrier
>> You guys are really focused on the nitty-gritty, where you came from and where you are.
David Kanter
>> That's rightAnd where you're going.
John Furrier
>> All right, so for people who want to get involved, what's the process?
David Kanter
>> Come to our website. You can sign up for a membership. A lot of our groups are open to the public. You can just sort of sign up for those. But you can engage with us on Twitter or X, LinkedIn. I think we have a YouTube channel. And we'll often be at conferences presenting what we're doing. We're a really friendly bunch. So, if you have some great ideas, Sign up and say hi.
John Furrier
>> Well, we love what you guys have done in the past, set the table. That's the foundation. Now the world all wants in on what is going on. AI factories are super hot. It's the computing industry revolution again, but it's the same game, but it looks different, more dense. A lot of subsystems are involved. They're closer together, kind of like the old days, but bigger. Oh yeah, bigger and better rack scale systems. Edge is coming super fast. Latency, all these factors.
David Kanter
>> Yeah, it's a recasting of all of our compute infrastructure in a wildly different way. And it's both exciting and a little bit terrifying to be in the eye of the storm. Right. I tell my members what we do is through consensus. Right. And so there is a lot of negotiation over what we should be doing. And consensus is deliberate. It's slow. It builds trust. But at the same time, everything's moving at a million miles an hour. And so, it's— but it is an absolute pleasure and an honor to get to have this role. And it's exciting to see what's going to be coming out in the next year.
John Furrier
>> Well, thanks for coming on, sharing the mission totally behind it. Open always wins. I've been saying that from day one. I've said on theCUBE probably the most of anything. Governance gets a lot of buzzwords these days in AI, but, open wins and that's where innovation lives.
John Furrier
>> Yeah
John Furrier
>> . Thanks for coming on. Appreciate it, David.
David Kanter
>> Thank you so much for your time.
John Furrier
>> I'm John Furrier. This is a Mixture of Experts here where people share their thoughts on the key issues facing what's being built on the innovation, obviously AI infrastructure, the hottest area, agents, physical AI and robotics, defense tech, all booming as part of this new AI revolution. And again, people want to know where things fit, where to buy it. Doing our part here. Thanks for watching.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, co-host of theCUBE, here with Dave Vellante, my co-host. Hello, I'm John Furrier, your host of theCUBE, here at theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street, part of our NYSE Wired program and community. This is our mixture of experts series where we bring in people who are experts in their field, doing great work and innovating. David Kanter's here. He's the co-founder of MLCommons and head of MLPerf. If you know all about machine learning, you know what that organization has done. Nonprofit doing really amazing work helping people figure out what's safe, what's real, what's not. David, great to see you. Thanks for coming on theCUBE. Saw you at AMD's event in San Francisco.
David Kanter
>> Absolutely a pleasure. It's great to be
John Furrier
>> here. You get to be mixed up with all the other expertsYeah, yeah. This is kind of a— originally was kind of a goof on AI when we did the series, but it's actually a great way to bring in our community and kind of make sure experts kind of share in the data. And one of the things that everyone loves about AI is that it's got a great utility. But pre the transformer technology, machine learning has been around for a long, long time. You know, fraud detection, every bank has it, supervised, unsupervised machine learning. It really was the genesis of what got all the deep tech nerds in the labs, looking at what's coming out. And then the field just went supernova from there. So it's really a valuable organization and lesson and also a template for the future. Explain what you do at MLCommons and MLPerf, how it all came together. What is it for people who don't know what it is and what it does and where it is?
David Kanter
>> Absolutely. So we got started in 2018, back pre-ChatGPT, pre-generative AI, and everyone was looking at we knew these AI models could do incredible things. We wanted to improve performance, to improve capabilities, but there was no standard way of measuring things. And so a group of us all came together from industry and academia to build that standard set of benchmarks to measure speed and energy efficiency, and that became MLPerf. And before that—
John Furrier
>> and by the way, that then became the cited benchmark stat in every presentation at that time.
David Kanter
>> Exactly right. And part of the thing that's really wonderful about getting to be involved in this group is we bring together everyone from all across the industry and through consensus we build these trusted standards like MLPerf. And, you know, at the time it would almost be as if you were buying a car and one guy says, hey, my car can do 0 to 60 in a second. The next guy says, my car has a turn signal. And the third guy says, I've got airbags. Which one do you want to buy? You don't really know. And so you need some way to compare them all. I live in San Francisco, so I might go for the airbags. But that was the genesis of MLPerf. And then we realized sort of the impact you can have to help drive the whole industry. And we said we should put this into a nonprofit. And then look for other ways that we can deploy our expertise in measurement and data to help make AI better, right? And so after we first did performance, we then zeroed in on measuring power efficiency, building large open datasets, and then over time we've started looking at benchmarks in risk and reliability of helping to make sure that the outputs of generative models are kind of in line with what we want.
John Furrier
>> The evolution of AI now is the number one conversation is safety, right? And then you see the Anthropic versus, say, OpenAI approach, fast and loose, more conservative. There's a general consensus, a lot of consensus around no one really knows what the hell that means. So take us through kind of what you guys are focused on now because you guys have the playbook on open AI. We see the success of open source. damn, it's the most successful trend ever in the computer industry. Look at what it's done. Now you've got OpenWeights. So you got a lot of open things happening. What are you guys focused on now? How is this translating into some of the conversations today?
David Kanter
>> Yeah, so I'd say one of the most critical things is when you're looking at anything, whether it's performance or risk or responsibility, it's about measuring it in the right way. Having written down what you're doing, what you're trying to accomplish, how much precision you have. And so for us, one of the things that we released last week was MLPerf Endpoints v0.7, which is a rethinking of our inference benchmarks for the modern era.
David Kanter
>> Right.
David Kanter
>> And we see this race to deploy as everyone's discovered that there's so many valuable things you can do with AI. How do we deploy it across the enterprise for consumers? We see building more data centers, needing more power and more performant systems. So we had to evolve our benchmarks to match that pace of innovation and to help customers really make the decisions that they need. You look at a Fortune 500 company.
David Kanter
>> Yeah.
David Kanter
>> They're not just saying, hey, I want AI for the C-suite. They're saying I have dozens or hundreds of applications that I'm going to deploy. Each one's different. How do I find the infrastructure that's going to pair up in the right way? And so I look at that as being our job is how do we build the tools that help empower those folks to make the decisions to help deploy AI?
John Furrier
>> So you're really kind of taking the DNA of MLPerf, MLCommons, machine learning.
John Furrier
>> Yeah
John Furrier
>> . Applying that to the AI growth wave, which is infrastructure and trying to help people navigate that. So I love that. The question that we're seeing now is I just had— I just wrote a post, went to lunch, actually, I wrote a post, but on agents and a lot of the fear with agents is you have more nondeterministic workloads. Again, that's cool, but you have workloads. And they're different. So now you have different conditions. What's the scope of some of the things you guys are getting your arms around in the open? Because a decision for Company A will be different than Company B because I might want to have more compute, less GPU, or pre-fill decode. all these things are kind of now coming into the systems. It's not a clear general-purpose benchmarking market. And so how do you guys think about that? What's the community doing on— is there any Data you can share, thoughts, personal thoughts?
David Kanter
>> Thoughts for sure. And stay tuned. We'll have data later this year for sure. But I think one of the things you really touched on is when we see blending AI inference with standard computing workloads through agentic flows, right? The world is your oyster, right? Before it was like, oh, maybe you're doing recommendation or translation. Well, Now you might be pairing that with hey, this is the conventional workflow that I have in my bank, but now I'm going to stick AI in here to accomplish my goal. Or, of course, the thing that we've seen the most demand for is coding.
John Furrier
>> Yeah, right
David Kanter
>> . And it sort of makes sense. The folks who are developing the tools are like, hey, wait, I can do what with this? Like, let's get some acceleration.
John Furrier
>> Hugging Face. Hello. Testbed went off the rails again. There are so many use cases where it could go off. Yeah, may or may not be related to anything other than the environment.
David Kanter
>> That's right. And one of the core insights in MLPerf actually was it's not just about speed, but it's about how fast you get to the right answer. And to the point you made earlier, you could achieve the same task with a lot of accelerated inference compute and maybe less conventional compute. Or maybe a different balance. And it depends on what you have. Imagine you want to pick what's the best South Indian restaurant in Manhattan? You could look at the top 10 if you're really confident of that top 10, or maybe you've got a system that's even more accurate and you just say, I really need the top 3 from this system, so I don't need to read all those reviews. Both of those will hopefully get me to a delicious meal. But what's the right path? That's really tricky. And so we're just in the starting stages.
John Furrier
>> That's a really good point. I think what you just said was compelling because most people look at the user experience of, say, ChatGPT, which most consumers experience as getting a good answer fast. The first answer, kind of search results. Yeah. Hey, that's great. Where do I find food? Boom. Answers. Reasoning is a different— it's not a search paradigm. Are you doing discovery?
John Furrier
>> That's right.
John Furrier
>> With multi-step reasoning, which is completely different. So it's a very nuanced point, but it changes the configuration of the data, what systems I might want to use. Maybe it's a complex answer. Maybe you have certain info needs that might require a Pareto curve of the variable routing, right? Who knows? Cha-ching on the tokens. So this is a cost trade-off. This is a math equation.
David Kanter
>> That's exactly right. And so part of our goal with MLPerf Endpoints is how do we present that trade-off so that if you're a CIO or a CISO, you can say, all right, I've got all these applications, these need to be really fast. These are for frontline. I need an answer for the customer quick. These other things may be back office are going to look different. And so how can you get the right infrastructure for everything? But I think to your point, the world of agentic, before you had AI walled off and it was its own thing, you had separate AI infrastructure people. But now, one of the things I used to say is to me, inference is a lot like salt in cooking. You don't eat pure salt usually.
John Furrier
>> Yeah
David Kanter
>> , but if you know from any recipe book, you add a little bit of salt and it makes almost everything taste better. And that's what we're seeing. So now it's not AI here and regular computing there. It's all braided together and all throughout. So almost everything that we have today in the enterprise is going to be recast with an agentic side. And so we're still in the early stages. And when I think about what we want to do, we have an agentic benchmark coming out later this year focusing on some of the things that we think are most popular, code development and software engineering, as well as Q&A and customer support. But I wouldn't be surprised if we have dozens of use cases in the future as we're discovering them live.
John Furrier
>> You know, you and I, when we were chatting at the AMD event in San Francisco, we were talking about some of the historical views. We've lived through many cycles of innovation. This one's obviously the most kick-ass ever because it's got everything popping. You got infrastructure up and down the stack. But in the old days when I was breaking into the business at Hewlett-Packard and before that IBM, PCs and servers, they all had the benchmarks. And that was really twofold. One, to do an industry service to kind of level the playing field on horses on the track, apples to apples, making sure everyone knows what's what. But it also helped customers scope what they wanted to buy. Yeah, it was really an economic beacon too for the, okay, I need a mid-range system for these desktops or whatever. When you get to AI, a lot of that's kind of going on, but it's not as simple. What's the biggest change in your mind today trying to rally the industry around MLCommons while looking at the aperture of use cases. Do you guys look at that as an opportunity? Do you— are there some first principles and then playbook tactics you guys are using? Because everyone wants the same thing. What do I buy it for? What and when? I don't want to waste any money. I don't want GPU cycles wasted. I want to use the right token, expensive tokens for the right models and let people do their job.
David Kanter
>> So I would say, one of the— when I look at what's changed, right, we've got a much broader audience and the rate of evolution is just incredible. And so I think for us that resolves down to we have to shift from the speed of hardware because the truth is MLPerf was started by hardware folks. To the speed of software, right? you're used to your apps on your phone getting updated every week or so, and it—
John Furrier
>> not when my battery's low though, right?
David Kanter
>> Of courseBut, my head of marketing drew out this great chart. And if you look at sort of leading edge frontier labs and capabilities, they're adding a new thing every 2 weeks. And so we have to shift to that sort of speed and get results and benchmarks that are going to stay current. We have to have benchmarks that are comprehensive, that really map out all the options a buyer is going to look at and compare them in sensible economic terms and then contextualize them. Right. it's because it's not just the infrastructure people anymore. It's everyone.
David Kanter
>> Yeah.
David Kanter
>> That contextualization is a big challenge. So we call that sort of the 4C challenge. We want things to be current.
David Kanter
>> Yeah.
David Kanter
>> Comparable, comprehensive, and then contextualized. And of course, you're part of that contextualization as well. You help illuminate the path for AI for many people.
John Furrier
>> Well, a lot of people want to know what's on the roadmap. They want to connect the dots. And again, back to the open piece, I think that is really the most important because you guys were grounded in the early days of AI, cloud native, Linux Foundation, CNCF, again, a unique approach, offered up some nice benefits with cloud native. So the open source equation is key to success. In a way, you're AI Commons now. ML's kind of too small in my mind because you're everything. You're helping everybody. You can call it Agent Commons, Physical AI Commons. basically it's—
John Furrier
>> You should be careful because we might bring you in for renaming ours.
John Furrier
>> No, no, but it's broad and you're doing great work. And also you're member funded. So explain that piece. This is not like you guys have a particular agenda. Talk about the scope and the mission, because I think this is also an important balancing piece.
David Kanter
>> Yeah, no, that's exactly right. So we are a member-driven organization. We have over 125 members on 6 out of 7 continents. We're still waiting for some folks in Antarctica to sign up, but one day, one day. And it's drawn from all pieces of the AI industry, and it's really focused on how can we make AI better for everyone through measurements.
David Kanter
>> Yeah.
David Kanter
>> And it's those members that fund us.
David Kanter
>> Right.
David Kanter
>> And so we're a nonprofit. There's a degree of transparency, of good governance that lets us do great work and that gives everyone the trust and the knowledge that we can pool our members' expertise to help shed light on what's going on and hopefully drive us down a path where AI is just going to do tremendous good. I think the applications in medicine or you've taken a Waymo in San Francisco. Yeah, it's almost magical, right?
John Furrier
>> Yeah, it's awesome. And if you look at the applications to the human society, again, back to the original mission of MLCommons and MLPerf was to make it better for people. That is what people want today in AI. Share what the activities are like for a member, what goes on behind the curtain. Every kind of project, open group has kind of different paths. What's it like? How do you guys engage? How do you get consensus? Take us through some of the sausage making and some of the behind the curtain.
David Kanter
>> Yeah. So the first thing is membership is open to everyone. We have individual members, we have academics. If you're interested, you can come get involved and help guide what we're building. And so for a lot of the members, they'll have representatives that show up and say, here is something that we think is important, we'd like to incorporate. Or yes, you've got a plan. The plan is 90% right. But if we tweak it just this way, it'll help us include our solutions in here. And so we can make the benchmarks broader.
John Furrier
>> And so you guys are more intentional in your focus than say, let a thousand flowers bloom, may the best projects win, which is kind of like a Linux Foundation. That's— yeah, that works for them. Yeah. But is that the same? You guys have a different approach?
David Kanter
>> I would say the Linux Foundation is in a lot of ways, sort of one of the original architects of this kind of an organization. And they're absolutely huge and have a huge number of projects. I think we're a lot more focused on We want to be doing things where we're uniquely suited. And so is it relevant to AI? Does measurement expertise play in? Does data play in? There's a set of things that we're really good at and align with what we do, and we want to focus on that. People ask me all the time, do you want to host an open source project? And I say there's other folks who are far more expert in doing that. The Linux Foundation has been doing it for 15 years.
John Furrier
>> You guys are really focused on the nitty-gritty, where you came from and where you are.
David Kanter
>> That's rightAnd where you're going.
John Furrier
>> All right, so for people who want to get involved, what's the process?
David Kanter
>> Come to our website. You can sign up for a membership. A lot of our groups are open to the public. You can just sort of sign up for those. But you can engage with us on Twitter or X, LinkedIn. I think we have a YouTube channel. And we'll often be at conferences presenting what we're doing. We're a really friendly bunch. So, if you have some great ideas, Sign up and say hi.
John Furrier
>> Well, we love what you guys have done in the past, set the table. That's the foundation. Now the world all wants in on what is going on. AI factories are super hot. It's the computing industry revolution again, but it's the same game, but it looks different, more dense. A lot of subsystems are involved. They're closer together, kind of like the old days, but bigger. Oh yeah, bigger and better rack scale systems. Edge is coming super fast. Latency, all these factors.
David Kanter
>> Yeah, it's a recasting of all of our compute infrastructure in a wildly different way. And it's both exciting and a little bit terrifying to be in the eye of the storm. Right. I tell my members what we do is through consensus. Right. And so there is a lot of negotiation over what we should be doing. And consensus is deliberate. It's slow. It builds trust. But at the same time, everything's moving at a million miles an hour. And so, it's— but it is an absolute pleasure and an honor to get to have this role. And it's exciting to see what's going to be coming out in the next year.
John Furrier
>> Well, thanks for coming on, sharing the mission totally behind it. Open always wins. I've been saying that from day one. I've said on theCUBE probably the most of anything. Governance gets a lot of buzzwords these days in AI, but, open wins and that's where innovation lives.
John Furrier
>> Yeah
John Furrier
>> . Thanks for coming on. Appreciate it, David.
David Kanter
>> Thank you so much for your time.
John Furrier
>> I'm John Furrier. This is a Mixture of Experts here where people share their thoughts on the key issues facing what's being built on the innovation, obviously AI infrastructure, the hottest area, agents, physical AI and robotics, defense tech, all booming as part of this new AI revolution. And again, people want to know where things fit, where to buy it. Doing our part here. Thanks for watching.