David Kanter, ML Commons
In this interview from theCUBE + NYSE Wired's Mixture of Experts series, David Kanter, co-founder of MLCommons and head of MLPerf at MLCommons, joins theCUBE's John Furrier to discuss how AI benchmarking is evolving to keep pace with agentic workloads and enterprise-scale deployment. Kanter traces MLCommons' origins back to 2018, when a coalition of industry and academic players set out to standardize how AI performance and efficiency are measured, creating MLPerf. He explains how that benchmark evolved from measuring raw speed and power efficiency into a broader focus on risk, reliability and safety as generative AI matured. Kanter also introduces MLPerf Endpoints v0.7, a newly released benchmark rethinking inference measurement for the agentic era, where AI is increasingly blended with conventional enterprise workflows rather than treated as a separate system. The conversation also explores MLCommons' shift from benchmarking hardware to keeping pace with software that evolves every few weeks, a challenge Kanter frames around what he calls the "4C" standard: benchmarks that are current, comparable, comprehensive and contextualized for buyers ranging from CIOs to individual developers. He details how MLCommons' member-driven, nonprofit structure — spanning more than 125 organizations across six continents — enables open, consensus-based governance distinct from broader open-source foundations like the Linux Foundation. Kanter also previews an upcoming agentic benchmark targeting code development and customer support, two of the fastest-growing AI use cases. From the early days of MLPerf's hardware benchmarks to today's push to measure trust and reasoning in agentic systems, Kanter makes the case that open, community-driven standards remain essential to guiding AI's next phase of growth.