This session from theCUBE and NYSE Wired Robotics and artificial intelligence
Infrastructure Leaders Series, the third annual infrastructure leaders event,
features Max Kan of SemiAnalysis. The discussion examines infrastructure topics
such as tokenomics, inference hardware, model distillation, agentic workloads
and the evolving dynamics between frontier and open source model development.
Kan examines token economics and AI infrastructure dynamics, focusing on agentic
workloads, token volume growth, model distillation, inference hardware
trade-offs and accelerator architecture segmentation. They emphasize that
accelerating global token production driven by multi-turn agentic workloads
requires monitoring frontier-model token share to forecast market shifts. They
note that low switching costs make model quality the principal competitive
factor and that hardware segmentation channels most volume to general-purpose
chips while premium accelerators such as Groq and Cerebras serve
latency-sensitive high-price niches. theCUBE Research frames the conversation
with hosts John Furrier and Gabe Valladares and explores implications for
developers, cloud providers and AI labs during the ongoing infrastructure
buildout. Key insights include: - Global token production accelerates as
multi-turn agentic workloads expand, making frontier-model token share a
critical forecasting metric. - Low switching costs elevate model quality as the
primary determinant of adoption. - Hardware segmentation directs volume to
general-purpose chips while premium accelerators such as Groq and Cerebras
address latency-sensitive high-price niches.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Max Kan, SemiAnalysis
This session from theCUBE and NYSE Wired Robotics and artificial intelligence
Infrastructure Leaders Series, the third annual infrastructure leaders event,
features Max Kan of SemiAnalysis. The discussion examines infrastructure topics
such as tokenomics, inference hardware, model distillation, agentic workloads
and the evolving dynamics between frontier and open source model development.
Kan examines token economics and AI infrastructure dynamics, focusing on agentic
workloads, token volume growth, model distillation, inference hardware
trade-offs and accelerator architecture segmentation. They emphasize that
accelerating global token production driven by multi-turn agentic workloads
requires monitoring frontier-model token share to forecast market shifts. They
note that low switching costs make model quality the principal competitive
factor and that hardware segmentation channels most volume to general-purpose
chips while premium accelerators such as Groq and Cerebras serve
latency-sensitive high-price niches. theCUBE Research frames the conversation
with hosts John Furrier and Gabe Valladares and explores implications for
developers, cloud providers and AI labs during the ongoing infrastructure
buildout. Key insights include: - Global token production accelerates as
multi-turn agentic workloads expand, making frontier-model token share a
critical forecasting metric. - Low switching costs elevate model quality as the
primary determinant of adoption. - Hardware segmentation directs volume to
general-purpose chips while premium accelerators such as Groq and Cerebras
address latency-sensitive high-price niches.
>> Hello, I'm John Furrier with theCUBE here at our Palo Alto studios for theCUBE and NYSE Wired, it's our third annual AI Infrastructure Leaders event covering the AI infrastructure build out and boom of course, physical AI right around the corner. I want to thank ScaleFlux for supporting us as well as all of our industry sponsors over the year. Really appreciate it. It's been a really growing community. Our first guest to kick it off is Max Kan, tokenomics technical lead at SemiAnalysis. Max kicking off the program, we got our big 180 folks coming to the event tonight. Little networking nerd fest here in Silicon Valley. Thanks for coming on theCUBE and NYSE Wired, third annual event.
Max Kan
>> Yeah, thank you for having me, John. Pleasure to be here.
John Furrier
>> So you guys just do great work, SemiAnalysis. In the industry, everyone kind of knows what you guys do, but you guys do deep research on all things AI, obviously from the infrastructure, what's going on at the Pareto curves, token economics, all the nuances around what's powering the AI infrastructure. Everybody wants to know about token economics. We saw the big paper that went out this week. Jensen put it out there, like 25 million views on X. Everyone's citing, everyone's jumping on the bandwagon. Open weights, you're seeing open source surge, you're seeing Anthropic OpenAI trying to go public. Robotics is booming, we think it's going to be a big year next year, so tokens are a huge part of the AI piece. What are you seeing in the market? What's your focus?
Max Kan
>> Yeah, I'd say here on the Tokenomics team, we're sort of at the top of SemiAnalysis full stack coverage. So at the very bottom, we have things like, Accelerator team that tracks where the accelerators are, Data Center team will track all the data centers. Tokenomics, you can maybe say the stated goal is to track all the tokens, but our kind of primary coverage area today is the hyperscalers and the AI labs. Obviously, open source has been a big deal for the past week or two, ever since the Kimi K3 blog post with all the crazy benchmark scores came out. And then we've seen this big fight on Twitter and elsewhere with Jensen and it kind of seems like the rest of the AI community really pushing for open source, even OpenAI signing on and kind of Anthropic on the other side. The way that I view that letter personally is I think it's mostly Jensen talking his book. He knows that he can't be sort of fully beholden to only two customers, Anthropic and OpenAI to buy all this compute. And so he very desperately and actively wants to nurture kind of an open source ecosystem. Whether or not open source AI is good or bad, I think probably not good overall. But I think a lot of the sort of discourse going on is mostly people talking their book and less like maybe moral principle discussions.
John Furrier
>> Frame the discussion for the folks that aren't inside the ropes because you have, kind of talking their book, their messaging, they have their strategies, AMD clearly playing the open card, NVIDIA playing open as well. They say a lot, hey we're open, and they're supporting open source. But they're kind of two different approaches. you look at both of them, so at the end of the day, do I really care if I got the best tokens per watt per dollar? I just want to get more tokens, and I don't want to pay through the nose for tokens if I don't need to, take us through this, frame up the landscape, the playing field, the arena. What's going on with token economics? Who's where, what's what, and who's leaning in, and who's moving the needle in terms of pushing the innovation for developers and companies?
Max Kan
>> Sure, yeah, so I would say at one end of the spectrum, you definitely have Anthropic, which has historically been sort of the most AI safety conscious AI lab. They've sort of, even back when I think GPT-2 first came out, when Dario was still working at OpenAI, he was saying it's potentially too dangerous to release to the public. We've got to make sure we do some rigorous safety testing here. Obviously that was maybe a bit hyperbolic, but I think a lot of his sort of original claims and fears have started playing out to a degree. At the same time, they're obviously a closed source lab that has incredibly high API margins today. I think if you're paying API price for Opus or Sonnet, there's a very good chance that's like 85 % plus gross margins for Anthropic. Maybe even better than traditional SaaS, which is very much the opposite of what people were sort of characterizing the AI build -out as just maybe six months ago. But anyway, so they're on one end. I'd say on the other end, you have the open source guys. this is primarily Chinese labs today. But obviously, there's a kind of moral and philosophical movement with open source where people say, you just want technological development to be free, you don't want the government or any one lab regulating everything or especially controlling everything. From the economic side, obviously the open source guys have worse margins. So I think if you want to take maybe the cynical business point of view, it's like you have Anthropic who wants to defend their lead in AI right now, they want to defend these 85 % plus gross margins, you have the open source guys who have less margins, they're trying to eat into that. But then obviously there's the more moral angle as well, which I think Anthropic does truly believe to a degree, I don't want to, I don't think they're purely being cynical here, but that's kind of the landscape.
John Furrier
>> Yeah, and you see a lot of people, if you're building something, AI native companies tend to do well, we saw Fireworks AI is crushing it. They're doing well, so there's a lot of demand for tokens. There's a lot of development going on. Now you got the suppliers and it's not obvious. We're supposed to get lower gross margins. We thought Anthropic would be going down, not up. How do the economics shift? Because when you start seeing that balance between general intelligence and specialized intelligence, you're going to have this kind of, not mutually exclusive environment, but more integrated environment. A lot of people are talking about small language models. Are people behaving differently with how they think about tokens? Given that the compute's now back on the table for inference, the role of the GPU. in your guys' analysis, how does the token piece fit into some of those economic discussions?
Max Kan
>> Sure, yeah. So I would say historically, basically all the revenue has been concentrated just on the smartest frontier models, which is why you see Anthropic and OpenAI have such crazy ARR numbers today, growing faster than, no business has ever grown this fast at this scale in the history of capitalism. It's honestly quite impressive. Right, so obviously though the open source models have gotten kind of remarkably good over the past maybe call it the past month and so the big question on everyone's mind is like is this finally going to change? Maybe sort of a precise kind of way toward it is sort of one metric I think people should start looking at is if you were to plot global token volumes over time is the percentage going to frontier models decreasing or increasing? Obviously if that number starts decreasing dramatically, it means that open source models are becoming increasingly popular. If it starts sort of, if it continues increasing, then it's kind of just more of the same old, same old. Admittedly it's sort of very hard to collect data to actually answer this question. We are trying very hard at SemiAnalysis. Hopefully we should have something out soon, maybe in the next couple months or so. Big reason why this is difficult is that Anthropic has never published a token production number so you kind of just back it out from their revenue and other disclosures and stuff.
John Furrier
>> So you guys are sifting through all that to try to piece it together.
Max Kan
>> Yeah, yeah, that's our job. Sort of sift through all the minutia and kind of distill it down into takeaways people can actually understand.
John Furrier
>> Max, I'm old enough to remember when the internet came, everyone's like, oh, look, the internet's going to be great. And that's when we saw that bubble firsthand. And one thing that was interesting, Mary Meeker was an analyst at Morgan Stanley at that time, and she had this online population metric, the total online population, because people were like, oh the internet, first the geeks get on there, the people who know the internet, and then the general, everyone else used it. So that was a key metric. So the question is, in tokens, is there a view on total token usage? Because as more people are using AI, that's kind of a similar online population of users. What is your view on the general purpose? Because everyone's on the consumer side, I see people in New York City using the chat all the time on OpenAI, but business people tend to go towards Anthropic. Gemini's got the embedded Google advantage of being embedded in the search. So you have all these different distribution pieces, but is the overall pie getting bigger?
Max Kan
>> Yeah.
John Furrier
>> And how does that factor into how does one squint through who gets what tokens? Is there a power law? these are kind of like big picture questions. What are your thoughts and reaction to that?
Max Kan
>> So I think token production is just very clearly going exponential. So the total pie. I would guess that total global ex-China tokens a day is maybe a little less than a quadrillion tokens a day. So this is obviously a monster number. Big reason or a big part of what's driving the exponential growth is the shift to agentic workloads. I think people, for those who don't understand how inference actually works, the basics. In general, agentic workloads consume exponentially more tokens than the chat workloads everyone was using a year ago because it's all inherently multi -turn and every time you ask your agent a follow -up or it uses some tool like web search or maybe it searches your code base, it has to reprocess all previous tokens in your conversation, which really causes total tokens produced slash processed to go up exponentially.
John Furrier
>> Again, going back to what we were talking about You mean like looping, they have to do multi -step reasoning, they have to do a lot more than just getting an answer.
Max Kan
>> Yeah, it's not quite the same as looping. I said when people say looping, they generally refer to like, I just put my agent in a loop and then tell it to keep doing the same thing over and over again.
John Furrier
>> Okay, so that's specific, okay.
Max Kan
>> Yeah, this is just more inherent to how technology actually works. It's like every single time you ask the agent a follow -up, or, and this is actually where the KV cache comes in too, like every single time you ask the agent a follow -up or the agent uses a tool, it sort of has to reload the key and value vectors for all previous tokens into its context in order to generate the next token. And this just causes exponential growth of token usage. It's also why people, a lot of people like to look at the blended price per million tokens over time and try to use that as some gauge of like, are open-source models becoming more popular or not? The one caveat I think anyone doing this analysis needs to consider is that workload shape absolutely matters. As people know, there are three main types of tokens. There are the cached input tokens, the base input tokens, and the output tokens. Those are priced very differently. And in general, agentic workloads, cached input tokens go way up, and also regular input tokens go way up. Those are much cheaper than the output tokens, and so you could see your blended price per million tokens go down, even though it's just a function of workload shape changing and not actually model usage changing. And so that's one thing people should look out for if they're doing this now.
John Furrier
>> So what you're getting at is that it's complicated in the sense of the token relationships to the workload task could be very complex or elementary and those token decisions make a difference. Sounds like there's a lot going on under the covers.
Max Kan
>> Yeah, yeah, yeah.
John Furrier
>> And one of the things at GTC this year, when Jensen put up the Pareto curves, saying to Dave Vellante and folks like, okay, now we're starting to see premium tokens because if you're on a Vera Rubin, you're going to be paying out the nose for those tokens. But I don't want to use those tokens to do calculations that I can get for maybe a better price. So context, not content, token routing is starting to come into it. You're starting to see a lot more token stuff under the covers. You mentioned agents. There's so much going on. How does that factor into the analysis? Because it's blended. Is there an intelligent algorithm in there? How do you see this evolving? because people are starting to talk about, okay, if I'm going to have this design system, I will use maximum premium tokens for this, but I want to make sure I'm not going to suck up the GPU to do basic stuff. I won't say low-end stuff, but basically it's lower-end calculations.What's your view on that piece?
Max Kan
>> So I would say actually on the Pareto curves, I think you definitely will be using Vera to generate cheap tokens, actually. I think Vera will give you sort of the cheapest tokens, whether it's per dollar or per watt, that you could possibly ask for. This sort of really premium token that Jensen was talking about at GTC, that's what Groq gives you. And this kind of goes down into the actual architecture of what does a Groq SRAM chip versus an NVIDIA Vera CPU look like. I'm not going into the details here, but TLDR is that Groq is super expensive, but it can generate tokens really fast. People have historically shown they're willing to pay a premium for fast tokens. We saw Anthropic at one point was charging 6x the price for just 2.5x the speed. and SemiAnalysis was happily paying that. Fortunately, now it's only actually 2x the price for still 2.5x the speed, so that was pretty good for our bottom line. But that's sort of like the Pareto curve Jensen was talking about.
John Furrier
>> In terms of performance, I know you guys spend a lot on tokens in your research.
Max Kan
>> Yeah.
John Furrier
>> Cerebras was pumping up their thing, they just did a deal with AMD, they claim to be the fastest. They have a big chip, a lot of SRAM involved. And then the diversity piece of the vendors, because some are saying Groq's okay but expensive, there might be other alternatives for inference, but that's working, but then it's like, okay, they have the IP, but there might be other solutions. How do you think about the alternatives, like a Cerebras, like an AMD, like an Nvidia, as people start to figure out where the costs are?
Max Kan
>> Yeah, I think if you believe like I do, that inference is going to become sort of the largest market ever, there will be a place for sort of all of these niche accelerators. The question is just how big. Personally, I think most of the volume is actually still going to just go to, at least it's not going to go to these SRAM chips, Groq and Cerebras, because they're more niche, they're much more expensive, really the only people, it's like people who are willing to pay 10x the price of current API prices will be willing to pay, like, OpenAI announced they're going to have 750 tokens per second on Cerebras in July. They're running out of time here. We'll see if they actually still hit that deadline. But that's probably going to cost like 10 times more than their regular API price. I think it's a very niche segment of the population is going to be using that. I don't even know if it's in budget for SemiAnalysis. I'd have to ask Dylan. So I think most of the volume will be going to the general purpose chips. I think it is an open question whether maybe a Positron or an Etched or one of these guys will be able to create potentially even a better general purpose inference chip than NVIDIA. I don't have super strong opinions there. I haven't spent too much time looking at them.
John Furrier
>> But token maxing, we saw that wave.It was like, hey, look at all the tokens I'm using. Now the conversation on X is value maxing. So there's really a movement towards the old school coding days. Less code is better. Is there a new metric on token to outcome? Are people starting to look at the revenue? Is there a new metric behind just the mechanism on token to value. Are you seeing, tracking anything there?
Max Kan
>> I don't think there's any one golden bullet metric you can use for every enterprise. I think it really depends. Every single company sort of has to look at their token spend honestly today and try to calculate the ROI on that for their own business the same way they would sort of see whether or not this extra employee headcount is worth it for their business. I don't think we'll ever have any one golden metric.
John Furrier
>> They're all different too. They have different business models, different objectives, different workloads, different agents.
Max Kan
>> Yeah, I think some people have floated the idea of maybe charging sort of per task or based on outcome instead of per token. I think that maybe makes sense in some use cases, but at the same time, I think really one of the benefits of having this general purpose technology is that OpenAI or Anthropic have no idea what sort of like all the great, crazy, useful, high ROI things people are going to use it for. And so I think token -based pricing will be here to stay for at least the next year or two.
John Furrier
>> Give your bull and bear use case view on OpenAI and Anthropic. What's the bull bear analysis? How do you look at that? What do the scenarios look like? How could the ball drop on either side?
Max Kan
>> Yeah, so I would say the bull case for these two companies, and actually, are you asking about sort of like the two companies as one unit?
John Furrier
>> No, no, no, two separate companies as individuals going public.
Max Kan
>> Oh, I see.
John Furrier
>> You said that they're the fastest growing companies in the history of capitalism, right? So, to me, I think there's a power law development, I think you're going to have the big guys doing general and then I think it's going to be specialism. But there's a lot of debate around, are these going to be viable companies? And the neoclouds are going through the same thing, where it's like, okay, they're getting some beachhead, they're backstopped by the big guys, you're seeing that kind of funding. So there's definitely, and his goal of getting alternatives out there is working, but will they be viable? That's the question. Some are leaning toward the bull case, like this is a game changer, We don't know yet what the economics look like, hasn't been modeled yet. The bear case is like, well, this is a bubble, it's going to pop. How can they sustain that cost on the CapEx? So you have kind of competing views and sentiment around these companies. Now, my opinion, I won't share, but you kind of get where I'm going. I'm kind of on the bull side. I think, scale matters and billions of users is a good thing.
Max Kan
>> Yeah, maybe I'll start with the bull case or sort of frontier closed source labs in general. and then the bear case for them, and then I'm going to talk about OpenAI and Anthropic separately. So I think in terms of the bull case for why at least one of OpenAI or Anthropic could become maybe the first $10 trillion company within the next couple of years, is if the percentage of global token volumes going to frontier models keeps increasing. And I think if that happens, what it means is that sort of the new economically viable tasks that are unlocked by increasingly smart frontier intelligence, sort of the TAM of that is growing faster than the easier tasks that will be switched over to these cheaper open source models. I think these are sort of the two key lines that everyone needs to watch. Historically, obviously, the answer is just that that top line has grown way faster than the bottom line. That's why Anthropic in particular, the ARR growth has significantly outpaced the rest of the industry, which has also been growing very quickly since the start of the year. I think really it depends on like are they going to be curing cancer and doing all these other crazy technological things because obviously regular software engineering is not going to need sort of the smartest model probably a year from now so what new use cases will be unlocked I think that's sort of the key question for the AI labs the frontier closed source AI labs as a business model obviously the bear case then is that people don't actually find these new use cases fast enough. And six months from now, the best Chinese open source model is sort of just sufficient for the majority of white collar work today. Then I think bubbles popped, it's all over.
John Furrier
>> So basically the adoption has to get there.People have to taste it, get addicted to it, understand it, use it, get value out of it quickly. And if that adoption value piece doesn't kick in.
Max Kan
>> Exactly. It has to be fast. And they have to find things that aren't just like summarizing emails or doing basic software engineering. It sort of has to be like fundamentally new tasks that are just not possible without the increase in the smart frontier intelligence. And then to this point, I think if you want to differentiate between OpenAI and Anthropic, really what matters is just who has the better model. I think we've seen that switching costs are incredibly low in this industry. You can maybe make some arguments like, oh, with the advent of Claude Code and Codex and sort of this thesis, the agent is the product, it's not just the model, you have to have this good harness too, and the switching costs are slightly higher. I don't really buy this. I would say personally, I've switched between kind of Claude Code and Codex two times in the past two weeks, just because it's not that hard to change your harness. And I think the reason why Anthropic really overtook OpenAI over these past seven months is that last November, when they came out with Opus 4.5, they just had by far the best agentic model in the world. OpenAI was looking extremely dire. Their models were basically complete trash until this April or May when they came out with GPT-5.5. And then that's also the point when we saw their revenues start accelerating again. So I think really for the AI lab.
John Furrier
>> It's like an F1 race. You don't know, you're in the lead, you drop back, but you're getting to a good point around the model. So a better model wins. So you think switching costs are low or high?
Max Kan
>> Super low. Super low.
John Furrier
>> You don't think that recompiling tokens and managing, not recompiling, but recasting the agentic workflows matter or is this irrelevant?
Max Kan
>> Are they already decoupled? So I would say, this is maybe a slightly more technical point, but there is an argument to be made that if you really invest your effort into developing a set of evals, so you can rigorously test model performance. and then you also customize your entire harness and all your context management system for a particular model, then the switching costs will be pretty high if you were to rip it out. But I think people aren't actually doing that today and I think investing the time to do that is honestly a mistake.
John Furrier
>> So the harness is a key.The key is the harness and making sure that you kind of decouple core operations from the token models.
Max Kan
>> Yeah, but I'd say harness engineering is more, that's sort of like OpenAI and Anthropic's job. Like Anthropic's job is to make sure that Claude Code still works incredibly well when Claude (Anthropic model family) comes out. Because the current harness is kind of optimized to the current generation models. They're going to have to change it dramatically. They actually had a pretty interesting blog post on Twitter where they talked about these huge changes they made to their context management system as the models got smarter over the past few months. I think obviously if an individual company was going to invest the same amount of effort that Anthropic puts into making the Claude Code harness into optimizing for today's models, then the switching cost would be very high.
John Furrier
>> I mean, they're basically building in low switching costs by the fact that they have to swap their models out. Inherently, that's what they have to do, is what you're getting at.
Max Kan
>> Yeah, or I would just say they can put in the upfront effort to make a good harness for this generation of model. It's not worth it for anyone else to put in that effort. they should just use the Anthropic and OpenAI harness.
John Furrier
>> So final question for you, just because I'm curious what your thoughts are, because I riff on this all the time and I don't really know the answer, but it seems distillation's a huge problem with the big models. Does that become a lab issue in terms of moat protection? Is there a movement afoot to say, okay, distillation just costs money? Is that a feature? Is it a bug? what are your thoughts on the risk of distillation? Me kind of going using theCUBE model to distill off of the main model and build my own kind of media model. people are looking at the distillation. I see the Chinese did that a lot.
John Furrier
>> Yeah.
John Furrier
>> Are people talking about this or what's your reaction and view on distillation, competitive advantage, or is it just annoyance to them and it's not a factor?
Max Kan
>> Yeah, I'd say it's definitely an open research question, sort of how important distillation is to, sort of like how good Kimi K2 is. I really wish some AI researchers would just try to answer this very quickly because I think it is worth noting that distillation has gotten much, much harder over the past year or two. Historically, distillation is actually done if you knew the log probs of all the next token predictions. Obviously, you don't get that for the models today. The Frontier Labs are also hiding all the reasoning traces. So really, if Moonshot really is distilling Fable, they're sort of doing it even... So obviously, they don't have the log probs. They know the actual probabilities for the next tokens. And they also don't even have the reasoning traces available. They only have the final output. Is it possible that you can still do some serious distillation with just this neutered piece of data? Maybe it's possible, I think this is an open research problem.
John Furrier
>> It's probably going to be trash anyway, but you're getting a little bit of nuggets out of there. Yeah. Okay, you guys are on the tokenomics team at SemiAnalysis, so you guys do great research. Put a plug in for what you're working on right now. What's the team focused on? What are some of the core things you're watching? All the horses on the track? What's the focus? What do you want to know?
Max Kan
>> Yeah, yeah, yeah.Yeah. First, I would say the one-line description of what Tokenomics is, is we study the economics of tokens. What I like to think about it is if we're actually going to pave the world with these AI factories as we undergo the largest infrastructure build of all time, it'd be really nice to know what are people using these factories for? Are they making money? What are the margins? These are the types of questions that we try to answer on the Tokenomics team. Some products I'm working on right now, we're trying to make our own benchmark. I think a lot of the existing benchmarks today are quite bad, and I think having really good private benchmarks lets you say interesting things about new model releases. I'm really trying to understand, can you break down global token volumes over time by different models, by different providers, going towards what I mentioned earlier. I think whether or not frontier token share is decreasing or increasing is just a super important thing that people need to track.
John Furrier
>> It's a hard job because I remember we used to track cloud market share, no one really reported cloud market share. We had to squint through earnings and oh, the lawsuit came out, they disclosed that piece of data. So there's a lot of ball hiding going on from the big labs right now.
Max Kan
>> Yeah, yeah, yeah. We read through every disclosure you could possibly imagine and we try to triangulate the pieces.
John Furrier
>> Hunting, putting that puzzle together, Max. Thanks for coming on theCUBE and the NYSE Wired for our third annual AI Leaders. Great stuff. Say hello to Dylan and the team. Appreciate you coming in and sharing your perspective.
Max Kan
>> Yeah, thanks for having me, John. It's a lot of fun.
John Furrier
>> I'm John Furrier with theCUBE. More live coverage here in Palo Alto. Got the big event here in Palo Alto, right here tonight. 180 industry friends getting together to talk about these questions. where are the economics? What's the size of the population? What key hardware is coming out? What's the next chip? Tons of action in AI infrastructure as the build-out continues. Once it's built out, you got to operate it. And you got to know where the money is. So that's what we're focused on. Thanks for watching.
>> Hello, I'm John Furrier with theCUBE here at our Palo Alto studios for theCUBE and NYSE Wired, it's our third annual AI Infrastructure Leaders event covering the AI infrastructure build out and boom of course, physical AI right around the corner. I want to thank ScaleFlux for supporting us as well as all of our industry sponsors over the year. Really appreciate it. It's been a really growing community. Our first guest to kick it off is Max Kan, tokenomics technical lead at SemiAnalysis. Max kicking off the program, we got our big 180 folks coming to the event tonight. Little networking nerd fest here in Silicon Valley. Thanks for coming on theCUBE and NYSE Wired, third annual event.
Max Kan
>> Yeah, thank you for having me, John. Pleasure to be here.
John Furrier
>> So you guys just do great work, SemiAnalysis. In the industry, everyone kind of knows what you guys do, but you guys do deep research on all things AI, obviously from the infrastructure, what's going on at the Pareto curves, token economics, all the nuances around what's powering the AI infrastructure. Everybody wants to know about token economics. We saw the big paper that went out this week. Jensen put it out there, like 25 million views on X. Everyone's citing, everyone's jumping on the bandwagon. Open weights, you're seeing open source surge, you're seeing Anthropic OpenAI trying to go public. Robotics is booming, we think it's going to be a big year next year, so tokens are a huge part of the AI piece. What are you seeing in the market? What's your focus?
Max Kan
>> Yeah, I'd say here on the Tokenomics team, we're sort of at the top of SemiAnalysis full stack coverage. So at the very bottom, we have things like, Accelerator team that tracks where the accelerators are, Data Center team will track all the data centers. Tokenomics, you can maybe say the stated goal is to track all the tokens, but our kind of primary coverage area today is the hyperscalers and the AI labs. Obviously, open source has been a big deal for the past week or two, ever since the Kimi K3 blog post with all the crazy benchmark scores came out. And then we've seen this big fight on Twitter and elsewhere with Jensen and it kind of seems like the rest of the AI community really pushing for open source, even OpenAI signing on and kind of Anthropic on the other side. The way that I view that letter personally is I think it's mostly Jensen talking his book. He knows that he can't be sort of fully beholden to only two customers, Anthropic and OpenAI to buy all this compute. And so he very desperately and actively wants to nurture kind of an open source ecosystem. Whether or not open source AI is good or bad, I think probably not good overall. But I think a lot of the sort of discourse going on is mostly people talking their book and less like maybe moral principle discussions.
John Furrier
>> Frame the discussion for the folks that aren't inside the ropes because you have, kind of talking their book, their messaging, they have their strategies, AMD clearly playing the open card, NVIDIA playing open as well. They say a lot, hey we're open, and they're supporting open source. But they're kind of two different approaches. you look at both of them, so at the end of the day, do I really care if I got the best tokens per watt per dollar? I just want to get more tokens, and I don't want to pay through the nose for tokens if I don't need to, take us through this, frame up the landscape, the playing field, the arena. What's going on with token economics? Who's where, what's what, and who's leaning in, and who's moving the needle in terms of pushing the innovation for developers and companies?
Max Kan
>> Sure, yeah, so I would say at one end of the spectrum, you definitely have Anthropic, which has historically been sort of the most AI safety conscious AI lab. They've sort of, even back when I think GPT-2 first came out, when Dario was still working at OpenAI, he was saying it's potentially too dangerous to release to the public. We've got to make sure we do some rigorous safety testing here. Obviously that was maybe a bit hyperbolic, but I think a lot of his sort of original claims and fears have started playing out to a degree. At the same time, they're obviously a closed source lab that has incredibly high API margins today. I think if you're paying API price for Opus or Sonnet, there's a very good chance that's like 85 % plus gross margins for Anthropic. Maybe even better than traditional SaaS, which is very much the opposite of what people were sort of characterizing the AI build -out as just maybe six months ago. But anyway, so they're on one end. I'd say on the other end, you have the open source guys. this is primarily Chinese labs today. But obviously, there's a kind of moral and philosophical movement with open source where people say, you just want technological development to be free, you don't want the government or any one lab regulating everything or especially controlling everything. From the economic side, obviously the open source guys have worse margins. So I think if you want to take maybe the cynical business point of view, it's like you have Anthropic who wants to defend their lead in AI right now, they want to defend these 85 % plus gross margins, you have the open source guys who have less margins, they're trying to eat into that. But then obviously there's the more moral angle as well, which I think Anthropic does truly believe to a degree, I don't want to, I don't think they're purely being cynical here, but that's kind of the landscape.
John Furrier
>> Yeah, and you see a lot of people, if you're building something, AI native companies tend to do well, we saw Fireworks AI is crushing it. They're doing well, so there's a lot of demand for tokens. There's a lot of development going on. Now you got the suppliers and it's not obvious. We're supposed to get lower gross margins. We thought Anthropic would be going down, not up. How do the economics shift? Because when you start seeing that balance between general intelligence and specialized intelligence, you're going to have this kind of, not mutually exclusive environment, but more integrated environment. A lot of people are talking about small language models. Are people behaving differently with how they think about tokens? Given that the compute's now back on the table for inference, the role of the GPU. in your guys' analysis, how does the token piece fit into some of those economic discussions?
Max Kan
>> Sure, yeah. So I would say historically, basically all the revenue has been concentrated just on the smartest frontier models, which is why you see Anthropic and OpenAI have such crazy ARR numbers today, growing faster than, no business has ever grown this fast at this scale in the history of capitalism. It's honestly quite impressive. Right, so obviously though the open source models have gotten kind of remarkably good over the past maybe call it the past month and so the big question on everyone's mind is like is this finally going to change? Maybe sort of a precise kind of way toward it is sort of one metric I think people should start looking at is if you were to plot global token volumes over time is the percentage going to frontier models decreasing or increasing? Obviously if that number starts decreasing dramatically, it means that open source models are becoming increasingly popular. If it starts sort of, if it continues increasing, then it's kind of just more of the same old, same old. Admittedly it's sort of very hard to collect data to actually answer this question. We are trying very hard at SemiAnalysis. Hopefully we should have something out soon, maybe in the next couple months or so. Big reason why this is difficult is that Anthropic has never published a token production number so you kind of just back it out from their revenue and other disclosures and stuff.
John Furrier
>> So you guys are sifting through all that to try to piece it together.
Max Kan
>> Yeah, yeah, that's our job. Sort of sift through all the minutia and kind of distill it down into takeaways people can actually understand.
John Furrier
>> Max, I'm old enough to remember when the internet came, everyone's like, oh, look, the internet's going to be great. And that's when we saw that bubble firsthand. And one thing that was interesting, Mary Meeker was an analyst at Morgan Stanley at that time, and she had this online population metric, the total online population, because people were like, oh the internet, first the geeks get on there, the people who know the internet, and then the general, everyone else used it. So that was a key metric. So the question is, in tokens, is there a view on total token usage? Because as more people are using AI, that's kind of a similar online population of users. What is your view on the general purpose? Because everyone's on the consumer side, I see people in New York City using the chat all the time on OpenAI, but business people tend to go towards Anthropic. Gemini's got the embedded Google advantage of being embedded in the search. So you have all these different distribution pieces, but is the overall pie getting bigger?
Max Kan
>> Yeah.
John Furrier
>> And how does that factor into how does one squint through who gets what tokens? Is there a power law? these are kind of like big picture questions. What are your thoughts and reaction to that?
Max Kan
>> So I think token production is just very clearly going exponential. So the total pie. I would guess that total global ex-China tokens a day is maybe a little less than a quadrillion tokens a day. So this is obviously a monster number. Big reason or a big part of what's driving the exponential growth is the shift to agentic workloads. I think people, for those who don't understand how inference actually works, the basics. In general, agentic workloads consume exponentially more tokens than the chat workloads everyone was using a year ago because it's all inherently multi -turn and every time you ask your agent a follow -up or it uses some tool like web search or maybe it searches your code base, it has to reprocess all previous tokens in your conversation, which really causes total tokens produced slash processed to go up exponentially.
John Furrier
>> Again, going back to what we were talking about You mean like looping, they have to do multi -step reasoning, they have to do a lot more than just getting an answer.
Max Kan
>> Yeah, it's not quite the same as looping. I said when people say looping, they generally refer to like, I just put my agent in a loop and then tell it to keep doing the same thing over and over again.
John Furrier
>> Okay, so that's specific, okay.
Max Kan
>> Yeah, this is just more inherent to how technology actually works. It's like every single time you ask the agent a follow -up, or, and this is actually where the KV cache comes in too, like every single time you ask the agent a follow -up or the agent uses a tool, it sort of has to reload the key and value vectors for all previous tokens into its context in order to generate the next token. And this just causes exponential growth of token usage. It's also why people, a lot of people like to look at the blended price per million tokens over time and try to use that as some gauge of like, are open-source models becoming more popular or not? The one caveat I think anyone doing this analysis needs to consider is that workload shape absolutely matters. As people know, there are three main types of tokens. There are the cached input tokens, the base input tokens, and the output tokens. Those are priced very differently. And in general, agentic workloads, cached input tokens go way up, and also regular input tokens go way up. Those are much cheaper than the output tokens, and so you could see your blended price per million tokens go down, even though it's just a function of workload shape changing and not actually model usage changing. And so that's one thing people should look out for if they're doing this now.
John Furrier
>> So what you're getting at is that it's complicated in the sense of the token relationships to the workload task could be very complex or elementary and those token decisions make a difference. Sounds like there's a lot going on under the covers.
Max Kan
>> Yeah, yeah, yeah.
John Furrier
>> And one of the things at GTC this year, when Jensen put up the Pareto curves, saying to Dave Vellante and folks like, okay, now we're starting to see premium tokens because if you're on a Vera Rubin, you're going to be paying out the nose for those tokens. But I don't want to use those tokens to do calculations that I can get for maybe a better price. So context, not content, token routing is starting to come into it. You're starting to see a lot more token stuff under the covers. You mentioned agents. There's so much going on. How does that factor into the analysis? Because it's blended. Is there an intelligent algorithm in there? How do you see this evolving? because people are starting to talk about, okay, if I'm going to have this design system, I will use maximum premium tokens for this, but I want to make sure I'm not going to suck up the GPU to do basic stuff. I won't say low-end stuff, but basically it's lower-end calculations.What's your view on that piece?
Max Kan
>> So I would say actually on the Pareto curves, I think you definitely will be using Vera to generate cheap tokens, actually. I think Vera will give you sort of the cheapest tokens, whether it's per dollar or per watt, that you could possibly ask for. This sort of really premium token that Jensen was talking about at GTC, that's what Groq gives you. And this kind of goes down into the actual architecture of what does a Groq SRAM chip versus an NVIDIA Vera CPU look like. I'm not going into the details here, but TLDR is that Groq is super expensive, but it can generate tokens really fast. People have historically shown they're willing to pay a premium for fast tokens. We saw Anthropic at one point was charging 6x the price for just 2.5x the speed. and SemiAnalysis was happily paying that. Fortunately, now it's only actually 2x the price for still 2.5x the speed, so that was pretty good for our bottom line. But that's sort of like the Pareto curve Jensen was talking about.
John Furrier
>> In terms of performance, I know you guys spend a lot on tokens in your research.
Max Kan
>> Yeah.
John Furrier
>> Cerebras was pumping up their thing, they just did a deal with AMD, they claim to be the fastest. They have a big chip, a lot of SRAM involved. And then the diversity piece of the vendors, because some are saying Groq's okay but expensive, there might be other alternatives for inference, but that's working, but then it's like, okay, they have the IP, but there might be other solutions. How do you think about the alternatives, like a Cerebras, like an AMD, like an Nvidia, as people start to figure out where the costs are?
Max Kan
>> Yeah, I think if you believe like I do, that inference is going to become sort of the largest market ever, there will be a place for sort of all of these niche accelerators. The question is just how big. Personally, I think most of the volume is actually still going to just go to, at least it's not going to go to these SRAM chips, Groq and Cerebras, because they're more niche, they're much more expensive, really the only people, it's like people who are willing to pay 10x the price of current API prices will be willing to pay, like, OpenAI announced they're going to have 750 tokens per second on Cerebras in July. They're running out of time here. We'll see if they actually still hit that deadline. But that's probably going to cost like 10 times more than their regular API price. I think it's a very niche segment of the population is going to be using that. I don't even know if it's in budget for SemiAnalysis. I'd have to ask Dylan. So I think most of the volume will be going to the general purpose chips. I think it is an open question whether maybe a Positron or an Etched or one of these guys will be able to create potentially even a better general purpose inference chip than NVIDIA. I don't have super strong opinions there. I haven't spent too much time looking at them.
John Furrier
>> But token maxing, we saw that wave.It was like, hey, look at all the tokens I'm using. Now the conversation on X is value maxing. So there's really a movement towards the old school coding days. Less code is better. Is there a new metric on token to outcome? Are people starting to look at the revenue? Is there a new metric behind just the mechanism on token to value. Are you seeing, tracking anything there?
Max Kan
>> I don't think there's any one golden bullet metric you can use for every enterprise. I think it really depends. Every single company sort of has to look at their token spend honestly today and try to calculate the ROI on that for their own business the same way they would sort of see whether or not this extra employee headcount is worth it for their business. I don't think we'll ever have any one golden metric.
John Furrier
>> They're all different too. They have different business models, different objectives, different workloads, different agents.
Max Kan
>> Yeah, I think some people have floated the idea of maybe charging sort of per task or based on outcome instead of per token. I think that maybe makes sense in some use cases, but at the same time, I think really one of the benefits of having this general purpose technology is that OpenAI or Anthropic have no idea what sort of like all the great, crazy, useful, high ROI things people are going to use it for. And so I think token -based pricing will be here to stay for at least the next year or two.
John Furrier
>> Give your bull and bear use case view on OpenAI and Anthropic. What's the bull bear analysis? How do you look at that? What do the scenarios look like? How could the ball drop on either side?
Max Kan
>> Yeah, so I would say the bull case for these two companies, and actually, are you asking about sort of like the two companies as one unit?
John Furrier
>> No, no, no, two separate companies as individuals going public.
Max Kan
>> Oh, I see.
John Furrier
>> You said that they're the fastest growing companies in the history of capitalism, right? So, to me, I think there's a power law development, I think you're going to have the big guys doing general and then I think it's going to be specialism. But there's a lot of debate around, are these going to be viable companies? And the neoclouds are going through the same thing, where it's like, okay, they're getting some beachhead, they're backstopped by the big guys, you're seeing that kind of funding. So there's definitely, and his goal of getting alternatives out there is working, but will they be viable? That's the question. Some are leaning toward the bull case, like this is a game changer, We don't know yet what the economics look like, hasn't been modeled yet. The bear case is like, well, this is a bubble, it's going to pop. How can they sustain that cost on the CapEx? So you have kind of competing views and sentiment around these companies. Now, my opinion, I won't share, but you kind of get where I'm going. I'm kind of on the bull side. I think, scale matters and billions of users is a good thing.
Max Kan
>> Yeah, maybe I'll start with the bull case or sort of frontier closed source labs in general. and then the bear case for them, and then I'm going to talk about OpenAI and Anthropic separately. So I think in terms of the bull case for why at least one of OpenAI or Anthropic could become maybe the first $10 trillion company within the next couple of years, is if the percentage of global token volumes going to frontier models keeps increasing. And I think if that happens, what it means is that sort of the new economically viable tasks that are unlocked by increasingly smart frontier intelligence, sort of the TAM of that is growing faster than the easier tasks that will be switched over to these cheaper open source models. I think these are sort of the two key lines that everyone needs to watch. Historically, obviously, the answer is just that that top line has grown way faster than the bottom line. That's why Anthropic in particular, the ARR growth has significantly outpaced the rest of the industry, which has also been growing very quickly since the start of the year. I think really it depends on like are they going to be curing cancer and doing all these other crazy technological things because obviously regular software engineering is not going to need sort of the smartest model probably a year from now so what new use cases will be unlocked I think that's sort of the key question for the AI labs the frontier closed source AI labs as a business model obviously the bear case then is that people don't actually find these new use cases fast enough. And six months from now, the best Chinese open source model is sort of just sufficient for the majority of white collar work today. Then I think bubbles popped, it's all over.
John Furrier
>> So basically the adoption has to get there.People have to taste it, get addicted to it, understand it, use it, get value out of it quickly. And if that adoption value piece doesn't kick in.
Max Kan
>> Exactly. It has to be fast. And they have to find things that aren't just like summarizing emails or doing basic software engineering. It sort of has to be like fundamentally new tasks that are just not possible without the increase in the smart frontier intelligence. And then to this point, I think if you want to differentiate between OpenAI and Anthropic, really what matters is just who has the better model. I think we've seen that switching costs are incredibly low in this industry. You can maybe make some arguments like, oh, with the advent of Claude Code and Codex and sort of this thesis, the agent is the product, it's not just the model, you have to have this good harness too, and the switching costs are slightly higher. I don't really buy this. I would say personally, I've switched between kind of Claude Code and Codex two times in the past two weeks, just because it's not that hard to change your harness. And I think the reason why Anthropic really overtook OpenAI over these past seven months is that last November, when they came out with Opus 4.5, they just had by far the best agentic model in the world. OpenAI was looking extremely dire. Their models were basically complete trash until this April or May when they came out with GPT-5.5. And then that's also the point when we saw their revenues start accelerating again. So I think really for the AI lab.
John Furrier
>> It's like an F1 race. You don't know, you're in the lead, you drop back, but you're getting to a good point around the model. So a better model wins. So you think switching costs are low or high?
Max Kan
>> Super low. Super low.
John Furrier
>> You don't think that recompiling tokens and managing, not recompiling, but recasting the agentic workflows matter or is this irrelevant?
Max Kan
>> Are they already decoupled? So I would say, this is maybe a slightly more technical point, but there is an argument to be made that if you really invest your effort into developing a set of evals, so you can rigorously test model performance. and then you also customize your entire harness and all your context management system for a particular model, then the switching costs will be pretty high if you were to rip it out. But I think people aren't actually doing that today and I think investing the time to do that is honestly a mistake.
John Furrier
>> So the harness is a key.The key is the harness and making sure that you kind of decouple core operations from the token models.
Max Kan
>> Yeah, but I'd say harness engineering is more, that's sort of like OpenAI and Anthropic's job. Like Anthropic's job is to make sure that Claude Code still works incredibly well when Claude (Anthropic model family) comes out. Because the current harness is kind of optimized to the current generation models. They're going to have to change it dramatically. They actually had a pretty interesting blog post on Twitter where they talked about these huge changes they made to their context management system as the models got smarter over the past few months. I think obviously if an individual company was going to invest the same amount of effort that Anthropic puts into making the Claude Code harness into optimizing for today's models, then the switching cost would be very high.
John Furrier
>> I mean, they're basically building in low switching costs by the fact that they have to swap their models out. Inherently, that's what they have to do, is what you're getting at.
Max Kan
>> Yeah, or I would just say they can put in the upfront effort to make a good harness for this generation of model. It's not worth it for anyone else to put in that effort. they should just use the Anthropic and OpenAI harness.
John Furrier
>> So final question for you, just because I'm curious what your thoughts are, because I riff on this all the time and I don't really know the answer, but it seems distillation's a huge problem with the big models. Does that become a lab issue in terms of moat protection? Is there a movement afoot to say, okay, distillation just costs money? Is that a feature? Is it a bug? what are your thoughts on the risk of distillation? Me kind of going using theCUBE model to distill off of the main model and build my own kind of media model. people are looking at the distillation. I see the Chinese did that a lot.
John Furrier
>> Yeah.
John Furrier
>> Are people talking about this or what's your reaction and view on distillation, competitive advantage, or is it just annoyance to them and it's not a factor?
Max Kan
>> Yeah, I'd say it's definitely an open research question, sort of how important distillation is to, sort of like how good Kimi K2 is. I really wish some AI researchers would just try to answer this very quickly because I think it is worth noting that distillation has gotten much, much harder over the past year or two. Historically, distillation is actually done if you knew the log probs of all the next token predictions. Obviously, you don't get that for the models today. The Frontier Labs are also hiding all the reasoning traces. So really, if Moonshot really is distilling Fable, they're sort of doing it even... So obviously, they don't have the log probs. They know the actual probabilities for the next tokens. And they also don't even have the reasoning traces available. They only have the final output. Is it possible that you can still do some serious distillation with just this neutered piece of data? Maybe it's possible, I think this is an open research problem.
John Furrier
>> It's probably going to be trash anyway, but you're getting a little bit of nuggets out of there. Yeah. Okay, you guys are on the tokenomics team at SemiAnalysis, so you guys do great research. Put a plug in for what you're working on right now. What's the team focused on? What are some of the core things you're watching? All the horses on the track? What's the focus? What do you want to know?
Max Kan
>> Yeah, yeah, yeah.Yeah. First, I would say the one-line description of what Tokenomics is, is we study the economics of tokens. What I like to think about it is if we're actually going to pave the world with these AI factories as we undergo the largest infrastructure build of all time, it'd be really nice to know what are people using these factories for? Are they making money? What are the margins? These are the types of questions that we try to answer on the Tokenomics team. Some products I'm working on right now, we're trying to make our own benchmark. I think a lot of the existing benchmarks today are quite bad, and I think having really good private benchmarks lets you say interesting things about new model releases. I'm really trying to understand, can you break down global token volumes over time by different models, by different providers, going towards what I mentioned earlier. I think whether or not frontier token share is decreasing or increasing is just a super important thing that people need to track.
John Furrier
>> It's a hard job because I remember we used to track cloud market share, no one really reported cloud market share. We had to squint through earnings and oh, the lawsuit came out, they disclosed that piece of data. So there's a lot of ball hiding going on from the big labs right now.
Max Kan
>> Yeah, yeah, yeah. We read through every disclosure you could possibly imagine and we try to triangulate the pieces.
John Furrier
>> Hunting, putting that puzzle together, Max. Thanks for coming on theCUBE and the NYSE Wired for our third annual AI Leaders. Great stuff. Say hello to Dylan and the team. Appreciate you coming in and sharing your perspective.
Max Kan
>> Yeah, thanks for having me, John. It's a lot of fun.
John Furrier
>> I'm John Furrier with theCUBE. More live coverage here in Palo Alto. Got the big event here in Palo Alto, right here tonight. 180 industry friends getting together to talk about these questions. where are the economics? What's the size of the population? What key hardware is coming out? What's the next chip? Tons of action in AI infrastructure as the build-out continues. Once it's built out, you got to operate it. And you got to know where the money is. So that's what we're focused on. Thanks for watching.