This interview examines Tensordyne's tape-out milestone, log-domain compute approach and rack-scale system strategy for agentic artificial intelligence workloads. It explores implications for inference compute, data center and edge deployments, and system-level trade-offs in compute density, memory architecture and interconnect design.
Marc Bolitho of Tensordyne, founder and chief executive officer, discusses the company's log-math architecture, tape-out progress and system-level vision for inference at scale. Bolitho describes how log-domain math converts multiplications into additions, enabling smaller compute engines, greater on-chip SRAM, lower power and higher tokens-per-second-per-watt, and they argue this approach allows multi-trillion-parameter, agentic models to run within a single rack at significantly reduced cost and energy. These insights inform AI accelerator design and inference compute strategies. John Furrier of theCUBE Research and Gabe Olave of theCUBE Research lead a technical conversation on compute density, SRAM trade-offs, interconnect partnerships and how Tensordyne positions itself against established players while targeting enterprise and neo cloud deployments. theCUBE analysts underscore implications for edge AI and data sovereignty as well as neo cloud differentiation.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: The AI Factory - Data Center of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Marc Bolitho, Tensordyne
This interview examines Tensordyne's tape-out milestone, log-domain compute approach and rack-scale system strategy for agentic artificial intelligence workloads. It explores implications for inference compute, data center and edge deployments, and system-level trade-offs in compute density, memory architecture and interconnect design.
Marc Bolitho of Tensordyne, founder and chief executive officer, discusses the company's log-math architecture, tape-out progress and system-level vision for inference at scale. Bolitho describes how log-domain math converts multiplications into additions, enabling smaller compute engines, greater on-chip SRAM, lower power and higher tokens-per-second-per-watt, and they argue this approach allows multi-trillion-parameter, agentic models to run within a single rack at significantly reduced cost and energy. These insights inform AI accelerator design and inference compute strategies. John Furrier of theCUBE Research and Gabe Olave of theCUBE Research lead a technical conversation on compute density, SRAM trade-offs, interconnect partnerships and how Tensordyne positions itself against established players while targeting enterprise and neo cloud deployments. theCUBE analysts underscore implications for edge AI and data sovereignty as well as neo cloud differentiation.
play_circle_outlineAgentic workloads driving massive inference demand and token consumption.
replyShare Clip
play_circle_outlineNVIDIA, Groq, Cerebras, Google Debate Heterogeneous Compute as Tensordyne's Log‑Domain Math Targets Tokens/sec/Watt and On‑Chip SRAM/HBM
replyShare Clip
play_circle_outlineClaims of order-of-magnitude cost and energy advantages versus current multi-rack setups.
replyShare Clip
play_circle_outlineTaped-Out Chip Powers 72-Accelerator, Quarter-Rack 30 kW Air-Cooled Pod for On-Prem, Telco, Edge Deployments
replyShare Clip
play_circle_outlineTensordyne's Sovereign Localized AI for NeoClouds and Telcos with Broadcom, HPE, Juniper — Chips Fall, Q1 2027 Beta
>> for compute and for accelerated computing. Marc Bolitho here, the CEO of Tensordyne Company. We've covered coming out of stealth. Marc, great to have you on theCUBE. Good to see you. John, great to be here. Thanks for having me at the Stock Exchange. It was great to cover the unveiling in the Palo Alto through the RK and Gilles in Paris, the RAISE Summit in Paris. But you guys are sitting in the middle of the innovation storm of the computer industry. if you look at the build-out, it is probably the biggest expansion economically, physically, just capabilities. The AI is essentially resetting the architecture from the old compute days. We saw the rise of the computer industry, the Internet revolution, SAS, cloud. But building on that cloud scale is the ongoing demand for compute, networking, and storage in a very dense package because it's mathematics now, right? So math is getting commoditized in a good way because it's spreading everywhere. But feeding the math with more math, that's inference. This is the difference between old school days when you have packets come in, it's processed. You have a whole other paradigm. It's mathematics. You hear about vectors. You hear about reasoning. And the rise of the agents brings up that this isn't like a search paradigm. It's a reasoning intelligence paradigm. That's the demand right now. Now, you guys have got a taped-out chip with some DNA from Juniper, which we covered. You guys got a different approach. Explain the approach of Tensordyne. You guys had four years of stealth. Probably you were itching to talk more about it, but now it's out there.
Marc Bolitho
>> Yeah, John, absolutely. if you look at this market, this is the largest market. We're going to approach a trillion -dollar TAM in inference alone. And right now, everything is being built upon these agentic workloads. So Anthropic plus OpenAI, they have a $100 billion ARR right now together. All that's being built because of agents. It's finally AI is doing real work for us. So whether it's coding or other workloads, it's useful and it's getting real jobs done. But these models are burning a ton of internal tokens. We're talking orders of magnitude more than just a chat model now to run agents. And so when you do that, you have a need for better speed. You have to have better speed when you're running these agentic workloads because you're not just communicating with a human anymore. Now you're doing tool calls, communicating with the next agent, and you've got to be able to process this very quickly. So what we did is early on when ChatGPT first came out, we were looking at our log math, our architecture we had for doing inference. And what we realized very quickly is that this architecture in log math was perfect for generative AI inference. And the reason for that is fundamentally when you're doing math in the log domain, you can do all of the trillions of multiplies in your matrix algebra. You can do all of those in a much smaller area of compute. And when you have that smaller area of compute on the chip, you can exchange that. You can exchange that for memory, exchange that for data flow and for networking, and you can create a very balanced system. So that's what we're doing, John, differently than anybody else.
John Furrier
>> It's how we're approaching the math. The log math piece. Let's explain that because I did a post last week on LinkedIn around. I get all the questions. I just had to plug in all my notes into Claude and it gave me a great one-pager to explain everything. You got the density of NVIDIA, and they got software to manage that. You see things like. I was learning about systolic arrays at the Hot Chips conference. And that's Google's approach. It's going through memory to multiple arrays to get that multi-step process. You got Cerebras. You got a variety of different approaches. It's all math. Yes. So mathematics is the key. And the floating point piece of the processor, which became the GPU for gaming and graphics, became the state of the art for AI. Talk about that math piece. Why is the math so important in this new equation? Because that's the difference in the AI generation versus the old compute days that we all kind of grew up with PCs and servers. Now it's different. You've got the mathematics piece.
Marc Bolitho
>> Yeah, so everyone else is working in floating point. Our thought early on was, if you're going to come up with a new solution for running AI, how do you differentiate? You've got to go to first principles. You have to focus on the math. So fundamentally AI is trillions of multiplications followed by trillions of additions. And so what we do is we work in the log domain so log of A times B is equivalent to log A plus log B which means we can now change every multiply into an add. When you do that you have a much smaller area on the die for doing all those multiplications and you have lower power and that lower power scales from the chip to the pod to the rack all the way to the data center. So that's fundamental.
John Furrier
>> Talk about what you guys hope to change in the industry, because you brought up a couple of key things that are constrained right now, power and speed. Everyone wants faster tokens per second per watt. Yes, that's fact. And now everyone's starting to see the benefits on the revenue side, tokens equal revenue. That's kind of well understood now. Why is it important to have that speed? And then what do you guys hope to change in the industry?
Marc Bolitho
>> Yeah, it's very important, as we said, to have that speed. today you see that to get into, say, the 800 to 1,000 output tokens per second per user. Companies are looking at disaggregated systems, so a Rubin for pre -fill, multiple Groq units, racks, say, eight Groq racks for the decode side, looking at, say, AWS doing the pre -fill with Trainium and multiple racks of Cerebras to do the decode. So what our architecture allows us to do is that because our compute die is so small, we can put more SRAM on the chip. And that means we have a highly utilized compute engine. So if you have a compute engine and you're only using 10 % of it, that's not very valuable. You have to be able to have an architecture that gets the most out of the silicon area that you have. And we do that with log math. We have an abundance of SRAM, as much SRAM as say a Groq or a Cerebras, but we also have high bandwidth memory. So we don't need multiple racks to run these 2 trillion parameter models. We can basically, in a nutshell, instead of having nine to 14 racks to run a 2 trillion parameter agentic model, we can run it in one rack. And we can run it at an order of magnitude less cost.
John Furrier
>> How does that change economics? Because you're accelerating the computing, but at lower cost and lower power.
Marc Bolitho
>> That's right. So, if you look at it today, to get these higher output speeds, the costs rise dramatically. And so you have a trade -off you have to make between intelligence and cost. And we need to get into an era that you can pick the best intelligence, you can pick the best model, and you can pick to have your responses come back very quickly and not sacrifice the cost. And so what we're seeing is in those higher output speeds, we have a 10x cost advantage.
John Furrier
>> Talk about the architecture compared to NVIDIA, because people say they hear NVIDIA, accelerated computing. That's their big narrative. What do you guys hope to define? If NVIDIA is accelerated computing, what's Tensordyne?
Marc Bolitho
>> Yeah, so we're definitely doing accelerated computing as well for AI. But think about it as a system that's future proof for the largest AI workloads and at the highest user speeds, at the lowest cost and the lowest energy. We haven't even mentioned energy yet because energy is very important as we look to scale out AI as well.
John Furrier
>> I mean, all the Neo clouds out there we're talking to, energy is the number one thing and finance. Those are the two bounding functions. Money, which we're seeing a lot of CapEx, great for now, but energy is going to be an ongoing sticky problem, especially when you look at connecting AI factories at, say, the edge or metro areas. We're going to see that progression very quickly, I think, from centralized factories to distributed computing. You mentioned the prefill-decode. Disaggregated serving is the hottest topic right now. Because the KV caches are getting so filled. Because there's now too much data coming in. When you have disaggregated serving, you're going to have disaggregated infrastructure. Yes. So you're going to have racks everywhere. Yes. Talk about that piece, because that's a huge determinant on the power side. You can't drop an NVL72 at a telco tower. That's right. Or maybe even an enterprise.
Marc Bolitho
>> Yeah, and so there's different ranges, absolutely. And so if you look at going from an enterprise or from a telco all the way up to a large data center, to do disaggregated inference today for these large MOEs, agentic models, 2 trillion parameter plus. Companies are deploying, say, nine racks of compute to do that. And to get those high user speeds, companies have said that's going to cost, say, $150 per million tokens and be 1.5 megawatts of power. We can do that in 120 kilowatts. So we can do that at a fraction.
John Furrier
>> What's the rack requirements? You mentioned 120 kilowatts. What's the baseline?
Marc Bolitho
>> So our pod requirement where we have 72 accelerators, think of 72 accelerators like a Blackwell also having 72 accelerators. We do that in a quarter of a rack with 72 accelerators at 30 kilowatts, and it's air-cooled. So you can fit into any existing infrastructure or data center. You don't need a liquid-cooled data center. You can work with stranded power or existing data centers. So it works for enterprises. It works for telcos. it works all the way up the chain.
John Furrier
>> We're seeing a lot of demand for, again, I mentioned distributed computing, but distributed intelligence. So, for example, if I'm an enterprise, I might have a little data center. Sure. Not little, but enough power to run in the classic sense of traditional data center. But in the modern era, it might not have the requirements for the monster systems. Okay, but I'm going to connect to, say, CoreWeave or Nscale, even AWS. I'm going to run an on -prem hybrid computing platform, so I'm going to be distributed. But I can have single -tenant -like experience on these Neo clouds and AI clouds, as well as the hyperscalers like AWS. We're already seeing AWS with VPCs handling that. So as this becomes distributed, you guys fit nicely. So explain where you fit, because if I'm an enterprise or I'm a telco, I'm not going to have a gigawatt to play with. That's right. I'm going to have a beefy edge or a data center that's kind of in scope, in a traditional sense.
Marc Bolitho
>> So think of being able to run multiple instances of a Kimi K3 model with hundreds, maybe a thousand concurrent users in a quarter rack at 30 kilowatts. It can be all on -prem and you can have your own fine tuning of your models, all your data stays on -prem and for larger workloads, you can either scale up to the extent of your on -prem or then, as you said, have distributed back to the cloud.
John Furrier
>> And what does that mean for say, agentic? Because a lot of the, whether it's a hospital or financial services or say, in retail, they don't have big data centers, but they want intelligence everywhere.
Marc Bolitho
>> Yes, what that means is that they can run agentic workloads in their own data centers. That's what we enable.
John Furrier
>> What's the token cost? Because again, a lot of people have been token maxing. Sure. I was talking to someone at Stripe. They're like, yeah, the token days are such a headache. It's like running a test dev environment on our production payment system. Yeah. people are using production -level tokens to essentially do stuff. That's essentially what's happening today. So, okay, the next step would be, okay, logical IT solution would be get me a box on -prem and just have unlimited development tokens and then scope it, then move it to production. Maybe it's a hybrid, Pareto curves might be different. But I don't want to blow token budgets. How is that factored into some of the conversations you have with large deployments? because that's a cost structure.
Marc Bolitho
>> And I think that's one of the things that have been preventing that build out and the scale out at the enterprise level is the cost of breaking into it. And so there has to be that cost -benefit analysis that's done for running these workloads. What we're seeing is with these 2 trillion parameter agentic models that we're going to be able to do this at an order of magnitude less cost than what the systems are that are going to be launching now in this generation. So that's going to really enable to expand the workloads to expand the work being done by enterprises across all of the space.
John Furrier
>> It's interesting, too. You guys have a nice tailwind with the trend that's happening now called specialized intelligence. So you've got general intelligence, AGI, everyone kind of chasing that kind of dream. But we're seeing AGI -like capabilities on the general side. the frontier models, they've crawled the Internet, but they haven't crawled the enterprise or domain-specific data. That's right. look at, say, a Verizon or AT&T or a big telco all over the world. They all have proprietary data. that's literally not unlocked yet. That's sitting there, locked down. So that's new data. They're looking at smaller models that specialize to their workloads. And then obviously use some of the general intelligence when they need to. But for the most part, you're seeing a mix and match of models coming in where this specialized intelligence, look at the success of say Fireworks AI, for instance, they're targeting these native AI developers that don't have the CapEx. So the developer will rely on existing build -out. This is where the enterprise starts to see, okay, I don't need to spend a billion dollars to get full intelligence for my company.
Marc Bolitho
>> That's right. And I think it also enables enterprises to say, I don't have to just work with the smaller, less intelligent models. I can use the largest open -source models. I can fine-tune those now with my data, and I can have those wherever I want, wherever it makes sense for me. on-prem, edge locations, or in the cloud.
John Furrier
>> All right. Positioning now here on Wall Street. Here at NVIDIA, they're on the board all the time. The market's going to open up in 10 minutes here. Everyone loves NVIDIA. They love Broadcom. These are the semiconductor companies. They always use TSMC for making chips. You guys are taped out, which means you have the chips. Yes. But it's not just a chip. It's a system. That's right. Explain where you fit vis-a-vis the competitive map as the analysts want to put you in a box. Okay, where are you? Right. Are you a dog or a cat? What kind of company are you? Because your growth and projections are going to be based on that.
Marc Bolitho
>> Yeah, you think of us as a rack-scalable inference compute provider. So we'll provide to hyperscalers, neoclouds, enterprises, sovereigns, anyone who wants to run AI workloads. Whether they're the smallest workloads to the largest MOEs, we can run those. But within that box, within that full rack, we have our own accelerator. So that's based on a lot of math that we do. So we have our own bespoke accelerator, our own bespoke data flow for running AI. And then it's very important because you can't fit these very large models on just one device. You have to be able to interconnect them all together. And so we utilize an interconnect that runs the backbone of the Internet. And that's our partnership with HPE Juniper Networks. And so we're using a cell-based fabric to connect 72 of our chips together in a near one microsecond latency and an incredibly reliable system. So that's amazing.
John Furrier
>> You can think of us as a system. one area I want to get your thoughts on, because I think this is going to be a growth area for you guys. I'm bullish on the edge. I think AI factories at the edge. I've been saying it for over a year now. It's so obvious. And NVIDIA won't admit it because they're too busy serving the big centralized factories. But when you start getting into the edge, you start looking at distributed computing, as we talked about, sovereignty comes up. You mentioned that word. This is a technical conversation, not just data and privacy. GDPR, during the cloud era is all about. Okay, keep the data in the country, with some geographical bounds, mostly by country. But sovereignty definition is being redefined. You're starting to hear enterprises saying they're taking a sovereign view because their sovereignty is their multinational domain. Now that could also sit in countries. So when you start having intelligence that can look at the data and the networks, sovereignty becomes technical. I'm sure you're having conversations. Expand on that piece because there's reference architectures being discussed, the European Union's looking at some things, the telcos certainly participate in any sovereign conversation because now they can not only keep the data in country or in the geographical boundaries, but they can keep the revenue. So when you start talking money, not IT mechanisms, the sovereign conversation gets elevated to, hey, let's keep all of our money in France if you're in France. Oh, by the way, let's use Mistral. You're starting to see these conversations whether that's a technical thing. That's a networking -based solution. That's intelligence -based. They all have to be there.
Marc Bolitho
>> It is. I think the customers that we talk to also include the major sovereigns as well. They're all looking to set up their own data centers, their own natural language models as well. And I see that as a huge opportunity. Everyone's looking for an alternative choice to run compute. And everyone is very interested in using less energy and also getting tokens at a much lower cost. So we're really excited to be able to provide that.
John Furrier
>> Yeah, and we're going to see a lot of conversations. So I know that HPE and Juniper are having a lot of those conversations, as is everybody. All right, you guys had some milestones. So the tapeout happened in May. That's right. That's a huge milestone. Yes. Now you have other milestones as the CEO, customer validation. That's right. Give us a taste and a feel for some of the conversations you're having with these large deployments, customers. Where is that conversation going? What are the constraints that you guys are attacking?
Marc Bolitho
>> Yeah, when we talk with Neoclouds right now, they're all getting the same equipment in and they're running the same workloads and they're fighting for a very, very thin margin. So there's not a lot of differentiation right now. So a lot of the Neoclouds are really looking for the next technical solution to come in and give them an advantage. Those are the ones that we're talking to today. Those are the ones that are early engagers with us. And if you look at the timeline, we'll have our chips back this fall. We'll be powering on the chip. And since we're using basically an existing interconnect, we'll be bringing that system up very quickly. And we expect to be able to demonstrate that system to all our customers and our partners in our beta systems in Q1.
John Furrier
>> So you're looking at a Q1 '27.
Marc Bolitho
>> Q1 '27.
John Furrier
>> Showing the tech. In action.
Marc Bolitho
>> Showing the benchmarks as well. And then we can move very quickly into production after that due to our relationships.
John Furrier
>> All right, so the question that everyone's going to ask is? Okay, we have a scarcity issue with supply. Yes, yes. How do you address that?
Marc Bolitho
>> Yeah, we have the opportunity, or we have the fortune of having a relationship with Broadcom. So Broadcom does the physical design of our device and has the relationships with TSMC for wafer capacity and also the memory suppliers for HBM. So as long as we're operating in lead times, we're very confident of being able to get that.
John Furrier
>> So Broadcom's your key on the supply side.
Marc Bolitho
>> Broadcom is a key, and it's, I think, a big superpower for a startup to have.
John Furrier
>> Since you brought up Broadcom, a big fan of Broadcom, obviously, and NVIDIA, both companies doing amazing work. Yes. But what makes Broadcom unique, and we talk about this all the time on theCUBE Pod, is NVIDIA is leaning in this way, too. You're starting to see they're changing their story a bit on the ecosystem side. Heterogeneous computing has been the standard. And if you look at all the open source success, continues to be great. Hugging Face just got bought by NVIDIA. That was a huge take off the table. But that was just the model directory. But that points to the validation of open source. Open Compute's coming up. That's gonna be a big show this year. We covered that show when it was an inaugural show in 2013 or 2012, I can't remember which year. But the RAC standards are in place. OCP was a big enabler of the rack scale. Open standards, the role of heterogeneous becomes important. Broadcom, that's one of their major things, saying, hey, we see a multi-vendor. How does that impact some of the buildout? Because at one level, the new clouds are so constrained, they just want to build as fast as they can. Do they even care about heterogeneous? What does that even mean?
Marc Bolitho
>> Yeah, I think if we look today and we talk about heterogeneous compute, it's trying to put different systems together and architectures together to get good pre -fill and good fast decode to work on these agentic models. And my view is that with this generation of compute that's coming out in 27, you don't need to do that. We have a system that can do that homogeneously in one system at a lower cost and lower power. I think going forward, where you're talking in the 29 -30 timeframe, I think there'll be some standards on rack size and such and power. And yeah, we already have our next generation system.
John Furrier
>> As an expert, you've been in the industry for a while. I think the pre -fill, decode, disaggregated serving is a telltale sign because the constraint is networking and memory, right? On dense systems. That's right. When you have disaggregated infrastructure, let's say the edge and say factories move into the enterprise with on -prem, smaller racks, standard racks, you have disaggregated infrastructure and disaggregated serving. This is true. Do you see that as a telltale sign? Do you agree with that view? And if you believe that they will have disaggregated infrastructure, in a way disaggregated serving is the standard.
Marc Bolitho
>> Yeah, I think the disaggregation that we're seeing today is a tell that there's a problem that hasn't been solved yet. And so that's really what we're coming with our system that can do one system, runs it all at the lowest cost and at the lowest power. So I think it's a tell.
John Furrier
>> Yeah, I'm sure you get a lot of orders. Marc, thanks for coming in. Appreciate the work you're doing. We're about to kick off the market in 30 seconds. The bell's going to ring. Perfect timing. We're here at the New York Stock Exchange where everyone's talking about the economics and the stock prices of all the chip companies. And if they have AI on their side, they're certainly in a good position. The market demand and market growth for AI infrastructure, compute, dense systems to serve up the new models of AI, which powers all the intelligence. You have to create it and you have to distribute it. And we're going to see a lot more of this content, this conversation. I'm John Furrier, host of theCUBE. Thanks for watching. Thank you.
>> for compute and for accelerated computing. Marc Bolitho here, the CEO of Tensordyne Company. We've covered coming out of stealth. Marc, great to have you on theCUBE. Good to see you. John, great to be here. Thanks for having me at the Stock Exchange. It was great to cover the unveiling in the Palo Alto through the RK and Gilles in Paris, the RAISE Summit in Paris. But you guys are sitting in the middle of the innovation storm of the computer industry. if you look at the build-out, it is probably the biggest expansion economically, physically, just capabilities. The AI is essentially resetting the architecture from the old compute days. We saw the rise of the computer industry, the Internet revolution, SAS, cloud. But building on that cloud scale is the ongoing demand for compute, networking, and storage in a very dense package because it's mathematics now, right? So math is getting commoditized in a good way because it's spreading everywhere. But feeding the math with more math, that's inference. This is the difference between old school days when you have packets come in, it's processed. You have a whole other paradigm. It's mathematics. You hear about vectors. You hear about reasoning. And the rise of the agents brings up that this isn't like a search paradigm. It's a reasoning intelligence paradigm. That's the demand right now. Now, you guys have got a taped-out chip with some DNA from Juniper, which we covered. You guys got a different approach. Explain the approach of Tensordyne. You guys had four years of stealth. Probably you were itching to talk more about it, but now it's out there.
Marc Bolitho
>> Yeah, John, absolutely. if you look at this market, this is the largest market. We're going to approach a trillion -dollar TAM in inference alone. And right now, everything is being built upon these agentic workloads. So Anthropic plus OpenAI, they have a $100 billion ARR right now together. All that's being built because of agents. It's finally AI is doing real work for us. So whether it's coding or other workloads, it's useful and it's getting real jobs done. But these models are burning a ton of internal tokens. We're talking orders of magnitude more than just a chat model now to run agents. And so when you do that, you have a need for better speed. You have to have better speed when you're running these agentic workloads because you're not just communicating with a human anymore. Now you're doing tool calls, communicating with the next agent, and you've got to be able to process this very quickly. So what we did is early on when ChatGPT first came out, we were looking at our log math, our architecture we had for doing inference. And what we realized very quickly is that this architecture in log math was perfect for generative AI inference. And the reason for that is fundamentally when you're doing math in the log domain, you can do all of the trillions of multiplies in your matrix algebra. You can do all of those in a much smaller area of compute. And when you have that smaller area of compute on the chip, you can exchange that. You can exchange that for memory, exchange that for data flow and for networking, and you can create a very balanced system. So that's what we're doing, John, differently than anybody else.
John Furrier
>> It's how we're approaching the math. The log math piece. Let's explain that because I did a post last week on LinkedIn around. I get all the questions. I just had to plug in all my notes into Claude and it gave me a great one-pager to explain everything. You got the density of NVIDIA, and they got software to manage that. You see things like. I was learning about systolic arrays at the Hot Chips conference. And that's Google's approach. It's going through memory to multiple arrays to get that multi-step process. You got Cerebras. You got a variety of different approaches. It's all math. Yes. So mathematics is the key. And the floating point piece of the processor, which became the GPU for gaming and graphics, became the state of the art for AI. Talk about that math piece. Why is the math so important in this new equation? Because that's the difference in the AI generation versus the old compute days that we all kind of grew up with PCs and servers. Now it's different. You've got the mathematics piece.
Marc Bolitho
>> Yeah, so everyone else is working in floating point. Our thought early on was, if you're going to come up with a new solution for running AI, how do you differentiate? You've got to go to first principles. You have to focus on the math. So fundamentally AI is trillions of multiplications followed by trillions of additions. And so what we do is we work in the log domain so log of A times B is equivalent to log A plus log B which means we can now change every multiply into an add. When you do that you have a much smaller area on the die for doing all those multiplications and you have lower power and that lower power scales from the chip to the pod to the rack all the way to the data center. So that's fundamental.
John Furrier
>> Talk about what you guys hope to change in the industry, because you brought up a couple of key things that are constrained right now, power and speed. Everyone wants faster tokens per second per watt. Yes, that's fact. And now everyone's starting to see the benefits on the revenue side, tokens equal revenue. That's kind of well understood now. Why is it important to have that speed? And then what do you guys hope to change in the industry?
Marc Bolitho
>> Yeah, it's very important, as we said, to have that speed. today you see that to get into, say, the 800 to 1,000 output tokens per second per user. Companies are looking at disaggregated systems, so a Rubin for pre -fill, multiple Groq units, racks, say, eight Groq racks for the decode side, looking at, say, AWS doing the pre -fill with Trainium and multiple racks of Cerebras to do the decode. So what our architecture allows us to do is that because our compute die is so small, we can put more SRAM on the chip. And that means we have a highly utilized compute engine. So if you have a compute engine and you're only using 10 % of it, that's not very valuable. You have to be able to have an architecture that gets the most out of the silicon area that you have. And we do that with log math. We have an abundance of SRAM, as much SRAM as say a Groq or a Cerebras, but we also have high bandwidth memory. So we don't need multiple racks to run these 2 trillion parameter models. We can basically, in a nutshell, instead of having nine to 14 racks to run a 2 trillion parameter agentic model, we can run it in one rack. And we can run it at an order of magnitude less cost.
John Furrier
>> How does that change economics? Because you're accelerating the computing, but at lower cost and lower power.
Marc Bolitho
>> That's right. So, if you look at it today, to get these higher output speeds, the costs rise dramatically. And so you have a trade -off you have to make between intelligence and cost. And we need to get into an era that you can pick the best intelligence, you can pick the best model, and you can pick to have your responses come back very quickly and not sacrifice the cost. And so what we're seeing is in those higher output speeds, we have a 10x cost advantage.
John Furrier
>> Talk about the architecture compared to NVIDIA, because people say they hear NVIDIA, accelerated computing. That's their big narrative. What do you guys hope to define? If NVIDIA is accelerated computing, what's Tensordyne?
Marc Bolitho
>> Yeah, so we're definitely doing accelerated computing as well for AI. But think about it as a system that's future proof for the largest AI workloads and at the highest user speeds, at the lowest cost and the lowest energy. We haven't even mentioned energy yet because energy is very important as we look to scale out AI as well.
John Furrier
>> I mean, all the Neo clouds out there we're talking to, energy is the number one thing and finance. Those are the two bounding functions. Money, which we're seeing a lot of CapEx, great for now, but energy is going to be an ongoing sticky problem, especially when you look at connecting AI factories at, say, the edge or metro areas. We're going to see that progression very quickly, I think, from centralized factories to distributed computing. You mentioned the prefill-decode. Disaggregated serving is the hottest topic right now. Because the KV caches are getting so filled. Because there's now too much data coming in. When you have disaggregated serving, you're going to have disaggregated infrastructure. Yes. So you're going to have racks everywhere. Yes. Talk about that piece, because that's a huge determinant on the power side. You can't drop an NVL72 at a telco tower. That's right. Or maybe even an enterprise.
Marc Bolitho
>> Yeah, and so there's different ranges, absolutely. And so if you look at going from an enterprise or from a telco all the way up to a large data center, to do disaggregated inference today for these large MOEs, agentic models, 2 trillion parameter plus. Companies are deploying, say, nine racks of compute to do that. And to get those high user speeds, companies have said that's going to cost, say, $150 per million tokens and be 1.5 megawatts of power. We can do that in 120 kilowatts. So we can do that at a fraction.
John Furrier
>> What's the rack requirements? You mentioned 120 kilowatts. What's the baseline?
Marc Bolitho
>> So our pod requirement where we have 72 accelerators, think of 72 accelerators like a Blackwell also having 72 accelerators. We do that in a quarter of a rack with 72 accelerators at 30 kilowatts, and it's air-cooled. So you can fit into any existing infrastructure or data center. You don't need a liquid-cooled data center. You can work with stranded power or existing data centers. So it works for enterprises. It works for telcos. it works all the way up the chain.
John Furrier
>> We're seeing a lot of demand for, again, I mentioned distributed computing, but distributed intelligence. So, for example, if I'm an enterprise, I might have a little data center. Sure. Not little, but enough power to run in the classic sense of traditional data center. But in the modern era, it might not have the requirements for the monster systems. Okay, but I'm going to connect to, say, CoreWeave or Nscale, even AWS. I'm going to run an on -prem hybrid computing platform, so I'm going to be distributed. But I can have single -tenant -like experience on these Neo clouds and AI clouds, as well as the hyperscalers like AWS. We're already seeing AWS with VPCs handling that. So as this becomes distributed, you guys fit nicely. So explain where you fit, because if I'm an enterprise or I'm a telco, I'm not going to have a gigawatt to play with. That's right. I'm going to have a beefy edge or a data center that's kind of in scope, in a traditional sense.
Marc Bolitho
>> So think of being able to run multiple instances of a Kimi K3 model with hundreds, maybe a thousand concurrent users in a quarter rack at 30 kilowatts. It can be all on -prem and you can have your own fine tuning of your models, all your data stays on -prem and for larger workloads, you can either scale up to the extent of your on -prem or then, as you said, have distributed back to the cloud.
John Furrier
>> And what does that mean for say, agentic? Because a lot of the, whether it's a hospital or financial services or say, in retail, they don't have big data centers, but they want intelligence everywhere.
Marc Bolitho
>> Yes, what that means is that they can run agentic workloads in their own data centers. That's what we enable.
John Furrier
>> What's the token cost? Because again, a lot of people have been token maxing. Sure. I was talking to someone at Stripe. They're like, yeah, the token days are such a headache. It's like running a test dev environment on our production payment system. Yeah. people are using production -level tokens to essentially do stuff. That's essentially what's happening today. So, okay, the next step would be, okay, logical IT solution would be get me a box on -prem and just have unlimited development tokens and then scope it, then move it to production. Maybe it's a hybrid, Pareto curves might be different. But I don't want to blow token budgets. How is that factored into some of the conversations you have with large deployments? because that's a cost structure.
Marc Bolitho
>> And I think that's one of the things that have been preventing that build out and the scale out at the enterprise level is the cost of breaking into it. And so there has to be that cost -benefit analysis that's done for running these workloads. What we're seeing is with these 2 trillion parameter agentic models that we're going to be able to do this at an order of magnitude less cost than what the systems are that are going to be launching now in this generation. So that's going to really enable to expand the workloads to expand the work being done by enterprises across all of the space.
John Furrier
>> It's interesting, too. You guys have a nice tailwind with the trend that's happening now called specialized intelligence. So you've got general intelligence, AGI, everyone kind of chasing that kind of dream. But we're seeing AGI -like capabilities on the general side. the frontier models, they've crawled the Internet, but they haven't crawled the enterprise or domain-specific data. That's right. look at, say, a Verizon or AT&T or a big telco all over the world. They all have proprietary data. that's literally not unlocked yet. That's sitting there, locked down. So that's new data. They're looking at smaller models that specialize to their workloads. And then obviously use some of the general intelligence when they need to. But for the most part, you're seeing a mix and match of models coming in where this specialized intelligence, look at the success of say Fireworks AI, for instance, they're targeting these native AI developers that don't have the CapEx. So the developer will rely on existing build -out. This is where the enterprise starts to see, okay, I don't need to spend a billion dollars to get full intelligence for my company.
Marc Bolitho
>> That's right. And I think it also enables enterprises to say, I don't have to just work with the smaller, less intelligent models. I can use the largest open -source models. I can fine-tune those now with my data, and I can have those wherever I want, wherever it makes sense for me. on-prem, edge locations, or in the cloud.
John Furrier
>> All right. Positioning now here on Wall Street. Here at NVIDIA, they're on the board all the time. The market's going to open up in 10 minutes here. Everyone loves NVIDIA. They love Broadcom. These are the semiconductor companies. They always use TSMC for making chips. You guys are taped out, which means you have the chips. Yes. But it's not just a chip. It's a system. That's right. Explain where you fit vis-a-vis the competitive map as the analysts want to put you in a box. Okay, where are you? Right. Are you a dog or a cat? What kind of company are you? Because your growth and projections are going to be based on that.
Marc Bolitho
>> Yeah, you think of us as a rack-scalable inference compute provider. So we'll provide to hyperscalers, neoclouds, enterprises, sovereigns, anyone who wants to run AI workloads. Whether they're the smallest workloads to the largest MOEs, we can run those. But within that box, within that full rack, we have our own accelerator. So that's based on a lot of math that we do. So we have our own bespoke accelerator, our own bespoke data flow for running AI. And then it's very important because you can't fit these very large models on just one device. You have to be able to interconnect them all together. And so we utilize an interconnect that runs the backbone of the Internet. And that's our partnership with HPE Juniper Networks. And so we're using a cell-based fabric to connect 72 of our chips together in a near one microsecond latency and an incredibly reliable system. So that's amazing.
John Furrier
>> You can think of us as a system. one area I want to get your thoughts on, because I think this is going to be a growth area for you guys. I'm bullish on the edge. I think AI factories at the edge. I've been saying it for over a year now. It's so obvious. And NVIDIA won't admit it because they're too busy serving the big centralized factories. But when you start getting into the edge, you start looking at distributed computing, as we talked about, sovereignty comes up. You mentioned that word. This is a technical conversation, not just data and privacy. GDPR, during the cloud era is all about. Okay, keep the data in the country, with some geographical bounds, mostly by country. But sovereignty definition is being redefined. You're starting to hear enterprises saying they're taking a sovereign view because their sovereignty is their multinational domain. Now that could also sit in countries. So when you start having intelligence that can look at the data and the networks, sovereignty becomes technical. I'm sure you're having conversations. Expand on that piece because there's reference architectures being discussed, the European Union's looking at some things, the telcos certainly participate in any sovereign conversation because now they can not only keep the data in country or in the geographical boundaries, but they can keep the revenue. So when you start talking money, not IT mechanisms, the sovereign conversation gets elevated to, hey, let's keep all of our money in France if you're in France. Oh, by the way, let's use Mistral. You're starting to see these conversations whether that's a technical thing. That's a networking -based solution. That's intelligence -based. They all have to be there.
Marc Bolitho
>> It is. I think the customers that we talk to also include the major sovereigns as well. They're all looking to set up their own data centers, their own natural language models as well. And I see that as a huge opportunity. Everyone's looking for an alternative choice to run compute. And everyone is very interested in using less energy and also getting tokens at a much lower cost. So we're really excited to be able to provide that.
John Furrier
>> Yeah, and we're going to see a lot of conversations. So I know that HPE and Juniper are having a lot of those conversations, as is everybody. All right, you guys had some milestones. So the tapeout happened in May. That's right. That's a huge milestone. Yes. Now you have other milestones as the CEO, customer validation. That's right. Give us a taste and a feel for some of the conversations you're having with these large deployments, customers. Where is that conversation going? What are the constraints that you guys are attacking?
Marc Bolitho
>> Yeah, when we talk with Neoclouds right now, they're all getting the same equipment in and they're running the same workloads and they're fighting for a very, very thin margin. So there's not a lot of differentiation right now. So a lot of the Neoclouds are really looking for the next technical solution to come in and give them an advantage. Those are the ones that we're talking to today. Those are the ones that are early engagers with us. And if you look at the timeline, we'll have our chips back this fall. We'll be powering on the chip. And since we're using basically an existing interconnect, we'll be bringing that system up very quickly. And we expect to be able to demonstrate that system to all our customers and our partners in our beta systems in Q1.
John Furrier
>> So you're looking at a Q1 '27.
Marc Bolitho
>> Q1 '27.
John Furrier
>> Showing the tech. In action.
Marc Bolitho
>> Showing the benchmarks as well. And then we can move very quickly into production after that due to our relationships.
John Furrier
>> All right, so the question that everyone's going to ask is? Okay, we have a scarcity issue with supply. Yes, yes. How do you address that?
Marc Bolitho
>> Yeah, we have the opportunity, or we have the fortune of having a relationship with Broadcom. So Broadcom does the physical design of our device and has the relationships with TSMC for wafer capacity and also the memory suppliers for HBM. So as long as we're operating in lead times, we're very confident of being able to get that.
John Furrier
>> So Broadcom's your key on the supply side.
Marc Bolitho
>> Broadcom is a key, and it's, I think, a big superpower for a startup to have.
John Furrier
>> Since you brought up Broadcom, a big fan of Broadcom, obviously, and NVIDIA, both companies doing amazing work. Yes. But what makes Broadcom unique, and we talk about this all the time on theCUBE Pod, is NVIDIA is leaning in this way, too. You're starting to see they're changing their story a bit on the ecosystem side. Heterogeneous computing has been the standard. And if you look at all the open source success, continues to be great. Hugging Face just got bought by NVIDIA. That was a huge take off the table. But that was just the model directory. But that points to the validation of open source. Open Compute's coming up. That's gonna be a big show this year. We covered that show when it was an inaugural show in 2013 or 2012, I can't remember which year. But the RAC standards are in place. OCP was a big enabler of the rack scale. Open standards, the role of heterogeneous becomes important. Broadcom, that's one of their major things, saying, hey, we see a multi-vendor. How does that impact some of the buildout? Because at one level, the new clouds are so constrained, they just want to build as fast as they can. Do they even care about heterogeneous? What does that even mean?
Marc Bolitho
>> Yeah, I think if we look today and we talk about heterogeneous compute, it's trying to put different systems together and architectures together to get good pre -fill and good fast decode to work on these agentic models. And my view is that with this generation of compute that's coming out in 27, you don't need to do that. We have a system that can do that homogeneously in one system at a lower cost and lower power. I think going forward, where you're talking in the 29 -30 timeframe, I think there'll be some standards on rack size and such and power. And yeah, we already have our next generation system.
John Furrier
>> As an expert, you've been in the industry for a while. I think the pre -fill, decode, disaggregated serving is a telltale sign because the constraint is networking and memory, right? On dense systems. That's right. When you have disaggregated infrastructure, let's say the edge and say factories move into the enterprise with on -prem, smaller racks, standard racks, you have disaggregated infrastructure and disaggregated serving. This is true. Do you see that as a telltale sign? Do you agree with that view? And if you believe that they will have disaggregated infrastructure, in a way disaggregated serving is the standard.
Marc Bolitho
>> Yeah, I think the disaggregation that we're seeing today is a tell that there's a problem that hasn't been solved yet. And so that's really what we're coming with our system that can do one system, runs it all at the lowest cost and at the lowest power. So I think it's a tell.
John Furrier
>> Yeah, I'm sure you get a lot of orders. Marc, thanks for coming in. Appreciate the work you're doing. We're about to kick off the market in 30 seconds. The bell's going to ring. Perfect timing. We're here at the New York Stock Exchange where everyone's talking about the economics and the stock prices of all the chip companies. And if they have AI on their side, they're certainly in a good position. The market demand and market growth for AI infrastructure, compute, dense systems to serve up the new models of AI, which powers all the intelligence. You have to create it and you have to distribute it. And we're going to see a lot more of this content, this conversation. I'm John Furrier, host of theCUBE. Thanks for watching. Thank you.