This discussion examines artificial intelligence factories and the evolving inference infrastructure for data centers. Darren Chien of Positron AI appears on theCUBE Research interview hosted by Furrier of theCUBE and Allen of theCUBE. Chien offers Positron AI's perspective on specialized silicon, memory-centric decode workloads and data center deployment. They address token economics, disaggregation of prefill versus decode, Positron's Asimov roadmap and subsequent generations, recent funding validation and how rack-scale systems and software compatibility enable AI factories.
Chien states inference is heterogeneous and not a one-chip market. Positron AI focuses on decode workloads to improve tokens per dollar and tokens per watt. Analysts of theCUBE highlight the importance of disaggregation, sovereign cloud and latency considerations and the need for Hugging Face compatible software stacks and fast turn-on rack solutions to accelerate deployment.
Subscribe for ongoing coverage of AI infrastructure, inference strategies and data center innovation.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Darren Chien, Positron AI
This discussion examines artificial intelligence factories and the evolving inference infrastructure for data centers. Darren Chien of Positron AI appears on theCUBE Research interview hosted by Furrier of theCUBE and Allen of theCUBE. Chien offers Positron AI's perspective on specialized silicon, memory-centric decode workloads and data center deployment. They address token economics, disaggregation of prefill versus decode, Positron's Asimov roadmap and subsequent generations, recent funding validation and how rack-scale systems and software compatibility enable AI factories.
Chien states inference is heterogeneous and not a one-chip market. Positron AI focuses on decode workloads to improve tokens per dollar and tokens per watt. Analysts of theCUBE highlight the importance of disaggregation, sovereign cloud and latency considerations and the need for Hugging Face compatible software stacks and fast turn-on rack solutions to accelerate deployment.
Subscribe for ongoing coverage of AI infrastructure, inference strategies and data center innovation.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, co-hosting theCUBE here with Gemma Allen, my co-host. Hello, I'm John Furrier, host of theCUBE. We are here at theCUBE's NYC studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is our AI Factory series. We've talked to the leaders who are making it happen. Of course, we've got an event tonight with a lot of investors in the hottest companies. Swan Company will be having some remarks at this meetup here in New York City. Again, we got Darren back. He was just back on theCUBE in our, the NYSE Wired program in July. He's the managing director of APAC for Positron, with big news just recently, big-time funding at a $5 billion valuation. It was like $575 million. Darren, great to see you again.
Darren Chien
>> Great to see you.
John Furrier
>> You got a spring in your step. You got some fresh funding, great validation on the valuation. Congratulations. Some great board members. Really stacking the team up since we last talked. Of course, it was probably in the works. You couldn't say anything. Now it's all out there.
Darren Chien
>> Now it's public. And yeah, great news. Like you mentioned, great validation. Fundraising is just a tool in our tool belt to really get to where we need to go. And so really happy to be here. Palo Alto was great. This is a different, different vibe.
John Furrier
>> Yeah. You know, I remember a year and a half ago, Mitesh and the team grinding away and the market continues to grow. The AI infrastructure If you look at what's happened just on the supply chain, obviously HBM memory, KV cache is exploding. Disaggregated serving is standard pretty much now in the clusters. You get different approaches, but the demand is just off the charts. Speak to that dynamic because it's just not stopping.
Darren Chien
>> You already saw the level of the demand a year and a half ago. I don't even know how much bigger it is now and how much bigger it is going to be. Even within our own specialized niche industry of chips, right? You used to think it was one chip to rule them all, maybe NVIDIA, maybe then it's NVIDIA and AMD. Now we have NVIDIA, AMD, Groq, Cerebras, all the frontier labs working on their own chip efforts. You have the startups like us trying to make a dent. The beautiful thing about it is that the market is so large that there is a wedge for all of us if we bring something differentiated to the table, bringing value to the customers, bringing scale to the customers. It's a great environment to be in.
John Furrier
>> I mean, if you look back 3 years ago, Dave and I would probably have commented on our CUBE Pod, it's a winner-take-all. NVIDIA is looking good, middle of the fairway. You got the hyperscalers buying up everything and then the rise of the Neo Cloud comes in. Okay, great. Still potentially a winner-take-all. And then I think our commentary shifted to rising tide lifts all boats. To use that example, that's kind of what happened. And if you look at— I want to get your thoughts on this because there's some data points and I like to get your analysis on this, because if you look at just the overall growth of inference, this aggregate survey I mentioned, these are signals that it's not a one-chip game. In fact, NVIDIA, they don't want to be called a GPU company. They're an AI infrastructure company. Now you've got footprint diversity coming into the equation with AI factories. Yeah, those monster factories pumping out tokens. Tokens equals revenue. I totally buy all that. But it's not just the one monster system. You guys see obviously different models. Smaller models can run on a lot less compute. agents are compute hungry. Yeah, they'll use GPUs. They'll use the Mellanox. Talk about that piece of the dynamics because I think this idea of networking with the KV cache Dynamo from NVIDIA speaks to an ongoing trend that you're going to have a lot of different systems that will perform high-performance token generation. But it's not the same thing. No, the design will be different. This is where the chip diversity comes in. You guys play into that, right?
Darren Chien
>> 100%. And listen, I spent 2 years at Lambda. When I started at Lambda, if you told me I wanted to use anything other than NVIDIA, I would have laughed you out of the room. I said, hey, Talk to someone else or give me what you're having. Now, 2 years later, we are in a space where there is demand for Cerebras, there is demand for Groq, there is demand for AMD. What that is telling us is that, yes, inference equals revenue, but not all inference is equal. You have premium inference. These are the tokens that you want the fastest, you want the smartest, you cannot fail. You are willing to pay 10 times more for maybe 5 times more serving speed. Then there's also the 80% of the inference, which is your code generation, your video generation. You can afford to wait. Wait a little bit, you really care about economics. And that's kind of the wedge that Positron AI plays in. Cerebras, fantastic technology, wafer-scale engine, my goodness, a technological innovation never before seen. But they serve really fast but really expensive tokens, right? We are focused on the workloads, specifically the decode workloads that require massive memory capacity, trillion parameter plus models, multiple trillion parameter plus models, and just ultra-high memory bandwidth utilization at a really great cost.
John Furrier
>> Yeah, talk about the economics, because I think when we, if you look at the history, let's take the computer industry, the PC industry, you had a 286 processor, 386 processor, each one had kind of like an entry-level, mid-range, flagship, the classic product positioning. Okay, you've kind of seen that on Pareto curves. You say, okay, to your point, yeah, I can wait a few seconds. Milliseconds, but some things require the best— robotics, autonomous vehicles, real workload, financial transactions. Talk about what that means from a chip standpoint. You mentioned Decode. Explain the nuances of this inference market because you're starting to get into product positioning of token generation. That's what's happening.
Darren Chien
>> Exactly. And so we talked about this in our last talk, prefill and decode, two workloads within inference that are kind of served by different processing powers. prefill is still compute heavy, still parallel, and that's when you upload your file, you input your chat into ChatGPT, and it can process that in parallel. Then you have decode, which is still memory bandwidth bound, right? You still need massive memory bandwidth utilization, massive memory capacity to fit the models to fit the KV caches. And so when we're looking at the workloads that Positron AI is positioned for, high-frequency trading is actually a use case that fits great on our hardware. We have Jump Trading who co-led our Series B. We have HRT who is getting involved on our Series C. That should tell you something. And then across the entire set of workloads, you mentioned Pareto curves, right? We're really pushing the Pareto curves of the optimal point matching throughput per dollar and interactivity up and to the right as it compares to some of the alternatives in the market.
John Furrier
>> So you guys are positioning Positron for the high-end piece of the market?
Darren Chien
>> Not necessarily at the high end, but the two metrics that we will win on are tokens per dollar and tokens per watt. So making it economical to generate these tokens and energy efficient to generate these tokens across all use cases and price performance layers. Correct. Within inference.
John Furrier
>> Inference.
Darren Chien
>> Okay, got it.
John Furrier
>> Okay. So I had to look it up on the internet here because I didn't know the exact date, but the, NVIDIA Groq deal was around December 2025. They were probably in the works a couple of months before. But let's just say that the fall/winter of 2025, that's literally 9 months ago, 10 months ago.
Darren Chien
>> Okay. not years.
John Furrier
>> So think about— I want you to share your thoughts on the importance of that deal because since then it went from NVIDIA to we can have a separate thing for inference. Look at the market since then. Explain what happened, because that was kind of like a very much a shot heard around the world in the chip game.
Darren Chien
>> Correct.
John Furrier
>> Because it opens up. Okay.
Darren Chien
>> Wow.
John Furrier
>> Again, back to your point about the wedge and room for everyone to play, that kind of was the template for everything we're seeing now. And agents and code generation, they're coming in, they have different needs. Compute becomes more important with agents. You can then free up those GPUs that are being used for things like that. Talk about that importance, because that Groq deal really was the moment the industry seemed, at least from my standpoint, to shift.
Darren Chien
>> Yeah. So to be clear, NVIDIA is the largest company in the world. It's one of the smartest companies in the world for a reason. As new chip startups or as anyone trying to play in this AI game at all, we have to understand where we fit in this dynamic. You brought up disaggregation, splitting prefill and decode. That is exactly a scenario in which we could look to NVIDIA and say, hey, listen, your chips are the best at compute-bound workloads. Training, unquestionable, you will use NVIDIA. But when we look at inference, let's see, perhaps we pair an NVIDIA Vera Rubin rack with a rack of Titan, and let's see, maybe we shift some of the prefill work to the accelerator that is more compute-heavy, and let's shift some of the decode work, where we can get more economical. we might not get as fast as an NVIDIA Rubin on a memory bandwidth perspective, but a memory bandwidth utilization and efficiency and capacity perspective. Let's pair with the NVIDIA ecosystem and see how we can make, 1+ 1= 3. That's the dream of any chip startup in today's environment.
John Furrier
>> And by the way, from a system standpoint, that makes total sense because NVIDIA has to make a choice. Do I want to bogart the whole thing and slow down progress? And you see Jensen, they're okay with this. It's not like they're against any of this. They're actually promoting the notion of more AI for everyone. In fact, they're making the market on the policy side. They're making the market on the ecosystem side. And they're actually saying bring more stuff to the table because they'll win too. There's supply constraints, but also capability.
Darren Chien
>> 100%. Listen, we are in the footsteps of giants here, right? NVIDIA laid the path. We would not be here if not for NVIDIA. But as you mentioned, they are making moves to partner with startups like ours. Another company that is going to be on this panel is going to be at the event later tonight, d-Matrix. They had some huge news of being, I think, integrated into the NVLink ecosystem with NVIDIA. So it's going to show that NVIDIA is not trying to do anything that is aggressive or do anything that's harmful for the ecosystem. They understand that the market is so large that we at Positron AI just need to come up with a proposal or a way of working together that makes sense for all of us.
John Furrier
>> Yeah, rising tide lifts all boats. There's plenty of beachfront for everybody to play and camp out and create innovation. Okay, talk about the funding and why that was driving so hard because again, the validation, $5 billion, you only raised what, $600 million, $575, what was the
Darren Chien
>> amount?
Darren Chien
>> total round was $875 million.
John Furrier
>> $875 million, so a little bit under a billion. You could have done better. Mitesh, come on, get it up to a billion, get across the threshold. Only kidding, obviously, that's plenty of capital. But the validation on the valuation is for a reason. What's the reason why? Why is the momentum so strong with Positron?
Darren Chien
>> Listen, it's a couple of things. One, our Series B fundraising, we raised around $230 million untranched. So in the bank, you're talking about the billion. I can see our bank, I see that billion number. So we feel pretty good about that. Two, we are actually shipping product today. Our first-gen Atlas system is deployed in a hyperscaler environment in OCI. They've been great partners for us. We've been able to demonstrate real traction, real shipping expertise, which is always the question mark when, the investors who participated in our round, they're the most sophisticated people. On the planet Earth. They looked at our team, they looked at our current product deployments. What really got them excited is our future products. So our Asimov silicon taping out this year, we expect to have chips back, sometime next year and we'll be shipping them in Titan servers, both air-cooled, liquid-cooled SKUs depending on what the customer wants end of next year. And so we are really excited for that to come out because that is going to be really the point in time in which you will see Positron AI is making a dent in the market.
John Furrier
>> Okay, so we know what's happening. Inference is the killer app. The systems are becoming more expansive on the capability side. We know what it means. There's demand that needs to be served and you guys got the products. Connect the dots on what's next. What is that vision for the future products? Without giving away any trade secrets, take us through what's the next hill you guys are going to climb?
Darren Chien
>> Listen, fundraising is good. It's not the end-all be-all. It's not the goal. The goal is more chips in the world serving tokens that everyone can access, right? The chips are a way to just democratize intelligence, which the unit of tracking that is tokens. But what we're really looking to use this fundraise and any future fundraisers for is to continue to iterate on our custom silicon. We're starting with Asimov. The next generation is Bradbury. The generation after that is Clarke. And I'll let you kind of guess what the following generation is.
John Furrier
>> You guys definitely got me thinking. Okay, so now let's get back to use cases. What's the top use case that you guys are selling to now and what is developing?
Darren Chien
>> So the main use cases that we see are code generation and video generation. Those are currently the new gen— the new use cases. Myself personally, the agentic use case is the most compelling one. I don't know if you've seen the latest news around this thing called Manus. My goodness, my flight here was booked using Manus. I didn't have to interact with the United app. I had a middle seat. I asked, hey, please change me to an aisle seat. It changed me to an aisle seat. These agentic use cases we've been hearing about for years are actually here. You and I, real people, not inside Fortune 500 companies, can use these agentic use cases today.
John Furrier
>> I was one of those people. I was in a room about a couple of years ago, and the question was, when will agents book flights for you? And I was one of the skeptics. I always admit when I'm right, obviously, and admit when I'm wrong. I said, I don't think it's gonna happen. And mainly because I didn't think it would be advanced enough to cross some of these transactional rails. And this is what's getting interesting. They're so smart now, they can actually do it. I thought you'd see some great automation, basically RPA on steroids to begin with, and then I would see reasoning. But I didn't, I said 5 years, maybe. It was 3. So now we have change the flight. What's going on in my PG&E bill? It's overdue, payment's due today. So a lot's involved. So trust is a huge factor in all these agents. So I was wrong on that one, but hey, better to have it now because Anthropic is getting access to all my stuff. And a little bit scary, but the upside is pretty damn good.
Darren Chien
>> I have my passport in there, my driver's license, my credit card. Luckily, there's some strong integrations. Anthropic is working with 1Password. Meta has taken some strong security and privacy precautions but there is risk as with everything. But the upside, like you mentioned, is just incredible.
John Furrier
>> Which one do you like better?
Darren Chien
>> Oh, Muse is a really fast, really delightful experience. But there's something about texting your agent on iMessage where I'm already living that is— Yeah, yeah. That
John Furrier
>> doesn't—>> And they got WhatsApp too. They got WhatsApp and iMessage.
Darren Chien
>> I'm an iMessage guy. we had Europeans in this morning. They use WhatsApp. WhatsApp, but iMessage for
John Furrier
>> us.>> I know, I am completely schizophrenic between the different— and of course you have the random Android user in there. Yeah, let's not talk about
Darren Chien
>> that.
Darren Chien
>> Let's not talk about those.
John Furrier
>> All right, so what are you most excited about right now? Obviously it's a great time for you guys. Congratulations to the team. We're big fans. Obviously we were following you guys since the beginning. well deserved. We need more chips. What's the most exciting thing going on now?
Darren Chien
>> Listen, I think the external validation— big fundraise number, big valuation number. Great board members. I could talk about our board members for this entire talk. The external validation is good, but unlike some of the other startups that exist, we really pride ourselves on internal validation. We want to make sure that we are taping out and shipping a chip and system that actually delivers on performance claims such that it drives value to our customers.
John Furrier
>> Darren, I love that attitude and I really appreciate that because one of the things I was very critical of during the last wave of tech innovation. Fundraising was a celebratory milestone, like, we made it. In this new era, the pragmatic view is, look, we just want gas in the tank. We're not really focused on that as a— it's a milestone, mile marker, moment in time. Yeah, you said it's like winning a game. The coaches celebrate for the night, but tomorrow we'll get back to work. That's an attitude change that I've seen in this new generation because I think it's mainly because, one, the expectation for the market is so strong that the reality is sinking in like, hey, we got $1 billion in the bank, let's have a big party, let's throw the— let's get some cool stuff for the office. Moves to no, no, we got more work to do. That's different. What's your take on that? And what's your— do you agree with that? Do you think that's a general sentiment or is there still kind of like, hey, we got the money in the bank, let's go to Tahoe and we could do an offsite meeting. But I'm saying, you know what I'm saying? I'm not stopping.
Darren Chien
>> No, no, no. The only thing I disagree with you on is that it is a generational change. No, I think we're pretty unique in this in that we don't view this, obviously we celebrate. It's great news. we're going to go out to dinner. But like you mentioned, it's not the tool. It is not the goal. It is a tool in the toolbox. Right. And I credit this to Mitesh and Thomas, CEO and CTO of Positron AI. they've seen this before. They've been building startups for the past decade. Right. The money that you get from investors is probably some of the most expensive dollars you'll ever take, right? They have expectations. You must deliver. And so if anything, it is a weight on our shoulders to make sure that we deliver.
John Furrier
>> you got to deliver.
Darren Chien
>> It's—
John Furrier
>> there's expectations.
Darren Chien
>> There's expectations.
John Furrier
>> Here's the gas in the tank. Get to your destination.
Darren Chien
>> The beautiful thing is with such a large fundraise, now we can get creative. How do we accelerate? How do we de-risk? Do we spend $15 million more on emulation hardware? Maybe we should if that gets us, 10% of the way there. we're obviously going to be expanding the team, new office in Austin, expanding our facilities in Spokane, doing some international expansion that I'm very excited about. Yeah, but the money is a tool that we're going to use.
John Furrier
>> It's a totally great attitude. Let's talk about the board members. You mentioned some of them. We checked them out. You got Gavin, you got Dylan Patel from SemiAnalysis, although he's got an analyst firm, so maybe he's got some inside baseball there. You got Jim Clark from Netscape, founder of Netscape. Talk about some of the board members, the new board members, and how has that changed the dynamic?
Darren Chien
>> So Gavin has been involved with the company for a while, but we're bringing him onto the board as part of this round. Gavin— no, no, who, who— I cannot give a better introduction to Gavin than he can himself, on his Twitter, on his podcast. So obviously his Twitter is huge. He's become a leading voice in AI. So having him join the board, we're really excited to draw on his expertise, see what he sees in the market. Same thing with Dylan. SemiAnalysis, SemiAnalysis Capital, you know, participated in the round. He sees everything right across the board, across the chip, but even down the layers. Right. Because when we're in the chip space, we need all of the inputs, the substrates, the assembly, the packaging, the wafer, the memory. And so having someone like Dylan who has a great insight into the entire market, but actually what I'm most, most excited about is Jim Clark. So you mentioned Netscape, co-founded with Marc Andreessen. The other company he was involved with, SGI, Silicon Graphics.
John Furrier
>> Yeah. Yeah.
Darren Chien
>> What better name can you have on the board of a silicon company, right?
John Furrier
>> Well, SiliconANGLE was another great, great name. Silicon Graphics, because silicon's back.
Darren Chien
>> Silicon is back.
John Furrier
>> And talk about the software stack, cuz one of the things that's emerging outta this is that layer you mentioned, code generation and GenTech. You're now starting to see a new kind of software stack. It's not like the old days, hey, buy a server and load Linux and run your stuff and put some virtualization on it. A new software layer, because as you pointed out, The new chips are creating new subsystems around this holistic system.
Darren Chien
>> Correct.
John Furrier
>> That's powering all the AI. Okay, great. We agree on that. But the software stack now has to adapt.
Darren Chien
>> Correct.
John Furrier
>> What's your vision on that? How does that play out?
Darren Chien
>> So we know that the customers who are buying Positron AI systems are plugging them into GPUs, TPUs, AMDs, Trainium. They don't want to have to deal with another compiler, another language, another runtime. Right. And so the vision for Positron AI and what we're trying to develop on the software side is basically Hey, Hugging Face Transformers library compatible, download them onto the system. We will spit out an API. It'll plug right into your existing fleet. Obviously a lot of work to be done there. But with this new agentic development that obviously the software engineers will be more familiar with is we can download, ingest, and run models in a matter of hours. I remember in Lambda's day, it was a team of engineers 24/7 round the clock for a week before we could get something up and running on Hopper systems. Now with the age of AI, age of agents, age of fast hardware, We're there.
John Furrier
>> Yeah, I interviewed Cognichip. They basically have built an AI model for the physics of years of learning on how to produce chips. You're starting to see AI become key. One other interesting thing I want to get your thoughts on, because I'm old school. I remember the old rack and stack days. It used to be the word was let's stand up some infrastructure. Now it's turn it on. So your point about ease of deployment becomes a huge factor. Talk about that piece, because that Hugging Face thing you just mentioned, is kind of in that same vein of, hey, how fast can we turn it on, not just stand up?
Darren Chien
>> Yep. So in terms of turning on these systems, right, you're starting to see the trend towards rack-scale systems. It comes to your data center in a rack, you plug it in, you turn it on, and it's good. What it shows me is that we are living in a world of silicon, which is hardware, but ultimately it powers software. What we're forgetting is there's a whole physical world that these servers, these data centers have to plug into. You know, it's bringing massive amount of jobs into country, into cities, countries, regions, United States, where it's bringing electricians, plumbers, construction to build up this infrastructure. And when we turn it on, we need energy. Whether it's solar, whether it's wind, whether it's nuclear, we need energy in the United States to power this intelligence.
John Furrier
>> Talk about sovereignty because you mentioned geographical boundaries. Enterprises view themselves as sovereign, not nations, but environments. So it's not sovereign nations, of course, there are sovereign nations, and that's geographical, good for telecom, things that we know. But when you look at an enterprise, they're multinational companies. They will soon be— you'll soon be a multinational company if you're not already kind of embedded. But sovereignty for an enterprise is, it's our world. How do we translate our workflows? Because they drive on the left side of the road in London, but on the right side in America. That's a governance problem. These are little things that are starting to come up as agents start traversing through domains. Namespaces become important. A lot of stuff comes up that's not obvious. Sovereign is a complex— how do you define sovereign now and what are the core issues?
Darren Chien
>> So the first thing I think of in sovereign is sovereign cloud, right? It's the idea that each country has recognized that AI is an issue of national security. I have my country's data, I have my country's processing power, I have my country's individual use case. What can I do to keep that compute inside the four walls of my border? And so we're starting to see some really interesting projects coming out of India, coming out of Korea, coming out of Malaysia. NAVER Cloud is one of the participants in our fundraising round. They're standing up a huge data center, which we're really excited to be a part of there. But it basically provisions compute in region such that either, when you're doing training, if you're talking NVIDIA or if you're doing inferencing, it's in the four walls of your country.
John Furrier
>> What's interesting about cloud sovereignty is that it used to be a privacy issue around data. But you mentioned tokens and tokens per watt, token economics, money's involved. So if a country wants to produce AI, that'll have an economic impact, money, revenue in country.
Darren Chien
>> Correct.
John Furrier
>> So you're going to start to see, we believe, an in-country dynamic of, okay, we're going to stand up the data centers, turn them on, produce value and generate intelligence, leverage the energy that stays in country. That's an economic— that is a leadership issue that's beyond is someone's privacy violated, a GDPR thing or something.
Darren Chien
>> Well, there's the economic benefits, but there's also a performance benefit of latency. Right? You talked about all of these super fast, super premium use cases of robotics, self-driving, real-world surgical advice. When you have the Meta glasses on and you're looking, what do I do? Where do I cut? Latency is a key driver there. And so when you're thinking about things like edge data centers or sovereign compute, locating the processing power to perform inferencing such that it gets you answers faster is a critical aspect of making sure that
John Furrier
>> this—>> I think that's why sovereignty expands beyond the definition, because if I'm an enterprise, a hospital, I want to have an operating room that's going to have computer vision. Why can't I tap doctors around the world real time? Correct. So I'm going to have something on property.
Darren Chien
>> Uh-huh. And we're working with a party who can do just that. These will come out shortly.
John Furrier
>> Darren, you're excited. The energy is awesome. Congratulations. Thank you, sir. Good stuff for you guys. Positron AI, again, another example. It's a startup. They're like a $5 billion company now. So again, this is the world we're living in. AI infrastructures continue to thunder away. Again, standing up and turning on intelligence will be a multi-year wave that's also going to accelerate code generation, agentic, and bring other compute, XPUs, GPUs to the table for the new systems that are coming. They're called AI factories. We're doing our part here on theCUBE. I'm John Furrier, your host of theCUBE. Thanks for watching.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, co-hosting theCUBE here with Gemma Allen, my co-host. Hello, I'm John Furrier, host of theCUBE. We are here at theCUBE's NYC studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is our AI Factory series. We've talked to the leaders who are making it happen. Of course, we've got an event tonight with a lot of investors in the hottest companies. Swan Company will be having some remarks at this meetup here in New York City. Again, we got Darren back. He was just back on theCUBE in our, the NYSE Wired program in July. He's the managing director of APAC for Positron, with big news just recently, big-time funding at a $5 billion valuation. It was like $575 million. Darren, great to see you again.
Darren Chien
>> Great to see you.
John Furrier
>> You got a spring in your step. You got some fresh funding, great validation on the valuation. Congratulations. Some great board members. Really stacking the team up since we last talked. Of course, it was probably in the works. You couldn't say anything. Now it's all out there.
Darren Chien
>> Now it's public. And yeah, great news. Like you mentioned, great validation. Fundraising is just a tool in our tool belt to really get to where we need to go. And so really happy to be here. Palo Alto was great. This is a different, different vibe.
John Furrier
>> Yeah. You know, I remember a year and a half ago, Mitesh and the team grinding away and the market continues to grow. The AI infrastructure If you look at what's happened just on the supply chain, obviously HBM memory, KV cache is exploding. Disaggregated serving is standard pretty much now in the clusters. You get different approaches, but the demand is just off the charts. Speak to that dynamic because it's just not stopping.
Darren Chien
>> You already saw the level of the demand a year and a half ago. I don't even know how much bigger it is now and how much bigger it is going to be. Even within our own specialized niche industry of chips, right? You used to think it was one chip to rule them all, maybe NVIDIA, maybe then it's NVIDIA and AMD. Now we have NVIDIA, AMD, Groq, Cerebras, all the frontier labs working on their own chip efforts. You have the startups like us trying to make a dent. The beautiful thing about it is that the market is so large that there is a wedge for all of us if we bring something differentiated to the table, bringing value to the customers, bringing scale to the customers. It's a great environment to be in.
John Furrier
>> I mean, if you look back 3 years ago, Dave and I would probably have commented on our CUBE Pod, it's a winner-take-all. NVIDIA is looking good, middle of the fairway. You got the hyperscalers buying up everything and then the rise of the Neo Cloud comes in. Okay, great. Still potentially a winner-take-all. And then I think our commentary shifted to rising tide lifts all boats. To use that example, that's kind of what happened. And if you look at— I want to get your thoughts on this because there's some data points and I like to get your analysis on this, because if you look at just the overall growth of inference, this aggregate survey I mentioned, these are signals that it's not a one-chip game. In fact, NVIDIA, they don't want to be called a GPU company. They're an AI infrastructure company. Now you've got footprint diversity coming into the equation with AI factories. Yeah, those monster factories pumping out tokens. Tokens equals revenue. I totally buy all that. But it's not just the one monster system. You guys see obviously different models. Smaller models can run on a lot less compute. agents are compute hungry. Yeah, they'll use GPUs. They'll use the Mellanox. Talk about that piece of the dynamics because I think this idea of networking with the KV cache Dynamo from NVIDIA speaks to an ongoing trend that you're going to have a lot of different systems that will perform high-performance token generation. But it's not the same thing. No, the design will be different. This is where the chip diversity comes in. You guys play into that, right?
Darren Chien
>> 100%. And listen, I spent 2 years at Lambda. When I started at Lambda, if you told me I wanted to use anything other than NVIDIA, I would have laughed you out of the room. I said, hey, Talk to someone else or give me what you're having. Now, 2 years later, we are in a space where there is demand for Cerebras, there is demand for Groq, there is demand for AMD. What that is telling us is that, yes, inference equals revenue, but not all inference is equal. You have premium inference. These are the tokens that you want the fastest, you want the smartest, you cannot fail. You are willing to pay 10 times more for maybe 5 times more serving speed. Then there's also the 80% of the inference, which is your code generation, your video generation. You can afford to wait. Wait a little bit, you really care about economics. And that's kind of the wedge that Positron AI plays in. Cerebras, fantastic technology, wafer-scale engine, my goodness, a technological innovation never before seen. But they serve really fast but really expensive tokens, right? We are focused on the workloads, specifically the decode workloads that require massive memory capacity, trillion parameter plus models, multiple trillion parameter plus models, and just ultra-high memory bandwidth utilization at a really great cost.
John Furrier
>> Yeah, talk about the economics, because I think when we, if you look at the history, let's take the computer industry, the PC industry, you had a 286 processor, 386 processor, each one had kind of like an entry-level, mid-range, flagship, the classic product positioning. Okay, you've kind of seen that on Pareto curves. You say, okay, to your point, yeah, I can wait a few seconds. Milliseconds, but some things require the best— robotics, autonomous vehicles, real workload, financial transactions. Talk about what that means from a chip standpoint. You mentioned Decode. Explain the nuances of this inference market because you're starting to get into product positioning of token generation. That's what's happening.
Darren Chien
>> Exactly. And so we talked about this in our last talk, prefill and decode, two workloads within inference that are kind of served by different processing powers. prefill is still compute heavy, still parallel, and that's when you upload your file, you input your chat into ChatGPT, and it can process that in parallel. Then you have decode, which is still memory bandwidth bound, right? You still need massive memory bandwidth utilization, massive memory capacity to fit the models to fit the KV caches. And so when we're looking at the workloads that Positron AI is positioned for, high-frequency trading is actually a use case that fits great on our hardware. We have Jump Trading who co-led our Series B. We have HRT who is getting involved on our Series C. That should tell you something. And then across the entire set of workloads, you mentioned Pareto curves, right? We're really pushing the Pareto curves of the optimal point matching throughput per dollar and interactivity up and to the right as it compares to some of the alternatives in the market.
John Furrier
>> So you guys are positioning Positron for the high-end piece of the market?
Darren Chien
>> Not necessarily at the high end, but the two metrics that we will win on are tokens per dollar and tokens per watt. So making it economical to generate these tokens and energy efficient to generate these tokens across all use cases and price performance layers. Correct. Within inference.
John Furrier
>> Inference.
Darren Chien
>> Okay, got it.
John Furrier
>> Okay. So I had to look it up on the internet here because I didn't know the exact date, but the, NVIDIA Groq deal was around December 2025. They were probably in the works a couple of months before. But let's just say that the fall/winter of 2025, that's literally 9 months ago, 10 months ago.
Darren Chien
>> Okay. not years.
John Furrier
>> So think about— I want you to share your thoughts on the importance of that deal because since then it went from NVIDIA to we can have a separate thing for inference. Look at the market since then. Explain what happened, because that was kind of like a very much a shot heard around the world in the chip game.
Darren Chien
>> Correct.
John Furrier
>> Because it opens up. Okay.
Darren Chien
>> Wow.
John Furrier
>> Again, back to your point about the wedge and room for everyone to play, that kind of was the template for everything we're seeing now. And agents and code generation, they're coming in, they have different needs. Compute becomes more important with agents. You can then free up those GPUs that are being used for things like that. Talk about that importance, because that Groq deal really was the moment the industry seemed, at least from my standpoint, to shift.
Darren Chien
>> Yeah. So to be clear, NVIDIA is the largest company in the world. It's one of the smartest companies in the world for a reason. As new chip startups or as anyone trying to play in this AI game at all, we have to understand where we fit in this dynamic. You brought up disaggregation, splitting prefill and decode. That is exactly a scenario in which we could look to NVIDIA and say, hey, listen, your chips are the best at compute-bound workloads. Training, unquestionable, you will use NVIDIA. But when we look at inference, let's see, perhaps we pair an NVIDIA Vera Rubin rack with a rack of Titan, and let's see, maybe we shift some of the prefill work to the accelerator that is more compute-heavy, and let's shift some of the decode work, where we can get more economical. we might not get as fast as an NVIDIA Rubin on a memory bandwidth perspective, but a memory bandwidth utilization and efficiency and capacity perspective. Let's pair with the NVIDIA ecosystem and see how we can make, 1+ 1= 3. That's the dream of any chip startup in today's environment.
John Furrier
>> And by the way, from a system standpoint, that makes total sense because NVIDIA has to make a choice. Do I want to bogart the whole thing and slow down progress? And you see Jensen, they're okay with this. It's not like they're against any of this. They're actually promoting the notion of more AI for everyone. In fact, they're making the market on the policy side. They're making the market on the ecosystem side. And they're actually saying bring more stuff to the table because they'll win too. There's supply constraints, but also capability.
Darren Chien
>> 100%. Listen, we are in the footsteps of giants here, right? NVIDIA laid the path. We would not be here if not for NVIDIA. But as you mentioned, they are making moves to partner with startups like ours. Another company that is going to be on this panel is going to be at the event later tonight, d-Matrix. They had some huge news of being, I think, integrated into the NVLink ecosystem with NVIDIA. So it's going to show that NVIDIA is not trying to do anything that is aggressive or do anything that's harmful for the ecosystem. They understand that the market is so large that we at Positron AI just need to come up with a proposal or a way of working together that makes sense for all of us.
John Furrier
>> Yeah, rising tide lifts all boats. There's plenty of beachfront for everybody to play and camp out and create innovation. Okay, talk about the funding and why that was driving so hard because again, the validation, $5 billion, you only raised what, $600 million, $575, what was the
Darren Chien
>> amount?
Darren Chien
>> total round was $875 million.
John Furrier
>> $875 million, so a little bit under a billion. You could have done better. Mitesh, come on, get it up to a billion, get across the threshold. Only kidding, obviously, that's plenty of capital. But the validation on the valuation is for a reason. What's the reason why? Why is the momentum so strong with Positron?
Darren Chien
>> Listen, it's a couple of things. One, our Series B fundraising, we raised around $230 million untranched. So in the bank, you're talking about the billion. I can see our bank, I see that billion number. So we feel pretty good about that. Two, we are actually shipping product today. Our first-gen Atlas system is deployed in a hyperscaler environment in OCI. They've been great partners for us. We've been able to demonstrate real traction, real shipping expertise, which is always the question mark when, the investors who participated in our round, they're the most sophisticated people. On the planet Earth. They looked at our team, they looked at our current product deployments. What really got them excited is our future products. So our Asimov silicon taping out this year, we expect to have chips back, sometime next year and we'll be shipping them in Titan servers, both air-cooled, liquid-cooled SKUs depending on what the customer wants end of next year. And so we are really excited for that to come out because that is going to be really the point in time in which you will see Positron AI is making a dent in the market.
John Furrier
>> Okay, so we know what's happening. Inference is the killer app. The systems are becoming more expansive on the capability side. We know what it means. There's demand that needs to be served and you guys got the products. Connect the dots on what's next. What is that vision for the future products? Without giving away any trade secrets, take us through what's the next hill you guys are going to climb?
Darren Chien
>> Listen, fundraising is good. It's not the end-all be-all. It's not the goal. The goal is more chips in the world serving tokens that everyone can access, right? The chips are a way to just democratize intelligence, which the unit of tracking that is tokens. But what we're really looking to use this fundraise and any future fundraisers for is to continue to iterate on our custom silicon. We're starting with Asimov. The next generation is Bradbury. The generation after that is Clarke. And I'll let you kind of guess what the following generation is.
John Furrier
>> You guys definitely got me thinking. Okay, so now let's get back to use cases. What's the top use case that you guys are selling to now and what is developing?
Darren Chien
>> So the main use cases that we see are code generation and video generation. Those are currently the new gen— the new use cases. Myself personally, the agentic use case is the most compelling one. I don't know if you've seen the latest news around this thing called Manus. My goodness, my flight here was booked using Manus. I didn't have to interact with the United app. I had a middle seat. I asked, hey, please change me to an aisle seat. It changed me to an aisle seat. These agentic use cases we've been hearing about for years are actually here. You and I, real people, not inside Fortune 500 companies, can use these agentic use cases today.
John Furrier
>> I was one of those people. I was in a room about a couple of years ago, and the question was, when will agents book flights for you? And I was one of the skeptics. I always admit when I'm right, obviously, and admit when I'm wrong. I said, I don't think it's gonna happen. And mainly because I didn't think it would be advanced enough to cross some of these transactional rails. And this is what's getting interesting. They're so smart now, they can actually do it. I thought you'd see some great automation, basically RPA on steroids to begin with, and then I would see reasoning. But I didn't, I said 5 years, maybe. It was 3. So now we have change the flight. What's going on in my PG&E bill? It's overdue, payment's due today. So a lot's involved. So trust is a huge factor in all these agents. So I was wrong on that one, but hey, better to have it now because Anthropic is getting access to all my stuff. And a little bit scary, but the upside is pretty damn good.
Darren Chien
>> I have my passport in there, my driver's license, my credit card. Luckily, there's some strong integrations. Anthropic is working with 1Password. Meta has taken some strong security and privacy precautions but there is risk as with everything. But the upside, like you mentioned, is just incredible.
John Furrier
>> Which one do you like better?
Darren Chien
>> Oh, Muse is a really fast, really delightful experience. But there's something about texting your agent on iMessage where I'm already living that is— Yeah, yeah. That
John Furrier
>> doesn't—>> And they got WhatsApp too. They got WhatsApp and iMessage.
Darren Chien
>> I'm an iMessage guy. we had Europeans in this morning. They use WhatsApp. WhatsApp, but iMessage for
John Furrier
>> us.>> I know, I am completely schizophrenic between the different— and of course you have the random Android user in there. Yeah, let's not talk about
Darren Chien
>> that.
Darren Chien
>> Let's not talk about those.
John Furrier
>> All right, so what are you most excited about right now? Obviously it's a great time for you guys. Congratulations to the team. We're big fans. Obviously we were following you guys since the beginning. well deserved. We need more chips. What's the most exciting thing going on now?
Darren Chien
>> Listen, I think the external validation— big fundraise number, big valuation number. Great board members. I could talk about our board members for this entire talk. The external validation is good, but unlike some of the other startups that exist, we really pride ourselves on internal validation. We want to make sure that we are taping out and shipping a chip and system that actually delivers on performance claims such that it drives value to our customers.
John Furrier
>> Darren, I love that attitude and I really appreciate that because one of the things I was very critical of during the last wave of tech innovation. Fundraising was a celebratory milestone, like, we made it. In this new era, the pragmatic view is, look, we just want gas in the tank. We're not really focused on that as a— it's a milestone, mile marker, moment in time. Yeah, you said it's like winning a game. The coaches celebrate for the night, but tomorrow we'll get back to work. That's an attitude change that I've seen in this new generation because I think it's mainly because, one, the expectation for the market is so strong that the reality is sinking in like, hey, we got $1 billion in the bank, let's have a big party, let's throw the— let's get some cool stuff for the office. Moves to no, no, we got more work to do. That's different. What's your take on that? And what's your— do you agree with that? Do you think that's a general sentiment or is there still kind of like, hey, we got the money in the bank, let's go to Tahoe and we could do an offsite meeting. But I'm saying, you know what I'm saying? I'm not stopping.
Darren Chien
>> No, no, no. The only thing I disagree with you on is that it is a generational change. No, I think we're pretty unique in this in that we don't view this, obviously we celebrate. It's great news. we're going to go out to dinner. But like you mentioned, it's not the tool. It is not the goal. It is a tool in the toolbox. Right. And I credit this to Mitesh and Thomas, CEO and CTO of Positron AI. they've seen this before. They've been building startups for the past decade. Right. The money that you get from investors is probably some of the most expensive dollars you'll ever take, right? They have expectations. You must deliver. And so if anything, it is a weight on our shoulders to make sure that we deliver.
John Furrier
>> you got to deliver.
Darren Chien
>> It's���
John Furrier
>> there's expectations.
Darren Chien
>> There's expectations.
John Furrier
>> Here's the gas in the tank. Get to your destination.
Darren Chien
>> The beautiful thing is with such a large fundraise, now we can get creative. How do we accelerate? How do we de-risk? Do we spend $15 million more on emulation hardware? Maybe we should if that gets us, 10% of the way there. we're obviously going to be expanding the team, new office in Austin, expanding our facilities in Spokane, doing some international expansion that I'm very excited about. Yeah, but the money is a tool that we're going to use.
John Furrier
>> It's a totally great attitude. Let's talk about the board members. You mentioned some of them. We checked them out. You got Gavin, you got Dylan Patel from SemiAnalysis, although he's got an analyst firm, so maybe he's got some inside baseball there. You got Jim Clark from Netscape, founder of Netscape. Talk about some of the board members, the new board members, and how has that changed the dynamic?
Darren Chien
>> So Gavin has been involved with the company for a while, but we're bringing him onto the board as part of this round. Gavin— no, no, who, who— I cannot give a better introduction to Gavin than he can himself, on his Twitter, on his podcast. So obviously his Twitter is huge. He's become a leading voice in AI. So having him join the board, we're really excited to draw on his expertise, see what he sees in the market. Same thing with Dylan. SemiAnalysis, SemiAnalysis Capital, you know, participated in the round. He sees everything right across the board, across the chip, but even down the layers. Right. Because when we're in the chip space, we need all of the inputs, the substrates, the assembly, the packaging, the wafer, the memory. And so having someone like Dylan who has a great insight into the entire market, but actually what I'm most, most excited about is Jim Clark. So you mentioned Netscape, co-founded with Marc Andreessen. The other company he was involved with, SGI, Silicon Graphics.
John Furrier
>> Yeah. Yeah.
Darren Chien
>> What better name can you have on the board of a silicon company, right?
John Furrier
>> Well, SiliconANGLE was another great, great name. Silicon Graphics, because silicon's back.
Darren Chien
>> Silicon is back.
John Furrier
>> And talk about the software stack, cuz one of the things that's emerging outta this is that layer you mentioned, code generation and GenTech. You're now starting to see a new kind of software stack. It's not like the old days, hey, buy a server and load Linux and run your stuff and put some virtualization on it. A new software layer, because as you pointed out, The new chips are creating new subsystems around this holistic system.
Darren Chien
>> Correct.
John Furrier
>> That's powering all the AI. Okay, great. We agree on that. But the software stack now has to adapt.
Darren Chien
>> Correct.
John Furrier
>> What's your vision on that? How does that play out?
Darren Chien
>> So we know that the customers who are buying Positron AI systems are plugging them into GPUs, TPUs, AMDs, Trainium. They don't want to have to deal with another compiler, another language, another runtime. Right. And so the vision for Positron AI and what we're trying to develop on the software side is basically Hey, Hugging Face Transformers library compatible, download them onto the system. We will spit out an API. It'll plug right into your existing fleet. Obviously a lot of work to be done there. But with this new agentic development that obviously the software engineers will be more familiar with is we can download, ingest, and run models in a matter of hours. I remember in Lambda's day, it was a team of engineers 24/7 round the clock for a week before we could get something up and running on Hopper systems. Now with the age of AI, age of agents, age of fast hardware, We're there.
John Furrier
>> Yeah, I interviewed Cognichip. They basically have built an AI model for the physics of years of learning on how to produce chips. You're starting to see AI become key. One other interesting thing I want to get your thoughts on, because I'm old school. I remember the old rack and stack days. It used to be the word was let's stand up some infrastructure. Now it's turn it on. So your point about ease of deployment becomes a huge factor. Talk about that piece, because that Hugging Face thing you just mentioned, is kind of in that same vein of, hey, how fast can we turn it on, not just stand up?
Darren Chien
>> Yep. So in terms of turning on these systems, right, you're starting to see the trend towards rack-scale systems. It comes to your data center in a rack, you plug it in, you turn it on, and it's good. What it shows me is that we are living in a world of silicon, which is hardware, but ultimately it powers software. What we're forgetting is there's a whole physical world that these servers, these data centers have to plug into. You know, it's bringing massive amount of jobs into country, into cities, countries, regions, United States, where it's bringing electricians, plumbers, construction to build up this infrastructure. And when we turn it on, we need energy. Whether it's solar, whether it's wind, whether it's nuclear, we need energy in the United States to power this intelligence.
John Furrier
>> Talk about sovereignty because you mentioned geographical boundaries. Enterprises view themselves as sovereign, not nations, but environments. So it's not sovereign nations, of course, there are sovereign nations, and that's geographical, good for telecom, things that we know. But when you look at an enterprise, they're multinational companies. They will soon be— you'll soon be a multinational company if you're not already kind of embedded. But sovereignty for an enterprise is, it's our world. How do we translate our workflows? Because they drive on the left side of the road in London, but on the right side in America. That's a governance problem. These are little things that are starting to come up as agents start traversing through domains. Namespaces become important. A lot of stuff comes up that's not obvious. Sovereign is a complex— how do you define sovereign now and what are the core issues?
Darren Chien
>> So the first thing I think of in sovereign is sovereign cloud, right? It's the idea that each country has recognized that AI is an issue of national security. I have my country's data, I have my country's processing power, I have my country's individual use case. What can I do to keep that compute inside the four walls of my border? And so we're starting to see some really interesting projects coming out of India, coming out of Korea, coming out of Malaysia. NAVER Cloud is one of the participants in our fundraising round. They're standing up a huge data center, which we're really excited to be a part of there. But it basically provisions compute in region such that either, when you're doing training, if you're talking NVIDIA or if you're doing inferencing, it's in the four walls of your country.
John Furrier
>> What's interesting about cloud sovereignty is that it used to be a privacy issue around data. But you mentioned tokens and tokens per watt, token economics, money's involved. So if a country wants to produce AI, that'll have an economic impact, money, revenue in country.
Darren Chien
>> Correct.
John Furrier
>> So you're going to start to see, we believe, an in-country dynamic of, okay, we're going to stand up the data centers, turn them on, produce value and generate intelligence, leverage the energy that stays in country. That's an economic— that is a leadership issue that's beyond is someone's privacy violated, a GDPR thing or something.
Darren Chien
>> Well, there's the economic benefits, but there's also a performance benefit of latency. Right? You talked about all of these super fast, super premium use cases of robotics, self-driving, real-world surgical advice. When you have the Meta glasses on and you're looking, what do I do? Where do I cut? Latency is a key driver there. And so when you're thinking about things like edge data centers or sovereign compute, locating the processing power to perform inferencing such that it gets you answers faster is a critical aspect of making sure that
John Furrier
>> this—>> I think that's why sovereignty expands beyond the definition, because if I'm an enterprise, a hospital, I want to have an operating room that's going to have computer vision. Why can't I tap doctors around the world real time? Correct. So I'm going to have something on property.
Darren Chien
>> Uh-huh. And we're working with a party who can do just that. These will come out shortly.
John Furrier
>> Darren, you're excited. The energy is awesome. Congratulations. Thank you, sir. Good stuff for you guys. Positron AI, again, another example. It's a startup. They're like a $5 billion company now. So again, this is the world we're living in. AI infrastructures continue to thunder away. Again, standing up and turning on intelligence will be a multi-year wave that's also going to accelerate code generation, agentic, and bring other compute, XPUs, GPUs to the table for the new systems that are coming. They're called AI factories. We're doing our part here on theCUBE. I'm John Furrier, your host of theCUBE. Thanks for watching.