Darren Chien of Positron AI, managing director for Asia Pacific, APAC, discusses
Positron's field-programmable gate array, FPGA-based Atlas product and its
upcoming Asimov chip and Titan system within the context of robotics and
artificial intelligence, AI infrastructure. Chien explains that inference is
fundamentally an economics and energy-efficiency problem and they emphasize
optimization for tokens per dollar and tokens per watt. They highlight the
distinction between compute-bound prefill and memory-bound decode, the
importance of context windows and Titan's air-cooled design that enables legacy
enterprise deployment without costly retrofits. The conversation addresses
memory capacity and memory bandwidth as critical factors for memory-optimized
systems and for enterprise scaling across Taiwan, Korea, Japan and Singapore.
This theCUBE Research conversation with hosts John Furrier and Gabe Olave occurs
at the NYSE Wired AI Leaders Summit. The hosts discuss market competition and
open models and underscore APAC's strategic role in silicon and supply chains.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Darren Chien, Positron AI
Darren Chien of Positron AI, managing director for Asia Pacific, APAC, discusses
Positron's field-programmable gate array, FPGA-based Atlas product and its
upcoming Asimov chip and Titan system within the context of robotics and
artificial intelligence, AI infrastructure. Chien explains that inference is
fundamentally an economics and energy-efficiency problem and they emphasize
optimization for tokens per dollar and tokens per watt. They highlight the
distinction between compute-bound prefill and memory-bound decode, the
importance of context windows and Titan's air-cooled design that enables legacy
enterprise deployment without costly retrofits. The conversation addresses
memory capacity and memory bandwidth as critical factors for memory-optimized
systems and for enterprise scaling across Taiwan, Korea, Japan and Singapore.
This theCUBE Research conversation with hosts John Furrier and Gabe Olave occurs
at the NYSE Wired AI Leaders Summit. The hosts discuss market competition and
open models and underscore APAC's strategic role in silicon and supply chains.
>> Welcome back everyone to theCUBE at NYSE Wired, third annual AI Leaders Summit here and CUBE Day here in Palo Alto. I'm John Furrier, your host. Of course, we have our event tonight, 180 leaders coming together for the pool party, our third annual, it's not really a pool party, so don't show up with swimming trunks as one person did the first year. Darren Chien's here, the managing director of APAC at Positron AI, fast growing inference chip and system for the killer app. Darren, great to have you coming on. We've been covering you guys, love what you do. So give us an update, APAC, how's that connect in? You guys are on escape velocity. What's up?
Darren Chien
>> No, great to be here, John. Thank you so much for having me. You've been watching the program for a while and great to be here. So since you last caught up with Mitesh, I believe it was a few months ago, we've had great commercial traction on our first generation product Atlas. In terms of APAC, all roads lead to Taiwan, really growing our relationships there, working with customers, investors, partners, to grow our relationships, not just in Taiwan, but all around APAC, Korea, Singapore, Japan.
John Furrier
>> A lot of international, I love the traction. The inference market we saw last year, obviously Nvidia buys the IP from Groq, they go over there, we see the rise of Cerebras. It's been pretty clear for insiders and now the general public that inference is the killer app of all of AI, obviously training infra is still going on, the large models, specialty intelligence is becoming big, we're starting to see that. Agents eat inference for breakfast, they need more tokens, they're inferring all the time. So inference will be the standard discussion. Where are we on the innovation side, cost, performance? What are the key things that are going on for people to understand where the puck is going for inference?
Darren Chien
>> For sure, for sure. Yeah, over the past decade or so, AI hardware has been mostly focused on training, getting the models ready for production. Now we're at a state where the models are ready. They're here, GPT-5.6, Fable, Kimi K3, whether it's closed source or open source, the models are here. Demand for tokens is insatiable. It's almost unlimited. there's a lot of work needed to be done. We see the cost for Fable, for GPT-5.6, for Kimi K3, it's pretty much only available to the people who can pay a pretty penny for it. There's a lot of work needed to be done on the hardware side, infrastructure side, software side, to make these optimizations, to allow these models to be accessible and available to everyone.
John Furrier
>> I have a lot of conversations about inference and one day it's, hey, Cerebras has got the fastest throughput, big chip, on the other hand I hear, hey, supply chain's tight, let's make CPUs work. So there's a lot of industry conversation around optionality and also architecture. What's your view on this? Because we're starting to see people, the demand is there, obviously. So inference will be a long game market, but now it's not like just buying one SKU. There's no one SKU, there's no one company. How do you guys look at that and how do you explain that and rationalize that?
Darren Chien
>> Yeah, we look at GPUs as a general purpose, accelerator, right? Great for training, great for inference. But when we look into the realm of inference, we can split that into pre -fill, decode, compute bound, memory bound workloads. And within that, there are different metrics that you can optimize for interactivity, throughput, cost, energy efficiency, right? Fundamentally at Positron AI, we believe that inference is an economics and an energy efficiency problem, right? We want to drive these tokens to be as cheap as possible so that we can generate as many of them as possible to as many people as possible. That's what we've built and designed our architecture from the ground up to do.
John Furrier
>> And explain where you guys are at on the product side.
Darren Chien
>> Yeah, so we have our first generation Atlas. It's an FPGA product. It's live shipping to customers. We're going to end the year close to 100 million in revenue just on that alone. It's an FPGA, so it's flexible, allowed us to get fast time to market. But what we're really looking forward to is our second generation Asimov chip, which will be shipped in a Titan system, that's going to be really focused on massive memory capacity and really high memory bandwidth utilization.
John Furrier
>> And what are those constraints that Titan solves?
Darren Chien
>> So Titan really solves those two metrics, memory capacity and realized memory bandwidth utilization, where with memory capacity, the more memory capacity you have, the larger model you can fit, longer context length, models can remember more about you, and you can serve more users. And then in terms of memory bandwidth utilization, that's in terms of tokens per second, generate more tokens faster.
John Furrier
>> All right, so I got you here because you're good at explaining. So explain two concepts, the pre -fill and decode, and then context window and why that's important. So explain what pre -fill and decode is for the folks that are learning, because this is material to how things work.
Darren Chien
>> Yeah, so pre -fill and decode are the two different parts of inference, or the two different parts that make up inference. On the pre -fill side, you can think of that as when you put in, you copy and paste your entire code repository, you put it into ChatGPT. When ChatGPT has to read through that entire thing, it doesn't have to do it one at a time because you've uploaded the whole thing. So that's a massively parallel problem, which is compute bound.
John Furrier
>> So that's something that - It calculates what you're basically prompting.
Darren Chien
>> Correct, it can see the entire thing all at once. It can analyze it in parallel. So that's what we call the compute bound operation. So within inference, there's still compute bound operations. However, on the decode side, that's when after you've put in the GitHub repo, it's analyzed it all, it now needs to spit out one word or one token at a time, what its response is, that is decode. And that is memory bound. And so that's why you need a lot of memory capacity and memory bandwidth in order to make that output or the response back to you fast and accurate.
John Furrier
>> So to put it in simple terms, here's the request and then it thinks.
Darren Chien
>> Correct.
John Furrier
>> And that's the tokens and that's the context.
Darren Chien
>> Correct.
John Furrier
>> Great definition, because we hear a lot of that in these conversations, oh, I'm going to use AMD, I'm going to use Vera, we all know what a CPU is, we've been around PCs and servers.
Darren Chien
>> Correct.
John Furrier
>> But the GPU, which does all the matrix multiplication and all the high speed stuff, that's the engine for thinking.
Darren Chien
>> That's compute. So raw compute, really great for training, really great for pre -fill, but not so, on the decode side, it's more of a memory bound operation.
John Furrier
>> That's where you need, that's the memory we see in supply chain, Micron stocks up, everyone wants more memory.
Darren Chien
>> Correct.
John Furrier
>> And SRAM and you got solid state, SRAM, DRAM, HBM. Solid state, solid time is killing it, it's all these, okay, got it. Now context window, we hear that, and that's kind of gone mainstream, like my context window for this model, and most people relate it to a bigger prompt. I could put a Word document in there, GitHub, code repo, so that's stuffing in the prefill basically. Right?
Darren Chien
>> Correct.Well, context window refers to basically how much the model can remember about you. So if you remember using ChatGPT when it first came out, you would have to remind it, hey, I am Darren Chien, I'm running Positron AI's APAC business. You would have to tell it that every single time. Now when I talk to ChatGPT, it knows that about me. Hey, I don't need to explain it to you, just you know this about me, just shape your response around that.
John Furrier
>> Yeah, I do that a lot with ChatGPT. Based on my chat history, tell me where my head's at right now.
Darren Chien
>> You don't even need to tell it based on my chat history because it remembers.
John Furrier
>> Yeah, it's good because it has all my editorial in there. But okay, so now get into that, how will you guys see it? Because now you guys are attacking the large scale piece. Now, we just laid out the inference equation, context, windows, prefill, decode. What are you guys optimizing for? How should people think of Positron?
Darren Chien
>> Yeah, that's a great, great, great, great question. And our answer is we believe that inference is fundamentally an economics problem, right? We need to make sure that these tokens are generated as cheaply and as widely available to users as possible. And so the way we compete is on optimizing the two metrics of tokens per dollar and tokens per watt. There are other competitors who will win on throughput or interactivity, but we believe that inference is fundamentally an economics problem.
John Furrier
>> On the enterprise side, you mentioned enterprise. Scaling is a huge deal. They're paying a lot for tokens. We see a lot of budgets kind of on hold. We saw some IBM blip in the earnings that was due to some spend. Obviously supply chain had impact there, but people got surprised with these big token bills. everyone's using cloud. Oh, it's just when it says I'm done and hit upgrade and 10 ,000 people spend another 89 or 150 bucks times 10 ,000, it adds up super fast. How do you scale to the enterprise? Because they don't have the huge CapEx budgets. even financial services, they're not a $15 billion CapEx, NeoCloud / neocloud (ambiguous), although JPMorgan Chase has a big budget, But the enterprises have to deal with their pre -existing stuff. They have a lot of compute, they got a lot of Dell, HPE, Lenovo, NetApp, and all the storage in there, and they have all that old IT infrastructure. How do they turn on their systems? How do you see that scaling?
Darren Chien
>> And that's actually a fantastic question, because it's exactly the problem that Positron is solving, which is that you have these legacy enterprises, they have access to air-cooled data center space that they have existing infrastructure in. they might have a rack or two available, but when you're talking about these, the latest and greatest of NVIDIA GPUs, for example, or these rack scale solutions, which require liquid cooling, you're not going to be able to do that in your existing data center. Our Titan systems are designed by default to be air cooled. We can do liquid cooled if you want, but since they're designed by default to be air cooled, we can slot those into those spaces. We can sell it to you as a server. We can sell it to you as an entire rack. And it really just lowers the barrier to entry for these enterprises. they don't need a retrofitted data center, they don't need to go looking and competing with Meta. They can look at their existing portfolio and add net new compute where they otherwise may not have had.
John Furrier
>> I talk to Mitesh about all this all the time because he and I talk about scaling and some of the market stuff, but talk about the origin story. What's the culture like? You have other founders in there. What's the DNA? Where'd this all start? What's the, how do you do the TLDR of the origin story?
Darren Chien
>> So the TLDR is a bunch of smart people came together, looked at AI, looked at where it was going, thought to themselves, what could we do to design an AI chip to serve the workloads of the future, but designed today and not 20 years ago.
John Furrier
>> You know what I love about the story of the company is that it's not like, it wasn't 10 years old. This has just happened. So they're in the middle of it. So they have a clean sheet of paper. They see the inflection around 2022, 2023. That was inside the ropes. 2023. Everyone saw that.
Darren Chien
>> Yeah, 2023, our founders are all have semiconductor background, from background in Groq, Mitesh and myself background at Lambda. So we have a really unique blend of both the chip side, the cloud side, and basically the user facing side.
John Furrier
>> Yeah, I love this conversation. And one of the things we're seeing right now is the biggest build out I've ever seen in my lifetime. I think it's not going to stop because the demand curve is still there. In fact, Dave and I were arguing on our CUBE Pod about the bubble. I said, no, there's no bubble in this demand. Pricing might be just supply chain fluctuations, excitement. Obviously pricing goes up, but at the end of the day, if there's demand to fill, it's there. But when you look at what has to happen, they're standing up infrastructure. So there's demand for the components. Partnerships have changed in the go -to -market. People are partnering much deeper. It takes a village to put out these large scale systems, but once it's built, you got to operate them. So what are your thoughts on the monetization? because we're seeing a massive surge on the neoclouds. You're seeing enterprise now putting their toe in the water with coding, agentics coming in full throttle and physical AI is right behind it. That's the sequence of the waves coming. Yeah. They're monster waves.
Darren Chien
>> Yeah. And it's actually really incredible to see because both at Lambda and at Positron, your potential customers could also be your suppliers, but could also be your competitors. And it just speaks to the pace of the industry. The demand is so unlimited that we should work together at what we're best at, right? So we are the best at designing chips. We could work with a partner who's the best at making them, packaging them. We could work with a partner who's the best at deploying them, or we could work with a partner who's the best at deploying ocean-based data centers, right? It's not like we're going to go in and deploy ocean-based data centers, but we are happy to work with those.
John Furrier
>> Space is coming around the corner, too.
Darren Chien
>> Space as well.
John Furrier
>> Well, I guess my question is interesting because if you take it to the go-to-market and partnership side on the biz dev, on the business development side, it used to take six months to a year to get the engineering teams together. I'm hearing stories of a week, a month, scrum, jump on a Zoom, get face to face, get embedded. How has the partnership equation changed? Because you guys got, you stick to your knitting semi, then you say, okay, this guy does cloud native software, control plane software, here's a GSI to implement, here's a builder who has facility. There's a lot of people that need to be involved at a deep co-design level. How has that changed in the past cycle?
Darren Chien
>> Well, it's changed in that, like you mentioned, it's just shortened the time to make something happen, right? Before you'd plan, you would identify different sites, you would want to de-risk it as much as possible. But because we have strong partners like Oracle, who are really strong technically, really strong engineering, really strong data center footprint, really strong deployments, we are able to go in, speak their language, tell them, hey, this is what we are missing. Can you fill X, Y, Z? They tell us, oh, we have XYZ, we just need ABC. We just find a middle ground and work together.
John Furrier
>> And they know hardware too.
Darren Chien
>> They know hardware as well.
John Furrier
>> They know memory. It's almost interesting how it's full circle. It's a systems game.
Darren Chien
>> It is.
John Furrier
>> All right, final question. APAC, which is your focus, what's going on there? Obviously we see a lot on Taiwan. Obviously TSMC's there. What else is going on? We're seeing telecom starting to get hyper-converged. We're starting to see a lot more in -country sovereignty is a huge discussion. Just what's the macro like in APAC right now?
Darren Chien
>> Well, we saw the open model support letter by Jensen the other day. All the companies in the APAC block signed on supporting that. I see that increasing, right? Because open models just distribute accessibility and availability of intelligence across the world. Speaking specifically in APAC, Taiwan, all roads to Silicon lead to Taiwan, right? We are investing heavily in the country. We have our strong investor base there, strong partnership base there. In terms of customers, right, we talked about air -cooled space, legacy data center space. That is plenty in APAC. There's something like nine to 10 gigawatts of air -cooled space that is waiting to be monetized in APAC.
John Furrier
>> You know what I love about this market, Darren, is that everyone talks about NVIDIA versus AMD. You got to give AMD an assist on that forcing function for Jensen's and everyone else's coalition around the open weights and open source, because they play the open card. So NVIDIA, hey, we're open too. So, and then that helps everyone. You have this massive rise. So it's really not about one company versus the other because everybody wins because the demand curve. What's your reaction to that? Do you believe that same thing?
Darren Chien
>> I love it. I love seeing open models, seeing Qwen3, Kimi K2, the open weights were just released. They had to shut, they had to turn away users because of how much demand there was. And you know what that does? That forces Anthropic to allow Claude to be more widely accessible. That forces ChatGPT to allow Sora to be widely available. If it was just one company there, they called the shots, they set the terms, you will have one model and you will like it whether you like it or not. We have a plethora of optionality and availability that is just better for us as everyday users.
John Furrier
>> And it speaks to capitalism and competition. With no competition, if AMD's not putting the pressure on NVIDIA, you guys don't come in with a new chip. Competition, this is where it's healthy.
Darren Chien
>> 100%, we saw actually that competition forces, whether it's conversations around security where we saw actually GLM, a Chinese open model, being deployed by Hugging Face to fight off the rogue ChatGPT hacking incident, right? It's that type of thing where when there are more widely distributed, more available models to everyone, it's better for the entire ecosystem.
John Furrier
>> If you had one player, we'd be basically still doing marketing documents and email writing. There's no motivation to innovate if there's no competition.
Darren Chien
>> Correct.
John Furrier
>> Look at NVIDIA, they're the leader. Look at you guys coming out in 2023, guns blazing, inference is the hottest product. Constraints drive engineering, drive entrepreneurship, competition drives action.
Darren Chien
>> Correct.
John Furrier
>> Congratulations, great to see you. Say hello to Mitesh and the team. Thanks for coming on. Third annual, we'll see you tonight. All right, I'm John Furrier here in theCUBE Studios in Palo Alto for the NYSE Wired. Third annual CUBE Pool Party tonight, but all day today talking to the leaders, making it happen. Competition, openness, that creates more innovation faster. At the end of the day, you're reining in the chaos. And at the end of the day, you have more innovation doing our part here at theCUBE. Thanks for watching.
>> Welcome back everyone to theCUBE at NYSE Wired, third annual AI Leaders Summit here and CUBE Day here in Palo Alto. I'm John Furrier, your host. Of course, we have our event tonight, 180 leaders coming together for the pool party, our third annual, it's not really a pool party, so don't show up with swimming trunks as one person did the first year. Darren Chien's here, the managing director of APAC at Positron AI, fast growing inference chip and system for the killer app. Darren, great to have you coming on. We've been covering you guys, love what you do. So give us an update, APAC, how's that connect in? You guys are on escape velocity. What's up?
Darren Chien
>> No, great to be here, John. Thank you so much for having me. You've been watching the program for a while and great to be here. So since you last caught up with Mitesh, I believe it was a few months ago, we've had great commercial traction on our first generation product Atlas. In terms of APAC, all roads lead to Taiwan, really growing our relationships there, working with customers, investors, partners, to grow our relationships, not just in Taiwan, but all around APAC, Korea, Singapore, Japan.
John Furrier
>> A lot of international, I love the traction. The inference market we saw last year, obviously Nvidia buys the IP from Groq, they go over there, we see the rise of Cerebras. It's been pretty clear for insiders and now the general public that inference is the killer app of all of AI, obviously training infra is still going on, the large models, specialty intelligence is becoming big, we're starting to see that. Agents eat inference for breakfast, they need more tokens, they're inferring all the time. So inference will be the standard discussion. Where are we on the innovation side, cost, performance? What are the key things that are going on for people to understand where the puck is going for inference?
Darren Chien
>> For sure, for sure. Yeah, over the past decade or so, AI hardware has been mostly focused on training, getting the models ready for production. Now we're at a state where the models are ready. They're here, GPT-5.6, Fable, Kimi K3, whether it's closed source or open source, the models are here. Demand for tokens is insatiable. It's almost unlimited. there's a lot of work needed to be done. We see the cost for Fable, for GPT-5.6, for Kimi K3, it's pretty much only available to the people who can pay a pretty penny for it. There's a lot of work needed to be done on the hardware side, infrastructure side, software side, to make these optimizations, to allow these models to be accessible and available to everyone.
John Furrier
>> I have a lot of conversations about inference and one day it's, hey, Cerebras has got the fastest throughput, big chip, on the other hand I hear, hey, supply chain's tight, let's make CPUs work. So there's a lot of industry conversation around optionality and also architecture. What's your view on this? Because we're starting to see people, the demand is there, obviously. So inference will be a long game market, but now it's not like just buying one SKU. There's no one SKU, there's no one company. How do you guys look at that and how do you explain that and rationalize that?
Darren Chien
>> Yeah, we look at GPUs as a general purpose, accelerator, right? Great for training, great for inference. But when we look into the realm of inference, we can split that into pre -fill, decode, compute bound, memory bound workloads. And within that, there are different metrics that you can optimize for interactivity, throughput, cost, energy efficiency, right? Fundamentally at Positron AI, we believe that inference is an economics and an energy efficiency problem, right? We want to drive these tokens to be as cheap as possible so that we can generate as many of them as possible to as many people as possible. That's what we've built and designed our architecture from the ground up to do.
John Furrier
>> And explain where you guys are at on the product side.
Darren Chien
>> Yeah, so we have our first generation Atlas. It's an FPGA product. It's live shipping to customers. We're going to end the year close to 100 million in revenue just on that alone. It's an FPGA, so it's flexible, allowed us to get fast time to market. But what we're really looking forward to is our second generation Asimov chip, which will be shipped in a Titan system, that's going to be really focused on massive memory capacity and really high memory bandwidth utilization.
John Furrier
>> And what are those constraints that Titan solves?
Darren Chien
>> So Titan really solves those two metrics, memory capacity and realized memory bandwidth utilization, where with memory capacity, the more memory capacity you have, the larger model you can fit, longer context length, models can remember more about you, and you can serve more users. And then in terms of memory bandwidth utilization, that's in terms of tokens per second, generate more tokens faster.
John Furrier
>> All right, so I got you here because you're good at explaining. So explain two concepts, the pre -fill and decode, and then context window and why that's important. So explain what pre -fill and decode is for the folks that are learning, because this is material to how things work.
Darren Chien
>> Yeah, so pre -fill and decode are the two different parts of inference, or the two different parts that make up inference. On the pre -fill side, you can think of that as when you put in, you copy and paste your entire code repository, you put it into ChatGPT. When ChatGPT has to read through that entire thing, it doesn't have to do it one at a time because you've uploaded the whole thing. So that's a massively parallel problem, which is compute bound.
John Furrier
>> So that's something that - It calculates what you're basically prompting.
Darren Chien
>> Correct, it can see the entire thing all at once. It can analyze it in parallel. So that's what we call the compute bound operation. So within inference, there's still compute bound operations. However, on the decode side, that's when after you've put in the GitHub repo, it's analyzed it all, it now needs to spit out one word or one token at a time, what its response is, that is decode. And that is memory bound. And so that's why you need a lot of memory capacity and memory bandwidth in order to make that output or the response back to you fast and accurate.
John Furrier
>> So to put it in simple terms, here's the request and then it thinks.
Darren Chien
>> Correct.
John Furrier
>> And that's the tokens and that's the context.
Darren Chien
>> Correct.
John Furrier
>> Great definition, because we hear a lot of that in these conversations, oh, I'm going to use AMD, I'm going to use Vera, we all know what a CPU is, we've been around PCs and servers.
Darren Chien
>> Correct.
John Furrier
>> But the GPU, which does all the matrix multiplication and all the high speed stuff, that's the engine for thinking.
Darren Chien
>> That's compute. So raw compute, really great for training, really great for pre -fill, but not so, on the decode side, it's more of a memory bound operation.
John Furrier
>> That's where you need, that's the memory we see in supply chain, Micron stocks up, everyone wants more memory.
Darren Chien
>> Correct.
John Furrier
>> And SRAM and you got solid state, SRAM, DRAM, HBM. Solid state, solid time is killing it, it's all these, okay, got it. Now context window, we hear that, and that's kind of gone mainstream, like my context window for this model, and most people relate it to a bigger prompt. I could put a Word document in there, GitHub, code repo, so that's stuffing in the prefill basically. Right?
Darren Chien
>> Correct.Well, context window refers to basically how much the model can remember about you. So if you remember using ChatGPT when it first came out, you would have to remind it, hey, I am Darren Chien, I'm running Positron AI's APAC business. You would have to tell it that every single time. Now when I talk to ChatGPT, it knows that about me. Hey, I don't need to explain it to you, just you know this about me, just shape your response around that.
John Furrier
>> Yeah, I do that a lot with ChatGPT. Based on my chat history, tell me where my head's at right now.
Darren Chien
>> You don't even need to tell it based on my chat history because it remembers.
John Furrier
>> Yeah, it's good because it has all my editorial in there. But okay, so now get into that, how will you guys see it? Because now you guys are attacking the large scale piece. Now, we just laid out the inference equation, context, windows, prefill, decode. What are you guys optimizing for? How should people think of Positron?
Darren Chien
>> Yeah, that's a great, great, great, great question. And our answer is we believe that inference is fundamentally an economics problem, right? We need to make sure that these tokens are generated as cheaply and as widely available to users as possible. And so the way we compete is on optimizing the two metrics of tokens per dollar and tokens per watt. There are other competitors who will win on throughput or interactivity, but we believe that inference is fundamentally an economics problem.
John Furrier
>> On the enterprise side, you mentioned enterprise. Scaling is a huge deal. They're paying a lot for tokens. We see a lot of budgets kind of on hold. We saw some IBM blip in the earnings that was due to some spend. Obviously supply chain had impact there, but people got surprised with these big token bills. everyone's using cloud. Oh, it's just when it says I'm done and hit upgrade and 10 ,000 people spend another 89 or 150 bucks times 10 ,000, it adds up super fast. How do you scale to the enterprise? Because they don't have the huge CapEx budgets. even financial services, they're not a $15 billion CapEx, NeoCloud / neocloud (ambiguous), although JPMorgan Chase has a big budget, But the enterprises have to deal with their pre -existing stuff. They have a lot of compute, they got a lot of Dell, HPE, Lenovo, NetApp, and all the storage in there, and they have all that old IT infrastructure. How do they turn on their systems? How do you see that scaling?
Darren Chien
>> And that's actually a fantastic question, because it's exactly the problem that Positron is solving, which is that you have these legacy enterprises, they have access to air-cooled data center space that they have existing infrastructure in. they might have a rack or two available, but when you're talking about these, the latest and greatest of NVIDIA GPUs, for example, or these rack scale solutions, which require liquid cooling, you're not going to be able to do that in your existing data center. Our Titan systems are designed by default to be air cooled. We can do liquid cooled if you want, but since they're designed by default to be air cooled, we can slot those into those spaces. We can sell it to you as a server. We can sell it to you as an entire rack. And it really just lowers the barrier to entry for these enterprises. they don't need a retrofitted data center, they don't need to go looking and competing with Meta. They can look at their existing portfolio and add net new compute where they otherwise may not have had.
John Furrier
>> I talk to Mitesh about all this all the time because he and I talk about scaling and some of the market stuff, but talk about the origin story. What's the culture like? You have other founders in there. What's the DNA? Where'd this all start? What's the, how do you do the TLDR of the origin story?
Darren Chien
>> So the TLDR is a bunch of smart people came together, looked at AI, looked at where it was going, thought to themselves, what could we do to design an AI chip to serve the workloads of the future, but designed today and not 20 years ago.
John Furrier
>> You know what I love about the story of the company is that it's not like, it wasn't 10 years old. This has just happened. So they're in the middle of it. So they have a clean sheet of paper. They see the inflection around 2022, 2023. That was inside the ropes. 2023. Everyone saw that.
Darren Chien
>> Yeah, 2023, our founders are all have semiconductor background, from background in Groq, Mitesh and myself background at Lambda. So we have a really unique blend of both the chip side, the cloud side, and basically the user facing side.
John Furrier
>> Yeah, I love this conversation. And one of the things we're seeing right now is the biggest build out I've ever seen in my lifetime. I think it's not going to stop because the demand curve is still there. In fact, Dave and I were arguing on our CUBE Pod about the bubble. I said, no, there's no bubble in this demand. Pricing might be just supply chain fluctuations, excitement. Obviously pricing goes up, but at the end of the day, if there's demand to fill, it's there. But when you look at what has to happen, they're standing up infrastructure. So there's demand for the components. Partnerships have changed in the go -to -market. People are partnering much deeper. It takes a village to put out these large scale systems, but once it's built, you got to operate them. So what are your thoughts on the monetization? because we're seeing a massive surge on the neoclouds. You're seeing enterprise now putting their toe in the water with coding, agentics coming in full throttle and physical AI is right behind it. That's the sequence of the waves coming. Yeah. They're monster waves.
Darren Chien
>> Yeah. And it's actually really incredible to see because both at Lambda and at Positron, your potential customers could also be your suppliers, but could also be your competitors. And it just speaks to the pace of the industry. The demand is so unlimited that we should work together at what we're best at, right? So we are the best at designing chips. We could work with a partner who's the best at making them, packaging them. We could work with a partner who's the best at deploying them, or we could work with a partner who's the best at deploying ocean-based data centers, right? It's not like we're going to go in and deploy ocean-based data centers, but we are happy to work with those.
John Furrier
>> Space is coming around the corner, too.
Darren Chien
>> Space as well.
John Furrier
>> Well, I guess my question is interesting because if you take it to the go-to-market and partnership side on the biz dev, on the business development side, it used to take six months to a year to get the engineering teams together. I'm hearing stories of a week, a month, scrum, jump on a Zoom, get face to face, get embedded. How has the partnership equation changed? Because you guys got, you stick to your knitting semi, then you say, okay, this guy does cloud native software, control plane software, here's a GSI to implement, here's a builder who has facility. There's a lot of people that need to be involved at a deep co-design level. How has that changed in the past cycle?
Darren Chien
>> Well, it's changed in that, like you mentioned, it's just shortened the time to make something happen, right? Before you'd plan, you would identify different sites, you would want to de-risk it as much as possible. But because we have strong partners like Oracle, who are really strong technically, really strong engineering, really strong data center footprint, really strong deployments, we are able to go in, speak their language, tell them, hey, this is what we are missing. Can you fill X, Y, Z? They tell us, oh, we have XYZ, we just need ABC. We just find a middle ground and work together.
John Furrier
>> And they know hardware too.
Darren Chien
>> They know hardware as well.
John Furrier
>> They know memory. It's almost interesting how it's full circle. It's a systems game.
Darren Chien
>> It is.
John Furrier
>> All right, final question. APAC, which is your focus, what's going on there? Obviously we see a lot on Taiwan. Obviously TSMC's there. What else is going on? We're seeing telecom starting to get hyper-converged. We're starting to see a lot more in -country sovereignty is a huge discussion. Just what's the macro like in APAC right now?
Darren Chien
>> Well, we saw the open model support letter by Jensen the other day. All the companies in the APAC block signed on supporting that. I see that increasing, right? Because open models just distribute accessibility and availability of intelligence across the world. Speaking specifically in APAC, Taiwan, all roads to Silicon lead to Taiwan, right? We are investing heavily in the country. We have our strong investor base there, strong partnership base there. In terms of customers, right, we talked about air -cooled space, legacy data center space. That is plenty in APAC. There's something like nine to 10 gigawatts of air -cooled space that is waiting to be monetized in APAC.
John Furrier
>> You know what I love about this market, Darren, is that everyone talks about NVIDIA versus AMD. You got to give AMD an assist on that forcing function for Jensen's and everyone else's coalition around the open weights and open source, because they play the open card. So NVIDIA, hey, we're open too. So, and then that helps everyone. You have this massive rise. So it's really not about one company versus the other because everybody wins because the demand curve. What's your reaction to that? Do you believe that same thing?
Darren Chien
>> I love it. I love seeing open models, seeing Qwen3, Kimi K2, the open weights were just released. They had to shut, they had to turn away users because of how much demand there was. And you know what that does? That forces Anthropic to allow Claude to be more widely accessible. That forces ChatGPT to allow Sora to be widely available. If it was just one company there, they called the shots, they set the terms, you will have one model and you will like it whether you like it or not. We have a plethora of optionality and availability that is just better for us as everyday users.
John Furrier
>> And it speaks to capitalism and competition. With no competition, if AMD's not putting the pressure on NVIDIA, you guys don't come in with a new chip. Competition, this is where it's healthy.
Darren Chien
>> 100%, we saw actually that competition forces, whether it's conversations around security where we saw actually GLM, a Chinese open model, being deployed by Hugging Face to fight off the rogue ChatGPT hacking incident, right? It's that type of thing where when there are more widely distributed, more available models to everyone, it's better for the entire ecosystem.
John Furrier
>> If you had one player, we'd be basically still doing marketing documents and email writing. There's no motivation to innovate if there's no competition.
Darren Chien
>> Correct.
John Furrier
>> Look at NVIDIA, they're the leader. Look at you guys coming out in 2023, guns blazing, inference is the hottest product. Constraints drive engineering, drive entrepreneurship, competition drives action.
Darren Chien
>> Correct.
John Furrier
>> Congratulations, great to see you. Say hello to Mitesh and the team. Thanks for coming on. Third annual, we'll see you tonight. All right, I'm John Furrier here in theCUBE Studios in Palo Alto for the NYSE Wired. Third annual CUBE Pool Party tonight, but all day today talking to the leaders, making it happen. Competition, openness, that creates more innovation faster. At the end of the day, you're reining in the chaos. And at the end of the day, you have more innovation doing our part here at theCUBE. Thanks for watching.