Jason Goodison of General Compute joins hosts John Furrier and Dave Vellante on theCUBE Research’s NYSE Wired artificial intelligence, AI Factory series to explain General Compute’s role in AI infrastructure deployment as a deployment partner for alternative accelerators. Goodison discusses deploying heterogeneous application-specific integrated circuits, ASICs for enterprise, cloud and edge customers. They cover model bring-up, software stacks and partnerships that enable bare-metal rentals and single-tenant service-level agreements, SLA for a widening set of inference and training workloads.
Goodison highlights technical drivers and market dynamics, and they emphasize the structural advantage of memory-heavy on-chip architectures and the operational benefits of splitting inference into pre-fill and decode workloads. They note that financing and supply constraints favor GPUs today, while specialized deployers are needed to mitigate ASIC depreciation and bring diverse silicon to production. theCUBE analysts underscore growing demand for heterogeneous stacks and new financing models that support deployment of diverse accelerators.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, host of theCUBE, here with Dave Vellante, my co-host. Hello, I'm John Furrier, host of theCUBE, here in our Palo Alto studio. Of course, we have theCUBE's NYSE studio connecting Silicon Valley to Wall Street, Wall Street to Silicon Valley. The NYSE Wired program is where technology and Wall Street intersect, of course. theCUBE's got the deep coverage tracking all the semis, all the AI infrastructure buildouts. This is our AI Factory series. We talk to the leaders who are building it and setting the table for the era of AI. Jason Goodison's here. He's the CTO and co-founder of General Compute. Thanks for coming in, popping into the studio today. Appreciate it.
Jason Goodison
>> Thank you for having me.
John Furrier
>> You're doing some pretty cool work. Again, the demand for AI infrastructure is off the charts, and again, the demand for intelligence is feeding in. Obviously the coding, which we're— everyone's seeing the value there. GenAI, physical AI edge right in line. So you can see the progression, the trajectories forming. There's just way too much demand. But the role of the ASICs and the silicon and software, we're seeing the different approaches from NVIDIA, Google, Cerebras, AMD all have kind of their bets.
Jason Goodison
>> Yeah.
John Furrier
>> CUDA, obviously with NVIDIA, that's the software evolution, programmable TPUs for Google. They got a different approach with the pods. You got Cerebras, the big wafer, the general purpose to custom and specialized compute have always been that spectrum. The harder you go to specialize, the harder it is to program. The use cases are more narrow. But with AI, that's all being bundled together. You're building out your venture and in this demand curve, A lot's going on. I want to unpack with you, but let's get into what you guys are doing right now. What's the state of your company? What's the thesis?
Jason Goodison
>> Yeah.
John Furrier
>> Where are you guys seeing the action?
Jason Goodison
>> Yeah. So fundamentally what we do is we like to see ourselves as the deployment arm for these alternative chips, alternatives to GPUs. So if you think about all of these up-and-coming chips, you've got Cerebras, SambaNova. those companies have been around a long time. You've got TensorDyn, Tenstorrent, Positron, d-Matrix. There's just— the list goes on.
John Furrier
>> Loads of new stuff coming.
Jason Goodison
>> And all of these engineers are just fantastic, right? They've built these incredible chips that work for different use cases. But we have a lot of the asset-heavy clouds of the world, the Nebius, the CoreWeave that are kind of locked into their chip supplier. So you think about CoreWeave, they have a lot of circular financing with NVIDIA. And so they only deploy NVIDIA. TensorWave only really deploys AMD, they might do other stuff in the future, I don't know. And then Fluidstack does a lot with Google TPU. So there's all of these other chips in the world that are really fantastic and the engineering teams there are exceptional. But there's no one that just takes on that debt and deploys them and then rents them out bare metal to an enterprise or to another cloud that needs the capacity. So that's what we do. So a few weeks ago we just announced a $400 million debt facility in combination with Upper90. Upper90, joined the cap table and we're really excited to work with them and productionize a lot of these awesome chips.
John Furrier
>> Yeah, there's a race. you can see the vertical integration, CoreWeave, Nscale, they just bought Anyscale. So you can start to see the disaggregated serving kicking in here at Hot Chips Stanford. I was poking around yesterday, top conversations, hey, the capability and capacity demand is high. The KV caches are getting stuffed. So you start to see the patterns. More data is coming in, more demand for the tokens, aka intelligence. So it's going to put pressure on the architecture. So, okay, the big guys, they got their partners, but there's a whole nother onboarding of these new Neo Clouds, Neo Labs, or someone's got some Bitcoin, they got some data center facilities. Those are different businesses, but they got the energy, they got the footprint. They want to bring that into this new build out. So it's almost as if there's a new breed. Of infrastructure opportunity. Sounds like that's what you're targeting.
Jason Goodison
>> Yeah, absolutely. So when you think about any major technological innovation, you start off by not really understanding the problem. And so you just use what you have at your disposal to solve the problem. So even when you think about Bitcoin back in the day, people just mined Bitcoin on GPUs. Nowadays, people mine Bitcoin on ASICs because as the workload stabilizes, you just get so many advantages to baking that into the silicon, right? And so what we're seeing, we're a little bit in the Wild West period right now in AI because there's a new model drop every month. An open source model comes out. There's a new attention mechanism. There's different infrastructure and architectures that are coming in the AI models. So we have all these chips and they're all racing. They all have their own bets that they made often 4, 5, 6 years ago when they were just designing the silicon for the first time. And it's not clear who is going to have the best structural advantage because we don't know how the architectures are shaping up. What I will say though is it seems overwhelmingly like we're moving to this world where memory-heavy chips have a fundamental advantage, chips that are able to keep the data on the chip. So think about all these data flow chips that don't have to go back and forth to HBM all the time. SambaNova was a great example of that. Those chips that can keep the memory on the chip and are reducing how much they have to move memory around seem to have a very structural advantage right now. And you're starting to see those chips do well.
John Furrier
>> Yeah, I just wrote a post on LinkedIn because I get this question all the time. Hey, what's the difference between NVIDIA and Google and everyone else? And Hot Chips again is highlighting this kind of where the engineering focus is. Mathematics isn't everything in this world, right? So the math is commoditizing.
Jason Goodison
>> Yeah.
John Furrier
>> But the data feeding the math is becoming the key value, which is why you're seeing these different approaches. And when you start bringing data into the equation, you bring in governance and compliance. Yeah. Or routing or other technical features. Talk about that piece because there are different approaches. Yeah. why should I send that data there when I can send it over there? The gap between capability and control becomes interesting and becomes a technical problem, not just a governance problem. Share your thoughts and vision on this because this seems to be one of the key areas where there's a technical opportunity to architect something that can give you the capabilities and capacity with control. The data flow.
Jason Goodison
>> Absolutely. so think about this. Let's say you're an enterprise and you have this proprietary data and this is essentially as the world is moving to infinite intelligence, you could call it. This is essentially your moat. This is your IP. You could work with an Anthropic or an OpenAI, but every single time you make an API request, you are sending that data to them. And so even if you trust those companies, there is some risk that, hey, in the future it might not be the same governance at Anthropic, but they still have all the data you've sent them. So even if you trust them now, you have to think about long term. How do I want to protect my company as intelligence becomes infinite? And one of the ways to do that is open source. So all of these open source, open weight models that are coming out often from China, and NVIDIA is gonna do some stuff on that soon too. A lot of companies wanna own the racks that these open weight models are running on. And so basically the data goes in and it comes out and they own the entire stack. They don't have to worry about someone's gonna take my moat, someone's gonna train a new model based on the proprietary data from my company. I'm not sure if that answered your question.
John Furrier
>> Well, if you look at CUDA, right? CUDA's whole thing from NVIDIA is programmability.
Jason Goodison
>> Yep.
John Furrier
>> So ASICs take huge cycles to get the next one going. So having programmability— True. —in the stack, how do you guys look at that? Because you have, like I said, CoreWeave, Nscale, they vertically integrate, they provide some SLA and some services. And then you got the approaches, hey, we're just going to provide raw intelligence. And feed that up to whoever wants it. Yeah. I wouldn't call it headless. I hate that word in this capacity, but there's a retail side of this business, which is developers. Yeah. There's almost no AWS for this world.
Jason Goodison
>> Absolutely. So, a few things there. So a lot of the ecosystem is moving towards open source. So you'll see a lot of these companies come out and they'll say, hey, we're going to be the software stack for all heterogeneous hardware and you compile kernels once on our stack and it works on all the different architectures. But imagine for a second, let's say you're SambaNova or you're Cerebras, a GPU is going to do part of the problem very well, pre-fill they call it, GPU does exceptionally. And these ASICs do decode exceptionally well. So if you pair them together for one solution, you get 40% better TCO, but you also get up to 10, 20 times faster AI. So it makes sense to do it, right? All of these decode silicon, we call them internally, it's a bit technical, but all of these decode silicons, they know that it's life or death, that if they can actually build a software ecosystem internally that integrates well with vLLM and SGLang and all of these inference engines, that's the whole business. They're all working on that right now. They know that. So there's no lack of people trying that. And then when you talk about model bring up too, it's a really good question because some of the compilers for these ASICs, they use different programming paradigms, right? And so the compilers can be extremely complex. So one thing that we're doing at General Compute is we actually hire model bring up experts from all of the different ASIC companies. someone starting tomorrow was at NVIDIA and AMD and Meta. We have someone that was at Cerebras and SambaNova starting in a week. And so we bring all of those experts internally and we work with the companies to actually build out a better software stack and bring up models internally. And then we can offer SLAs on them too, because nobody wants to rent a machine for, $100,000 a month. And then it turned out to be a brick because they can't actually run the models that even matter in the first place.
John Furrier
>> Or the market shifted. So we saw that with training to inference, great clusters for training. They didn't have the pre-fill decode problem. They're just training. Exactly. Inference gets interesting. And again, my takeaway from Hot Chips so far is that, and I've been kind of circling around this one, I get your reaction to that the whole disaggregated serving concept is not so much vertical integration, it's just more efficiency. Exactly. And talk about the reasons, because this is a very nuanced technical point, but pre-fill and decode do something very good when you separate them because of the demand and complexity of the pre-fill. Yeah. Now the KV cache has to get smarter. It gets fatter, gets bigger, bloats up a little bit. Talk about what this means because this demand curve's not going away. Yeah. So talk about this pre-fill decode dynamic and it's not so much up the stack, it's more of really around the resource.
Jason Goodison
>> Yeah. So fundamentally when you think of inference, so that's when you actually ask the AI a question and it spits out an answer, you can actually split that into two workloads. One of them is called pre-fill. That one is about prompt processing. So let's say I've asked a really long question. I have 100,000 tokens in that question. Those tokens or those words need to be turned into numbers to actually run through the math. It's math, right? Those have to be turned into numbers. And then decode is when I'm autoregressively, they call it, which is generating one token at a time, right? Those are actually two different problems. And so they're taxing the models, right?
John Furrier
>> The decode taxes the models differently in math terms.
Jason Goodison
>> Right, right. And so it takes what came out of pre-fill and then it uses that in the autoregressive fashion to generate tokens one at a time. And so if you think about a GPU, fundamentally it's a graphics processing unit. And if you think, what do you need if you're generating graphics? Well, imagine you had one core, just for simplicity's sake, mapped to one pixel on the screen, and you're like, I need to know what color this pixel should be, 300 times a second. Well, actually it's really easy if I do a bunch of matrix multiplication for one pixel and then I just have a bunch of different cores and I do it all independently. I do it all at the same time in parallel, right? And so actually prefill for these AI models is something kind of similar where I can process all of the prefill or the context window, all of the prompt, each word I can process independently. So it maps really well onto that GPU, right? Because I'm just doing a bunch of matrix multiplications. Maps one-to-one, perfect. With decode, I take that prefill, they call it KV cache, right? That's just the context but in numbers. Well, now I have to store that right next to the actual silicon as well as the weights of the model. And the weights of the model are huge now. We're getting up to 3 trillion, 5 trillion parameter models. And then also a million context length windows. And so if I want to process 10 people, 20, 100 people at the same time, I've got 100 KV caches and I also have the model weights. And so it just blows up the memory problem. And so look, I can do all of my pre-fill independently on a GPU, maps perfectly. The decode though is where the GPU really breaks down.
John Furrier
>> talk about the consequences. Again, we're getting in the weeds a little bit here, but I think it's important for people to understand that if you screw up the KV cache, You got to reboot everything. So as you're getting through the multi-step processing, whether you're doing it through some sort of systolic array or whatever matrix multiplication, whatever process you're using, which is very layered, very complex, it screws up if you have to reset. Yeah. And you start that, and that's a GPU monolithic problem. Yeah. So the answer is, okay, put some compute here. And I think people misunderstood Cerebras when they first started, when I first met the team years ago. They're like, oh, it'll never work. Inference, it's just purpose-built inference. And it turns out they took on the memory wall. Good bet. Yeah, that was a good bet for Cerebras. But now guess what? You can integrate that in. Yeah, it's not a standalone. So you're starting to see the architecture approach is different. What's your technical view on this? Because it's kind of like a systems architecture game. Yeah, not a I got NVIDIA for this or Cerebras for that or SambaNova for this.
Jason Goodison
>> It's more than an architecture game because even if you find the best chip in the world, can that actually scale to production capacity? Because you know very well that there's constraints at TSMC, there's constraints with Micron or SK Hynix or whoever your memory provider is. So let's say $250 billion of NVIDIA silicon gets deployed next year. It could be double, it could be triple. And then the ASICs market might only be able to deploy like $1 billion. Right. And so it's a small market, but it's growing very quickly and it'll be more than $1 billion, but call it even $10 billion. It's a fraction of what NVIDIA is going to do, but it's growing very quickly. But you had to create those alliances with Broadcom or Intel, which is what SambaNova is doing early so that you could actually build the production capacity. Right. So it's that. It's also the fact that the memory problem is deeper than you would think because let's say you've got Cerebras. Cerebras has 44 gigabytes of SRAM. So they have that one big wafer and they put all of the memory right into it. Well, models are bigger than 44 gigabytes. So now we have to think about how are we going to wire these things together to actually split the model across a bunch of wafers,
John Furrier
>> right?
John Furrier
>> And it kicks ass at inference too. So it's a great use case. So you plug— wire it up. It's a great use
Jason Goodison
>> case.
Jason Goodison
>> But I guess what I'm trying to get to is that the bigger the model, it might actually change what chip you want to use too, right? So there's going to be different chips that are actually going to excel at different model types and different architectures. And we're just seeing an explosion of AI models. So it's not clear which one is going to win. It's getting bigger.
John Furrier
>> It's just the beginning. Let me ask you a question on the— question I get a lot, which is, hey, what's going on with all these new ASICs coming out? We see Positron, d-Matrix. So some custom accelerator market, ASICs, but they're not just accelerators, they're playing another role. What's the view of the market as these new entrants come in with this demand curve? How do you think that's going to play out? New formation? How does— because you guys are doing this. This is what you're doing. Yeah, we're doing it. How is that going to play out?
Jason Goodison
>> What's your vision of this? Well, I think people are really curious about ASICs ever since the OpenAI Cerebras deal and the IPO of Cerebras. People are starting to accept it. So when we started the company, this was pre all of that. And so we told people, we said Cerebras is going to be a big deal. And people would pooh-pooh it.
John Furrier
>> Yeah, they would pooh-pooh that. I've heard people, it'll never work. It's too big. It's never done before. Yeah.
Jason Goodison
>> And industry leaders and experts would do that. we would kind of get snubbed to some degree. But now everybody is on board, they understand it. And so we're even seeing the demand from enterprise of like, hey I know you've showed us this chip and you taught us how this chip works, but what about that chip? what about Etched? What about Cerebras? What about this and that? And so we just bring— we kind of connect the supply to the demand and we deploy it. And so you can take on less risk. If you're an enterprise and you really want to try a Cerebras or an Etched or any of these companies, you could either, pay millions of dollars and hopefully be able to actually deploy it and manage it
John Furrier
>> correctly.
John Furrier
>> want
Jason Goodison
>> to.
Jason Goodison
>> They don't want to take the risk. It's a huge
John Furrier
>> risk.
John Furrier
>> So are you targeting them for customers? We're targeting everyone that wants fast. All right. What do you guys— let's talk about your momentum. Take a minute to explain the momentum you have. Of course, there's a macro trend that's your friend. So that's going to be good for you guys. I think there's going to be a whole other class of build-out companies. I think you're a highlight of that. What happens next? The enterprise, they don't have billions in CapEx now. They'll use services. They'll do on-prem. They'll put in maybe a smaller cluster with FPGAs or some other low-cost, high-performance configuration and connect to a service that can give them single-tenant-like capability. Of course. That makes total sense to me. What are you guys targeting?
Jason Goodison
>> Well, there's 3 kinds of customer profiles, right? One of them would be you've got the frontier labs. The frontier labs just need a ridiculous amount of capacity. They're going to deploy some of their own. They're going to rent some from other people. FluidStack does a lot of stuff with Google and Anthropic. There's a big market there, but they care about training and they care about inference and they're experts at model compilation and everything. They have their own people on staff and they're just, they have infinite pockets to just deal with these problems. Then you have AI application companies and enterprise. So call it Cursor, Perplexity of the world, the OpenCodes of the world. They have a lot of demand, token demand. They're actually more comfortable most of the time paying for an SLA. So they're saying, I want to have X number of tokens per second. I want to be able to process X number of requests per second, things like that. And then the third category is the asset-light inference cloud. So for example, you've got Baseten, the Fireworks, Together AI. Some of these guys are starting to move down the stack, but they have essentially infinite demand, right? So any capacity that becomes available to them, they're gonna be able to connect that to a buyer. And so they own a lot of the end-user customer relationships already, and their end users are growing so fast, they just need more supply. And they're also, because they're asset-light, they're not locked into any—
John Furrier
>> Asset-light meaning they're not spending a lot of CapEx to do it and scale and sort of— Correct.
Jason Goodison
>> Before we did. Yeah, sorry, I should have explained that. When we say asset-light, we mean that they're not deploying, buying hardware and deploying it themselves, but they're renting capacity from other people, right? And then they're on-selling that and they're making their own SLAs and they have their own value-added services on top of that.
John Furrier
>> So they have— Fireworks is kicking ass. Everyone looks at those numbers, they're killing it. They're doing great. I think that's a big market. There's two approaches that I call the vertically integrated and then okay, meet that asset light demand, which is I got customers. Yeah, I will qualify what you have and then integrate in.
Jason Goodison
>> Yeah. Just a thought experiment. Let's say you raised $100 million of equity right now. You could either go and buy X number of machines and deploy them, or you could rent like 5 times X number of machines and then you could make 5x the revenue. So it makes a lot of sense for everyone to kind of pick what they're experts at and then do that. And we have a lot of these asset-light guys that are doing fantastic.
John Furrier
>> I really like what you guys are doing. In March at GTC, we saw all the Pareto curves. Jensen did his thing. I then wrote a post the next month, April, because Jensen's like, we're bounded by energy, the 5-layer cake he puts out there, which is totally legit. I wrote a post that, no, no, it's bounded by energy and money. That was the first post that kind of went out and was about a company and we featured Argentum AI, which was trying to figure out the financial code because as you pointed out, this has been documented on Bloomberg and other places, the circular financing, which I think some people try to throw shade on NVIDIA, but they're just doing a great job to help build the infrastructure. Talk about the financing aspect, though, because you're taking an approach to bet on the asset-light market that's in demand. Yep. And that financing— you got to get facilities, energy, and then you got to get the finance. Talk about this funding function of finance. Yeah, because then fast forward to this month, Jensen was in New York with the CEO of Goldman, KKR, $500 billion. They're taking care of the physical plant, my word. But the physical buildout. But there's still now a financial market developing. You're in the middle of this. Yeah. What's your vision on how that plays out? Because risk management is now in play on both sides.
Jason Goodison
>> Yeah, it's 100% true. There's a very well-understood debt market for GPUs. So if you want to go buy a bunch of GPUs, you can raise debt. And what NVIDIA will do is they'll underwrite the purchase of it. So they'll say, hey, if you cannot rent these machines or sell tokens on these machines, we'll rent them back for you. And what that allows you to do is go to the banks and say, hey, look, this is basically a guaranteed deal.
John Furrier
>> It's a AAA bond right there. It's like— and it's sometimes pledged. Yeah. So that's like every bank's like, I'm in. Exactly.
Jason Goodison
>> And we're talking about the company that's worth, over $4 trillion, they're not going to default. They're going to come through. And there's also infinite demand. So you will get a customer. Taking it a step farther, you talk about an ASIC. It's like, well, nobody really understands what the depreciation lifecycle on that ASIC is. Nobody really understands the residual value of that ASIC after the end of its life. There's no secondary market for it, because people haven't really adopted it or diffused it into the economy yet. So, it's fundamentally a much harder thing to get people to bet on. You're probably going to get quite a low interest rate if you're deploying GPUs. If you're deploying an ASIC, it's going to be much higher and you're going to have a much smaller pool of capital to do it from. So that is fundamentally what our business solves. A lot of these ASIC companies, I'm sure you chat with them here all the time, just absolutely exceptional people, like
John Furrier
>> incredible—
John Furrier
>> Great tech and there's demand for what they have.
Jason Goodison
>> There's demand for what they have and they've spent so much time on the technology, but the deployment piece is actually how NVIDIA is running circles around them. NVIDIA has great tech too, but it's not as good as a lot of these other ASIC companies, but they're adopted way more. And CUDA is not a good enough excuse anymore. the ecosystem is opening up. You can write your own compiler and AI
John Furrier
>> can—
John Furrier
>> I mean, CUDA is just a software model that makes the ASICs last longer until the next rev. So you can level up if the market changes. Whatever nuance is key, that could be replicated, baked into the platform.
Jason Goodison
>> It can be replicated for sure, especially with AI coding capabilities. You should be able to get something at least workable. But the reason that they're not being diffused more into the economy is the fact that there's just no one deploying them. And that's why we're stepping in.
John Furrier
>> Well, Jason, I'm really jealous of you. You're a young gun. I'm aging out over the years. You've got a long runway here. But you brought up the depreciation things because Jensen said something Dave Vellante and I were— and Brian were talking about is the analysts haven't modeled, and he put it in kind of quotes, they haven't spreadsheeted out what this is going to be. So a lot of people don't know what depreciation means because they don't know what the reuse is. So if we assume scarcity, architecturally smart engineers are using older chips Talk about that from a tech perspective, because the old idea was, oh, that's a chip. The next one comes out, the value drops. You can depreciate. That makes total sense in the old way. But in the new world where you have diversity of clusters, you have diversity of capabilities, there's a reuse market that keeps the prices up, which changes the modeling on the financial spreadsheets of valuation. So what's your view on this? It's kind of a random question, but it's one that everyone's asking, well, can you hedge that? Well, there's a futures market, but then again, if it's depreciating. So there's a whole conversation around the thesis of will the hardware and software be worthless, more or less in the future, or will it have staying power and durability? If you assume, okay, big clusters, small clusters, edge, physical AI, a chip today could be put into a robot maybe. Yeah, there's all kinds of supply chain functionality discussion.
Jason Goodison
>> I think you're thinking about it the right way. I'm sure you saw the deal with CoreWeave where they signed A100s through, I believe, 2029. And that is a very old chip at this point. And I think what's happening is you're in a supply constrained market. Some workloads are more valuable than others. And let's say I could run on an A100, and I'm making these numbers up, but at 10 tokens a second, or I could run on a, NVIDIA Cerebras combination with, 2,000 tokens per second. I'm obviously going to put my high-value workloads on the Cerebras rack, but there's probably a bunch of stuff I could just put on the A100 overnight and not think about, probably internal things, right? So there, I think it depends.
John Furrier
>> It's a TCO calculation at that point.
Jason Goodison
>> It is. It's like, what am I running?
John Furrier
>> It's policy-based, it's resource-based. Exactly. Intelligence could manage that. You put some AI in there. Yeah.
Jason Goodison
>> Well, I'm also always thinking about too, okay, what is the revenue per megawatt you can get? So let's say you've got an A100 and you have a megawatt of it deployed. Theoretically, you can make X amount on it. And then if you could upgrade to a new Blackwell generation or the Vera Rubin generation, you'd have to rework and put CapEx into the facility in order to actually be able to run those machines. But you'd get, I don't know, X times 10 or X times 100 in revenue. So I think the calculation there is really interesting. But the fact is, look, we're all so supply constrained that everything that's in production right now, we're just going to use it. And as we upgrade and build new data centers and upgrade old data centers, we will plug in new stuff and things will depreciate. And I don't think people will be using A100s forever.
John Furrier
>> There's just not enough sample size, Jason, on this. So it's— I think it's an open question. I think that's going to be one we're going to watch certainly in the middle of it. All right, final question. What are you optimizing for now? Give us a taste of what's coming. I know you got some deals brewing you can't talk about right now. What's going on? Set the direction.
Jason Goodison
>> Where's the company heading? The company is headed towards being the heterogeneous ASIC deployment arm. So everything that is not already being handled by your CoreWeave and your Nebius, There's a lot of fantastic chips out there. Everyone wants to try them. They have different use cases. We are going to be deploying those for customers and we'll be doing some announcements in the next few weeks, I believe. And I'm really excited to talk about that.
John Furrier
>> Maybe I can come back and we can talk about that. Yeah, we'll definitely do it. General Compute, not doing general purpose computing as we know it. General Compute is providing the scale for what we see as a democratization on the ASIC side as more entrants come in, more capabilities, again, more infrastructure demand continues to thunder away. I'm John Furrier, your host of theCUBE. Thanks for watching.
>> Palo Alto Studio Connection, Silicon Valley and Wall Street. I'm John Furrier, host of theCUBE, here with Dave Vellante, my co-host. Hello, I'm John Furrier, host of theCUBE, here in our Palo Alto studio. Of course, we have theCUBE's NYSE studio connecting Silicon Valley to Wall Street, Wall Street to Silicon Valley. The NYSE Wired program is where technology and Wall Street intersect, of course. theCUBE's got the deep coverage tracking all the semis, all the AI infrastructure buildouts. This is our AI Factory series. We talk to the leaders who are building it and setting the table for the era of AI. Jason Goodison's here. He's the CTO and co-founder of General Compute. Thanks for coming in, popping into the studio today. Appreciate it.
Jason Goodison
>> Thank you for having me.
John Furrier
>> You're doing some pretty cool work. Again, the demand for AI infrastructure is off the charts, and again, the demand for intelligence is feeding in. Obviously the coding, which we're— everyone's seeing the value there. GenAI, physical AI edge right in line. So you can see the progression, the trajectories forming. There's just way too much demand. But the role of the ASICs and the silicon and software, we're seeing the different approaches from NVIDIA, Google, Cerebras, AMD all have kind of their bets.
Jason Goodison
>> Yeah.
John Furrier
>> CUDA, obviously with NVIDIA, that's the software evolution, programmable TPUs for Google. They got a different approach with the pods. You got Cerebras, the big wafer, the general purpose to custom and specialized compute have always been that spectrum. The harder you go to specialize, the harder it is to program. The use cases are more narrow. But with AI, that's all being bundled together. You're building out your venture and in this demand curve, A lot's going on. I want to unpack with you, but let's get into what you guys are doing right now. What's the state of your company? What's the thesis?
Jason Goodison
>> Yeah.
John Furrier
>> Where are you guys seeing the action?
Jason Goodison
>> Yeah. So fundamentally what we do is we like to see ourselves as the deployment arm for these alternative chips, alternatives to GPUs. So if you think about all of these up-and-coming chips, you've got Cerebras, SambaNova. those companies have been around a long time. You've got TensorDyn, Tenstorrent, Positron, d-Matrix. There's just— the list goes on.
John Furrier
>> Loads of new stuff coming.
Jason Goodison
>> And all of these engineers are just fantastic, right? They've built these incredible chips that work for different use cases. But we have a lot of the asset-heavy clouds of the world, the Nebius, the CoreWeave that are kind of locked into their chip supplier. So you think about CoreWeave, they have a lot of circular financing with NVIDIA. And so they only deploy NVIDIA. TensorWave only really deploys AMD, they might do other stuff in the future, I don't know. And then Fluidstack does a lot with Google TPU. So there's all of these other chips in the world that are really fantastic and the engineering teams there are exceptional. But there's no one that just takes on that debt and deploys them and then rents them out bare metal to an enterprise or to another cloud that needs the capacity. So that's what we do. So a few weeks ago we just announced a $400 million debt facility in combination with Upper90. Upper90, joined the cap table and we're really excited to work with them and productionize a lot of these awesome chips.
John Furrier
>> Yeah, there's a race. you can see the vertical integration, CoreWeave, Nscale, they just bought Anyscale. So you can start to see the disaggregated serving kicking in here at Hot Chips Stanford. I was poking around yesterday, top conversations, hey, the capability and capacity demand is high. The KV caches are getting stuffed. So you start to see the patterns. More data is coming in, more demand for the tokens, aka intelligence. So it's going to put pressure on the architecture. So, okay, the big guys, they got their partners, but there's a whole nother onboarding of these new Neo Clouds, Neo Labs, or someone's got some Bitcoin, they got some data center facilities. Those are different businesses, but they got the energy, they got the footprint. They want to bring that into this new build out. So it's almost as if there's a new breed. Of infrastructure opportunity. Sounds like that's what you're targeting.
Jason Goodison
>> Yeah, absolutely. So when you think about any major technological innovation, you start off by not really understanding the problem. And so you just use what you have at your disposal to solve the problem. So even when you think about Bitcoin back in the day, people just mined Bitcoin on GPUs. Nowadays, people mine Bitcoin on ASICs because as the workload stabilizes, you just get so many advantages to baking that into the silicon, right? And so what we're seeing, we're a little bit in the Wild West period right now in AI because there's a new model drop every month. An open source model comes out. There's a new attention mechanism. There's different infrastructure and architectures that are coming in the AI models. So we have all these chips and they're all racing. They all have their own bets that they made often 4, 5, 6 years ago when they were just designing the silicon for the first time. And it's not clear who is going to have the best structural advantage because we don't know how the architectures are shaping up. What I will say though is it seems overwhelmingly like we're moving to this world where memory-heavy chips have a fundamental advantage, chips that are able to keep the data on the chip. So think about all these data flow chips that don't have to go back and forth to HBM all the time. SambaNova was a great example of that. Those chips that can keep the memory on the chip and are reducing how much they have to move memory around seem to have a very structural advantage right now. And you're starting to see those chips do well.
John Furrier
>> Yeah, I just wrote a post on LinkedIn because I get this question all the time. Hey, what's the difference between NVIDIA and Google and everyone else? And Hot Chips again is highlighting this kind of where the engineering focus is. Mathematics isn't everything in this world, right? So the math is commoditizing.
Jason Goodison
>> Yeah.
John Furrier
>> But the data feeding the math is becoming the key value, which is why you're seeing these different approaches. And when you start bringing data into the equation, you bring in governance and compliance. Yeah. Or routing or other technical features. Talk about that piece because there are different approaches. Yeah. why should I send that data there when I can send it over there? The gap between capability and control becomes interesting and becomes a technical problem, not just a governance problem. Share your thoughts and vision on this because this seems to be one of the key areas where there's a technical opportunity to architect something that can give you the capabilities and capacity with control. The data flow.
Jason Goodison
>> Absolutely. so think about this. Let's say you're an enterprise and you have this proprietary data and this is essentially as the world is moving to infinite intelligence, you could call it. This is essentially your moat. This is your IP. You could work with an Anthropic or an OpenAI, but every single time you make an API request, you are sending that data to them. And so even if you trust those companies, there is some risk that, hey, in the future it might not be the same governance at Anthropic, but they still have all the data you've sent them. So even if you trust them now, you have to think about long term. How do I want to protect my company as intelligence becomes infinite? And one of the ways to do that is open source. So all of these open source, open weight models that are coming out often from China, and NVIDIA is gonna do some stuff on that soon too. A lot of companies wanna own the racks that these open weight models are running on. And so basically the data goes in and it comes out and they own the entire stack. They don't have to worry about someone's gonna take my moat, someone's gonna train a new model based on the proprietary data from my company. I'm not sure if that answered your question.
John Furrier
>> Well, if you look at CUDA, right? CUDA's whole thing from NVIDIA is programmability.
Jason Goodison
>> Yep.
John Furrier
>> So ASICs take huge cycles to get the next one going. So having programmability— True. —in the stack, how do you guys look at that? Because you have, like I said, CoreWeave, Nscale, they vertically integrate, they provide some SLA and some services. And then you got the approaches, hey, we're just going to provide raw intelligence. And feed that up to whoever wants it. Yeah. I wouldn't call it headless. I hate that word in this capacity, but there's a retail side of this business, which is developers. Yeah. There's almost no AWS for this world.
Jason Goodison
>> Absolutely. So, a few things there. So a lot of the ecosystem is moving towards open source. So you'll see a lot of these companies come out and they'll say, hey, we're going to be the software stack for all heterogeneous hardware and you compile kernels once on our stack and it works on all the different architectures. But imagine for a second, let's say you're SambaNova or you're Cerebras, a GPU is going to do part of the problem very well, pre-fill they call it, GPU does exceptionally. And these ASICs do decode exceptionally well. So if you pair them together for one solution, you get 40% better TCO, but you also get up to 10, 20 times faster AI. So it makes sense to do it, right? All of these decode silicon, we call them internally, it's a bit technical, but all of these decode silicons, they know that it's life or death, that if they can actually build a software ecosystem internally that integrates well with vLLM and SGLang and all of these inference engines, that's the whole business. They're all working on that right now. They know that. So there's no lack of people trying that. And then when you talk about model bring up too, it's a really good question because some of the compilers for these ASICs, they use different programming paradigms, right? And so the compilers can be extremely complex. So one thing that we're doing at General Compute is we actually hire model bring up experts from all of the different ASIC companies. someone starting tomorrow was at NVIDIA and AMD and Meta. We have someone that was at Cerebras and SambaNova starting in a week. And so we bring all of those experts internally and we work with the companies to actually build out a better software stack and bring up models internally. And then we can offer SLAs on them too, because nobody wants to rent a machine for, $100,000 a month. And then it turned out to be a brick because they can't actually run the models that even matter in the first place.
John Furrier
>> Or the market shifted. So we saw that with training to inference, great clusters for training. They didn't have the pre-fill decode problem. They're just training. Exactly. Inference gets interesting. And again, my takeaway from Hot Chips so far is that, and I've been kind of circling around this one, I get your reaction to that the whole disaggregated serving concept is not so much vertical integration, it's just more efficiency. Exactly. And talk about the reasons, because this is a very nuanced technical point, but pre-fill and decode do something very good when you separate them because of the demand and complexity of the pre-fill. Yeah. Now the KV cache has to get smarter. It gets fatter, gets bigger, bloats up a little bit. Talk about what this means because this demand curve's not going away. Yeah. So talk about this pre-fill decode dynamic and it's not so much up the stack, it's more of really around the resource.
Jason Goodison
>> Yeah. So fundamentally when you think of inference, so that's when you actually ask the AI a question and it spits out an answer, you can actually split that into two workloads. One of them is called pre-fill. That one is about prompt processing. So let's say I've asked a really long question. I have 100,000 tokens in that question. Those tokens or those words need to be turned into numbers to actually run through the math. It's math, right? Those have to be turned into numbers. And then decode is when I'm autoregressively, they call it, which is generating one token at a time, right? Those are actually two different problems. And so they're taxing the models, right?
John Furrier
>> The decode taxes the models differently in math terms.
Jason Goodison
>> Right, right. And so it takes what came out of pre-fill and then it uses that in the autoregressive fashion to generate tokens one at a time. And so if you think about a GPU, fundamentally it's a graphics processing unit. And if you think, what do you need if you're generating graphics? Well, imagine you had one core, just for simplicity's sake, mapped to one pixel on the screen, and you're like, I need to know what color this pixel should be, 300 times a second. Well, actually it's really easy if I do a bunch of matrix multiplication for one pixel and then I just have a bunch of different cores and I do it all independently. I do it all at the same time in parallel, right? And so actually prefill for these AI models is something kind of similar where I can process all of the prefill or the context window, all of the prompt, each word I can process independently. So it maps really well onto that GPU, right? Because I'm just doing a bunch of matrix multiplications. Maps one-to-one, perfect. With decode, I take that prefill, they call it KV cache, right? That's just the context but in numbers. Well, now I have to store that right next to the actual silicon as well as the weights of the model. And the weights of the model are huge now. We're getting up to 3 trillion, 5 trillion parameter models. And then also a million context length windows. And so if I want to process 10 people, 20, 100 people at the same time, I've got 100 KV caches and I also have the model weights. And so it just blows up the memory problem. And so look, I can do all of my pre-fill independently on a GPU, maps perfectly. The decode though is where the GPU really breaks down.
John Furrier
>> talk about the consequences. Again, we're getting in the weeds a little bit here, but I think it's important for people to understand that if you screw up the KV cache, You got to reboot everything. So as you're getting through the multi-step processing, whether you're doing it through some sort of systolic array or whatever matrix multiplication, whatever process you're using, which is very layered, very complex, it screws up if you have to reset. Yeah. And you start that, and that's a GPU monolithic problem. Yeah. So the answer is, okay, put some compute here. And I think people misunderstood Cerebras when they first started, when I first met the team years ago. They're like, oh, it'll never work. Inference, it's just purpose-built inference. And it turns out they took on the memory wall. Good bet. Yeah, that was a good bet for Cerebras. But now guess what? You can integrate that in. Yeah, it's not a standalone. So you're starting to see the architecture approach is different. What's your technical view on this? Because it's kind of like a systems architecture game. Yeah, not a I got NVIDIA for this or Cerebras for that or SambaNova for this.
Jason Goodison
>> It's more than an architecture game because even if you find the best chip in the world, can that actually scale to production capacity? Because you know very well that there's constraints at TSMC, there's constraints with Micron or SK Hynix or whoever your memory provider is. So let's say $250 billion of NVIDIA silicon gets deployed next year. It could be double, it could be triple. And then the ASICs market might only be able to deploy like $1 billion. Right. And so it's a small market, but it's growing very quickly and it'll be more than $1 billion, but call it even $10 billion. It's a fraction of what NVIDIA is going to do, but it's growing very quickly. But you had to create those alliances with Broadcom or Intel, which is what SambaNova is doing early so that you could actually build the production capacity. Right. So it's that. It's also the fact that the memory problem is deeper than you would think because let's say you've got Cerebras. Cerebras has 44 gigabytes of SRAM. So they have that one big wafer and they put all of the memory right into it. Well, models are bigger than 44 gigabytes. So now we have to think about how are we going to wire these things together to actually split the model across a bunch of wafers,
John Furrier
>> right?
John Furrier
>> And it kicks ass at inference too. So it's a great use case. So you plug— wire it up. It's a great use
Jason Goodison
>> case.
Jason Goodison
>> But I guess what I'm trying to get to is that the bigger the model, it might actually change what chip you want to use too, right? So there's going to be different chips that are actually going to excel at different model types and different architectures. And we're just seeing an explosion of AI models. So it's not clear which one is going to win. It's getting bigger.
John Furrier
>> It's just the beginning. Let me ask you a question on the— question I get a lot, which is, hey, what's going on with all these new ASICs coming out? We see Positron, d-Matrix. So some custom accelerator market, ASICs, but they're not just accelerators, they're playing another role. What's the view of the market as these new entrants come in with this demand curve? How do you think that's going to play out? New formation? How does— because you guys are doing this. This is what you're doing. Yeah, we're doing it. How is that going to play out?
Jason Goodison
>> What's your vision of this? Well, I think people are really curious about ASICs ever since the OpenAI Cerebras deal and the IPO of Cerebras. People are starting to accept it. So when we started the company, this was pre all of that. And so we told people, we said Cerebras is going to be a big deal. And people would pooh-pooh it.
John Furrier
>> Yeah, they would pooh-pooh that. I've heard people, it'll never work. It's too big. It's never done before. Yeah.
Jason Goodison
>> And industry leaders and experts would do that. we would kind of get snubbed to some degree. But now everybody is on board, they understand it. And so we're even seeing the demand from enterprise of like, hey I know you've showed us this chip and you taught us how this chip works, but what about that chip? what about Etched? What about Cerebras? What about this and that? And so we just bring— we kind of connect the supply to the demand and we deploy it. And so you can take on less risk. If you're an enterprise and you really want to try a Cerebras or an Etched or any of these companies, you could either, pay millions of dollars and hopefully be able to actually deploy it and manage it
John Furrier
>> correctly.
John Furrier
>> want
Jason Goodison
>> to.
Jason Goodison
>> They don't want to take the risk. It's a huge
John Furrier
>> risk.
John Furrier
>> So are you targeting them for customers? We're targeting everyone that wants fast. All right. What do you guys— let's talk about your momentum. Take a minute to explain the momentum you have. Of course, there's a macro trend that's your friend. So that's going to be good for you guys. I think there's going to be a whole other class of build-out companies. I think you're a highlight of that. What happens next? The enterprise, they don't have billions in CapEx now. They'll use services. They'll do on-prem. They'll put in maybe a smaller cluster with FPGAs or some other low-cost, high-performance configuration and connect to a service that can give them single-tenant-like capability. Of course. That makes total sense to me. What are you guys targeting?
Jason Goodison
>> Well, there's 3 kinds of customer profiles, right? One of them would be you've got the frontier labs. The frontier labs just need a ridiculous amount of capacity. They're going to deploy some of their own. They're going to rent some from other people. FluidStack does a lot of stuff with Google and Anthropic. There's a big market there, but they care about training and they care about inference and they're experts at model compilation and everything. They have their own people on staff and they're just, they have infinite pockets to just deal with these problems. Then you have AI application companies and enterprise. So call it Cursor, Perplexity of the world, the OpenCodes of the world. They have a lot of demand, token demand. They're actually more comfortable most of the time paying for an SLA. So they're saying, I want to have X number of tokens per second. I want to be able to process X number of requests per second, things like that. And then the third category is the asset-light inference cloud. So for example, you've got Baseten, the Fireworks, Together AI. Some of these guys are starting to move down the stack, but they have essentially infinite demand, right? So any capacity that becomes available to them, they're gonna be able to connect that to a buyer. And so they own a lot of the end-user customer relationships already, and their end users are growing so fast, they just need more supply. And they're also, because they're asset-light, they're not locked into any—
John Furrier
>> Asset-light meaning they're not spending a lot of CapEx to do it and scale and sort of— Correct.
Jason Goodison
>> Before we did. Yeah, sorry, I should have explained that. When we say asset-light, we mean that they're not deploying, buying hardware and deploying it themselves, but they're renting capacity from other people, right? And then they're on-selling that and they're making their own SLAs and they have their own value-added services on top of that.
John Furrier
>> So they have— Fireworks is kicking ass. Everyone looks at those numbers, they're killing it. They're doing great. I think that's a big market. There's two approaches that I call the vertically integrated and then okay, meet that asset light demand, which is I got customers. Yeah, I will qualify what you have and then integrate in.
Jason Goodison
>> Yeah. Just a thought experiment. Let's say you raised $100 million of equity right now. You could either go and buy X number of machines and deploy them, or you could rent like 5 times X number of machines and then you could make 5x the revenue. So it makes a lot of sense for everyone to kind of pick what they're experts at and then do that. And we have a lot of these asset-light guys that are doing fantastic.
John Furrier
>> I really like what you guys are doing. In March at GTC, we saw all the Pareto curves. Jensen did his thing. I then wrote a post the next month, April, because Jensen's like, we're bounded by energy, the 5-layer cake he puts out there, which is totally legit. I wrote a post that, no, no, it's bounded by energy and money. That was the first post that kind of went out and was about a company and we featured Argentum AI, which was trying to figure out the financial code because as you pointed out, this has been documented on Bloomberg and other places, the circular financing, which I think some people try to throw shade on NVIDIA, but they're just doing a great job to help build the infrastructure. Talk about the financing aspect, though, because you're taking an approach to bet on the asset-light market that's in demand. Yep. And that financing— you got to get facilities, energy, and then you got to get the finance. Talk about this funding function of finance. Yeah, because then fast forward to this month, Jensen was in New York with the CEO of Goldman, KKR, $500 billion. They're taking care of the physical plant, my word. But the physical buildout. But there's still now a financial market developing. You're in the middle of this. Yeah. What's your vision on how that plays out? Because risk management is now in play on both sides.
Jason Goodison
>> Yeah, it's 100% true. There's a very well-understood debt market for GPUs. So if you want to go buy a bunch of GPUs, you can raise debt. And what NVIDIA will do is they'll underwrite the purchase of it. So they'll say, hey, if you cannot rent these machines or sell tokens on these machines, we'll rent them back for you. And what that allows you to do is go to the banks and say, hey, look, this is basically a guaranteed deal.
John Furrier
>> It's a AAA bond right there. It's like— and it's sometimes pledged. Yeah. So that's like every bank's like, I'm in. Exactly.
Jason Goodison
>> And we're talking about the company that's worth, over $4 trillion, they're not going to default. They're going to come through. And there's also infinite demand. So you will get a customer. Taking it a step farther, you talk about an ASIC. It's like, well, nobody really understands what the depreciation lifecycle on that ASIC is. Nobody really understands the residual value of that ASIC after the end of its life. There's no secondary market for it, because people haven't really adopted it or diffused it into the economy yet. So, it's fundamentally a much harder thing to get people to bet on. You're probably going to get quite a low interest rate if you're deploying GPUs. If you're deploying an ASIC, it's going to be much higher and you're going to have a much smaller pool of capital to do it from. So that is fundamentally what our business solves. A lot of these ASIC companies, I'm sure you chat with them here all the time, just absolutely exceptional people, like
John Furrier
>> incredible—
John Furrier
>> Great tech and there's demand for what they have.
Jason Goodison
>> There's demand for what they have and they've spent so much time on the technology, but the deployment piece is actually how NVIDIA is running circles around them. NVIDIA has great tech too, but it's not as good as a lot of these other ASIC companies, but they're adopted way more. And CUDA is not a good enough excuse anymore. the ecosystem is opening up. You can write your own compiler and AI
John Furrier
>> can—
John Furrier
>> I mean, CUDA is just a software model that makes the ASICs last longer until the next rev. So you can level up if the market changes. Whatever nuance is key, that could be replicated, baked into the platform.
Jason Goodison
>> It can be replicated for sure, especially with AI coding capabilities. You should be able to get something at least workable. But the reason that they're not being diffused more into the economy is the fact that there's just no one deploying them. And that's why we're stepping in.
John Furrier
>> Well, Jason, I'm really jealous of you. You're a young gun. I'm aging out over the years. You've got a long runway here. But you brought up the depreciation things because Jensen said something Dave Vellante and I were— and Brian were talking about is the analysts haven't modeled, and he put it in kind of quotes, they haven't spreadsheeted out what this is going to be. So a lot of people don't know what depreciation means because they don't know what the reuse is. So if we assume scarcity, architecturally smart engineers are using older chips Talk about that from a tech perspective, because the old idea was, oh, that's a chip. The next one comes out, the value drops. You can depreciate. That makes total sense in the old way. But in the new world where you have diversity of clusters, you have diversity of capabilities, there's a reuse market that keeps the prices up, which changes the modeling on the financial spreadsheets of valuation. So what's your view on this? It's kind of a random question, but it's one that everyone's asking, well, can you hedge that? Well, there's a futures market, but then again, if it's depreciating. So there's a whole conversation around the thesis of will the hardware and software be worthless, more or less in the future, or will it have staying power and durability? If you assume, okay, big clusters, small clusters, edge, physical AI, a chip today could be put into a robot maybe. Yeah, there's all kinds of supply chain functionality discussion.
Jason Goodison
>> I think you're thinking about it the right way. I'm sure you saw the deal with CoreWeave where they signed A100s through, I believe, 2029. And that is a very old chip at this point. And I think what's happening is you're in a supply constrained market. Some workloads are more valuable than others. And let's say I could run on an A100, and I'm making these numbers up, but at 10 tokens a second, or I could run on a, NVIDIA Cerebras combination with, 2,000 tokens per second. I'm obviously going to put my high-value workloads on the Cerebras rack, but there's probably a bunch of stuff I could just put on the A100 overnight and not think about, probably internal things, right? So there, I think it depends.
John Furrier
>> It's a TCO calculation at that point.
Jason Goodison
>> It is. It's like, what am I running?
John Furrier
>> It's policy-based, it's resource-based. Exactly. Intelligence could manage that. You put some AI in there. Yeah.
Jason Goodison
>> Well, I'm also always thinking about too, okay, what is the revenue per megawatt you can get? So let's say you've got an A100 and you have a megawatt of it deployed. Theoretically, you can make X amount on it. And then if you could upgrade to a new Blackwell generation or the Vera Rubin generation, you'd have to rework and put CapEx into the facility in order to actually be able to run those machines. But you'd get, I don't know, X times 10 or X times 100 in revenue. So I think the calculation there is really interesting. But the fact is, look, we're all so supply constrained that everything that's in production right now, we're just going to use it. And as we upgrade and build new data centers and upgrade old data centers, we will plug in new stuff and things will depreciate. And I don't think people will be using A100s forever.
John Furrier
>> There's just not enough sample size, Jason, on this. So it's— I think it's an open question. I think that's going to be one we're going to watch certainly in the middle of it. All right, final question. What are you optimizing for now? Give us a taste of what's coming. I know you got some deals brewing you can't talk about right now. What's going on? Set the direction.
Jason Goodison
>> Where's the company heading? The company is headed towards being the heterogeneous ASIC deployment arm. So everything that is not already being handled by your CoreWeave and your Nebius, There's a lot of fantastic chips out there. Everyone wants to try them. They have different use cases. We are going to be deploying those for customers and we'll be doing some announcements in the next few weeks, I believe. And I'm really excited to talk about that.
John Furrier
>> Maybe I can come back and we can talk about that. Yeah, we'll definitely do it. General Compute, not doing general purpose computing as we know it. General Compute is providing the scale for what we see as a democratization on the ASIC side as more entrants come in, more capabilities, again, more infrastructure demand continues to thunder away. I'm John Furrier, your host of theCUBE. Thanks for watching.