This discussion examines artificial intelligence, commonly abbreviated as AI, factories with a focus on neoClouds performance GPU goodput and security. Jordan Nanos of SemiAnalysis joins theCUBE Research hosts at the NYSE Wired theCUBE studio to discuss the ClusterMAX rating system and AI infrastructure considerations.
Nanos explains how neoClouds differentiate by performance reliability and speed and outlines hands-on testing methods for measuring GPU goodput conducting security audits and validating failure recovery. They emphasize that speed and rapid capacity delivery are non negotiable for neoCloud economics and that strong demand from OpenAI and Anthropic drives GPU purchases and supply constraints. Nanos also warns of significant security gaps in provider stacks and urges embargoed patching and clear ownership of the operational chain.
theCUBE analysts highlight monitoring goodput and recovery time as key financial and operational predictors and discuss implications for data center design GPU procurement and vendor risk management. Watch the full conversation for detailed testing methodology vendor analysis and practical guidance for operators and investors in the AI ecosystem.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Jordan Nanos, SemiAnalysis
This discussion examines artificial intelligence, commonly abbreviated as AI, factories with a focus on neoClouds performance GPU goodput and security. Jordan Nanos of SemiAnalysis joins theCUBE Research hosts at the NYSE Wired theCUBE studio to discuss the ClusterMAX rating system and AI infrastructure considerations.
Nanos explains how neoClouds differentiate by performance reliability and speed and outlines hands-on testing methods for measuring GPU goodput conducting security audits and validating failure recovery. They emphasize that speed and rapid capacity delivery are non negotiable for neoCloud economics and that strong demand from OpenAI and Anthropic drives GPU purchases and supply constraints. Nanos also warns of significant security gaps in provider stacks and urges embargoed patching and clear ownership of the operational chain.
theCUBE analysts highlight monitoring goodput and recovery time as key financial and operational predictors and discuss implications for data center design GPU procurement and vendor risk management. Watch the full conversation for detailed testing methodology vendor analysis and practical guidance for operators and investors in the AI ecosystem.
>> I'm Gemma Allen with NYSE Wired. This is AI Factories, where we talk all things infrastructure layer, fueling the next wave of technology. Joining me now for a conversation on exactly that is Jordan Nanos, semiconductor and AI infrastructure analyst at SemiAnalysis. Welcome Jordan.
Jordan Nanos
>> Great to be here. Thanks for having me.
Gemma Allen
>> So you operate in an interesting space. I was very excited for this conversation because I love the opportunity in the world of AI factories where we hear everyone talk a good game every day in, day out on this show. Zoom out a little bit and talk about the more macro picture right?
Jordan Nanos
>> Yeah.
Gemma Allen
>> Interesting time, so much money being spent, a lot of skepticism, a lot of excitement. The markets are manic, quite frankly. Maybe just to start, talk to me a little bit about your work, what you've been thinking on this last month or two even, because I think before that, it seems like it's irrelevant now in the world of tech, and we can maybe go from there.
Jordan Nanos
>> Yeah, definitely. So, primary thing that I work on at SemiAnalysis is called ClusterMAX. It's a rating system for all of the neoClouds in the industry. As you said, AI factories are incredibly important. neoClouds are what a lot of the AI labs are using to build the AI infrastructure that they're going to need to train models, deploy them at scale. Biggest thing we're seeing right now is just incredible demand. So that's causing a lot of little cracks to form where people start to need to build new technologies and start to figure out what the problems at scale are that are different than what happened in the seed round when they were starting up the neoCloud, for example.
Gemma Allen
>> Let's start on the cracks. neoClouds. Jensen said this year at GTC, it's hard to talk about infrastructure and not talk about Jensen, so let's just go there first, that there's a new Neocloud every day. How do you truly differentiate? A lot of it is just demand-based. Have we really separated the men from the boys in that space? Do we know for sure? What are your thoughts? It seems like it's a space that's getting so much cash injection. Do you think that the numbers add up? What's your thought on the economics of this model?
Jordan Nanos
>> Yeah, so obviously supply-demand, and I think at a basic level, demand is there. We're seeing absolutely massive ARR growth from OpenAI and Anthropic in particular, but also really strong growth from a lot of enterprise companies and stuff like Gemini or Grok from Google and xAI, as well as a lot of the startups that are just raising incredibly large rounds, and they need to spend that money on compute. And where that plays out is that everybody started out by picking a dance partner where the big labs, OpenAI, Anthropic, they were finding individual neoClouds that could go faster for them than the hyperscalers could. And then they just needed to buy from everybody. And so you can look at how OpenAI is buying from Microsoft, how they're buying from CoreWeave, how they're going down the list, and even how they're doing self -build. And then you can look at Anthropic and you can see Project Rainier with Trainium at AWS. They've got TPUs with Google. They've got GPUs with NVIDIA. They've got a CoreWeave deal. There's all sorts of ways in which these guys can figure out ways to spend money so that they can get access to GPUs and then return it at a significant rate. If they were not returning massive cash flows from these GPUs, they wouldn't be buying them right now. But we're seeing them go for more than what anybody can provide. And I guess that's just driving supply. And I think there's a lot of different ways in which supply plays out. Obviously, NVIDIA, it's been great for them. But there's limits in terms of how many balance sheets you can put GPUs on, how much data center, power, land, cooling equipment you can actually get set up in a certain amount of time, how many GPUs you can get access to. And as those dynamics play out, we start to see a lot of the business level stuff dictating who's being successful rather than the technical merits, which I think is kind of what you implied in the question that if demand continues to be so strong, it's really not going to matter who's got better reliability or who's got better performance, which is obviously what we spend almost all of our time focused on in the ClusterMAX rating system. And much to our chagrin, a lot of people that are just focused on building data centers of relatively low quality as fast as they can are being successful right now.
Gemma Allen
>> I want to get into ClusterMAX and I want to talk about the technical performance side of it. But first, you said something interesting there, right? Relationships have been formed. People are getting into bed together. We've seen this really compound, actually, over the last 6, 12 months alone. Yeah. What are your thoughts, though, on the true moat of that model? How do you think Nebius versus the CoreWeave versus a hyperscaler, truly adds some sort of competitive advantage three years from now? What is the one metric you think is non -negotiable there?
Jordan Nanos
>> Yeah, it speeds them up. At the end of the day, the amount that a neoCloud can respond to a demand signal in the market is going to dictate whether they're going to return capital on what they've invested in GPUs. So at this point, frankly, xAI is the leader in speed from our tracking. They built Colossus 1 and 2 in absolutely record time. Many of the other NeoClouds, CoreWeave, Nebius, they've had their problems. They've had their delays in construction and getting air permits and all of this stuff to turn on a lot of these sites. And you can't start recognizing revenue until you onboard those customers. I think we've seen the hyperscalers pour in a ton of money into capital. they're buying land, they're buying a ton of equipment, they're hiring construction contractors all over the world. And this is to the tune of like a trillion dollars of CapEx or more next year. So they've got baseload figured out. And the neoclouds, the rest of these projects, which in some cases involve these hyperscalers, like Google's project with Blackstone, for example, they're becoming neo -clouds of a sort themselves, where Microsoft's got to go procure capacity from Lambda or from Nscale and they've got to find that little bit of flex on top so that they can continue to respond to the demand signals, which are so strong from the market. And what that turns into is that the returns on capital for the base load are, they're okay, but the really strong returns are for people who can deliver you GPUs three months, one month now as opposed to a project that takes 18 monthsto pass all the permitting and start construction and then finally roll stuff in later.
Gemma Allen
>> Talk about ClusterMAX. I know you talk a lot about good put right this kind of metric of managing performance versus cost versus output. First of all break that down. But also i'm interested to understand what is your evaluation based upon? Like what like talk me through how you come up with these assessments?
Jordan Nanos
>> Yeah so there's two big components. One is just talking to the buyers about their experience.They run at a bigger scale than we can during our testing, which is the second component. And so the buyers from these markets are very informative in terms of what their experience was for support, for reliability, for performance, things like that. We also do our own hands -on testing to validate what we're seeing because we can get mixed signals from those providers and a lot of the buyers as well. So our testing process is kind of three phases. We start with an audit, we focus a lot on security right now, making sure that they are providing a secure cluster. Really terrible results there. We are putting out an article next week going through some of that. The second component is performance, where we test to see that for NVIDIA GPUs or AMD GPUs, we know what the performance expectations are, and so we're just testing that they meet these expectations. We're not really differentiating at a level of you're 2 % faster on a training workload, your micro benchmark on networking is 5 % faster. These things come out in the wash in many cases when you're working hand in hand with a good provider. But what we notice obviously is when performance is just way below the actual expectation, which has happened many times. And then to your point, good point, the third component, which is the most critical for our hands on testing is reliability. What we do there is we simulate a series of failures on the hardware and we check to see how the provider reacts. So this is a component of first of all, having the monitoring and software systems in place to be able to identify a failure. And then second of all, being able to actually repair or just replace a node, for example, a switch, a cable, things that we can simulate are going to happen in the real world. Interestingly, our testing process is about a week on a cluster of just four nodes. And we see real hardware failures too. So yeah, the overall experience of, roughly speaking, goodput, the definition is how much good work you can do with the GPUs that you have, how much throughput is good, which is to say if I'm training a model, I'm running it for a month, and I get a peak amount of performance when all my GPUs are online, that's great to know. It's great to know that metric. But it's also important to know how things perform when GPUs start failing, which they do. And a lot of people have a dependency on the providers especially with these new GB200 and GB300 racks where there's a big scale-up domain and one failure can effectively take down the full rack of 72 GPUs. This is like thousands of dollars an hour or hundreds and yeah, just a rack is about five million dollars capital expense upfront. So divide that by your three-year contractlike a lot.
Gemma Allen
>> What are you seeing from the perspective of correction time right? A node fails, the GPU fails, which you're saying happens all the time. There's multiple layers of ownership there, though, right? Yeah, you have one provider who's providing the service, but there is a whole technology ecosystem feeding that failure.
Jordan Nanos
>> Yeah.
Gemma Allen
>> Who's doing it? What needs to happen for it to be managed well?
Jordan Nanos
>> Well, I think in an interesting way, it's not always one provider. In NeoCloud world, you have the people that you sign the contract with for the GPUs. Sometimes you might have somebody who's playing a broker role. Then you have the people who actually like operate the bare metal cluster. These are still software engineers who don't live near the site usually. And then there's the data center technicians who are like actually swapping parts when something fails. Many times they depend on an OEM. So there's like the OEM technicians that come in. It's this whole chain of people that you need to trust. And so roughly speaking, our best, most reliable experiences have been with providers that own that entire chain. Like the guy who is a data center technician who is swapping a cable or a drive has like equity in the company that you signed a contract with and really cares about them being successful. There's a lot of other situations where you're four layers removed from those people, the broker, the neoCloud that's selling you GPUs is kind of trying to hide who it actually is because otherwise they think you're going to go circumvent them and go around them. That's not a great way to start a relationship. Now there's many cases where construction company hands off to data center operations company which employs the technicians who then hand off to the SREs who build the cluster who then you work with. And those are there's many models that are great for that as well and work well. But generally speaking when we do that testing we expect to see some sort of ownership some sort of response time some sort of intelligent response from organic intelligence like a human not some automated chatbot that gives us assurances like. This is when your stuff's going to get fixed. This is how long it's going to take. This is what we found happened. It's not necessarily a secret shopper experience. But when people are unprepared and they haven'tyou know reviewed the criteria that we're using or asked what we're going to be up to a lot of them get caught by surprise.
Gemma Allen
>> I love it. You are though identifying some of the used car salesmen in the process too right which is an important part because we hear a lot about the world of neoCloud the financing programs that kind of sub -layer, right, where there is a lot of question over how that can actually truly be efficient. Like, longer term. So back to what you said at the beginning, I know you said it's not going to be released until next week, but security. It's an interesting point, though, right? Are there, some blind assumptions being made, by, technical leaders that you feel are somewhat, highly misinformed? When you say it's surprisingly bad, give me some context to that.
Jordan Nanos
>> Yeah, it's surprising in some ways, unsurprising in others when you get more experience working with these guys. I just described that chain, right? And you can kind of imagine that when that goes wrong, it's like broken telephone trying to get something fixed and things don't work quite the way you expect or on the timeline you expect. So a real example is many of the labs are bringing their CISO onto these calls because it's a counterparty risk who you decide to go with. And you need to trust these people. Simple example is just keeping software up to date on the cluster. We're seeing two dynamics right now. One is that all these frontier models, when you look at Project Glasswing from Anthropic or everything with OpenAI and their Codex model and that announcement that they made with Hugging Face where they showed how the model was autonomously hacking Hugging Face's data sets, infrastructure to try to get answers to an eval during training without any human involved. Really scary, interesting talk from Black Hat that everybody should go watch. Oh, wow. Gives you a taste for the future here. But, yeah, the point is that these models are finding zero days in software. They're finding existing vulnerabilities that humans don't know about, and they're exploiting them. And so as models get better, we only expect more of that to happen, more sophisticated exploits of the software that we have, that we use today. But the second thing is, when these CVEs come out describing this vulnerability and they tell people to patch it, it's up to your provider to roll out a fix, to be keeping track of when things are coming out. In some cases, being in an embargo program with companies like NVIDIA or AMD so they get advanced notice that these disclosures are going to happen and they can prepare a patch so that on day zero, when that vulnerability gets disclosed, they can already start the rollout of upgrading your software so that you're not getting exploited. A lot of these providers that we check have stuff that's not like a month or two months old, but like three years old that they haven't upgraded. And this just means it's a ticking time bomb until somebody comes along and starts looking for your model weights or your data sets or your RL environments that you really care about keeping private, for example. Or data exfiltration isn't the only thing. think about ransomware, think about these crypto mining hackers that take over clusters. We hear about all sorts of bad stuff that happens from poor security.
Gemma Allen
>> What are your thoughts on this narrative that open -weight models are less secure?
Jordan Nanos
>> Okay, so that's the second part of it. They are certainly less guardrailed, if I can use that word. I'm not sure if it is a word, but, yeah, the point is, they can be used. So we try to build POC exploits in order to explain to providers exactly how somebody can use these, zero days in their, or, sorry, these vulnerabilities that are in their environment. and how they would be exploited because a lot of them will push back and say, oh, this old software version, it doesn't matter. I don't need to upgrade it. My customer said it's okay. And in some respects, okay, it's the customer's decision. But in other respects, they're just saying that. And so, yeah, we try to demonstrate these things. And you can't even ask Claude about a security issue. You can't ask it to check your own cluster. It's going to deny you immediately. o1 is a little bit better, but we have to use a lot of these open weight models because they don't have guardrails on them that reject any sort of research into security. Even if you're in the security program and approved by Anthropic or OpenAI like we are in some cases.
Gemma Allen
>> Wow!
Jordan Nanos
>> So that dynamic totally exists where it's a bit of a wild west with the Chinese models. Where I would say people don't necessarily have the same cause for concern is that in our experience they are still remarkably, significantly worse at exploiting these security vulnerabilities than the leading frontier American models. I would say China has not at least demonstrated a focus on cybersecurity in the open weight models.
Gemma Allen
>> I guess for now though, right? That is a broader concern. Okay, so Wild Wild West. you could also say from the perspective of macroeconomics the entire industry is a bit of a wild west right, if you think About how we create any sort of universal agreement on what depreciation looks like for GPU costs five years from now, right? There's so much capital flowing into the space you're a bank. You're lending money against these bets How are you truly evaluating it right that is a conversation that is kind of floating out there that everyone's kind of skirting around What are your thoughts from the perspective of how you do actually have some sort of unique performance measure that can feed financial predictability? What does that look like?
Jordan Nanos
>> Yeah. So I think first of all you need to segment the market in terms of how people make these decisions. The top level is like Anthropic or OpenAI or Google or Meta or Microsoft just taking full sites. and the way in which they do these deals is completely different than the way a new Silicon Valley startup who got funding for GPUs is going out and trying to get one tenant like one section of a bigger cluster from a CoreWeave Nebius or Crusoe or Lambda or Together any of these neoClouds that are out there and then there's the kind of like bottom end of the market, which is you or me paying for tokens or people at home that are doing development on a single GPU at a time. And there's clearly this like backwardation in the pricing curve right now where if you are willing to put up money and prepay for stuff, you can get a significant discount on the total contract value, but you got to wait like six months for stuff to get installed. Right. If you want stuff now, you pay a significant premium. and if you want tokens, like tokens on a per GPU hour basis, come at a significant premium on top of that and on demand, people only want tokens or want to have a GPU for an hour or two, they pay a significant premium as well. So I think the market is shaping up where, kind of like what I was saying earlier, many companies need to have this long term base load committed capacity of GPUs or tokens and then they need to do the engineering work to understand what their demand is going to grow like in the future and properly plan to bring stuff online or have the cash set aside or the relationships available to kind of flex up and down and get access to the stuff they need. But generally speaking, we see people buy more GPUs, not less. We see very few people giving stuff back. And there's limits that are being reached in the supply chain of how much can be produced and how much can be turned on. We expect those to continue for a long time. I think we are the industry experts on understanding how much can be produced on the supply side of this curve There are real limits to that. There's only so many wafers from TSMC. There's only so much HBM only so much DRAM so there's limits and Unless the models get worse I find it really hard to understand scenarios where demand is gonna trail off and fall off a cliff.
Gemma Allen
>> So last question, let's talk about something I mentioned to you before the show.I heard Dylan Patel commenting on the Anthropic and OpenAI evaluations, how the market has responded, some of those kind of fuzzy metrics that have been used to determine revenue. both these companies are predicted to go public between now and I guess 2027 at some really staggering ARR. What are your thoughts? What comes to top of mind to you when you think aboutthis?
Jordan Nanos
>> Yeah historically it's super strange to see people take a uh a weekly run rate and multiply it by 52 or something like that and call that the ARR. Um but yeah when you're Sam and you're Dario you write your own rules right? Yeah I think investors want to understand the growth rate of the business, and I think it's really hard to contend with the fact that these businesses are growing incredibly fast. Let's put OpenAI and Anthropic aside for a second. I was talking to a software company earlier this week who was in the middle of raising their Series B. They've had incredible growth. They were doing it during a big conference. They go out to the investors at a certain price. They finish the conference. They've got a massive pipeline, and they go, Look, guys, I don't think we need the money right now. Now, I can give you a Salesforce export at the end of this week, and we can put a multiple on top of that and reprice the round. But maybe let's just check in in two months and see what happens. And so everybody says, okay, let's see how growth goes for the next two months. You're not net profitable, but you got a lot of runway. It's all good. And then two months come, and the growth just continues. So I think in this case, it's kind of valid to have both metrics. You'd like to know what the current run rate is multiplied by whatever factor you care about. And look, Anthropic entered this year projecting 100 million ARR. I think a lot of people were a little bit questioning whether they would come through on that. And they're going to come through way before December on that. So they're going to cross 100 quite soon by our modeling. Anyway, look, these businesses are incredible. Like I said earlier, they turn on more GPUs, they get more revenue. It's almost a direct line from all of our tracking of their data centers and their chips installations.
Gemma Allen
>> Can we do a quick hot takes, couple of questions at the end? Sure. Okay. Hyperscaler, biggest kind of, favorite hyperscaler?
Jordan Nanos
>> Favorite hyperscaler, right now, when it comes to construction, AWS is the fastest, but can't stand EFA, so I'm gonna go Oracle. Oracle. They've pioneered a lot with the multi-plane RoCE networking. Love that.
Gemma Allen
>> Might be a good time to buy Oracle stock too, right?
Jordan Nanos
>> This is not an endorsement of their stock.
Gemma Allen
>> Joking, joking.
Jordan Nanos
>> Buy the research, everybody.
Gemma Allen
>> Yes, happy with that. Chip company outside of NVIDIA.
Jordan Nanos
>> Oh, man. Startup or real product?
Gemma Allen
>> Real project.
Jordan Nanos
>> Okay. TPUs are amazing. A lot of people, there's so much demand for TPUs. They're really cool. A lot of cool stuff on the roadmap too. I'll put Trainium in there as well. I've had some good experience doing micro benchmarking with Trainium. Love the profiling. Chip Startup, I don't know. I can't pick. Don't want to give too strong an endorsement. Really impressive what a lot of them are doing. Need to see these guys produce tokens, okay? Not these deals. These cool marketing videos. Produce some tokens. Chip Startups, let's go.
Gemma Allen
>> Favorite NeoCloud?
Jordan Nanos
>> Favorite NeoCloud? CoreWeave's been the top of the Platinum tier rankings. They're a great experience. Yeah, we'll go with that.
Gemma Allen
>> Let's end with an easy one. Favorite tech leader.
Jordan Nanos
>> Favorite tech leader. Jensen.
Gemma Allen
>> in NVIDIA this year, we asked folks, who's the bigger celebrity, Jesus or Jensen? You know what the answer was?
Jordan Nanos
>> Who's Jesus?
Gemma Allen
>> Jordan Nanos, thank you so much for joining us on theCUBE and NYSE Wired.
Jordan Nanos
>> Okay, great to be here.
Gemma Allen
>> I'm Gemma Allen, coming to you from theCUBE Studio at the New York Stock Exchange. This is AI Factories, one of our programs with NYSE Wired. Thanks for watching.
>> I'm Gemma Allen with NYSE Wired. This is AI Factories, where we talk all things infrastructure layer, fueling the next wave of technology. Joining me now for a conversation on exactly that is Jordan Nanos, semiconductor and AI infrastructure analyst at SemiAnalysis. Welcome Jordan.
Jordan Nanos
>> Great to be here. Thanks for having me.
Gemma Allen
>> So you operate in an interesting space. I was very excited for this conversation because I love the opportunity in the world of AI factories where we hear everyone talk a good game every day in, day out on this show. Zoom out a little bit and talk about the more macro picture right?
Jordan Nanos
>> Yeah.
Gemma Allen
>> Interesting time, so much money being spent, a lot of skepticism, a lot of excitement. The markets are manic, quite frankly. Maybe just to start, talk to me a little bit about your work, what you've been thinking on this last month or two even, because I think before that, it seems like it's irrelevant now in the world of tech, and we can maybe go from there.
Jordan Nanos
>> Yeah, definitely. So, primary thing that I work on at SemiAnalysis is called ClusterMAX. It's a rating system for all of the neoClouds in the industry. As you said, AI factories are incredibly important. neoClouds are what a lot of the AI labs are using to build the AI infrastructure that they're going to need to train models, deploy them at scale. Biggest thing we're seeing right now is just incredible demand. So that's causing a lot of little cracks to form where people start to need to build new technologies and start to figure out what the problems at scale are that are different than what happened in the seed round when they were starting up the neoCloud, for example.
Gemma Allen
>> Let's start on the cracks. neoClouds. Jensen said this year at GTC, it's hard to talk about infrastructure and not talk about Jensen, so let's just go there first, that there's a new Neocloud every day. How do you truly differentiate? A lot of it is just demand-based. Have we really separated the men from the boys in that space? Do we know for sure? What are your thoughts? It seems like it's a space that's getting so much cash injection. Do you think that the numbers add up? What's your thought on the economics of this model?
Jordan Nanos
>> Yeah, so obviously supply-demand, and I think at a basic level, demand is there. We're seeing absolutely massive ARR growth from OpenAI and Anthropic in particular, but also really strong growth from a lot of enterprise companies and stuff like Gemini or Grok from Google and xAI, as well as a lot of the startups that are just raising incredibly large rounds, and they need to spend that money on compute. And where that plays out is that everybody started out by picking a dance partner where the big labs, OpenAI, Anthropic, they were finding individual neoClouds that could go faster for them than the hyperscalers could. And then they just needed to buy from everybody. And so you can look at how OpenAI is buying from Microsoft, how they're buying from CoreWeave, how they're going down the list, and even how they're doing self -build. And then you can look at Anthropic and you can see Project Rainier with Trainium at AWS. They've got TPUs with Google. They've got GPUs with NVIDIA. They've got a CoreWeave deal. There's all sorts of ways in which these guys can figure out ways to spend money so that they can get access to GPUs and then return it at a significant rate. If they were not returning massive cash flows from these GPUs, they wouldn't be buying them right now. But we're seeing them go for more than what anybody can provide. And I guess that's just driving supply. And I think there's a lot of different ways in which supply plays out. Obviously, NVIDIA, it's been great for them. But there's limits in terms of how many balance sheets you can put GPUs on, how much data center, power, land, cooling equipment you can actually get set up in a certain amount of time, how many GPUs you can get access to. And as those dynamics play out, we start to see a lot of the business level stuff dictating who's being successful rather than the technical merits, which I think is kind of what you implied in the question that if demand continues to be so strong, it's really not going to matter who's got better reliability or who's got better performance, which is obviously what we spend almost all of our time focused on in the ClusterMAX rating system. And much to our chagrin, a lot of people that are just focused on building data centers of relatively low quality as fast as they can are being successful right now.
Gemma Allen
>> I want to get into ClusterMAX and I want to talk about the technical performance side of it. But first, you said something interesting there, right? Relationships have been formed. People are getting into bed together. We've seen this really compound, actually, over the last 6, 12 months alone. Yeah. What are your thoughts, though, on the true moat of that model? How do you think Nebius versus the CoreWeave versus a hyperscaler, truly adds some sort of competitive advantage three years from now? What is the one metric you think is non -negotiable there?
Jordan Nanos
>> Yeah, it speeds them up. At the end of the day, the amount that a neoCloud can respond to a demand signal in the market is going to dictate whether they're going to return capital on what they've invested in GPUs. So at this point, frankly, xAI is the leader in speed from our tracking. They built Colossus 1 and 2 in absolutely record time. Many of the other NeoClouds, CoreWeave, Nebius, they've had their problems. They've had their delays in construction and getting air permits and all of this stuff to turn on a lot of these sites. And you can't start recognizing revenue until you onboard those customers. I think we've seen the hyperscalers pour in a ton of money into capital. they're buying land, they're buying a ton of equipment, they're hiring construction contractors all over the world. And this is to the tune of like a trillion dollars of CapEx or more next year. So they've got baseload figured out. And the neoclouds, the rest of these projects, which in some cases involve these hyperscalers, like Google's project with Blackstone, for example, they're becoming neo -clouds of a sort themselves, where Microsoft's got to go procure capacity from Lambda or from Nscale and they've got to find that little bit of flex on top so that they can continue to respond to the demand signals, which are so strong from the market. And what that turns into is that the returns on capital for the base load are, they're okay, but the really strong returns are for people who can deliver you GPUs three months, one month now as opposed to a project that takes 18 monthsto pass all the permitting and start construction and then finally roll stuff in later.
Gemma Allen
>> Talk about ClusterMAX. I know you talk a lot about good put right this kind of metric of managing performance versus cost versus output. First of all break that down. But also i'm interested to understand what is your evaluation based upon? Like what like talk me through how you come up with these assessments?
Jordan Nanos
>> Yeah so there's two big components. One is just talking to the buyers about their experience.They run at a bigger scale than we can during our testing, which is the second component. And so the buyers from these markets are very informative in terms of what their experience was for support, for reliability, for performance, things like that. We also do our own hands -on testing to validate what we're seeing because we can get mixed signals from those providers and a lot of the buyers as well. So our testing process is kind of three phases. We start with an audit, we focus a lot on security right now, making sure that they are providing a secure cluster. Really terrible results there. We are putting out an article next week going through some of that. The second component is performance, where we test to see that for NVIDIA GPUs or AMD GPUs, we know what the performance expectations are, and so we're just testing that they meet these expectations. We're not really differentiating at a level of you're 2 % faster on a training workload, your micro benchmark on networking is 5 % faster. These things come out in the wash in many cases when you're working hand in hand with a good provider. But what we notice obviously is when performance is just way below the actual expectation, which has happened many times. And then to your point, good point, the third component, which is the most critical for our hands on testing is reliability. What we do there is we simulate a series of failures on the hardware and we check to see how the provider reacts. So this is a component of first of all, having the monitoring and software systems in place to be able to identify a failure. And then second of all, being able to actually repair or just replace a node, for example, a switch, a cable, things that we can simulate are going to happen in the real world. Interestingly, our testing process is about a week on a cluster of just four nodes. And we see real hardware failures too. So yeah, the overall experience of, roughly speaking, goodput, the definition is how much good work you can do with the GPUs that you have, how much throughput is good, which is to say if I'm training a model, I'm running it for a month, and I get a peak amount of performance when all my GPUs are online, that's great to know. It's great to know that metric. But it's also important to know how things perform when GPUs start failing, which they do. And a lot of people have a dependency on the providers especially with these new GB200 and GB300 racks where there's a big scale-up domain and one failure can effectively take down the full rack of 72 GPUs. This is like thousands of dollars an hour or hundreds and yeah, just a rack is about five million dollars capital expense upfront. So divide that by your three-year contractlike a lot.
Gemma Allen
>> What are you seeing from the perspective of correction time right? A node fails, the GPU fails, which you're saying happens all the time. There's multiple layers of ownership there, though, right? Yeah, you have one provider who's providing the service, but there is a whole technology ecosystem feeding that failure.
Jordan Nanos
>> Yeah.
Gemma Allen
>> Who's doing it? What needs to happen for it to be managed well?
Jordan Nanos
>> Well, I think in an interesting way, it's not always one provider. In NeoCloud world, you have the people that you sign the contract with for the GPUs. Sometimes you might have somebody who's playing a broker role. Then you have the people who actually like operate the bare metal cluster. These are still software engineers who don't live near the site usually. And then there's the data center technicians who are like actually swapping parts when something fails. Many times they depend on an OEM. So there's like the OEM technicians that come in. It's this whole chain of people that you need to trust. And so roughly speaking, our best, most reliable experiences have been with providers that own that entire chain. Like the guy who is a data center technician who is swapping a cable or a drive has like equity in the company that you signed a contract with and really cares about them being successful. There's a lot of other situations where you're four layers removed from those people, the broker, the neoCloud that's selling you GPUs is kind of trying to hide who it actually is because otherwise they think you're going to go circumvent them and go around them. That's not a great way to start a relationship. Now there's many cases where construction company hands off to data center operations company which employs the technicians who then hand off to the SREs who build the cluster who then you work with. And those are there's many models that are great for that as well and work well. But generally speaking when we do that testing we expect to see some sort of ownership some sort of response time some sort of intelligent response from organic intelligence like a human not some automated chatbot that gives us assurances like. This is when your stuff's going to get fixed. This is how long it's going to take. This is what we found happened. It's not necessarily a secret shopper experience. But when people are unprepared and they haven'tyou know reviewed the criteria that we're using or asked what we're going to be up to a lot of them get caught by surprise.
Gemma Allen
>> I love it. You are though identifying some of the used car salesmen in the process too right which is an important part because we hear a lot about the world of neoCloud the financing programs that kind of sub -layer, right, where there is a lot of question over how that can actually truly be efficient. Like, longer term. So back to what you said at the beginning, I know you said it's not going to be released until next week, but security. It's an interesting point, though, right? Are there, some blind assumptions being made, by, technical leaders that you feel are somewhat, highly misinformed? When you say it's surprisingly bad, give me some context to that.
Jordan Nanos
>> Yeah, it's surprising in some ways, unsurprising in others when you get more experience working with these guys. I just described that chain, right? And you can kind of imagine that when that goes wrong, it's like broken telephone trying to get something fixed and things don't work quite the way you expect or on the timeline you expect. So a real example is many of the labs are bringing their CISO onto these calls because it's a counterparty risk who you decide to go with. And you need to trust these people. Simple example is just keeping software up to date on the cluster. We're seeing two dynamics right now. One is that all these frontier models, when you look at Project Glasswing from Anthropic or everything with OpenAI and their Codex model and that announcement that they made with Hugging Face where they showed how the model was autonomously hacking Hugging Face's data sets, infrastructure to try to get answers to an eval during training without any human involved. Really scary, interesting talk from Black Hat that everybody should go watch. Oh, wow. Gives you a taste for the future here. But, yeah, the point is that these models are finding zero days in software. They're finding existing vulnerabilities that humans don't know about, and they're exploiting them. And so as models get better, we only expect more of that to happen, more sophisticated exploits of the software that we have, that we use today. But the second thing is, when these CVEs come out describing this vulnerability and they tell people to patch it, it's up to your provider to roll out a fix, to be keeping track of when things are coming out. In some cases, being in an embargo program with companies like NVIDIA or AMD so they get advanced notice that these disclosures are going to happen and they can prepare a patch so that on day zero, when that vulnerability gets disclosed, they can already start the rollout of upgrading your software so that you're not getting exploited. A lot of these providers that we check have stuff that's not like a month or two months old, but like three years old that they haven't upgraded. And this just means it's a ticking time bomb until somebody comes along and starts looking for your model weights or your data sets or your RL environments that you really care about keeping private, for example. Or data exfiltration isn't the only thing. think about ransomware, think about these crypto mining hackers that take over clusters. We hear about all sorts of bad stuff that happens from poor security.
Gemma Allen
>> What are your thoughts on this narrative that open -weight models are less secure?
Jordan Nanos
>> Okay, so that's the second part of it. They are certainly less guardrailed, if I can use that word. I'm not sure if it is a word, but, yeah, the point is, they can be used. So we try to build POC exploits in order to explain to providers exactly how somebody can use these, zero days in their, or, sorry, these vulnerabilities that are in their environment. and how they would be exploited because a lot of them will push back and say, oh, this old software version, it doesn't matter. I don't need to upgrade it. My customer said it's okay. And in some respects, okay, it's the customer's decision. But in other respects, they're just saying that. And so, yeah, we try to demonstrate these things. And you can't even ask Claude about a security issue. You can't ask it to check your own cluster. It's going to deny you immediately. o1 is a little bit better, but we have to use a lot of these open weight models because they don't have guardrails on them that reject any sort of research into security. Even if you're in the security program and approved by Anthropic or OpenAI like we are in some cases.
Gemma Allen
>> Wow!
Jordan Nanos
>> So that dynamic totally exists where it's a bit of a wild west with the Chinese models. Where I would say people don't necessarily have the same cause for concern is that in our experience they are still remarkably, significantly worse at exploiting these security vulnerabilities than the leading frontier American models. I would say China has not at least demonstrated a focus on cybersecurity in the open weight models.
Gemma Allen
>> I guess for now though, right? That is a broader concern. Okay, so Wild Wild West. you could also say from the perspective of macroeconomics the entire industry is a bit of a wild west right, if you think About how we create any sort of universal agreement on what depreciation looks like for GPU costs five years from now, right? There's so much capital flowing into the space you're a bank. You're lending money against these bets How are you truly evaluating it right that is a conversation that is kind of floating out there that everyone's kind of skirting around What are your thoughts from the perspective of how you do actually have some sort of unique performance measure that can feed financial predictability? What does that look like?
Jordan Nanos
>> Yeah. So I think first of all you need to segment the market in terms of how people make these decisions. The top level is like Anthropic or OpenAI or Google or Meta or Microsoft just taking full sites. and the way in which they do these deals is completely different than the way a new Silicon Valley startup who got funding for GPUs is going out and trying to get one tenant like one section of a bigger cluster from a CoreWeave Nebius or Crusoe or Lambda or Together any of these neoClouds that are out there and then there's the kind of like bottom end of the market, which is you or me paying for tokens or people at home that are doing development on a single GPU at a time. And there's clearly this like backwardation in the pricing curve right now where if you are willing to put up money and prepay for stuff, you can get a significant discount on the total contract value, but you got to wait like six months for stuff to get installed. Right. If you want stuff now, you pay a significant premium. and if you want tokens, like tokens on a per GPU hour basis, come at a significant premium on top of that and on demand, people only want tokens or want to have a GPU for an hour or two, they pay a significant premium as well. So I think the market is shaping up where, kind of like what I was saying earlier, many companies need to have this long term base load committed capacity of GPUs or tokens and then they need to do the engineering work to understand what their demand is going to grow like in the future and properly plan to bring stuff online or have the cash set aside or the relationships available to kind of flex up and down and get access to the stuff they need. But generally speaking, we see people buy more GPUs, not less. We see very few people giving stuff back. And there's limits that are being reached in the supply chain of how much can be produced and how much can be turned on. We expect those to continue for a long time. I think we are the industry experts on understanding how much can be produced on the supply side of this curve There are real limits to that. There's only so many wafers from TSMC. There's only so much HBM only so much DRAM so there's limits and Unless the models get worse I find it really hard to understand scenarios where demand is gonna trail off and fall off a cliff.
Gemma Allen
>> So last question, let's talk about something I mentioned to you before the show.I heard Dylan Patel commenting on the Anthropic and OpenAI evaluations, how the market has responded, some of those kind of fuzzy metrics that have been used to determine revenue. both these companies are predicted to go public between now and I guess 2027 at some really staggering ARR. What are your thoughts? What comes to top of mind to you when you think aboutthis?
Jordan Nanos
>> Yeah historically it's super strange to see people take a uh a weekly run rate and multiply it by 52 or something like that and call that the ARR. Um but yeah when you're Sam and you're Dario you write your own rules right? Yeah I think investors want to understand the growth rate of the business, and I think it's really hard to contend with the fact that these businesses are growing incredibly fast. Let's put OpenAI and Anthropic aside for a second. I was talking to a software company earlier this week who was in the middle of raising their Series B. They've had incredible growth. They were doing it during a big conference. They go out to the investors at a certain price. They finish the conference. They've got a massive pipeline, and they go, Look, guys, I don't think we need the money right now. Now, I can give you a Salesforce export at the end of this week, and we can put a multiple on top of that and reprice the round. But maybe let's just check in in two months and see what happens. And so everybody says, okay, let's see how growth goes for the next two months. You're not net profitable, but you got a lot of runway. It's all good. And then two months come, and the growth just continues. So I think in this case, it's kind of valid to have both metrics. You'd like to know what the current run rate is multiplied by whatever factor you care about. And look, Anthropic entered this year projecting 100 million ARR. I think a lot of people were a little bit questioning whether they would come through on that. And they're going to come through way before December on that. So they're going to cross 100 quite soon by our modeling. Anyway, look, these businesses are incredible. Like I said earlier, they turn on more GPUs, they get more revenue. It's almost a direct line from all of our tracking of their data centers and their chips installations.
Gemma Allen
>> Can we do a quick hot takes, couple of questions at the end? Sure. Okay. Hyperscaler, biggest kind of, favorite hyperscaler?
Jordan Nanos
>> Favorite hyperscaler, right now, when it comes to construction, AWS is the fastest, but can't stand EFA, so I'm gonna go Oracle. Oracle. They've pioneered a lot with the multi-plane RoCE networking. Love that.
Gemma Allen
>> Might be a good time to buy Oracle stock too, right?
Jordan Nanos
>> This is not an endorsement of their stock.
Gemma Allen
>> Joking, joking.
Jordan Nanos
>> Buy the research, everybody.
Gemma Allen
>> Yes, happy with that. Chip company outside of NVIDIA.
Jordan Nanos
>> Oh, man. Startup or real product?
Gemma Allen
>> Real project.
Jordan Nanos
>> Okay. TPUs are amazing. A lot of people, there's so much demand for TPUs. They're really cool. A lot of cool stuff on the roadmap too. I'll put Trainium in there as well. I've had some good experience doing micro benchmarking with Trainium. Love the profiling. Chip Startup, I don't know. I can't pick. Don't want to give too strong an endorsement. Really impressive what a lot of them are doing. Need to see these guys produce tokens, okay? Not these deals. These cool marketing videos. Produce some tokens. Chip Startups, let's go.
Gemma Allen
>> Favorite NeoCloud?
Jordan Nanos
>> Favorite NeoCloud? CoreWeave's been the top of the Platinum tier rankings. They're a great experience. Yeah, we'll go with that.
Gemma Allen
>> Let's end with an easy one. Favorite tech leader.
Jordan Nanos
>> Favorite tech leader. Jensen.
Gemma Allen
>> in NVIDIA this year, we asked folks, who's the bigger celebrity, Jesus or Jensen? You know what the answer was?
Jordan Nanos
>> Who's Jesus?
Gemma Allen
>> Jordan Nanos, thank you so much for joining us on theCUBE and NYSE Wired.
Jordan Nanos
>> Okay, great to be here.
Gemma Allen
>> I'm Gemma Allen, coming to you from theCUBE Studio at the New York Stock Exchange. This is AI Factories, one of our programs with NYSE Wired. Thanks for watching.