This conversation examines ClusterMAX 3.0 and the evolving landscape of neocloud artificial intelligence infrastructure. Jordan Nanos of SemiAnalysis brings practical benchmarking expertise in GPU clusters and neocloud services. Nanos discusses ClusterMAX 3.0’s holistic ranking methodology, which evaluates compute, storage, networking, security and reliability through live cluster testing and customer interviews. The discussion with Gemma Allen of theCUBE covers managed clusters, inference endpoints, vertical integration and provider comparisons including Nebbius and CoreWeave.
Key takeaways include ClusterMAX’s emphasis on real-world customer experience and the trade-off between deployment speed and operational quality. Nanos highlights that vertically integrated providers tend to deliver higher reliability and support. They note that Nebbius demonstrates strong monitoring and support during testing and that chip startups face significant supply chain and data center operational challenges, as observed by theCUBE analysts. The conversation previews upcoming ClusterMAX updates 3.1 and 4.0 and addresses implications for AI infrastructure strategy and provider selection.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Jordan Nanos, SemiAnalysis
This conversation examines ClusterMAX 3.0 and the evolving landscape of neocloud artificial intelligence infrastructure. Jordan Nanos of SemiAnalysis brings practical benchmarking expertise in GPU clusters and neocloud services. Nanos discusses ClusterMAX 3.0’s holistic ranking methodology, which evaluates compute, storage, networking, security and reliability through live cluster testing and customer interviews. The discussion with Gemma Allen of theCUBE covers managed clusters, inference endpoints, vertical integration and provider comparisons including Nebbius and CoreWeave.
Key takeaways include ClusterMAX’s emphasis on real-world customer experience and the trade-off between deployment speed and operational quality. Nanos highlights that vertically integrated providers tend to deliver higher reliability and support. They note that Nebbius demonstrates strong monitoring and support during testing and that chip startups face significant supply chain and data center operational challenges, as observed by theCUBE analysts. The conversation previews upcoming ClusterMAX updates 3.1 and 4.0 and addresses implications for AI infrastructure strategy and provider selection.
>> Welcome back to theCUBE Studio here at the New York Stock Exchange. I'm Gemma Allen, co-host of NYSE Wired: AI Factories. And joining me now for a conversation about a newly released benchmark, ClusterMAX 3.0, is SemiAnalysis' very own Jordan. Jordan, welcome back to the show.
Jordan Nanos
>> Thanks for having me. Excited to be here.
Gemma Allen
>> So I've learned a few things since you were on, which honestly feels like a blink ago, but I know it was almost 6, 7 weeks ago now. One which humored me greatly is that you, Jordan, are actually Dylan Patel's man crush. I listened to you guys on the podcast. I didn't know I was talking to the most— what he referred to as stable and best looking dude in SemiAnalysis. So I wish I'd known that before. But jokes aside, you guys have obviously been very busy. You have released this— are you calling it a benchmark, a report? ClusterMAX 3.0. But I know a lot of research and analysis has gone into it. So maybe just to start, help me understand the scope of this, bring it to life here.
Jordan Nanos
>> Yeah. So ClusterMAX is a rating system or ranking system as part of the effort to assess the quality of service that different neo clouds provide to their customers. We do run a lot of benchmarks, but definitely the scope of it is more holistic in that we assess a lot of the business decisions that they've been making. As well as a lot of technical details. We look into things like performance of compute, storage, networking, the setup of the orchestration software. We focus a lot on security now, as well as reliability. And so in all of these cases, what we're doing is hands-on testing with the clusters where we work collaboratively with the providers for about a week, sometimes two. We get a small cluster and we try to just really get the experience that their customers would be experiencing themselves and tease out a lot of the details of things that we're apt to see or hear about when we actually talk to customers themselves of these neoclouds and get their feedback on their experience as well.
Gemma Allen
>> So when we talked last time, we talked about this economy that we're in, right? This demand economy, and we know that it generates certain behaviors and in some respects when we think about the performance and the longevity and the moat of the world of neoclouds, especially some of the newbies and the accelerators, 10 years from now, we haven't really separated the men from the boys yet, right? It's not so certain. What are your thoughts in terms of what you've seen in terms of who is delivering customer experience that's based on true customer needs and delivering solutions, and who is kind of just selling access? Help me first, start with that.
Jordan Nanos
>> Yeah, so I think there's different levels or different types of services that these providers actually give to their customers. At the basic level, you can think of a lot of these guys as being in the business of data center construction. They sell a finished product, which is like an empty shell. And then another layer on that would be bare metal, where people actually own the GPUs. They worry about procurement, design, they get them installed, they burn them in. And many people will buy just bare metal GPUs at massive scale from there. In fact, that's the section of the market where most of the transactions are happening in terms of transaction volume. It's also where most of the revenue is happening. And then on top of bare metal, which is really where we start to do our testing with ClusterMAX, you have managed clusters and inference endpoints. These are both things that we research a ton at SemiAnalysis, but the core of ClusterMAX is about that managed cluster experience because this is the background by which the biggest companies in AI in the world like OpenAI or Anthropic were buying their first managed clusters. And like you said, I think, yeah, you can separate the market in terms of who is actually offering services across all of those different tiers, which is to say vertically integrated clouds are the ones that we see the most high-quality services from, the people who have full control over everything from the production of these custom facilities that are built for the latest and greatest NVIDIA GPUs or custom chips. They deploy them, they operate them, they manage a lot of the software, and they really get embedded with the teams at their customers to make sure that they're getting great performance out of the cluster.
Gemma Allen
>> We talked the last time, about this whole idea of, throughput, goodput, right? And what that means. And maybe there's been some misalignment or some, misunderstanding as to how that performance is ranked. But one thing which does seem to be abundantly clear day in, day out in this show with all the folks we talk to across both the buyers and the builders is speed, right? Like speed is everything. What are your thoughts on that? Like how core, how absolutely centric is that to your overall outcome here? Is it really a speed test alone? Help me understand that, or even debate that.
Jordan Nanos
>> Yeah. And well, what's happened in the market is that there's so much demand for these GPUs that there's no scenario where people just have GPUs they've purchased speculatively, but they're waiting to rent to people and you can turn them on tomorrow. Everybody is making prepayments. They're putting orders in ahead of schedule and they're waiting months in most cases to have these facilities built and the GPUs installed. And so it is about speed, but it's also about quality, which is to say, as you're deploying, let's say, 18,000 GPUs, you might be deploying tranches of 2,000 or 4,000 at a time and then handing over each of those to a given customer. Some are all to the same customer where they just take tranches at a time and others to different customers as you make progress on your buildout. Any delay for one little tranche, it just cascades and everything else is delayed. And so it is about speed, but it's not like you can cut corners in terms of how you're setting up your power, your cooling, your internet, your reliability systems like monitoring. we know about these example deployments where GPUs have shown up at a data center site way ahead of schedule. And they have to have them rolled in and racked, but the white space in the data center isn't set up, the walls aren't even painted, there's huge air quality, dust issues, and this impacts the performance of the GPUs when people turn them on. You have to come in later and clean everything out, you have to clean the fiber ends, you have to assess if these cooling systems actually work, you have to redo almost the whole install in some cases. And so it is about speed, but you can't sacrifice quality when you have speed.
Gemma Allen
>> And do you think that folks that are winning this race are centralizing those kind of core services as add-ons? We talk a lot about performance engineering and efficiency and all of the nuts and bolts that happens from the perspective of data centers. But there are kind of different variations of what you buy, right, and what you get. Talk a little bit about that additional slice of the market. Like, what are you seeing from the perspective of TAM and how is it actually helping them rise up in your rankings?
Jordan Nanos
>> Yeah, people make good and bad decisions in terms of the upfront design that they get locked into and they can't really change. The most obvious example is in networking with NVIDIA GPUs. NVIDIA sells standard scale-out networking. A lot of customers will change this. They'll go with some networking switches from a vendor that sells RoCE, RDMA over Ethernet switches that maybe don't perform as well, or it takes longer for them to support the latest libraries. Others have custom networks, like we always pick on AWS with EFA. Customers just don't like it. They don't want to have to wait to support the latest open source releases of a lot of these software frameworks and worry about the network supporting it. They just want it to work. And then there's a huge variation in terms of how people design networks for AMD GPUs where they don't have a standard networking offering. AMD doesn't have a scale-out network switch or a scale-up network switch, frankly. And then all of the chip startups, it's kind of different when you're talking about, yeah, the long tail of all the different startups that we're gonna see more of over time. You can look at things like storage, you can look at monitoring systems and reliability and how that impacts goodput as well. But if you make a decision of how you design the network, even assuming that you're using NVIDIA components, that can have an impact on performance. And people like one over the other. It really just depends on the customer's requirements. And the last thing to say is just that customers want to be involved in this process. If they are ordering GPUs from you and they're waiting 6 months and you haven't completed the order yet, people who have opinions about the network, they want that to be reflected in the design of the system that you're actually going and deploying, because this is much more of a partnership relationship than it is a negotiation between a customer and a provider.
Gemma Allen
>> For sure. So let's talk about the ranking for a second. Let's actually look at this report. So you have some interesting revelations there. Like, let's talk about Nebius. As an example, right? So you're talking to Roman, you're like, hey, you're doing a great job, buddy. What exactly is changing in their underlying performance, right? Or their unique customer experience? What drove those numbers for you?
Jordan Nanos
>> Well, on the business level, Nebius has really just been the default, cloud of choice for the mid-tier of the market for quite a while. there was, I think, a gap left with CoreWeave when CoreWeave got just a lot bigger and focused on the absolute biggest customers in the world. And so customers who want 512 to 1,000 to 2,000 GPUs, they're increasingly going with Nebius. The reason why we think is technical, which is to say Nebius was some of the first— they were the first to turn on B200 HGX instances in December of last year. A lot of people are still installing B200s for the first time now. They've got GB200 support from scratch. They build a lot of these data centers themselves. They have the on-site technicians and people trained. Their monitoring stack is really solid. They had no health checks this time last year. They are, or they were just implementing them and kind of building a dataset from this. This has almost completely changed. They have super reliable clusters. They also have quite strong support. A lot of people have told us that their experience with their support engineers is really strong and they're quite technical and they really know what they're doing. And they don't see a bunch of, they see strong retention on these support engineers where they don't have to work with a new person every 2 months or something. And that's our experience too when we're doing our testing, is we've always just enjoyed working with the Nebbius technical team. So yeah, they're on track to just keep doing well. And the signal for people is really like, who are you going to trust for your big Vera Rubin deployments next year? And Nebbius is right up there with CoreWeave as the top provider to consider for
Gemma Allen
>> that.So you mentioned that, talking to their customers— I probably should have started with this, so I'm going to loop back around here, but In terms of the scale of this research, right, you've spoken to 200 customers across 77, cloud players, or I should call them AI clouds, however you term them, right? Neo-cloud. Some of them don't want to be called neo-clouds anymore. I don't actually know what the safe word is these days. But, so 77 and then it kind of filters down, right? So how long, how extensive is this research? Say for Nebbius, how many customers, for example, Did you guys speak to how— what does the testing look like?
Jordan Nanos
>> Okay, so the hands-on testing that we do with Nebbius takes about a week or two. It depends on the size of the cluster. With them, it was like a little over 2 weeks because we tested multiple clusters, both the GB300 instance, which is like a full rack, actually 2 racks, and then like a smaller B300 cluster. In terms of what customers we're talking to, This is everybody in the industry. It cuts across all sorts of different industries, frankly. So it's startups in Silicon Valley that are doing the typical thing where they're trying to just train a language model and make it to that next round of funding because they are outperforming the frontier on some specific task, or they're much more efficient, or things like that. There's people that are building agents on top of these models where they need custom smaller models. There's people exploring all sorts of other industries, let's say weather prediction or voice or autonomous vehicle or video generation or material science. Drug discovery. There's all sorts of different startups that all have a lot of funding and are really in need of GPUs where they spend like 60, 70, sometimes 80% of their funding all on just one big GPU cluster from a provider. And then of course there's lots of people across government entities. There's typical Fortune 500 banks and telcos and people that give their feedback. They're entering the market with the neoclouds, but it's driven a lot by the startup companies that are not small and not small valueand are the biggest customers of these neoclouds today in terms of who we get feedback from.
Gemma Allen
>> Wow. I wanna ask you about something that kind of sparked my interest when I listened to you on the podcast, on your podcast recently with Dylan, which by the way was a great episode. I encourage everyone to listen to it. It was funny. and that's about accelerators, right? He obviously has some skepticism. Criticism in that space about these huge commits that are being shared in the market, you know, folks committing $1 billion for chips that don't yet exist. I'm interested to get your perspective because, again, call me naive, but you would imagine that there's some level of due diligence done in this space. But where do you think that these kind of gaps are emerging? And what do you think the reality is in terms of what's being sold and promised versus what's being delivered?
Jordan Nanos
>> Well, I think the gap is just time for a lot of these people. And I don't actually believe these criticisms of the chip startups that they're just like total smoke and mirrors and fake and it doesn't work. It's not going to power on. We have not seen that for a long time in chip startup land where people just tape out a dud, probably close to 5+ years for a big notable one that I can think of. And so what we're seeing is there's a class of chip startups that's 10+ years old. This is Cerebras, Groq, SambaNova, and a few other ones that did not work out. And then there's this new class which is under 5 years old, let's say, or around 5 years old. And a lot of them are bringing products to market this year or next year. And I think, we can name off, but these are companies like Etched and Positron and MatX and Tenstorrent. And OLIX and Tensordyne. There's all sorts of them. And I think in general, I would expect that the upper bound of achievement would be roughly the OpenAI Jalapeño custom chip program, which went from RTL to tape-out in 7 months and is outperforming Blackwell already. Very strong signal that you can do it for yourself if you have the money and the talent on the team. And I think a lot of these people do. And then there's a secondary consideration, which is, How do you actually deploy this stuff? I think the first hurdle for all of these chip startups was always, can you take what's on paper, actually put it into an application, or produce something faster or cheaper than what others can do? And the simplest way people do this today is by producing tokens on an API endpoint. And we saw this with Groq and Cerebras and SambaNova. Like, this was their go-to-market strategy. Log into the API, mint API key, start issuing tokens. Worked out for all of them. They're all doing fine. I expect others to follow that exact same playbook for their initial capacity. And if they are not doing a public endpoint, I would expect it's typically because they have a dedicated customer that's bought out all the initial capacity and they just don't need to. And they're focused on delivering the needs of these customers. I do believe them. When they say this. Now, the challenge that they all have to contend with is the fact that if they believe they're the NVIDIA killer, they need to do supply chain and data center operations. And if you're going to actually produce 100 megawatts of these chips and turn them on in a data center, this means you have to compete with NVIDIA and others in the markets of getting land and power and construction companies signed up and then training technicians and allocation of HBM or CoWoS in the supply chain going all the way back to the fabs. This is a huge challenge, it's a huge operations challenge and they're going to be put to the test there in some ways, I think more critically than writing the software stack, designing a good chip and actually taping it out.
Gemma Allen
>> So I want to stay on that for my last question to you. You mentioned LPU and the inference chip, right? And the fact that this competitive landscape, it's getting dicey. I think dicey is probably the best word to use. It's interesting, you've got hyperscalers, Neoclouds, GPU providers, and a lot of companies like OpenAI, for example, building their own infrastructure, right? What are your thoughts on the competitive advantage of that, especially for some of those frontier labs? You've got Meta, Google. I know you've had some interesting thoughts on them in the past, but they are at the end of the day, they control so much access to compute already, right? They're gatekeepers. They have a particular leg up on an industry, maybe not OpenAI because they're obviously a frontier, relative newbie, but they're a huge titan now. What do you think from the perspective of this kind of competitive landscape when you see this happen? There was a time in tech where people kind of stayed in their zone. In this world, it seems like everyone is everything.
Jordan Nanos
>> Yeah. And I think a lot of people, including myself, have that personal experience where the power of using coding agents today means that a lot of other, a lot of knowledge is now at your fingertips and accessible where you can start making contributions and learning about things that you thought might be outside of your comfort zone. I think this applies to a lot of companies as well, where in the past you would say, stay in your lane, you're a social media company, stay in your lane, you're like an AI research lab. But now they can produce chips, they can build data centers, they can certainly contend with others in this fully integrated AI future where you've got everything from a mobile application to the model itself to the chip that it runs on to the data center facility it runs in. And I think a lot of people— well, the biggest companies in the world have the capital to approach this problem seriously. There's a lot of capital expense that's required to just get started when it comes to building a data center or taping out a chip. We're talking about hundreds of millions of dollars, but probably billions of dollars to take it seriously. So only certain companies can even access that sort of stuff. And I think you also need a bunch of capital and weight and brand to be able to attract the talent, the team that actually executes on a lot of this stuff.
Gemma Allen
>> So ClusterMAX 3.0, what's the iteration here? What's the cadence? Do you release this annually, quarterly? Do you update it? Talk us through what comes next for this.
Jordan Nanos
>> Yeah, 3.1, we're doing testing right now. That's just a minor revision where we're retesting providers that have had major improvements in their software stack or other operations, as well as getting new providers that we've either missed or are launching a service onto the list. It's a minor update, which means again, we're not retesting everybody we've tested before. And we're just testing the same chips, we're targeting the B300 or GB300s with 800 gig networking. 4.0, the next major version, that's probably less than a year away. But we'll be testing for that in the spring or summer next year. And the focus will be on Vera Rubin, the next generation of NVIDIA GPUs and everything else that launches around that same time. Wow.
Gemma Allen
>> Okay. We certainly look forward to following your analysis. It's very central to the conversations we have here every day on NYSE Wired AI Factories. Jordan, great to see you again and thanks so much for joining us.
Jordan Nanos
>> Thanks for having me. Great to be here.
Gemma Allen
>> I'm Gemma Allen here at theCUBE Studio at the New York Stock Exchange. This is NYSE Wired's AI Factories. Thanks for watching.
>> Welcome back to theCUBE Studio here at the New York Stock Exchange. I'm Gemma Allen, co-host of NYSE Wired: AI Factories. And joining me now for a conversation about a newly released benchmark, ClusterMAX 3.0, is SemiAnalysis' very own Jordan. Jordan, welcome back to the show.
Jordan Nanos
>> Thanks for having me. Excited to be here.
Gemma Allen
>> So I've learned a few things since you were on, which honestly feels like a blink ago, but I know it was almost 6, 7 weeks ago now. One which humored me greatly is that you, Jordan, are actually Dylan Patel's man crush. I listened to you guys on the podcast. I didn't know I was talking to the most— what he referred to as stable and best looking dude in SemiAnalysis. So I wish I'd known that before. But jokes aside, you guys have obviously been very busy. You have released this— are you calling it a benchmark, a report? ClusterMAX 3.0. But I know a lot of research and analysis has gone into it. So maybe just to start, help me understand the scope of this, bring it to life here.
Jordan Nanos
>> Yeah. So ClusterMAX is a rating system or ranking system as part of the effort to assess the quality of service that different neo clouds provide to their customers. We do run a lot of benchmarks, but definitely the scope of it is more holistic in that we assess a lot of the business decisions that they've been making. As well as a lot of technical details. We look into things like performance of compute, storage, networking, the setup of the orchestration software. We focus a lot on security now, as well as reliability. And so in all of these cases, what we're doing is hands-on testing with the clusters where we work collaboratively with the providers for about a week, sometimes two. We get a small cluster and we try to just really get the experience that their customers would be experiencing themselves and tease out a lot of the details of things that we're apt to see or hear about when we actually talk to customers themselves of these neoclouds and get their feedback on their experience as well.
Gemma Allen
>> So when we talked last time, we talked about this economy that we're in, right? This demand economy, and we know that it generates certain behaviors and in some respects when we think about the performance and the longevity and the moat of the world of neoclouds, especially some of the newbies and the accelerators, 10 years from now, we haven't really separated the men from the boys yet, right? It's not so certain. What are your thoughts in terms of what you've seen in terms of who is delivering customer experience that's based on true customer needs and delivering solutions, and who is kind of just selling access? Help me first, start with that.
Jordan Nanos
>> Yeah, so I think there's different levels or different types of services that these providers actually give to their customers. At the basic level, you can think of a lot of these guys as being in the business of data center construction. They sell a finished product, which is like an empty shell. And then another layer on that would be bare metal, where people actually own the GPUs. They worry about procurement, design, they get them installed, they burn them in. And many people will buy just bare metal GPUs at massive scale from there. In fact, that's the section of the market where most of the transactions are happening in terms of transaction volume. It's also where most of the revenue is happening. And then on top of bare metal, which is really where we start to do our testing with ClusterMAX, you have managed clusters and inference endpoints. These are both things that we research a ton at SemiAnalysis, but the core of ClusterMAX is about that managed cluster experience because this is the background by which the biggest companies in AI in the world like OpenAI or Anthropic were buying their first managed clusters. And like you said, I think, yeah, you can separate the market in terms of who is actually offering services across all of those different tiers, which is to say vertically integrated clouds are the ones that we see the most high-quality services from, the people who have full control over everything from the production of these custom facilities that are built for the latest and greatest NVIDIA GPUs or custom chips. They deploy them, they operate them, they manage a lot of the software, and they really get embedded with the teams at their customers to make sure that they're getting great performance out of the cluster.
Gemma Allen
>> We talked the last time, about this whole idea of, throughput, goodput, right? And what that means. And maybe there's been some misalignment or some, misunderstanding as to how that performance is ranked. But one thing which does seem to be abundantly clear day in, day out in this show with all the folks we talk to across both the buyers and the builders is speed, right? Like speed is everything. What are your thoughts on that? Like how core, how absolutely centric is that to your overall outcome here? Is it really a speed test alone? Help me understand that, or even debate that.
Jordan Nanos
>> Yeah. And well, what's happened in the market is that there's so much demand for these GPUs that there's no scenario where people just have GPUs they've purchased speculatively, but they're waiting to rent to people and you can turn them on tomorrow. Everybody is making prepayments. They're putting orders in ahead of schedule and they're waiting months in most cases to have these facilities built and the GPUs installed. And so it is about speed, but it's also about quality, which is to say, as you're deploying, let's say, 18,000 GPUs, you might be deploying tranches of 2,000 or 4,000 at a time and then handing over each of those to a given customer. Some are all to the same customer where they just take tranches at a time and others to different customers as you make progress on your buildout. Any delay for one little tranche, it just cascades and everything else is delayed. And so it is about speed, but it's not like you can cut corners in terms of how you're setting up your power, your cooling, your internet, your reliability systems like monitoring. we know about these example deployments where GPUs have shown up at a data center site way ahead of schedule. And they have to have them rolled in and racked, but the white space in the data center isn't set up, the walls aren't even painted, there's huge air quality, dust issues, and this impacts the performance of the GPUs when people turn them on. You have to come in later and clean everything out, you have to clean the fiber ends, you have to assess if these cooling systems actually work, you have to redo almost the whole install in some cases. And so it is about speed, but you can't sacrifice quality when you have speed.
Gemma Allen
>> And do you think that folks that are winning this race are centralizing those kind of core services as add-ons? We talk a lot about performance engineering and efficiency and all of the nuts and bolts that happens from the perspective of data centers. But there are kind of different variations of what you buy, right, and what you get. Talk a little bit about that additional slice of the market. Like, what are you seeing from the perspective of TAM and how is it actually helping them rise up in your rankings?
Jordan Nanos
>> Yeah, people make good and bad decisions in terms of the upfront design that they get locked into and they can't really change. The most obvious example is in networking with NVIDIA GPUs. NVIDIA sells standard scale-out networking. A lot of customers will change this. They'll go with some networking switches from a vendor that sells RoCE, RDMA over Ethernet switches that maybe don't perform as well, or it takes longer for them to support the latest libraries. Others have custom networks, like we always pick on AWS with EFA. Customers just don't like it. They don't want to have to wait to support the latest open source releases of a lot of these software frameworks and worry about the network supporting it. They just want it to work. And then there's a huge variation in terms of how people design networks for AMD GPUs where they don't have a standard networking offering. AMD doesn't have a scale-out network switch or a scale-up network switch, frankly. And then all of the chip startups, it's kind of different when you're talking about, yeah, the long tail of all the different startups that we're gonna see more of over time. You can look at things like storage, you can look at monitoring systems and reliability and how that impacts goodput as well. But if you make a decision of how you design the network, even assuming that you're using NVIDIA components, that can have an impact on performance. And people like one over the other. It really just depends on the customer's requirements. And the last thing to say is just that customers want to be involved in this process. If they are ordering GPUs from you and they're waiting 6 months and you haven't completed the order yet, people who have opinions about the network, they want that to be reflected in the design of the system that you're actually going and deploying, because this is much more of a partnership relationship than it is a negotiation between a customer and a provider.
Gemma Allen
>> For sure. So let's talk about the ranking for a second. Let's actually look at this report. So you have some interesting revelations there. Like, let's talk about Nebius. As an example, right? So you're talking to Roman, you're like, hey, you're doing a great job, buddy. What exactly is changing in their underlying performance, right? Or their unique customer experience? What drove those numbers for you?
Jordan Nanos
>> Well, on the business level, Nebius has really just been the default, cloud of choice for the mid-tier of the market for quite a while. there was, I think, a gap left with CoreWeave when CoreWeave got just a lot bigger and focused on the absolute biggest customers in the world. And so customers who want 512 to 1,000 to 2,000 GPUs, they're increasingly going with Nebius. The reason why we think is technical, which is to say Nebius was some of the first— they were the first to turn on B200 HGX instances in December of last year. A lot of people are still installing B200s for the first time now. They've got GB200 support from scratch. They build a lot of these data centers themselves. They have the on-site technicians and people trained. Their monitoring stack is really solid. They had no health checks this time last year. They are, or they were just implementing them and kind of building a dataset from this. This has almost completely changed. They have super reliable clusters. They also have quite strong support. A lot of people have told us that their experience with their support engineers is really strong and they're quite technical and they really know what they're doing. And they don't see a bunch of, they see strong retention on these support engineers where they don't have to work with a new person every 2 months or something. And that's our experience too when we're doing our testing, is we've always just enjoyed working with the Nebbius technical team. So yeah, they're on track to just keep doing well. And the signal for people is really like, who are you going to trust for your big Vera Rubin deployments next year? And Nebbius is right up there with CoreWeave as the top provider to consider for
Gemma Allen
>> that.So you mentioned that, talking to their customers— I probably should have started with this, so I'm going to loop back around here, but In terms of the scale of this research, right, you've spoken to 200 customers across 77, cloud players, or I should call them AI clouds, however you term them, right? Neo-cloud. Some of them don't want to be called neo-clouds anymore. I don't actually know what the safe word is these days. But, so 77 and then it kind of filters down, right? So how long, how extensive is this research? Say for Nebbius, how many customers, for example, Did you guys speak to how— what does the testing look like?
Jordan Nanos
>> Okay, so the hands-on testing that we do with Nebbius takes about a week or two. It depends on the size of the cluster. With them, it was like a little over 2 weeks because we tested multiple clusters, both the GB300 instance, which is like a full rack, actually 2 racks, and then like a smaller B300 cluster. In terms of what customers we're talking to, This is everybody in the industry. It cuts across all sorts of different industries, frankly. So it's startups in Silicon Valley that are doing the typical thing where they're trying to just train a language model and make it to that next round of funding because they are outperforming the frontier on some specific task, or they're much more efficient, or things like that. There's people that are building agents on top of these models where they need custom smaller models. There's people exploring all sorts of other industries, let's say weather prediction or voice or autonomous vehicle or video generation or material science. Drug discovery. There's all sorts of different startups that all have a lot of funding and are really in need of GPUs where they spend like 60, 70, sometimes 80% of their funding all on just one big GPU cluster from a provider. And then of course there's lots of people across government entities. There's typical Fortune 500 banks and telcos and people that give their feedback. They're entering the market with the neoclouds, but it's driven a lot by the startup companies that are not small and not small valueand are the biggest customers of these neoclouds today in terms of who we get feedback from.
Gemma Allen
>> Wow. I wanna ask you about something that kind of sparked my interest when I listened to you on the podcast, on your podcast recently with Dylan, which by the way was a great episode. I encourage everyone to listen to it. It was funny. and that's about accelerators, right? He obviously has some skepticism. Criticism in that space about these huge commits that are being shared in the market, you know, folks committing $1 billion for chips that don't yet exist. I'm interested to get your perspective because, again, call me naive, but you would imagine that there's some level of due diligence done in this space. But where do you think that these kind of gaps are emerging? And what do you think the reality is in terms of what's being sold and promised versus what's being delivered?
Jordan Nanos
>> Well, I think the gap is just time for a lot of these people. And I don't actually believe these criticisms of the chip startups that they're just like total smoke and mirrors and fake and it doesn't work. It's not going to power on. We have not seen that for a long time in chip startup land where people just tape out a dud, probably close to 5+ years for a big notable one that I can think of. And so what we're seeing is there's a class of chip startups that's 10+ years old. This is Cerebras, Groq, SambaNova, and a few other ones that did not work out. And then there's this new class which is under 5 years old, let's say, or around 5 years old. And a lot of them are bringing products to market this year or next year. And I think, we can name off, but these are companies like Etched and Positron and MatX and Tenstorrent. And OLIX and Tensordyne. There's all sorts of them. And I think in general, I would expect that the upper bound of achievement would be roughly the OpenAI Jalapeño custom chip program, which went from RTL to tape-out in 7 months and is outperforming Blackwell already. Very strong signal that you can do it for yourself if you have the money and the talent on the team. And I think a lot of these people do. And then there's a secondary consideration, which is, How do you actually deploy this stuff? I think the first hurdle for all of these chip startups was always, can you take what's on paper, actually put it into an application, or produce something faster or cheaper than what others can do? And the simplest way people do this today is by producing tokens on an API endpoint. And we saw this with Groq and Cerebras and SambaNova. Like, this was their go-to-market strategy. Log into the API, mint API key, start issuing tokens. Worked out for all of them. They're all doing fine. I expect others to follow that exact same playbook for their initial capacity. And if they are not doing a public endpoint, I would expect it's typically because they have a dedicated customer that's bought out all the initial capacity and they just don't need to. And they're focused on delivering the needs of these customers. I do believe them. When they say this. Now, the challenge that they all have to contend with is the fact that if they believe they're the NVIDIA killer, they need to do supply chain and data center operations. And if you're going to actually produce 100 megawatts of these chips and turn them on in a data center, this means you have to compete with NVIDIA and others in the markets of getting land and power and construction companies signed up and then training technicians and allocation of HBM or CoWoS in the supply chain going all the way back to the fabs. This is a huge challenge, it's a huge operations challenge and they're going to be put to the test there in some ways, I think more critically than writing the software stack, designing a good chip and actually taping it out.
Gemma Allen
>> So I want to stay on that for my last question to you. You mentioned LPU and the inference chip, right? And the fact that this competitive landscape, it's getting dicey. I think dicey is probably the best word to use. It's interesting, you've got hyperscalers, Neoclouds, GPU providers, and a lot of companies like OpenAI, for example, building their own infrastructure, right? What are your thoughts on the competitive advantage of that, especially for some of those frontier labs? You've got Meta, Google. I know you've had some interesting thoughts on them in the past, but they are at the end of the day, they control so much access to compute already, right? They're gatekeepers. They have a particular leg up on an industry, maybe not OpenAI because they're obviously a frontier, relative newbie, but they're a huge titan now. What do you think from the perspective of this kind of competitive landscape when you see this happen? There was a time in tech where people kind of stayed in their zone. In this world, it seems like everyone is everything.
Jordan Nanos
>> Yeah. And I think a lot of people, including myself, have that personal experience where the power of using coding agents today means that a lot of other, a lot of knowledge is now at your fingertips and accessible where you can start making contributions and learning about things that you thought might be outside of your comfort zone. I think this applies to a lot of companies as well, where in the past you would say, stay in your lane, you're a social media company, stay in your lane, you're like an AI research lab. But now they can produce chips, they can build data centers, they can certainly contend with others in this fully integrated AI future where you've got everything from a mobile application to the model itself to the chip that it runs on to the data center facility it runs in. And I think a lot of people— well, the biggest companies in the world have the capital to approach this problem seriously. There's a lot of capital expense that's required to just get started when it comes to building a data center or taping out a chip. We're talking about hundreds of millions of dollars, but probably billions of dollars to take it seriously. So only certain companies can even access that sort of stuff. And I think you also need a bunch of capital and weight and brand to be able to attract the talent, the team that actually executes on a lot of this stuff.
Gemma Allen
>> So ClusterMAX 3.0, what's the iteration here? What's the cadence? Do you release this annually, quarterly? Do you update it? Talk us through what comes next for this.
Jordan Nanos
>> Yeah, 3.1, we're doing testing right now. That's just a minor revision where we're retesting providers that have had major improvements in their software stack or other operations, as well as getting new providers that we've either missed or are launching a service onto the list. It's a minor update, which means again, we're not retesting everybody we've tested before. And we're just testing the same chips, we're targeting the B300 or GB300s with 800 gig networking. 4.0, the next major version, that's probably less than a year away. But we'll be testing for that in the spring or summer next year. And the focus will be on Vera Rubin, the next generation of NVIDIA GPUs and everything else that launches around that same time. Wow.
Gemma Allen
>> Okay. We certainly look forward to following your analysis. It's very central to the conversations we have here every day on NYSE Wired AI Factories. Jordan, great to see you again and thanks so much for joining us.
Jordan Nanos
>> Thanks for having me. Great to be here.
Gemma Allen
>> I'm Gemma Allen here at theCUBE Studio at the New York Stock Exchange. This is NYSE Wired's AI Factories. Thanks for watching.