Emilio Andere, Wafer | theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
>> Palo Alto studio connecting Silicon Valley and Wall Street. I'm John Furrier, host of theCUBE, here with Dave Vellante, my co-host. We are here in theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is the NYSE Wired program and open community, and this is our AI Factory series. where we talk to the leaders who are making it happen. Obviously, AI factories continuing to power the massive buildout of the AI infrastructure, which is accelerating new architectures, new systems designs, accelerating the software paradigm, bringing in all kinds of new agentic and/or, AI-related applications. Our next guest is Emilio Andere, co-founder and CEO of Wafer, Bay Area-based. Emilio, great to have you on.
Emilio Andere
>> Thank you for having me.
John Furrier
>> Cool. First of all, give us some of the stats. When you guys were founded.
John Furrier
>> Yeah.
John Furrier
>> What's the status of the funding? Did you do a pre-seed at $200 million? Give us that.
Emilio Andere
>> We were actually a pretty simple pre-seed. It was Y Combinator. Okay, all right. So, me and my co-founder, we were roommates in college at UChicago, and we decided to start Wafer about a year and a half ago, and we did YC a couple of weeks after, so that was our sort of pre-seed. A pre-seed. Then we raised a seed, we raised about $4 million, about 3 months after that. And then most recently, about a month ago, actually a couple weeks ago, we announced our $40 million Series A, so just to grow the team and keep expanding the
John Furrier
>> technology.So pretty fast
Emilio Andere
>> progression.Yeah, it's been pretty
John Furrier
>> fast.All right, so what's the status now? Take what you guys do, explain the value proposition. What's your main problem that you're going
Emilio Andere
>> after?Yeah, totally. So when companies want to use LLMs and want to run on GPUs, it's actually quite hard to get those GPUs to run very fast for any LLM. Let's say you want to use DeepSpeed, Qwen, whatever you would want to use. It's really hard to actually get those GPUs and extract their maximum performance. So the compute, the sort of hardware is there, but in terms of software, it's really hard to write those kernels. It's hard to do the end-to-end optimization so that your LLM actually is extracting the best performance per dollar invested into that GPU.
Emilio Andere
>> Yeah.
Emilio Andere
>> So what we do is we sort of found this scarcity in the number of GPU engineers that could do this work. So we decided that it was very obvious to us that the, yeah, what you should be doing to optimize your GPU should be with AI itself. We call this kind of AI to optimize AI. It's like, why would you use scarce humans that know how to actually do this low-level work? Yeah. When agents are getting so good at coding and you can finally use these agents as performance engineers, GPU engineers.
John Furrier
>> You're right, for deployed AI.
Emilio Andere
>> Exactly.
John Furrier
>> Essentially.
Emilio Andere
>> Exactly.
John Furrier
>> I mean, back in the old days, not to date myself because you're the young gun kicking ass, taking names in the AI world, we would buy a server and load Linux on it. There was a software stack well understood that gets you up and running. You got an OS and then you put your stuff on there and then you're off to the races, connect to the network. Here, AI, it's hard to really identify what the stack is for a GPU.
John Furrier
>> Yeah.
John Furrier
>> And it's not just the GPU either. There's other stuff going around it. So one of the things we observed when we started the series 2 years ago was, I love this AI factories concept, ship a rack to someone. What the hell do you run on it? And then now we're seeing all kinds of disaggregated serving. The caches are getting bloated. So the entire performance engine is getting clogged up. Is that kind of the area you guys are going in and trying to nail that piece?
Emilio Andere
>> Yeah, yeah, exactly. So it's getting super complicated. And as you said, one of the other reasons it's getting super complicated is because there's other chips coming into play. So now you not only have to figure out the stack for NVIDIA, you also have to figure it out for AMD and you have to figure it out for Google TPUs and for Trainium. And most of the traditional inference companies that came out in 2021, that are now giants, they are basically just fully on NVIDIA and that's just because the industry was there at that time. But today we're in a much more heterogeneous world. As you said, people are combining Prefill and Decode chips in very interesting ways. You saw AMD do that with Cerebras. You saw NVIDIA do that with Groq. There's clearly a lot of stuff going
John Furrier
>> on.the Groq acquisition or the AccuHire, however you want to frame it, yeah, really kind of changed the direction of the industry. I'd love to get your comments on that because at that time Cerebras was not viewed as a serious player in inference, although they had the big chip. We covered them. We actually liked what they were doing. Half the industry thought it wouldn't work. You had other people trying to go after inference. But NVIDIA set the agenda just by buying Groq, a separate thing. Yeah, kind of changed the inference game, kind of almost segments training out. What's your opinion? Did I get that right? How would you describe that moment? Because it really was not well reported that that actually changed the trajectory of the innovation.
Emilio Andere
>> It did. Yeah. And you did get that right. I think it was similar to NVIDIA's Mellanox acquisition many years before. I think it was honestly just Jensen and the team being just visionaries. Of just like, hey, clearly there's a gap in how well GPUs can serve certain workloads. And clearly there's a market for people who are willing to pay more money for faster tokens. Just give me better latency for every one of my applications and I will pay you more per token that you generate. And he was able to identify that market and just obviously with the acquisition, make it even more serious.
John Furrier
>> Yeah, because on the last earnings call, they actually are now quantifying dollar value per gigawatt metrics. So he's already moving the metrics. You got to give NVIDIA props. They have actually accelerated not just on the financing side, but on the business side. They're pushing the envelope. Oh yeah, they're even in politics. So they're shaping the agenda. You got to give them props for that. All right. So the question really now is, okay, AI factories, we all see the big gigawatt factories, Crusoe Nscale, monster campuses. But that's going to be the big token factories, almost like oil refineries with tokens, right? Think about it that way. Now you have people at the enterprise going, hey, I can't afford $10 billion of CapEx. I just have some infrastructure. Maybe I'll do a deal with Nscale or CoreWeave, but I want to get an on-premise capacity, but I'm not going to have the footprint. So that's going to change the scope of my software stack.
Emilio Andere
>> Yeah, yeah.
John Furrier
>> Do you solve that problem? Could you be a solution for them?
Emilio Andere
>> Yeah. So our current solution and the thing that's been having the most growth is just going to customers and saying, we can give you the best performance in the market because we have these agents that we can deploy that will tune every part of the stack and then serving them the best performance per dollar. But we do it on our own GPUs. So we would go to companies like CoreWeave, Nebius and get compute from them and then just serve them as the wafer cloud. Right. I think there is a very interesting future where we do on-prem deployments because the demand is there.
John Furrier
>> So you're optimizing the raw power from the people who actually have gear.
John Furrier
>> Exactly.
John Furrier
>> Seems to work.
Emilio Andere
>> Yeah. So we're above the sort of— we rent out these long-term commitments from the neo clouds like the ones you mentioned and then serve AI-native startups, very fast inference.
John Furrier
>> Do you see a bifurcation between, say, I want to be vertically integrated. I'm going to just get my GPUs. I'm just going to serve, be a service provider and also offer raw takeout. Takeout.
Emilio Andere
>> Well, yeah. Yeah. So it's a good question. I think I do, but I do think the people that verticalize will win. I think it's a very important strategy. You basically, as you know, probably better than most, there's this— you have the inference providers there at the layer with the demand. You have the Fireworks, Baseten of the world, and they traditionally don't own GPUs themselves. They just rent it out from these data center companies, but they're trying to go to that level because they want to increasingly own their own compute. So they're kind of going down. And then you have the traditional data center companies like CoreWeave and Nebius, and they're trying to go up because they realize that there's better margins on tokens than just reselling compute. So they're like, hey, now we're a token factory. So you have this middle layer that in my opinion will get squeezed. And we used to do this. We used to sell our engine to other inference companies. But then we realized that we were going to get squeezed out. It was time to just go serve tokens directly and then verticalize at some point or we would get squeezed. So we do want to end up doing—
John Furrier
>> and there's also a lot of stuff going on in that serving. that's why I brought up disaggregated serving. I think that's also a tell to disaggregate infrastructure at the edge. So you say, okay, that's a symptom of the growth, but that also speaks to how hard it is to manage those resources. So if you need to do resource management observability, you can't just outsource that. Yeah, you got to have it in the stack. You agree?
Emilio Andere
>> Yeah, I do agree. It's getting much harder. Like there's so much increasingly complex stuff that's going on in these AI deployments. If we talked about disaggregating Prefill and Decode, different chips coming along, different serving engines, it's just getting so complicated that it just stops making sense for companies to, at least at the smaller scale, do it themselves. And instead it's— all right, so let's just say I'm a customer.
John Furrier
>> Yeah, Cube Cloud. We have a Cube Cloud on AWS with a lot of data. I want to stand up, turn on an AI factory. Yeah, I don't have any GPU engineers. Maybe I can't even get a hold of GPUs. I want to deploy a software stack that ties all my existing stuff together. Would I call you guys for that? What would it be?
Emilio Andere
>> Yeah, you would basically call us and you would say, hey, here's the workload I want to run. Here's the model I want to run for this workload. And here's the chip I want to run for that workload. And that's a huge permutation space. that's a huge space. And we deploy our agents to find the best kernels, the best serving engine. It basically— you can think of it as generating the entire stack just for your use case. You can finally do this because we're not bound by human ability.
John Furrier
>> You take the template inbound logistics of what I want.
John Furrier
>> Exactly.
John Furrier
>> And it's almost like cloud configuring EC2 or whatever.
Emilio Andere
>> Instances. Exactly. And we can promise you that we're giving you the best performance out of the chips that
John Furrier
>> we—who's the competition? Is Amazon a frenemy or are
Emilio Andere
>> they—Amazon is so big. We don't think about them today.
John Furrier
>> Are you targeting startups and what's your— well, obviously you guys are growing fairly fast. What's the landing zone? Where's your
Emilio Andere
>> beachhead?We are targeting startups. AI-native startups that are growing super fast. One of our largest customers, for example, is Vercel.
Emilio Andere
>> Yeah.
Emilio Andere
>> And they use us a lot for very fast inference and offering it to their customers. We also deal with a lot of voice agents because they care a lot about low latency. If you're over the phone with a voice agent and it takes like 3 seconds responding, you're kind of suspicious about the whole experience. So it's really important to have a good low latency experience for every user.
John Furrier
>> it's interesting, in 2015, if you asked me if there would ever be another hyperscaler, I would have said never, because at that time the barrier to entry was so high. And then in comes the GPU changeover. Now you have the neoclouds. I think there's a huge opportunity for the same reason why AWS was successful, for why I think you're going to be successful and Fireworks and other ones and everyone else, because this generation has the same dynamics. The reason why we were born in the cloud is the alternative was to buy a machine, put it in a shared cage, rack and stack it and then boot it up. before we even know what we're going to build? So now you start to see this AI native culture They don't yet truly know, but when they— once they get product market fit, they have to go fast. And so they don't have time to hire an engineer and do all the things.
Emilio Andere
>> Yeah.
John Furrier
>> Are you tapping into that?
Emilio Andere
>> Yeah. Yeah. So the whole— I think AWS is like one argument of why AWS gets so big is EC2, right? Yeah. You were finally able to pay for machines by the minute without having to worry about the machine and all this configuration. I think this—
John Furrier
>> by the way, no one ever heard of Airbnb or Twitter or Box or Dropbox.
Emilio Andere
>> Exactly.
John Furrier
>> They were hanging out. They were coming out of it.
Emilio Andere
>> Yeah, exactly.
John Furrier
>> It's our dorm rooms.
Emilio Andere
>> Yeah, so it enabled a huge generation of new companies. So I think this is especially true. I think a lot of our business today is just also just selling you by the token. Like, hey, you don't have to worry about any of the sort of GPU per hour deployments. Just come to us and we'll serve you the best performance by the token. Similar to that EC2 model of just like pay us by the token instead of having to reserve this capacity.
John Furrier
>> So you can tune So my— if my parameters were, hey, I don't really need full frontier general intelligence. Yes, I've got specialized domain expertise. I want a relatively small energy budget.
Emilio Andere
>> Yeah.
John Furrier
>> Budget.
Emilio Andere
>> Yeah.
John Furrier
>> And I would feed that into your agents.
Emilio Andere
>> Exactly. Exactly. And then so we focus on the frontier sort of open source models of today, the ones you've probably heard of from other folks, the DeepSeek, Kimi of the world that are amazing. they're really getting close to the performance of the Opuses and Sonnets of the world from both
John Furrier
>> open—What's the best practice? What's a good thing that's happening that you could point to? Because a lot of people get caught in the noise, but there's a lot of activity where people are taking models, open source and training their existing stuff, refining it. What are some of the best practices of people who want to— who know their data? Yeah, but might not be up to speed on what leaderboard has what, this, where's the open source vulnerabilities. So there's a lot in there you got to wade through.
John Furrier
>> There is a lot in there.
John Furrier
>> What's the best practice for people to get up and running as fast as possible?
Emilio Andere
>> I think maybe something that's counter to what you usually hear, maybe not, but I usually tell people just try open source models, just try them out of the box before deciding that you want to fine tune it. I think fine tuning is a big project. I think getting your data clean, it's a big effort and there's these companies and—
John Furrier
>> Define big.
Emilio Andere
>> It's, yeah, scope-wise, just scope out. Yeah, call it like, I would say, If you're doing it yourself, call it 5 to 10 engineers depending on the company over a period of a couple months, maybe 1 or 2 quarters. That's pretty long. It takes a long time. It's generally because you won't get it right the first time. So companies that are great like Fireworks or Baseten will help you do this. We will also help you do this. But what I tell customers is, just try the open source Kimi K3, cuz it might just be that with prompts you're able to avoid all that fine tuning. You don't have to pay the compute cost for fine-tuning and you can just use a cheaper open-source model. I think we're getting to a world where people are not really fine-tuning OpenAI and Anthropicanymore.
John Furrier
>> They've already nailed the internet. They crawled everything.
Emilio Andere
>> Exactly.
John Furrier
>> But they don't crawl my data if I'm an enterprise.
John Furrier
>> Exactly.
John Furrier
>> And the enterprise also has IT people that aren't necessarily skilled. So I think you're on this interesting path of this, I call it forward-deployed performance engineer, but it's an agent. And that's what a forward deploy engineer does.
Emilio Andere
>> Yeah, but it's an agent.
John Furrier
>> The enterprises need help because they're used to the old school load Linux, connect to the network, deploy it. But the demand is so high for the enterprise.
John Furrier
>> You can't move
John Furrier
>> fast.They can't move fast. Well, they don't have the skills. They need the agents. Okay, what are you working on now that's cool that you can talk about?
Emilio Andere
>> Yeah, yeah, yeah. So we're, I think the coolest thing about having agents that can do the work of forward deployed performance engineers is that we're doing, I think for the first time, what we're calling continual inference optimization. Which is we will learn from your traffic patterns over time to make sure that you are always running the most optimized version of the LLM that you want to run. So if you get a traffic spike every 9:00 AM, because that's when people start working on your product, you're a company like Vercel. Yeah. So a lot of people start coding at 9:00 AM. Then we will figure out those spikes and we'll figure out where exactly we need to do what things in the stack. So that you're running at the best performance per dollar at every point in time. So we're seeing these things where people get deployed on us and we give them the best performance in the market on these open source models. But we tell them we can get you 30 to 50% better performance about 2 months in once we've learned your stack and really mapped out where you need the spikes, where we can change your—
John Furrier
>> that literally is having an engineer on staff.
John Furrier
>> Exactly. But it's agents.
John Furrier
>> All right. Let me ask you this question, because what I love about NVIDIA, they're so good on these Pareto curves. Last GTC. And then Jensen put out the Vera Rubin curve and he said now you have four Pareto curves basically, and they have different price performance levels. I mean, this is pretty obvious. People talk about this all the time, but you don't want basic prompts going into the tier one, most expensive tokens.
John Furrier
>> Yeah.
John Furrier
>> So we're getting into a kind of a policy game here. What are your thoughts on this? You guys currently doing this? Because if I want to have the best tokens per watt value price, yeah, because token prices are dropping, which is great, but I don't I might not need the heavy-duty reasoning and processing.
Emilio Andere
>> Yeah.
John Furrier
>> I might want to say, here's my budget. Budget management meets tuning.
Emilio Andere
>> Totally. That's a big part of how we work with customers where they're like, hey, I just want to spend this amount and you need to get your agents to find me the absolute best performance that is possible with that amount of money. That amount of money might mean that they can only run on H100s. And maybe they won't run on B200s or
John Furrier
>> B300s.It could be
Emilio Andere
>> they're—or it could be AMD, which is
John Furrier
>> cheaper.I mean, we, as I always say, beauty is in the eye of the beholder. Whatever your environment is, you have the constraints of your own environment.
John Furrier
>> Exactly.All right. Why the name Wafer? Give us the story behind the Wafer, because I was thinking, okay, we're going to talk semis, wafers. It's great to get to know your company. So I love what you're doing, but why
Emilio Andere
>> Wafer?Was that— yeah. So there's like two answers. I think the one that I like to say is that because we will eventually build a chip, we really believe in full verticalization, right? Like at some point we really believe in this idea of hyper-optimization. Every customer should be extremely tuned to what they need in the AI stack. And I think at some point that will mean, at some point, many, many years from now, that will mean, can you actually build a custom chip for every important workload that a customer wants to run? So that's kind of an internal joke inside the company, not really within any quarterly timeline.
John Furrier
>> It's a north star.
Emilio Andere
>> It's a north star.
John Furrier
>> All right, so I gotta ask you, since you're here, we can put you in the Mixture of Experts series. It's good. So good. All right, so where— when does vertical integration not work? Because one of the benefits of horizontal scalability is data access, but also there's really great advantages of vertically integrating. Yeah, on performance. How do you guys think about that? How should people think architecturally around having kind of a horizontal layer but also really vertically integrating?
Emilio Andere
>> I think the way we think about it is the main sort of moat of cloud companies is scale economies and purchasing power. And the more vertically integrated you are in those two dimensions, the better, economics and sort of product you can give people because the cheaper you can produce your goods for and the faster and better you can make your processes to produce those goods. So in an industry like cloud, it makes a lot of sense to verticalize because just like That is the way to create
John Furrier
>> a—The domain expertise are there too.
Emilio Andere
>> Exactly. And the domain expertise, which is part of that process power. When you have that verticalization, you can have those agents at some point. Those agents right now are in the software stack.
Emilio Andere
>> Right.
Emilio Andere
>> And they're helping the GPUs run faster. But at some point we get to deploy agents to figure out the right temperature in data centers. Right.
John Furrier
>> That's a great point, Emilio. In fact, I was commenting on Nscale when they were in the building for their investor meeting. They bought my friend's company, Anyscale.
Emilio Andere
>> Yeah.
John Furrier
>> So I've been following those guys coming out of Cal, and they were originally in the Kubernetes space playing around with all the kind of microservices. But what they're actually doing is basically managing all the resources in real time for any workload, if there's any disaggregated serving or any kind of resource management. So that speaks to the fact that these workloads are tapping into multiple subsystems, not just the GPU.
Emilio Andere
>> Totally. Totally. If you want to extract maximum performance, also a benefit of verticalization, you have to look at every single component now, all the way down to at some point the data center, all the way down to hopefully a chip
John Furrier
>> soon.And you see Kubernetes as a standard layer for
Emilio Andere
>> that?We do. We do. Yeah. I think everybody kind of understands that interface and it's been around for so long.
John Furrier
>> All What's next? You got the funding, only $40 million. So. By the way, I'm not one of those people who look at the funding as a validation. I think lean and mean and then scale up. The funding. Sometimes overfunding can be bad, but I don't mean to bring that as a negative, it's a positive. Yeah. But you guys are growing, you're gonna build out, what's your plans?
John Furrier
>> Yeah.
John Furrier
>> What are you optimizing for now?
Emilio Andere
>> Yeah, just scale at this point. The demand is there. It's just clearly people are seeing the benefits.
Emilio Andere
>> Yeah.
Emilio Andere
>> Of having the equivalent of like dozens of performance engineers that are just agents doing the work that they don't want to do and that traditionally larger inference companies haven't been able to give them. So we just scale that technology. Like the demand is there and just grow the team, get more GPU use.
John Furrier
>> What are you looking for for the team? Put a plug in for openings, areas.
Emilio Andere
>> I'll plug in for openings. Yeah, totally.
John Furrier
>> What are you looking to hire?
Emilio Andere
>> We're looking to hire members of technical staff. So generally exceptional people that have programmed. generally we just look for people that have done really interesting work with computers, people that are very driven, the usual. Yeah. And obviously they'll want to work very hard.
John Furrier
>> Yeah.And align with the tribe, the vibe of the tribe, which is AI native.
Emilio Andere
>> Exactly.
John Furrier
>> yeah.Yeah.
Emilio Andere
>> Totally.100%. Okay, cool. So we end up hiring a lot of, I think the people coming out of college end up being really good, sort of really, really good archetypes for the people that are just using agents in ways that you're like, whoa, I didn't even know this was like a thing people were doing.
John Furrier
>> Exactly.Incredible young guns. Yeah. Emilio, thanks for coming in.
Emilio Andere
>> Appreciate
John Furrier
>> you.Congratulations again. AI Factory. It's very complicated under the hood, but it's only getting better as the infrastructure learns how to run at large-scale, hyperscale performance, but also the edge. You have different form factors, GPUs, CPUs, XPUs, all memory systems all tied together. You got to kind of know how it works, and hiring a performance engineer agent seems to be a great path. That's theCUBE here in New York City. Thanks for watching.
>> Palo Alto studio connecting Silicon Valley and Wall Street. I'm John Furrier, host of theCUBE, here with Dave Vellante, my co-host. We are here in theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is the NYSE Wired program and open community, and this is our AI Factory series. where we talk to the leaders who are making it happen. Obviously, AI factories continuing to power the massive buildout of the AI infrastructure, which is accelerating new architectures, new systems designs, accelerating the software paradigm, bringing in all kinds of new agentic and/or, AI-related applications. Our next guest is Emilio Andere, co-founder and CEO of Wafer, Bay Area-based. Emilio, great to have you on.
Emilio Andere
>> Thank you for having me.
John Furrier
>> Cool. First of all, give us some of the stats. When you guys were founded.
John Furrier
>> Yeah.
John Furrier
>> What's the status of the funding? Did you do a pre-seed at $200 million? Give us that.
Emilio Andere
>> We were actually a pretty simple pre-seed. It was Y Combinator. Okay, all right. So, me and my co-founder, we were roommates in college at UChicago, and we decided to start Wafer about a year and a half ago, and we did YC a couple of weeks after, so that was our sort of pre-seed. A pre-seed. Then we raised a seed, we raised about $4 million, about 3 months after that. And then most recently, about a month ago, actually a couple weeks ago, we announced our $40 million Series A, so just to grow the team and keep expanding the
John Furrier
>> technology.So pretty fast
Emilio Andere
>> progression.Yeah, it's been pretty
John Furrier
>> fast.All right, so what's the status now? Take what you guys do, explain the value proposition. What's your main problem that you're going
Emilio Andere
>> after?Yeah, totally. So when companies want to use LLMs and want to run on GPUs, it's actually quite hard to get those GPUs to run very fast for any LLM. Let's say you want to use DeepSpeed, Qwen, whatever you would want to use. It's really hard to actually get those GPUs and extract their maximum performance. So the compute, the sort of hardware is there, but in terms of software, it's really hard to write those kernels. It's hard to do the end-to-end optimization so that your LLM actually is extracting the best performance per dollar invested into that GPU.
Emilio Andere
>> Yeah.
Emilio Andere
>> So what we do is we sort of found this scarcity in the number of GPU engineers that could do this work. So we decided that it was very obvious to us that the, yeah, what you should be doing to optimize your GPU should be with AI itself. We call this kind of AI to optimize AI. It's like, why would you use scarce humans that know how to actually do this low-level work? Yeah. When agents are getting so good at coding and you can finally use these agents as performance engineers, GPU engineers.
John Furrier
>> You're right, for deployed AI.
Emilio Andere
>> Exactly.
John Furrier
>> Essentially.
Emilio Andere
>> Exactly.
John Furrier
>> I mean, back in the old days, not to date myself because you're the young gun kicking ass, taking names in the AI world, we would buy a server and load Linux on it. There was a software stack well understood that gets you up and running. You got an OS and then you put your stuff on there and then you're off to the races, connect to the network. Here, AI, it's hard to really identify what the stack is for a GPU.
John Furrier
>> Yeah.
John Furrier
>> And it's not just the GPU either. There's other stuff going around it. So one of the things we observed when we started the series 2 years ago was, I love this AI factories concept, ship a rack to someone. What the hell do you run on it? And then now we're seeing all kinds of disaggregated serving. The caches are getting bloated. So the entire performance engine is getting clogged up. Is that kind of the area you guys are going in and trying to nail that piece?
Emilio Andere
>> Yeah, yeah, exactly. So it's getting super complicated. And as you said, one of the other reasons it's getting super complicated is because there's other chips coming into play. So now you not only have to figure out the stack for NVIDIA, you also have to figure it out for AMD and you have to figure it out for Google TPUs and for Trainium. And most of the traditional inference companies that came out in 2021, that are now giants, they are basically just fully on NVIDIA and that's just because the industry was there at that time. But today we're in a much more heterogeneous world. As you said, people are combining Prefill and Decode chips in very interesting ways. You saw AMD do that with Cerebras. You saw NVIDIA do that with Groq. There's clearly a lot of stuff going
John Furrier
>> on.the Groq acquisition or the AccuHire, however you want to frame it, yeah, really kind of changed the direction of the industry. I'd love to get your comments on that because at that time Cerebras was not viewed as a serious player in inference, although they had the big chip. We covered them. We actually liked what they were doing. Half the industry thought it wouldn't work. You had other people trying to go after inference. But NVIDIA set the agenda just by buying Groq, a separate thing. Yeah, kind of changed the inference game, kind of almost segments training out. What's your opinion? Did I get that right? How would you describe that moment? Because it really was not well reported that that actually changed the trajectory of the innovation.
Emilio Andere
>> It did. Yeah. And you did get that right. I think it was similar to NVIDIA's Mellanox acquisition many years before. I think it was honestly just Jensen and the team being just visionaries. Of just like, hey, clearly there's a gap in how well GPUs can serve certain workloads. And clearly there's a market for people who are willing to pay more money for faster tokens. Just give me better latency for every one of my applications and I will pay you more per token that you generate. And he was able to identify that market and just obviously with the acquisition, make it even more serious.
John Furrier
>> Yeah, because on the last earnings call, they actually are now quantifying dollar value per gigawatt metrics. So he's already moving the metrics. You got to give NVIDIA props. They have actually accelerated not just on the financing side, but on the business side. They're pushing the envelope. Oh yeah, they're even in politics. So they're shaping the agenda. You got to give them props for that. All right. So the question really now is, okay, AI factories, we all see the big gigawatt factories, Crusoe Nscale, monster campuses. But that's going to be the big token factories, almost like oil refineries with tokens, right? Think about it that way. Now you have people at the enterprise going, hey, I can't afford $10 billion of CapEx. I just have some infrastructure. Maybe I'll do a deal with Nscale or CoreWeave, but I want to get an on-premise capacity, but I'm not going to have the footprint. So that's going to change the scope of my software stack.
Emilio Andere
>> Yeah, yeah.
John Furrier
>> Do you solve that problem? Could you be a solution for them?
Emilio Andere
>> Yeah. So our current solution and the thing that's been having the most growth is just going to customers and saying, we can give you the best performance in the market because we have these agents that we can deploy that will tune every part of the stack and then serving them the best performance per dollar. But we do it on our own GPUs. So we would go to companies like CoreWeave, Nebius and get compute from them and then just serve them as the wafer cloud. Right. I think there is a very interesting future where we do on-prem deployments because the demand is there.
John Furrier
>> So you're optimizing the raw power from the people who actually have gear.
John Furrier
>> Exactly.
John Furrier
>> Seems to work.
Emilio Andere
>> Yeah. So we're above the sort of— we rent out these long-term commitments from the neo clouds like the ones you mentioned and then serve AI-native startups, very fast inference.
John Furrier
>> Do you see a bifurcation between, say, I want to be vertically integrated. I'm going to just get my GPUs. I'm just going to serve, be a service provider and also offer raw takeout. Takeout.
Emilio Andere
>> Well, yeah. Yeah. So it's a good question. I think I do, but I do think the people that verticalize will win. I think it's a very important strategy. You basically, as you know, probably better than most, there's this— you have the inference providers there at the layer with the demand. You have the Fireworks, Baseten of the world, and they traditionally don't own GPUs themselves. They just rent it out from these data center companies, but they're trying to go to that level because they want to increasingly own their own compute. So they're kind of going down. And then you have the traditional data center companies like CoreWeave and Nebius, and they're trying to go up because they realize that there's better margins on tokens than just reselling compute. So they're like, hey, now we're a token factory. So you have this middle layer that in my opinion will get squeezed. And we used to do this. We used to sell our engine to other inference companies. But then we realized that we were going to get squeezed out. It was time to just go serve tokens directly and then verticalize at some point or we would get squeezed. So we do want to end up doing—
John Furrier
>> and there's also a lot of stuff going on in that serving. that's why I brought up disaggregated serving. I think that's also a tell to disaggregate infrastructure at the edge. So you say, okay, that's a symptom of the growth, but that also speaks to how hard it is to manage those resources. So if you need to do resource management observability, you can't just outsource that. Yeah, you got to have it in the stack. You agree?
Emilio Andere
>> Yeah, I do agree. It's getting much harder. Like there's so much increasingly complex stuff that's going on in these AI deployments. If we talked about disaggregating Prefill and Decode, different chips coming along, different serving engines, it's just getting so complicated that it just stops making sense for companies to, at least at the smaller scale, do it themselves. And instead it's— all right, so let's just say I'm a customer.
John Furrier
>> Yeah, Cube Cloud. We have a Cube Cloud on AWS with a lot of data. I want to stand up, turn on an AI factory. Yeah, I don't have any GPU engineers. Maybe I can't even get a hold of GPUs. I want to deploy a software stack that ties all my existing stuff together. Would I call you guys for that? What would it be?
Emilio Andere
>> Yeah, you would basically call us and you would say, hey, here's the workload I want to run. Here's the model I want to run for this workload. And here's the chip I want to run for that workload. And that's a huge permutation space. that's a huge space. And we deploy our agents to find the best kernels, the best serving engine. It basically— you can think of it as generating the entire stack just for your use case. You can finally do this because we're not bound by human ability.
John Furrier
>> You take the template inbound logistics of what I want.
John Furrier
>> Exactly.
John Furrier
>> And it's almost like cloud configuring EC2 or whatever.
Emilio Andere
>> Instances. Exactly. And we can promise you that we're giving you the best performance out of the chips that
John Furrier
>> we—who's the competition? Is Amazon a frenemy or are
Emilio Andere
>> they—Amazon is so big. We don't think about them today.
John Furrier
>> Are you targeting startups and what's your— well, obviously you guys are growing fairly fast. What's the landing zone? Where's your
Emilio Andere
>> beachhead?We are targeting startups. AI-native startups that are growing super fast. One of our largest customers, for example, is Vercel.
Emilio Andere
>> Yeah.
Emilio Andere
>> And they use us a lot for very fast inference and offering it to their customers. We also deal with a lot of voice agents because they care a lot about low latency. If you're over the phone with a voice agent and it takes like 3 seconds responding, you're kind of suspicious about the whole experience. So it's really important to have a good low latency experience for every user.
John Furrier
>> it's interesting, in 2015, if you asked me if there would ever be another hyperscaler, I would have said never, because at that time the barrier to entry was so high. And then in comes the GPU changeover. Now you have the neoclouds. I think there's a huge opportunity for the same reason why AWS was successful, for why I think you're going to be successful and Fireworks and other ones and everyone else, because this generation has the same dynamics. The reason why we were born in the cloud is the alternative was to buy a machine, put it in a shared cage, rack and stack it and then boot it up. before we even know what we're going to build? So now you start to see this AI native culture They don't yet truly know, but when they— once they get product market fit, they have to go fast. And so they don't have time to hire an engineer and do all the things.
Emilio Andere
>> Yeah.
John Furrier
>> Are you tapping into that?
Emilio Andere
>> Yeah. Yeah. So the whole— I think AWS is like one argument of why AWS gets so big is EC2, right? Yeah. You were finally able to pay for machines by the minute without having to worry about the machine and all this configuration. I think this—
John Furrier
>> by the way, no one ever heard of Airbnb or Twitter or Box or Dropbox.
Emilio Andere
>> Exactly.
John Furrier
>> They were hanging out. They were coming out of it.
Emilio Andere
>> Yeah, exactly.
John Furrier
>> It's our dorm rooms.
Emilio Andere
>> Yeah, so it enabled a huge generation of new companies. So I think this is especially true. I think a lot of our business today is just also just selling you by the token. Like, hey, you don't have to worry about any of the sort of GPU per hour deployments. Just come to us and we'll serve you the best performance by the token. Similar to that EC2 model of just like pay us by the token instead of having to reserve this capacity.
John Furrier
>> So you can tune So my— if my parameters were, hey, I don't really need full frontier general intelligence. Yes, I've got specialized domain expertise. I want a relatively small energy budget.
Emilio Andere
>> Yeah.
John Furrier
>> Budget.
Emilio Andere
>> Yeah.
John Furrier
>> And I would feed that into your agents.
Emilio Andere
>> Exactly. Exactly. And then so we focus on the frontier sort of open source models of today, the ones you've probably heard of from other folks, the DeepSeek, Kimi of the world that are amazing. they're really getting close to the performance of the Opuses and Sonnets of the world from both
John Furrier
>> open—What's the best practice? What's a good thing that's happening that you could point to? Because a lot of people get caught in the noise, but there's a lot of activity where people are taking models, open source and training their existing stuff, refining it. What are some of the best practices of people who want to— who know their data? Yeah, but might not be up to speed on what leaderboard has what, this, where's the open source vulnerabilities. So there's a lot in there you got to wade through.
John Furrier
>> There is a lot in there.
John Furrier
>> What's the best practice for people to get up and running as fast as possible?
Emilio Andere
>> I think maybe something that's counter to what you usually hear, maybe not, but I usually tell people just try open source models, just try them out of the box before deciding that you want to fine tune it. I think fine tuning is a big project. I think getting your data clean, it's a big effort and there's these companies and—
John Furrier
>> Define big.
Emilio Andere
>> It's, yeah, scope-wise, just scope out. Yeah, call it like, I would say, If you're doing it yourself, call it 5 to 10 engineers depending on the company over a period of a couple months, maybe 1 or 2 quarters. That's pretty long. It takes a long time. It's generally because you won't get it right the first time. So companies that are great like Fireworks or Baseten will help you do this. We will also help you do this. But what I tell customers is, just try the open source Kimi K3, cuz it might just be that with prompts you're able to avoid all that fine tuning. You don't have to pay the compute cost for fine-tuning and you can just use a cheaper open-source model. I think we're getting to a world where people are not really fine-tuning OpenAI and Anthropicanymore.
John Furrier
>> They've already nailed the internet. They crawled everything.
Emilio Andere
>> Exactly.
John Furrier
>> But they don't crawl my data if I'm an enterprise.
John Furrier
>> Exactly.
John Furrier
>> And the enterprise also has IT people that aren't necessarily skilled. So I think you're on this interesting path of this, I call it forward-deployed performance engineer, but it's an agent. And that's what a forward deploy engineer does.
Emilio Andere
>> Yeah, but it's an agent.
John Furrier
>> The enterprises need help because they're used to the old school load Linux, connect to the network, deploy it. But the demand is so high for the enterprise.
John Furrier
>> You can't move
John Furrier
>> fast.They can't move fast. Well, they don't have the skills. They need the agents. Okay, what are you working on now that's cool that you can talk about?
Emilio Andere
>> Yeah, yeah, yeah. So we're, I think the coolest thing about having agents that can do the work of forward deployed performance engineers is that we're doing, I think for the first time, what we're calling continual inference optimization. Which is we will learn from your traffic patterns over time to make sure that you are always running the most optimized version of the LLM that you want to run. So if you get a traffic spike every 9:00 AM, because that's when people start working on your product, you're a company like Vercel. Yeah. So a lot of people start coding at 9:00 AM. Then we will figure out those spikes and we'll figure out where exactly we need to do what things in the stack. So that you're running at the best performance per dollar at every point in time. So we're seeing these things where people get deployed on us and we give them the best performance in the market on these open source models. But we tell them we can get you 30 to 50% better performance about 2 months in once we've learned your stack and really mapped out where you need the spikes, where we can change your—
John Furrier
>> that literally is having an engineer on staff.
John Furrier
>> Exactly. But it's agents.
John Furrier
>> All right. Let me ask you this question, because what I love about NVIDIA, they're so good on these Pareto curves. Last GTC. And then Jensen put out the Vera Rubin curve and he said now you have four Pareto curves basically, and they have different price performance levels. I mean, this is pretty obvious. People talk about this all the time, but you don't want basic prompts going into the tier one, most expensive tokens.
John Furrier
>> Yeah.
John Furrier
>> So we're getting into a kind of a policy game here. What are your thoughts on this? You guys currently doing this? Because if I want to have the best tokens per watt value price, yeah, because token prices are dropping, which is great, but I don't I might not need the heavy-duty reasoning and processing.
Emilio Andere
>> Yeah.
John Furrier
>> I might want to say, here's my budget. Budget management meets tuning.
Emilio Andere
>> Totally. That's a big part of how we work with customers where they're like, hey, I just want to spend this amount and you need to get your agents to find me the absolute best performance that is possible with that amount of money. That amount of money might mean that they can only run on H100s. And maybe they won't run on B200s or
John Furrier
>> B300s.It could be
Emilio Andere
>> they're—or it could be AMD, which is
John Furrier
>> cheaper.I mean, we, as I always say, beauty is in the eye of the beholder. Whatever your environment is, you have the constraints of your own environment.
John Furrier
>> Exactly.All right. Why the name Wafer? Give us the story behind the Wafer, because I was thinking, okay, we're going to talk semis, wafers. It's great to get to know your company. So I love what you're doing, but why
Emilio Andere
>> Wafer?Was that— yeah. So there's like two answers. I think the one that I like to say is that because we will eventually build a chip, we really believe in full verticalization, right? Like at some point we really believe in this idea of hyper-optimization. Every customer should be extremely tuned to what they need in the AI stack. And I think at some point that will mean, at some point, many, many years from now, that will mean, can you actually build a custom chip for every important workload that a customer wants to run? So that's kind of an internal joke inside the company, not really within any quarterly timeline.
John Furrier
>> It's a north star.
Emilio Andere
>> It's a north star.
John Furrier
>> All right, so I gotta ask you, since you're here, we can put you in the Mixture of Experts series. It's good. So good. All right, so where— when does vertical integration not work? Because one of the benefits of horizontal scalability is data access, but also there's really great advantages of vertically integrating. Yeah, on performance. How do you guys think about that? How should people think architecturally around having kind of a horizontal layer but also really vertically integrating?
Emilio Andere
>> I think the way we think about it is the main sort of moat of cloud companies is scale economies and purchasing power. And the more vertically integrated you are in those two dimensions, the better, economics and sort of product you can give people because the cheaper you can produce your goods for and the faster and better you can make your processes to produce those goods. So in an industry like cloud, it makes a lot of sense to verticalize because just like That is the way to create
John Furrier
>> a—The domain expertise are there too.
Emilio Andere
>> Exactly. And the domain expertise, which is part of that process power. When you have that verticalization, you can have those agents at some point. Those agents right now are in the software stack.
Emilio Andere
>> Right.
Emilio Andere
>> And they're helping the GPUs run faster. But at some point we get to deploy agents to figure out the right temperature in data centers. Right.
John Furrier
>> That's a great point, Emilio. In fact, I was commenting on Nscale when they were in the building for their investor meeting. They bought my friend's company, Anyscale.
Emilio Andere
>> Yeah.
John Furrier
>> So I've been following those guys coming out of Cal, and they were originally in the Kubernetes space playing around with all the kind of microservices. But what they're actually doing is basically managing all the resources in real time for any workload, if there's any disaggregated serving or any kind of resource management. So that speaks to the fact that these workloads are tapping into multiple subsystems, not just the GPU.
Emilio Andere
>> Totally. Totally. If you want to extract maximum performance, also a benefit of verticalization, you have to look at every single component now, all the way down to at some point the data center, all the way down to hopefully a chip
John Furrier
>> soon.And you see Kubernetes as a standard layer for
Emilio Andere
>> that?We do. We do. Yeah. I think everybody kind of understands that interface and it's been around for so long.
John Furrier
>> All What's next? You got the funding, only $40 million. So. By the way, I'm not one of those people who look at the funding as a validation. I think lean and mean and then scale up. The funding. Sometimes overfunding can be bad, but I don't mean to bring that as a negative, it's a positive. Yeah. But you guys are growing, you're gonna build out, what's your plans?
John Furrier
>> Yeah.
John Furrier
>> What are you optimizing for now?
Emilio Andere
>> Yeah, just scale at this point. The demand is there. It's just clearly people are seeing the benefits.
Emilio Andere
>> Yeah.
Emilio Andere
>> Of having the equivalent of like dozens of performance engineers that are just agents doing the work that they don't want to do and that traditionally larger inference companies haven't been able to give them. So we just scale that technology. Like the demand is there and just grow the team, get more GPU use.
John Furrier
>> What are you looking for for the team? Put a plug in for openings, areas.
Emilio Andere
>> I'll plug in for openings. Yeah, totally.
John Furrier
>> What are you looking to hire?
Emilio Andere
>> We're looking to hire members of technical staff. So generally exceptional people that have programmed. generally we just look for people that have done really interesting work with computers, people that are very driven, the usual. Yeah. And obviously they'll want to work very hard.
John Furrier
>> Yeah.And align with the tribe, the vibe of the tribe, which is AI native.
Emilio Andere
>> Exactly.
John Furrier
>> yeah.Yeah.
Emilio Andere
>> Totally.100%. Okay, cool. So we end up hiring a lot of, I think the people coming out of college end up being really good, sort of really, really good archetypes for the people that are just using agents in ways that you're like, whoa, I didn't even know this was like a thing people were doing.
John Furrier
>> Exactly.Incredible young guns. Yeah. Emilio, thanks for coming in.
Emilio Andere
>> Appreciate
John Furrier
>> you.Congratulations again. AI Factory. It's very complicated under the hood, but it's only getting better as the infrastructure learns how to run at large-scale, hyperscale performance, but also the edge. You have different form factors, GPUs, CPUs, XPUs, all memory systems all tied together. You got to kind of know how it works, and hiring a performance engineer agent seems to be a great path. That's theCUBE here in New York City. Thanks for watching.