This interview examines artificial intelligence, abbreviated AI, infrastructure
and a robotics partnership that expands secure serverless token factories for
enterprise adoption. Recorded at theCUBE and NYSE Wired Robotics and AI
Infrastructure Leaders Series during the third annual AI Leaders Summit, the
session addresses GPU orchestration, multi-tenancy, privacy and developer
experience for neo-cloud operators and chief information security officers,
abbreviated CISOs. Eiman Ebrahimi of Protopia and Haseeb Budhani of Rafay
Systems join hosts John Furrier and Dave Vellante for an interview produced by
theCUBE Research. Ebrahimi outlines how Protopia and Rafay Systems collaborate
to enable virtual private token factories improve multi-tenant GPU orchestration
and reconcile privacy and compute constraints; they describe technical
approaches to model compatibility and agent-driven token demand. Budhani details
how eliminating the isolation tax enables operators to monetize idle capacity;
they emphasize the business model and economic impacts while preserving
developer experience. Key takeaways include the business and technical impacts
of true serverless multi-tenancy and an upstream privacy layer. Budhani asserts
that removing the isolation tax enables capacity monetization. Ebrahimi
demonstrates that Stained Glass transforms inputs to reduce plaintext exposure.
Analysts John Furrier and Dave Vellante emphasize that security, observability
and economics are essential to scale AI infrastructure adoption.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Eiman Ebrahimi, Protopia & Haseeb Budhani, Rafay
This interview examines artificial intelligence, abbreviated AI, infrastructure
and a robotics partnership that expands secure serverless token factories for
enterprise adoption. Recorded at theCUBE and NYSE Wired Robotics and AI
Infrastructure Leaders Series during the third annual AI Leaders Summit, the
session addresses GPU orchestration, multi-tenancy, privacy and developer
experience for neo-cloud operators and chief information security officers,
abbreviated CISOs. Eiman Ebrahimi of Protopia and Haseeb Budhani of Rafay
Systems join hosts John Furrier and Dave Vellante for an interview produced by
theCUBE Research. Ebrahimi outlines how Protopia and Rafay Systems collaborate
to enable virtual private token factories improve multi-tenant GPU orchestration
and reconcile privacy and compute constraints; they describe technical
approaches to model compatibility and agent-driven token demand. Budhani details
how eliminating the isolation tax enables operators to monetize idle capacity;
they emphasize the business model and economic impacts while preserving
developer experience. Key takeaways include the business and technical impacts
of true serverless multi-tenancy and an upstream privacy layer. Budhani asserts
that removing the isolation tax enables capacity monetization. Ebrahimi
demonstrates that Stained Glass transforms inputs to reduce plaintext exposure.
Analysts John Furrier and Dave Vellante emphasize that security, observability
and economics are essential to scale AI infrastructure adoption.
>> Welcome back to theCUBE here in Palo Alto, California. I'm John Furrier, host of theCUBE. This is our third annual AI Leaders Summit, all day interviews here in our theCUBE NYSE Wired studios, and then the party after 180 leaders coming together to meet, talk about all the latest trends. This next segment focuses on the partnership between Protopia AI and Rafay Systems that really expands the addressable market for the AI infrastructure. We have both CEOs here. Eiman from Protopia, good to see you.
Eiman Ebrahimi
>> Great to see you as well.
John Furrier
>> Haseeb, good to see you again.
Haseeb Budhani
>> Good to see you again.
John Furrier
>> CEO of Rafay Systems. Guys, the addressable market for enterprise AI is growing super fast. The adoption is now kicking in. This token factory partnership you guys have that really opens the door for more operations, more development, explain the relationship between the two companies and really what this does to increase the addressable market for the enterprise.
Haseeb Budhani
>> Yeah, happy to do that.So, as you know, at Rafay, we sell an orchestration stack that AI infrastructure providers use to deliver a cloud experience. As part of our stack, we provide a token factory where my customers can deliver open source models to their downstream customers, enterprises, for example. A key part of our business and my customers' business is how do we deliver a serverless experience where we don't need to allocate GPUs on a per enterprise basis. So you can run many, many enterprises on the same infrastructure in a secure fashion. The gap so far has been how do we make sure the contexts are essentially multi-tenant? How do you make sure those are secure? Protopia has built a technology set that allows us to provide essentially a virtual private token factory to each enterprise that is engaging with my Neocloud customers. So it's a great technology, sits on top of a platform and delivers a better experience to the enterprise and makes the enterprise CISO happy because they get the right controls and governance in place.
John Furrier
>> Even on your end, token factory, what does that mean? How does that connect?
Eiman Ebrahimi
>> Yeah, I think one of the biggest realizations for us just being in market talking to the same types of customers that run token factories is that often they're thinking about how do they increase the amount of supply that they have and so there's this constant consideration that they have a supply problem but when you dig into how are you operating the underlying factory and how are you actually serving the tokens to your customers you see that there's a lot of over allocated capacity passive meaning just because something is booked if it's not busy that's revenue that they're leaving on the table and so when we started working with Rafay the aspect that Haseeb was just describing of the end enterprises wanting to have privacy of their information comes together with the desire from the operator side to maximize revenueand the easiest supply they're ever going to get access to is the supply that they already have if they are able to run true serverless without having to over -allocate, say, a node instead of a slice, or having to separate things physically.
John Furrier
>> So, I can see all these optimizations coming in my head's kind of connecting the dots, but let's zoom out and look at the problem. Most enterprises have unique needs. We all have been covering the topic of the data's still locked away, and that unlock represents a huge opportunity. a lot of regulated industries in the enterprise, but a lot of them haven't taken advantage of the shared GPU infrastructure because of many reasons. Is that the core problem, is that one of the problems? Because they don't have the huge capex, they just need service.
Haseeb Budhani
>> So I think most enterprises we all know, they are signing up for contracts with Anthropic or OpenAI or something like that. I think a big trend in the market right now is, okay, I'm going to use my cloud, but when it's 2pm and I get the message that says no more tokens, come back at 6pm, that happens to all of us. Okay, I want to be able to use open source models. Right, so there's a slew of new clouds who are jumping to that opportunity, which is they're now using their infrastructure to deliver tokens as a service, using open source models that are available in the market that we all are aware of. The challenge now is, okay, I have the infrastructure, I have the Token Factory from Rafay, so I have two options. I can pre-allocate these GPUs to the enterprise, which means they may use it, they may not, which means I'm leaving money on the table. Or I can come up with a...
John Furrier
>> You're basically provisioning.
Haseeb Budhani
>> Exactly, I provision and I walk away, right? If they're not being used, it's a waste of money. But the user, no, then the enterprise is saying, why am I paying for this, I'm not using it presently. What if there's a way to do a true serverless solution? The serverless part at the infrastructure level, we've solved that, and that we've done for a while. What Protopia has done is taken it to the next level. Right, so now we have a slice of essentially a confidential GPU, right? So I can now have better security on top of the GPUs that I already have, which means as a provider, I get to monetize every second of my infrastructure, but for the enterprise, it's actually a better deal. They get a better price point, because the infrastructure is now shared holistically, versus on a time-sliced basis. And that's a value, right? They like that because they look at the total cost of ownership, they look at the capex, they don't have the big money. It's too expensive, otherwise, right? And this allows them to have a secure solution, which actually delivers a better price point. That's what everybody wants.
John Furrier
>> You mentioned privacy. He mentioned privacy at the beginning. Is privacy or compute the bottleneck for adoption in the enterprise?
Eiman Ebrahimi
>> why is it one or the other? The problem is that they're tied together. The inability to perform private or confidential inference, whichever, however you want to name it, the inability to do that will result in needing to physically isolate compute that becomes a supply issue because there's not enough to physically separate for every given, say, business unit inside of an enterprise or every data owner that has sensitive information. So these two topics are actually tied at the hip and being able to solve one then really enables technologies like solutions like the token factories that Rafay has to be fully utilized by the operator inside an AI factory. And another important thing here is that we've seen this topic not just in NeoCloud and sovereign AI deployments, but also when you have many different data owners inside the same organization and you've got the organization trying to operate what is essentially inside the organization, an AI factory for those different data owners. It's the same problem. If business unit A and business unit B can't expose data in a multi-tenant system because they already have data silos where it's stored and when it's going to compute it also needs to be separated, they're going to land in the same issue.
John Furrier
>> Yeah, in the discussion around your relationship, there was a discussion on isolation tax. Explain what that means because a lot of people don't want to rewrite their applications. Or they don't want to waste resources they paid for that aren't being utilized. What does isolation tax mean, and what does that mean to the enterprise from an application or workload standpoint?
Haseeb Budhani
>> Yeah, look, the best way to have security is you own everything, right? The next level is we provide some level of abstraction, so they have a VM or a Kubernetes cluster, et cetera. Eventually, what everybody wants is to be able to drive their infrastructure at every second of every day for some user, right? and being able to context switch that is a hard problem. So from an infrastructure level, we get it all the way up to, you can run serverless requests across multiple GPUs that are shared across multiple enterprises. The problem that Eiman is describing is a real one, which is within an enterprise or across enterprises, I want a separation between my data and that guy's data. And that is a gap right now, it is what it is. So the only solution for that is we pre -allocate GPUs for customers, and that's the tax. So if you have a way to consume that infrastructure every second of every day, no tax. But if you're going to be in a position where
John Furrier
>> Unused resources.
Haseeb Budhani
>> That's right. If you're going to have four, six hours of non -usage, you're just paying for infrastructure that you're not utilizing. In that window of time, somebody else could use it, and you should get a better deal.
John Furrier
>> Smart money's going to be on this, in the enterprise, because they're going to focus on, where's the value? And what's the cost? They're going to look at that, wait, it's not being used, squeeze more cost out of that. Let's talk model and compatibility. One of the things that everyone talks about is, hey, I still got to use general intelligence, that's OpenAI, they have APIs, we've been using them. You're going to have specialism, specialty models, specialized intelligence, obviously domain expertise in the enterprise is going to be highly domain specific.
Haseeb Budhani
>> Yes, sir.
John Furrier
>> Talk about the compatibility issue to the bigger models.
Haseeb Budhani
>> The good news is, most of these models now, at least the ones that we are providing as part of our catalog, they're OpenAI-compatible. So as a developer, you don't need to switch. You don't need to unlearn or relearn anything. It's basically the same experience. The requirement for a vendor like us is, be it the models that NVIDIA's publishing, the Nemotron models or the NIM models, or the open source models like Kimi, et cetera. So what we have done is we sort of brought them all together in one place where the enterprise or the NeoCloud can pick and choose, I want to use this one or this one or this one. But then on top of that, there has to be a layer which is sort of essentially independent of the model itself to be able to provide some sort of data isolation. I'm using that phrase very loosely, of course, Eiman. And that's where this comes together. Right, so we now independently deliver a marketplace of models that are open, if you will. Right, open source, open models. And then now there's a data isolation layer on top, which is independent of the model. So this is the right way to think about it, because you said something earlier which is most important, right? the security and the lower cost, both have to be true for enterprises to adopt this at scale. And that's what's possible here.
John Furrier
>> Can we talk about the Stained Glass technology, you guys call it, what does it do? What's the secret sauce there and how does that tie in? Because you got multi -tenancy, you got privacy and compute, efficiency on the cost side of the value. Where is the Stained Glass technology? How does that fit in?
Eiman Ebrahimi
>> Yeah, Stained Glass, when you think of it just as the technology part of it, is focused on enabling models, whatever model we're talking about, to be able to make inferences on non -raw representations of the inference data and the context that goes to them. So today, in almost every single system that's out there, whether it's the frontier models, whether it's the open models being deployed on-prem, being deployed by a vendor on inference platforms, when the request arrives at the input on Ingress, there is a point where the context and everything about the prompt will go to raw information, plain text, and it will live that way through the entire inference data path. and what you see just recently without necessarily naming names there's multiple occurrences of even the best of vendors having data leaks that occur where information ends up being copied from one place to another when it shouldn't have information showing up in plain text entire code bases on systems where they shouldn't have been this isn't necessarily because the vendor is doing something poorly these are some of the best vendors out there But the fact is that there's all sorts of hidden data leak vectors across these systems. So we call these various operational surfaces that exist. And what Stained Glass is meant to do is to create an inference privacy layer that's way upstream from that and transforms the data from that raw representation into a, we call it a stochastic re-representation of the data that can be consumed by the target model without ever needing to go back to the original form. And when you do that upstream, it allows for all of those operational surfaces to not have to expose plain raw information. So if a leak like what we just saw a few weeks ago does happen again, and it will, and it is happening right now somewhere in many of these systems, right? The data doesn't need to show up as a leak in plain text, which is what's happening today. And so what the users of Private Token Factories that Rafay is building with us are going to have the ability to show to their customers is that at that upstream point, your data is being transformed out of plain text. And so wherever it may land, even with the best of everyone's efforts, then it's not going to be in plain text.
John Furrier
>> So it's a protection mechanism on the data.
Eiman Ebrahimi
>> That's correct.
John Furrier
>> Also people can maybe relate to this when I say, don't put company information in OpenAI. That's another, maybe is that related or is that just more of the same kind of upstream behavior?
Eiman Ebrahimi
>> I think there's two components to that. The part of it that we're focused on is complementing all of the zero data retention policies that the vendors are trying to provide. And I use the word trying very deliberately in that it's very difficult to say that zero data retention was absolute in any given system. And as evidenced by the fact that these sorts of leaks happen very frequently. And so what we're doing is saying for any such leakage that may happen even when you have zero data retention, it's not going to be in plain text. Now there's another component to that where people say don't put information in something because they wonder about whether or not the vendor is training on their data, which is a separate story. What we're talking about is not necessarily saying your vendor is trying to do something that they said to you they weren't going to do. we're just pointing out the fact that leaks happen.
John Furrier
>> I love, we're in Palo Alto, but also theCUBE, NYSE Wired, and New York Stock Exchange. It's fun to speak Silicon Valley and Wall Street, and I want to ask you guys a question, because the business impact of what you're doing is interesting. When I walk the halls of the NYSE, I talk to the traders, I talk to the analysts, even talk to some of the CNBC folks all the time, they're all asked the same questions. Is there a bubble? We address that very easily. But then when they start getting into questions: who's going to make the money? Where's the value? Which stock is going to go public? Who's going to hit escape velocity first? We're in an era where whoever builds the best AI infrastructure that aligns with how people think and work wins. That's clear, that's documented. What's the business impact? Because what you're saying to me increases the cloud addressable, cloud AI addressable market. So you see the neo -clouds like CoreWeave, Nscale, they're building out super fast. Nscale, talking to those guys, they're crushing it, right? But they have to monetize, so they have to get business value. So where does that fit in? Because you got NVIDIA trying to put together a partnership network, OpenAI's got an ecosystem. How does this make the addressable market and the economics and the money-making side work?
Haseeb Budhani
>> I think the key thing for all of us in the AI infrastructure space is we need to be able to deliver infrastructure that enterprises can consume. The same way they were so comfortable consuming AWS or GCP or Azure, the same things apply here, right? So the same controls, the same guardrails, the same quota management, the same auditability, observability, all of that has to be true, right? At every layer of the stack. AI is a different use case, of course, than the general purpose compute, right? So this is what we're all going after. you and I talked about this before, in this AI infrastructure, is this a bubble? I don't see it because the demand is incredible.
John Furrier
>> There's no bubble discussion. You build and you go.
Haseeb Budhani
>> Bubble discussion is ridiculous. Yeah, there's so much demand. If there was just infrastructure sitting around, that's a different conversation, but you're deploying, people are ready. They'll buy everything that's being shipped. There you go, right? That's basically what we're seeing. So the question now is, how do you bring the enterprises who are presently using an AWS environment, because they feel comfortable with AWS, how do we give them the same comfort? These are the steps we take. So at our level, we invest heavily in multi -tenancy, and you are very familiar with our stack, but this relationship takes it to the next level, So these are my words, of course not Eiman's, but this data level sort of segregation that is now possible using the Stained Glass technology that these guys have built, makes it easier for my Neo Cloud customer or my enterprise customer to say, hey man, I can bet on this. So the more comfort we can provide with solid technology, that solves the security problem, without taking away from the experience for the developer.
John Furrier
>> This is what's going to expand the opportunity. It also is a highlight of the partnership model we're seeing in this new era because you mentioned upstream. We saw NVIDIA do that, they put KV cache, now you got the storage vendors coming in. They have density, they're at a low level, but when you go upstream, up the stack, it's like, is my data secure? Can the application work with confidence? Can I audit it? Is there traceability? these are words that we know, observability.
Haseeb Budhani
>> That's right.
John Furrier
>> That's cloud, right? So this cloud game kind of comes in. Eiman, I want to ask you on your end, because you deal with a lot of the data and AI at the platform level. Which verticals are kind of ripe right now? healthcare used to be viewed as, oh, they're slow, HIPAA's so old, you can't move, it's antiquated, slow. They got data, though. See, healthcare's popping. You got banking, government. government's the hottest area. Defense tech, public-private, private's now leading the way. Just saw SpaceX do another launch. They launch USA, it was really SpaceX yesterday when they were launching. You got Sovereign AI coming. All this is kind of enterprise-y.
Eiman Ebrahimi
>> These ones are popping. I think when you think about where the fastest moving folks have been, that's one part of where they've had the ability to use their data like you were talking about earlier, unlock their data faster and go faster, as long as they're able to meet the ROI of the use cases that they're going after. I think focusing again on what we're unlocking here is we're focused on how do you unlock use cases that weren't meeting the ROI because either they couldn't use the data that was most interesting to them, or if they tried to use it, it would become very expensive very quickly. When you think about that, the places where you find newfound movement are actually in those same places that you just mentioned within banking, within sovereign plays, within government. And that's a heavy focus for both of our companies. I do want to quickly tie back into the previous question that you asked about economics for the operators of these systems. so the customers of Rafay Systems or these sort of token factories, for those operators, the entire model of economics for those systems is dependent on high degrees of multi-tenancy. And that's one very critical component that is unlocked by this collaboration because being able to have true serverless, where you are just shipping off requests to wherever there is compute available in that moment, depends on the ability to be able to reason about what got exposed in those environments. And having the collaboration that we've been building upon is enabling that because again, the economics don't hold up at scale unless you have this or the end user will say, okay, I'm going to go back to wherever it is that I can actually send the data because nobody's going to peek at that.
John Furrier
>> Multi-tenancy absolutely jumps out as the table stakes for enterprise architecture, clearly. It used to be siloed before, hey build out a cluster, serve a workload. Now you have multiple workloads, or tenants basically. A workload's kind of a tenant.
Eiman Ebrahimi
>> Yeah, absolutely,
John Furrier
>> they go to run things. Okay, so multi-tenancy check, we're watching that closely. Agents, this is the future, they're going to be acting very microservices-like. They're going to touch a lot of resources, they're going to touch a lot of data. You got agents coming in, they're going to be more critical that they have privacy and data-bounded functions and intelligence.
Eiman Ebrahimi
>> I'm very happy you brought it up because agents are actually one of the biggest opportunities for these joint customers because agents are where that massive spike in token consumption is going to exist, but those tokens need to be served somewhere, right? And so we've seen, for instance, NVIDIA's put a lot of work into the ability to have the workspace of the agent be more secure with products like OpenShift and everything that's gone into that. But then the dependency on an LLM endpoint coming out of that harness is where this collaboration really unlocks the ability for an operator to be able to capture that demand.
John Furrier
>> It's almost like you're building, he uses the buzzword here, private AI factories. In a way, private cloud was not a thing, then became a thing, hybrid. It's the same kind of movie here. We're seeing the same thing play out from the cloud days, but at a different scale level, economics are involved, you've got privacy, data security, resilience needs, and also uncertainty around what's being developed. You want unlimited tokens, but you don't want to pay for it, but then scope it, scale it. Yeah, well let's coin a phrase today, virtual private token factory.
Haseeb Budhani
>> How about that? All right, let's go with that. VPTF, I guess so. Yeah, but this is kind of what's happening. Oh, very much so. The way I think about this is, 20 some years ago we would basically use our cell phones and the minutes were counted and we'd get a bill for inbound, outbound minutes of calls. This is where we are with tokens, right? We're token counting, accounting inbound, outbound. I'm convinced that soon enough we'll be buying, no, we buy cell phone plans. Unlimited plans. We'll have unlimited plans for tokens, right? And there'll be an individual plan and we'll have whatever tiers and enterprises will have plans. This is where the world is.
John Furrier
>> Is there a family plan in there?
Haseeb Budhani
>> I could use it.
John Furrier
>> My son goes through so many tokens,
Haseeb Budhani
>> I think we should be buying a company. But this is where clearly the world is going, and this is a tangent, but the telcos are very eager to find a way to monetize this because they have the consumer base, right?
John Furrier
>> Telcos, they are moving at glacial speeds. They have all the pieces. They got the power, they got the networking, they got the facilities.
Haseeb Budhani
>> Yeah, and the user base. They have 80, 100 million users who are basically, well, they're the customer, right? Right, and if there's a way to basically increase the ARPU, the average revenue per user, by selling them tokens at the edge of their telco network, that's an opportunity. Right, which I'm convinced is going to happen.
John Furrier
>> They have to do it, otherwise Starlink will come over the top and AWS will have some satellites. Guys, wrapping it up, I'm a little bit long here, but I'm really glad we got that multi-tenancy out, and I love that upstream. This is the future of ecosystems, these kinds of partnerships. How would you talk about your partnership going forward? What's next? Obviously it takes a village, a lot of people are involved in these deals. It's not simply hire one company, roll out AI infrastructure at scale.
Haseeb Budhani
>> Very much so.Yeah, look, to build one AI factory, eight, nine, ten different vendors come together to make it happen. That's just how it is. And the good news is, I think NVIDIA's kind of set a good stage for all of us. They do it this way. They're a very ecosystem-friendly company, and all of us, at least, I'm trying to emulate what they're doing and it's working. And what we're doing is, we're trying to educate our customer base. I get it that you want to optimize and monetize tokens. Here's another way for you to make more money. If you take these security solutions to your customer base, they will actually end up paying you more money, which means your margins are higher, but the customer's better off because they have a more secure model.
John Furrier
>> By the way, not to put a plug in for NVIDIA, but I will, Jensen, I love when he introduced KV cache a few years ago. I'm like, that's the killer app, and it took about a year. AI Factory was the year before these things come out. This year, he talked about monetization. He normally talks on stage. He's talking about physical AI, brings the robots out, talks about KV cache, Pareto curves, but this year he really highlighted the economics. This is kind of where the rubber's meeting the road.
Eiman Ebrahimi
>> Yeah, and this is straight at the heart of those economics because these systems, they were designed for this very large scale multi-tenant delivery. Without it, the economics don't hold. And so while these systems have been put into production effectively, what Haseeb is bringing up is really important. There are multiple different components that go into solving for the choke points of that system in order for the economics to be delivered. And we've found these are two very critical pieces of that story in the software stack. And so really excited about the partnership and looking forward to expanding it in both of our customer bases.
John Furrier
>> Well as people get more and more familiar with adopting software defined environments or intelligent systems injecting intelligence, with all that work, really makes it scale. Guys, thanks so much for coming on. Our third annual pool party event tonight. We'll see you there, and of course, thanks for participating in our program here. Thanks for coming on. Swim trunks or no? No swim trunks. Well, you're good. We can keep the streak going. We had one the first year. It's hot, you can go dip in the pool, we'll see. Thanks for coming on. Thank you. See you guys. All right, we got the pros, the CEOs, who are basically making it happen. This is what the new ecosystem looks like. Injecting intelligence into the enterprise requires a lot of work, takes a lot of players to put it together, and again, open but scalable, secure, multi -tenancy, you got the agents and multi -tenancy, that's the big theme here. Stay tuned for more after this short break.
>> Welcome back to theCUBE here in Palo Alto, California. I'm John Furrier, host of theCUBE. This is our third annual AI Leaders Summit, all day interviews here in our theCUBE NYSE Wired studios, and then the party after 180 leaders coming together to meet, talk about all the latest trends. This next segment focuses on the partnership between Protopia AI and Rafay Systems that really expands the addressable market for the AI infrastructure. We have both CEOs here. Eiman from Protopia, good to see you.
Eiman Ebrahimi
>> Great to see you as well.
John Furrier
>> Haseeb, good to see you again.
Haseeb Budhani
>> Good to see you again.
John Furrier
>> CEO of Rafay Systems. Guys, the addressable market for enterprise AI is growing super fast. The adoption is now kicking in. This token factory partnership you guys have that really opens the door for more operations, more development, explain the relationship between the two companies and really what this does to increase the addressable market for the enterprise.
Haseeb Budhani
>> Yeah, happy to do that.So, as you know, at Rafay, we sell an orchestration stack that AI infrastructure providers use to deliver a cloud experience. As part of our stack, we provide a token factory where my customers can deliver open source models to their downstream customers, enterprises, for example. A key part of our business and my customers' business is how do we deliver a serverless experience where we don't need to allocate GPUs on a per enterprise basis. So you can run many, many enterprises on the same infrastructure in a secure fashion. The gap so far has been how do we make sure the contexts are essentially multi-tenant? How do you make sure those are secure? Protopia has built a technology set that allows us to provide essentially a virtual private token factory to each enterprise that is engaging with my Neocloud customers. So it's a great technology, sits on top of a platform and delivers a better experience to the enterprise and makes the enterprise CISO happy because they get the right controls and governance in place.
John Furrier
>> Even on your end, token factory, what does that mean? How does that connect?
Eiman Ebrahimi
>> Yeah, I think one of the biggest realizations for us just being in market talking to the same types of customers that run token factories is that often they're thinking about how do they increase the amount of supply that they have and so there's this constant consideration that they have a supply problem but when you dig into how are you operating the underlying factory and how are you actually serving the tokens to your customers you see that there's a lot of over allocated capacity passive meaning just because something is booked if it's not busy that's revenue that they're leaving on the table and so when we started working with Rafay the aspect that Haseeb was just describing of the end enterprises wanting to have privacy of their information comes together with the desire from the operator side to maximize revenueand the easiest supply they're ever going to get access to is the supply that they already have if they are able to run true serverless without having to over -allocate, say, a node instead of a slice, or having to separate things physically.
John Furrier
>> So, I can see all these optimizations coming in my head's kind of connecting the dots, but let's zoom out and look at the problem. Most enterprises have unique needs. We all have been covering the topic of the data's still locked away, and that unlock represents a huge opportunity. a lot of regulated industries in the enterprise, but a lot of them haven't taken advantage of the shared GPU infrastructure because of many reasons. Is that the core problem, is that one of the problems? Because they don't have the huge capex, they just need service.
Haseeb Budhani
>> So I think most enterprises we all know, they are signing up for contracts with Anthropic or OpenAI or something like that. I think a big trend in the market right now is, okay, I'm going to use my cloud, but when it's 2pm and I get the message that says no more tokens, come back at 6pm, that happens to all of us. Okay, I want to be able to use open source models. Right, so there's a slew of new clouds who are jumping to that opportunity, which is they're now using their infrastructure to deliver tokens as a service, using open source models that are available in the market that we all are aware of. The challenge now is, okay, I have the infrastructure, I have the Token Factory from Rafay, so I have two options. I can pre-allocate these GPUs to the enterprise, which means they may use it, they may not, which means I'm leaving money on the table. Or I can come up with a...
John Furrier
>> You're basically provisioning.
Haseeb Budhani
>> Exactly, I provision and I walk away, right? If they're not being used, it's a waste of money. But the user, no, then the enterprise is saying, why am I paying for this, I'm not using it presently. What if there's a way to do a true serverless solution? The serverless part at the infrastructure level, we've solved that, and that we've done for a while. What Protopia has done is taken it to the next level. Right, so now we have a slice of essentially a confidential GPU, right? So I can now have better security on top of the GPUs that I already have, which means as a provider, I get to monetize every second of my infrastructure, but for the enterprise, it's actually a better deal. They get a better price point, because the infrastructure is now shared holistically, versus on a time-sliced basis. And that's a value, right? They like that because they look at the total cost of ownership, they look at the capex, they don't have the big money. It's too expensive, otherwise, right? And this allows them to have a secure solution, which actually delivers a better price point. That's what everybody wants.
John Furrier
>> You mentioned privacy. He mentioned privacy at the beginning. Is privacy or compute the bottleneck for adoption in the enterprise?
Eiman Ebrahimi
>> why is it one or the other? The problem is that they're tied together. The inability to perform private or confidential inference, whichever, however you want to name it, the inability to do that will result in needing to physically isolate compute that becomes a supply issue because there's not enough to physically separate for every given, say, business unit inside of an enterprise or every data owner that has sensitive information. So these two topics are actually tied at the hip and being able to solve one then really enables technologies like solutions like the token factories that Rafay has to be fully utilized by the operator inside an AI factory. And another important thing here is that we've seen this topic not just in NeoCloud and sovereign AI deployments, but also when you have many different data owners inside the same organization and you've got the organization trying to operate what is essentially inside the organization, an AI factory for those different data owners. It's the same problem. If business unit A and business unit B can't expose data in a multi-tenant system because they already have data silos where it's stored and when it's going to compute it also needs to be separated, they're going to land in the same issue.
John Furrier
>> Yeah, in the discussion around your relationship, there was a discussion on isolation tax. Explain what that means because a lot of people don't want to rewrite their applications. Or they don't want to waste resources they paid for that aren't being utilized. What does isolation tax mean, and what does that mean to the enterprise from an application or workload standpoint?
Haseeb Budhani
>> Yeah, look, the best way to have security is you own everything, right? The next level is we provide some level of abstraction, so they have a VM or a Kubernetes cluster, et cetera. Eventually, what everybody wants is to be able to drive their infrastructure at every second of every day for some user, right? and being able to context switch that is a hard problem. So from an infrastructure level, we get it all the way up to, you can run serverless requests across multiple GPUs that are shared across multiple enterprises. The problem that Eiman is describing is a real one, which is within an enterprise or across enterprises, I want a separation between my data and that guy's data. And that is a gap right now, it is what it is. So the only solution for that is we pre -allocate GPUs for customers, and that's the tax. So if you have a way to consume that infrastructure every second of every day, no tax. But if you're going to be in a position where
John Furrier
>> Unused resources.
Haseeb Budhani
>> That's right. If you're going to have four, six hours of non -usage, you're just paying for infrastructure that you're not utilizing. In that window of time, somebody else could use it, and you should get a better deal.
John Furrier
>> Smart money's going to be on this, in the enterprise, because they're going to focus on, where's the value? And what's the cost? They're going to look at that, wait, it's not being used, squeeze more cost out of that. Let's talk model and compatibility. One of the things that everyone talks about is, hey, I still got to use general intelligence, that's OpenAI, they have APIs, we've been using them. You're going to have specialism, specialty models, specialized intelligence, obviously domain expertise in the enterprise is going to be highly domain specific.
Haseeb Budhani
>> Yes, sir.
John Furrier
>> Talk about the compatibility issue to the bigger models.
Haseeb Budhani
>> The good news is, most of these models now, at least the ones that we are providing as part of our catalog, they're OpenAI-compatible. So as a developer, you don't need to switch. You don't need to unlearn or relearn anything. It's basically the same experience. The requirement for a vendor like us is, be it the models that NVIDIA's publishing, the Nemotron models or the NIM models, or the open source models like Kimi, et cetera. So what we have done is we sort of brought them all together in one place where the enterprise or the NeoCloud can pick and choose, I want to use this one or this one or this one. But then on top of that, there has to be a layer which is sort of essentially independent of the model itself to be able to provide some sort of data isolation. I'm using that phrase very loosely, of course, Eiman. And that's where this comes together. Right, so we now independently deliver a marketplace of models that are open, if you will. Right, open source, open models. And then now there's a data isolation layer on top, which is independent of the model. So this is the right way to think about it, because you said something earlier which is most important, right? the security and the lower cost, both have to be true for enterprises to adopt this at scale. And that's what's possible here.
John Furrier
>> Can we talk about the Stained Glass technology, you guys call it, what does it do? What's the secret sauce there and how does that tie in? Because you got multi -tenancy, you got privacy and compute, efficiency on the cost side of the value. Where is the Stained Glass technology? How does that fit in?
Eiman Ebrahimi
>> Yeah, Stained Glass, when you think of it just as the technology part of it, is focused on enabling models, whatever model we're talking about, to be able to make inferences on non -raw representations of the inference data and the context that goes to them. So today, in almost every single system that's out there, whether it's the frontier models, whether it's the open models being deployed on-prem, being deployed by a vendor on inference platforms, when the request arrives at the input on Ingress, there is a point where the context and everything about the prompt will go to raw information, plain text, and it will live that way through the entire inference data path. and what you see just recently without necessarily naming names there's multiple occurrences of even the best of vendors having data leaks that occur where information ends up being copied from one place to another when it shouldn't have information showing up in plain text entire code bases on systems where they shouldn't have been this isn't necessarily because the vendor is doing something poorly these are some of the best vendors out there But the fact is that there's all sorts of hidden data leak vectors across these systems. So we call these various operational surfaces that exist. And what Stained Glass is meant to do is to create an inference privacy layer that's way upstream from that and transforms the data from that raw representation into a, we call it a stochastic re-representation of the data that can be consumed by the target model without ever needing to go back to the original form. And when you do that upstream, it allows for all of those operational surfaces to not have to expose plain raw information. So if a leak like what we just saw a few weeks ago does happen again, and it will, and it is happening right now somewhere in many of these systems, right? The data doesn't need to show up as a leak in plain text, which is what's happening today. And so what the users of Private Token Factories that Rafay is building with us are going to have the ability to show to their customers is that at that upstream point, your data is being transformed out of plain text. And so wherever it may land, even with the best of everyone's efforts, then it's not going to be in plain text.
John Furrier
>> So it's a protection mechanism on the data.
Eiman Ebrahimi
>> That's correct.
John Furrier
>> Also people can maybe relate to this when I say, don't put company information in OpenAI. That's another, maybe is that related or is that just more of the same kind of upstream behavior?
Eiman Ebrahimi
>> I think there's two components to that. The part of it that we're focused on is complementing all of the zero data retention policies that the vendors are trying to provide. And I use the word trying very deliberately in that it's very difficult to say that zero data retention was absolute in any given system. And as evidenced by the fact that these sorts of leaks happen very frequently. And so what we're doing is saying for any such leakage that may happen even when you have zero data retention, it's not going to be in plain text. Now there's another component to that where people say don't put information in something because they wonder about whether or not the vendor is training on their data, which is a separate story. What we're talking about is not necessarily saying your vendor is trying to do something that they said to you they weren't going to do. we're just pointing out the fact that leaks happen.
John Furrier
>> I love, we're in Palo Alto, but also theCUBE, NYSE Wired, and New York Stock Exchange. It's fun to speak Silicon Valley and Wall Street, and I want to ask you guys a question, because the business impact of what you're doing is interesting. When I walk the halls of the NYSE, I talk to the traders, I talk to the analysts, even talk to some of the CNBC folks all the time, they're all asked the same questions. Is there a bubble? We address that very easily. But then when they start getting into questions: who's going to make the money? Where's the value? Which stock is going to go public? Who's going to hit escape velocity first? We're in an era where whoever builds the best AI infrastructure that aligns with how people think and work wins. That's clear, that's documented. What's the business impact? Because what you're saying to me increases the cloud addressable, cloud AI addressable market. So you see the neo -clouds like CoreWeave, Nscale, they're building out super fast. Nscale, talking to those guys, they're crushing it, right? But they have to monetize, so they have to get business value. So where does that fit in? Because you got NVIDIA trying to put together a partnership network, OpenAI's got an ecosystem. How does this make the addressable market and the economics and the money-making side work?
Haseeb Budhani
>> I think the key thing for all of us in the AI infrastructure space is we need to be able to deliver infrastructure that enterprises can consume. The same way they were so comfortable consuming AWS or GCP or Azure, the same things apply here, right? So the same controls, the same guardrails, the same quota management, the same auditability, observability, all of that has to be true, right? At every layer of the stack. AI is a different use case, of course, than the general purpose compute, right? So this is what we're all going after. you and I talked about this before, in this AI infrastructure, is this a bubble? I don't see it because the demand is incredible.
John Furrier
>> There's no bubble discussion. You build and you go.
Haseeb Budhani
>> Bubble discussion is ridiculous. Yeah, there's so much demand. If there was just infrastructure sitting around, that's a different conversation, but you're deploying, people are ready. They'll buy everything that's being shipped. There you go, right? That's basically what we're seeing. So the question now is, how do you bring the enterprises who are presently using an AWS environment, because they feel comfortable with AWS, how do we give them the same comfort? These are the steps we take. So at our level, we invest heavily in multi -tenancy, and you are very familiar with our stack, but this relationship takes it to the next level, So these are my words, of course not Eiman's, but this data level sort of segregation that is now possible using the Stained Glass technology that these guys have built, makes it easier for my Neo Cloud customer or my enterprise customer to say, hey man, I can bet on this. So the more comfort we can provide with solid technology, that solves the security problem, without taking away from the experience for the developer.
John Furrier
>> This is what's going to expand the opportunity. It also is a highlight of the partnership model we're seeing in this new era because you mentioned upstream. We saw NVIDIA do that, they put KV cache, now you got the storage vendors coming in. They have density, they're at a low level, but when you go upstream, up the stack, it's like, is my data secure? Can the application work with confidence? Can I audit it? Is there traceability? these are words that we know, observability.
Haseeb Budhani
>> That's right.
John Furrier
>> That's cloud, right? So this cloud game kind of comes in. Eiman, I want to ask you on your end, because you deal with a lot of the data and AI at the platform level. Which verticals are kind of ripe right now? healthcare used to be viewed as, oh, they're slow, HIPAA's so old, you can't move, it's antiquated, slow. They got data, though. See, healthcare's popping. You got banking, government. government's the hottest area. Defense tech, public-private, private's now leading the way. Just saw SpaceX do another launch. They launch USA, it was really SpaceX yesterday when they were launching. You got Sovereign AI coming. All this is kind of enterprise-y.
Eiman Ebrahimi
>> These ones are popping. I think when you think about where the fastest moving folks have been, that's one part of where they've had the ability to use their data like you were talking about earlier, unlock their data faster and go faster, as long as they're able to meet the ROI of the use cases that they're going after. I think focusing again on what we're unlocking here is we're focused on how do you unlock use cases that weren't meeting the ROI because either they couldn't use the data that was most interesting to them, or if they tried to use it, it would become very expensive very quickly. When you think about that, the places where you find newfound movement are actually in those same places that you just mentioned within banking, within sovereign plays, within government. And that's a heavy focus for both of our companies. I do want to quickly tie back into the previous question that you asked about economics for the operators of these systems. so the customers of Rafay Systems or these sort of token factories, for those operators, the entire model of economics for those systems is dependent on high degrees of multi-tenancy. And that's one very critical component that is unlocked by this collaboration because being able to have true serverless, where you are just shipping off requests to wherever there is compute available in that moment, depends on the ability to be able to reason about what got exposed in those environments. And having the collaboration that we've been building upon is enabling that because again, the economics don't hold up at scale unless you have this or the end user will say, okay, I'm going to go back to wherever it is that I can actually send the data because nobody's going to peek at that.
John Furrier
>> Multi-tenancy absolutely jumps out as the table stakes for enterprise architecture, clearly. It used to be siloed before, hey build out a cluster, serve a workload. Now you have multiple workloads, or tenants basically. A workload's kind of a tenant.
Eiman Ebrahimi
>> Yeah, absolutely,
John Furrier
>> they go to run things. Okay, so multi-tenancy check, we're watching that closely. Agents, this is the future, they're going to be acting very microservices-like. They're going to touch a lot of resources, they're going to touch a lot of data. You got agents coming in, they're going to be more critical that they have privacy and data-bounded functions and intelligence.
Eiman Ebrahimi
>> I'm very happy you brought it up because agents are actually one of the biggest opportunities for these joint customers because agents are where that massive spike in token consumption is going to exist, but those tokens need to be served somewhere, right? And so we've seen, for instance, NVIDIA's put a lot of work into the ability to have the workspace of the agent be more secure with products like OpenShift and everything that's gone into that. But then the dependency on an LLM endpoint coming out of that harness is where this collaboration really unlocks the ability for an operator to be able to capture that demand.
John Furrier
>> It's almost like you're building, he uses the buzzword here, private AI factories. In a way, private cloud was not a thing, then became a thing, hybrid. It's the same kind of movie here. We're seeing the same thing play out from the cloud days, but at a different scale level, economics are involved, you've got privacy, data security, resilience needs, and also uncertainty around what's being developed. You want unlimited tokens, but you don't want to pay for it, but then scope it, scale it. Yeah, well let's coin a phrase today, virtual private token factory.
Haseeb Budhani
>> How about that? All right, let's go with that. VPTF, I guess so. Yeah, but this is kind of what's happening. Oh, very much so. The way I think about this is, 20 some years ago we would basically use our cell phones and the minutes were counted and we'd get a bill for inbound, outbound minutes of calls. This is where we are with tokens, right? We're token counting, accounting inbound, outbound. I'm convinced that soon enough we'll be buying, no, we buy cell phone plans. Unlimited plans. We'll have unlimited plans for tokens, right? And there'll be an individual plan and we'll have whatever tiers and enterprises will have plans. This is where the world is.
John Furrier
>> Is there a family plan in there?
Haseeb Budhani
>> I could use it.
John Furrier
>> My son goes through so many tokens,
Haseeb Budhani
>> I think we should be buying a company. But this is where clearly the world is going, and this is a tangent, but the telcos are very eager to find a way to monetize this because they have the consumer base, right?
John Furrier
>> Telcos, they are moving at glacial speeds. They have all the pieces. They got the power, they got the networking, they got the facilities.
Haseeb Budhani
>> Yeah, and the user base. They have 80, 100 million users who are basically, well, they're the customer, right? Right, and if there's a way to basically increase the ARPU, the average revenue per user, by selling them tokens at the edge of their telco network, that's an opportunity. Right, which I'm convinced is going to happen.
John Furrier
>> They have to do it, otherwise Starlink will come over the top and AWS will have some satellites. Guys, wrapping it up, I'm a little bit long here, but I'm really glad we got that multi-tenancy out, and I love that upstream. This is the future of ecosystems, these kinds of partnerships. How would you talk about your partnership going forward? What's next? Obviously it takes a village, a lot of people are involved in these deals. It's not simply hire one company, roll out AI infrastructure at scale.
Haseeb Budhani
>> Very much so.Yeah, look, to build one AI factory, eight, nine, ten different vendors come together to make it happen. That's just how it is. And the good news is, I think NVIDIA's kind of set a good stage for all of us. They do it this way. They're a very ecosystem-friendly company, and all of us, at least, I'm trying to emulate what they're doing and it's working. And what we're doing is, we're trying to educate our customer base. I get it that you want to optimize and monetize tokens. Here's another way for you to make more money. If you take these security solutions to your customer base, they will actually end up paying you more money, which means your margins are higher, but the customer's better off because they have a more secure model.
John Furrier
>> By the way, not to put a plug in for NVIDIA, but I will, Jensen, I love when he introduced KV cache a few years ago. I'm like, that's the killer app, and it took about a year. AI Factory was the year before these things come out. This year, he talked about monetization. He normally talks on stage. He's talking about physical AI, brings the robots out, talks about KV cache, Pareto curves, but this year he really highlighted the economics. This is kind of where the rubber's meeting the road.
Eiman Ebrahimi
>> Yeah, and this is straight at the heart of those economics because these systems, they were designed for this very large scale multi-tenant delivery. Without it, the economics don't hold. And so while these systems have been put into production effectively, what Haseeb is bringing up is really important. There are multiple different components that go into solving for the choke points of that system in order for the economics to be delivered. And we've found these are two very critical pieces of that story in the software stack. And so really excited about the partnership and looking forward to expanding it in both of our customer bases.
John Furrier
>> Well as people get more and more familiar with adopting software defined environments or intelligent systems injecting intelligence, with all that work, really makes it scale. Guys, thanks so much for coming on. Our third annual pool party event tonight. We'll see you there, and of course, thanks for participating in our program here. Thanks for coming on. Swim trunks or no? No swim trunks. Well, you're good. We can keep the streak going. We had one the first year. It's hot, you can go dip in the pool, we'll see. Thanks for coming on. Thank you. See you guys. All right, we got the pros, the CEOs, who are basically making it happen. This is what the new ecosystem looks like. Injecting intelligence into the enterprise requires a lot of work, takes a lot of players to put it together, and again, open but scalable, secure, multi -tenancy, you got the agents and multi -tenancy, that's the big theme here. Stay tuned for more after this short break.