Leaders in robotics and artificial intelligence infrastructure discuss
device-scale foundation models, architecture alternatives to transformers and
enterprise deployment challenges. The conversation addresses model efficiency,
edge inference, hybrid on-device and cloud routing, and in-car multimodal
intelligence use cases to inform enterprise strategies around cost, data
sovereignty and privacy. Ramin Hasani of Liquid AI is co-founder and chief
executive officer and a leader in developing Liquid Foundation Models for
device-scale deployment. Hasani explains model architecture choices,
memory-optimized deployment and edge inference approaches. They highlight how
smaller Liquid Foundation Models combined with fine-tuning enable specialized
vertical solutions while reducing memory and compute footprints. Analysts on
theCUBE note transformer limits and key-value cache scaling drive exploration of
new architectures. They identify continuous learning, edge networking and
production data flywheels as central elements for enterprise deployment. The
discussion provides practical considerations for deploying foundation models at
device scale, including cost management, model efficiency, privacy and
regulatory compliance. Hosts John Furrier and Dave Vellante with analyst Howie
Xu conduct the theCUBE Research interview to surface implications for robotics
and AI infrastructure and deployment strategies.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Ramin Hasani, Liquid Ai
Leaders in robotics and artificial intelligence infrastructure discuss
device-scale foundation models, architecture alternatives to transformers and
enterprise deployment challenges. The conversation addresses model efficiency,
edge inference, hybrid on-device and cloud routing, and in-car multimodal
intelligence use cases to inform enterprise strategies around cost, data
sovereignty and privacy. Ramin Hasani of Liquid AI is co-founder and chief
executive officer and a leader in developing Liquid Foundation Models for
device-scale deployment. Hasani explains model architecture choices,
memory-optimized deployment and edge inference approaches. They highlight how
smaller Liquid Foundation Models combined with fine-tuning enable specialized
vertical solutions while reducing memory and compute footprints. Analysts on
theCUBE note transformer limits and key-value cache scaling drive exploration of
new architectures. They identify continuous learning, edge networking and
production data flywheels as central elements for enterprise deployment. The
discussion provides practical considerations for deploying foundation models at
device scale, including cost management, model efficiency, privacy and
regulatory compliance. Hosts John Furrier and Dave Vellante with analyst Howie
Xu conduct the theCUBE Research interview to surface implications for robotics
and AI infrastructure and deployment strategies.
>> I'm John Furrier, your host here in our Palo Alto studios at theCUBE and the NYSE Wired program. Our third annual conference where we're broadcasting the summit, which is an all day media session. And then we're going to have a party tonight, the pool party, and we're going to be having 180 leaders here, but of course we have our CUBE alumni back. Ramin Hasani, co -founder and CEO of Liquid AI, pioneering the Liquid Foundation models that he's pioneered. But also it's really the inference engine for small footprints. Howie Xu is my co -host analyst here, but also pioneering AI as well. Ramin, thanks for coming on theCUBE.
Ramin Hasani
>> Thank you, appreciate it.Thanks for having me.
John Furrier
>> So Howie and I have been having many chats over the years from super cloud days to now the AI. We've been speculating that specialty models are going to come in pretty strong to be a power law and that the big models will do their thing. They'll have billions of users, but you're going to start to see the emergence of different form factors. You're in the heart of it. Last time we chatted, you were making some great progress. What's the update?
Ramin Hasani
>> Yeah, great. So as we have been promoting, there are very different ways to build AIs. So we've been building the liquid foundation models. These are smaller models that can go outside of data centers. You can host them on a Raspberry Pi, from as small a processor as a Raspberry Pi, to hosting them on laptops, on PCs, and basically bring the power of a specialization or a specialized kind of foundation models to the devices. And then through those devices or those models that we put on devices, LFMs, Liquid Foundation Models, you can route intelligence very smartly between different mediums, right? So there's closed source models, there are larger open source models, there are hosted, sovereign versions of the AIs around you. So you have a lot of options for intelligence right now. And then there is the device category as well. So we are enabling that device category so that we can actually be the the entry point for intelligence, for personalized intelligence so that we can actually route intelligence between different mediums. So, what is the goal here? The goal here is to reduce the cost of intelligence, right? Because that's number one goal for enterprises, right? You are having attractive models that are creating a lot of, addiction for developers, that's what I say. And developers want to always work with the highest capable model. But not all requests need to go through to the cloud, right? To the largest version of a model. So you can actually deploy some of those use cases to specialized models. And then through that intelligent routing and intelligent kind of navigation of intelligence, you would be able to actually massively reduce the cost of inference among other properties.
Howie Xu
>> So you mentioned the cost being a key thing that enterprises are thinking about. But what about the rest of that? privacy, speed, the compliance. Do you see that being key use cases for people to come to you or cost is the number one priority?
Ramin Hasani
>> I think for enterprises, especially when they have a large set of developers, I would say cost is the primary kind of factor. Privacy for, let's say, consumers, it's not as big of a deal as for an enterprise because governance and privacy also for enterprises are extremely important. when they want to like use their AI, They don't want to send all of their proprietary information to the cloud and the biggest language models. Therefore, sovereign solutions like ours, like device intelligence plus, let's say, on -prem intelligence that we power, this is becoming very attractive. Of course, privacy is another factor that comes in. It's also very sector dependent. Depending on what sector you enter, the stress on privacy becomes very, very different.
John Furrier
>> The past year, since we last chatted, the memory price has jacked through the roof, okay? Compute now becomes much more of a centerpiece around the prefill-decode component of agents. So now the whole system architecture kind of, I won't say flips on its head, but it does open up more constraints. How do you fit into that solution? Because devices are memory sensitive, obviously footprint model sizes aren't the monster. rack scale, how's that impacting you guys?
Ramin Hasani
>> You know, efficiency has been a first class citizen as a model development stack that we have in-house at Liquid. We have always been thinking about the objective function is not just maximizing intelligence where we are building efficient AI. We were thinking about what is the computational footprint of this system, from the compute, literally how fast it can perform computations on the tokens. And then another axis, which is where does this intelligent system go? So for example, if you want to host on CPUs, you don't have an abundance of memory, right? And memory costs are increasing, let's say for mobile companies, for laptop companies, and for let's say consumer market in general, consumer electronic market, this is becoming really defining their profit margins. So memory is extremely important for us. And because we treated memory and this computational footprint as a first class citizen, we are building models that are not sacrificing on quality when we are talking about deployment on devices. Reliability of the models is very, very high from the performance, but from the memory footprint, we are substantially kind of reducing that cost. An example of this thing would be, recently we have done a deal with Mercedes-Benz, Mercedes-Benz for powering in-car intelligence. We are literally bringing Liquid Foundation Models, like audio language models. These are like multimodal intelligence systems. We are bringing them inside the cars to power in-car intelligence. So we are talking about you can talk to your car, You can control this. The first version of this thing would be a primary use case that you have, close the door, open the door, oh, it is hot. You can really have conversations with your model and also control anything that happens inside the car. But the evolution of this thing is something that I'm very, very excited about. And guess what is the size of this model that goes in there and powers the entire in-car intelligence? The size is 600 megabytes. Right. So it can actually fit onto like the cheapest kind of automotive processors that are mounted inside the car, right? So this is kind of the world that we wanted to enable now, hopefully this year, we are talking about production.
Howie Xu
>> So the small model, based on my observation, right, I've been working on a variety of the AI applications in the last few years, small model by and large never worked. Let me cut to the chase of it, right? The model became bigger and bigger, lately, like Kimi from Moonshot and of course, Anthropic, OpenAI, the model became bigger and bigger, more data, pre -training data, more parameters. Small model, once you get down to, eight gig, so, but it's sort of just small model, just step function decreasing performance, always, right? It never worked. Why is that? And then what did you do to make that fundamentally different?
Ramin Hasani
>> Great question.So think about what is it that these small models are bad at? It's generalization. General behavior. They cannot solve your kid's homework at night and at the same time solve the most sophisticated physics problems that you would have that you want to solve, let's say, the world's biggest problems, right? They cannot be like that kind of horizontally kind of enabling a lot of different kind of verticals. What they can do, they can be fine -tuned. You can customize them. The nice thing about this thing is that customizing a small model, the cost of it is going to be shockingly low. The bet that we put on was that can we really bring computational graphs that are very efficient onto processors like CPUs, neural processing units, NPUs, or consumer GPUs outside of data centers to ensure that we can do computation there? They have understanding of language, to some extent, you can add different languages to them, depending on the size that they have. Now you can specialize them. So the specialization process happens through a process called customization or fine tuning, right? You can take these models and fine tune them. It's not going to cost you that much. And you can get them specialized to really solve a dedicated problem really, really well.
Howie Xu
>> So you don't believe in small model for generalization.
Ramin Hasani
>> You don't quite believe in general purpose behavior.
Howie Xu
>> General intelligence.
Ramin Hasani
>> Exactly. It's not going to be like a small model. One small model is not going to be generally intelligent.
Howie Xu
>> So you are betting on small model plus fine tuning for specialized vertical use cases.
Ramin Hasani
>> Correct.
Howie Xu
>> And then you seem to also bet against Transformer. Can you talk more about it?
Ramin Hasani
>> Absolutely. So we have always thought about if you want to bring intelligence into the physical world, you have to understand those constraints. One of the constraints of the transformer architecture from the start was that the more information they process, the cost of computation exponentially kind of grows.
Howie Xu
>> Context window.
Ramin Hasani
>> Exactly. The longer the context window, you have exponentially harder time to perform computation. And this becomes both from the computation footprint and memory footprint. This becomes very prohibitive.
Howie Xu
>> KV cache is the bottleneck.
Ramin Hasani
>> Exactly. The KV cache problem really shows up. So anyway, everybody was talking about adding more efficient compression of KV cache, but we thought, okay, so let's explore the space of possibilities on the architecture front, not just putting a bet on a single architecture that we invented ourselves at MIT, but we wanted to explore what are other computational building blocks that can give rise to general purpose behavior. They can have scaling laws. They're not limited computational building blocks and they can actually scale to the scale of 50 to 100 billion parameters. This is kind of the regime that we are operating in Let's identify those operators and then test them on the hardware like build hybrid architectures that allow you to bring intelligence onto hardware That matches the properties and constraints of that device. So this way we kind of enable some sort of a hybrid world which is like a post -transformer kind of era.
Howie Xu
>> One more technical question.So basically, you are betting that some model architecture, so that you increase the context window or whatnot, the memory footprint requirement is not exponentially high, right? It's going to be linear, something like that. Why don't you do, why does the industry not do it for the large language model?
Ramin Hasani
>> They do. Now they do. So if you look at the Kimi K3, like the latest models that came out of the lab, MiniMax or China? Moonshot. Moonshot. So that model is actually from the architecture point of view, it has a lot of efficiencies. If you look at the original transformer architecture and the Kimi K3, you're going to see a massive difference between the architecture.
Howie Xu
>> They borrowed your concept.
Ramin Hasani
>> They are bringing this concept because it is absolutely essential. If you really want to bring intelligence at this level and distribute it to the entire world, efficiency is king. You have to work on efficiency at the core. So everybody's working on efficiency right now.
Howie Xu
>> Okay, technical question, to piggyback on that.
John Furrier
>> Small models, you guys optimize on that, congratulations.
Ramin Hasani
>> Thank you.
John Furrier
>> But now you have disaggregated infrastructure at the edge, because now that you're at the device level, you can have it either on PCs, wearables, sitting next to a telecom operators tower. what's the networking concept behind it? Because KV cache works great because it's tying all that together in the large rack scale. But when you're out in the edge, you don't have necessarily all that KV cache -like environment, but you might want to network to another device somewhere else in a distributed computing paradigm. What's the networking thoughts behind managing the data?
Ramin Hasani
>> That's a great question. So at the moment, what we are thinking is that let's bring the intelligence units on top of the devices and then the network problem is something that we have to really like like we have a lot of innovations to be done like this space is like actually like it's opening up like really really fast and I think networking between devices is something that we want to capitalize on it as a next kind of trajectory of this device AIs that we are talking about today we're counting on, let's say, single processors that are inside the cars and inside, let's say, mobile phones, and they are pretty powerful. You can have multiple models inside one mobile phone even. You can have multiple models on your laptop.
John Furrier
>> You can have a model to do the networking.
Ramin Hasani
>> Exactly, 100%. But then it becomes, again, the bandwidth problem, coverage, and it becomes more of a network kind of issue. And I think in the future we would definitely have systems that are going to be doing distributed computing in a way that we envision how blockchain technology as long as you can actually impact these as well.
John Furrier
>> One of the things I want to ask you guys, because before we went on camera, we were talking about you're an MIT PhD. He's a Stanford dropout, went to VMware with the chairman of the department, two PhDs. I'm the dumbest person here. So I have to ask you guys, if we were going to start a lab today, what would you guys do? Because there's a lot of opportunities out there. I know you got your deal now, but just as a thought exercise, you're involved in all of it too. What is the biggest problem that would be great for an entrepreneur to work on right now? Is it the networking side? Is it just another lab? Is it small language models? What would you guys do?
Howie Xu
>> You go first.You are the guest.
Ramin Hasani
>> Look, I think there are a lot of opportunities. For example, one that I'm very excited about is AI for science. So there's a lot of opportunities for, let's say AI for material discovery, AI for biology, AI for robotics, getting intelligence into physical world. So these are kind of the places that I think there's a lot of opportunities. All these kind of, this wave of LLMs, they enabled, let's say, a cloud -based intelligence kind of solution for everyone. I think the physical world is unexplored. Then there are data sets that are actually out there that we do not know of. They're not human understandable. Working on data sets is also another kind of area that we can help enterprises also using that data that they have to the way that they can actually bring sovereign intelligence to themselves, that's another kind of data would be like another area that I would really stress on. Maybe you can complete this.
Howie Xu
>> Well, I'm biased, right? I work for Gen Digital. We are consumer software company, right? For cybersecurity, financial awareness, personal assistance. From that point of view, I wanted to see more personal assistants that really work, right? Everyone should deserve an EA, executive assistant. Everyone should have a personal coach, right? Everyone should have someone to help them for the education, all that kind of thing. Personalized education, the medical, all those services, period. However, at the same time, I was so deeply worried that for the time being, for the foreseeable future, we are probably going to see more AI slop than necessary, right, the internet is going to be full of AI slop, in the coming year. and how do we deal with that?
John Furrier
>> Is that misinformation or hallucinations or just vanilla content?
Howie Xu
>> No, no, no. It's actually maybe even good content, but just not as useful. A lot of the videos on TikTok or whatnot. It's not bad. It's not malicious, but is that what human society needs, that, a million X, such content? I'm not sure, right? Even coding, right? I'm a big fan of vibe coding. I do that myself, but do we really need so many outputs?
John Furrier
>> So more noise.
Howie Xu
>> More noise. More noise is entering the system. Yeah, yeah, yeah. Your program has always been talking about signal out of the noise. You are going to extract more signals, but at the same time, let's face it, noise is going to be 100x, a millionx in the coming years. That I'm worried. On the flip side, I think there will be fundamental, fundamental things, right? AI for science, AI for longevity, AI for getting kids educated in a better way, more literacy, more food, less people dying too early, that kind of thing. To me, that's the fundamental.
John Furrier
>> Well, I like what you're working on, because you're working on the user experience shift. Because remember, we've had graphical user interfaces since I was in college in the 80s. That goes away in this era, because it's all how we think and work, which is natural language, whether it's voice or text. So that's changed, so I think you nailed it. On your side, I think what I'm excited about is the fact that companies could literally create their brain. The organizational brain is not one thing. It's a time series database. It's a graph database. It's a small language model. When you need it, you tap that model. You could have general intel, hey, I know the general Internet. I go to the large language. So I think we're going to have this integrated brain -like metaphor.
Ramin Hasani
>> I agree.
John Furrier
>> Where the output is utility.
Ramin Hasani
>> Absolutely. Absolutely. And then, when we talk about production AI, when we talk about building the brain of organizations, like the enterprises, we're thinking about, we always think about a model getting shipped into production and it is now getting used. what happens after when the model is actually in production? You're going to have a data flywheel. You're going to have something that actually manages like that kind of requests. How is this model, should this model be like a fixed weight neural network, or it should be a liquid neural network? So you got to think about like.
John Furrier
>> It's an organism, it's not a mechanism.
Ramin Hasani
>> Exactly, and you got to, I think there's a new category, like for us, from a science point of view, they're talking about this is a commercial term that everybody uses, recursive self -improvement, which is kind of like the scary thing that Frontier Labs are talking about. But I think I like the technical term that we call it continual learning.
John Furrier
>> I was talking to a friend, you'd appreciate this from me, I was talking to a friend, think of algae, it's in the ocean, it's algae, it's there, but it's not an organism.
Howie Xu
>> It's not evolving.
John Furrier
>> That's true. And that's why AI, if you think about that further, this comes up more technical, but there's a whole shift towards deterministic.
Ramin Hasani
>> Yes.
John Furrier
>> You look at physical AI where safety's on the line, you can't be wrong when you've got a car.
Howie Xu
>> You can't be wrong on a - OpenAI Hugging Face thing.
John Furrier
>> You can't be, and if you have a workload, you can go from non -deterministic to deterministic is because once you realize that's how it works, you can repeat it, that's deterministic. You guys' take on that because I think this is something that's coming up a lot in agents and physical AI where you can have a brain, but then when you lock in on a task, just don't get it wrong.
Howie Xu
>> That's true. I think continuous learning is really a big thing because I've worked on AI for years, right? In the old days, right? Well, seven, eight years ago, it's all about data drift, the feature drift, how do you deal with those challenges? Now, we have a new level of challenges. And plus, this OpenAI, Hugging Face thing, I think it's a,
John Furrier
>> Yeah, the leakage.
Howie Xu
>> It's a beginning of a new era for AI.
John Furrier
>> That leakage problem is just another vector or a surface area that, that's confidential computing, that's a privacy problem, that's a cyber problem. And that's another. That's not just a data problem.
Ramin Hasani
>> I would look at it also like from the, you know, like from new businesses, is also a direction that businesses can take, you know, cybersecurity as a whole, it's one of those areas that we need a lot more help. And a lot of big brains, like from the MITs and Stanford's and all the entrepreneurs that are coming out of like big labs, I think this is a space that I think is worth exploring.
John Furrier
>> Well, great to have you on, Howie. Great to have you on as my co -host and doubling as a guest, cause you're an expert. Got a mixture of experts here.
Howie Xu
>> By the way, just one last thing, I have a Norton Neo AI Browser. Do you think it would be able touse your small model effectively at the edge?
Ramin Hasani
>> 100%.I think this is one for the, let's say enterprise workload. And I think for personal kind of computing, I think this is one of those places.
Howie Xu
>> For privacy, for security.
Ramin Hasani
>> 100%, because you can build a brain around these type of models? I would love to work with you on this.
John Furrier
>> We're doing biz dev here on theCUBE, getting deals done. This is the market we're in. We are in an emerging AI infrastructure build out, which is also an enablement to accelerate intelligence. And when you have that intelligence built, you have to capture. So building intelligence, injecting it, and capturing that value all right here on theCUBE, doing our part. Thanks for watching. I'm John Furrier, host here with Howie Xu. Thanks for watching.
>> I'm John Furrier, your host here in our Palo Alto studios at theCUBE and the NYSE Wired program. Our third annual conference where we're broadcasting the summit, which is an all day media session. And then we're going to have a party tonight, the pool party, and we're going to be having 180 leaders here, but of course we have our CUBE alumni back. Ramin Hasani, co -founder and CEO of Liquid AI, pioneering the Liquid Foundation models that he's pioneered. But also it's really the inference engine for small footprints. Howie Xu is my co -host analyst here, but also pioneering AI as well. Ramin, thanks for coming on theCUBE.
Ramin Hasani
>> Thank you, appreciate it.Thanks for having me.
John Furrier
>> So Howie and I have been having many chats over the years from super cloud days to now the AI. We've been speculating that specialty models are going to come in pretty strong to be a power law and that the big models will do their thing. They'll have billions of users, but you're going to start to see the emergence of different form factors. You're in the heart of it. Last time we chatted, you were making some great progress. What's the update?
Ramin Hasani
>> Yeah, great. So as we have been promoting, there are very different ways to build AIs. So we've been building the liquid foundation models. These are smaller models that can go outside of data centers. You can host them on a Raspberry Pi, from as small a processor as a Raspberry Pi, to hosting them on laptops, on PCs, and basically bring the power of a specialization or a specialized kind of foundation models to the devices. And then through those devices or those models that we put on devices, LFMs, Liquid Foundation Models, you can route intelligence very smartly between different mediums, right? So there's closed source models, there are larger open source models, there are hosted, sovereign versions of the AIs around you. So you have a lot of options for intelligence right now. And then there is the device category as well. So we are enabling that device category so that we can actually be the the entry point for intelligence, for personalized intelligence so that we can actually route intelligence between different mediums. So, what is the goal here? The goal here is to reduce the cost of intelligence, right? Because that's number one goal for enterprises, right? You are having attractive models that are creating a lot of, addiction for developers, that's what I say. And developers want to always work with the highest capable model. But not all requests need to go through to the cloud, right? To the largest version of a model. So you can actually deploy some of those use cases to specialized models. And then through that intelligent routing and intelligent kind of navigation of intelligence, you would be able to actually massively reduce the cost of inference among other properties.
Howie Xu
>> So you mentioned the cost being a key thing that enterprises are thinking about. But what about the rest of that? privacy, speed, the compliance. Do you see that being key use cases for people to come to you or cost is the number one priority?
Ramin Hasani
>> I think for enterprises, especially when they have a large set of developers, I would say cost is the primary kind of factor. Privacy for, let's say, consumers, it's not as big of a deal as for an enterprise because governance and privacy also for enterprises are extremely important. when they want to like use their AI, They don't want to send all of their proprietary information to the cloud and the biggest language models. Therefore, sovereign solutions like ours, like device intelligence plus, let's say, on -prem intelligence that we power, this is becoming very attractive. Of course, privacy is another factor that comes in. It's also very sector dependent. Depending on what sector you enter, the stress on privacy becomes very, very different.
John Furrier
>> The past year, since we last chatted, the memory price has jacked through the roof, okay? Compute now becomes much more of a centerpiece around the prefill-decode component of agents. So now the whole system architecture kind of, I won't say flips on its head, but it does open up more constraints. How do you fit into that solution? Because devices are memory sensitive, obviously footprint model sizes aren't the monster. rack scale, how's that impacting you guys?
Ramin Hasani
>> You know, efficiency has been a first class citizen as a model development stack that we have in-house at Liquid. We have always been thinking about the objective function is not just maximizing intelligence where we are building efficient AI. We were thinking about what is the computational footprint of this system, from the compute, literally how fast it can perform computations on the tokens. And then another axis, which is where does this intelligent system go? So for example, if you want to host on CPUs, you don't have an abundance of memory, right? And memory costs are increasing, let's say for mobile companies, for laptop companies, and for let's say consumer market in general, consumer electronic market, this is becoming really defining their profit margins. So memory is extremely important for us. And because we treated memory and this computational footprint as a first class citizen, we are building models that are not sacrificing on quality when we are talking about deployment on devices. Reliability of the models is very, very high from the performance, but from the memory footprint, we are substantially kind of reducing that cost. An example of this thing would be, recently we have done a deal with Mercedes-Benz, Mercedes-Benz for powering in-car intelligence. We are literally bringing Liquid Foundation Models, like audio language models. These are like multimodal intelligence systems. We are bringing them inside the cars to power in-car intelligence. So we are talking about you can talk to your car, You can control this. The first version of this thing would be a primary use case that you have, close the door, open the door, oh, it is hot. You can really have conversations with your model and also control anything that happens inside the car. But the evolution of this thing is something that I'm very, very excited about. And guess what is the size of this model that goes in there and powers the entire in-car intelligence? The size is 600 megabytes. Right. So it can actually fit onto like the cheapest kind of automotive processors that are mounted inside the car, right? So this is kind of the world that we wanted to enable now, hopefully this year, we are talking about production.
Howie Xu
>> So the small model, based on my observation, right, I've been working on a variety of the AI applications in the last few years, small model by and large never worked. Let me cut to the chase of it, right? The model became bigger and bigger, lately, like Kimi from Moonshot and of course, Anthropic, OpenAI, the model became bigger and bigger, more data, pre -training data, more parameters. Small model, once you get down to, eight gig, so, but it's sort of just small model, just step function decreasing performance, always, right? It never worked. Why is that? And then what did you do to make that fundamentally different?
Ramin Hasani
>> Great question.So think about what is it that these small models are bad at? It's generalization. General behavior. They cannot solve your kid's homework at night and at the same time solve the most sophisticated physics problems that you would have that you want to solve, let's say, the world's biggest problems, right? They cannot be like that kind of horizontally kind of enabling a lot of different kind of verticals. What they can do, they can be fine -tuned. You can customize them. The nice thing about this thing is that customizing a small model, the cost of it is going to be shockingly low. The bet that we put on was that can we really bring computational graphs that are very efficient onto processors like CPUs, neural processing units, NPUs, or consumer GPUs outside of data centers to ensure that we can do computation there? They have understanding of language, to some extent, you can add different languages to them, depending on the size that they have. Now you can specialize them. So the specialization process happens through a process called customization or fine tuning, right? You can take these models and fine tune them. It's not going to cost you that much. And you can get them specialized to really solve a dedicated problem really, really well.
Howie Xu
>> So you don't believe in small model for generalization.
Ramin Hasani
>> You don't quite believe in general purpose behavior.
Howie Xu
>> General intelligence.
Ramin Hasani
>> Exactly. It's not going to be like a small model. One small model is not going to be generally intelligent.
Howie Xu
>> So you are betting on small model plus fine tuning for specialized vertical use cases.
Ramin Hasani
>> Correct.
Howie Xu
>> And then you seem to also bet against Transformer. Can you talk more about it?
Ramin Hasani
>> Absolutely. So we have always thought about if you want to bring intelligence into the physical world, you have to understand those constraints. One of the constraints of the transformer architecture from the start was that the more information they process, the cost of computation exponentially kind of grows.
Howie Xu
>> Context window.
Ramin Hasani
>> Exactly. The longer the context window, you have exponentially harder time to perform computation. And this becomes both from the computation footprint and memory footprint. This becomes very prohibitive.
Howie Xu
>> KV cache is the bottleneck.
Ramin Hasani
>> Exactly. The KV cache problem really shows up. So anyway, everybody was talking about adding more efficient compression of KV cache, but we thought, okay, so let's explore the space of possibilities on the architecture front, not just putting a bet on a single architecture that we invented ourselves at MIT, but we wanted to explore what are other computational building blocks that can give rise to general purpose behavior. They can have scaling laws. They're not limited computational building blocks and they can actually scale to the scale of 50 to 100 billion parameters. This is kind of the regime that we are operating in Let's identify those operators and then test them on the hardware like build hybrid architectures that allow you to bring intelligence onto hardware That matches the properties and constraints of that device. So this way we kind of enable some sort of a hybrid world which is like a post -transformer kind of era.
Howie Xu
>> One more technical question.So basically, you are betting that some model architecture, so that you increase the context window or whatnot, the memory footprint requirement is not exponentially high, right? It's going to be linear, something like that. Why don't you do, why does the industry not do it for the large language model?
Ramin Hasani
>> They do. Now they do. So if you look at the Kimi K3, like the latest models that came out of the lab, MiniMax or China? Moonshot. Moonshot. So that model is actually from the architecture point of view, it has a lot of efficiencies. If you look at the original transformer architecture and the Kimi K3, you're going to see a massive difference between the architecture.
Howie Xu
>> They borrowed your concept.
Ramin Hasani
>> They are bringing this concept because it is absolutely essential. If you really want to bring intelligence at this level and distribute it to the entire world, efficiency is king. You have to work on efficiency at the core. So everybody's working on efficiency right now.
Howie Xu
>> Okay, technical question, to piggyback on that.
John Furrier
>> Small models, you guys optimize on that, congratulations.
Ramin Hasani
>> Thank you.
John Furrier
>> But now you have disaggregated infrastructure at the edge, because now that you're at the device level, you can have it either on PCs, wearables, sitting next to a telecom operators tower. what's the networking concept behind it? Because KV cache works great because it's tying all that together in the large rack scale. But when you're out in the edge, you don't have necessarily all that KV cache -like environment, but you might want to network to another device somewhere else in a distributed computing paradigm. What's the networking thoughts behind managing the data?
Ramin Hasani
>> That's a great question. So at the moment, what we are thinking is that let's bring the intelligence units on top of the devices and then the network problem is something that we have to really like like we have a lot of innovations to be done like this space is like actually like it's opening up like really really fast and I think networking between devices is something that we want to capitalize on it as a next kind of trajectory of this device AIs that we are talking about today we're counting on, let's say, single processors that are inside the cars and inside, let's say, mobile phones, and they are pretty powerful. You can have multiple models inside one mobile phone even. You can have multiple models on your laptop.
John Furrier
>> You can have a model to do the networking.
Ramin Hasani
>> Exactly, 100%. But then it becomes, again, the bandwidth problem, coverage, and it becomes more of a network kind of issue. And I think in the future we would definitely have systems that are going to be doing distributed computing in a way that we envision how blockchain technology as long as you can actually impact these as well.
John Furrier
>> One of the things I want to ask you guys, because before we went on camera, we were talking about you're an MIT PhD. He's a Stanford dropout, went to VMware with the chairman of the department, two PhDs. I'm the dumbest person here. So I have to ask you guys, if we were going to start a lab today, what would you guys do? Because there's a lot of opportunities out there. I know you got your deal now, but just as a thought exercise, you're involved in all of it too. What is the biggest problem that would be great for an entrepreneur to work on right now? Is it the networking side? Is it just another lab? Is it small language models? What would you guys do?
Howie Xu
>> You go first.You are the guest.
Ramin Hasani
>> Look, I think there are a lot of opportunities. For example, one that I'm very excited about is AI for science. So there's a lot of opportunities for, let's say AI for material discovery, AI for biology, AI for robotics, getting intelligence into physical world. So these are kind of the places that I think there's a lot of opportunities. All these kind of, this wave of LLMs, they enabled, let's say, a cloud -based intelligence kind of solution for everyone. I think the physical world is unexplored. Then there are data sets that are actually out there that we do not know of. They're not human understandable. Working on data sets is also another kind of area that we can help enterprises also using that data that they have to the way that they can actually bring sovereign intelligence to themselves, that's another kind of data would be like another area that I would really stress on. Maybe you can complete this.
Howie Xu
>> Well, I'm biased, right? I work for Gen Digital. We are consumer software company, right? For cybersecurity, financial awareness, personal assistance. From that point of view, I wanted to see more personal assistants that really work, right? Everyone should deserve an EA, executive assistant. Everyone should have a personal coach, right? Everyone should have someone to help them for the education, all that kind of thing. Personalized education, the medical, all those services, period. However, at the same time, I was so deeply worried that for the time being, for the foreseeable future, we are probably going to see more AI slop than necessary, right, the internet is going to be full of AI slop, in the coming year. and how do we deal with that?
John Furrier
>> Is that misinformation or hallucinations or just vanilla content?
Howie Xu
>> No, no, no. It's actually maybe even good content, but just not as useful. A lot of the videos on TikTok or whatnot. It's not bad. It's not malicious, but is that what human society needs, that, a million X, such content? I'm not sure, right? Even coding, right? I'm a big fan of vibe coding. I do that myself, but do we really need so many outputs?
John Furrier
>> So more noise.
Howie Xu
>> More noise. More noise is entering the system. Yeah, yeah, yeah. Your program has always been talking about signal out of the noise. You are going to extract more signals, but at the same time, let's face it, noise is going to be 100x, a millionx in the coming years. That I'm worried. On the flip side, I think there will be fundamental, fundamental things, right? AI for science, AI for longevity, AI for getting kids educated in a better way, more literacy, more food, less people dying too early, that kind of thing. To me, that's the fundamental.
John Furrier
>> Well, I like what you're working on, because you're working on the user experience shift. Because remember, we've had graphical user interfaces since I was in college in the 80s. That goes away in this era, because it's all how we think and work, which is natural language, whether it's voice or text. So that's changed, so I think you nailed it. On your side, I think what I'm excited about is the fact that companies could literally create their brain. The organizational brain is not one thing. It's a time series database. It's a graph database. It's a small language model. When you need it, you tap that model. You could have general intel, hey, I know the general Internet. I go to the large language. So I think we're going to have this integrated brain -like metaphor.
Ramin Hasani
>> I agree.
John Furrier
>> Where the output is utility.
Ramin Hasani
>> Absolutely. Absolutely. And then, when we talk about production AI, when we talk about building the brain of organizations, like the enterprises, we're thinking about, we always think about a model getting shipped into production and it is now getting used. what happens after when the model is actually in production? You're going to have a data flywheel. You're going to have something that actually manages like that kind of requests. How is this model, should this model be like a fixed weight neural network, or it should be a liquid neural network? So you got to think about like.
John Furrier
>> It's an organism, it's not a mechanism.
Ramin Hasani
>> Exactly, and you got to, I think there's a new category, like for us, from a science point of view, they're talking about this is a commercial term that everybody uses, recursive self -improvement, which is kind of like the scary thing that Frontier Labs are talking about. But I think I like the technical term that we call it continual learning.
John Furrier
>> I was talking to a friend, you'd appreciate this from me, I was talking to a friend, think of algae, it's in the ocean, it's algae, it's there, but it's not an organism.
Howie Xu
>> It's not evolving.
John Furrier
>> That's true. And that's why AI, if you think about that further, this comes up more technical, but there's a whole shift towards deterministic.
Ramin Hasani
>> Yes.
John Furrier
>> You look at physical AI where safety's on the line, you can't be wrong when you've got a car.
Howie Xu
>> You can't be wrong on a - OpenAI Hugging Face thing.
John Furrier
>> You can't be, and if you have a workload, you can go from non -deterministic to deterministic is because once you realize that's how it works, you can repeat it, that's deterministic. You guys' take on that because I think this is something that's coming up a lot in agents and physical AI where you can have a brain, but then when you lock in on a task, just don't get it wrong.
Howie Xu
>> That's true. I think continuous learning is really a big thing because I've worked on AI for years, right? In the old days, right? Well, seven, eight years ago, it's all about data drift, the feature drift, how do you deal with those challenges? Now, we have a new level of challenges. And plus, this OpenAI, Hugging Face thing, I think it's a,
John Furrier
>> Yeah, the leakage.
Howie Xu
>> It's a beginning of a new era for AI.
John Furrier
>> That leakage problem is just another vector or a surface area that, that's confidential computing, that's a privacy problem, that's a cyber problem. And that's another. That's not just a data problem.
Ramin Hasani
>> I would look at it also like from the, you know, like from new businesses, is also a direction that businesses can take, you know, cybersecurity as a whole, it's one of those areas that we need a lot more help. And a lot of big brains, like from the MITs and Stanford's and all the entrepreneurs that are coming out of like big labs, I think this is a space that I think is worth exploring.
John Furrier
>> Well, great to have you on, Howie. Great to have you on as my co -host and doubling as a guest, cause you're an expert. Got a mixture of experts here.
Howie Xu
>> By the way, just one last thing, I have a Norton Neo AI Browser. Do you think it would be able touse your small model effectively at the edge?
Ramin Hasani
>> 100%.I think this is one for the, let's say enterprise workload. And I think for personal kind of computing, I think this is one of those places.
Howie Xu
>> For privacy, for security.
Ramin Hasani
>> 100%, because you can build a brain around these type of models? I would love to work with you on this.
John Furrier
>> We're doing biz dev here on theCUBE, getting deals done. This is the market we're in. We are in an emerging AI infrastructure build out, which is also an enablement to accelerate intelligence. And when you have that intelligence built, you have to capture. So building intelligence, injecting it, and capturing that value all right here on theCUBE, doing our part. Thanks for watching. I'm John Furrier, host here with Howie Xu. Thanks for watching.