Mansour Karam, Aria Networks | theCUBE + NYSE Wired: Robotics & AI Infra Leaders
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: The AI Factory - Data Center of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: The AI Factory - Data Center of the Future. Signing in with LinkedIn ensures a professional environment.
>> Palo Alto Studio, connecting Silicon Valley and Wall Street. I'm John Furrier, the host of theCUBE here with Gabe Olave, my co-host. theCUBE, NYSE Wired Robotics and AI Infra Leaders Series is brought to you by ScaleFlux. Welcome back to theCUBE here in Palo Alto, California. I'm John Furrier, host of theCUBE. This is our third annual poolside party and AI summit all day here in our Palo Alto studio. My co-host, Howie Xu, is here breaking down all the action. We got an entrepreneur here, Mansour Karam, founder and CEO of Aria Networks. Great to see you again. We were just talking off camera about all the history of networking, and now we got the AI projects. Great to see you. We were sitting together in Paris at the RAISE Summit.>> In Versailles.
John Furrier
>> Great event. The AI infrastructure has the biggest build-out we've ever seen. Every piece of infrastructure is in demand. The demand curve is off the charts. And over the past 3 years, networking has emerged finally as the number one issue and technology that's driving everything. We look at all the AI factories, KV Cache, we talk about InfiniBand prior life, but now it's come full circle where if you don't have good networking, the AI really doesn't work very well.>> That's exactly right.>> This is where we're at.>> Well, yeah, it's actually really a fun time if you're a networking guy like I am. I've been in networking all my career. it's a renaissance of sorts. But yes, AI factories are a brand new use case with very different requirements, in fact, stringent requirements. And as you said, the network is at the center of everything. And in fact, that's the impetus for why we started Aria Networks. If you think about it, the networking is maybe 10% of the spend depending on which cluster it is or which infrastructure, 5 to 15%. Yet it touches every component. Right? And so if the network is not working well, is not optimized, it affects every metric. It affects the performance of the cluster as a whole. And depending on the use case, the impact can be dramatic.
John Furrier
>> Yeah. And there's dollars now associated with it. The GPUs are millions and millions of dollars. If they're not being fed the data, it's game over or disruption, loss of money, direct impact to user experience. everything. Kind of crumbles.>> Correct. It's not millions, millions. It's billions.>> Billions. Billions.>> Right.
John Furrier
>> So I'm not talking per GPU, but it's a system now. It's not just a server and a switch. the switches are all in the high-density environment. So it's interesting you're co-hosts, but you're also both experts in networking. This is fundamentally the biggest thing. I guess my question for you and Hari too is what's been the biggest change in networking? Obviously, we see the KV Cache relevance that's now 3 years in. People are seeing that, but now you're looking at the edge in other areas, disaggregated inference. We see disaggregated infrastructure, networking with intelligence now has the classic intersection opportunity saying, oh, you got networks for AI and you got AI for networks. We've heard that a lot. What does that mean in practice?>> That's right. So, well, the first thing that has changed in networking is that we just see the speeds and feeds get bigger and bigger to support the type of throughput that you need and also the latency, right? Specifically, in AI clusters. So that's kind of at the foundation. You need switches of that caliber to support those use cases. So that's number one. But number two, what's really critical is that you want to optimize the performance of these networks. You want to be able to detect issues very quickly. You want to detect inefficiencies. You can have the load on the network be inefficiently spread out. And so for that, when you talk about AI for networking, it's the notion of specializing AI for the networking use case so that intuitively you have a capability that is always detecting issues, resolving them on the fly, right? And then alerting users when they need to be involved. For example, if they need to go change a transceiver or make a modification to their network. And so a lot of what we're doing here with Aria is taking AI and specializing it the same way for every other domain. You want to specialize AI to the domain. We specialize AI to the networking use case in order to deliver on this requirement. And we can dig into what that means, right? Because it is not just a matter of slapping an LLM on top of an already existing architecture, which a lot of times we see that, right?
John Furrier
>> You can't crawl the internet for that one because general intelligence, they crawl the internet, but you have a lot of domain expertise in networking. There's nuanced situations, there's configurations.>> Well, it starts with collecting telemetry, right? And there are two aspects to collecting telemetry. You want to collect it at the proper resolution. Typical networking software or equipment collects telemetry at the second or 30-second interval, right? This is what we did in the past and in past experiences. It was totally fine if your goal is just to validate configuration, right? But if you want to be at the speed of AI, optimizing the network for AI, 1 second is like a century in human scale, right? So many things could happen within a second. And so you need to be collecting telemetry at the microsecond resolution. So that's number 1, you will need it at the right resolution. And number 2, the network, and we can talk about all the different networks in an AI factory, but it spans very different domains. You have the front-end network, the back-end network. You not only have the switches, but you also have the network cards. You have networking happening on the host. You have the cables, the >> transceivers.You have new technologies like photonics coming to the mix.>> You have photonics coming to the mix.
John Furrier
>> All kinds of speeds and feeds. If you look at the NVL rack, Howie, it's the 72, half the rack is switches.
John Furrier
>> well, yeah.
John Furrier
>> it's a monster machine, but A bulk of the density are switches.>> Yeah. And if you have, for any of those cables, if you have a higher bit error rate, you can imagine what it does, right? So a lot of these queries when they come in, especially for inference, also for training, but for inference it's critical because the user is waiting. They get divided into many, many branches and then you have so many opportunities for tail latency to occur because you're waiting for the one that is gonna take the longest, right? And so if you have one of those thousands that waits, then everyone else is waiting.
John Furrier
>> So I'm going to call up my inner Andy Bechtolsheim and say, it's a physics problem.
John Furrier
>> Okay.
John Furrier
>> Because physics are involved here. So he's always a skeptic on physics. That'll never work. Or he leans in. So we're dealing with a lot of physical latency. You see photonics now coming in. You got copper. NVIDIA relies on copper. That's the big thing for them because they want to have that short distance. But Can you eliminate that? How many hops are in those switches? This is like old, this is networking stuff, but this would slow stuff down. You lose some heat and energy.>> That's right. And so you're going to have all these components and it gets more and more complex. But then what's critical is to have an ability to get telemetry to measure all of these components end to end at the right resolution. And that's step number one.>> Is the telemetry mostly to reduce the operational cost or is it more than that?>> Well, the telemetry is to see what's going on, right? Because if you don't have a picture of what's going on, then you never know if you're solving the right problem. You cannot root cause effectively or fix the problem effectively. So it starts by collecting telemetry, but then you want to extract knowledge from this telemetry, extract the right signal so that you know what's going on. And you need to do that at massive scale. And then from there, you need to have an ability to go in and make the right modifications. And that's where AI comes in. But again, AI first, there is that foundation of telemetry and then you have to be very sophisticated in how you're layering this AI in order for it to work for the networking use >> case.So what's the AI behind the scenes? Is it like the large language model from OpenAIs of the world plus fine-tuning or you have your own model on device, whatnot?>> Well, it's all of the above. So I'll tell you what you need to have. It's kind of think of the nervous system, right? In a human. If you put your finger on a flame, right? Your body will react. You don't even have to think about it. So that's one layer, right? That is the deterministic extreme quick reaction if something happens, that's happening at the lowest layers. Then there are multiple layers that go above where you're essentially taking on a bigger problem and you have an ability to reason about it at a larger scale. And you do that all the way to when you have to deploy sophisticated LLMs in the way you're reasoning about the data center as a whole, and then communicating that back and forth with the operator or maybe agents, because a lot of these softwares are going to be agentic. So we have agentic interfaces and we expect that other agents will be able to use the softwares. Basically, that's the world we're evolving towards.>> Yeah.>> So your shirt says networks that think, right? So that's the AI part. What about the speeds and feeds? You mentioned earlier, right? In the old days, it's more about speeds and feeds. Is that a solved problem or you're kind of leveraging AI to actually make the speeds and feeds a better deal for the customers? How do you think about it? >> Rightspeeds and feeds generally, it's a physics problem.>> So that's not a key problem.>> For us, we are a system company, so we're not building chips. But we are always looking at leveraging all the best chips, the best optics that exist for a given use case, right? So it's about how we put all of these things together. And then again, where AI comes in is that you may see problems or you may become aware of problems and AI can extract that information at the critical seconds before the problem occurs, right? Or in some cases many minutes before. Optics have a very specific pattern in the way they fail. You want to have an ability to go and detect that optic is about to fail before it fails. So somebody replaces it before it becomes an issue in production, right? So yes, there is a role for AI in terms of just managing the infrastructure, right? But at the end, there is the software and then there is the hardware. We do like to say that we co-design it all together, right? Because at the end, if you want to extract telemetry at the microsecond resolution, you have to be really deep into the hardware, right? In some cases we have agents that are sitting on ASICs, right? And basically embedded code that is extracting telemetry. So it's critical that both the software and the hardware work really well together. But there is a separation.
John Furrier
>> Remember the old days, there was a company called Thinking Machines. That's AI for everything that we're seeing embedded in. How's the company doing? Give a quick update on Aria Networks, how you guys are doing, the momentum, secret sauce. Obviously telemetry sounds like a super big part of it.>> Yeah, this deep networking layer, this co-designed hardware and software, including the telemetry and all of these agentic layers. That's part of the secret sauce. we're doing great. I haven't seen infrastructure evolve at the speeds that it has, right? So in every other company or project I was involved in, it took years before first PO. It took a very long time before we shipped anything into production.>> Scope the problem.>> We remember Big Switch days, right?
John Furrier
>> But there's a lot of engineering, but it's super fast now. What problem are you going after? Is it taking the latency out, energy, what's the focus?>> Right, so that's actually a great question. Our software, we've designed it in a way that is adaptive to whatever optimization you need. So it continuously and dynamically optimizes to your use case. And so it depends on the metric. So we discussed the fact that we're moving away from generic clusters towards more specialized clusters. And for every one of those clusters, you may have a different use case, a different metric. For example, if you're doing high throughput inference, then your metric is really tokens, dollars per million tokens, right? And so what you want is, these are simple inference use cases and you wanna be able to do them as many as possible, as cheaply as possible. Compare that to a high interactivity inference. There you have a lot more sophisticated inference job that requires a lot more complex computing and KV-Cache transfers. In some cases, you are separating the prefill from the decode. So it's a lot more complex. But then there, what's really critical is the number of tokens per second per user, because you have for every user, depending on how many tokens they get, per second, right? This is how fast the query is gonna get back to them or the response is gonna get back to them. And on the other hand, you may have a cluster that is the basic limit is the power, right? So there you wanna be able to extract as many tokens per watt as possible. So these are different clusters.>> So they're different. Totally situational. Based upon, for example, we think fast, as humans, the way we work, if you're in that stream, I might be using premium tokens, I might need premium resource. It sounds like policy, like the old days, policy networking. Well, I mean, it's policy.>> That's right.>> And it's happening.>> So it's in a sense, it's kind of, you have a different performance profile that you're optimizing towards.>> Right.>> And so, the job for the software is to adapt to these situations, right? And where AI gets super, it can get super intuitive, right? It's perfectly adapted to those types of use cases, right? And so it can adapt to the different use case intuitively.
John Furrier
>> So it's a networking concept, but you're combining all the old school techniques and thoughts like switch versus a router because you're doing a lot of things. The software is making the call, right? Is that right?>> Yeah, think of, It's an overused analogy, but it's like self-driving cars. You're not reinventing the wheels or the suspension, right? Or the fact that it has this many doors. It's a bit like that. So you mentioned InfiniBand. It's clear now that Ethernet is going to be the protocol of choice, right? And that has been the case for 30 years. So you can use the same exact management tools as before because The formats haven't changed. The standards haven't changed. So we're not changing that.
John Furrier
>> The rise of photonics has come up a lot in networking in certain use cases. Ethernet as the standard.>> Yeah, Ethernet as the standard. So the components themselves, other than them being a lot more, having more speeds and feeds, there are a lot of things that don't change. The protocols are pretty much the same, right? Whether it's routing or switching, But exactly to your point, right, the software layers that you're bringing in to essentially deliver on the requirement, have to leverage the latest and greatest >> technology.So in summary, in simple terms, you're injecting intelligence.
John Furrier
>> Totally.
John Furrier
>> To the networking.>> Yeah, exactly. We're leveraging the latest and greatest technologies, including obviously AI, including all this fine-grained telemetry, right? Putting it all together and bringing it to bear to deliver on this use case. Right?>> So several decades ago, right? Ethernet, TCP/IP won out over ATM, InfiniBand, that kind of technology because the alternative is too complex, right? Very hard to configure, make sense out of it. Are you saying that, look, AI will add sophistication to the network, the Networks that Think, but it's in the doing that, being done in the autonomous way, in the automagic way. So it's not, the sophistication without the complexity.>> Right, that's exactly right. But one thing I want to clarify is that I'm not a big fan of self-operating or self-driving networks because at the end—>> You still need a human first, AI assist.>> Networking, think of an autopilot in a plane, you can't get rid of the pilot, right? And I think it's the same in networking. When it's so critical, it needs to run. You don't want the network to self-drive off a cliff. You want the operator.>> And you got security concerns as well. We saw the Hugging Face stuff that's going out like, oops, that's a vector.>> 100%. And so this is where Networks that Think essentially think there is an operator or operators, but it helps them think, it helps them reason. And you're giving them all of the data. You're making them feel like it's at their fingertips in ways that were never possible before. And that is really the power of AI. It's kind of like how you use AI in a sense. It's not replacing me, or at least not yet, but it's empowering me incredibly. And I think it's the same here, adapted to the networking use case.
John Furrier
>> Well, soon you'll be able to hike the Dish in Palo Alto and talk to your machine that does all the coding for you, runs all your business. That's where AI is coming. We see some funny posts where people— how people work and think are different. There's no GUI anymore. There's a lot more—
John Furrier
>> That's exactly right.
John Furrier
>> Automation, a lot more intelligence.>> Yeah, so we don't have the static dashboard like you would have in typical management solutions for networking like we had in the past. You start with an interface. Well, first of all, you should assume that agents are using your product, but also as an operator, it's only the alerts that matter and then you're having the ability to ask it questions and then it will pull up whatever visuals are helpful to you in the task that you're performing at the time.>> So in your model, humans or the operators are still in the driver's seat. And then—>> Yes.>> But they have a lot more sophistication when it comes to telemetry, intelligence, >> visibility.Yeah, that's the idea. It's orders of magnitude more power in terms of their ability to see everything and then to control
John Furrier
>> everything.Well, Mansour, we're going to have to get some more time with you. Congratulations on—
John Furrier
>> Thank you. >> —Aria Networks. We love the intelligent networks, the edge. We haven't even talked about edge computing. There is a lot more we can talk about. This won't be the last. Thanks for coming on.>> Appreciate it.
John Furrier
>> Thank you, MansourThank you so much. I'm John Furrier, host of theCUBE. This is our AI Summit with theCUBE and NYSE Wired. Again, we are here in Palo Alto, full media day, a big event tonight, 180 practitioners and industry insiders and entrepreneurs talking about how the AI industry is going to be built out and how they're moving the needle and also what it's going to enable as the value capture comes around after the intelligence is injected into the network and the applications. We're doing our part here on theCUBE. Thanks for watching.
>> Palo Alto Studio, connecting Silicon Valley and Wall Street. I'm John Furrier, the host of theCUBE here with Gabe Olave, my co-host. theCUBE, NYSE Wired Robotics and AI Infra Leaders Series is brought to you by ScaleFlux. Welcome back to theCUBE here in Palo Alto, California. I'm John Furrier, host of theCUBE. This is our third annual poolside party and AI summit all day here in our Palo Alto studio. My co-host, Howie Xu, is here breaking down all the action. We got an entrepreneur here, Mansour Karam, founder and CEO of Aria Networks. Great to see you again. We were just talking off camera about all the history of networking, and now we got the AI projects. Great to see you. We were sitting together in Paris at the RAISE Summit.>> In Versailles.
John Furrier
>> Great event. The AI infrastructure has the biggest build-out we've ever seen. Every piece of infrastructure is in demand. The demand curve is off the charts. And over the past 3 years, networking has emerged finally as the number one issue and technology that's driving everything. We look at all the AI factories, KV Cache, we talk about InfiniBand prior life, but now it's come full circle where if you don't have good networking, the AI really doesn't work very well.>> That's exactly right.>> This is where we're at.>> Well, yeah, it's actually really a fun time if you're a networking guy like I am. I've been in networking all my career. it's a renaissance of sorts. But yes, AI factories are a brand new use case with very different requirements, in fact, stringent requirements. And as you said, the network is at the center of everything. And in fact, that's the impetus for why we started Aria Networks. If you think about it, the networking is maybe 10% of the spend depending on which cluster it is or which infrastructure, 5 to 15%. Yet it touches every component. Right? And so if the network is not working well, is not optimized, it affects every metric. It affects the performance of the cluster as a whole. And depending on the use case, the impact can be dramatic.
John Furrier
>> Yeah. And there's dollars now associated with it. The GPUs are millions and millions of dollars. If they're not being fed the data, it's game over or disruption, loss of money, direct impact to user experience. everything. Kind of crumbles.>> Correct. It's not millions, millions. It's billions.>> Billions. Billions.>> Right.
John Furrier
>> So I'm not talking per GPU, but it's a system now. It's not just a server and a switch. the switches are all in the high-density environment. So it's interesting you're co-hosts, but you're also both experts in networking. This is fundamentally the biggest thing. I guess my question for you and Hari too is what's been the biggest change in networking? Obviously, we see the KV Cache relevance that's now 3 years in. People are seeing that, but now you're looking at the edge in other areas, disaggregated inference. We see disaggregated infrastructure, networking with intelligence now has the classic intersection opportunity saying, oh, you got networks for AI and you got AI for networks. We've heard that a lot. What does that mean in practice?>> That's right. So, well, the first thing that has changed in networking is that we just see the speeds and feeds get bigger and bigger to support the type of throughput that you need and also the latency, right? Specifically, in AI clusters. So that's kind of at the foundation. You need switches of that caliber to support those use cases. So that's number one. But number two, what's really critical is that you want to optimize the performance of these networks. You want to be able to detect issues very quickly. You want to detect inefficiencies. You can have the load on the network be inefficiently spread out. And so for that, when you talk about AI for networking, it's the notion of specializing AI for the networking use case so that intuitively you have a capability that is always detecting issues, resolving them on the fly, right? And then alerting users when they need to be involved. For example, if they need to go change a transceiver or make a modification to their network. And so a lot of what we're doing here with Aria is taking AI and specializing it the same way for every other domain. You want to specialize AI to the domain. We specialize AI to the networking use case in order to deliver on this requirement. And we can dig into what that means, right? Because it is not just a matter of slapping an LLM on top of an already existing architecture, which a lot of times we see that, right?
John Furrier
>> You can't crawl the internet for that one because general intelligence, they crawl the internet, but you have a lot of domain expertise in networking. There's nuanced situations, there's configurations.>> Well, it starts with collecting telemetry, right? And there are two aspects to collecting telemetry. You want to collect it at the proper resolution. Typical networking software or equipment collects telemetry at the second or 30-second interval, right? This is what we did in the past and in past experiences. It was totally fine if your goal is just to validate configuration, right? But if you want to be at the speed of AI, optimizing the network for AI, 1 second is like a century in human scale, right? So many things could happen within a second. And so you need to be collecting telemetry at the microsecond resolution. So that's number 1, you will need it at the right resolution. And number 2, the network, and we can talk about all the different networks in an AI factory, but it spans very different domains. You have the front-end network, the back-end network. You not only have the switches, but you also have the network cards. You have networking happening on the host. You have the cables, the >> transceivers.You have new technologies like photonics coming to the mix.>> You have photonics coming to the mix.
John Furrier
>> All kinds of speeds and feeds. If you look at the NVL rack, Howie, it's the 72, half the rack is switches.
John Furrier
>> well, yeah.
John Furrier
>> it's a monster machine, but A bulk of the density are switches.>> Yeah. And if you have, for any of those cables, if you have a higher bit error rate, you can imagine what it does, right? So a lot of these queries when they come in, especially for inference, also for training, but for inference it's critical because the user is waiting. They get divided into many, many branches and then you have so many opportunities for tail latency to occur because you're waiting for the one that is gonna take the longest, right? And so if you have one of those thousands that waits, then everyone else is waiting.
John Furrier
>> So I'm going to call up my inner Andy Bechtolsheim and say, it's a physics problem.
John Furrier
>> Okay.
John Furrier
>> Because physics are involved here. So he's always a skeptic on physics. That'll never work. Or he leans in. So we're dealing with a lot of physical latency. You see photonics now coming in. You got copper. NVIDIA relies on copper. That's the big thing for them because they want to have that short distance. But Can you eliminate that? How many hops are in those switches? This is like old, this is networking stuff, but this would slow stuff down. You lose some heat and energy.>> That's right. And so you're going to have all these components and it gets more and more complex. But then what's critical is to have an ability to get telemetry to measure all of these components end to end at the right resolution. And that's step number one.>> Is the telemetry mostly to reduce the operational cost or is it more than that?>> Well, the telemetry is to see what's going on, right? Because if you don't have a picture of what's going on, then you never know if you're solving the right problem. You cannot root cause effectively or fix the problem effectively. So it starts by collecting telemetry, but then you want to extract knowledge from this telemetry, extract the right signal so that you know what's going on. And you need to do that at massive scale. And then from there, you need to have an ability to go in and make the right modifications. And that's where AI comes in. But again, AI first, there is that foundation of telemetry and then you have to be very sophisticated in how you're layering this AI in order for it to work for the networking use >> case.So what's the AI behind the scenes? Is it like the large language model from OpenAIs of the world plus fine-tuning or you have your own model on device, whatnot?>> Well, it's all of the above. So I'll tell you what you need to have. It's kind of think of the nervous system, right? In a human. If you put your finger on a flame, right? Your body will react. You don't even have to think about it. So that's one layer, right? That is the deterministic extreme quick reaction if something happens, that's happening at the lowest layers. Then there are multiple layers that go above where you're essentially taking on a bigger problem and you have an ability to reason about it at a larger scale. And you do that all the way to when you have to deploy sophisticated LLMs in the way you're reasoning about the data center as a whole, and then communicating that back and forth with the operator or maybe agents, because a lot of these softwares are going to be agentic. So we have agentic interfaces and we expect that other agents will be able to use the softwares. Basically, that's the world we're evolving towards.>> Yeah.>> So your shirt says networks that think, right? So that's the AI part. What about the speeds and feeds? You mentioned earlier, right? In the old days, it's more about speeds and feeds. Is that a solved problem or you're kind of leveraging AI to actually make the speeds and feeds a better deal for the customers? How do you think about it? >> Rightspeeds and feeds generally, it's a physics problem.>> So that's not a key problem.>> For us, we are a system company, so we're not building chips. But we are always looking at leveraging all the best chips, the best optics that exist for a given use case, right? So it's about how we put all of these things together. And then again, where AI comes in is that you may see problems or you may become aware of problems and AI can extract that information at the critical seconds before the problem occurs, right? Or in some cases many minutes before. Optics have a very specific pattern in the way they fail. You want to have an ability to go and detect that optic is about to fail before it fails. So somebody replaces it before it becomes an issue in production, right? So yes, there is a role for AI in terms of just managing the infrastructure, right? But at the end, there is the software and then there is the hardware. We do like to say that we co-design it all together, right? Because at the end, if you want to extract telemetry at the microsecond resolution, you have to be really deep into the hardware, right? In some cases we have agents that are sitting on ASICs, right? And basically embedded code that is extracting telemetry. So it's critical that both the software and the hardware work really well together. But there is a separation.
John Furrier
>> Remember the old days, there was a company called Thinking Machines. That's AI for everything that we're seeing embedded in. How's the company doing? Give a quick update on Aria Networks, how you guys are doing, the momentum, secret sauce. Obviously telemetry sounds like a super big part of it.>> Yeah, this deep networking layer, this co-designed hardware and software, including the telemetry and all of these agentic layers. That's part of the secret sauce. we're doing great. I haven't seen infrastructure evolve at the speeds that it has, right? So in every other company or project I was involved in, it took years before first PO. It took a very long time before we shipped anything into production.>> Scope the problem.>> We remember Big Switch days, right?
John Furrier
>> But there's a lot of engineering, but it's super fast now. What problem are you going after? Is it taking the latency out, energy, what's the focus?>> Right, so that's actually a great question. Our software, we've designed it in a way that is adaptive to whatever optimization you need. So it continuously and dynamically optimizes to your use case. And so it depends on the metric. So we discussed the fact that we're moving away from generic clusters towards more specialized clusters. And for every one of those clusters, you may have a different use case, a different metric. For example, if you're doing high throughput inference, then your metric is really tokens, dollars per million tokens, right? And so what you want is, these are simple inference use cases and you wanna be able to do them as many as possible, as cheaply as possible. Compare that to a high interactivity inference. There you have a lot more sophisticated inference job that requires a lot more complex computing and KV-Cache transfers. In some cases, you are separating the prefill from the decode. So it's a lot more complex. But then there, what's really critical is the number of tokens per second per user, because you have for every user, depending on how many tokens they get, per second, right? This is how fast the query is gonna get back to them or the response is gonna get back to them. And on the other hand, you may have a cluster that is the basic limit is the power, right? So there you wanna be able to extract as many tokens per watt as possible. So these are different clusters.>> So they're different. Totally situational. Based upon, for example, we think fast, as humans, the way we work, if you're in that stream, I might be using premium tokens, I might need premium resource. It sounds like policy, like the old days, policy networking. Well, I mean, it's policy.>> That's right.>> And it's happening.>> So it's in a sense, it's kind of, you have a different performance profile that you're optimizing towards.>> Right.>> And so, the job for the software is to adapt to these situations, right? And where AI gets super, it can get super intuitive, right? It's perfectly adapted to those types of use cases, right? And so it can adapt to the different use case intuitively.
John Furrier
>> So it's a networking concept, but you're combining all the old school techniques and thoughts like switch versus a router because you're doing a lot of things. The software is making the call, right? Is that right?>> Yeah, think of, It's an overused analogy, but it's like self-driving cars. You're not reinventing the wheels or the suspension, right? Or the fact that it has this many doors. It's a bit like that. So you mentioned InfiniBand. It's clear now that Ethernet is going to be the protocol of choice, right? And that has been the case for 30 years. So you can use the same exact management tools as before because The formats haven't changed. The standards haven't changed. So we're not changing that.
John Furrier
>> The rise of photonics has come up a lot in networking in certain use cases. Ethernet as the standard.>> Yeah, Ethernet as the standard. So the components themselves, other than them being a lot more, having more speeds and feeds, there are a lot of things that don't change. The protocols are pretty much the same, right? Whether it's routing or switching, But exactly to your point, right, the software layers that you're bringing in to essentially deliver on the requirement, have to leverage the latest and greatest >> technology.So in summary, in simple terms, you're injecting intelligence.
John Furrier
>> Totally.
John Furrier
>> To the networking.>> Yeah, exactly. We're leveraging the latest and greatest technologies, including obviously AI, including all this fine-grained telemetry, right? Putting it all together and bringing it to bear to deliver on this use case. Right?>> So several decades ago, right? Ethernet, TCP/IP won out over ATM, InfiniBand, that kind of technology because the alternative is too complex, right? Very hard to configure, make sense out of it. Are you saying that, look, AI will add sophistication to the network, the Networks that Think, but it's in the doing that, being done in the autonomous way, in the automagic way. So it's not, the sophistication without the complexity.>> Right, that's exactly right. But one thing I want to clarify is that I'm not a big fan of self-operating or self-driving networks because at the end—>> You still need a human first, AI assist.>> Networking, think of an autopilot in a plane, you can't get rid of the pilot, right? And I think it's the same in networking. When it's so critical, it needs to run. You don't want the network to self-drive off a cliff. You want the operator.>> And you got security concerns as well. We saw the Hugging Face stuff that's going out like, oops, that's a vector.>> 100%. And so this is where Networks that Think essentially think there is an operator or operators, but it helps them think, it helps them reason. And you're giving them all of the data. You're making them feel like it's at their fingertips in ways that were never possible before. And that is really the power of AI. It's kind of like how you use AI in a sense. It's not replacing me, or at least not yet, but it's empowering me incredibly. And I think it's the same here, adapted to the networking use case.
John Furrier
>> Well, soon you'll be able to hike the Dish in Palo Alto and talk to your machine that does all the coding for you, runs all your business. That's where AI is coming. We see some funny posts where people— how people work and think are different. There's no GUI anymore. There's a lot more—
John Furrier
>> That's exactly right.
John Furrier
>> Automation, a lot more intelligence.>> Yeah, so we don't have the static dashboard like you would have in typical management solutions for networking like we had in the past. You start with an interface. Well, first of all, you should assume that agents are using your product, but also as an operator, it's only the alerts that matter and then you're having the ability to ask it questions and then it will pull up whatever visuals are helpful to you in the task that you're performing at the time.>> So in your model, humans or the operators are still in the driver's seat. And then—>> Yes.>> But they have a lot more sophistication when it comes to telemetry, intelligence, >> visibility.Yeah, that's the idea. It's orders of magnitude more power in terms of their ability to see everything and then to control
John Furrier
>> everything.Well, Mansour, we're going to have to get some more time with you. Congratulations on—
John Furrier
>> Thank you. >> —Aria Networks. We love the intelligent networks, the edge. We haven't even talked about edge computing. There is a lot more we can talk about. This won't be the last. Thanks for coming on.>> Appreciate it.
John Furrier
>> Thank you, MansourThank you so much. I'm John Furrier, host of theCUBE. This is our AI Summit with theCUBE and NYSE Wired. Again, we are here in Palo Alto, full media day, a big event tonight, 180 practitioners and industry insiders and entrepreneurs talking about how the AI industry is going to be built out and how they're moving the needle and also what it's going to enable as the value capture comes around after the intelligence is injected into the network and the applications. We're doing our part here on theCUBE. Thanks for watching.