Mansour Karam, Aria Networks | theCUBE + NYSE Wired: Robotics & AI Infra Leaders
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for AI Factories - Data Centers of the Future
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AI Factories - Data Centers of the Future.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. If you don’t think you received an email check your
spam folder.
Sign in to AI Factories - Data Centers of the Future.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: AI Factories - Data Centers of the Future. Signing in with LinkedIn ensures a professional environment.
(INTRO)>> . Welcome back to theCUBE here in Palo Alto, California.>> I'm John Furrier, host of theCUBE. This is our third annual Poolside Party AI Summit all day here in our Palo Alto studio. My co -host, Howie, she was here breaking down all the action. We got an entrepreneur here, Mansour Karam, founder and CEO of Aria Networks. Great to see you again. We were just talking off camera about all the history of networking, and now we got the AI partners. Great to see you. We were sitting together in Paris for the RAISE Summit. In Versailles. Great event. The AI infrastructure has the biggest build out we've ever seen. Every piece of infrastructure is in demand. The demand curve is off the charts, and over the past three years, networking has emerged finally as the number one issue and technology that's driving everything. We look at all the AI factories, KVCache, we talk about InfiniBand, Prior Life, but now it's come full circle where if you don't have good networking, the AI really doesn't work very well.>> That's exactly right. This is where we're at. Well yeah, it's actually really a fun time if you're a networking guy like I am. I've been in networking all my career. it's this renaissance of sorts. But yes, AI factories are a brand new use case with very different requirements, in fact, stringent requirements. And as you said, the network is at the center of everything. And in fact, that's the impetus for why we started Aria Networks. If you think about it, the networking is maybe 10 % of the spend, depending on which cluster it is or which infrastructure, 5 % to 15%. Yet it touches every component. And so if the network is not working well, is not optimized, it affects every metric. It affects the performance of the cluster as a whole. And depending on the use case, the impact can be dramatic.
John Furrier
>> And there's dollars now, it says we have the GPUs are millions and millions of dollars that are not being fed the data. It's game over or disruption, loss of money, direct impact to the user experience. Everything kind of crumbles.>> Correct, it's not millions and millions, Billions, it's billions. Billions, billions. So that's not per GPU.>> But it's a system now, it's not just a server and a switch. The switches are all in the high density environment. So, it's interesting you're co-hosts, but you're also both experts in networking. This is fundamentally the biggest thing. I guess my question for you and for you too is, what's been the biggest change in networking? Obviously we see the KVCache relevance, that's now three years in, people have seen it. But now you're looking at the edge in other areas, is disaggregated inference we see, disaggregated infrastructure. Networking with intelligence now has the classic adjacency opportunity saying, oh you got networks for AI and you got AI for networks, we've heard that a lot. What does that mean in practice?>> That's right, so, the first thing that has changed in networking is that we just see the speeds and feeds get bigger and bigger to support the type of throughput that you need and also the latency, specifically in AI clusters. So that's kind of at the foundation. You need switches of that caliber to support those use cases. So that's number one. But number two, what's really critical is that you want to optimize the performance of these networks. You want to be able to detect issues very quickly. You want to detect inefficiencies. You can have the load on the network be inefficiently spread out. And so for that, when you talk about AI for networking, it's the notion of specializing AI for the networking use case, so that intuitively you have a capability that is always detecting issues, resolving them on the fly, right? And then alerting users when they need to be involved. For example, if they need to go change a transceiver or make a modification to their network. And so, a lot of what we're doing here with Aria is taking AI and specializing it the same way for every other domain. You want to specialize AI to the domain. You specialize AI to the networking use case in order to deliver on this requirement. And we can kind of dig into what that means, right? Because it is not just a matter of slapping an LLM on top of an already existing architecture, which a lot of times we see that, right?
John Furrier
>> They didn't crawl the internet book on that one because in general intelligence they crawl the internet. But you have a lot of domain expertise in networking, there's nuanced situations, there's configurations.>> Well it starts with collecting telemetry, right? And there are two aspects to collecting telemetry. You want to collect it at the proper resolution. Typical networking software or equipment collects telemetry at a second or 30-second interval. This is what we did in the past and in past experiences. it was totally fine if your goal is just to validate configuration, right? But if you want to be at the speed of AI, optimizing the network for AI, one second is like a century in human scale, right? So many things could happen within a second. And so you need to be collecting telemetry at the microsecond resolution. So that's number one. You need it at the right resolution. And number two, the network, and we can talk about all the different networks in an AI factory, but it spans very different domains. You have the front end network, the back end network. You not only have the switches, but you also have the network cards. You have networking happening on the host. You have the cables, the transceivers. And you have new technologies, like photonics coming into the mix. >> Yeah, photonics.All kinds of speeds and feeds. If you look at the NVL rack, Howie, it's the 72. Half the rack is switches. it's a monster machine, but the bulk of the density is switches.>> Yeah, and for any of those cables, if you have a higher bit error rate, you can imagine what it does, right? So a lot of these queries when they come in, especially for inference, also for training, but for inference it's critical because the user is waiting. They get divided into many, many branches, and then you have so many opportunities for tail latency to occur because you're waiting for the one that is going to take the longest, right? And so if you have one of those thousands that waits, then everyone else is waiting.
John Furrier
>> So I'm going to call up my inner Andy Bechtolsheim and say it's a physics problem. Okay, because physics is involved here. So he's always a skeptic on physics. That'll never work, or he just leans in. So we deal with a lot of physics. Latency, you see photonics now coming in. You got copper, NVIDIA relies on copper. That's the big thing for them because they want to have that short distance, but can you eliminate that? How many hops are in those switches? This is like old, this is networking stuff, but this would slow stuff down. You get to lose some heat and energy.>> That's right. And so you're going to have all these components, and it gets more and more complex. But then what's critical is to have an ability to get telemetry to measure all of these components end -to -end at the right resolution. And that's step number one.
John Furrier
>> Is the telemetry mostly to reduce the operational cost, or it's more than that?>> Well, the telemetry is to see what's going on, right? Because if you don't have a picture of what's going on, then you never know if you're solving the right problem. You cannot root cause effectively or fix the problem effectively. So it starts by collecting telemetry, but then you want to extract knowledge from the telemetry, extract the right signal, so that you know what's going on, and you need to do that at massive scale. And then from there, you need to have an ability to go in and make the right modifications. And that's where AI comes in. But again, AI, first, there is that foundation of telemetry, and then you have to be very sophisticated in how you're layering this AI in order for it to work for the networking use case. >> What's the AI behind the scenes?Is it like the large language models from the OpenAIs of the world, plus fine tuning, or you have your own model, on device, whatnot?>> Well, it's all of the above. I'll tell you what you need to have. It's kind of like, think of the nervous system in a human. If you put your finger on a flame, your body will react. You don't even have to think about it. So that's one layer. That is the deterministic, extreme quick reaction if something happens. That's happening at the lowest layers. Then there are multiple layers that go above where you're essentially taking on a bigger problem and you have an ability to reason about it at a larger scale. and you do that all the way to when you have to deploy sophisticated LLMs in the way you're reasoning about the data center as a whole and then communicating that back and forth with the operator or maybe agents because a lot of these softwares are going to be agentic. So we have agentic interfaces and we expect that other agents will be able to use the software. Basically, that's the world we're evolving towards. words.>> So your shirt says, Networks that Think, right? So that's the AI part. What about the speeds and feeds? You mentioned earlier, right? In the old days, it's more about speeds and feeds. Is that a solved problem or you're kind of leveraging AI to actually make the speeds and feeds, a better deal for the customers? How do you think about it? Right.
John Furrier
>> I mean, speeds and feeds, generally, it's a physics problem.
John Furrier
>> So that's not the key problem?
John Furrier
>> Well, for us, we are a system company, so we're not building chips. But we are always looking at leveraging all the best chips, the best optics that exist for a given use case, right? So it's about how we put all of these things together. And then again, where AI comes in is that, you may see problems, or you may become aware of problems, and AI can extract that information and, critically, seconds before the problem occurs, right? or in some cases, many minutes before. Optics have a very specific pattern in the way they fail. You want to have an ability to go and detect that an optic is about to fail before it fails, so somebody replaces it before it becomes an issue in production. So yes, there is a role for AI in terms of just managing the infrastructure, but at the end, there is the software and then there is the hardware. We do like to say that we co-design it all together, because at the end, if you want to extract telemetry at the microsecond resolution, you have to be really deep into the hardware. In some cases we have agents that are sitting on ASICs and basically embedded code that is extracting telemetry. So it's critical that both the software and the hardware work really well together. But there is a separation.>> Remember the old days, there was a company called Thinking Machines. That's AI for everything that we're seeing that embedded in. How's the company doing? Give a quick update on Aria Networks. How you guys are doing, the momentum, secret sauce. I see telemetry sounds like a super big part of it.
John Furrier
>> Yeah, this deep networking layer, right? This co-designed hardware and software, including the telemetry and all of these agentic layers. That's part of the secret sauce. we're doing great. Right, I haven't seen infrastructure evolve at the speeds that it has, right? So in every other company or project I was involved in, it took years before first PO. Yeah. It took a very long time before we shipped anything. and scoped the problem.
John Furrier
>> We remember Big Switch days.
John Furrier
>> Yeah, yeah. >> But there's a lot of engineering, there's a ton of engineering, but it's super fast now. What problem are you going after? Is it taking the latency out, energy, what's the focus?>> Right, so that's actually a great question. Our software, we've designed it in a way that is adaptive to whatever optimization you need. So it continuously and dynamically optimizes to your use case. And so, depends on the metric. So we discussed the fact that we're moving away from generic clusters towards more specialized clusters. And for every one of those clusters, you may have a different use case, a different metric. For example, if you're doing high throughput inference, then your metric is really dollars per million tokens. And so what you want is, these are simple inference use cases and you want to be able to do them as many as possible, as cheaply as possible. Compare that to a high interactivity inference. There you have a lot more sophisticated inference job that requires a lot more complex computing and KV cache transfers. In some cases, you're separating the pre -fill from the decode. So it's a lot more complex. But then there, what's really critical is the number of tokens per second per user because you have for every user, depending on how many tokens they get per second, this is how fast the query is going to get back to them or the response is going to get back to them. And so, on the other hand, you may have a cluster that the basic limit is the power, right? So there you want to be able to extract as many tokens per watt as possible. So these are different clusters. So they're different, they're situational. Totally situational.>> Based upon, for example, we think fast as humans, the way we work, if you're in that stream, I might be using premium tokens, I might need premium resource. It sounds like policy, like the old days, policy-based networking.>> It's policy. That's right. It's happening. So in a sense, you have a different performance profile that you're optimizing towards. And so the job for the software is to adapt to these situations, and where AI can get super intuitive. It's perfectly adapted to those types of use cases. And so it can adapt to the different use case intuitively.>> So it's a networking concept, but you're combining all the old school techniques and thoughts like a switch versus a router, because you're doing a lot of things. The software's making the call, right? Is that right?>> Yeah, think of it as, it's an overused analogy, but it's like self -driving cars. you're not reinventing the wheels or the suspension, right? Or the fact that it has this many doors. It's a bit like that. So you mentioned, InfiniBand, it's clear now that Ethernet is going to be the protocol of choice, right? And that has been the case for 30 years. So you can use the same exact management tools as before because the formats haven't changed, the standards haven't changed. So we're not changing that.>> The rise of photonics has come up a lot, and networking, certain use cases. Ethernet as a standard.>> Yeah, Ethernet as a standard. So the components themselves, other than them being a lot more, having more speeds and feeds, there are a lot of things that don't change. The protocols are pretty much the same, whether it's routing or switching. But exactly to your point, the software layers that you're bringing in to essentially deliver on the requirementhave to leverage the latest and greatest technologies.>> So you're, in summary, in simple terms, you're injecting intelligence to the networking.>> Yeah, exactly, we're leveraging the latest and greatest technologies, including obviously AI, including all this fine -grained telemetry, right? Putting it all together and bringing it to bear to deliver on this use case, right?>> So several decades ago, Ethernet, TCP/IP, won out, the ATM, the PMI; InfiniBand, the kind of the technology, because the alternative is too complex, right? You know, very hard to configure, make sense out of it. Are you saying that, look, AI will add sophistication to the network, the networks that think, but it isn't doing that, being there in the autonomous way, in the automagic way. So it's the sophistication without the complexity.
John Furrier
>> Right, that's exactly right. But one thing I want to clarify is that I'm not a big fan of self -operating or self -driving networks, because at the end...>> You still need a human first, AI assist.>> Networking, think of an autopilot in a plane. You can't get rid of the pilot, right? And I think it's the same in networking. When it's so critical, right, it needs to run. You don't want the network to self -drive off a cliff. You want the operator...>> And you got security concerns as well. Oh, you saw the Hugging Face stuff that's going on, ooh, oops, that's a vector.
John Furrier
>> 100%, and so this is where networks that think, essentially, there is an operator, there are operators, but it helps them think, it helps them reason, and you're giving them all of the data, you're making them feel like it's at their fingertips in ways that were never possible before. And that is really the power of AI. It's kind of like how you use AI. In a sense, it's not replacing me, or at least not yet, but it's empowering me incredibly, right? And I think it's the same here, adapted to the networking use case.>> Well, assuming you've had to hike the Dish in Palo Alto and talk to your machine, that does all the coding for you, runs all your business. that's where AI is coming. We see some funny posts where people, the work, how people work and think are different. there's no GUI anymore. There's a lot more automation, a lot more intelligence.>> Yeah, so we don't have the set dashboard like you would have in typical management solutions for networking, like we had in the past. You start with an interface, well first of all you should assume that agents are using your product, but also as an operator, it's only the alerts that matter, and then you're having the ability to ask it questions. And then it will pull up whatever visuals are helpful to you in the task that you're performing at the time.>> So in your model, humans are the operators that are still in the driver's seat and then... >> Yes. >> But they have a lot more sophistication when it comes to telemetry, intelligence, visibility.>> Yeah, that's the idea. It's orders of magnitude more power in terms of their ability to see everything and then to control everything.>> Well, Mansour, we're going to have to get some more time with you. Congratulations on Momentum. We love the intelligent networks. The Edge, we haven't even talked about The Edge coming. There's a lot more we can talk about. This won't be the last. Thanks for going on. Thank you, Mansour.
John Furrier
>> Thank you so much. >> I'm John Furrier, here on theCUBE? This is our AI Summit with theCUBE at NYSE Wired. Again, we are here in Palo Alto. Full media day, a big event tonight, 180 practitioners and industry insiders and entrepreneurs talking about how AI is just going to be built out and how they're moving the needle, and also what it's going to enable as the value capture comes around after the intelligence is injected into the network and the applications. We're doing our part here on theCUBE. Thanks for watching. Thank you.
(INTRO)>> . Welcome back to theCUBE here in Palo Alto, California.>> I'm John Furrier, host of theCUBE. This is our third annual Poolside Party AI Summit all day here in our Palo Alto studio. My co -host, Howie, she was here breaking down all the action. We got an entrepreneur here, Mansour Karam, founder and CEO of Aria Networks. Great to see you again. We were just talking off camera about all the history of networking, and now we got the AI partners. Great to see you. We were sitting together in Paris for the RAISE Summit. In Versailles. Great event. The AI infrastructure has the biggest build out we've ever seen. Every piece of infrastructure is in demand. The demand curve is off the charts, and over the past three years, networking has emerged finally as the number one issue and technology that's driving everything. We look at all the AI factories, KVCache, we talk about InfiniBand, Prior Life, but now it's come full circle where if you don't have good networking, the AI really doesn't work very well.>> That's exactly right. This is where we're at. Well yeah, it's actually really a fun time if you're a networking guy like I am. I've been in networking all my career. it's this renaissance of sorts. But yes, AI factories are a brand new use case with very different requirements, in fact, stringent requirements. And as you said, the network is at the center of everything. And in fact, that's the impetus for why we started Aria Networks. If you think about it, the networking is maybe 10 % of the spend, depending on which cluster it is or which infrastructure, 5 % to 15%. Yet it touches every component. And so if the network is not working well, is not optimized, it affects every metric. It affects the performance of the cluster as a whole. And depending on the use case, the impact can be dramatic.
John Furrier
>> And there's dollars now, it says we have the GPUs are millions and millions of dollars that are not being fed the data. It's game over or disruption, loss of money, direct impact to the user experience. Everything kind of crumbles.>> Correct, it's not millions and millions, Billions, it's billions. Billions, billions. So that's not per GPU.>> But it's a system now, it's not just a server and a switch. The switches are all in the high density environment. So, it's interesting you're co-hosts, but you're also both experts in networking. This is fundamentally the biggest thing. I guess my question for you and for you too is, what's been the biggest change in networking? Obviously we see the KVCache relevance, that's now three years in, people have seen it. But now you're looking at the edge in other areas, is disaggregated inference we see, disaggregated infrastructure. Networking with intelligence now has the classic adjacency opportunity saying, oh you got networks for AI and you got AI for networks, we've heard that a lot. What does that mean in practice?>> That's right, so, the first thing that has changed in networking is that we just see the speeds and feeds get bigger and bigger to support the type of throughput that you need and also the latency, specifically in AI clusters. So that's kind of at the foundation. You need switches of that caliber to support those use cases. So that's number one. But number two, what's really critical is that you want to optimize the performance of these networks. You want to be able to detect issues very quickly. You want to detect inefficiencies. You can have the load on the network be inefficiently spread out. And so for that, when you talk about AI for networking, it's the notion of specializing AI for the networking use case, so that intuitively you have a capability that is always detecting issues, resolving them on the fly, right? And then alerting users when they need to be involved. For example, if they need to go change a transceiver or make a modification to their network. And so, a lot of what we're doing here with Aria is taking AI and specializing it the same way for every other domain. You want to specialize AI to the domain. You specialize AI to the networking use case in order to deliver on this requirement. And we can kind of dig into what that means, right? Because it is not just a matter of slapping an LLM on top of an already existing architecture, which a lot of times we see that, right?
John Furrier
>> They didn't crawl the internet book on that one because in general intelligence they crawl the internet. But you have a lot of domain expertise in networking, there's nuanced situations, there's configurations.>> Well it starts with collecting telemetry, right? And there are two aspects to collecting telemetry. You want to collect it at the proper resolution. Typical networking software or equipment collects telemetry at a second or 30-second interval. This is what we did in the past and in past experiences. it was totally fine if your goal is just to validate configuration, right? But if you want to be at the speed of AI, optimizing the network for AI, one second is like a century in human scale, right? So many things could happen within a second. And so you need to be collecting telemetry at the microsecond resolution. So that's number one. You need it at the right resolution. And number two, the network, and we can talk about all the different networks in an AI factory, but it spans very different domains. You have the front end network, the back end network. You not only have the switches, but you also have the network cards. You have networking happening on the host. You have the cables, the transceivers. And you have new technologies, like photonics coming into the mix. >> Yeah, photonics.All kinds of speeds and feeds. If you look at the NVL rack, Howie, it's the 72. Half the rack is switches. it's a monster machine, but the bulk of the density is switches.>> Yeah, and for any of those cables, if you have a higher bit error rate, you can imagine what it does, right? So a lot of these queries when they come in, especially for inference, also for training, but for inference it's critical because the user is waiting. They get divided into many, many branches, and then you have so many opportunities for tail latency to occur because you're waiting for the one that is going to take the longest, right? And so if you have one of those thousands that waits, then everyone else is waiting.
John Furrier
>> So I'm going to call up my inner Andy Bechtolsheim and say it's a physics problem. Okay, because physics is involved here. So he's always a skeptic on physics. That'll never work, or he just leans in. So we deal with a lot of physics. Latency, you see photonics now coming in. You got copper, NVIDIA relies on copper. That's the big thing for them because they want to have that short distance, but can you eliminate that? How many hops are in those switches? This is like old, this is networking stuff, but this would slow stuff down. You get to lose some heat and energy.>> That's right. And so you're going to have all these components, and it gets more and more complex. But then what's critical is to have an ability to get telemetry to measure all of these components end -to -end at the right resolution. And that's step number one.
John Furrier
>> Is the telemetry mostly to reduce the operational cost, or it's more than that?>> Well, the telemetry is to see what's going on, right? Because if you don't have a picture of what's going on, then you never know if you're solving the right problem. You cannot root cause effectively or fix the problem effectively. So it starts by collecting telemetry, but then you want to extract knowledge from the telemetry, extract the right signal, so that you know what's going on, and you need to do that at massive scale. And then from there, you need to have an ability to go in and make the right modifications. And that's where AI comes in. But again, AI, first, there is that foundation of telemetry, and then you have to be very sophisticated in how you're layering this AI in order for it to work for the networking use case. >> What's the AI behind the scenes?Is it like the large language models from the OpenAIs of the world, plus fine tuning, or you have your own model, on device, whatnot?>> Well, it's all of the above. I'll tell you what you need to have. It's kind of like, think of the nervous system in a human. If you put your finger on a flame, your body will react. You don't even have to think about it. So that's one layer. That is the deterministic, extreme quick reaction if something happens. That's happening at the lowest layers. Then there are multiple layers that go above where you're essentially taking on a bigger problem and you have an ability to reason about it at a larger scale. and you do that all the way to when you have to deploy sophisticated LLMs in the way you're reasoning about the data center as a whole and then communicating that back and forth with the operator or maybe agents because a lot of these softwares are going to be agentic. So we have agentic interfaces and we expect that other agents will be able to use the software. Basically, that's the world we're evolving towards. words.>> So your shirt says, Networks that Think, right? So that's the AI part. What about the speeds and feeds? You mentioned earlier, right? In the old days, it's more about speeds and feeds. Is that a solved problem or you're kind of leveraging AI to actually make the speeds and feeds, a better deal for the customers? How do you think about it? Right.
John Furrier
>> I mean, speeds and feeds, generally, it's a physics problem.
John Furrier
>> So that's not the key problem?
John Furrier
>> Well, for us, we are a system company, so we're not building chips. But we are always looking at leveraging all the best chips, the best optics that exist for a given use case, right? So it's about how we put all of these things together. And then again, where AI comes in is that, you may see problems, or you may become aware of problems, and AI can extract that information and, critically, seconds before the problem occurs, right? or in some cases, many minutes before. Optics have a very specific pattern in the way they fail. You want to have an ability to go and detect that an optic is about to fail before it fails, so somebody replaces it before it becomes an issue in production. So yes, there is a role for AI in terms of just managing the infrastructure, but at the end, there is the software and then there is the hardware. We do like to say that we co-design it all together, because at the end, if you want to extract telemetry at the microsecond resolution, you have to be really deep into the hardware. In some cases we have agents that are sitting on ASICs and basically embedded code that is extracting telemetry. So it's critical that both the software and the hardware work really well together. But there is a separation.>> Remember the old days, there was a company called Thinking Machines. That's AI for everything that we're seeing that embedded in. How's the company doing? Give a quick update on Aria Networks. How you guys are doing, the momentum, secret sauce. I see telemetry sounds like a super big part of it.
John Furrier
>> Yeah, this deep networking layer, right? This co-designed hardware and software, including the telemetry and all of these agentic layers. That's part of the secret sauce. we're doing great. Right, I haven't seen infrastructure evolve at the speeds that it has, right? So in every other company or project I was involved in, it took years before first PO. Yeah. It took a very long time before we shipped anything. and scoped the problem.
John Furrier
>> We remember Big Switch days.
John Furrier
>> Yeah, yeah. >> But there's a lot of engineering, there's a ton of engineering, but it's super fast now. What problem are you going after? Is it taking the latency out, energy, what's the focus?>> Right, so that's actually a great question. Our software, we've designed it in a way that is adaptive to whatever optimization you need. So it continuously and dynamically optimizes to your use case. And so, depends on the metric. So we discussed the fact that we're moving away from generic clusters towards more specialized clusters. And for every one of those clusters, you may have a different use case, a different metric. For example, if you're doing high throughput inference, then your metric is really dollars per million tokens. And so what you want is, these are simple inference use cases and you want to be able to do them as many as possible, as cheaply as possible. Compare that to a high interactivity inference. There you have a lot more sophisticated inference job that requires a lot more complex computing and KV cache transfers. In some cases, you're separating the pre -fill from the decode. So it's a lot more complex. But then there, what's really critical is the number of tokens per second per user because you have for every user, depending on how many tokens they get per second, this is how fast the query is going to get back to them or the response is going to get back to them. And so, on the other hand, you may have a cluster that the basic limit is the power, right? So there you want to be able to extract as many tokens per watt as possible. So these are different clusters. So they're different, they're situational. Totally situational.>> Based upon, for example, we think fast as humans, the way we work, if you're in that stream, I might be using premium tokens, I might need premium resource. It sounds like policy, like the old days, policy-based networking.>> It's policy. That's right. It's happening. So in a sense, you have a different performance profile that you're optimizing towards. And so the job for the software is to adapt to these situations, and where AI can get super intuitive. It's perfectly adapted to those types of use cases. And so it can adapt to the different use case intuitively.>> So it's a networking concept, but you're combining all the old school techniques and thoughts like a switch versus a router, because you're doing a lot of things. The software's making the call, right? Is that right?>> Yeah, think of it as, it's an overused analogy, but it's like self -driving cars. you're not reinventing the wheels or the suspension, right? Or the fact that it has this many doors. It's a bit like that. So you mentioned, InfiniBand, it's clear now that Ethernet is going to be the protocol of choice, right? And that has been the case for 30 years. So you can use the same exact management tools as before because the formats haven't changed, the standards haven't changed. So we're not changing that.>> The rise of photonics has come up a lot, and networking, certain use cases. Ethernet as a standard.>> Yeah, Ethernet as a standard. So the components themselves, other than them being a lot more, having more speeds and feeds, there are a lot of things that don't change. The protocols are pretty much the same, whether it's routing or switching. But exactly to your point, the software layers that you're bringing in to essentially deliver on the requirementhave to leverage the latest and greatest technologies.>> So you're, in summary, in simple terms, you're injecting intelligence to the networking.>> Yeah, exactly, we're leveraging the latest and greatest technologies, including obviously AI, including all this fine -grained telemetry, right? Putting it all together and bringing it to bear to deliver on this use case, right?>> So several decades ago, Ethernet, TCP/IP, won out, the ATM, the PMI; InfiniBand, the kind of the technology, because the alternative is too complex, right? You know, very hard to configure, make sense out of it. Are you saying that, look, AI will add sophistication to the network, the networks that think, but it isn't doing that, being there in the autonomous way, in the automagic way. So it's the sophistication without the complexity.
John Furrier
>> Right, that's exactly right. But one thing I want to clarify is that I'm not a big fan of self -operating or self -driving networks, because at the end...>> You still need a human first, AI assist.>> Networking, think of an autopilot in a plane. You can't get rid of the pilot, right? And I think it's the same in networking. When it's so critical, right, it needs to run. You don't want the network to self -drive off a cliff. You want the operator...>> And you got security concerns as well. Oh, you saw the Hugging Face stuff that's going on, ooh, oops, that's a vector.
John Furrier
>> 100%, and so this is where networks that think, essentially, there is an operator, there are operators, but it helps them think, it helps them reason, and you're giving them all of the data, you're making them feel like it's at their fingertips in ways that were never possible before. And that is really the power of AI. It's kind of like how you use AI. In a sense, it's not replacing me, or at least not yet, but it's empowering me incredibly, right? And I think it's the same here, adapted to the networking use case.>> Well, assuming you've had to hike the Dish in Palo Alto and talk to your machine, that does all the coding for you, runs all your business. that's where AI is coming. We see some funny posts where people, the work, how people work and think are different. there's no GUI anymore. There's a lot more automation, a lot more intelligence.>> Yeah, so we don't have the set dashboard like you would have in typical management solutions for networking, like we had in the past. You start with an interface, well first of all you should assume that agents are using your product, but also as an operator, it's only the alerts that matter, and then you're having the ability to ask it questions. And then it will pull up whatever visuals are helpful to you in the task that you're performing at the time.>> So in your model, humans are the operators that are still in the driver's seat and then... >> Yes. >> But they have a lot more sophistication when it comes to telemetry, intelligence, visibility.>> Yeah, that's the idea. It's orders of magnitude more power in terms of their ability to see everything and then to control everything.>> Well, Mansour, we're going to have to get some more time with you. Congratulations on Momentum. We love the intelligent networks. The Edge, we haven't even talked about The Edge coming. There's a lot more we can talk about. This won't be the last. Thanks for going on. Thank you, Mansour.
John Furrier
>> Thank you so much. >> I'm John Furrier, here on theCUBE? This is our AI Summit with theCUBE at NYSE Wired. Again, we are here in Palo Alto. Full media day, a big event tonight, 180 practitioners and industry insiders and entrepreneurs talking about how AI is just going to be built out and how they're moving the needle, and also what it's going to enable as the value capture comes around after the intelligence is injected into the network and the applications. We're doing our part here on theCUBE. Thanks for watching. Thank you.