We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For SC24
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for SC24.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to SC24
Please sign in with LinkedIn to continue to SC24. Signing in with LinkedIn ensures a professional environment.
Head of HPC and AI Infrastructure Product MarketingNVIDIA
The panelists discuss the WARRP reference architecture created by WEKA for AI inferencing using Run:ai and NVIDIA's software frameworks. The goal is to help customers understand and implement AI in production environments. They also touch on the importance of deploying AI models in real-time for maximum impact. The discussion highlights the evolution of AI adoption in enterprises and the role of open-source models in controlling data and reducing costs. The panelists stress the need for efficient GPU utilization as AI models become more prominent. They also t...Read more
exploreKeep Exploring
What is the WARRP and how can it help customers struggling with RAG inferencing?add
What were some of the use cases discussed at climate week that were being powered by AI?add
What are some considerations to keep in mind when deploying AI models in production?add
What are the speaker's hopes for the future of enterprise SuperComputing and AI implementation?add
>> Good morning, nerd fam, and welcome back to Atlanta, Georgia. We're here at SuperComputing at 2024, just kicking off day one of our three days of coverage on theCUBE. My name's Savannah Peterson. Delighted to have this power packed panel. My goodness. Dion, Shimone, and Ronan. Thank you so much for taking the time this morning. It's a busy day, it's already buzzing.
Dion Harris
>> Ah, man.
Savannah Peterson
>> You guys are smiling. I'm loving the vibes. Are you having a good time so far?
Shimon Ben-David
>> Amazing.
Dion Harris
>> It's been a great start so far. I mean, SuperComputing is always an exciting show and-
Savannah Peterson
>> It is....
Dion Harris
>> it is not just disappointing this year. Incredible stuff happening.
Savannah Peterson
>> Yes, it is an exciting show. I mean, I know we were talking about how we're all nerds over here before we went live. And I feel like this is one of the great shows for nerds. And we all get to nerd out with our nerdy friends, and there's some really exciting announcements. Speaking of, Shimone, big announcement out of WEKA this morning.
Shimon Ben-David
>> Yeah.
Savannah Peterson
>> You want to tell us a little bit about the WARRP?
Shimon Ben-David
>> Exactly. Yeah. I'm glad that we can finally talk about it. It's been in the making for a long time. And finally, we can tell the story.
Savannah Peterson
>> Yes.
Shimon Ben-David
>> So, we created this reference architecture called WARRP, WEKA AI RAG Reference Platform. What we're seeing is that, as you know, WEKA is a data platform, a high-performance data platform, but we were just talking about it. We are seeing customers still struggling with how to implement RAG inferencing. It has a lot of moving components. Honestly, there's no real blueprint or protocols defined yet for that. And we created this environment that shows all of the layers that are needed. We're actually heavily using Run:ai and then video stack also, the GPUs, but also the software frameworks, NEMS, NeMo, and more. And we built this reference architecture that the customer can now use to first of all, learn and see how is an AI reference platform, how is an AI in fencing environment looks like in production. There's a big difference between a POC environment and a production environment. So, how does it look in production as it scales out and in cross to multiple clouds and back? And then obviously show the WEKA value in that.
Dion Harris
>> Right.
Savannah Peterson
>> Which is super exciting. Congratulations, by the way. I know this just came out this morning, so very much breaking news. What does that mean for partners like you?
Dion Harris
>> Well, again, if I take a step back, first, I like to introduce terms so that I know the audience understands.
Savannah Peterson
>> Everyone likes to introduce the terms that trains our a AI too.
Dion Harris
>> RAG is retrieval augment degeneration. And it's the way that you can customize foundational models to basically incorporate proprietary data, or data that you care about, that you want represented in your AI models. So basically as it relates to NVIDIA, we've been down this path of trying to be a proponent of AI and help customers adopt AI. And so I think by providing more guidelines, blueprints, templates, APIs that make it easy to plug and play and leverage these tools, that's really what we're trying to do here at the show, and just in general. So working, obviously with WEKA and Run:ai, it's a great example of doing exactly just that.
Savannah Peterson
>> What I'm hearing from both of you that I think is really awesome, is not only are we learning together 'cause we all are on this adoption curve, quite at an accelerated rate. But we're teaching together, is what I just heard from you.
Dion Harris
>> Yeah, absolutely. Absolutely.
Shimon Ben-David
>> Absolutely.
Savannah Peterson
>> Which is pretty cool.
Shimon Ben-David
>> Absolutely.
Savannah Peterson
>> How do you feel-
Ronen Dar
>> I think it feels-
Savannah Peterson
>> Yeah, go for it.
Ronen Dar
>> It reminds me what happened in the training market, right? We're speaking here about inference, but four or five years ago, we had in a similar situation for training workloads. Right?
Shimon Ben-David
>> Exactly.
Ronen Dar
>> AI, how are you training AI models, deep learning models? How do you build your AI infrastructure? How do you build your software stack on top of it? So, there were no best practices. People were learning that. So, I think we are in a similar situation right now for the inference market. And then blueprints like wow, and reference architecture really are a good move, good to progressing the entire community forward in the right direction. So, that's amazing.
Savannah Peterson
>> That was an endorsement and a half, if I ever heard it. It's exciting. I just heard wow too. that sounds. How long have y'all been working together, actually? Now I'm curious about this.
Shimon Ben-David
>> So, Run:ai and WEKA have been working for several years already.
Ronen Dar
>> Yeah, six, seven years.
Shimon Ben-David
>> And at least from the WEKA side and NVIDIA from day one.
Dion Harris
>> Yeah, yeah, yeah. So we've been involved with a lot of these guys early on, because like I said, NVIDIA is a platform company. And so, we want to really work with our extended ecosystem to make sure that as they empower their customers, we can work together to basically make sure these solutions are easy to adopt and they're ready to be deployed. And so, a lot of solutions that they bring to market perfectly complement that.
Ronen Dar
>> Yeah, exactly. And we at Run:ai, we're working together with NVIDIA for sure, with WEKA for six, seven years, but with NVIDIA, I think for four or five years. So we started early, we started I think around DJXs, right? NVIDIA started to sell DJXs, right? It's like a new offering in the market. And NVIDIA wanted to build the ecosystem, the ISV software vendors around the DJX. And we were part of that ecosystem, so the relationship with NVIDIA started there. And it was from a business perspective, and it moves forward in the last four, five years. It's amazing, amazing partnership that we have with NVIDIA.
Savannah Peterson
>> So, y'all knew we were going to be hitting this hype curve moment right now then, is what I'm hearing.
Ronen Dar
>> No doubt.
Dion Harris
>> We saw it coming a long time.
Savannah Peterson
>> You've been preparing, just waiting for the rest of us to catch up.
Dion Harris
>> So yo know what's funny, I would say maybe 2018, for example, I remember Jensen was doing this special address. And he basically had this pie chart that said Old HPC, the new HPC. And he was describing this new sort of workload mix. 45% was going to be AI. And I never forget it. I remember he showed the chart and then you could hear the audience, it was just like, nah, I don't think that's going to happen. But it's so interesting to see if you fast-forward in the last four or five years, if you walk the show floor, there's so much discussion about AI. It's having so much .
Savannah Peterson
>> I haven't heard that acronym all today.
Dion Harris
>> I mean, come on. Exactly. Who's heard of AI? . But there's so much transformative impact that it's having in science. So I mean, we're here at HPC at SuperComputing, so it's having an incredible impact on science. But now also enterprises and other organizations are adopting it, and so these types of solutions help simplify and streamline that process. So, very cool.
Savannah Peterson
>> Yeah, go for it, Ronan.
Ronen Dar
>> I think the vision was there since 2006, right?
Dion Harris
>> Yeah.
Ronen Dar
>> 2006. And data started to build CUDA to enable more developers to build applications on top of GPUs, right? And they aimed for scientific computing. So the vision was there, scientific computing, and turned out 10 years afterwards, that AI is a big scientific computing application that's going to change the world. And so, the vision was-
Savannah Peterson
>> Going to change the world, hopefully save the planet. There's so many things.
Dion Harris
>> Absolutely.
Savannah Peterson
>> Not to be casual about it, but I feel like...
Dion Harris
>> No, no there's a lot of interesting discussion right now. And so, I just came back from climate week a couple of weeks ago. And it was really fascinating to see all the incredible use cases from climate and weather forecasting, to grid management optimization, to renewables research. I mean, it was all these incredible use cases that were all being powered by AI. And so, it's really cool to be at the center of that and see all the great work happening, for sure.
Shimon Ben-David
>> I can tell you that by the way, that from our side, we're seeing that we're doing AI for many years. So, I installing these large-scale AI scientific projects and some production large-scale productions. And it was almost like it was a niche a few years ago. The vision was there, but it was almost like reserved for a group of people. They knew what CUDA is, the is.
Savannah Peterson
>> Such a great point. I'm really glad you're bringing this up. Yes.
Shimon Ben-David
>> It was very specialized. And what we're seeing that's very different in the last two, two and a half years is now it's going into the enterprise. My grandmother can now talk AI with me. My little daughter is arguing with her AI phone. So when it gets to that level of the enterprises, that's where technology really booms. And that's where everybody can benefit from it.
Dion Harris
>> Yeah.
Savannah Peterson
>> I totally agree with you there. I feel like as nerds, we're kind of having a moment 'cause now everybody knows what we do, or at least the industry we work in. And I think the scientific... I'm glad we brought up scientific applications and Climate Week. I'd love to chat with you more about that, as can make it. I think it really helps make AI real for a lot of people-
Dion Harris
>> For sure, for sure....
Savannah Peterson
>> because it's not just this hypothetical thing that we're all doing in our little nerd lands. It's very much this.
Shimon Ben-David
>> So the barrier to entry actually, is much lower. In the past there are too, to know what I'm doing, all of the ML provisioning, all of the training of the models. Now, I just go and download models to my laptop even, and I'm running it.
Savannah Peterson
>> Which is pretty wild.
Dion Harris
>> And it kind of ties back to the announcement today because NVIDIA has been sort of championing this new API-driven model, which is, how do you make this drag and drop easy for people to implement? And so, we've implemented something called NVIDIA Inference Microservices, which are NIMS. And so, those are basically containerized foundational models that customers can download, fine-tune, tweak, use RAG workflows to help make them their own, and make them easily adoptable. And so what's really cool about, like I said, when you look at the suite of products, we have something called NIMO Retriever, which is basically a RAG workflow automation technique. And then there's also something called cuVS, which is our CUDA Vector Search Optimization Library. And so it accelerates a lot of the vector databases like Milvus, which is incorporated in this new workflow that we're highlighting through this announcement. So, it's really incredible to see NVIDIA, we see have all these different tools and libraries and acceleration techniques. So, it's great to see them all being adopted and worked into actual production level workflows.
Savannah Peterson
>> It's really exciting.
Dion Harris
>> Yeah.
Shimon Ben-David
>> It is.
Savannah Peterson
>> AI workloads are huge, but if we're able to leverage models like that or download things or even, I mean, shoot, you can even play with stuff in your browser at this point, which is pretty crazy to think about out loud. How does today's announcement and your collective collaboration going to impact user experience?
Shimon Ben-David
>> So I think what we aim to do... Sorry, .
Ronen Dar
>> Go ahead.
Shimon Ben-David
>> What we aim to do is, we aim to simplify. We wanted to show to educate. I think that was the first topic that we touched on. How is an AI RAG inferencing environment, how does it look like? What are the components, what are the layers? And then we wanted to make sure that the customers have an easy blueprint to start their journey on. And this blueprint is composed out of multiple layers. And we each participate in each of these layers, contributing and actually amplifying the other's capabilities.
Ronen Dar
>> Right. Exactly. And I think now enterprises understand that AI is a transformative technology. It's out there, the value, the potential for value is there, right? It's now about just getting to that value. And I think it's about experimentation right now, experimenting with LLMs and experimenting with closed-source LLMs. But we also see big demand right now for the open-source, LLMs, right?
Shimon Ben-David
>> Huge.
Ronen Dar
>> Putting open-source, LLMs, open-weight LLMs on infrastructure, on enterprises' own GPUs for multiple reasons. They want to control their data, they want to control their IP, they want controlled costs. So then open-source models provide a path forward in that sense. But then it becomes difficult, then it becomes difficult. How do I deploy those models on my own infrastructure, right? NIMS come to up. How do I orchestrate it? How do I scale it? Scaling it efficiently becomes a big, big challenge. And people are speaking about GPU utilization. When you scale your application, GPU utilization, the cost of those LLMs, the cost to serve those LLMs becomes a real problem. So GPU utilization, as we all move forward and LLMs become more and more important, GPU utilization will become more and more important. Just increasing GPU utilization and reducing the cost of serving LLMs.
Shimon Ben-David
>> Yeah. And what we found, by the way, when we went through that journey of describing WARRP, creating WARRP, building it, we saw that obviously as I mentioned, there's a lot of moving parts, a lot of frameworks, orchestration, data challenges, whether you are scaling or not. And not all of them are actually the GPUs. So we are hitting some, we measure our efficiency by times to token, cost per token, token throughput. There's out of token economics, customers,
Dion Harris
>> Tokenomics is what we call it.
Savannah Peterson
>> We're in a new era of tokenomics. We had our blockchain era, now we're in this era of tokenomics.
Shimon Ben-David
>> Exactly. It's the new blockchain ear.
Savannah Peterson
>> Oh, no, totally is.
Shimon Ben-David
>> Now we need to explain to people, what is a token?
Ronen Dar
>> What is a token? Yeah.
Shimon Ben-David
>> But enterprises are now in the business of measuring their ROI on their AI investment. So the more tokens-
Savannah Peterson
>> Which is not a cheap investment in most cases.
Shimon Ben-David
>> The more efficient you can be, the better. And what we saw is that constantly we're scaling the environment when we built WARRP. We're scaling the environment in terms of GPUs and then with the bottleneck. And then the bottleneck was not actually the GPUs, it was our chain server that was actually bombarded. So, we're using Run:ai to just automatically scale to another chain. Suddenly, the GPUs are on bottlenecks. So there's all of these scale out, scale up, scale in considerations that eventually allows you to get the token economics much better.
Dion Harris
>> Yeah, yeah. And I think we're in this sort of evolution. We talked about it earlier around training was the first big hurdle, I should say, where people were trying to understand how they can build and train models. And they use foundational models. They said, "Okay, well that's not specifically to what I need." Then they started saying, "Okay, I'll kind of use RAG or fine-tuning to customize the models." And now we're at a point, I think he was alluding to earlier, they're deploying these in production. And so, you go from what we call AI boundary to building developing models, to the AI factory where you're building and producing tokens. And that has a whole different set of considerations, utilization, productivity, making sure that you're being able to use Kubernetes at scale effectively and efficiently. So, it really introduces a new set of considerations when people are deploying these models in production. And that's where we are today. And that's the exciting part about it, because the point of AI isn't to train models.
Savannah Peterson
>> No, I know.
Dion Harris
>> That's a means to an end.
Savannah Peterson
>> And there's so much emphasis on that.
Dion Harris
>> .
Savannah Peterson
>> And it's like, guys, wait, just let it out there in the wild.
Dion Harris
>> Exactly. The means is really to get it in inferences and to get it deployed in production where you can really make these real-time impacts.
Savannah Peterson
>> Well, an inference is what makes AI real-time.
Dion Harris
>> Absolutely.
Savannah Peterson
>> When it comes to realizing some of these-
Shimon Ben-David
>> It's the value.
Savannah Peterson
>> Exactly.
Shimon Ben-David
>> Everything else is just plumbing.
Savannah Peterson
>> All this big bloated stuff doesn't matter. It's like running on a treadmill, but never actually getting out there and doing the marathon. I'm glad you brought that up. I'm one of the bigger inference nerds on the team as well. But I think that makes the biggest difference. That's what's going to make it for your grandma or for your family. That's what's going to make all of this real. We're playing with it like it's a toy right now, and in certain cases in certain applications, but I think that's really what's going to make that big broad impact. Okay. Well, this is fun. Now that we're here and we're in this part of the pathway, what do you hope your collaborative efforts, but even our industry as a whole, what challenges do you hope we solve collectively with all this new tech?
Dion Harris
>> Yeah, I mean, I think quite simply, is to help ease that transition to production like we've been describing. If everyone is sort of crossing that chasm to adopt AI, look for new use cases, look for opportunities to get efficiencies, to get new revenue opportunities, we want to help make that transition as seamless and as smooth as possible.
Savannah Peterson
>> Seamless and as smooth as possible.
Shimon Ben-David
>> That's really what I'm hoping for.
Savannah Peterson
>> Yes, please.
Shimon Ben-David
>> Big vision, obviously, as I mentioned, everything is a plumbing. Even when I show WARRP to customers, prospects, partners, we're showing all of the layers of WARRP. Everything under the application is plumbing. I'm sorry, we're all plumbers.
Savannah Peterson
>> Yeah.
Dion Harris
>> No. Yeah, yeah. That's right. No, no, you're right.
Savannah Peterson
>> Incredibly talented plumbers.
Dion Harris
>> Right, right, right. Exactly. We all need plumbers. .
Shimon Ben-David
>> But it's how efficient can you be? But then what challenges do we solve? So, you mentioned climate. Climate is huge for us, and we know that for NVIDIA as well, if we can actually model and solve climate changes using that, we already are seeing a lot of research in life science. So there's already... It's funny, there's a lot of GPUs acceleration in life science, which is not AI. Cryo-EM, genomic sequencing, digital pathology. But then we're starting to see more and more AI in life science as well. So, cancer research, drug research, and more.
Savannah Peterson
>> Detection.
Shimon Ben-David
>> So, these are the big things.
Savannah Peterson
>> There's so many exciting things.
Shimon Ben-David
>> And I'll tell you, my wife's private dream, a system that will just clean the house. So, kind of like the robot the Jetson's-
Dion Harris
>> Exactly. What's in it do for me? Exactly.
Savannah Peterson
>> That's what she wants to-
Ronen Dar
>> But I think we're seeing already, we're seeing already value from AI. It's happening, right?
Dion Harris
>> Absolutely.
Ronen Dar
>> It's happening. It's going to change everything, probably, right? That's what will happen. But we see already the software industry is being changed right now. Every software tool right now needs to have AI. Developers right now are coding with Copilot, with , with AI tools that are programming for themselves. So, I think the software industry is one of the first industries that are actually being changed right now because of AI.
Dion Harris
>> Absolutely.
Ronen Dar
>> So, that's interesting.
Savannah Peterson
>> It's such an exciting time. Wow. Okay. We are tearing through time here. I have one last question for you, since you're all fabulous guests, and I'm sure we'll have you on at SuperComputing next year. And I want you each to answer this so you can thumb wrestle over who goes first. What do you hope to say when we're sitting down in St. Louis next year that you can't yet say today? I mean, there's been a pretty big change from the last time we were here to now. So, what do you think our next little leap is going to be?
Dion Harris
>> So I'll do my own selfish plug. I hope that we can really celebrate the work that is happening. Let me be clear first. This SuperComputing, I see Grace Hopper, I see Hopper making an incredible splash and incredible contribution to all these different research use cases, the Gordon Bell prizes. They had some incredible work done on Grace Hopper. So, that's been a really-
Savannah Peterson
>> Glad you brought that up....
Dion Harris
>> proud moment here this year.
Savannah Peterson
>> Yeah.
Dion Harris
>> Next year at SuperComputing, I hope we're having some of these same conversations about Blackwell, particularly some of these inference abuse cases. If you walk the show floor, you see a lot of Blackwell NVL72, Gb200 on the floor. So I think as we start to hit that production ramp, we'll see systems coming online, we'll see researchers leveraging that capability. And there'll be a lot to celebrate there as well.
Savannah Peterson
>> I think you are right about that. What about you two?
Dion Harris
>> Fingers crossed.
Ronen Dar
>> So actually, for me, it's my first supercomputer. Funny enough-
Savannah Peterson
>> Welcome to the fam.
Ronen Dar
>> Yeah, yeah.
Dion Harris
>> Welcome. Yeah.
Savannah Peterson
>> I'm surprised to hear that, but also kind of makes sense from a timeliness perspective.
Ronen Dar
>> Yeah. We are trying AI, we've been here for-
Shimon Ben-David
>> ....
Ronen Dar
>> three, four years.
Dion Harris
>> Yeah, yeah. Exactly.
Ronen Dar
>> Yeah. But for me, it's the first time. So actually, the energies here are amazing, right? You can feel the energies, you can feel exciting stuff are happening with AI. So I think next year, let's see more energies for us. I'm willing to see more and more AI applications changing the world. So, waiting to see that.
Savannah Peterson
>> Yeah.
Shimon Ben-David
>> I'll give you my 2 cents.
Savannah Peterson
>> Let's go for it.
Shimon Ben-David
>> So, I'm here for a few years already. And again, I think my first SuperComputing was really HPC-oriented, everything large SuperComputing centers, simulations. Slowly, it started being into some of the AI conversation. This is I think the first or second really AI-centric SuperComputing. What I hope to see next year, actually it's called SuperComputing, but I would like to see enterprise SuperComputing. I would like to see the proliferation of AI through enterprises. What we are creating today is to allow actually enterprises to break the barriers of entry just starting implementing AI. And more than that, I would like to see customers in the enterprises worldwide. So currently, there's a lot of, again, GPUs are a rare resource and you need to move the data to the GPUs. So, what we're building is to help customers break these barriers and just use them wherever they are and get the value in an enterprise-secure fashion.
Dion Harris
>> Yeah. Very cool.
Savannah Peterson
>> Well stated. I can't wait to discuss all three of those predictions next year.
Dion Harris
>> There you go.
Savannah Peterson
>> Ronan, Shimone, Dion, thank you so much for hanging out with me today. This has been fun.
Dion Harris
>> Thank you for having us.
Savannah Peterson
>> I feel smarter. Your energy is great. It's just like the energy here on the show floor. I hope you're getting as much energy from these exciting conversations as I am here in Atlanta, Georgia at SuperComputing 2024. My name's Savannah Peterson. You're watching theCUBE, the leading source for enterprise tech news.