We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For SC24
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for SC24.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to SC24
Please sign in with LinkedIn to continue to SC24. Signing in with LinkedIn ensures a professional environment.
Generative AI has driven tech and business changes. Enterprises are shifting from exploring AI to implementing it for real-world applications to boost ROI and utilize data sets. Key AI trends for 2025: AI inference, secure language models, and preparing for an exascale future. Mainstream enterprises focus on AI inference over training large language models. Traditional large enterprises adopt AI projects by adjusting pre-trained models with data sources. RAG deployment can be complex. Shimon Ben-David discusses scaling infrastructure challenges in exascale en...Read more
exploreKeep Exploring
What are the key AI trends poised to reshape enterprises in 2025?add
What is the barrier of entry into a gen AI environment compared to a traditional AI environment and how are enterprises looking to benefit from implementing LLMs through 2025?add
What is the trend in storage capacity and data generation observed this year, and what new challenges are emerging as a result of this trend?add
What features were WEKA Storage Technologies built with in order to ensure scalability and accommodation for future growth and changes in infrastructure?add
What is the current outlook on artificial intelligence and its potential impact on industries in the future?add
>> Generative AI has captured the narrative and has catalyzed a lot of rapid shifts in technology and business priorities. Now, the initial wave of enterprise AI centered on experimentation, a lot of governance and training and concerns about legal and compliance issues. But really the focus has been on training large language models. But as 2025 approaches, enterprises are moving from exploration to implementation and they're refining their AI models for real-world applications to drive ROI. And importantly, they're getting ready to tap the potential of the proprietary enterprise data sets. And in this conversation, WEKA's Jonathan Martin and Shimon Ben-David unpack the key AI trends poised to reshape enterprises in 2025 from AI inference, the rise of small, secure and sovereign language models to the vital importance of power, performance density, and preparing for an exascale future. Now together we're going to dive into the challenges and opportunities that lie ahead for AI scale. Welcome gentlemen, to our pregame coverage of Supercomputing 2024. It's great to have you on the program.>> Great to be here.
Shimon Ben-David
>> Great to be here. Thank you.
Dave Vellante
>> Let's get into it. So guys, training has been all the rage. It's driven billions of dollars in CapEx. As the big tens of billions, hundreds of billions of dollars really is the big five LLM vendors, they rush after what we call the holy grail of artificial general intelligence. But for mainstream enterprises, inference is going to be the name of the game and it's expected to dominate the future of AI workloads. Jonathan, let's start with you. How do you see the current state of gen AI and what should our audience expect in the coming months and years ahead in this topic of training the inference?>> Sure. So I think you're right on the money in your opening around 2024 being a big focus for the AI explosion or the first wave of the AI explosion really being around large scale training of models, build out of GPU clouds and a significant amount of infrastructure and large enterprises really focusing on governance, on compliance, on risk, on bias, on the policy end of AI, as we move into the backend of 24, we are absolutely seeing that traditional large enterprise is beginning to technically foray into deploying their first AI projects at scale. So we're starting to see financial services, manufacturing traditional large enterprise who typically in 2024 haven't been deploying at scale beginning to do so, but they're doing it in a very, very different way. They're taking existing pre-trained models, they're fine-tuning those models, they're augmenting those models with other sources of data and they're beginning to run those in production for the very first time as we go into 2025.
Dave Vellante
>> And I think that's the key is their sources of data that data's not going to just seep into the public domain, and that's really where the competitive advantage is. Shimon, I wonder if you could discuss the critical elements of this second wave that Jonathan just talked about. How is this going to affect in your view, the priorities and what tech trends should we expect are going to unfold as a result?
Shimon Ben-David
>> Brilliant. I think as Jonathan mentioned, the barrier of entry into a gen AI environment is actually lower than what it was for a traditional AI environment. So if in the past you had to train your models and you had to have a whole practice of data scientists around it. Today, it's very easy to just take models, existing models, pre-trained LLMs and run them in your environment. Now saying easy is actually a bit of a lie. It's actually easier than what it was before. Eventually, enterprises that will implement it are looking for an outcome and we're seeing that going through 2025, they will actually look for a better ROI on their investment because they need to benefit from these LLMs and they need to do it in a way that actually allows them to get more revenue than not doing it. So talking about the overall arching theme, it's how do I as an enterprise manage to get a better ROI? The way to do it is enterprise are now exploring whether to continue using or start using inferencing services as a service in cloud environments or maybe build their own enterprise inferencing environment in a GPU cloud or on-prem. If they're actually going to do it, how are they going to do it? Because this is relatively a new field, it's out there for two years already, but that's not a lot in terms of creating blueprints. So there's actually a lot of exploration still whether to do it with different frameworks, which networking. I think GPUs is actually... we're seeing NVIDIA dominating that market. We're seeing that trend actually increasing with maybe additional players coming into play. As you mentioned, how do you augment the data that your LLMs are familiar with and how do you benefit from that as well?
Dave Vellante
>> I wanted to ask you guys about retrieval augmented generation, and we did a survey this summer and I was surprised at what a low percentage of enterprises had actually embraced RAG and put it into deployment. And when I talked to some of the folks that took the survey, they said, "Yeah, it's not as simple as everybody thinks." Do you guys have a point of view on that in terms of, in the one hand there's the cost aspects of what you pay, but also just the time it takes and the skills? Do you see that accelerating and changing?
Shimon Ben-David
>> Definitely. Maybe I'll take a stab at it. So we are actually building these environments and we're actually interacting with customers that are doing that at a POC level and at the production level. And it's very easy to get a RAG pipeline POC environment. It's much more complicated to take that POC environment and deploy it in a way that is production-ready, secure, scalable, safe, and in a cost-effective way. As I mentioned, there's a lot of know-how regarding the pipeline itself, which utilities, which frameworks, and how to stitch all of these environments together. I have to say that NVIDIA is doing a really good job at providing scalable utilities such as Nemes and the NeMo stack to simplify these environments. And we are seeing... we're actually deploying a few of these environments as well, and we're seeing where it saves in complexity. Having said that, it's still something that's not trivial. It's still something that customers are exploring. There is new frameworks every day, and honestly, this is where customers are still looking for these blueprints on how to do it, what's the best outcome? How can we talk with organization that have that knowledge and familiarity and can guide us through this?
Dave Vellante
>> Let's talk about exascale. We're talking massive scale. That's what you guys are all about. Just for the audience exascale, it's a quintillion calculations per second. So think 10 to the power of 18. It's just mind-boggling, six orders of magnitude, more than a trillion. So it's just incredible. Now, we always talk about you can't have good AI without good data. GPT-4 was trained on half a petabyte of data, but organizations have far more, J.P. Morgan, for instance, it was reported has 150 petabytes of data, and I'm sure there are many, many organizations have greater amount of data. So as that data growth curve continues, it's bending exponentially. So exascale computing, Jonathan, it's no longer a niche. It's becoming a necessity, isn't it? What are you seeing in your customer base with respect to exascale and how real is it beyond those most isolated government and supercomputing labs?>> Yeah, so I'd say this is the year where also exascale has become very real. So at the start of this year, we had no customers that were over an exabyte of storage capacity. We'll end this year with five customers over an exabyte and one customer almost 10 exabytes. The realm of particularly text to image and text to video is burgeoning very, very quickly, and that generates absolutely massive volumes of data. Because of that, what you've found is that there are a whole bunch of new phrases and new metrics that are beginning to appear. Last year was very much just AI at all costs. This year, increasingly there is a cost, and that cost tends to be at times quite eye-watering. And so the finance people, the governance people are getting involved and more and more organizations are beginning to get focused on a new set of metrics to measure the impact of the infrastructure they're deploying. So people are beginning to use phrases and metrics like measuring the number of tokens that they're generating for every dollar that they're spending on infrastructure, beginning to measure the number of tokens they're generating for every watt that they're using, beginning to generate or to measure the number of tokens for every hour of processing time that they're running. And a massive amount of that is getting focused on things like what's called the power density of the environment for a rack of equipment. How dense can you make that equipment? How can you maximize the performance that it has at the lowest power consumption? As these environments get bigger, and last year big data centers were being measured in tens or maybe hundreds of petabytes. I think next year you're going to see a number of data centers at large corporations being measured in the tens of exabytes. Power density, performance density become incredibly important, and doing simplicity at scale is a brand new challenge that many organizations are doing for the first time. Not only are vendors selling exabytes of storage for the first time, organizations are deploying it and implementing exabytes of storage for the first time. So it really is a new panacea for everybody that's involved in this space.
Dave Vellante
>> Thank you, Jonathan. Shimon, I wonder if you can address what are the technical challenges of providing infrastructure at this type of exascale environment?
Shimon Ben-David
>> Brilliant. I think there's a way to categorize it by a few levels. So first of all, there's the hardware complexity. Exascale means a certain amount of footprint, racks, servers, controllers, cables, network switches, network topologies, heating, cooling. So first of all, there's the massive amount of footprint that you need to handle, and that's a big problem. Data center capacity is a real challenge in modern environments. Powering all of these racks is a real challenge. So the more you can shrink these environments, obviously the better. But first of all, and I would say even that this is almost the easy problem, that the next problem is actually the logical problem of how do you scale your data environment, your storage environment to accommodate for these exascale capacities. That's something simply that not a lot of storage solution can do. We've seen customers that had a storage environment, a data environment that simply could not scale anymore in terms of capacity just like raw gigabytes, terabytes, petabytes that it could handle. And also there's another aspect to it that sometimes we , but that's just the number of objects, files, iNodes that the system can handle. Eventually everything is represented as a logical entity and there's tables and memories, there's data structures that needs to be accommodated for. There's algorithms that need to process those in an efficient way. At some point we're seeing usually that data environments can scale in terms of capacity and performance, but then at some point there's the diminishing return where they either cannot scale anymore or worse, they even decrease in performance and efficiency. So that's a massive challenge actually, that logical challenge. I'll give an example with WEKA, what we did is we actually created the environment in a way that all of the data structures and algorithm are computed base. So there is almost no tables in memory. So we have a very thin and effective dense environment. And also the data structures are built in a way that the environment can scale and not increase in terms of memory on the compute nodes and on the storage nodes. Also, there's the scale. When we look at scale at modern scale, there's also the how do we scale between multiple data centers because we are seeing that in 2024 and definitely in 2025, organizations are not going to be single geolocated anymore. There's usually I can get compute and I can get GPUs in multiple environments. Maybe I already as an organization, I'm acquiring data in one environment and I'm processing it in another environment and then I'm inferencing on it in a third environment. So there's all of this notion of how do I scale globally also? So again, if you're looking at WEKA with the ability to move the data seamlessly at petabyte and exabyte scale between this location that facilitates for that. So the combination of a very dense footprint with everything built to scale from day one, and this is part of the being a modern environment and with the data movement actually accommodates for these future scale challenges.
Dave Vellante
>> Speaking of future scale, Shimon, the way you're describing this, I think about exponential growth and sometimes it's hard to fathom what the future is going to look like even in the near term next 2, 3, 4, 5 years. How do you make sure that what your customers deploy today are going to actually accommodate the needs in the next couple of years?
Shimon Ben-David
>> I think that's a great question and that's something that we're seeing as part of the customer exploring new environments, new frameworks, exploring gen AI. We're seeing the way that they're currently... as I think I mentioned, the way that they're currently doing things is a day one implementation. We're only scratching the surface of the art of the possible, of how their environment would look in the future as they grow, as they scale in terms of compute and capacity and between different data centers. WEKA was built to scale in all of these dimensions, capacity, performance, footprint. It was built to be able to accommodate for new hardware environments, new CPUs, new GPUs, new accelerators. It was built to accommodate for cloud environment. It was built to also shrink and expand as needed to accommodate for better power utilization, for better ROI on the environment if it's a cloud environment for example. So we're future-proofing the environment by looking at where customers are actually going to be in the next 2, 3, 4 years and making sure that WEKA can actually accommodate that and more. And we actually... maybe I'll finish with this. We're actually also finding ourselves guiding customers because we have a lot of accumulated knowledge regarding a lot of large-scale AI projects in production. And we're seeing how these large-scale AI projects in production, which are a year or two ahead of the market even are doing things. We actually take that and we help customers accommodate for their new environments using that experience.
Dave Vellante
>> Thank you, Shimon. Jonathan, the narrative today, of course, we talked a lot about it here around ROI, the cost, and I like to say that enterprises are kind of hitting singles today, but they're excited to really lean in. You've basically got five giant LLM vendors, two of which are open source, and the economics of that space are brutal. But when you think about getting beyond training and taking that proprietary data that we were talking about within organizations and start driving inference and edge type of use cases, from a business perspective, Jonathan, how are your customers do you think going to be leveraging inference in the future? What's the value that that can bring to organizations? Maybe you could paint a picture for us.>> Yeah, I think the values is phenomenal. People ask me all the time, "Are we done with this AI thing? Is the bottle over?" And I keep going back to, it kind of feels like the internet in 1994 where your vision of the internet or your perception of the internet was you got on your 28.8 modem and dialed into a bulletin board service and that was the internet. And so it's very hard to think 20 or 30 years down the line that you can walk around with a bit of plastic in your pocket with some of the world's knowledge on it. And that's where we are with AI. It's incredibly, incredibly early and the market is evolving incredibly quickly. So last year, obviously a big focus on the monolithic model. This year, I think more progressive customers are really thinking, how do I use an ensemble of models? Those models, some of them are very general, some of them may leverage ChatGPT-like interfaces, but other models that are incredibly specific. And it's a combination of large generic models and a slew of smaller, very, very specialist models that will work in an ensemble, in an orchestra together to deliver an outcome. And I think for the more progressive customers that we're working with that is the nut that they're trying to crack. If they can crack that, then I think you're going to find in just a few years there is going to be two types of companies on the planet. There's going to be companies that are AI native at the core that have built a solid data pipeline, and they are industrializing the process of ingesting large volumes of data and transforming that data through an ensemble of models into tokens and into insights. And then there are going to be other companies that are not. We saw the impact of the internet on transforming existing industries and creating brand new industries. I think we're going to see exactly the same with AI over the next couple of years.
Dave Vellante
>> And it could actually be a double whammy. I've been saying that I think you're going to have both organizational top-down command and control implementations of AI, plus you're going to have citizens AI where personal productivity is going to be driven by people close to the action who are going to learn how to deploy AI. That's to me is what an AI native company looks like. So I think it's internet plus PC cycle combined. It's going to be incredible productivity booms. But I want to talk about power and cooling because that seems to be the big elephant in the room, if we will, or potential blocker. I mean, these systems are just insanely dense. You talked about that earlier, the miles and miles of cabling. You walk into a data center and just put your hand near these things and you can fry an egg on them. So what are you hearing from customers in terms of the sustainability challenge? People are moving data centers closer to where the power is. Jonathan, what are you hearing from customers and governments about this concern? And then maybe Shimon, you could talk about how WEKA is architecting to address this issue.>> Yeah. So you're absolutely right on the money with that. I think again, the numbers are eye-watering. ChatGPT-3.5 took about $4.6 million of power to train the model. Just five months later, ChatGPT-4 took $100 million of power to train the model. Every time you're on one of these new generative sites and you're typing in a prompt to create an image, every time that image gets created, it's the same power consumed as a full charge of an iPhone. So it's maybe no surprise that people are beginning to wake up that while AI may solve the sustainability and environmental challenges on the planet, it's probably also equally possible that it's going to be the thing that melts it. So because of that, you're seeing more and more and more organizations, we're certainly seeing in more and more RFPs a focus on tokens per kilowatt, tokens per watt, and tokens per hour focus. We also see though, is there's a bit of an impedance mismatch between the board or the executive team, maybe have a sustainability target to hit and yet the AI practitioners that are implementing these are kind of running and delivering AI at any cost at the moment. So it's what we'd call the sustainability conundrum or the AI sustainability conundrum. It's an initiative that we kicked off about two years ago here. As you may or may not know, our series D was led by Generation, by Al Gore's fund, ex-Vice President of the United States. And so there's a big focus here on how do we deliver all of the benefits of AI, but do it in a way that is sustainable and is not going to melt the planet.
Dave Vellante
>> Thank you. Shimon, what can you do? You optimize through software? What are the tips and tricks that you're applying to actually solve this problem technically?
Shimon Ben-David
>> Loads. So first of all, what is using off-the-shelf component? So we are a software environment, so we didn't take any hardware shortcuts. So when you look at the WEKA bill of material on how you deploy a WEKA environment, these are just servers that... and we do everything through software, all of the data distribution, all of the data protection, everything is through software. So by that, we actually removed a lot of hardware components that we simply don't need. We don't need JBODs, JBOD storage class memory, NVRAMs, NVDIMs, RAID controllers. We consume much less network cables, connectivity, no fiber channel right? So by removing a lot of the hardware, the necessary hardware from these petabyte, exabyte scale system, we actually consume significantly less power. More than that, if we're looking at the rack performance density, if we're looking at the WECA system certain capacity performance compared to alternatives, that could be 10, 15x of a difference that consumes significantly less power. So it's significantly more efficient. I would also say that we have the ability to, and we're seeing customers utilizing it, obviously as I mentioned, to run on different cloud environments, and we do know that some cloud environments are more power efficient than others. We even had this idea around having a sustainable API where you can burst into a data center that is now powered by sustainable energy and then burst back when to another data center is needed. But even more than that, the ability to run on a cloud environment and scale according to what you need at that point. So for example, if I need a certain capacity performance, I am utilizing a WEKA at a certain size and then as I need more, I scale that environment to additional instances, servers. And also I'm obviously scaling the compute. And as I scale down the compute because I don't need... I completed the computations, I can scale down the storage environments again. So I'm using power in the most efficient way. And I would say that there's also another secret weapon that we're using when applied, and that's our converged mode. And you can think about it as our zero footprint storage. And again, where it applies and we see customers using it's actually converging WEKA on the same GPU servers because if you can now run... if you have a slew of GPU servers, 10, a hundred, a thousand more, and these GPU servers already contain components that are being used to compute and store data, today we see a lot of customers just moving around data between different storage environments, copying the data, wasting time and effort. Imagine that you could have run and you can, WEKA on these environments alongside in a safe way with these GPU environments just carving up a few slim sliver of resources from the same GPU servers, now you have a zero footprint environment that provides a high scalable, performant, resilient file system, WEKA, at no additional footprint. So there's actually, we see customers converging on hundreds of GPU nodes getting the performance that they need with no additional footprint. It came to be more efficient than that even.
Dave Vellante
>> Awesome. The Supercomputing event, it used to be the exclusive domain, is kind of a nichey domain of the supercomputing-powered labs and governments, but it's become the Super Bowl of AI and WEKA has a dominant presence there. Jonathan, tell the audience what's happening at Supercomputing 2024 in Atlanta. What do you guys have going on? I know there's lots of action at the show.>> Lots of action. So we have some really very exciting announcements coming up, a number of industry firsts touching on a lot of the things that we've talked about here, touching on things like inferencing, doing things differently at exascale, helping people transform their power profile, building simplicity into doing these environments at scale. So we've got some really, really exciting announcements, some exciting industry firsts. Obviously we have a massive booth presence. You won't be able to miss the purple on the show floor there. And then we socialize when we're at these things. So Wednesday evening, I'll do the gratuitous plug. We're running WEKAFest, got about 1200 people coming to that event. Jimmy Eat World's playing. We've got DJs from Ibiza. It's going to be certainly the party of the week at Supercomputing. So if you're interested in coming to WEKAFest, come check us out on the purple booth center of the floor.
Dave Vellante
>> Fantastic gentlemen, thanks so much for your time. We'll see you in Atlanta. We are going to be featuring WEKA on theCUBE. We've got tons of action, three days, wall-to-wall coverage. Appreciate your time and have a great show next week.>> Thanks so much. Safe travels.
Dave Vellante
>> All right, you too. And thank you for watching. We'll see you next time. This is Dave Vellante for theCUBE.