We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For SC24
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for SC24.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
SC24. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to SC24
Please sign in with LinkedIn to continue to SC24. Signing in with LinkedIn ensures a professional environment.
TheCUBE's coverage of day three at an AI show at SuperComputing continues with discussions about the world's most efficient AI computing accelerator for inference. d-Matrix, a company focused on generative AI inference at high speeds, efficiency, and cost-effectiveness, introduced a PCI card for AI workloads. The company collaborates with ecosystem partners to offer flexible solutions to customers, with a focus on user experience, cost, and energy efficiency. The accelerator card aims to handle large models with batched throughput and low latency for interact...Read more
exploreKeep Exploring
What are the key focuses when building a computing platform for efficiency?add
What is the approach taken by the company in terms of selling and implementing their computing platform?add
What are the key signs in a customer's environment that indicate they should consider using d-Matrix's product?add
What are the key aspects of the process of training and using a machine learning model for inference?add
What are the key components needed for generative AI inference?add
>> Welcome back everyone to theCUBE's live coverage. Day three, where the show is really cranking, we still got the adrenaline pumping. I'm John Furrier, host of theCUBE with Savannah Peterson, Dave Vellante, Kristen Nicole Martin. We've all been here on the ground getting all the data, a lot of action. Obviously, it's an AI show, high-performance computing. It's really evolved, and you can see the clear lines of sight into how the cloud's going to play into it, how the data center's going to erupt in huge opportunities as the large supercomputing is coming to the masses. Sid is back on theCUBE. He was in our studio for our AI NYSE Wired community event. He's back, founder and CEO of d-Matrix. Sid, great to see you.
Sid Sheth
>> Yeah, I'm back.>> Thanks for coming back on.
Sid Sheth
>> Pleasure to be back, thank you for having me.>> First of all, I remember the conversation we had in Palo Alto at our event, AI Leaders, you were on. You were talking, setting the table for how the data center's going to change significantly. Making AI inference affordable, making the scale, changing the architecture. You got a lot of funding in 2023. Now, you launched here at SuperComputing.
Sid Sheth
>> That's right.>> Okay, let's get right to the hard news first, and we'll get into what we talked about in the underlying scale, but what's the news? The d-Matrix, what's the company do and what's the news?
Sid Sheth
>> Well, we are super excited at SuperCompute to announce the world's most efficient AI computing accelerator for inference. We built this product with inference and inference only in mind. When we started the company back in 2019, we essentially looked at the landscape of AI compute out there and made a bet that inference computing would be the largest computing opportunity of our lifetime. For obvious reasons, I mean you can train models only so many times and then you got to deploy them, help businesses operationalize with the help of AI, and that all requires inference. The leap of faith at the time, of course was this was 2019. It was not very clear when we would have the big watershed moment in AI, and that happened in '22 with ChatGPT. And then as soon as we went past that moment, everyone the world over was talking about what does it really take to deploy really fast inference and cost-effective inference? And d-Matrix as a company is really built for generative AI inference being done very quickly, very fast. We talk about three key things that we focus on. One is, it's all about doing more with less. Doing more inference in less time, doing more inference with less dollars, doing more inference with less power, less energy. The more with less paradigm is really what efficiency is all about, and we have built a computing platform that is efficient on all fronts, whether it is energy efficiency, cost efficiency, or speed efficiency.>> Talk about the product, what you guys are selling. Is it the data center? What is the product? How do people use it? How is it deployed and consumed? Price? Give the specifics of the actual product.
Sid Sheth
>> Yes, so it's a computing platform, it's an acceleration platform. It's a PCI card that we have built. Today, we have started with the PCI form factor. We can go to other form factors like OAM, but we build a silicon at d-Matrix. We build the entire computing silicon at d-Matrix. We've built the acceleration cards to package the silicon together, and we built a software stack that goes along with it to essentially map AI workloads onto the silicon. We sell the whole unit along with the software and then we work with partners in the server ecosystem that would essentially put these accelerator cards into their servers, and then we can go to data center operators who build out the racks and the cooling. But the big difference in the way we go to market is the fact that we take a very collaborative approach with the ecosystem. We are not going in to the data center saying, "Hey, we've got a whole rack," or we are coming in with services or anything of that sort. We build the cards and then we let the customers decide which server vendor they want to work with, how they want to build their racks, what kind of cooling they want there or liquid. And that's the flexibility that is so important when it comes to building out an inference computing cluster, because inference again is all about efficiency and cost.>> And choice too.
Sid Sheth
>> And choice, exactly. It's a lot about choice, and we offer all of that by working collaboratively with the ecosystem. We essentially give customers plenty of choices.>> Okay, explain. First off, I get the cards, the PCI cards, the hardware you put into the devices. So inference, I'm going to say open and choice become a big theme. I know you guys are looking at plug and play and making sure you can work with everybody. How does that work? Take me through the nuances. So let's just say I'm a customer. Or well, first of all, are your customers the end user? Are your customers the OEM manufacturers? Split that out for me, who's the buyer of this?
Sid Sheth
>> We work through the chain. We don't necessarily work directly with our direct touchpoints because I think inference is still early so it's very important to understand the end use cases and making sure that whatever we've built really solves a pain point at the end customer. We actually prefer to work with the end customers, understanding their pain points, what are they trying to really accelerate when they want to use inference? And not all inferences build the same.>> You're selling directly to the end user customer.
Sid Sheth
>> We are directly selling to the end user, and then we work backwards from the end use case.>> Got it, okay.
Sid Sheth
>> And then work with the right server partner or the right data center, cloud hosting provider, et cetera.>> Just a classic, we've got a great platform, you buy it, you put it in your environment, and then of course there might be other opportunities for you to work and partner.
Sid Sheth
>> Right.>> OEM and other thing. Okay, got it. All right, so let's think through. If I'm a customer and I want to have a Dell, I want to have other vendors in there, it's just on the network? Or how does that connect and how do you manage that?
Sid Sheth
>> We already are working with multiple system integrators, whether it's take the standard OEM folks or the ODM folks. We have many options already on the system integration side. We announced three at the show: Liquid, we announced GigaIO, we announced Supermicro as three of our partners and we have others which we will be announcing as we move forward here. But there'll be plenty of system integration partners. And again, we are flexible. The way we have built the PCI card, we do not need any special server configuration for that card to plug into. It is something that is already available as a server configure, pretty standard server config from pretty much a lot of the system integrators, so options are already there. Really, it comes down to what is the end user application that the customer is trying to solve for, and how we can make them better at that.>> When we were in Palo Alto, when we had our digital twin event with the NYSE Wired community in theCUBE, you talked about transforming AI workloads.
Sid Sheth
>> Yeah.>> Specifically about unifying the architecture around acceleration of inference. Explain that again, and tie that to the news and how the product works. Okay, you got my attention. Inference is the killer app, but there's different inference use cases and scenarios. I might need different kind of inference for this. And there's a lot of data involved, so whoever's downstream or upstream on the data.
Sid Sheth
>> Unsent work.>> That's good, I can pack this.
Sid Sheth
>> Absolutely.>> But take me through the unification and then how did the product work? Now I'm going to deploy it.
Sid Sheth
>> Yeah. The stuff I was touching upon when we talked about this in the past was around a unified model architecture, which was a transformer architecture, we talked about that. That emerged in 2017 and by, I call it 2020, 2021 leading to ChatGPT, the transformer architecture essentially had proliferated into multitude of different workloads. It showed that it is a very scalable architecture. It's a multi-modal architecture so it works with different modalities, video search, text, et cetera, images. The underlying engine of generative AI inference today runs on transformer-based models. We have built a solution that is highly optimized for transformers. In that sense, we are taking advantage of a common architecture and building an inference computing platform that would accelerate all of those models. That gives us broad reach with our solution. Now, coming to the use cases, what we announced at the show was an accelerator card that is really, really good with what we call low latency batch throughput. What I mean by that is we hear a lot of people talking about, "Oh, my latency is really great, but it only works with a single user."
Or we hear about GPUs, talking a lot about batching, but they never talk about latency. What we have announced is a product that can do batching if it has to. It can certainly do extremely low latency with single users if it has to, if that's the use case. But really what we think the use cases are veering to are the low latency batched use cases, which is, let me give you an example. A good example would be something like video generation. Somebody wants to interact with video and create a video in real time. Right now, if you see the way video generation happens, you prompt a model, it takes minutes for the model to generate a video, typically generates 10 seconds of video. You may not like it, you come back, you re-prompt the model, you go away. That way of dealing with video is just not working. Where we want to go to is a highly interactive model where you are dealing and interacting with video in real time like you would interact with a chatbot or ChatGPT for instance, where it's interactive. Where you prompt the model, you get a response, you don't like the response, you prompt it back and it responds back in real time. We are not there with video, and the reason for that is memory, bandwidth, memory compute, memory capacity is a real problem. GPUs today don't solve that problem. What we have done is we have essentially created a class of memory that is tightly coupled to compute that allows for these kind of interactive video applications to get created. You typically will never have a single user working with these applications. You have a batch of collection of users, so multi-user low latency. You want to create a use case where you can have multiple users all interacting with a model simultaneously, but not having to sacrifice the latency for each user.>> So you're an accelerating situation, not offload or.
Sid Sheth
>> We are accelerating. We would be accelerating the video influence use case.>> Got it, all right. And what are some of the other use cases? Okay, let me ask the question differently. What problem do I have if I'm the customer that I, what signs in my environment tell me that I need to talk to you and get your product? Is it the latency, is it the batch piece? Or is it just I got certain configurations? When do I know to call d-Matrix up?
Sid Sheth
>> That's an excellent->> For saving the day.
Sid Sheth
>> That's an excellent question.>> Basically, because you've become, you're an accelerant.
Sid Sheth
>> So if you're dealing with any kind of use case that you need, first of all, number one priority for you would be if you're dealing with a use case where you have multiple users where they're all interacting with a model, but the interactivity with the model is not where you need it to be. Right?>> Like slow.
Sid Sheth
>> It's very slow.>> Right, exactly.
Sid Sheth
>> Very slow. And you say, you know what?>> AI can't be slow.
Sid Sheth
>> Yeah, AI cannot be slow, and it's affecting my user experience. Users don't want to stay on the application because they don't like the user experience. That would be the first time you pick up the phone and call us, right? Say, "I've got a problem. I've got many users trying to access this and really the user experience sucks and can you help me make it better?"
So we come in there. Now, all of that, eventually we need to go solve that problem first. That's problem number one. Problem number two would be once we solve that problem, okay, how much does it really cost me to serve each user? Cost is a very big ingredient.>> Huge.
Sid Sheth
>> Right? Because you want many, many users to be able to use this. And as more users use it, it becomes, the cost escalates and you need to have a fundamentally different way of doing compute that keeps costs very low. We have solved that problem too, just the way we build the platform. And then the third thing would be energy efficiency. So you say, okay, great. Now I have a great user experience. Because I have a great user experience, a lot of users want to use it. Guess what? I'm going to be spending a lot of power and energy because I'm trying to support many users. Are you energy efficient? Can you do more influencing in a given power envelope compared to a GPU? And we can do that too. So I think it's one, two, and three. Number one, user experience. Number two, cost. Number three, energy efficiency.>> That's a really great way to describe because a lot of times people don't know where they're paying. They have a lot of pain right now. I say good pain by the way. No pain, no gain as they say in sports. So that's a good call out. I want to ask you about one other thing that's come up on theCUBE that is going to be part of re:Invent, a lot of the conversations at Amazon Web Services conference. Resource allocation also is an issue because now you have multiple systems mumbling to each other, and the cost piece is tied to that too. So it's not about cost optimization, it's just knowing about just overall cost. I mean, even go back in the old school routing days, you have least-cost routing. I mean, quality of service. I mean, we're in a quality of service game. We're accelerating. Can't be slow. Inference is a killer app and the answer or a prompt can't be slow because the next one's coming.
Sid Sheth
>> Right.>> It's a fatal flaw of AI. Networking and speed, and is the job complete? Never mind the fact that the GPU's got to do all that work too.
Sid Sheth
>> Right.>> So what's the resource conversation? Are people talking about that because you accelerate, great, but what's required? I mean, what other things do I have to do around the solution? Or is it truly not disruptive in sense I got to spend more over here? Or maybe the question is what's the requirements for you guys to have this, and what's my forecast of my job? I might have to do other things. That's the way I would be thinking about. What's your answer to that?
Sid Sheth
>> Yeah. The other thing that we talked about when we announced our product is how we do really well with 100 billion parameter models inside a single rack. That's the target. What we are seeing is 100 billion in a rack. We can do models that are 100 billion in a rack and we can do them better than anyone else on the three things that we talked about, which is the user experience, which is interactivity, cost, and power efficiency, and energy efficiency. Every enterprise customer that we have spoken to, most of them at least feel that their AI journeys are starting with models in that size bucket. They feel 100 billion is plenty and that can fit in a single rack. One of the reasons we did that was as you go across racks and you need a lot of hardware and you need to scale out the complexity of the problem grows. Do you really need to take on that complexity? What are you really solving with that complexity? In our opinion, you don't really need that complexity for most customers because most customers are like, "Hey look, my AI journey starts with models that are a lot smaller than 100 billion. If you can fit that in a rack, I don't need to go across many, many different racks."
Now that we cannot go across racks, we can certainly go solve that problem, and we can do models that are much larger than 100 billion, but what is the target customer opportunity around those much larger models? And we think it is not as big, at least right now.>> So you put your card in the machine where the models are.
Sid Sheth
>> Right.>> And then it does its job. That's it.
Sid Sheth
>> That's it. It sits inside that rack and it does its job and we think it serves a lot of customers.>> I got to ask you, because this comes up a lot, certainly in this crowd at SuperComputing '24. I mean, this is, everyone knows what inference is. But sometimes people don't really know where in the... An interaction inference does. People can understand training. Take a bunch of data, train and figure it out. I'm oversimplifying it obviously. When you get to inference, what is inference? Is it just a query, a prompt? Is it doing other things? How would you explain to an average person in tech what inference is? Is it getting smarter doing some training? Is it reinforced learning? Is it just a prompt? Is it getting a better answer? What does inference do? How would you describe inference?
Sid Sheth
>> I like to use an analogy that I use quite often and hopefully that helps your listeners, which is it's not very different from us human beings getting educated and applying our knowledge. We spent the first 20 years of our lifetime going to school and college. Depending on where you go to school in college, you can spend a lot of money doing it, but you acquire learning for the first 20 years.>> That's training.
Sid Sheth
>> And that's training. That's training. And then the next 40 years of your life, all you're really doing is applying that knowledge. And of course, you're learning stuff along the way. That's what they call fine-tuning. You're specializing in certain areas, but most of the time, you're really applying everything that you acquired in the first 20 years. You're monetizing that education. That is inference. And clearly, we spend a lot more time doing "application" of our knowledge or working, which is inference than the training part. So the inference opportunity is obviously much larger than the training part. And really, that's really the same analogy that applies with AI computing. You train a model, you make it learn, and then you use it to monetize the data or make decisions with the data that you already have.>> I actually use that example, Sid. I leveraged that example you said on theCUBE in the last event. Someone said, "Yeah, unless you're in academia, then you're constantly being trained."
I go, "Well, those are the models that's just, no one uses."
I couldn't resist. But let's get back to the inference. First of all, great explanation. So we're inferring, which means I'm calculating connecting dots and my brain's working. What's the technical version of that? What's going on that makes the acceleration needed? Is it the math? Is it calculating? How is inference happening? How can you explain that piece, and why is it hard?
Sid Sheth
>> Yeah. It is hard for multiple reasons, right? First of all, I think, so let's take a step back, right? The underlying math is it is in many senses, it's very similar to training. The only difference here is training has these recursive passes where you do a forward pass on the model where you take the data, you train the model. You hope that the model has learned. In most cases, it won't learn in the first pass so you recursively go back, you correct a few things, and you start again. And you keep doing that over and over again, right? Until the model starts converging and you get to a point where you know that the model has learned. In inference, you don't need to do that. Inference is basically what you do is take the data and you're essentially running it through the model in a forward pass, and all the model has to do is make a bunch of decisions. There's a lot of math. There's underlying math. There is matrix math. There is what we call non-linear math that is the neural network mathematics of getting to a decision. That is what inference.>> And that's what's cranking out the GPUs.
Sid Sheth
>> There's just a lot of compute. There's a lot of compute as you are doing a lot of those mathematical operations. But there is a lot of access to memory because what you're doing is you're essentially taking data, you're taking context, you're storing that context, you're remembering the query. When you query a model, the model is remembering every query that you put out there. It's keeping a state of all of that, so you need a lot of access to memory. Bandwidth, memory capacity, and a lot of compute. Three things.>> Sid, great to have you back on theCUBE again. You got the silicon, which is really critical. That's great. That's a differentiator. Sure, you got tons of patents, all that fun. I'm sure you had to defend that. But again, this is just going to be more headroom. PCI card today, tomorrow, what's next?
Sid Sheth
>> We're going to grow from here. I think the d-Matrix story has always been about very tight compute and memory integration because we really feel the generative AI inference problem is all about three things again, I want to repeat them. Lots of compute, access to memory bandwidth, and memory capacity. The best platforms and solutions will be the one that put compute and memory together in a very creative way.>> Tightly coupled.
Sid Sheth
>> Tightly coupled, and we are going to keep growing our road map in that direction. We have lots more to come in the coming year.>> Congratulations, d-Matrix launches here at SC24. We've got the scoop here on theCUBE. I'm John Furrier, your host. Thanks for watching. We'll be right back.