We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Marketing Leaders. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Marketing Leaders.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For theCUBE + NYSE Wired: Marketing Leaders
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for theCUBE + NYSE Wired: Marketing Leaders.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Marketing Leaders. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Marketing Leaders.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: Marketing Leaders
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: Marketing Leaders. Signing in with LinkedIn ensures a professional environment.
Julie Choi, chief marketing officer at Cerebras Systems, joins theCUBE's CMO Leaders Series Summit to discuss the advancements in Artificial Intelligence (AI) and their implications for the industry. She shares her insights alongside theCUBE Research and analysts, spotlighting Cerebras' third-generation wafer-scale engine and its revolutionary impact on AI training and inference. Choi delves into the dynamic developments happening within the AI space, specifically the partnerships with notable companies like Mistral and Perplexity, and how Cerebras' technolog...Read more
exploreKeep Exploring
What are the key specifications of the Cerebras wafer-scale engine and why did the speaker decide to join Cerebras?add
What is the Cerebras wafer-scale engine and its specifications?add
What updates can you provide on the performance and recent developments of the company, including partnerships and new product offerings?add
What breaking news is being shared from Cerebras regarding their support of Perplexity's new model, Sonar?add
What use cases is Cerebras winning on constantly currently, and what use cases are upcoming in the future?add
What tasks are being won on currently and what is the next area of focus for Cerebras?add
What factors contributed to the success and innovation of DeepSeek, particularly in terms of their use of resources and algorithmic capabilities?add
What achievement or innovation was discussed related to Cerebras and their team?add
>> Hello, everyone. Welcome back to theCUBE. I'm John Furrier, host of theCUBE. We are here in our Palo Alto office at the center of Silicon Valley,. Of course, we have the NYSE studio connecting Wall Street and Silicon Valley. This is our CMO Leaders Series Summit. TheCUBE and the NYSE Wired community, Brian Baumann, the founder and his team is here, and one of the major principals of the Wired community, Julie Choi, she is CMO of Cerebras, part of the early formation of the Wired community. We're now almost in our eighth month of doing the media. You're back. Thanks for coming in.
Julie Choi
>> Hey, John. It's so good to be here. Thank you so much, theCUBE and NYSE Wired for having me. It's an absolute pleasure to be here.>> Well, all thanks aside, I'm super excited you're here. One, Cerebras is doing extremely well. Congratulations to the entire team on the business performance, the technology performance. You guys have great benchmarks. You have news we're going to get to with Perplexity, again, on top of some other news we covered last week. But also, you got some awesome props here for us. This is what I've been talking about on LinkedIn, the god chip, the big chips are back. This is a way for the wafer engine. Can you share, first of all, what you have in front of you?
Julie Choi
>> Yeah, absolutely. This is our claim to fame. This is the Cerebras wafer-scale engine, third generation. This is really at the heart of Cerebras AI training and inference performance. This is 46,250 square millimeters of silicon. It's 900,000 cores, 4 trillion transistors. It's the largest chip ever created, let alone, only created for AI. This is actually why I decided to join Cerebras a little over a year ago.>> We saw each other at Supercomputing. Andrew, the founder, was on theCUBE. A lot's changed since then. Give us the quick update because business performance has been really strong for you guys, congratulations, and you got all the adoption. I saw some benchmarks, the speed's off the charts. Give us an update on what's going on, and we'll get into the news with Perplexity today. Give us an update on some of the performance.
Julie Choi
>> Yeah. Since we last talked, our inference has just gone through the roof. We've added support for DeepSeek, DeepSeek-R1 70B, and we just announced that a few weeks ago, about 48 hours after that week when it really hit the news. We're really proud to be able to offer DeepSeek-R1 70B in private preview. Just last week, we announced with our partner, Mistral, the fastest AI assistant now, Le Chat. Mistral's Le Chat is running on Cerebras. These amazing partners, it's just been non-stop since the beginning of 2025 and we're excited for this year. I think it's off to the races. >> Le Chat, I saw that top of the charts usage-wise. Congratulations on that. Then, today, a deal with Perplexity. Talk about that deal, because I saw the benchmarks on the Perplexity ran a blog. You guys had a release right after. This is just, again, speed is everything in this game. Speed wins. Talk about speed. You guys are playing the speed game right now with all this innovation. What's the news with Perplexity? Give us the update.
Julie Choi
>> Yeah. We're so thrilled today to share here some of the breaking news from Cerebras, is that we're supporting Perplexity in the inference of their new model, Sonar. Sonar is an adapted Llama, so it's Llama 3.3 70B that Perplexity has adapted for their advanced search experiences. Cerebras is basically powering the speed. One of the things that I'm really excited about is to really see the speed of fast inference finding its way into every consumer's life. Partners like Perplexity, they really allow us to bring the magic of AI through advanced search experiences that we've never seen.>> I saw Perplexity ran a blog. I'm reading, Meet the new Sonar, a blazing fast model optimized for Perplexity search. Two things to point out here. One is I love the fact that search is back, because everyone thought, "Oh, Google's never going to get dismantled." Perplexity, ChatGPT, all these engines, La Web are being used now more and more than some Google search. But in the blog post, they got graphs, so they got charts. Sonar has got better speeds than Claude, GPT-4o Mini and Claude 3.5. But then, they got my favorite thing that shows the speed, which reminds me of when everyone goes, does a speed check on their internet, "How fast is my internet?" It just blows away everything else, Gemini and Claude. Take us through, why the blazing speed? What's the secret sauce?
Julie Choi
>> I believe the performance is, and keep me honest, I think it was 1200 tokens per second.>> Yep, 1200 tokens per second.
Julie Choi
>> This is top speed for serving this type of model.>> 140, by the way, Gemini, and 75 for Claude.
Julie Choi
>> Right. 1200 versus these other numbers, this is fastest in the world. Now, the answer to how this is possible, it's back to this beautiful piece of hardware. It really is the way that this architecture was designed with memory and compute co-located on this one beautiful piece of silicon. When you co-locate memory and compute, what that enables is for the data to move very quickly on that piece of silicon. We, basically, have 7,000 times more memory bandwidth than even the most advanced GPUs. That's 7,000 times advantage in memory bandwidth, which is coming from our architecture. This is what's translating into these world-leading speeds.>> This is propelling the AI infrastructure game to go faster, which will unlock all these startups dying to build applications, the POC purgatory at the enterprise, we're going to see that unleash. I think the DeepSeek style innovation speaks to what's around the system, how you build these systems, so I think this is, again, another big trend, so congratulations.
Julie Choi
>> Thank you so much, John. Yeah, we're so excited.>> We get the CMO leaders in here. You're a CMO and all this great news a little bit kind of blending it together. Your career, you were at Intel Mosaic ML, which is state-of-the art now. Naveen is over there at Databricks. You were pioneering AI with Mosaic. We've seen, at HPE and then Intel, your career, you know this game. Share your perspective of the magnitude of the opportunity that the new AI infrastructure will enable, because this is continuing to go very, very fast. We've seen the pre, you entered in with Mosaic, the pre-AI that's now thriving and growing super fast. Scope how big this is from your perspective.
Julie Choi
>> Yeah. I think that this is the single most significant moment for developers, because what we're seeing now as these foundation models are being exposed as APIs. Not slow APIs but truly state-of-the-art open source model APIs, and they're being offered to developers through platforms like Cerebras. What this is unlocking is a level of innovation that we haven't seen since probably 2008, when Apple launched iOS. I remember that moment. It was an earlier time in my career. I was one of Yahoo's first developer product marketers, and I just remember 2008, it's like, "Everyone was talking about social web, right?">> Web 2.0.
Julie Choi
>> Yeah, and I don't even think Facebook had IPOed at that point. Mark Zuckerberg was trying to figure out how to make sense of social. This idea of apps and developer innovation, this is iOS, it unlocked mobile. That was that wave, 2008, I'm kind of dating myself, 17 years ago. Here we are. We are in the AI era. What's different now is these tools, these APIs, once again, are going to empower an innovation that we have yet to see.>> What's great about that example is that iOS, obviously, was the iPhone proprietary to Apple, not proprietary but closed, if you will. They had apps on the phone to incubate the experience, so we're kind of seeing that with AI, but the marketplace of apps came out, the App Store. Now, you got the global internet, you're starting to see some applications that get you used to it, but we got this huge marketplace of applications coming.
Julie Choi
>> Yeah. What I'm really excited about with AI and with platforms like Cerebras that are enabling the fastest AI APIs for developers, I think we're going to see the first pure play, independent developer billionaire. This is going to be someone with an idea and they're going to just be able to build a business using AI and they're not going to need a marketplace. I think this is what we're walking into, is this type of incredibly exciting time for engineers and developers and just creative people.>> Yeah, and I think that DeepSeek style, I call it, is an innovation playbook not a one thing, because that's the cycle we're seeing. I think there'd be a lot of adoptions. I have to ask you, as the CMO at Cerebras, and obviously, the CMO Leaders Summit, everyone's watching and the conversation is how do I leverage AI for my business, for my career, for my team, for my stakeholders, partners? As people start to rethink their processes and their mechanisms, they're moving, but the game is still the same. You want to get customers to engage with you, do business with you, you want to retain those customers, you want to have a happy life while doing it. This is now, okay, but the process, the things might change. What do you see there? What models and what techniques are marketers looking at that you think should be front and center?
Julie Choi
>> I think that what we see with our enterprise customers, I'm just going to bring up one of them, Mayo Clinic for example. They're really embracing their core, which is, of course, patient care. They're embracing AI as a way to deliver the best healthcare the world has ever seen. We announced a model with them in January at JP Morgan Health Conference. Mayo Clinic has trained their own genomic AI model that will help diagnose rheumatoid arthritis much more effectively. Rheumatoid arthritis is close to my heart, because someone in my family has struggled with that disease. They did not have the right medicine for that disease for 10 years. It was like trying to treat it with the wrong disease.>> They were guessing, basically.
Julie Choi
>> Exactly. This happens 40% of the time. I think that enterprises, amazing leaders like Mayo Clinic, are so bold because they're embracing AI as an enabler. It is an innovation that can actually help patients, like my family member, find the right medicine not 10 years later. That's the kind of work that I think we as CMOs, we're so lucky to be able to do the storytelling with partners like that.>> On the application front, with you seeing successes out of the gate now, could you point to some use cases of other applications that are leveraging some of the horsepower and the speed that Cerebras and the new systems are building or are providing?
Julie Choi
>> My favorite example of application right now is, for sure, Le Chat. The Le Chat app, which I think I texted you over the weekend, it is blazing fast. This thing, if you haven't tried it out, please download it from the Apple Store or the Android store, it's Mistral Le Chat. I think that that application is now, it's going to overtake perhaps ChatGPT. It's already, I think, number one in the app store, and the reason is it takes less than a second to get really quality answers. That's one example of a consumer application that I think is going to really benefit from these advanced models, but also the speed of the inference of our hardware.>> Talk about the speed, because I've been talking on LinkedIn and you and I have been chatting on each other's threads on LinkedIn around this, because you brought up to me speed matters and you guys are fast, you guys do a lot of speed. We're starting to see that in other use cases, not just speed of performance of a chip, but whether it's an application, business process, outcomes are getting compressed in terms of time. They're accelerating time to value and the output of businesses. Talk about the speed here, because when you look at an LLM, how fast to the naked eye, how can I determine what's fast or slow besides having a counter on? Is it seconds? What's the speed calculation on the models?
Julie Choi
>> Absolutely. When we look at the model that Mistral has powering Le Chat in the side-by-side demo, their query completes in 1.3 seconds versus OpenAI ChatGPT, which is running on other hardware, is completing in 47 seconds. That speed difference actually, to the naked eye, maybe one second or 47 seconds->> That's 40 seconds, 47 is a lot of seconds.
Julie Choi
>> It's a lot of seconds, and that particular prompt is to write a game. It's like write a game of Snake using Rust. That prompt's a little bit more complex to show off the advantage of the speed. But when we think about the next wave beyond chatbots, we're talking about agents. Agentic AI, it's like a chain of multiple AI models. When we look at the next chapter for AI, it's all about Agentic AI, and we can't be waiting 47 seconds times five models. That'll lead to two minutes.>> Also, the time, I think, you gave is the most complex today, but that's reasoning. That's coming down the pike. I have to ask you, last time we met six months ago and sooner, inference was the big story and it still is. Has inference changed in terms of the definition? Now, we're seeing training and inference at the edge. Can you share any thoughts of what's going on around the dynamic between training and inference? There seems to be bundling together of that, where it's still inference, but it's not just inference. Can you share your thoughts on that dynamic?
Julie Choi
>> Yes, absolutely. I think inference is going to be a really important piece of accelerating the model's ability to reason. Some people call this test time inference, test time. When the model is being called to do its thing, the more time it has to think actually it can reason through the answer that it's about to give. Now, the first reasoning model, of course, was OpenAi's o1, this was released in September, and that was like a breakthrough. I remember my friend, Mitesh Agrawal, we were on the island, we were at this AI summit, this really cool thing that the New York Stock Exchange invited us to. He's like, "Julie, did you see o1?" We were so excited about, this is the beginning of what we see AI transforming into AGI, because it's the reasoning that the model needs to actually be super intelligent. That is all powered at time of inference. You do need more inference compute to basically increase the amount of reasoning that a model does at testing time. Inference computing is going to be absolutely vital, I think, to super intelligence. Then, of course, training is still important. We need to be able to pre-train these models from scratch and continue to have hardware that can support bigger and bigger models, because as we've learned over the past decade, bigger models are just better for AI.>> Then, the speed, the training and inference is the core, so relationships happens, is there any use cases that you're seeing now that Cerebras is winning constantly? Then, what use cases are coming? Because we see tasks now, you mentioned super intelligence. What tasks are you guys winning on now, and then, what's the next area you see Cerebras winning at?
Julie Choi
>> Absolutely. I think, right now, our business is booming among AI-first founders, and so like the startup community, they absolutely love the speed. They're using it for instant code apps, instant customer experience apps, instant avatars, so anything where low latency, high throughput is what they need, that's where we're winning. We're also seeing instant voice, and so anywhere where voice needs to be transcribed or delivered, that's also a great area. And then, reasoning. People who are really interested in reasoning, the phone is also ringing.>> It's interesting. You're a CMO, we're talking about how AI could be used for CMOs. You are an AI CMO in the sense you have the AI infrastructure, so you kind of have the cheat codes to how AI is going on, you can see it. What advice would you give other CMOs? Because you get a lot of experience and now you're in the front edge of the wave of innovation with Cerebras because you're leading the performance. How would you prepare a CMO or share your thoughts to a friend, colleague on how to get ready? Readiness, process, team formation, alignment to the management team, what would be your vision?
Julie Choi
>> I think I have the best job in the C-suite. To be honest with you, for me, CMO, it's, of course, chief marketing officer, but I also think of it as chief magic officer. I would just encourage my fellow marketers to really embrace the wonder of this AI era we're in. Embrace the tools. You can come up with a press release in three seconds using something like Le Chat now. If you can come up with a press release template in three seconds, you can customize it. You can get more news out than you ever could before. I just think that's really cool.>> You can make moments, magical moments. I have to say this has been a great week for you guys. Two big magical moments, Le Chat, Perplexity, those are big moments and it's a huge accomplishment, and so I want to congratulate you on that.
Julie Choi
>> Yes. Thank you.>> But it's just the beginning, there's more. What's next? What else do you have up your sleeve? What's coming down the road from Cerebras? Can you share any hints?
Julie Choi
>> Yeah. We're taking the show on the road. In February, we're actually going up to Seattle and everyone watching the show, please stop by if you're in town. We're having Llama Palooza, which is our owned event for Llama Development community, so we're doing that. We're having a big party in San Jose in March. I'll send you the info and just celebrate. We want to bring the community together and continue to learn. From a product perspective, we are definitely going to continue to ship the fastest inference for the best models so that developers can thrive.>> Final question for you, for the folks that got all excited or energized or aware of DeepSeek's moment, I call it the DeepSeek style, because they innovated in a clever way. There was no real technical breakthrough, it was just managing constraints. That woke up the enterprise, AI world, because ChatGPT was like the consumer, "Hey, look, things can happen." DeepSeek, I think, sends a path direction, "Hey, there's real innovation on how you handle your resources, your chips, how you configure things." Using Cerebras, you see Perplexity getting massive numbers, Le Chat to the top of the charts, you can configure things. That was, to me, is going to be the ChatGPT moment here in the enterprise, which is the DeepSeek-style moment that you can engineer or be clever in how you design your system.
Julie Choi
>> Absolutely. I think that the modular approach, the super resourceful approach that DeepSeek took, they were so efficient with how they used what they had. Then, actually, these are among the most brilliant algorithmic scientists in the world, the people at DeepSeek. We see their papers at NeurIPS every year. When you have that kind of algorithmic genius and you pair it with that resourcefulness that can make the most of anything, I think that's why we see this amazing innovation that they were able to share in open source. I think that's just the name of the game.>> Yeah. That's going to open up the developer community. That's why I think it's going to be a breakthrough on the developer side, because as you talk about the hardware, performance levers are AI infrastructure and software and the algorithms, right?
Julie Choi
>> Yes.>> I think the software has to run on something. Look at Perplexity. Fast, good as it is, great as it is, now, it's greater.
Julie Choi
>> Absolutely.>> Julie, great to have you on again. Great to see you. Congratulations on a spectacular week-
Julie Choi
>> Thank you.... >> for Cerebras and to your team. Pass on the good wish from theCUBE. Again, it's a great week for moments for you guys. Again, two huge accomplishments, very big, big wave. Thanks so much for coming on.
Julie Choi
>> Thanks, John.>> Okay. We are here at the CMO Leaders, because the world's changing. Fast speed is very critical, competitive advantage, but also user experience, AI infrastructure and software running together, bringing it all down here in theCUBE. Thanks for watching.