Scott Stephenson of Deepgram, CEO and co-founder, joins theCUBE Research conversation hosted by John Furrier and Gabe Olave to explain Deepgram's approach to voice agents and Flux-TTS in the artificial intelligence era.
Stephenson explains Flux-TTS, the company's prior Flux speech-to-text momentum, adaptive speech models that preserve conversational context, low-latency runtimes and robotics-inspired system architectures that enable natural turn-aware voice interactions for developers and enterprises. They discuss developer workflows, integration patterns and enterprise deployment options.
Key takeaways include Deepgram's scale with over 200,000 developers and more than $100 million annual recurring revenue ARR. Stephenson highlights product differentiators: real-time contextual voice generation with sub-100 millisecond latency, high throughput and adaptive models that self-improve for specific use cases. They note flexible deployment options via a public application programming interface API, a virtual private cloud VPC and the Amazon Web Services Bedrock marketplace, and that a $200 starter credit enables rapid prototyping. A research-heavy organizational model accelerates production-grade voice infrastructure.
Watch for technical demonstrations and practical guidance on implementing Flux-TTS, designing low-latency voice agents and deploying voice solutions at enterprise scale.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Mixture of Experts Series. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Mixture of Experts Series.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For theCUBE + NYSE Wired: Mixture of Experts Series
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for theCUBE + NYSE Wired: Mixture of Experts Series.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
theCUBE + NYSE Wired: Mixture of Experts Series. If you don’t think you received an email check your
spam folder.
Sign in to theCUBE + NYSE Wired: Mixture of Experts Series.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to theCUBE + NYSE Wired: Mixture of Experts Series
Please sign in with LinkedIn to continue to theCUBE + NYSE Wired: Mixture of Experts Series. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Scott Stephenson, Deepgram
Scott Stephenson of Deepgram, CEO and co-founder, joins theCUBE Research conversation hosted by John Furrier and Gabe Olave to explain Deepgram's approach to voice agents and Flux-TTS in the artificial intelligence era.
Stephenson explains Flux-TTS, the company's prior Flux speech-to-text momentum, adaptive speech models that preserve conversational context, low-latency runtimes and robotics-inspired system architectures that enable natural turn-aware voice interactions for developers and enterprises. They discuss developer workflows, integration patterns and enterprise deployment options.
Key takeaways include Deepgram's scale with over 200,000 developers and more than $100 million annual recurring revenue ARR. Stephenson highlights product differentiators: real-time contextual voice generation with sub-100 millisecond latency, high throughput and adaptive models that self-improve for specific use cases. They note flexible deployment options via a public application programming interface API, a virtual private cloud VPC and the Amazon Web Services Bedrock marketplace, and that a $200 starter credit enables rapid prototyping. A research-heavy organizational model accelerates production-grade voice infrastructure.
Watch for technical demonstrations and practical guidance on implementing Flux-TTS, designing low-latency voice agents and deploying voice solutions at enterprise scale.
>> John Furrier, host of theCUBE here in theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is our Mixture of Experts series where we talk to the leaders, making it happen in the era of AI, all the technology innovation that's happening. We are in a market shift. We've never seen this kind of shift before. And of course, AI is the powered the AI infrastructure, moving up and down the stack and a lot of impact to how we work and live. Scott Stephenson, who is the CEO and co -founder of Deepgram, a very innovative company, innovating on the voice side, which is the interface now, as we see AI get moved beyond text into voice. And obviously behind that, you got a lot more going on in the stack with agents, et cetera. Scott, great to see you again. Last time we chatted we were at re:Invent. Thanks for coming back on theCUBE.
Scott Stephenson
>> Yeah, absolutely, great to see you again.
John Furrier
>> I got to ask you, so I love voice, but the voice piece has become so popular. You look at the adoption, even on the large frontier models, voice is hugely important. We're seeing demos all the time of voice activating coding. So obviously we've crossed over from text to voice, but the user interface, some are calling it headless. I kind of don't like the term, but I get it. but we're now in a new era of interfaces. This is what you guys do. Just your take real quick on how you guys fit into that.
Scott Stephenson
>> Yeah, we're seeing a lot of deployment of this technology. So we got over 2 ,000 products that are built on top of Deepgram, including Sierra, Decagon, a bunch of products from AWS, Genesys, NICE, LiveKit, Pipecat, just so much adoption in the voice agent space. And then also as a voice interface to just make it easier for people to use the products that they use every day, rather than you having to figure out how to adapt to a computer, the computer can just talk to you the way that you normally talk.
John Furrier
>> Well folks watching can check out the AWS re:Invent last year's interview. A lot's changed. Talk about the momentum, because you have news hitting today. New products, updates, announcements, and momentum numbers.
Scott Stephenson
>> Yeah, yeah.
John Furrier
>> Tell us what the hard news is.
Scott Stephenson
>> Yeah, so we have over 200 ,000 developers using Deepgram. We're over 100 million ARR now, growing very quickly. And we're announcing a new model today, which is Deepgram's latest voice generation model called Flux TTS. So we had a very successful model launch end of last year, Flux Speech -to -Text, which has become the default standard for people that are trying to understand what's happening in voice in real time. And now voice agents are becoming so popular, they want the same type of model, but for voice generation. So a model that understands context over turns is just a voice that you love to talk to. And so that's the model that we're releasing today.
John Furrier
>> And this is what people can see now with some of these prototypes and products out there where it's basically a computer-generated person or a voice from text.
John Furrier
>> Yeah.
John Furrier
>> And sometimes they're a little bit jittery. It stutters or whatever. You can tell it's a machine.
John Furrier
>> Yeah.
John Furrier
>> You guys are taking it to the next level. Explain the tech behind that.
Scott Stephenson
>> Yeah, so when you have a voice conversation going on, there's a lot of conversational negotiations. So whose turn is it to talk? Is it your turn? My turn? We give all these little cues. And what we would say in Deepgram is this is just extra context, basically, right? So when people are thinking about a voice agent, they might think, well, I'm going to take the voice and I'm going to turn it into text. And then I'm going to feed that text to an LLM. And then that LLM is going to generate text. And then I'm going to feed that to text-to-speech. But that's missing all the context, right? So you'd like to know, well, is it my turn? Is it your turn? what did I say previously? Is the conversation getting really positive? Is it getting really negative? Do I need to pull it back, et cetera? So these are all conversation, negotiation, context cues that are now included in the model, and it's all working in real time in Flux. So all of this works at lower than, as low as 80 milliseconds.
John Furrier
>> Talk about the AI impact, because natural language processing would parse some strings, turn it into voice, we've seen that old school way. There's demos of it could read a Substack or a SiliconANGLE, read the post. There's more there. Talk about the application and where it goes. Where is the intelligence piece? What does the model do? What are some of the use cases? Does it reason? take us through. Okay, if I have speech -to -text and text -to -speech, what are some of the use cases?
Scott Stephenson
>> Yeah, it's thinking. It's saying, hey, I as a model just said this five turns ago. Should I adjust the way that I say my next thing? You just said this thing the way that you said it. well, how should I adjust my response? This is something that humans are doing the whole time. The previous generation of TTS up until this date, that's not what it did. It just took the text that you put in, and you hit enter, you go, and then it says whatever it's going to say. So typically it sounded a lot like an audio book or something like that, right? Like it was performing.
John Furrier
>> It's obvious.
Scott Stephenson
>> Yeah, and you're wait a second, that isn't like the way that I would want it to be said. Well, you need real -time context feeding into the model in order to get it to speak like a human would or in a way that you would love to talk with it.
John Furrier
>> this is essentially why I love the application of robotics. You're starting to see a lot of kind of conversational rapport that isn't just chatbot -based.
Scott Stephenson
>> Yes.
John Furrier
>> It's got to actually think like a human. What's required to do that? you've got to nail your piece. So you've got the developers banging away on this. What are they doing? What are some of the things you see emerging from this? Because I can imagine, just for me, someone interacting with a post on SiliconANGLE and maybe asking a question to the writer.
John Furrier
>> Yeah.
John Furrier
>> which could be back end. AI could write that code maybe down the road. So I could see all kinds of machinations. What are some of the things people are going to do with this?
Scott Stephenson
>> Yeah, so the way that these systems are developing is that there is a sort of extremely real -time system that just has to be connected with the individual that you're talking to and understanding their needs and responding really quickly. And in robotics, there is an equivalent here. This is called System 1. This is what they would call it in robotics. But then also there's a slower system that's more asynchronous sort of in the background. Maybe you have a queue, maybe you're reasoning, et cetera. And so you have two side -by -side systems that are running. And so the models that we're talking about today, they're running in that System 1. They're like figuring out everything that's going on. But it still allows you to use maybe a larger reasoning model in the background or to pull information from your CRM in order to populate the conversation and make it more aware of the context from your previous calls and that type of thing. So I would say it's still early days in how all this is developing. But robotics, like you mentioned, is a great place to sort of look for, hey, if it's working on a bipedal robot walking around, accepting instructions and moving around, that's a very crucial job that it's performing. You can use a lot of those architectures in voice agents as well in order to make it work and feel natural and fluent.
John Furrier
>> Take me through some of the engagement models. I'm a developer, what do I do? How do I engage? What are some of the commercial partnerships? How do, are there credits? I can see this becoming very popular quickly. I know you've got great traction. It's already over 100 million ARR. But even more developer traction, I think, will come on board. What's the, how do I engage?
Scott Stephenson
>> Yeah, so we have an API. We have a public API that anybody can sign up to. You can go to Deepgram.com. You can check out the product that's there. We have speech -to -text. We have text -to -speech. We have a full voice agent. We have a playground. We have a way for you to just test all of this out. So you don't have to program in order to see how well it works. But once you want to build something that's a demo or scale it, then you can just vibe code something with Claude Code or Codex or any of those things and say, hey, use Deepgram. Here's my API key. And then go build something with it and get something great. That'll work with maybe Pipecat or LiveKit or something like that. But we provide $200 of free credit. People can just sign up now and utilize that. And it's enough to build a lot and get started and really understand the power of this.
John Furrier
>> So they can actually get their hands dirty and build something quickly with the $200 credit. All right. You mentioned Amazon. We also met at Amazon. I know you have a relationship with Amazon. If I have stuff in Amazon, I have a lot of data laying in there. Let's just say I have a lot of video.
John Furrier
>> Yeah.
John Furrier
>> Okay, with a lot of text and a lot of speech. How do I connect that? because this is like I want to get the prototype. I don't want to have to kind of do all this heavy lifting.
John Furrier
>> Yeah.
John Furrier
>> And what if I have a data lake on, say, Databricks or something on-premises? What is the data connection? Is it just an API ingestion? What's the interface? How does that all work?
Scott Stephenson
>> Yeah, so you can use the public endpoint for API ingestion. You could use it in your own VPC like in AWS or GCP or other partners that we have. and then also, we have great market partnerships with the infrastructure providers, where we partner with them in order to provide our real time models in their platform. So yes you could run it in your own virtual private cloud, that your VMs that you have with them, but also they have something that they pioneered with the LLMs and text models, but you can use that for voice now as well, where your data stays there, the models are run by AWS, so they're already your trusted platform provider. And so, yeah, we partner with them.
John Furrier
>> So do I hit your API for all of this, or can I bring your model to my data? Can I build my own model?
Scott Stephenson
>> Yeah.
John Furrier
>> Take me through some of those options.
Scott Stephenson
>> Our models are adaptive, so this is one of our differentiators at Deepgram. So, sure, you can use Deepgram in Bedrock, the general models in AWS, or you could use the general models in the API, but you could also work with our team, so we have a great success team at Deepgram and Data Team where if you say, hey, I want the models to automatically improve and self -adapt to what's happening in my voice agents or my use case, then the models will do that. We're unique in this offering at Deepgram.
John Furrier
>> So you'll meet the users where they are.
John Furrier
>> Yeah.
John Furrier
>> Whatever their environment is.
Scott Stephenson
>> Yeah, it's not just one size fits all, take it or leave it, we can meet them where they are.
John Furrier
>> Talk about the product that you guys are announcing today. What's the secret sauce there? What's the real upside from your perspective? knowing what you know. Someone's going, okay, this sounds pretty cool.
John Furrier
>> Yeah.
John Furrier
>> speech to text, text to speech. I got both now. I got both sides of the coin. What's the big upside and the secret sauce?
Scott Stephenson
>> Yeah.Yeah, sounding good is you have to be this tall to ride, basically. So the voices have to sound good. The voices sound amazing. We did third -party testing for naturalness and expressiveness. We topped the charts on those. But that's an entry point. You have to trust it through that. But when you're a B2B user building a product that's going to scale to hundreds of thousands of users, millions of users, then you need low latency, you need high throughput, you need reliability, you need the ability to ship it in your VPC or use it in the cloud, scale up, scale down, to follow the sun as it moves around the world. This is the product, the underlying runtime product in Deepgram that allows you to run it. So Flux TTS, yes, it has a great voice, low latency, high throughput, it sounds great, people love to talk to it, but it also has that runtime under it that's supporting it.
John Furrier
>> What's the language support?
Scott Stephenson
>> Yep.
John Furrier
>> Multi -language. You said follow the sun. Imagine, okay, Asia Pacific is a tough market to crack.
John Furrier
>> Yep.
John Furrier
>> Are you supporting multiple languages?
Scott Stephenson
>> Yeah, right now it's five languages and accents, focusing on English and the different accents, but we'll be expanding that very rapidly. We typically aim to support over 100 languages with a rapid follow -on, So we get it great with English, and then we roll it out to everything else in a very short follow -up.
John Furrier
>> Is there a lot of training involved? I'm just curious. when you perfect the model, take us through some of the process. You've got to train the model.
John Furrier
>> Yeah.
John Furrier
>> And then you have the inference side of it and then the reasoning. Take us through that process.
Scott Stephenson
>> Yeah, so there's an amazing data team at Deepgram. And what we do is learn from the world around us. So we take the best from open source,
Scott Stephenson
>> sure.
Scott Stephenson
>> but we also invent internally in Deepgram, and then we find amazing voice actors and voices that are out there and put them in a studio and get them in their natural state and pick voices that people love and then go down that path to capture the data. That's just step one, where you get really good data, and then you get that now for multiple accents, multiple languages, et cetera. But then at Deepgram, we have a frontier research team that then takes that data and explores so many different model variants. I lose track of how many we have. But we have to put them through the wringer. How much compute does it use? Does it hit the right cost curve? Does it have low latency? Can it do high throughput? Will it ship well on these GPUs, et cetera, that type of thing? So the team goes through that. And we actually, at Deepgram, we have an automated process that does this. If you had to do all those different model variants to find the most efficient model to do it, it would be kind of insane.
John Furrier
>> Well, you're hyper -focused on voice and text. That's really one good start. Explain, you mentioned research. How are you guys organized? You put a plug in for the company. How big is the company? Obviously, you've got great revenue traction. You have a labs team. You said you've got a data team. How are you guys laid out as a company? And what are the key focuses for the second half of the year? I'm sure re:Invent's coming up and all these big events, news is coming down the pike.
Scott Stephenson
>> Yeah, about 60 to 70 % of the company is in R &D in some way, shape, or form. So whether it's our research team, machine learning engineering, our engineering product deployment team, product team, our applied engineering. So a large chunk of the company is based on building an amazing product and then deploying it into the world. And we have a great go -to -market team at DPM and success team as well. But yeah, we are a research company that productizes as fast as possible. So we're doing frontier research at Deepgram in audio. We're the first to bring end-to-end deep learning to speech. We did that 11 years ago when we started as a company. And now, voice is having its moment in the last year, and we're in a really great position.
John Furrier
>> The world's come to your doorstep. you guys had the vision, a great vision. if you just connect the dots, just a little early, but now it's prime time. And everyone's using the new models. you're seeing chat interface all the time. it's voice.
John Furrier
>> Yep.
John Furrier
>> This is game over, in my opinion.
Scott Stephenson
>> Yeah, if there's a good voice agent out there and you're like, wow, I like interacting with this, it's probably powered by Deepgram.
John Furrier
>> Yeah. Scott, thanks for coming on. I really appreciate it. Love what you guys do. Continue to build faster.
Scott Stephenson
>> Yeah.
John Furrier
>> And then we'll be customers.
Scott Stephenson
>> Let's do it.
John Furrier
>> Great to see you.Thanks for coming on.
Scott Stephenson
>> Thanks.
John Furrier
>> All right. This is the kind of innovation the AI native culture is seeing. Rapid deployment, new kinds of methods of interface, and the AI infrastructure is being built out. We saw all kinds of new announcements, because of more funding as the AI infrastructure at cloud scale continues to increase, the speed at which these things can execute is going to go faster and faster. This is the velocity of the AI era, and of course we're doing our part here in theCUBE. I'm John Furrier, host of theCUBE. Thanks for watching.
>> John Furrier, host of theCUBE here in theCUBE's NYSE studio. Of course, we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is our Mixture of Experts series where we talk to the leaders, making it happen in the era of AI, all the technology innovation that's happening. We are in a market shift. We've never seen this kind of shift before. And of course, AI is the powered the AI infrastructure, moving up and down the stack and a lot of impact to how we work and live. Scott Stephenson, who is the CEO and co -founder of Deepgram, a very innovative company, innovating on the voice side, which is the interface now, as we see AI get moved beyond text into voice. And obviously behind that, you got a lot more going on in the stack with agents, et cetera. Scott, great to see you again. Last time we chatted we were at re:Invent. Thanks for coming back on theCUBE.
Scott Stephenson
>> Yeah, absolutely, great to see you again.
John Furrier
>> I got to ask you, so I love voice, but the voice piece has become so popular. You look at the adoption, even on the large frontier models, voice is hugely important. We're seeing demos all the time of voice activating coding. So obviously we've crossed over from text to voice, but the user interface, some are calling it headless. I kind of don't like the term, but I get it. but we're now in a new era of interfaces. This is what you guys do. Just your take real quick on how you guys fit into that.
Scott Stephenson
>> Yeah, we're seeing a lot of deployment of this technology. So we got over 2 ,000 products that are built on top of Deepgram, including Sierra, Decagon, a bunch of products from AWS, Genesys, NICE, LiveKit, Pipecat, just so much adoption in the voice agent space. And then also as a voice interface to just make it easier for people to use the products that they use every day, rather than you having to figure out how to adapt to a computer, the computer can just talk to you the way that you normally talk.
John Furrier
>> Well folks watching can check out the AWS re:Invent last year's interview. A lot's changed. Talk about the momentum, because you have news hitting today. New products, updates, announcements, and momentum numbers.
Scott Stephenson
>> Yeah, yeah.
John Furrier
>> Tell us what the hard news is.
Scott Stephenson
>> Yeah, so we have over 200 ,000 developers using Deepgram. We're over 100 million ARR now, growing very quickly. And we're announcing a new model today, which is Deepgram's latest voice generation model called Flux TTS. So we had a very successful model launch end of last year, Flux Speech -to -Text, which has become the default standard for people that are trying to understand what's happening in voice in real time. And now voice agents are becoming so popular, they want the same type of model, but for voice generation. So a model that understands context over turns is just a voice that you love to talk to. And so that's the model that we're releasing today.
John Furrier
>> And this is what people can see now with some of these prototypes and products out there where it's basically a computer-generated person or a voice from text.
John Furrier
>> Yeah.
John Furrier
>> And sometimes they're a little bit jittery. It stutters or whatever. You can tell it's a machine.
John Furrier
>> Yeah.
John Furrier
>> You guys are taking it to the next level. Explain the tech behind that.
Scott Stephenson
>> Yeah, so when you have a voice conversation going on, there's a lot of conversational negotiations. So whose turn is it to talk? Is it your turn? My turn? We give all these little cues. And what we would say in Deepgram is this is just extra context, basically, right? So when people are thinking about a voice agent, they might think, well, I'm going to take the voice and I'm going to turn it into text. And then I'm going to feed that text to an LLM. And then that LLM is going to generate text. And then I'm going to feed that to text-to-speech. But that's missing all the context, right? So you'd like to know, well, is it my turn? Is it your turn? what did I say previously? Is the conversation getting really positive? Is it getting really negative? Do I need to pull it back, et cetera? So these are all conversation, negotiation, context cues that are now included in the model, and it's all working in real time in Flux. So all of this works at lower than, as low as 80 milliseconds.
John Furrier
>> Talk about the AI impact, because natural language processing would parse some strings, turn it into voice, we've seen that old school way. There's demos of it could read a Substack or a SiliconANGLE, read the post. There's more there. Talk about the application and where it goes. Where is the intelligence piece? What does the model do? What are some of the use cases? Does it reason? take us through. Okay, if I have speech -to -text and text -to -speech, what are some of the use cases?
Scott Stephenson
>> Yeah, it's thinking. It's saying, hey, I as a model just said this five turns ago. Should I adjust the way that I say my next thing? You just said this thing the way that you said it. well, how should I adjust my response? This is something that humans are doing the whole time. The previous generation of TTS up until this date, that's not what it did. It just took the text that you put in, and you hit enter, you go, and then it says whatever it's going to say. So typically it sounded a lot like an audio book or something like that, right? Like it was performing.
John Furrier
>> It's obvious.
Scott Stephenson
>> Yeah, and you're wait a second, that isn't like the way that I would want it to be said. Well, you need real -time context feeding into the model in order to get it to speak like a human would or in a way that you would love to talk with it.
John Furrier
>> this is essentially why I love the application of robotics. You're starting to see a lot of kind of conversational rapport that isn't just chatbot -based.
Scott Stephenson
>> Yes.
John Furrier
>> It's got to actually think like a human. What's required to do that? you've got to nail your piece. So you've got the developers banging away on this. What are they doing? What are some of the things you see emerging from this? Because I can imagine, just for me, someone interacting with a post on SiliconANGLE and maybe asking a question to the writer.
John Furrier
>> Yeah.
John Furrier
>> which could be back end. AI could write that code maybe down the road. So I could see all kinds of machinations. What are some of the things people are going to do with this?
Scott Stephenson
>> Yeah, so the way that these systems are developing is that there is a sort of extremely real -time system that just has to be connected with the individual that you're talking to and understanding their needs and responding really quickly. And in robotics, there is an equivalent here. This is called System 1. This is what they would call it in robotics. But then also there's a slower system that's more asynchronous sort of in the background. Maybe you have a queue, maybe you're reasoning, et cetera. And so you have two side -by -side systems that are running. And so the models that we're talking about today, they're running in that System 1. They're like figuring out everything that's going on. But it still allows you to use maybe a larger reasoning model in the background or to pull information from your CRM in order to populate the conversation and make it more aware of the context from your previous calls and that type of thing. So I would say it's still early days in how all this is developing. But robotics, like you mentioned, is a great place to sort of look for, hey, if it's working on a bipedal robot walking around, accepting instructions and moving around, that's a very crucial job that it's performing. You can use a lot of those architectures in voice agents as well in order to make it work and feel natural and fluent.
John Furrier
>> Take me through some of the engagement models. I'm a developer, what do I do? How do I engage? What are some of the commercial partnerships? How do, are there credits? I can see this becoming very popular quickly. I know you've got great traction. It's already over 100 million ARR. But even more developer traction, I think, will come on board. What's the, how do I engage?
Scott Stephenson
>> Yeah, so we have an API. We have a public API that anybody can sign up to. You can go to Deepgram.com. You can check out the product that's there. We have speech -to -text. We have text -to -speech. We have a full voice agent. We have a playground. We have a way for you to just test all of this out. So you don't have to program in order to see how well it works. But once you want to build something that's a demo or scale it, then you can just vibe code something with Claude Code or Codex or any of those things and say, hey, use Deepgram. Here's my API key. And then go build something with it and get something great. That'll work with maybe Pipecat or LiveKit or something like that. But we provide $200 of free credit. People can just sign up now and utilize that. And it's enough to build a lot and get started and really understand the power of this.
John Furrier
>> So they can actually get their hands dirty and build something quickly with the $200 credit. All right. You mentioned Amazon. We also met at Amazon. I know you have a relationship with Amazon. If I have stuff in Amazon, I have a lot of data laying in there. Let's just say I have a lot of video.
John Furrier
>> Yeah.
John Furrier
>> Okay, with a lot of text and a lot of speech. How do I connect that? because this is like I want to get the prototype. I don't want to have to kind of do all this heavy lifting.
John Furrier
>> Yeah.
John Furrier
>> And what if I have a data lake on, say, Databricks or something on-premises? What is the data connection? Is it just an API ingestion? What's the interface? How does that all work?
Scott Stephenson
>> Yeah, so you can use the public endpoint for API ingestion. You could use it in your own VPC like in AWS or GCP or other partners that we have. and then also, we have great market partnerships with the infrastructure providers, where we partner with them in order to provide our real time models in their platform. So yes you could run it in your own virtual private cloud, that your VMs that you have with them, but also they have something that they pioneered with the LLMs and text models, but you can use that for voice now as well, where your data stays there, the models are run by AWS, so they're already your trusted platform provider. And so, yeah, we partner with them.
John Furrier
>> So do I hit your API for all of this, or can I bring your model to my data? Can I build my own model?
Scott Stephenson
>> Yeah.
John Furrier
>> Take me through some of those options.
Scott Stephenson
>> Our models are adaptive, so this is one of our differentiators at Deepgram. So, sure, you can use Deepgram in Bedrock, the general models in AWS, or you could use the general models in the API, but you could also work with our team, so we have a great success team at Deepgram and Data Team where if you say, hey, I want the models to automatically improve and self -adapt to what's happening in my voice agents or my use case, then the models will do that. We're unique in this offering at Deepgram.
John Furrier
>> So you'll meet the users where they are.
John Furrier
>> Yeah.
John Furrier
>> Whatever their environment is.
Scott Stephenson
>> Yeah, it's not just one size fits all, take it or leave it, we can meet them where they are.
John Furrier
>> Talk about the product that you guys are announcing today. What's the secret sauce there? What's the real upside from your perspective? knowing what you know. Someone's going, okay, this sounds pretty cool.
John Furrier
>> Yeah.
John Furrier
>> speech to text, text to speech. I got both now. I got both sides of the coin. What's the big upside and the secret sauce?
Scott Stephenson
>> Yeah.Yeah, sounding good is you have to be this tall to ride, basically. So the voices have to sound good. The voices sound amazing. We did third -party testing for naturalness and expressiveness. We topped the charts on those. But that's an entry point. You have to trust it through that. But when you're a B2B user building a product that's going to scale to hundreds of thousands of users, millions of users, then you need low latency, you need high throughput, you need reliability, you need the ability to ship it in your VPC or use it in the cloud, scale up, scale down, to follow the sun as it moves around the world. This is the product, the underlying runtime product in Deepgram that allows you to run it. So Flux TTS, yes, it has a great voice, low latency, high throughput, it sounds great, people love to talk to it, but it also has that runtime under it that's supporting it.
John Furrier
>> What's the language support?
Scott Stephenson
>> Yep.
John Furrier
>> Multi -language. You said follow the sun. Imagine, okay, Asia Pacific is a tough market to crack.
John Furrier
>> Yep.
John Furrier
>> Are you supporting multiple languages?
Scott Stephenson
>> Yeah, right now it's five languages and accents, focusing on English and the different accents, but we'll be expanding that very rapidly. We typically aim to support over 100 languages with a rapid follow -on, So we get it great with English, and then we roll it out to everything else in a very short follow -up.
John Furrier
>> Is there a lot of training involved? I'm just curious. when you perfect the model, take us through some of the process. You've got to train the model.
John Furrier
>> Yeah.
John Furrier
>> And then you have the inference side of it and then the reasoning. Take us through that process.
Scott Stephenson
>> Yeah, so there's an amazing data team at Deepgram. And what we do is learn from the world around us. So we take the best from open source,
Scott Stephenson
>> sure.
Scott Stephenson
>> but we also invent internally in Deepgram, and then we find amazing voice actors and voices that are out there and put them in a studio and get them in their natural state and pick voices that people love and then go down that path to capture the data. That's just step one, where you get really good data, and then you get that now for multiple accents, multiple languages, et cetera. But then at Deepgram, we have a frontier research team that then takes that data and explores so many different model variants. I lose track of how many we have. But we have to put them through the wringer. How much compute does it use? Does it hit the right cost curve? Does it have low latency? Can it do high throughput? Will it ship well on these GPUs, et cetera, that type of thing? So the team goes through that. And we actually, at Deepgram, we have an automated process that does this. If you had to do all those different model variants to find the most efficient model to do it, it would be kind of insane.
John Furrier
>> Well, you're hyper -focused on voice and text. That's really one good start. Explain, you mentioned research. How are you guys organized? You put a plug in for the company. How big is the company? Obviously, you've got great revenue traction. You have a labs team. You said you've got a data team. How are you guys laid out as a company? And what are the key focuses for the second half of the year? I'm sure re:Invent's coming up and all these big events, news is coming down the pike.
Scott Stephenson
>> Yeah, about 60 to 70 % of the company is in R &D in some way, shape, or form. So whether it's our research team, machine learning engineering, our engineering product deployment team, product team, our applied engineering. So a large chunk of the company is based on building an amazing product and then deploying it into the world. And we have a great go -to -market team at DPM and success team as well. But yeah, we are a research company that productizes as fast as possible. So we're doing frontier research at Deepgram in audio. We're the first to bring end-to-end deep learning to speech. We did that 11 years ago when we started as a company. And now, voice is having its moment in the last year, and we're in a really great position.
John Furrier
>> The world's come to your doorstep. you guys had the vision, a great vision. if you just connect the dots, just a little early, but now it's prime time. And everyone's using the new models. you're seeing chat interface all the time. it's voice.
John Furrier
>> Yep.
John Furrier
>> This is game over, in my opinion.
Scott Stephenson
>> Yeah, if there's a good voice agent out there and you're like, wow, I like interacting with this, it's probably powered by Deepgram.
John Furrier
>> Yeah. Scott, thanks for coming on. I really appreciate it. Love what you guys do. Continue to build faster.
Scott Stephenson
>> Yeah.
John Furrier
>> And then we'll be customers.
Scott Stephenson
>> Let's do it.
John Furrier
>> Great to see you.Thanks for coming on.
Scott Stephenson
>> Thanks.
John Furrier
>> All right. This is the kind of innovation the AI native culture is seeing. Rapid deployment, new kinds of methods of interface, and the AI infrastructure is being built out. We saw all kinds of new announcements, because of more funding as the AI infrastructure at cloud scale continues to increase, the speed at which these things can execute is going to go faster and faster. This is the velocity of the AI era, and of course we're doing our part here in theCUBE. I'm John Furrier, host of theCUBE. Thanks for watching.