Soyoung Lee, co-founder and head of go-to-market for TwelveLabs, joins John Furrier of SiliconANGLE Media, Inc., on theCube at AWS Summit NYC 2025. The discussion centers around the innovative strides TwelveLabs takes in the field of multimodal artificial intelligence and video understanding, highlighting its impact across various industries and potential future applications.
In this insightful conversation, Lee shares insights on how TwelveLabs revolutionizes video understanding through their AI and multimodal foundation models. Hosted by Furrier of theCube, the discussion delves into the complexities of video data, the company's groundbreaking application programming interface product, and the numerous applications of intelligent video tasks within different sectors.
The discourse elaborates on key takeaways and innovative insights from the video. According to Lee, TwelveLabs breaks new ground by providing solutions to complex video data challenges, enabling more nuanced understanding and interaction with video content. It emphasizes the importance of AI in transforming industries such as media, sports, and safety, underscoring the significance of partnerships and technology integration, according to analyses provided by Furrier.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
AWS Mid-Year Leadership Summit 2025. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For AWS Mid-Year Leadership Summit 2025
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for AWS Mid-Year Leadership Summit 2025.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
AWS Mid-Year Leadership Summit 2025. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Sign in to gain access to AWS Mid-Year Leadership Summit 2025
Please sign in with LinkedIn to continue to AWS Mid-Year Leadership Summit 2025. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Soyoung Lee, TwelveLabs | AWS Summit NYC 2025
Soyoung Lee, co-founder and head of go-to-market for TwelveLabs, joins John Furrier of SiliconANGLE Media, Inc., on theCube at AWS Summit NYC 2025. The discussion centers around the innovative strides TwelveLabs takes in the field of multimodal artificial intelligence and video understanding, highlighting its impact across various industries and potential future applications.
In this insightful conversation, Lee shares insights on how TwelveLabs revolutionizes video understanding through their AI and multimodal foundation models. Hosted by Furrier of theCube, the discussion delves into the complexities of video data, the company's groundbreaking application programming interface product, and the numerous applications of intelligent video tasks within different sectors.
The discourse elaborates on key takeaways and innovative insights from the video. According to Lee, TwelveLabs breaks new ground by providing solutions to complex video data challenges, enabling more nuanced understanding and interaction with video content. It emphasizes the importance of AI in transforming industries such as media, sports, and safety, underscoring the significance of partnerships and technology integration, according to analyses provided by Furrier.
In this exclusive conversation from AWS Summit NYC, Soyoung Lee, Co-Founder and Head of Go-to-Market at TwelveLabs, joins theCUBE’s John Furrier to explore how the company is tackling one of AI’s most complex frontiers: multimodal video understanding. Twelve Labs is building powerful foundation models that can “see,” “hear,” and “understand” videos like humans do. With native support for video, audio and text inputs, the company enables enterprises to extract deep context from vast video libraries using APIs or integrations via Amazon Bedrock.
Lee shar...Read more
exploreKeep Exploring
What is Twelve Labs and what do they do?add
What analogy is being made regarding the function and impact of multimodal models in relation to OpenAI?add
What are some challenges organizations face when trying to extract insights from video content?add
>> Welcome back everyone to theCube, here on the show floor of AWS Summit 2025. I'm John Furrier, host of theCube. This is like a reinvent moment that's not the end of the year. It feels like the whole year happened the first half of the year. Whether it was at HumanX where I interviewed on a panel, this next guest Soyoung Lee, co-founder and head of go-to-market at Twelve Labs. Soyoung, good to see you again, this time on my show theCube.
Soyoung Lee
>> I know. We talk about it.>> Welcome back. Well, we were really on theCube but we hosted a great panel at HumanX earlier in the year. Phenomenal conference, huge industry participation. Pretty big panel. I think it could have been smaller, but you really stood out because we honed in on multimodal. You're doing a lot of stuff with video, which we love, but also it's a data problem that has an enterprise vibe to it because you're doing a lot of large scale. Explain to the folks what you guys are doing and the problems that you're solving.
Soyoung Lee
>> Yeah, absolutely. So Twelve Labs is a research and product company. We describe it as AI and I can see here and understand videos just like humans. The technical term for it is we build multimodal foundation models for video understanding and so natively the model itself can ingest video, it can listen to the raw audio, the dialogue, also see everything visually that's happening, but connect all of those contexts holistically over time as things change over video and be able to almost create memory and context from it. Our actual product would be APIs that then allow developers and enterprises to be able to run this model on their videos. So they would typically have large video libraries or constant stream of video and be able to implement these intelligent video tasks on top of the model's understanding of their content.>> I mean not to be a simpleton, but you're like open AI for video. Open AI exposes an API for people to work on it and then charge back. We're seeing a lot of people do that with some of these big models. You're doing that for multimodal media, right?
Soyoung Lee
>> Yes, and that's a really great analogy. I think the explosion of language models and how they understand and generate data and information has changed the way that we access knowledge and documents and communicate. Very similarly we see a lot of applications today now that are now being built on top of multimodal models that fundamentally give consumers and business buyers new ways to interact with their data.>> Scope the problem space that you're in because not to say that other models are simpler to execute, nothing's easy in AI, but is language, is language. Media multimodal is a whole other ballgame of technology science problems to solve, scope the order of magnitude of the complexity that you guys are taking on, it's not trivial.
Soyoung Lee
>> Absolutely. So video is so... Just simply plug video is just really complicated. So if you think about cameras and manufacturing lines, a sports game, a film, TV show, even consumer video advertisement, all of the things that you see and hear inside these videos are so different. The inside you need to pull, the use case you need to power from each is very different. And so for any organization whose business relies on understanding insight from video, they've tried a lot of different things. So we've had computer vision models that will actually cut up video into frames and it will take an image and extract objects and tag objects in a frame. So it will tell you this is a cup, but it won't tell you if I spilled water or if I drank the water or if I gave you this water. And that's context you need to accumulate as you understand time. The other approach would be you can also take video, you can do automated speech recognition to take a transcript, run that through a language model and if it's a dialogue only video, like a podcast, that could get you sufficient understanding of it. But the fact is there's a lot of things that go missing when you're missing out on the tone, the way that somebody's saying something or emotions and things like that. So to be able to holistically understand the video itself and then do that at scale across petabytes of a library, it's a very exciting opportunity.>> It's interesting you mentioned the cup example because that's actually illustrates how hard it is. I mean you got to know all variations of multiple dimensions, time, space. I mean it's a phenomenal... Okay, so this is not for the faint of heart. Where are you with the progress? Because explain where you are because a good point about podcasts, our videos are very simple. Your algorithms will kill this split problem. Oh yeah, John, easy to do. But when you add in full motion video cinematography, this isn't like, hey, Soyoung, make me a video of two people talking at the store. You're getting into real video action at scale. Where are you on the progress? Can you share where you're at? How are people using it? What's the low hanging fruit use cases? Because I'm sure it's popular.
Soyoung Lee
>> Yeah, I'd say a really common one that we're seeing where we've gotten a lot of adoption is we actually work with some of the largest media organizations in the world, whether they are studios, they're sports leagues, sports teams, and what all of these businesses have in common is that they have accumulated decades of incredibly valuable footage and this is an archive from which they continue to tell stories in new ways from existing assets. It's how they continue to maintain and grow their brand and find new revenue opportunities. The sheer liability to search that library and access, find things has been really challenging because traditionally you would've had to actually manually tag content as somebody is scrubbing through it and taking on that spreadsheet.>> Yeah, exactly. So let me ask you a question. I'm nerding out here for a second, but tagging at scale is hard. I can see that being instantly worked on. So check, but that's not it.
Soyoung Lee
>> No.>> There's more to it. What database do you... Graphs would be great. Graph of graphs, so you start getting into architecture. How do you as an engineering would think about it? Like, okay, I got to think, okay, I have moments in time, I have schedule, I have full motion video that's got multiple dimensions. How do you store? Is it graphs, is it multiple databases? What's the underpinnings? You don't have to give away the secret sauce, but what's some of the technology decisions you made?
Soyoung Lee
>> It's a great question. So just like human memory, we have everything that we see and experience we store in memory and for some magical reason we're able to visualize exact moments in our heads instantly. For our models we basically use vector embeddings and so we have Marengo, which is our embedding model that powers search and retrieval tasks. And so our Marengo model would take any data, it could be image, but most of the time our use cases are video related, so video, image, audio and even text. And it would create numerical representations that capture all of its context. What's really powerful about embeddings is that they're super light and you can store them in a vector database. I know S3 just announced a vector store which allows you to do everything from almost instantaneous semantic search or even image to video search. So if you have for instance, an image of a product and you're trying to identify that in a library of video, you could also use that image as a query.>> So embeds are light. Are you looking at S3 Vector?
Soyoung Lee
>> We have some... I mean we've started to actually last year partner with all of the major vector databases that were available. And the reason was because we wanted to make it really easy for our customers to use their existing vector databases and just bring in the model. And so they could take our embeddings and build a search engine super easily on their choice of vector database. Alternatively, we also provide an end-to-end search engine, which you can have the embeddings stored under the Twelve Lab side and the great thing about that is then you just call our search API when you pass the query or search query to us. We've done a lot of optimization on the search engine side of things, which means we can actually enable more accurate results to be achieved if the building of that search engine becomes more complex.>> Vector embeds are great for clustering too, understanding. The math is perfect.
Soyoung Lee
>> Absolutely. Some of the really cool use case that we're starting to see is recommendation engines, personalized recommendations for streaming platforms and things like that where it's not a retrieval task.>> Okay, so I have to ask you, obviously I mentioned I love video, we love video, we've been doing it for 15 years here on theCUBE, but we're dialogue, it's a little bit different, but how would someone engage with you? Do I just bring my data to the table? What if I already have vector embeds, if I'm embedding my content, do I just transfer the database over? Do you read embed? How does all that work?
Soyoung Lee
>> So we have some different ways. We just announced at Twelve Labs both of our models, our video language model, Pegasus, and our embedding model, Marengo, are now available through the Amazon Bedrock platform, which means you don't have to bring the data anywhere, as long as it's in S3 you can call the models through Bedrock so you can maintain where your data is while you're leveraging the models in a serverless way. The easiest way to start, we do have a playground that offers a UI for anyone who's technical or non-technical to come to the Twelve Labs site, sign up for a free account and just see what's possible. And if you're actually getting good results for your use case, then you can transition over to Bedrock and then deploy that at scale really easily.>> And what's the use case customer profile? They take the videos in and if they have data, they're training off you or they're using your model to train theirs and infer? I'm trying to understand specifically what's the first three steps to get going?
Soyoung Lee
>> First step, index videos. So in this step you're creating the embeddings from your videos. Second step you are integrating the search API. So you're building some workflow and that's dependent on the customer's use case. For instance, if you are a sports league that has a really large library of footage and you're trying to give your internal production teams access to that library to be able to freely search in natural language, then you connect that search bar within a content management platform with a search API. So that we've seen done in 20 minutes. And then the third is just you run it in production, you apply that at scale.>> All right, well let's talk about customers. How many customers do you have? Where are you guys on of the journey? You're heading the go-to-market, what's the focus? What are you optimizing for? What is some of the progress? Where's the progress bar on that?
Soyoung Lee
>> Yeah, so today we have over 50,000 developer users. In terms of our enterprise verticals, we focus on three. So one is media, sports and entertainment. The second is advertising and ad tech, which also video is becoming so critical, the understanding is becoming critical to that industry. And then the third is safety. How do we ensure safety, be able to help find evidence and things like that across a wide pool of content.>> Safety as in public safety, police departments?
Soyoung Lee
>> Yes.>> Things of that nature. Crime scenes, pictures, not safety as in cameras on the streets too or no?
Soyoung Lee
>> I mean it also does apply to personal homes as well. If you're trying to do this with your own system, and we do have->> That's a married house right there as I... He does my videos. He has three kids.
Soyoung Lee
>> Very cool. So those are our enterprise miracles and we also have an incredibly strong partnership team. The reason why partnership is so important to Twelve Labs is because we focus on building really powerful, easy to use models, but we understand that for customers it's actually not just the model that's needed. They may have their own evidence management platform, they may have their own content management platform, they may have their own hyperscaler like AWS where they store all of their content and stream it from. And so having already integrated into all of these ecosystems that exist makes it really easy for a customer to just be, oh let me just get this on Bedrock.>> It's a data thing for them. It's a data platform play. Exactly. So it's not like, hey download an app and free me of... It's really an enterprise engagement.
Soyoung Lee
>> Exactly. So that's how we've been able to reduce the friction of being an API. At the meantime, how do we now have velocity on the business value creation?>> All right, so put a plug in film, are you guys looking to hire? What kind of people do you hire? What's the mission statement? What's the north star? What's the culture like?
Soyoung Lee
>> Let me start with culture.>> Yeah, go with culture first. Dealers choice, pick your...
Soyoung Lee
>> We just had a company where I saw where we defined this. One is whatever it takes. And I think that really embodies, so we're about 110 people today, but every single person in the company really embodies that spirit. Other reasons, because we have very ambitious mission, which is let's build a AI that can understand the world just like humans and give humans and machines the ability to see and understand the world in new ways. And that can apply to so many things. So how do we get there from where we are as a smaller startup is the spirit will go above and beyond and having sheer grit and how can we be smart? How can we listen and derive insight from customers better and things like that.>> Well I'm super excited for you. I love your venture. Super technical too, and right on the cutting edge. I think multimodal is clearly the growth area, but it's hard. You're doing the work. Obviously big Amazon relationship here, customer of yours, good.
Soyoung Lee
>> We're super excited.>> You get credits from Amazon? You don't have to answer that question.
Soyoung Lee
>> We're always happy to talk about that.>> Love credits, it's like cash. Thank you for coming on theCUBE. Thanks and great to see you again?
Soyoung Lee
>> Yeah.>> I'll see you HumanX next year.
Soyoung Lee
>> I know.>> We might have theCUBE there next year, so we'll talk to Stefan, the team. I'm John Furrier here at the AWS Summit. This is also media week at our NYSE studio. We've been broadcasting all week at theCUBE there. Again, AI and cloud happening here. Thanks for watching.