Evolving the Digital Landscape with Twelve Labs’ Jae Lee
Jae Lee, CEO of Twelve Labs, discusses integrating video, audio, image, and text in AI models and its implications for search and content understanding.
Transcript
This is Textron tv. Hi everyone. Welcome back here to Techstrong tv.
Got a new company I wanna introduce you to. Uh, and, and, and also their co-founder and CEO. Say hello to Jay Lee.
Jay is the co-founder and CEO of a company called 12 Labs, like a dozen 12 labs. Hey, Jay. Jay, welcome to Tech Drunk tv.
How are you, Matt? Hi, Ellen. I'm doing good.
Thanks for having me. Thank you. Excited to jam.
Yeah, we're, we're excited to have you on here, man. Let's see what we could do. So Jay, you're based in, in San Francisco?
That's right. And as I mentioned, you're the co-founder and CEO of 12 Labs. We're gonna talk about 12 Labs, but before we dive into 12 Labs, let's talk a little bit about Jay.
You know, how did you, uh, come to do, to do this? Yeah. Um, I come from kind of like, you know, traditional AI research background.
Uh, very geeky kid from, from early on. Um, so a little brief background around like how I grew up, I guess. I, um, was born in Seoul, uh, moved to Knoxville, Tennessee.
Um, wow. As I, um, yeah, I joined, uh, my uncle at the time, so he, he was getting his PhD, uh, in Beijing Statistics when Oh, Ut is in Knoxville. Yeah.
Yeah. So UT is, is University of Tennessee, only in Tennessee. But, yeah.
Uh, so, you know, I, early on, you know, having kind of, uh, spent a lot of time with a guy who was going through his PhD programs, I, you know, started, uh, kind of thinking about like statistics and what it means to kind of like build some sort of an algorithm or model that can kind of understand, understand our surroundings right. From very early on. Um, so, you know, my passion for learning from data kind of came about from, from the early ages, right.
Just kind of like looking over my uncle's shoulders during Met lab coding and so on and so forth. Kind of like, even though I don't understand very cool any of it, um, Uhhuh, but, uh, I would just kind of look at his textbooks. Um, but yeah, I really geeked out on computer science when I, uh, went to a boarding school in New Hampshire.
Uh, I'm, you know, you're, you know, from New York, so Yeah. Yeah. Northern New England area quite well.
I really geeked out and then went to Berkeley, um, uc, Berkeley for, for computer science. And, uh, got to work with, uh, some really cool folks, uh, that kind of helped build the, the modern AI industry. Right.
Um, yeah. And, uh, after Berkeley, I worked for a couple companies, but, um, you know, decided to go back and, and serve the country that I was born in. Um, so I, uh, joined the Korean Cyber Command, uh, it, which is very similar to Israeli's, you know, unit 8,200 or US Cyber 80.
Yeah, Yeah. Uh, and, uh, there, uh, we've worked on multimodal pd, understanding research, uh, so we can get into what that is. Uh, but I met my co-founders who are very similar, um, as me, uh, very geeky also looking towards, you know, going into, uh, uh, careers in, in academia, becoming a researcher or professor, uh, in ai.
Um, yeah. And, uh, you know, before I know it, we were working together like day and night and, uh, you know, starting 12 labs and, and building this company out properly. Wow.
What an American story, man. Right? This is, this is, you know, not to get all political or anything, but you know what, but this is that classic Horatio algebra cla, you know, you were afforded an opportunity to let your mind expand.
Yeah. You know, even before you knew what you were expanding into. Right.
And, you know, and, and then the world, the world turns and you just find your play. You know, you find your style, you, you find yourself in a place where, you know, there's something big going on AI in this case, right. Kinda world changing, maybe.
That's right. And, You know, and, and then you, you look at that, but you do your service, you do your, you know, you do the right thing, but then you see these opportunities knock and, and that, and that's really, I think, the American part of it. Right.
You see an opportunity, you, you, you take your shot. Yeah. Not everyone winds up to be Bill Gates or Elon Musk or Jeff Bezos, but you take your shot and you, and you follow your, your heart and your passion and, and that's exciting.
Yeah. Also, wanting to just push the boundary, uh, yeah. In AI and do our share of, of work.
Um, and hopefully it gets recognized. And that's why we used Ab you know, what, we were talking off car, I was telling little bit. So I've been talking to Entrepr, I've been doing media for almost 10 years now, but I, I've been in tech 30 years.
During that time, I used, for instance, I was working in a company, we were in the, uh, SoftBank Venture Capital tsu. So I was doing a lot of business development with like all the, all the, uh, portfolio companies. I, so I came in con I I met thousands of co-founders and founders, entrepreneurs, was out in Boulder when, uh, Techstars got started.
And I've never met a successful entrepreneur founder who didn't think that, at least in some small way, if not a big way, what they were working on, what they were doing was somehow gonna change the world. We're somehow gonna make the world better. That's right.
If you don't have that passion, get out of the kitchen. Right. Because it's hard.
No one says it's easy. Right. I've done it.
I, I've started, you know, co-founded probably five different companies, you know, venture back and work and, and, um, there are times where you say, why am I doing this to myself? I can go get a job and make good money. Right?
Um, yeah. But you do it because you have that passion. Yeah.
It's excruciatingly painful. Um, yeah. Sometimes, but to do it, It's just for money, right?
Yeah. So it's sometimes the money's, you don't do it just for the money. You can't, you're not gonna be successful if you just do it for the money.
You gotta have the passion of what you're doing, the conviction of the passion of what you're doing. Anyway, so tell us, what was, what was it about 12 Labs? What's the mission that you're so passionate about here?
Yeah, so 12 Labs is an ai, you know, research and product company that I founded with four other, um, co-founders. So we have five co-founders, uh, kind of, um, uh, not normal, but four dudes and one do that, and youth met at the Korean cyber Command, right? So what we do is, uh, we've realized that, um, you know, the industry has been kind of reframing video understanding problem into something easier, like image understanding, or maybe speech understanding.
And our kind of original thesis was that, okay, uh, we're seeing incredible capabilities on the language model side of things, but, you know, um, the world that we live in is a very visual world, right? So we somehow need a, a visual counterpart to, to, to language models, right? Um, so we basically train video foundation models that are kind of watching hundreds of millions of hours of content and try to learn how to map human language onto whatever's happening within video content.
So now, if a model is able to learn that ability now, it suddenly presents us with a lot of this hidden capabilities that allow developers to search for things really quickly within video archives or classify content, um, based on your Yeah. Own defined taxonomies, or even summarize or do a lot of like video question answering video chat per se. So we provide access to our models via APIs to developers and enterprises that are usually, uh, dealing with petabytes worth of video content.
So, you know, historically, uh, you know, video understanding has been a, a hair on fire problem for for many who deals with Right? Um, uh, uh, lots of content, right? So, you know, media, entertainment, sports advertising to, uh, public safety.
Absolutely. I, I wanna get a better definition of video understanding for our audience. Yeah, yeah.
What do you, what is, how would you define it? So, I would basically define video understanding as solving an AI algorithms problem, uh, to, uh, basically come up with a novel model or algorithms that can actually scale to many different downstream tasks that requires understanding of content within video. So meaning, uh, instead of just relying on kind of like the, the narrow purpose-built computer vision models where it tells you, oh, there's a watch, or there's a human face.
This model is something that you can actually, that, that can serve as a new interface for people to interact with video content. So instead of, uh, maybe for, for search, um, if you were to, uh, find a specific scene that you remember from the movie matrix by just describing it, you are able to pull, pull out that specific moment within like a two hour long content, right? Um, and that requires some, uh, level of understanding, right?
Sure. It does. Relying on relying on object based tags or remembering an actor's line word by word and just typing it and, and try to figure out on the transcription space Yeah.
Where, where that is. I get it. So that makes sense.
Yeah. 'cause you know, I don't know, maybe a month or two ago, we actually had someone from Sony, I forgot the division within Sony right now, but they, they, uh, so Sony has the cameras in most phones, the, some Sony, not the camera itself, but there's some Sony technology and a lot of phones and cell phones, and they have put in, or they've now developed like a chip that's more like what you said before, it could say, okay, there, there's a red watch over here, a black car over there, someone's face here. And based upon that kind of tag things, for lack of a better word, right?
So if I want to search, show me the scene with the red watch, it'll go to the red watch, you know, at some point. But this is much more fluid than that, if you will. That's right.
Yeah. Because that I think, is a relatively easier problem. Now, what gets really complex is let's you define the subject a, uh, a guy wearing a dude wearing a red watch, but then he does X, Y, and Z, and then ends up doing, uh, Z one or something like that, right?
Like this very complex query that requires like temporal understanding of what's happening and how things are progressing. Video within video content, you really can't solve it by just having object level tags or like, just transcription, right? And, and we're trying to solve that.
So, and you mentioned petabytes, right? That's enough to get people's attention. So this is, this is in essence a big data situation, right?
Uh, where we're talking just indu me a little bit. So are what is, is what you're doing actually going through all those petabytes of, of video and indexing it? That's right.
Like recognizing it, indexing it, and then when I do my search, it, it pulls it up. Yeah, that's right. So basically when you want to serve industries that are dealing with just a ton of millions and millions of hours of content, and that is like their, you know, uh, almost like a treasure, right?
Because that the content that they produce, It's their crowd jewel, their ip. Yeah. Um, so it requires, um, companies like us to not only push the boundaries in, you know, having models, uh, displaying like incredible video understanding capabilities, but so, uh, puts, uh, a physical limitation to the, the size of models, uh, that we can train, right?
So, you know, some of the, the, the really large ones, uh, you know, there's this concept of scaling model where if you make models bigger and bigger and more data, it gets better, but it also becomes incredibly expensive to serve. Uh, and you know, whenever you put in a video to like a trillion perimeter models, uh, it's very expensive, right? So our focus, uh, is, okay, so our customers have infinitely growing video archive, and they need to be able to search through this growing archive and also summarize work and, and, and do a lot of like video question answering tasks on their archetype.
Um, how do we solve it? We, you know, that, that, you know, kind of physical restriction allows us to have a very sharp mission statement, which is to, uh, provide really performant, small sized models that actually, you know, make sense economically to process petabytes, whether they without breaking in our customer stack. That makes sense.
It does. Makes sense. Yeah.
Fascinating stuff. So let me, before I jump back into the technology, let me come back to 12 Labs a bit. So you got four other, you know, five founders here altogether based in San Francisco.
Is this venture backed at this point, or, Yeah, so 12 Labs is a series a, uh, company. We have an office in APEC office in Seoul, South Korea, and we are headquartered here in San Francisco. Um, the company is about now 70 people, um, wow.
Across South Korea and, and, uh, San Francisco. And, uh, we've raised, um, close to a hundred million dollars, um, in total, uh, from investors like NEA, index Ventures, radical Ventures, along with our enterprise partners like Nvidia, Intel, Samsung, uh, and that's another awesome angels really, um, Who are, that's fantastic. Great founders And yeah, What a great story.
And maybe I'm wrong, but my, my feeling around a lot of the AI stuff, and we, we talk obviously about AI an awful lot here, um, and talk to a lot of people is, look, we're just, we're at the beginning of this journey, not at the end, Right? So, you know, you haven't finished your mission, you have, you know, what you're doing just started now is right. It's a fine beginning as they say, but where do we go from here?
Yeah. So I think, uh, there are a couple problems that we should solve now is, you know, the world is running out of text data, right? So I think, um, we're probably going to be seeing a lot more progress in multimodal ai.
So multimodal basically here means not just text, right? A model that can understand, uh, your video and by understanding video, hopefully it understand image and audio, um, as well as PDFs and other unstructured data. Um, so, you know, making sure that we have, uh, a model architecture that can effectively learn from data that's algorithm text, and, uh, hopefully being able to do that makes models, uh, generalized better, right?
Meaning, you know, it, it has more robust world representation, right? Um, instead of just relying on text. Um, so I think we'll continue to see a huge progress in multimodal ai.
And the second, I think, uh, you know, starting with, uh, the launch of PET GBT, the world was just going through the gen phrase, gen gen AI phase, right? So, you know, oh, it seems like the models are capable, so let's try to apply it this way, that way, right? Instead of kind of like really focusing on solving the hair buyer problems.
It might not necessarily be the most exciting or futuristic or sci-fi use case, but really trying to focus and, and solve problems that yeah, that have existed for a long time and now we actually have, uh, a solution that can address it, right? So we're probably gonna see a lot of like, groundedness in enterprise, uh, and, and, and other companies, um, that are adopting this technology. So from like the POC, like multiple POC phases to now, like finally production.
Um, second, I think, you know, uh, the excitement around models, um, have, uh, basically generated this en created this environment where all this like, you know, try to push and push and push and make these models bigger and bigger. Um, but you know, now the general audience is also understanding, huh, there is only like now marginal gains, uh, in just pouring cache onto compute or just throw data without, um, uh, a careful planning. So I think we're going to see a lot of movements around, okay, this is like a really capable models scaling is required to see what kind of capabilities that we can pull out of these models by making it bigger.
But now we have like a pretty useful model. How can we then now make it smaller and smaller without losing some of these capabilities, um, to make sure that, uh, uh, that it can be adopted widely, and most of the cases cost this the problem, right? By making these models smaller and hopefully small enough so that we can put it on device, um, or on edge devices, um, that'll be fantastic.
And I think the, the industry's moving towards that too. Love it. Jay.
Unfortunately, we're about out of time here, but this, this, this is a great story. It's, it, I mean, your personal story is a great story. The corporate story is a great story.
And then the, the whole, this whole area is, is amazing. You know, look, strikes home to us, I mean, we're, we don't have petabytes, but we're probably sitting on 20 terabytes of video data on you to text TV and on the backend. And, and this is a real world problem for us, right?
How the heck hundred percent that 20 terabytes, I think is about eight or 9,000 videos like this some longer. Yeah. But how do you effectively wield that, right.
And, and use this video and weaponize it, if you will. And, and there's not a lot out there because unless you have transcripts, which, you know, an indexing thing can then look at, it's, it's just a big glob. It's 20 terabytes of a glob.
Yeah. And you've spent a lot of time producing this, and hopefully Four or five years. Yeah.
You Can, yeah. You can capitalize on every single moment that you captured, right? If you want to, um, you should, you should have the, the, the, the right tools to be able to do that, right?
And so how people need to talk Do that. We talk to your people. Yeah.
We'll check it Out. Yeah. Happy to help out.
Yeah, no, I'd love to talk to you about it off camera. Alrighty. Jay Lee, co-founder and CEO of 12 labs.
Check them out if you're into keeping up on what's going on here. It's fascinating video search and understanding using ai. We're gonna take a break here on Text Trunk tv.
We'll be back in a bit.