Achieving 99.999% Accuracy for Visual AI – Digital CxO Podcast EP110
In this episode, Amanda Razani speaks with Dr. Jason Corso, co-founder of Voxel 51, about what it will take to achieve almost 100% accuracy in Visual AI, use cases for it, and how this technology could impact the future.
Transcript
Hello, and welcome to the digital CXO podcast. I'm Amanda Ani, and I'm excited to be here today with Dr. Jason Corso.
He is the co-founder of Voxel 51, as well as he is a professor of robotics and electrical engineering and computer science at the University of Michigan. How are you doing today? Hey, Amanda.
I'm doing great. How are you doing? Doing well.
Can you share a little bit about Voxel 51? What are y'all doing over there? Sure.
So Voxel 51 is a, uh, AI software company that builds a package called 51 and 51 teams for enterprise. That essentially, uh, you know, our, our target user base are builders of visual AI that seek to understand more about them, their model performance and their dataset quality so that they can ultimately build visual, you know, make visual AI a reality. Um, we've been open sourced since August of 2020.
We have about 3 million installs. Uh, we do set up the software in a way that is hyper flexible, right? We don't sort of tell you what to do or how to do it.
It's, think of it kind of like building blocks for the developer in the visual AI space. And, um, you know, ultimately it's been used across various industries from, you know, automotive industries to agriculture, to, um, consumer products, advertising, sports, you know, you, you name it. We have users either an open source or enterprise.
Wonderful. Thanks for sharing. So our topic of today is achieving almost 100% accuracy for visual ai.
So tell me a little bit about that and where you come in with that and we can move on from there. Sure. com or you know, various leaderboards for computer vision and visual ai, uh, sometimes the numbers will be jarring, right?
Like the numbers might be something like 60% or 70%, and, and yet that's a, that's a decade of research that went into that level of performance or decades of research. Really, if you, if you think about the predeep learning era too, too, um, but if you think about what's necessary for actually deploying visual AI into practice, right? You, you, you have an autonomous vehicle or you're doing a, you know, a quality assurance on a manufacturing line, 70%, frankly, not good enough.
80% is frankly not good enough. Uh, even 90 percent's really not good enough, right? Like, if you only miss nine out of every 10 pedestrians that come into your, in your view in an autonomous vehicle, you're gonna have some problems, right?
So, so we think of this, uh, as, you know, an analogous to the classical five nines concept, right? 999%, uh, comes from. It's, it's from network security and network uptime, basically, uh, in, in various, you know, pre-cloud, in the pre-cloud era.
Uh, and it, it is super critical to, uh, achieve that level of performance so that we can begin to, uh, assume robustness in various visual eye capabilities and then take the harder problems to humans downstream, right? Uh, where, whether or not that is, I mean, I, I like the QA on the assembly line or the fact manufacturing line. Like it's a really good example, right?
Like if you can basically have the, the humans, the domain experts who understand the product, understand how the circuit board needs to look or whatever, the, you know, I just made a French press coffee, right? Like how the French press needs to be assembled, right? Like, and only give them the corner cases, their time will be used more effectively, and the overall efficiency of the manufacturing line will also be used more effectively to, to do that.
We, we wanna get to, to this notion of five nines in, in visual ai. So Once you have this very accurate, uh, visual ai, what are some use cases where this could be integrated? Well, I, I mean, I think this, this really talks to you or gets at the, the, the popular terms that are, that are being used these days.
Uh, right? So you have visual AI is one of them. Then we have this notion of physical ai, right?
That, you know, ultimately, like the fu what, what does the future of, um, robots look like? Either, you know, take, take robots out of the factory and put them into the hospital, put them into the classroom, but to put them into the home, what can they do? And then even post physical ai.
Now, we've also begun to hear about age agentic ai, right? Like, not only can these robots or these physical AI things see through visual AI and listen through, you know, language ai, but can they also begin to reason about goal, the goals that are either given to them through conversation or talk to them in a training phase to then go and understand what they should be doing to increase, increase efficiency or to help the elderly who's fallen down or give the right medicine or what have you, right? 99%.
Uh, visual AI is one of the elements, one of the key elements of those foundations, right? Like o other elements are vast low cost compute. We're getting there, right?
The recent releases we've seen in, uh, January of 25, like from Nvidia about I think the system's called digit, like it's moving this notion of high power compute that in a more portable edge based manner. I think that's, that's one of the elements of the future. Another element is, um, you know, like the ability to do speech recognition and speech generation robustly, uh, to be able to interact with humans on the humans' terms.
Uh, and we're seeing a huge amount of progress in that, right? The whisper technology from OpenAI is, is like one of the latest in speech recognition, very capable, uh, still need some work around speech or separation. Like if, if there's a robot to make, you know, amongst many humans, and like I, I have three kids in my h in my home, so like many kids around making a lot of noise like that, that gets a little bit more difficult.
But once you have these pieces that are robust at the lower level, then we really can, uh, we can begin to think about more practical physical AI and physically agent AI in, in the, in other applications that, and I think that's ultimately the, like the, the true impact, positive impact that AI will have to everyday life is not gonna come until we, we've built those things first. Uh, and or at least demonstrate that they're robust and safe and and secure. We're starting to see some self-driving cars around.
When do you think it's going to be the case that almost every car is a self-driving car out on the road? How far away from that? Yeah, that's a hard question to really pit a pin a time on.
And anyone who has been, uh, confident enough to give a time has been wrong in the past. I think, um, think, think about it another way, right? Like when, when elevators were first, uh, created, um, there were humans in the elevator to operate the elevator, and the elevator was still fully autonomous, it was just being operated by a human because that was the social, um, contract if you will.
There, there was this comfort level. So, um, I don't know exactly I should look this up. I don't know when it became more common to have, um, fully autonomous, no human operator in the elevator than the, you know, than not.
But if you think about the progression in automobiles in the last decade, I think it's been really powerful, really amazing, right? Like we, we moved from classical cruise control, for example, to, I mean, different companies call it different things, but essentially like radar base or adaptive cruise control that can maintain a speed unless there's a vehicle up ahead, uh, can even, uh, uh, adjust the speed whether or not you're the, the traffic based, based on the traffic density, for example. Um, you know, so we are getting a lot closer to nearly autonomous, um, you know, uh, uh, vehicles.
And I, I don't think, I think the social comfort level will be the final hurdle to overcome even once we have the technology that's capable and it's real. So, and humans, you know, we humans are really hard to understand a lot. So it, you know, hard to, hard to predict when that will be, but I I actually think it's gonna be sooner, sooner than you think technologically, but later than you think socially.
Mm-hmm. So, I know it's hard for humans to trust technology fully in that way, and technology can of course make mistakes, but do you feel like if every car on the vehicle was every car on the road was self-driving, we'd have a lot less accidents? Oh, I, I do think there are both.
There are, there's clear evidence that that would be true, right? Like, um, that first of all, like these autonomous vehicles won't have to worry about the radical unpredictability of, uh, uh, what typically happens on the roads, like at least roads that are out of, uh, human occupied occupied areas like highways and things like that, right? Um, but I think the, there are other elements that if, if we, if we had infrastructure that was built or updated to support autonomy, I think we'd already be there with less, you know, in some sense with less traffic.
You know, I think there was a study in the nineties somewhere in California where, um, they built a small or built or like cordoned off like a portion of highway to do some tests, and they were, they did, um, these, like, I forget the exact name for them, but when you have autonomous vehicles that are able to like, go very close to each other as they're going down the road and they have these like, you know, a dozen or so, I mean, I mean, I think the fuel efficiency was demonstrated to be ridiculously higher. Um, and there were no accidents recorded in that study, you know? And even though there were humans in the cars at that point just for safety.
Um, so, so anyway, I think the evidence is there that we will see, um, like the, through better communication between the vehicles, through more predictability between what's happening via that communication and or just less, less, um, unpredictability from human drivers. I, I do think you'll, you, we will see safer roads, uh, at, at that point. Yeah.
Yeah. You bring up two great points. There is not just the safety issue of that, but also the energy efficiency.
So I, um, hadn't thought about that as much, but um, that should help with, um, energy and gasoline use too, if, if all the vehicles are driving as efficiently as possible. Uh, 100%. And, and I think, you know, not, not only can it be already demonstrated that that's the case, um, I mean, I, I think it, it's when you have roads that are that, that are more safer than, than we unlock the capability to, um, redirect other investments that are being made is maybe what the thought that I'm thinking, right?
Like, um, I, I know that, that there are some companies who focused their like, like autonomous driving trucking companies who focus on energy efficiency now even, right? Like, how do you do better cruise control knowing the weight that's in your bed when you have a, a hill coming up, for example, rather than it's gonna be a flat terrain, right? So I think fuel, the fuel efficiency is a, in some sense, like a fiscally responsible way to push toward infrastructure change, in my view, for example.
Um, yeah. Yeah. So infrastructure is next is important.
It's an important step for that, for this to happen. Do you think visual AI would soon be integrated into more cities? In other words, smart cities with, um, visual AI being integrated into the street lamps and the street signs and, you know, cameras and that kind of thing to determine traffic and all that kind of stuff?
Yeah, I, I think that that, uh, progression has already begun. You know, there are already modern cities, even in the US that have hundreds or thousands of cameras. Most of them were installed for security and public safety purposes initially.
But once one realizes that oh eight, if we actually model traffic flow through, we, we have all this, this data now and we can connect other visual AI to understand, um, you know, a general way of saying that is like patterns of life, right? Both from vehicles and events and humans and pedestrians and so on. Um, then we can, um, optimize for various use cases, both safety, uh, and, and energy responsibility or even just access, right?
Like we want to bring, we have these beautiful cities in the United States that, uh, downtown areas that are often not that occupied. Um, you know, so we can, if people feel safer, they feel the cost will be less, even just figure out how to do better parking, for example. All that will come through visual ai, AI mechanisms that better understand and utilize the data that are available.
So all this is underway in some level or other. What do you envision knowing how fast technology is advancing? What do you envision for the future, say 10 years from now?
Um, w well, I think it's, I mean, it's, it's such a, it's an great question. 'cause this is such a great time to, to like be thinking about what's, what's coming, right? Um, I mean, I do think the, we, we are, we have experienced a critical point in the development of, uh, of AI broadly like visual ai, physical AI and so on.
And that that did come from the demonstration that these like super large scale scale models connected with vast amounts of compute and vast amounts of data, were able to, um, pull out elements from our human world that, that allow a better translation between the computer world and human world. Um, I think this came initially in my SA area, like the world of like, um, computer vision, video understanding and so on. Like, you know what my, my, my research group for example, at, at Michigan Focus has focused on video to text for over a decade, right?
You know, an image to text. You know, my, my National Science Foundation career award was on the image captioning problem and like 2008, right? Like, so I've seen this, this field grow.
Uh, and I think when CLIP came out, you know, 20 18, 20 19, we saw this notion that, oh, wait, like if you have image data, visual data, and you have language data and they're connected and training against each other, and you're doing this at scale, there is immensely powerful structure in the UN in the model that's learned from those data. And, um, we are now be beginning to see that structure being leveraged for downstream tasks. So in the field, oftentimes these mo these models that do these types of things at scale are, are being called foundation models now.
Um, and when you can take a, a foundation model off the shelf and just make use of it for a, a different task for which it wasn't anticipated or it wasn't trained for, you may have to fine tune or do rag like retrieval augmented generation to go and like augment what was originally in that, originally in that data for a downstream task. Like that work is, is already light years ahead of where it would've been if you had to start from scratch, which starting from scratch would've been what data do I need to train this model? What's the right architecture for, for my model, right?
And then you go and spend months or years getting that data. That's why like, you know, in some sense maybe the 2010s where like the decade of large data sets in machine learning research, machine learning research, right? Like everyone, everyone who wanted like to get high citations on their papers were like, oh, I can just get a, make a big data set and I'll go get a lot of citations.
And I think that that was, that was last decade. And now we're in this era of what can I do with these foundation models? And this is, I think we're still really nascent on that, on what's possible, but I think that that inflection point, that critical change happened when we saw what things like clip were capable of, and that, that structure there is, is just super powerful.
And you see like foundation models, like Microsoft had this, uh, paper from CDPR last year, Lawrence two, which is one of those like sort of downstream models that build up on top of a lot of data, a lot of foundation elements, foundational elements, and then you can query it with pretty intricate questions, uh, about the visual data and the performance is, is quite, quite strong. All right. Well thanks for sharing that.
So if there was one key takeaway you could leave our audience with today, what would that be? Uh, the key takeaway that I, that I like to end, uh, sessions like this would be that even though there's such amazing progress in AI and visual AI and so on, um, I think it's important to always remember that we do this for, for humans. We are humans.
We're building for humans, right? Like, how can we make our society a more balanced, more equitable, safer place for everyone? All right.
Absolutely. Well, thank you so much for coming on the show and sharing your insights with us today. Thanks, Amanda.
Happy to do it. All right. And thanks to our audience, stay tuned.
There's more.