Voxel51 CEO Brian Moore on the Current State and Future of Visual AI Adoption
In this Techstrong.ai Leadership Insights Interview, Voxel51 CEO Brian Moore dives into the current state of visual artificial intelligence (AI) adoption and the journey ahead.
Transcript
Hello, and welcome to the latest edition of the Techstrong AI Leadership Series. I'm your host, Mike Ard today with Brian Moore, CEO for Voxel 51, and we're talking about why a lot of these visual AI projects seem to be failing. Brian, welcome to the show.
Thanks for having me, Mike. We see these use cases all the time. I think most commonly people are seeing, uh, vision applications and everything from factory floors to cars, but a lot of the efforts underway seem to be still experimental and people are struggling.
What's your assessment of what's going on here? Yeah, definitely. So first of all, just to generalize it to all of ai, I think that's kind of the state of play in 2025.
You know, we've had studies from places like Harvard Business Review sharing that something like 80 to 90% of AI initiatives within enterprises are not yet reaching production. You could call that a failure. I would just call that kind of expected or par for the course.
You know, this is new technology. There's lots of rapid innovation, there's a lot of excitement to try new things and build proof of concepts. Unsurprisingly, uh, something that you cobble together in a few weeks or even months, is unlikely to meet the needs of the production environment that you need to deploy that into.
And that's perhaps, you know, uh, most poignant in something like visual ai where we're talking about deploying, you know, robots or vehicles, uh, or automations that have to act in the physical world and deal with all the different sort of nuances, edge cases, strange scenarios that might crop up. So yeah, there's definitely a need to invest, you know, kind of the typical 80% of the time to get that last 20% of the way to production. Uh, but the good news is that folks are aware of that, uh, and companies like ourselves are building technology to help assist, uh, you know, practitioners of visual AI address those key, uh, needs, which we can dive deeper into and get that model ready for everything that the production Yeah, the real world, uh, will throw at it.
All right. Well, to your point on that maturity curve, where are we when it comes to vision ai? Because, um, I guess there are some unique challenges there, but what are they?
Yeah, so the interesting thing about visual ai, and by visual ai, I mean, anything that has to do with image or video or 3D uh, lidar radar data, um, that's an absolutely immense data source. Something like 90% of all of the bits that go through routers on the internet today are actually visual in nature. Uh, so the vision AI challenge is at least two orders of magnitude larger than the challenge, uh, of building models that can, for example, process text, right?
So it's kind of expected that it'll take some, you know, additional time and effort to get these things, uh, ready for production. Of course, the good news is that with all the investment going into accelerated computing infrastructure, you know, data centers, power r and d, all the things you hear about in the news, uh, those advancements are coming. The, the promise of being able to feed larger scales of data into these systems, uh, is also coming.
And so I would definitely expect to see continued progress on some of the kind of bulletin board vision AI use cases that everyone's familiar with, self-driving cars, humanoid robots. But maybe most, uh, exciting to us are kind of the more incremental advancements, you know, automating specific scenarios like maybe defect detection in manufacturing context, uh, or building purpose built expert systems that can, for example, you know, uh, detect the fall, uh, of a, a human in a healthcare context, uh, or automatically, you know, process, uh, a camera feed to make a decision about whether a part is, uh, high quality or low quality. Those kind of things, uh, are much more short term.
Uh, and we're seeing those types of technologies get to production, which is very exciting for the, the vision AI field overall. Mm-hmm. I think everybody's excited about the use cases, but it seems to me they're also running into issues around, well, what does it actually cost to run something in a production environment when you add up all the infrastructure and resources required?
So do we need to be smarter about what projects we're gonna pick with an eye towards what's gonna go into production sooner than later? Yeah, I, I think it's a great call out. So, you know, uh, one of the exciting thing that's happening in the AI space is the progress of these, uh, so-called foundation models.
The large models, you know, the GPTs, uh, coming from hyperscalers and, and those models are, uh, have a broad expertise of knowledge. Uh, however, large models are expensive to run. Uh, and so what you can expect to see is those models knowledge getting distilled into smaller expert models that are more efficient, uh, at solving, you know, specific tasks.
Uh, and so, you know, that's what's actually getting, uh, into production, uh, in vision AI especially, is these distilled models that are purpose-built for specific use cases that can run at a much more cost-effective, uh, price point. As you kind of sort that out, who's gonna build those smaller distilled models for organizations? Is that some data science team that they hire?
Or are there specialist organizations that are emerging who are gonna basically make those things available as a service? How does this kind of manifest? Yeah, so what we're seeing is that enterprises that, um, are having the most success in visual AI are ones that bring the development of these sort of fine tuned systems in-house.
They treat, uh, their AI strategy as a core part of their company's competitive advantage. Uh, and so they want to bring as much of that development in-house as possible. That definitely means using off the shelf models, uh, data sets and so forth, uh, to sort of, you know, turbocharge their, their development.
Uh, but they see their ability to develop, uh, a high quality data set, uh, and model that's an expert in their use case, uh, as being critical to their, uh, company strategy. I'm, what are the skills that are available as it relates to this? And I'm asking the question because, well, we're already having a hard time just finding your everyday run of the mill data scientist genius, and how many of them are actually cognizant of visual AI and, um, what does the pool of talent look like?
Yeah, so for context, uh, a little bit about myself. So I have a PhD in machine learning. Uh, voxel was founded by myself and my co-founder, Jason, actually over 10 years ago, uh, initially doing consulting, uh, in, back then it wasn't called visual ai, but rather computer vision.
Uh, and so computer vision as a field, it's actually been around for quite a long time. In fact, even Nvidia as a company, uh, got its start and spent many decades focused on computer graphics, the kind of, you know, uh, low level computer vision, uh, algorithms that are necessary to, you know, build graphics engines, video games, so forth, right? So there's a rich history and, and expertise in, in the, in the market that exists in computer vision.
Uh, and, you know, so that's the good news. There's lots of, uh, you know, capability out there. And then, uh, what we're seeing is that whenever there's advancements in sort of overall AI technology, uh, those models, those architectures can be deployed not only for language use cases, but also for vision use cases.
And so the visual AI field definitely benefits from all of the advancements that are happening in, you know, large language models. Uh, as an example, the, the transformer architecture that Google created a couple years ago, uh, was very important both in language but also in vision. Mm-hmm.
Of course, it takes a village to kind of build these applications and there's developers involved and, um, data engineers and all kinds of folks, but, um, how do I meld them together? I think a lot of organizations I talk to are struggling 'cause the data science folks have one culture and the developers have a different culture and they're not quite in sync with each other about how to not just build an app but maintain it and update it. That's a great point.
So, you know, historically what we've seen is that some, uh, companies have decided to kind of create separation between what they call their data team and then their model or product team, which might result in kind of a separation of duties where the data org is responsible for building data infrastructure, maybe gathering data, and then throwing it over the wall to the product teams that then make use of that data kind of a one-way street type of modality. Uh, however, you know, especially when we're talking about model failures, what we're seeing is that the ability to overcome model failures and actually get a system into production, like we were talking about earlier, that comes really from the iteration cycle, being able to understand, okay, I built a data set, I trained the first version of my model, where is it succeeding and where is it failing? And inevitably, if it has a failure, maybe an edge case or a certain type of bias that you've discovered, you need to address that, uh, through better data.
It's not sort of a one one-way street, it's an iterative process, and you're in the best position to make that kind of, uh, action to improve your model, uh, if your data and model teams are working closely together. So, as an example, uh, our software 51 that we deploy to enterprises, it kind of puts the data at the center of the entire development process of visual ai, allowing data teams and model teams to collaborate together in one place. And when they see a model's performance or lack thereof, the underlying data is always one click away.
So if just a click of a button, a user who's trying to, you know, evaluate a model can understand, oh, I see this is why the model's performing poorly. I can see that there's, uh, something I didn't expect about my data. A bunch of it is low light or low quality maybe, or it's having problems in the self-driving yeast case, uh, you know, understanding, you know, sort of crowded intersections and low light conditions.
If I'm able to go gather more examples of those problem areas, that'll be the most effective way to improve that model's performance. Mm-hmm. How readily accessible is that kind of data?
I think, you know, you hear people talking about how all the data that's publicly available aren't even sucked up. So do we have enough of this visual data to train the models going forward? Yeah, that's a great point.
So, so that, that kind of, um, quote, uh, is most often used about large language models, people are saying that, you know, the reason that let's say GPT five had a smaller delta than one might have hoped over GT four is that we've already ingested all of the data that's available on the internet. That's definitely not true of visual ai. Uh, like I said before, there's, you know, 90% of all of the bits that go through routers on the internet are visual in nature.
Uh, and there's definitely a vast amount of untapped data in visual, uh, it's visual in nature, uh, that is yet to be fed to all of these models. Having said that, uh, one trend that we're seeing with our customers is, you know, kind of by definition where you need to spend all of your time are on the edge cases, uh, or failure modes of a system. And those are rare, they're hard to acquire.
And so companies like let's say Tesla are in a, a good position, uh, where they can, for example, trigger, uh, anytime there's a, a hard braking event in a vehicle, they can capture that scene and feed that data back to headquarters and use that to specifically address that, you know, sort of failure mode. So being able to connect to your development process, to the products that you're putting in the real world, that's a great way to gather more data. On the other side, we're seeing, you know, synthetic data as an example.
You know, we partnered with Nvidia, uh, to make their, um, neural reconstruction, uh, models available to our customers. That's a situation where you can generate a synthetic version of the scene, and then you can play with things like, Hey, I've got this scene. Uh, what if I swapped out that FedEx truck for a UPS truck?
What would happen then? Or what if I took this sunny scene and I wanted to consider how the model would perform if it was instead snowy or rainy? Uh, you can perform those types of, you know, uh, you know, uh, synthetically generated tunings, uh, you know, uh, from a model standpoint, which obviously gives you the ability to plug those gaps that may be hard to acquire, uh, real data for now, I would say that that's kind of a, you know, uh, up and coming technology, uh, and there's definitely interesting questions to be answered about, you know, how do you evaluate how much real versus synthetic data you need, uh, to build, you know, a production ready model.
Uh, but it's definitely something that our customers are excited to, to tap into. It also seems to me the tolerance for being wrong in these apps is a lot less, shall we say, than it is in your typical, um, chat GPT type of application where, you know, if the thing hallucinates on some sort of summarization, I'll notice, but, you know, if it's not the end of the world, then I'll shrug. But it feels like with the visual ones, that those applications are a little more, shall we say, mission critical.
Is that fair? Yeah, I think what you're getting out there is that, you know, and this is kind of maybe obvious with hindsight, but the key to getting these systems into production is choosing the right use cases. And to your point, the right use cases, especially early in the development cycle, are ones where the, the, the system or the use case can tolerate a failure.
So yeah, if you're, if the task is to summarize, uh, some content, uh, for a human to take action on, and if it's not quite right, you know, there's a human in the loop already, and so maybe it's okay, right? Uh, on the other end of the spectrum, you could see something like a self-driving car where, you know, it makes kind of intuitive sense that it needs to be, at least in order of magnitude safer than a human driver in order for us to kind of accept, uh, any sort of failures that might happen. Uh, but the good news, like I was saying before, is that in addition to the sort of, you know, bulletin board use cases for visual ai like fully self-driving or fully humanoid robots, there's a lot of, uh, smaller more sort of focused tasks like detecting defects or, you know, uh, automatically processing, you know, user imagery for insurance claims where there's a lot of summarization and sort of constrained environment tasks that can reach production level and will, you know, uh, add lots of value to us while we continue to push towards those bulletin board use cases.
Kind of similar to how everyone spends some fraction of their time talking about what a GI will look like when in reality, uh, those use cases like, you know, automating customer calls, uh, in service centers or, you know, summarizing, uh, knowledge work, uh, for enterprise, those are the real value creation in the short term. And to your point about that, you know, everybody talks about, well, who moved my cheese and am I gonna get laid off? But when you look at those use cases you're talking about, some of them are things that we probably would never have done in the first place, and many more of them are things that well, nobody really enjoys doing in the first place, and we typically don't do it all that well.
So is that part of the thinking about where to make use of something like vision ai? Definitely. Right.
And, and just to put a point on maybe a macro trend that's happening right now and how it's impacting the visual AI space, uh, you know, we're talking a lot these days about onshoring manufacturing for various, you know, political reasons, uh, you know, so forth, which we won't go into here, but, you know, there's the sense that, well, uh, there's, there's certain tasks, sort of menial tasks in factories and so forth that may be, you know, some fraction of US workers aren't interested in doing, or it doesn't make sense to do, uh, at sort of like human level price points. Perfect use case for vision ai, right? We can come in and invest, uh, as we're onshoring manufacturing in building automation, uh, and they, you know, building out factories that'll put us in a, you know, a competitive advantage compared to, you know, even our, uh, offshore, uh, competition there.
Well, let me ask you this then. Um, is this really a separate discipline in the sense that there will be a separate ecosystem for it, or, you know, are you at all concerned that the open ais of the world and everybody else who's in that space is just gonna, you know, just add this to their portfolio of services? Yeah, so that, that's, um, kind of what I was getting at before when I was making a distinction between, you know, what a foundation model, uh, can do, uh, versus, you know, what's actually required to get that particular use case fully automated and in production.
Uh, you know, there's, there's interesting conversations happening right now. Let's take on the language side for a second around, you know, hey, if, uh, if it is true that these large language models have quote unquote PhD level knowledge in all fields, then why do I need to go to a healthcare provider? Why can't I just go to chat GPT and have it solved, you know, provide my diagnoses and provide a plan of action?
Well, there's not, it's not just as simple as providing the knowledge. There's expectations around, you know, quality of care and certification and so forth that come out. And for that reason, it's not gonna make sense for a single company, certainly in the short term, to provide sort of expert level guaranteed certified services in all these use cases.
Uh, and I would expect the same thing to happen in vision ai. You know, it'll make sense for vertical specific companies that deeply understand use cases and customer needs and so forth to take a technology, a general purpose technology off the shelf, and build a solution for a specific vertical. Uh, not to mention the, the, the point I mentioned earlier around how it's not cost effective to take that general purpose model and plug it in directly.
That may be sufficient to build a proof of concept to show that something can be done. But ultimately to drive margins up, cost down, you're gonna have to invest in building more expert, you know, fine tuned systems. And that clearly has to be done by a, you know, an entity that's focused on that one vertical and ready to make that commitment.
All right. Well, folks, you heard it here. There's a lot of things to be excited about in the AI era, but maybe all the cool kids are starting to hang out in the visual AI table because that's where the new and interesting applications are gonna be.
Brian, thanks for being on the show. Thanks, Mike. All right.
And thank you all for watching the latest episode of The Techstrong that AI Leadership series. You can find this episode and others on our website. We invite you to check them all out.
Until then, we'll see you next time.