The Debate Over OpenAI: Transparency, Licensing, and the Future of AI Development with Anaconda’s Peter Wang
OpenAI models are central to debates on definitions, open-source software, licensing, and model transparency. Concerns include data source clarity and dependence on proprietary systems. Peter Wang of Anaconda stresses stakeholder collaboration and clear development processes for effective AI.
Transcript
Hello and welcome to the latest edition of Techstrong AI Leadership Insights series. I'm your host, Mike Bazar. Today we're with Peter Wang, who is Chief AI Officer for Anaconda.
And we're gonna be talking about well open AI models. And you would be surprised there's a lot of controversy on this subject. Hey Peter, welcome to the show.
Thanks for having me. Really glad to be here. We've seen people toss around the term open AI models left, right, and center.
Some of them are partially open, some of them are all the way open, some of them have different licenses. What would you be thinking about here? Well, I think, um, you know, we should, we should definitely be very clear about what we mean by it.
'cause if, if we're not precise and clear about the meanings, then we can think ourselves and convince ourselves with things that are actually not true. Um, and the term open source AI is rooted in, of course, the term open source software. And the reason why the term open source software became important was because software, you know, used to be source code that got compiled to opaque binaries that people could not understand what was inside them.
People didn't know if they would always do the right thing. People couldn't fix their own bugs, they couldn't innovate on them. You'd have to have the source code in order to allow all those things to happen.
And, you know, fast forward 40 years, we're at a point now where people sort of take it for granted that most of the source code we use, most of the software use, we can look at the source code, uh, and now we go to AI models. We have now the same question. We have to, I think, recover the origins of that, of that word and that phrase and the motivations in order to think clearly about what it means for an AI model to be open and what are the virtues and values we want from such open models.
There's also, it seems like different kinds of licensing terms we used for different models for that matter. We see the same thing in software, but there's also the notion of whether or not the weights are actually open as well, because I mainly get access to, uh, the model, but I don't know how it was built. Yeah, well, there's the training process that produced the weights.
There's the data that was crunched through the training process. Those are different things actually. And then there's the weight themselves.
Now, in almost all cases, uh, no, sorry, in many cases, the weights are open to you and that's why people use the term open weight models to distinguish from open source, because almost none of the models will tell you the actual source data that went into them. So I don't wanna be too much of a stickler on the, on the details of this, but it is really, really important if we're gonna say open source, that implies you can see the source of it. But if you don't tell me the data, then I don't know the source of it.
You can give me numbers, you can give me a giant, you know, several hundred billion parameters in weights. That's very helpful and that's great, but that doesn't actually tell me what went into it. Um, so I think that's where, you know, there's definitely a, a spectrum of what people make available.
So are are any of these things truly open in the sense of how we think of that term? Or are they all just different degrees of proprietary? There's only a few that are really open.
Um, not very many. And, uh, they include some of the models from IBM, uh, where IBM is, is pretty transparent about the data that went into it. Um, there's one from Allen Institute, uh, and from playlist they just released one.
And I think maybe news research or Prime Prime, um, prime Mentorlike. Um, but for the most part, almost all of the open weight models do not reveal what their sources are. Um, they, they don't wanna talk about it at all.
So what's the danger there? Am I gonna get locked in? Because I also see that it seems like there are standard APIs now, so can I swap out these models regardless of how open or closed they may be anyway?
Well, um, we need to think about what are the things that people actually want in terms of, well, in, in, in terms of anything, right? Besides free. And besides not being locked in or having the option of substitutability, um, there's many other things that people actually want from these things because these models right now, we use them for, you know, gen AI for like text and images and video.
That's all great. But if you're using these things as serious, like the engines inside a, a real computational system that looks at customer data, looks at business data makes really, you know, meaningful, impactful decisions or helps you, you know, predict the future, you really want to know what's happening there. And, and you want not just have like a guarantee that things will be free forever.
You actually wanna understand what went into it and actually have the ability to change that and to modify it. Uh, if you don't have the source weights and training data, uh, sorry, if you don't have the source training data and the training regimen, then you actually, you could do some fine tuning. But at the current state of technology, you have no way of guaranteeing to yourself that, you know, there's not gonna be something smuggled inside there.
And every single day, every single week, we see these instances where, yeah, these models start popping out stuff that wasn't their trading day that people didn't quite expect. So we're such a in or such an early stage of the industry that, you know, the, there's lawsuits flying right now and people will start catching liability for these things and you know, that's, that's gonna be a problem for potentially for users. So I think, you know, we're, we're right now people are in a little bit of a YOLO mentality.
Um, but, but I don't think that's the way it will be in the long run. We've also seen the rise of AI agents and a lot of those AI agents seem to be able to swap out from one LLM to the next, or at least that's mm-hmm. Promise.
Um, so will the LLMs and the models become more disposable in time? How's that gonna evolve? I think there will be less and less distinction.
There's been really fascinating research over the last couple of years that, um, well, I guess it sort of tells something we already knew, which is that for the most part doesn't matter what actual model you use in terms of the code, you know, of the transformers and this, that, and the other, or your training regimen in the limit. Most models that are produced by kind of any, sorry, most weights produced by any of these models, um, they're really an expression, a compression, some, you know, representation of the, of the original training data. So if you have, uh, different, the same training data set, different models, they'll produce weights that are actually, uh, within a rotation, isomorphic to each other.
Sorry for big words there, but just to say, you actually get very, very similar weights coming out of similar source data. So the bigger the models are, the more likely they are to actually be kind of the same model, even if they're made by different teams, even with slightly different data sets. And even if the code for the models are somewhat different, at the end of the day, the alt the weights, like the giant pile of numbers that we use to then generate the outputs, those weights are actually very, very, very similar.
So this is a long-winded way of saying, emphatically to your question, emphatically yes, right? That if you try to build really big ones, what go, what goes into a big one? All the data in the world, well, if you take all the data in the world, they basically all look the same.
You know, they all contain all sorts of different same ideas. Not, not different, but all the same ideas, the same source, the same various things. So you end up with kind of the same to the end where you get differences.
Where we have differentiation possible, and therefore differentiable markets, is if you go to smaller subsets of those things, if you go to very narrow, very specific, the high saliency, high quality, high purity kinds of data sets, if you use those to fine tune, if you use those as agents inside a larger framework, you use those as, you know, thinkers inside a chain of thought or, uh, a mixture of experts. Now you can actually have a somewhat differentiated product. So there's also different sizes of LLMs, there's different context windows and, um, they cost differently, right?
I talked to some folks and they're experiencing what we call token shock. They're like, mm-hmm. So many of these things require inputs and outputs, and they're paying for things as on both ends of this and it quickly adds up.
Um, but do we need big LLMs for everything? Or can we be smarter about which LLMs we're using to be more cost effective? And dare I say, do we need something that feels like finops for LLMs?
Well, um, the answer is yes, we can expect them to get smaller. I'm sorry, you asked the big do we need really big ones? No, we don't.
We do expect them to get smaller. We do expect more people to use small ones and harness them in pipelines or in, you know, kind of, uh, somewhat dynamic fashion. Uh, and what we've seen again, or just over the last eight to 12 months, uh, is really I think a lot of people buying into the mentality or converting to, to, to the mentality that, to pre-training, to post-training or what they call test time scaling, like those kinds of things.
It really is a spectrum of how much, how much do you want, where do you wanna put all your energy, right? And, uh, you can pre-trade a gigantic model, then you have to quantize it down to something like a fraction of the size, what was the point of all that? And then you do this like test time scaling where you let it think for longer, why not have a less precise model then, you know, let it think for a little bit longer.
And then you, you're balancing your costs in that way. So absolutely, I think we'll see smaller models and we'll see people trying a ver a variety of different, uh, reasoning architectures, uh, to, um, to get great performance. And one of the reasons why the toad costs are expensive, just keep in mind one of the reasons they're expensive is because, uh, people are still running these most part on GPUs, which are supply constrained or have a significant markup on them, but actually the work that you need to do to, um, do token inference, token generation and inference on smaller models, you can run that on A CPU, you can run that on a Mac mini.
So that at that point is not just a more efficient model, but you break into being able to run on a category of hardware that doesn't have the price premium of a top of line NVIDIA server grade GPU People are also trying to figure out, well, it feels like it takes a village to build an AI application, right? I got data scientists and developers and software engineers and data engineers and um, uh, assuming that they're all on some common platform, but how do I operationalize that? And, and from an enterprise perspective at scale in a way that, um, you know, is cost effective?
How do you think this is all going to come together in the future? Well, it's, it's, I I, I think right now, it, it takes a village to do it, right? And that doesn't stop many people from trying and end up doing it wrong.
And even if you have a village, you can still end up doing it wrong because we're in such early stages of this stuff. You know, when a company, uh, as large and visible as X AI ends up having some of the gfas like they just had this week, right? One rogue engineer does something to a system prompt, and then you end up with a, a, you know, sort of a real controversial sort of headline, um, you know, the best practices for enterprises.
I, I think that, that that's still being figured out. And in terms of actual ROI for things, you know, it's, it's, it's sort of a dirty secret right now, but the ROI is still there. It's still yet to be seen for some of these folks.
People believe it's there, I think it's there, but we have to figure out how to get there. And, uh, but the real value ultimately is when AI doesn't require a team, when it doesn't need all these experts to carefully craft this thing. When you can actually empower a single end user or someone just, you know, in whatever line of business, um, not particularly technologically sophisticated, but they have wisdom about their problem.
They have information about their data and their customers and the business reality and context. So what AI can do as a true digital transformation, the AI transformation of businesses, is to bring that person's insight in, in the most effective way possible and then connect it to an infrastructure a backend that is well plumbed, well thought out, that is enabled to say, here's the best practices. No, if you don't know what Python is, you should probably not go try to find which of the million models a hugging face is the right model for your application.
Here's the golden path, right? And when you deploy a model, you don't have to know how to detect vulnerabilities in the software pipeline for that model. We have a platform that'll help you manage that.
And we have central it, which used to just deal with a few, you know, hundred internal developers. Now they're faced with literally tens of thousands of end users, business end users deploying LLM agents. How do they end up not having, not ripping their hair out?
This is all really what stuff that we've been hearing from our customers as they've been on the vanguard of deploying ai. That's why we built the Anacon ai, the enterprise AI platform, which we just released this week. It's really to allow all these different stakeholders to play their role well and have visibility and have traceability and provenance across the entire process of building an actual operationalized enterprise grade ai.
So we may not get to some magical Uber platform that automatically democratizes everything overnight, but it sounds like what you're describing is there will be, uh, swim lanes where everybody can more easily collaborate with each other within the context of the same project. And, and so doing reduce a lot of the friction we're currently seeing. Yeah, it's, um, the way I think about it's like, you know, any kind of complex system, right?
Um, it can, it can fail. It can fail if any one of the parts fails, right? So if you wanna have secure, reproducible and reliable ai, if any of the steps go wrong, then the thing doesn't work.
So you have to do each of the steps, right? But to find a single person or build a magic team that knows how to do all the steps right, is virtually impossible. It's like hunting a unicorn.
Um, it's a very similar problem to what we saw with people trying to operationalize data science and ML ops the same problem. AI makes it even bigger and harder. So the goal here is to create success lanes, right?
So people who are good at understanding the data, they can make their data available, they put it into registry, people can know what data sources they can use, what models are, are, are well tested and vetted by the organization for what kinds of use cases. So you're looking at internal model catalog as opposed to trying to scroll through, you know, piled millions of open source models. You can just bring ones in and, and then the people who take a model and they wanna do a domain problem, they're able to do that part of it and then again kick it over the fence to the next guy.
So it's really about every stage of that factory floor doing its part right and being connected well to the next stage. Uh, it is not, not trying to be some uber thing that solves everything, it's just setting standards and doing the right thing at each of the steps so you have a better chance of, of successful up the whole process. Alright?
You're heard in here folks, they say good engine based on separation of concerns and that's right. If we have a platform that enforces that, we might get to where we're going faster. Hey Peter, thanks for being on the show.
Thank you so much, Michael. All right. And thank you all for watching the latest episode of Techstrong AI Leadership Insights you can find on this episode and others on our website.
We invite check them all out till then, we'll see you next time.