Open Source AI Models with Kusari’s Ben Cotton
Ben Cotton, head of community for Kusari, delves into the ongoing debate surrounding the definition of “open AI models.” The discussion highlights challenges in applying the traditional open source concept to AI, particularly due to the inclusion of data and model weights. Cotton emphasizes that while efforts are underway to create an open source AI definition, the complexity of AI systems means that new terminologies and frameworks — such as different degrees of openness — may be necessary for the community to address biases, licensing, and transparency issues effectively.
Transcript
Hey guys, thanks to Throw. We're here with Ben Cotton, who's head of community for Qari, and we're talking about this debate that's going on about, well, what exactly does it mean to have an open AI model? Because, uh, a lot of folks are saying, Hey, it's not enough just to say I can inspect something and call it open.
And others are saying, well, I have licensing terms that are restrictive, and maybe they're not nearly as open as we think of in terms of open source. So it's getting complicated out there. Ben, welcome to show.
Thank you. It's good to be here. Alright, so frame this debate a little bit for us.
What exactly or, or is the argument that we're having amongst ourselves and how might it impact people in a way they should pay attention to? Yeah, so, you know, right now there is no, uh, generally accepted definition for what open source AI means. Um, you know, there've been a lot of companies that are trying to, um, you differentiate their models by saying this is open source and kind of latch onto the, the credibility that open source has gained in the software world over the last few decades.
And so you have things like meta's lama model, which, um, you know, they say they call open source, but it has restrictions on, you know, if you, uh, are making more than X amount of money, uh, you can't use this model without paying for it. Um, which is a very clear violation of the open source definition for software. So the open source initiative, who is the curator, the maintainer of the open source definition, uh, they've been working to put together an open source AI definition.
0, It would seem simple enough. But, um, I assume we're not forming a parliament here and having a big meeting. So what are the challenges, the defining that framework and, and, and where are we?
Yeah, so the, The main challenge is that, you know, we're kind of taking this term that's well accepted in software and applying it to something that is software plus configuration, which is, you know, model weights and also data. Um, you know, you can have an AI model without any of the training data, uh, but it's not something that people can actually use, right? And so, you know, the idea is if I want to go to, you know, a chat bot, I, you know, if I want to use an open source chat bot, well, you know, I need to access the training data because that's a big factor in what the end ultimate output is going to be, right?
We've seen, um, you know, even years ago, um, you know, the training data matters. A, uh, Amazon started to use AI in their hiring process, realized they were just reinforcing existing biases from, you know, decades of hiring and scrapped it. Uh, we've seen police departments use, um, AI for facial recognition and, you know, crime detection and, um, suspect identification.
And it turns out that it was, um, you know, ha struggling to tell different black people apart. And it was, you know, again, reinforcing these societal biases. Uh, and so without having access to that training data, it's really hard to know what the ultimate, you know, outcome you're getting.
Can you rely on it? Um, and so, you know, data is, you know, it's content, it's governed by different sets of laws. There's a lot of times, um, personally identifiable information.
So if you're having like, say, a medical chat bot, um, it might be fed with the medical history of, you know, many thousands of people, well, they don't necessarily want to share that, you know, that the full raw level. Um, and so, you know, the OSI has tried to strike a balance of, all right, you, you know, they say, well, you have to release the model data if you can, and if you can't, uh, you know, do the best you can to describe it how you got it, and, you know, maybe people can go either grab, grab it for themselves because it's a paid product, or, you know, they can put together a data set that's substantially similar. Um, the criticism of course is, well, that leaves a really big loophole.
You just, you know, you can find a reason to say, oh, we can't share this. Sorry. Uh, but you still get to call yourself open source.
And that's where, uh, you know, a lot of longtime open source contributors and advocates have pushed back on what, uh, the OSI is developing Mm-Hmm. It almost sounded like there's two things that work here. One is I kind of need to be able to understand the degree at which something is open, and I, so will there be like a scoring system for that?
Because I may have a full boat open source environment, or I may have a model that has, that's technically open source, but has some restrictions. So are we gonna create something of a rating system? So the, the open source AI definition as proposed is, uh, basically a binary yes or no, um, in the same way that the open source definition is for software.
Uh, and, you know, I understand that approach. I am not convinced that it will be the long-term way forward, because like you said, you know, there are, there are degrees of openness. Um, and, you know, I think preserving the meaning of open source, um, as a, here's what it that means is good, but maybe we need to develop, uh, a more rich vocabulary, right?
Of, you know, open weights versus open data v versus fully open. Um, you know, there are probably better terms that are, uh, you know, somebody can come up with given some time. But, you know, I think there does probably need to be a, you know, a nutrition label kind of thing of, you know, it's open in this way, but not in this way.
Um, just because AI systems are so complicated. And one thing that I've been thinking about recently as the OSI, um, but leadership and people working on this are, you know, discussing, you know, the, the draft that's available and the criticisms. And, you know, the, one of the key differences is for software, the OSI doesn't certify or make a determination that a individual software package is open source or not.
It evaluates licenses. So if I start a new software project and I pick a license that's approved by the OSI as open source, my project is open source, you know, I don't need to submit it for review ai, it's, you know, we're talking about systems here, right? So it's the software plus the model weights, plus the training data.
And, you know, all of a sudden that, you know, breaks down because now the OSI is looking at, all right, how do we evaluate on a system by system basis because there's not necessarily an overarching license that applies to all of the components. Um, and so that's been one of the challenges of, you know, trying to develop not only the definition, but then a process for evaluating that, uh, in, you know, in real life as people submit models for evaluation. And the second nuance here is the transparency of the model, because it may be theoretically open, but if I don't understand, uh, what data was used to train it, and that data may be proprietary, I can't assess it for bias.
And at the same time, I might also, uh, run into issues around how the weights were applied, which kind of also may lead to bias. So, um, is there gonna be kind of a, a grid for how we evaluate these things and the people understand all those issues? And for that matter, also, different models cost different things to run.
So how do we kind of decide which model to use when it feels like it's kinda hit or miss? Yeah, it's, you know, it's really hard right now, and I think, you know, from my own criticisms of the open source AI definition, I think having any definition, even if I, you know, disagree with, you know, calling it open source, uh, just having a definition out there that we're all, you know, using to mean the same thing is going to help, um, you know, it's, it's hard right now because, you know, there are certain, there's a certain degree of you just, uh, trust in the model provider, right? Most people who are using a chat bot or, uh, image generation tool or, you know, using machine learning for data analysis don't, uh, you know, they can't go in and, you know, really evaluate, um, you know, how the model was trained or how it works, you know, they don't have the expertise.
Um, you know, just in the same way I've been running Linux at home, uh, and at work for almost 20 years, and I couldn't go in and debug something in the Linux kernel. It's fully open source. All the source code is there.
Uh, I just don't have the, the expertise there. Um, and so I think that's both the benefit and the risk of the open source AI definition is now we have something where we can say, all right, this is open source. So, you know, there's the hope, um, you know, the trust that somebody along the line has said, all right, this is not, you know, obviously malicious and it's, you know, trained with, you know, some reasonable data and it kind of claims what it it says to be.
Um, and so it does sort of, you know, provide that, uh, that cognitive escape hatch for people. It's like, all right, well, I'll pick one that's open source. Um, but it also, because it doesn't necessarily require bold training data set to be available, uh, it does pro, you know, still preserve some opacity, um, especially if you are a consumer, an end consumer of the AI output and not a AI model developer.
Um, mm-Hmm. And I think that is, you know, potentially another avenue where, you know, we might need sort of multiple definitions down the line is if you're developing a new AI model, you actually maybe don't care what the version, the old version the, the upstream was trained on, because you're not using the output, you're using the model itself. And so to you, in that use case, the data, the training data really doesn't matter because you're going to take the model code or the model weights, change it in some way and provide your own data to produce your AI model, but as a consumer, you're not.
So it, you know, these are two very distinct use cases in a way that I don't think the conversation has really addressed that, you know, sort of differing need. Um, you know, and one, the data is irrelevant and the other, the data is everything almost. Hmm.
Um, as you kind of think about this for a minute, what is your sense? Are we just being too careless with the, uh, terminology for open source? Or are people deliberately going out of their way to kind of say they're open when they're not?
I think it's, there's a mixture of both. I think there are people who, um, you know, I, I, I don't work for meta. I don't have, uh, any conversations with people inside, so I'm, you know, kind of guessing, but I think, you know, meta is well established enough that they know what open source means.
They have people in the company who understand it. And I think the use of open source as applied to LAMA is probably a deliberate, well, there's no definition to say this isn't. And so, you know, I don't know that it's malicious necessarily, but it's, um, willfully, uh, applying the license when, you know, or the term when the OSI has even, you know, stepped in and said, yeah, absent a definition, this is clearly not open source anyway.
Mm-Hmm. Um, How much of this is there just concerned somebody's gonna try to fork the model on them and, and then lose control of it? And we've seen that happen in the past with companies where, um, somebody gets mad about something and this the prerogative, the open source community to create a fork, but a lot of companies who are in that space get a little agitated when that happens, shall we say?
So how much of this, they're just trying to kinda prevent a fight before it happens. I think there's, there's a lot of that, you know, especially because their license does have that, you know, uh, that trip wire for when you make a certain amount of revenue, um, all of a sudden now you need to, you know, pay for a, a license for that model. Um, you know, meta invested a lot of money in developing and training this model, and they want to, you know, they want to kind of build a moat around it as a business, you know, decision.
I think that's very reasonable. Um, you know, I, I don't fault them for trying to make money. That's, you know, what companies do.
But on the other hand, you know, uh, I do think there is some amount of, you know, nobody has, um, you know, nobody is owed a commercially viable o uh, meaning of open source. You know, we've seen companies that, you know, make a lot of money selling services and, uh, you know, consulting and hosting and things on top of open source software and do very well. We've seen other companies realize that, you know, Hey, I, I have this open source project that I've built my company on, and Amazon is making a service out of it.
I don't want them to do that. Let me change the license. Um, and, you know, there's, that applies in the AI space just as much and, you know, maybe even more right now because there's such a land grab, um, to try and be the dominant player, uh, and, you know, secure the funding and, um, the revenue streams while it's still, you know, a pretty nascent thing, in least in terms of, you know, the consumer view.
Mm-Hmm. How will the open source community kind of react to all this, do you think? In my experience so far, um, it sometimes takes them a while to rally the troops, but once the troops are rallied, the response can be pretty, uh, shall we say, uh, aggressive, You know, um, open source communities are, uh, are very interesting.
Um, you know, one of the, the analogies I really like to use is there's that scene in Finding Nemo when you have all the tuna caught in the net and they're all swimming in different directions, and then NEMO gets them all to swim downward and they can bust out of the net. Um, you know, I always hesitate, and I fall into the shop a lot, but like, you know, there is no one open source community. There are many small communities, and even within a community, everyone's sort of acting in, in, in their own interests.
You know, not as a pejorative, but just, you know, they're participating because they're interested in the technology and they want to drive it this direction, or they're trying to learn this thing, or their employer is paying them that, um, it takes a while to sort of develop consensus. Um, but you're right, you know, once that consensus is built, it tends to be pretty, uh, self-reinforcing, you know, the open source initiative, as far as I'm aware, does have a trademark on open source, uh, when applied to software. But really the enforcement of the open source definition has been a peer pressure kind of effort.
You know, we as, uh, you know, a community, you know, advocate for, um, using the term in compliance with the open source definition and responding to people, uh, when they don't. Um, if you've looked at the, the tracker for, um, Winamp, which was the software was recently released and they called it open source, and it's very clearly not, um, there have been a lot of issues, uh, and people have sort of brigade in typical internet fashion to, you know, let them know, Hey, yeah, this is an open source, and stop calling it that. Um, you know, right now it's hard to tell, uh, what, where the momentum will take it.
Um, you know, the OSI is firmly behind the draft definition, and there are people who are fully supportive of it. Uh, there are some who are very unsupportive of it, and it's hard to get a read on if they'll eventually be like, well, it's an imperfect definition, but it's better than no definition. Or if they will, you know, continually continue to hold out against it.
Uh, so there are really kind of two ways that it could go in the future. 0 really is insufficient. 0 that is more strict about the data sharing requirement.
Uh, or, you know, you know, option one B is that people will see, okay, well, while philosophically this is wrong, the practical impact has been, you know, ne negligible. So let's get on board because this is better than nothing. Um, if the two camps can't come to some consensus, I think there's a risk that the open source AI definition will just be irrelevant.
Um, it will just be a doc and on a website that some people point to, and the industry as a whole, uh, chooses to ignore. All right, folks, you heard it here. This is, as they say, developing story, but that's pretty clear that the open source AI community's looking for a few good nemos to help lead the fight.
Hey, Ben, thanks for being on the show. Thanks for having me, Mike. It's been great.
All right. And thank you for all watching The latest edition of the Techstrong AI video series can find this episode, others on our website, we invite you to check them all out. Until then, we'll see you next hour.