Embracing Sustainable AI Mini-Models with Thomas Wolf | EcoTech Insights
Hugging Face, a leading AI company, has doubled its valuation to $4.5 billion after a $235 million investment from major tech partners. In this exclusive interview, sustainability analyst Bonnie Schneider talks with Thomas Wolf, co-founder and chief science officer, about transforming from a gaming startup to a key AI hub. They discuss the company’s shift towards smaller, sustainable AI models, its community-driven culture, and efforts to make AI accessible and environmentally conscious.
Transcript
Hi everyone, I'm Bonnie Schneider and today I am pleased to be joined by Thomas Wolf, who is the Chief Science Officer and co-founder of Hugging Face. Thomas, it's so great to have you here. Hi Bonnie.
Yeah, happy to be here. Well, tell us about your background and your journey with starting this company. I Was first a researcher in, in physics, quantum physics.
I worked a bit in the US at the time in the Bay Area as well in the Lawrence Berkeley National Laboratory. I, I joined force with Glen and Julia, my co-founders to create hugging face. And in the beginning, hugging face was a game company.
So that's why we have this kind of un serious name that we are, that we have to have, I think today, which contrasts maybe with many companies in the field. And we pivoted this idea of building a platform for open source, open science, open knowledge and ai. I wanna take a deeper dive into the mini models that you mentioned because this seems to be a trend across, uh, many I AI companies, particularly in the aspect of sustainability.
Can you explain the models that you're introducing and why you think this is part of a, a bigger picture for sustainability? Yeah, I think in many case, I mean it sounds quite obvious, but surprisingly we still have to say it. You don't need a model that can solve very complex mathematical problems.
I don't know the reman conjecture to just help you in your daily automated task, right? You just want something that can batch translate something or can classify some text or can answer some simple question about your documentation. All these things you actually don't need the GPT for 1 trillion parameters model.
They can be done really well by, by a smaller model. And I think it's something that's not pushed a lot by close source providers because they mostly would like to simplify their stack and also they want you to pay by token, but if you have a small model, so first it's quite cheaper to run. And the second is that you can even have this model run locally.
So instead of, you know, having this model running on, on data centers, uh, a set of models that you released this year, actually this year, last month is, is called a small lm, which a set of of much, much smaller model that you can try on the Hagi face hub. And they work really well for many tasks and they can run in browser. So you don't even have to send your data out of your computer, it can actually just directly run inside your browser.
It's small enough, it's smaller than a, than all. They always stream from Instagram for instance, like a couple of hundred megabytes basically. And I think this product can be really useful for many tasks and for all of these tasks that we automate with smaller model, the energy that's required is much smaller.
For instance, where you have like solar energy in the mix, you have much more freedom to um, tailor them so that the costs both environmental and monetary is, is lower. What has been the reception for these smaller models from the people that you're talking to? Really great.
Yeah. A lot of excitement are actually surprised because I think a lot of the reason we didn't use them is that people thought it was not possible. It was not possible to have small model with interesting capabilities.
And that's the discovery of the past 10 months maybe, or at least at hugging faced the past eight months is if you're careful if you apply some of the learning from the, the large model, if you adapt some of this learning, you can actually have this model with very interesting capabilities. So great reception I would say both from public perception also from uh, companies. They're very interested also in the data privacy Id having modeled that run locally is is very interesting for many of them.
So yeah, that's why we are doubling down actually on this direction. We built up the team around small model even more. Your company has grown so quickly, so fast.
Uh, can you just tell us a little bit about what that's been like for you? I would say yeah, so fast, I don't know because it's still, it's already seven years I feel like a dinosaur, but definitely it's been a wild ride. Some breakthrough were maybe more in adoption like the widespread adoption of ai.
It is still something really incredible in terms of size. We've been very careful not to grow too fast. We know that open source AI is still a kind of a long term battle.
I think just like Linux Unix versus kind of Microsoft discussion in the early 2000, uh, really believe that in the end all these models will, a lot of them will just be open source and it'll be like in software where you run basically Unix everywhere. Right now hanging face is a little bit over 230 people, which may sound reasonable, but actually for 80 old company that's raised as much as we did, it's fairly small and we want to keep it fairly small. We've been managing to go through these various waves of craziness and still stay really excited about what's happening.
How are you building this community where it's just part of the, the culture and it's growing so much at Honey Face? I think it's quite natural, which is surprising is we don't have any community manager like you could think, right? With a company that's so centralized around this notion.
A part of it is we think if you try to lead by example, if you try to at least do things that you think are right, so for example when I was saying we, we train model and we try to be as open as possible on how we build them, how much it cost us in energy, where did we build them. I'm very conscious about this. I think it gives people a positive example of what you can do.
And what we saw is actually community is usually very excited about that and they join, they want to participate. A lot of people in AI actually want to do good and they would like AI to be well integrated and to be a positive force in the field. And a lot of people are also conscious about all the other challenge we have in today's world.
And so if you show kind of positive example and not just kind of a dystopian future basically, uh, people are actually very attracted about with with that. I think that's one of the thing about our community is that it's a reflection of the mission of hugging face, which is interesting in a way, which means you don't really, you don't have to think about the community as something you consciously craft or do, but if you have some specific values, very often they will be reflected in the community that surround you and That that's probably part of the success that you're saying it's leadership and then the community reflecting that same value system. Yeah, I think the same probably in the other large venture project like Wikipedia where you feel like this kind of early mission values around sharing knowledge and and welcoming everyone and you and trying to have like civilist or like very way to discuss and also tackle complex problem and, and energy consumption is one of this complex problem, right?
Because the best would be to not train any model and not consume any energy, right? But the most interesting I think is when you have some discussion and you just don't take one of the sides too strongly. One of the side that say we should not do any AI because it's bad for like consumption reason or other.
And the other side that say we don't care about the environment, AI is gonna solve everything. And usually we be a bit somewhere in the middle where we like understand both things and we say, okay, maybe can, we can find a way to train AI that's still interesting in terms of performance and what it can do, but also take into account the cost and how it should be done. So trying this middle ground is a lot how we try to navigate the AI world today.
Can You tell me about some of the partnerships that you're working on right now? Yeah, thousands of partnerships, which is funny 'cause we have a rather small team, as I said 200 people, but our Slack has maybe several thousand people and it's kind of Slack. But we have partnership with so many companies and, and, and research lab and, and industries that it feels like we're sometime a company that's much bigger than we are.
So some of the one we have. So, so a lot of them are very visible in terms of advertisements. So recently we had, we announced partnership with Google Cloud and we have partnership with most of the cloud providers basically out there in terms of getting make, helping them give easier access to all the models that are on the hub.
And we also have many partnership with more recent research institutions. For instance, one I was just working on this morning is one with with ETH and, and the Swed community where they have a very new large cluster and they're interested in pushing for more multilingual. 'cause one of the challenges in AI today is also how English is kind of dominant and then some lower resource languages struggle to get the good performances or, or definitely get less attention.
So we are focusing there on trying to bring some very high quality but also very large uh, pre-training data set in many languages. So I think that's quite interesting and that's something we, that's something personally coming month. It's actually almost ready.
So I think next month or so there, there will be a big release around this. Yeah, it just seems that when you're training AI to understand the nuances of language, that seems like it would be challenging because the meanings are different as you go from language to language, even just expressions and things. So that just seems kind of challenging.
Yeah, yeah. It's super interesting because some words don't have translation and you know, it's really common in language. You have a word.
And actually there is no direct equivalent in both in English, but uh, in, in many languages. And AI is a, is able to do this kind of translation. That's also one of the, one of the very interesting things about AI I would say is, is kind of this, not replacing people but acting as a link between people and translation is obviously a place where if you can translate well, it actually helps a lot communication between people, either actually as a translator but also even as a tutor.
It's also super interesting in terms of what possibilities for education. There's a lot of very interesting use cases there. So I'm quite excited about this one.
Also exploring, which is also a collaboration between us, uh, and startup called, called phys is around training models for, um, material, uh, discovery. Basically these are AI models. They're inspired by the most recent transformers and they try to model interaction of atoms or molecules and I try to help design, for instance, better batteries or better electrodes or better materials to, to how you say that, to capture carbon dioxides and all these things.
And here it's maybe one of the area, if you think about it, where AI could be the most impactful, which is helping us make new scientific discoveries that solve well the challenge that we have today. And uh, that's if we discover some specific material that make batteries and times more efficient or that then captures carbon, it times better on what we have today. The impact will be tremendous.
Even much more than a chatbot. And we already see the impact of yeah, chat buds, right? I'm very excited about that.
But obviously it's quite more, the road is probably still longer And that's kind of leaning into your background in, into science. Yeah. Yeah.
It's funny because I almost did my PhD in GFT, which is exactly that, but before G years It's all full circle. Yeah. Can you tell me what's coming, what you're working on now?
We're almost, we have more quarter left of the year and going into 2025. What's ahead, I know everyone's probably asking you that, but I'm really curious for hugging face what you're working on now that you're progressing into the future and you're excited about. So I would say, uh, we have very interesting things coming in in open source robotics.
So very cheap low cost robots that you can build yourself and that can help you help you for task like folding laundry or emptying your dishwasher. So this is something that's quite fun to do as well. To be honest, we have this project around on a material prediction.
I'm, I'm quite excited about this one. There was basically a breakthrough that not a lot of people saw at the beginning of the year because it's quantum chemistry and I don't think a lot of people followed that but breakthrough where AI started to really work for this, everything around LLM and basically opening the black box. I think it's really necessary today in this world.
So, so we have a lot of things coming on there explaining how to train the best LL M1 and with a special focus on this small model. Really excited about this as well and I think we keep pushing that for the coming year, smaller and smaller model and basically do, most of you can, most of the thing you can have with LGBT but just running live in your browser locally or whatever you want to choose where you think is the best for you. And we have some workers on multimodality, we've been pushing a lot on speech because I think it's almost ready, it's something that can almost work.
Um, by which I mean talking to your uh, talking to your computer, yeah. Which was the computer actually understanding and reacting in, in a very low latency way. We have a new project video which is also super interesting, a bit more exploratory, but it's quite exciting to see this one coming out In AI in general.
Everyone's excited about it. A lot of money is going towards it, some people are wary of it. Is this just a, a bubble, is this a phase or what, where do you see AI as a movement going in terms of positives and negatives?
'cause you're on the inside of it. Yeah, I think it's shifting progressively. I think there was a lot of interest in model building in the beginning and, and the most of the most famous companies are, are still model builders I would say.
Like open ai, entropy, all these, I think progressively people, we understand that AI is much more than that and just when the early internet, maybe the first focus was on the people who are making building website for Azure and that was the first thing people thought, oh, this is where all the money will go on the web like building website for Azure. And now people think, okay, building LLM is where all the money in AI will be. I really think this is wrong.
And people will just see that this, this is gonna be just commoditized and just like building website is, is far, far from being the most profitable thing you can do on the internet today. We'll find much more interesting way to use AI and to integrate it everywhere. And that's where I think some of the largest company that we don't know yet will appear.
Just like we have, I dunno, Airbnb like this huge company and at the early internet, I think nobody thought that this would be maybe large strip know this would be the largest company built on the web, but you need this time. So I think we only now a little bit in this moment where people understand, okay, this is maybe not all the thing you can do with ai. Maybe you actually can use it and maybe integrate it in smartly at this or this location or in this, on this vertical or creating this new thing that we never really thought about.
This will be the use case that will actually change the world or that will be the most well known effect of AI in a couple of years. You Started off by saying that hugging face is such a unique name, it's so warm and friendly, but it started off because you were in gaming. Are you happy that you kept the name now that the company's evolved?
Yeah, very happy. I think it's also a good reminder for us that whatever we do, we should not take ourselves too seriously as well. I mean, it's also nice to enjoy what we do and that's one of the reason I feel very lucky to be here.
Even after some time in the field that I still very excited about the project we work on. I still find them really interesting. Maybe one lesson for people who are in ai, it's still quite a fun field.
What we do is really something amazing when you think about it as a little kid, right? You can have this machine that talk to you like they are almost like little characters and they, and we keep this like freshness of being amazed, uh, by AI and not take this too, too seriously. At least that's what we try to do at Target phase.
Yeah, That's great. Well, Thomas Wolf, it's been a pleasure talking to you, co-founder, chief science officer at ve. Thanks so much for joining me.
Thanks Bernie.
