The Future Will be Open: Open as is Open Source AI With Dr. Ibrahim Haddad at AIE 2024
Over the past two decades, open source software — and its collaborative development model — has disrupted multiple industries and technology sectors, including Internet/web, telecom and consumer electronics. Today, large-scale open source projects in new technology sectors such as blockchain and artificial intelligence are driving the next wave of disruption in an even broader span of verticals ranging from finance, energy and automotive to entertainment and government. In this talk, Dr. Haddad will discuss the efforts of the LF AI & Data Foundation in supporting the development, harmonization, and acceleration of open source AI and Data projects. He will provide a preview of some of the industry challenges in these domains and present trends that will manifest in the next few years.
Transcript
Hi, everyone. Thank you very much for the invitation to speak at the event. My name is Ibrahim Madda and I am the executive director of the LFA and Data Foundation.
And the talk I will discuss about the future and how the future looks for open source ai. However, I will start by talking about everything we're doing in LFAN data. I will speak about generative AI activities.
I'll speak about the survey we've conducted on generative ai, and then I have an open invitation for collaboration at the end. So I hope you enjoy my talk and let's kick it out. My first opening slide, which is the one that I, I really enjoy talking about, uh, in all my presentation, is a capture of the growing ecosystem of open source ai.
So what you see on the screen is the L of AI and data landscape, and it captures over 360 projects that are key and critical in the open source AI ecosystem. These projects together, they represent the work of hundreds of thousands of developers coming from thousands of organizations. Can you imagine that these projects together every single week produce one new million line of code net every single week?
It's amazing the pace of innovation. However, despite all of these activities and the growing, uh, projects, there are a lot of challenges in the ecosystem. The five challenges you see on the screen are in relation to fragmentation across these different categories.
There are limited capabilities and limited collaboration opportunities across these different projects. There are governance challenges in relation to who runs the project and who manages, uh, innovation across the project. Uh, in addition to challenges with respect to managing the project assets and the fact that a lot of these projects initially started as internal activities of development to their companies to develop specific product requirement.
And over time, companies realize that the value is actually not in the platform or the framework or the library. The value is actually in the data and the models and the apps that leverage the data and the models and decided to open source. So then we ended up with a lot of redundancy in the ecosystem, a lot of fragmentation and a lot of activities that look very similar and project producing similar output.
So with that LA and data was created to address these different challenges in the ecosystem. Today, we gather over 70 member organization. Uh, you see most of 'em on the screen.
Um, this is a little outdated, but we have a few new members that are not there yet. Um, and the members we have are great, uh, slice of technology across North America, across Europe, uh, in Asia, and 10 stock is in China. We have, uh, several board members from China.
We have many companies that presented in LFA and data from China. Uh, we have opo, uh, for Paradigm, uh, ZTE, Huawei, Alibaba Ant Group, uh, and several others. Uh, in addition to even some organizations that are nonprofits, large universities that are activists.
So we've grown our memberships from nine founding members to over 70 organizations today. And we proudly host over 60 technical projects. In fact, many of them, uh, at least 10 of these projects came, uh, from China.
And what you see on the screen is a collective effort of thousands of developers, uh, putting together their efforts and creating these projects. And we host them across three different levels of incubation. At the base layer, we have sandbox projects followed by incubation level project and very mature, widely adopted and deployed at large scale projects that we call graduated projects.
Where do these projects come from? They come from industry leaders that you see on the screen that trust the Linux Foundation and LF AI and data to help their projects grow and become de facto standard in their own categories. Um, and in fact, uh, you may be able to recognize all of these logos.
They are the top tech companies across the globe, uh, some of them from North America, from Europe, from China, from uh, South Korea, uh, from Japan, uh, and several other geographics. Uh, so we are able to prove to these different organizations that we are the right entity for hosting and cultivating a strong developer community and encouraging collaboration and growing these projects in the adoption space. One great way to show that is the commit growth.
So what you see on the screen are basically the last three years from the fourth quarter of 2020 to the fourth quarter of 2023, we've had over 255,000 commits across our project. You can see the trend is coming up. Uh, similarly, we are able to add a new developer every single month to our collective developer community.
And as you can see, we have over 30 active, uh, 30,000 active developers or contributor across our projects from an organization engagement. This is actually very interesting. When we look our our projects, we realize that there are, um, over, um, 400 companies that are participating with us in contributing code to the project in the past year alone, which is really massive.
And if we look for the past five years, we've had over 650 companies that contributed code to the projects we host. So you're able to see growth in code, you're able to see growth in the number of developer active in the code, but also equally important, you're able to see more organizations joining the development in these projects. And we have created over the years a great program that enables projects to grow and increase their contribution to the project.
We offer multiple support services for our communities, and I would certainly urge you and invite you to consider LF AI and data for hosting your AI related project. And please reach out if you'd like to discuss our programs. Uh, furthermore.
So with all these different programs, our work is focused on addressing the key ecosystem challenges. We are minimizing segmentation. We are minimizing fragmentation across the industry.
We are increasing collaboration and integration across these different projects that we host. And between the projects that we host and the industry projects, we are a not-for-profit organization. And we offer a technical open governance for all of our projects.
And we are supporting the projects to grow functionally in terms of feature set and, um, providing what the industry is missing in terms of gaps. And we provide safe haven for all the project assets as they are hosted in the foundation. So there are a lot of opportunities for open source AI development.
These are opportunities in relation to research and innovation when you collaborate with others. Others can be commercial companies, can be r and d companies, can be government entities and research labs. Uh, there are a lot of opportunities for collaboration and knowledge sharing, which is very critical in such a really important domain.
There are opportunities for additional transparency and accountability. There are multiple ways to democratize access. So when you think of open source ai, not a lot of companies are creating AI technologies.
Plus every single company out there will be a user of these different technologies. So by opensourcing AI technologies, we are able to democratize access to this new innovation and new technologies and allow companies and even individuals who don't necessarily have expertise in AI to use and leverage these technologies for what they're building. Uh, there are definitely a lot of opportunities for federated LLMs, and I'll talk about that later on.
Uh, and certainly open source is a major force in minimizing fragmentation by rallying industry players across the set winning of technical open source projects. So why you should participate in open source AI development. Certainly there are many, many reasons why organizations and even individuals or even students and researcher and open scientists should participate in open source AI development.
Um, these are, um, very true for any type of technology and even more true in the sense of AI for reasons related to transparency and neutrality. Uh, I will certainly not go through each one of 'em. Uh, this will become a very long slide.
Uh, I leave it to you to agree, um, read it later on. But basically there are a ton of, um, participation, um, arguments if you wish, on why an individual or a company should be part of ai. And very similar, uh, with respect to data.
Uh, there are a lot of licenses today that help support open sourcing data sets. Uh, I would like to point out the CDLA license that was created by, um, the next foundation, uh, and its global community of organizations that came together and decided, uh, and realized that there's really, uh, no specific license targeted specifically for data. Of course, there are the open source, uh, OSI approved licenses, but these are more meant for source code and not data.
So we have A-T-C-D-L-A license today to adjust that gap. Um, we are able to offer data governance, uh, a governance model, uh, and structure for, uh, data sets and, uh, uh, data lakes. Uh, we have the ability to provide, um, diverse data sets, uh, that are under open source license.
And certainly there are dozens, if not hundreds of, uh, different tooling targeted for data, uh, to help you, uh, anonymized data sort dag, uh, and apply multiple techniques, uh, that will help you prepare data, uh, for the training purposes of a machine learning model. So all of this together, um, when we look at all the different activities, the next natural step for us in LFAN data was to launch a specific activity for generative ai. Um, and, uh, we had a proposal, uh, published, uh, to the board from one of our members, uh, Jerry Tan, I would like to say hi to you if you're listening to this.
So Jerry submitted, uh, a proposal to kick off a generative AI effort, uh, that was approved by the board, uh, in about four or five months ago, uh, towards the summer, um, of 2023. And today we have have over 100 people across over 40 companies that are participating in degenerative ai comms. It is a non-profit organization similar to, um, um, the talent, uh, organization, which is LAN data.
We are a neutral organization with credible and we apply open governance principles. So all of these different peoples and over 40 companies are focused on five different areas, emerging l and m architectures, curating open source projects, and focusing on security, uh, and privacy and ethical ai. We have a specific activity focused on models, we call it model centric, that looks at hosting models, collaboration on building foundational models, and hosting ai, generative AI tooling.
We have another thread focused on data, the fourth focus on education outreach. And the last one, which I'm personally very active in, is called the openness model framework, which is really a framework that will help us evaluate how open a certain model is given the license of the many component that makes it, uh, a model. Um, so we are getting to publish the paper, uh, of our proposal, um, and we will be, uh, announcing this on our website, on our social media, uh, most likely by mid-February.
So, uh, you should be able to, uh, hear about it, uh, no matter what social media or platform, uh, you are using. And very recently, we, about a week ago, uh, mid-January, we is made available the results of degenerative AI survey that the links foundation and LFAN data conducted, uh, towards the end of 2023. Uh, so the report, it's called 2023 Gen AI report, is available from the Linux Foundation website, and I would urge you to download it, and you're able to download the report and you're able to download all the different graphics used in it and use them in your presentation.
Uh, it offers a very deep detailed across all the different people, our, our organizations that we have talked to. But for the purpose of this talk, and because we have very limited amount of time, uh, I will only showcase three or four findings and results from this survey. The first one is, as you can see on the screen, open source ai, specifically Gen AI is considered better at supporting collaboration, innovation and ease of integration than proprietary solutions.
Certainly, we knew this because we know open source and we've been working on open source for decades, but this is a great validation that not just us, we see it this way, but also the people in different commercial entities that actually are decision makers in the adoption of generative AI technology see on the same path. They realize that for generative AI to be successful and to be more adopted by them, it has to be open source to allow more collaboration, innovation, and ease, uh, of integration as shown on the screen. The second survey result that I would like to highlight is, as you can see, open source gen AI leads or results in increased data control and transparency.
So when you make your, uh, generative AI efforts code and data open and available under open opensource license, absolutely that leads to additional transparency and gives you as well control, uh, on the data because you have access to it and you're able to use it as well. The third result here is certainly neutrality. We understand this very well at the Linx Foundation.
We are a neutral, not-for-profit organization. Neutrality is one of our key value proposition, and certainly neutrality is a key aspect of generative AI governance. 95% of the organization surveyed, they agree that neutrality is very critical and important, generative ai, and certainly this last key slide that I would like to share with you from the report is the importance of openness.
63% of the people from these different organization were a hundred percent firm on the importance of openness when it comes to generative AI and systems. Therefore, you know, we understand openness is very important and neutrality is very important. Transparency is very key, and we are able in LF AI and data and degenerative AI commons to present our stakeholders and our general community with generative AI efforts focused on data models, open model framework, and several other activities that are neutral open for participation and through transparent open governance, so are able to hit all these different keystone.
Furthermore, there are so many benefits to open source LLM benefits in terms of sharing and collaborating on research and innovation, collaborating on knowledge sharing, being transparent, democratizing access as previously mentioned. Uh, and certainly there's a surge in the federated LLM market in terms of, uh, decentralized LMS and the ability, uh, as well to increase privacy and data control, uh, even though you don't own, uh, the data. Uh, so there are so many opportunities and there are already a wealth of number of startups and organization focused, uh, on many key areas within that space.
Uh, and one of the last thing I would like to leave you is, uh, since we have been talking about open source collaboration and transparency and the neutrality and integration, and all of these different key principles in open development are three principles. The first one is, as a company, you cannot hire all the smart people in the world. You can be the biggest company, you can be the richest company.
You cannot go and hire everyone. However, you can collaborate with everybody, even with smart people who are working for your competitors, and you can collaborate with them under the large umbrella of opensource. And this is really important, uh, and I see this every day where we have multiple companies that come together under the next foundation umbrella and collaborate every single day on creating and expanding open source technologies while on the marketplace.
They compete on selling products and services. And open source is the unifier, all of these individuals and companies. The second principle that we need to embrace is in relation to the fact that open source r and d creates a lot of value.
There are billions of dollars worth of open source r and d available at everybody's fingerprints. And the key here is that organizations with their internal r and d, certainly they, they innovate and they create new technologies and create new values, but also a lot of that is also driven by their reliance on open source technology and innovation. So even though they are not creating, uh, these technologies, which lead us to the third principle, even though you're not creating a project or you're not creating a technology, you can benefit from it immensely.
Um, there there is, um, all the key technologies that you can think of. There's always a founder for this technology, but there are thousands of companies that today rely on these technologies, whether it's no js, whether it's Kubernetes, uh, whether it's the Ker, uh, or Onyx or tors or any of these key technological open source projects. They were created and initially founded by a given founder.
And today, hundreds and thousands of organizations are benefiting from them, although they did not create that piece of technology. And I think once we all come together and embrace these core principles and come together and collaborate on open source ai, we are able to make miracles. Uh, imagine there are over a hundred thousand developers collaborating every single week and creating one new million line of code in open source AI ecosystem.
My goal is to help more and more people get into that ecosystem and have more thousands of developers joining in LFA and data. We started with a handful of developers in 2018, and today we are over a hundred thousand developers that have contributed to our projects. So I would invite you to participate with LFA and Data, and you can participate with us by joining us as a member.
You can participate by hosting your projects with us. You can participate by joining one of our many committees that we have. We have committees focused on security in, uh, machine learning systems.
We have, uh, committees focused on, uh, data operations. We have committee focused on trusted and responsible ai, uh, certainly degenerative ai, uh, BI business intelligence and ai, and several others. Uh, if you are, uh, looking into ai, certainly open source AI is your starting point.
Uh, so thank you very much, uh, for your time today. I appreciate the invitation from the conference founders, uh, and certainly, um, you can reach out, uh, to me personally. org.
Uh, I wish you a great conference and thank you very much.