GenAI Innovations with Hakim Hacid at OSS Seattle 2024
The Technology Innovation Institute’s Hakim Hacid discusses the institute’s role within the Abu Dhabi Advanced Technology Research Council, focusing on generative AI and its application in various sectors. Dr. Hacid emphasizes the importance of leveraging existing LLMs and building ecosystems around them, highlighting the need for human involvement, especially in critical areas like security. He also touches on the challenges of bias and safety in AI models, underscoring the necessity of interdisciplinary collaboration to address these issues.
Transcript
This is Techron tv. Hey everybody. Welcome.
Mitch Ashley here with, uh, Techron Research at the Open Source Summit. As I've mentioned, we have such a variety of, uh, topics and folks that are involved in open source. It's one of the real great things about this kind of a conference.
It isn't just down one avenue of open source. And I'm joined today by a really special guest. Hi, um, Hakeem Kasi.
Yes. Welcome. With the, uh, technology Innovation Institute.
Correct. Good to you. From Abu Dhabi.
Tell us about the institute first, a little bit about, folks may not know about that, and then we'll get into your area of, uh, specialty. Absolutely. So, uh, coming from the Technology Innovation Institute, that is the research arm of the A TRC, the Abu Dhabi Advanced Technology Research Council.
So either A TRC, we are actually three entities. So we have Aspire, that is the business development and the program management of A TRC. Then we have the TII that is the research arm of, uh, uh, A TRC.
And then we have Venture one that is more into the commercialization and the, uh, business side for, uh, for a TS. Interesting. So you're kind of at the front end of the research going into maybe things that are Absolutely.
So the research is done on our side, and then with, when it's mature enough, it goes to venture one that does the commercialization and the business side around. And the TI we have, uh, 10 research institutes and where variety of technological, uh, areas. So we have, for example, can be robotics.
Uh, we have crypto, uh, we have biotech with space, and then we have the ai just to list This AI there. You get to the, which is your, Which is, uh, my, my What's your title with which your responsibilities there? So I'm the executive director and the acting chief researcher, uh, currency.
So I'm leading the force around generated AI and, uh, specifically on, uh, building or improving the Falcon motor that is coming. Fromt, Fascinat, I, I've got a lot of questions for you. A lot of things I'd like to learn from you.
Sure. Um, and thank you for, you know, doing research in this area. A lot of folks are, are grappling with not really understanding LLMs.
Um, do I need to develop my own? How does that fit into my own software development cycle? Sure.
Um, should I just rely, should I look, look at the large, very large LLMs, or maybe I should look small, um, or domain specific? How do you kind of break down that Yeah. That area of generative AI so that people can understand this is the right path to go down under different situations?
Yeah, so you have different perspectives, of course, when it comes to the, to the LLM, but there is one thing that is important to keep in mind. Building an LLM is not a cheap operation, so it costs a lot. So, uh, starting like with the thoughts to build my own LLM as a developer, for example, may not be reasonable.
So, uh, to build those other alarms, you have many, you need to have millions of dollars because in terms of compute, it's extremely Deep not for the fade apart. Absolutely. So it needs Yeah.
Or or the underfunded. Absolutely. So, uh, usually, so for big entities who are, who want their own l lms, that's dual.
So the data is there. Most of the data is public. They can build their own LLM, but when it comes to the development to the open source community, you would say it may not be worth building an LLM from scratch.
So taking existing l LMS and then, uh, building on top of that by doing some fine tuning, for example, it is much cheaper and much more straight to the point. So, uh, we have actors like TAI, uh, who are building those models, the core model, and then we have the developers who can come and then add some specializations of those at lamb through, uh, through, uh, the fine tuning. And you have all the applications that would come on top of that.
So my recommendation is more if you do not have the budget to build, uh, your own NA, I mean, just take the a lambs that are existing there, especially on the open source work, they are accessible. You can see what's going on inside the LLM, which is transparent. Uh, take one of those build, uh, get your data, fine, tune it, and then build your application on top Of it.
I think many of us know about hugging phase as a source for LLM. Is that the best source for open source or kind of individual projects are the right people? Yeah, you have a lot of actually, uh, locations on the internet that, uh, that are sort of providing those open source hagging face is something that we have been using now for, for for long.
But you have small projects that are coming on. There is a huge diversity actually, uh, where, where you can get those s We have also the hyperscalers, like AWS that, that are offering those, uh, those models. So you have different parts on the internet where you can get those more.
Fantastic. I'm, I just having a conversation with, um, kind of DevOps architect person earlier talking about, uh, we, we seem to be in the era of more copilots than you know, what to do with when it comes to generative ai, which kind of seems like the entry level for adding ai, generative AI to a product or technology. What do you see as the next evolutionary step that companies would take if they've gone down the chatbot or the copilot route?
Yeah, one of the things to the needs that are there all the time, and we need to find a solution for that, for it, is basically the reasoning capability of these models. So nowadays we are more into this generative dimension, uh, where the model is added to generate either text or image or, or, or a video or a sound or something like that. But you don't have really a huge reasoning capability.
So down the road, this is absolutely critical for all the applications that that, that, that we are building. And if you want to get the LLMs within the critical infrastructure, the reasoning has to come there. So this is why most of the community around the LMS is exploring, uh, this, this, this, this, this path to find ways on how to integrate this reasoning to make sure that the LMS will not just generate content, but somehow they need to think before they generate the content.
Yeah. Well, does that help with, well, hallucinations on one end of the spectrum, but also more deterministic outcomes? Because there are situations where there may be a lot of probable outcomes, but I need kind of reliable Certainty.
Yeah, we need, we need a lot of certainty. And then, you know, like what we see today is mainly the, the model itself, but we believe that the model itself is not the only ingredient to solve all the problems that we're facing today. So we need to build an ecosystem around those, uh, models that will help actually understanding better, for example, the prompt of the user, right?
Mm-Hmm. So the model itself alone may have some understanding, but you need to build around that some small applications, small systems that will be actually helping into decomposing the front to build this sort of a better understanding of that. What is that part of the, I don't know if you consider it training, but isn't that part of engineering?
What kind of prompts to expect and how to respond to them so you know how to interact with it either interactively or through an API. So You, you can do that, but you have also another part that is, uh, basically you, you have some parts of the prompts that could be understood better by an external system, right? Mm-Hmm.
So, uh, using, for example, knowledge grass, right? So you can reason on the pro before even you send it to the LLM to rewrite it and to do this kind of stuff. So the model itself may not be the answer for everything, but the ecosystem that will be built around that will be actually supporting and improving this, these models.
And I, I will not be surprised to be honest, all this proprietary models, because we do this comparison sort of straightforward comparison in terms of, uh, quality saying, you know, the closed models are much more interesting in terms of quality, but, uh, as a matter of fact, we don't know what is hiding behind those closed models, right? So, and my intuition says that, you know, you don't have only the, the model that is behind that. So you may have an ecosystem that is composed of the model plus different small applications, small subsystems that support actually the LLM to reach that quality, that, that, that are close source models reaching to That.
Interesting. Yeah. I, I think that's maybe for the folks that are newer to generative ai, the assumption is that, hey, if I can take my own unique data that's specific to my business, you train the model or build that into an LLM, that's one of the ways that I can do something that's differentiated from my competitors, for example.
But it isn't always necessary to go to generative AI route. I mean, you can do other things, ml, other kinds of extra assistance. Very good question actually.
So, uh, most of the people, I think they have, uh, extremely high expectations from gene to ai. They see it as a solution for everything. I do not necessarily agree on that.
'cause you have problems today that you can still solve with traditional machine learning, right? So I think everybody is going this way because there is also, uh, high expectations, uh, so from it, and there is also, I guess some, uh, wrong marketing that is done around Very high visibility, right? Absolutely.
Used to call that the airline magazine phenomena. Like, my boss read this in the magazine airplane, so we must need to do it. Absolutely.
So everybody wants to integrate that. Uh, the use cases, honestly, most of the use cases as of today can be sold by the, uh, traditional ai. So there are some use cases that become interesting, other, for example, the, the call centers use cases interesting enough that it's, it's a chart is in a direction.
So you can integrate a lot of, uh, generative ai. You can have things in the parts of the healthcare system, the legal systems, and you can have many things that you cannot solve with the traditional one. But when it comes to I, uh, predicting the weather, probably you don't need to integrate, uh, straight away the LLL.
We already have non-deterministic models for that. Exactly. That's, that's, that's, uh, that's the thing.
So getting these models into production will still take time. Getting them into, uh, critical infrastructure will take time. Mm-Hmm.
It'll take, it'll happen, but it'll, it still needs some time in, in terms of research. And the generative AI is not sort of the answer for all the problems that we have. It'll come at some point where it'll solve some problems.
And al problems are probably, uh, we don't need to deploy a, a monster like an LLM just to solve the same and predictive task at the end of The, so don't rush down the generative AI path, just 'cause this is the new shining thing. Absolutely. There Lotions.
Yeah. Talk a little bit about, I'm really interested in, you know, we're, we're, we're trying to, um, increase velocity of how we deliver software, improve the security of it, um, be able to respond to both business needs as well as security, whatever it might be. How does introducing AI into that workflow work?
Because, um, AI depends on data Yeah. But in its own unique way, as opposed to traditional database or data resource we might access. What are the unique characteristics of ai, whether you're doing AI ml or generative AI as you're thinking about delivering it with other software, either at the same time or at its own interval?
I think the, one of the domains that, uh, we, we see is sort of, uh, a big impact when it comes to generative AI is the software industry. So, uh, but I, I don't see it as replacing co literally the way we are doing today. Uh, the, the software of like productization and this kind of things, the LLM is actually learning from data that is out there generated by someone, is someone, is is the human, the Human right?
Absolutely. So, uh, in terms of efficiency, things would be much better. I think AI would be helping out there.
But when it comes to, uh, security for example, there is a need to have humans who will be involved with the ai. So who will be delivering much faster, better quality stuff. But when it comes to critical situations like the security, especially the security part, I think we still need to keep the human in the loop just to make sure the machine is doing the right thing.
Yeah. I wouldn't go to auto copilot quite yet. Uh, not yet.
Let me ask you kind of a theoretical question. Do we create a feedback loop that's kind of non-productive if at some point we have so much generat generated, uh, software, uh, available to train models against that we're act the, the LMS are using not human created software, but generated AI generated software. Now we have AI generating more AI generated software on top of it.
Does that create a kind of bad feedback loop? Well, if, if we rely on the studies that have been done so far, uh, one of the studies that, uh, our colleagues have done in the, in the research center where we, when we are this generat generated data is not actually adding a lot of value. So this is why This already exists out there.
Absolutely. And, and, and then you, you have an sort of a, a phenomena where actually a system is generating things and then it generates feedback and things that it, it has generated. So everything becomes artificial.
And then you need to measure actually how close that artificial thing to the reality, right? So if we keep like the system generating data and then evaluating some feedback, that data probably it'll completely reduce its value that, that feedback loop, right? Again, this is why we need the human to be in the loop, right?
At some point, AI is good at having it fully sort of independent, may not create value, actually underserved. I Think of it as kind of the two mirrors on opposite walls that you see the infinite Continuing Into the Absolutely. At some point you feel that you are doing things, but you are not adding value at the end of the day, actually.
And, and then you may even like theoretically, like reach to a point where you are changing completely what you have learned in the another, because it starts learning on things that are superficial at the end of it. What's your, what's your thoughts on general intelligence, you know, AI reaching that point, you hear, oh, the next version of chat two PT or whatever, where this far away from that happening? What's your view on it?
I, I think there is, there is, uh, an interesting progress that is going on, but honestly, I'm not, uh, confident that we are close to this gentleman, right? So we still, I think the AI is still at, uh, at its infancy. We do not really understand countries what's going just inside the neural network.
Mm-Hmm. We don't know how it works, right? Mm-Hmm.
So you have few papers that came out like a few days back. We have spent time building huge models with billions and billions of parameters, right? So the papers and this, this research work has demonstrated that with 50% less parameters, we do exactly the same.
Uh, sort of, you reach exactly the same quality that you reach with a hundred percent of the parameter adding. So we still do not understand much things inside. And again, the fact that these models are able only to generate not really to reason on things here, makes us a little bit far from having an ai that is how we say that is taking over the world, right?
So we're not there yet, maybe five, 10 years down the road we reached to that point. And maybe the path is not necessarily the LS maybe it's something else. Could be something else, you know, um, Netflix does popularize the, we introduced the idea of the three body problem, the people of three objects in space influencing each other, that you can't create mathematics Yeah.
To predict where, what will happen with those things. Does that kind of, do we face that problem with AI that we may never really ever be able to understand how it's determined what the output outcome is? I think there is progress.
There is a hope that at some point we will reach, uh, that understanding. Uh, but for that we need a sort of an interdisciplinary approach. So we need to have people from different perspectives working together to analyze and understand the limitations.
And then eventually the, the, the frontiers of that, of, of, of the problems in which we are. So probably it'll take time. Uh, it's not an easy problem, but there are a lot of people working on that around the world, and things are going extremely fast.
Uh, if you see the, just if you look at the statistics, the, the level of adoption of this AI and the interest number of papers we see coming out every week is extremely huge. So this maybe gives, uh, a positive hope. I would say that we'll progress quickly on this kind of problems and we'll build quickly a better understanding, maybe not the full understanding, but a better understanding pretty fast on that thing.
Good. What, what's, what problem is capturing your attention right now? What are you thinking working on?
Well, we have the, we have the reasoning. We look at this, uh, pretty close-knit. Uh, we explore it in different ways, uh, from different perspectives.
Uh, and then we have the safety and the bias of these models. Big topic. Yeah.
We have both of 'em are, yeah, exactly. Like the previous generations of these models, we have been focusing on, uh, building the model. It's like demonstrating to the world that, you know, we can build the model.
And I think all the folks who are working in the, in the area of the lms, were trying to demonstrate the same thing. We can do it for now. We have done it, but we never thought about the safety, right?
So, Uh, yeah. Just 'cause the Wright brothers could fly doesn't mean we're all gonna jump on board. Absolutely.
You could go someplace quite yet. Exactly. Not yet.
Yeah, absolutely. So I think it's an, uh, an important problem and it's, it's, it's time to look at this, uh, safety issue. Uh, because we, we are reaching a point where the quality is increasing, quality of these models is, is increasing.
And there are some use cases where we cannot integrate this technology. So, and the integration comes with safety. There is no safety, there is no integration anyway.
So it's important to build into this. Uh, you have the bias. So, uh, the bias, you know, like, uh, it comes mainly from the data in my opinion.
And the data that we're using today is mainly coming from sort of one location of the world. So maybe it's, it's interesting to look into how to get some, uh, cultural sort of, uh, minority also considerably this kind of model, uh, and have their word to say somehow in this. Fascinating.
It's been a pleasure talking with you about, it sounded like we could like go out to dinner, have drinks afterwards, cigars talk about this for, you know, hours and hours. I've enjoyed doing that. You, So Thanks a lot, HASI.
Very nice to meet you. And, uh, appreciate your work and research and helping advance the state they are. And solve those tough, well work on those tough problems.
Absolutely. With a lot of others. Smart.
We continuous thank you very much for your time. Thanks for the discussion. I hope we would have that opportunity again.
You bet. Fantastic. Thank you very much.
So where are you gonna hear a conversation like this? Right? You, you, you don't walk up to the water cooler and suddenly bump into a hot team every day.
So it's a great pleasure and a privilege to be talking with folks like Kim. We'll be back with more interviews, so stay tuned with us on Textron tv.