Building AI Applications with Red Hat’s Tushar Katarki
Tushar Katarki, senior director of product management for the OpenShift Core Platform at Red Hat, dives into what application developers really need to get started building artificial intelligence (AI) applications.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Thar Kaki, who's senior director for product Management for OpenShift Core, and we're talking about how AI will change the requirements for the workstations that developers are using to build applications as we go forward.
Tushar, welcome to show. Thank you. Thanks for having me.
What are the implications of having all these AI models running around and all this data that we're expecting developers to use to build the next generation of applications? What kind, what challenges should we expect? I mean, the implications, let's start with that and the opportunities.
I mean, I, I just think of so many different, uh, ways in which these things can, uh, help, um, users and, um, end users. I mean, you can think of everything from, um, customer service to, you know, answering, you know, um, like, uh, for example, oh, I have questions on tax, uh, or I have questions on, they, you know, uh, my new enrollment period is coming up for, uh, insurance, and I have so many questions. I don't read the whole document.
Or even if I did, uh, I want to some ready made questions, that kind of stuff. So I think there are so many opportunities here, uh, for, to be taken advantage of. I think the challenge, uh, in terms of, uh, is, uh, I would say, uh, threefold, uh, one is, uh, access to the models itself.
I mean, whether you want to, uh, access the model, uh, or as a service, or you want to host your own model regardless, you need to know which model. Uh, and, uh, you know, obviously there are very large gender purpose models, uh, such as chat, GPT, you know, especially the four series. Uh, but also, uh, Lama with the 70 B now.
And, uh, there is a talk of, uh, even larger, uh, model coming from, uh, meta. And then there are other large models out there. So there is that, but there are also, uh, I mean, on only one hand, there are very generalized.
On the other hand, uh, they do not have the specific information that you are looking for. Uh, and, uh, they are expensive to, um, serve, uh, as they call, um, and manage, you know, uh, and I think so, uh, I think there is, uh, there is that. But then also, uh, there is also like, uh, a question of, you know, what are the license associated with this, right?
Like, you know, and what is the indemnification associated with that? So there are a lot of questions that you gotta think about. Uh, who will I get sued if I use these particular models?
What is the indemnification that I'm offered? Uh, you know, uh, will I get sued because these don't have the right license? And so on, so forth.
So I think there is that, there is the consideration of how to choose the model. Then there's a consideration of, okay, but these models do not have information, private information, proprietary information, or other information that I would like, uh, these models to have because they are private, uh, proprietary, and so on and so forth. So, how do I infuse these models?
How do I, uh, you know, infuse this data and this source of knowledge into these models? And that also, you know, that are fine tuning techniques. That is a rag, uh, a retrieval, augmented generative AI that is, uh, popular out there.
But with all these, there's significant amount of expertise that is still required. A data science expertise. Maybe you don't need data science expertise, as in how much is needed to do pre-training, but still significant of expertise is needed to pick the right kind of fine tuning techniques.
And then on top of that, you need, uh, data. You know, all of these, um, techniques require data to be in certain formats. Um, you know, and this is, needs to be a very high quantity, uh, sometimes even human generated.
Uh, and then finally, uh, you know, it needs a decent amount of compute. Again, the order of magnitude needed for pre-training is significantly higher. But even for fine tuning, you need a significant amount of compute access to compute.
Um, and, uh, with RAG you also still need the, uh, large language model service, uh, but also you need to have, um, you know, uh, knowledge of things like lang chain or, uh, things like vector dbs and embedding models, uh, to chunk the data. So there's a lot there. So I think the opportunity obviously, uh, here is how do you kind of simplify it?
How do you simplify it and bring it to what I would call the developer, or even better yet, uh, the domain expert, right? You know, if we really want to, cause I mean, there has been talk of democratizing AI for many years now, as you know, Mike. Um, you know, but, uh, the question to ask is how, how do we bring it to everybody, right?
Like when I say e everybody, not just from a consumer point of view, but those people who are, uh, you know, training these models or, uh, you know, modifying the model's behavior or customizing the model's behavior for specific use cases, for aligning those use cases. How do you do it? I think, I think that's the big challenge.
So, in short, it kind of sounds a little bit like the wild, wild west out there. Is there any guidance for developers about how to navigate all this? Or is it just gonna be trial and error for a while?
Yeah, I mean, uh, definitely. I think, uh, the guidance really is, um, you know, obviously have your end objectives, uh, in mind, uh, you know, uh, start with that obviously, and, but work backwards, right? And working backwards really means, uh, that, okay, ultimately, uh, you will need to build a, uh, some kind of a ai intelligent AI based application generating AI application.
But to get to serve some specific purpose, you know, and to do that, you need to figure out, okay, where do I start? How do I start experimenting? How do I do it in a irad fashion?
This, in some ways, is no different than what developers are used to. Uh, but, uh, you know, when it comes specifically to, uh, you know, uh, large language models, um, you know, starting small, uh, starting again, as I mentioned, there are small models there, medium sized models, and then there is large, large models, right? You know, start with this small model, start experimenting with that, and potentially do that on something that you can do on your, you know, powerful I would call a laptop or a workstation.
I mean, I, I really don't think that, uh, you could just use any consumer grade, um, you know, laptop, at least the current technology doesn't allow it, but let's say a, uh, apple, Mac, M two, M three, something like that, or which has the GPU integrator, but there are others from Dell and, uh, you know, HPE, but a powerful, uh, desktop, uh, start with that. Um, you know, Dell also has a series called the, um, copilot infused ones. Uh, maybe they, they could be, uh, useful here, HPS something similar.
Uh, but start with that, and then start with the small models and, and start with some fine tuning techniques. Um, you know, obviously, um, you know, uh, and, and, and, and, and, and, and, and, and kind of feel the pain, so to speak. I mean, I, I don't know, sound it away, but you know, unless you do it, you don't, uh, feel the pain that a weird is fine tuning techniques.
But one thing I would also, uh, say, you know, maybe it's a plug, but, you know, go ahead and take a look at, uh, something that Red Hat announced at, um, uh, red Hat Summit earlier this year in May. Uh, we announced a, a community project called Instruct Lab. Uh, Mike, I don't know if you have a chance to take, take a look at that, but in strt lab really simplifies, um, what I just described.
I mean, you don't really need to know a lot of details about fine tuning and a lot of details about synthetic data generation and so on and so forth. It just distills it into a set of three or four easy commands for you to run. You don't have to be a domain, uh, a data science expert to use it.
So I would definitely encourage folks, and that can be something that can be done on your laptop, uh, you know, uh, and once you're kind of satisfied, I mean, and think about, you know, in the world of software development, there is this notion of inner loop and outer loop. I mean, uh, if, if, if, if you, uh, recall that Mike, I mean, you know, so kind of do the inner loop and development on your laptop, and then once, uh, you are satisfied iteratively with that process, um, you know, and, and instructor lab certainly allows you to do that. Then you can then think about how do I do the outer loop to do a full high fidelity model, which is infused with, uh, not just a small incremental amount of data, but the entire set of data that you want to infuse with.
So I think the inner loop is, you know, um, you know, small models, um, infused with small amounts of data using not that much amount of, uh, compute requirements on a kind of a laptop and a desktop, whereas the out outlook being a full fledged, um, you know, I have lots of data to add, and I want that to be custom, uh, the model to be customized. And by the way, I, I, I want it to be full fidelity model so that I can start, uh, you know, including that in an intelligent app, start serving, um, in a chat bot or something. Uh, so I think, uh, that would be, uh, one thing that I would plug here.
How often are developers gonna embed a model inside the app versus maybe just making an API call to something that gives them an output that they then incorporate inside an app? What's the balance of, uh, you know, the level of skill required? Yeah, I mean, I would say, you know, a great question.
I, I would say, if you kind of step back and think about the microservices approach, uh, you know, to cloud native microservices, uh, approach to modern applications, I would put the model and the apps using it kind of separately. I mean, there are two separate services. Uh, I mean, it is rather, uh, there are ready, there's several readymade, uh, open source technologies out there.
Uh, and they're also available as part of our, uh, redash product offerings, for example, with OpenShift AI and, uh, with red AI that allow you to serve and provide you a, uh, a restful API, uh, including HGTP and GRPC, for example. So you would rather, you would put the model, you'd give that service a, a, a service and restful API service on top of that, and then you would write your application as a separate thing that then talk to it, right? Like, you know, that that would be my typical pattern.
I mean, uh, and not embed the model inside the application itself. I mean, again, there is never say never in software, right? Like with everything.
But I would say that would be the, uh, that would be the more modern approach. That way other services, different services can take, uh, make use of that modular does, does not confide to this one particular, all the, all the goodness of why you would want to do microservices, right? Like, you know, it keeps it modular, it allows it to be independently developed and updated.
Uh, you know, um, you know, it can be reused for different use cases, um, incremental, um, yeah, uh, you know, changes are easier. All those goodness, uh, scaling is, you can independently scale them and so on and so forth. So I would definitely kind of keep them separate.
Aren't we underestimating the costs involved in this? And there's a lot of folks out there who are doing proof of concepts, but then when they get into thinking about deploying in a production environment, they realize, holy cow, these things are expensive. I mean, I think that's a great question.
Uh, I mean, this is where, this goes back to what I was trying to say early, the choosing the model or the model service becomes, uh, really important. And because, uh, no, as you rightly said, people don't think about how expensive these large models are to serve. So, uh, you want to think about in two ways.
One is do what is the minimal size model that I can get away with? I mean, do I really need all the capabilities of a a hundred billion parameter model or a 150 or a 200 billion parameter model? Um, you know, if I'm just doing some simple, uh, you know, uh, summarization tasks, for example, or if I'm just doing some simple, uh, q and a tasks, et cetera, right?
Like a handful assistant, do I need all that power? Or, and, uh, that's one, right? So that's where I was kind of saying, you know, get whatever the smallest model you can get away with.
I think the current thinking is kind of the seven B, eight b, uh, size, uh, 7 billion, 8 billion size is kind of the sweet spot. Uh, we call them the workhorses. Uh, you know, uh, you know, and so I think that's one.
The other is how can a monetize that, right? Like, you know, how can I, you know, running the model constantly means it's the Azure problem of, uh, computer science is how do you increase, increase the utilization of that model, right? You know, how do you monetize it among different applications, not just one so that you know, when it is being served, it's also being used so you can, uh, spread the costs.
The third thing really is, you know, uh, you know, uh, being able to gen use things like, I mean, there is, you know, off late, uh, agent workflow, uh, workflows and, uh, you know, AI agents and assistants have become, uh, popular as in you can take different small models and, you know, that are purposeful and they're fine tuned for your specific different needs, and they are all being served, and then you have a separate application that can take advantage of those different moving parts. I think that's one of the ways to control costs. I think this is known in some ways, it's kind of with the added, uh, understanding that these are neural networks and therefore they do need compute.
Uh, you know, and in some ways neural networks are brute push, uh, you know, uh, and that's just, you know, something that increases the cost, but one once beyond that, you get beyond that. There's ways to control the costs, uh, in the ways at, at store end. How will the workflows come together?
I'm asking the question because most developers are kind of part of something that feels like a DevOps team, and then there's a data science team running around doing ML ops, and are these things gonna converge eventually? And, and developers and data scientists will get joined at the hip? How?
Yeah, I mean, I, I, I think, I think so. I mean, you know, like when you think about, so I think that that's a great question. I would answer that in several ways, right?
One is, so data scientists, right? Like, you know, uh, so, so if you go back before the Gen AI revolution, uh, if you step back, uh, I mean, there was a lot of, uh, need for data science as in, because everything was kind of purpose built. Uh, um, you know, uh, you know, uh, but, you know, post chat GPT and Transformers and Chat GPT, especially with this modern era of Gen EI and LLM, I think one thing that has happened is that that three training of the large language model, uh, can be done by data scientists, and they can be kind of separate, but then once the large language model is available, there are a host of things that you can do with it, uh, that, uh, becomes the domain of, you know, data, uh, of developers and DevOps and all that.
So I would think that, I mean, obviously, you know, having some amount of data science expertise, uh, in your, uh, in your group is always, uh, important. But it doesn't have to be, you know, data scientists as in like, you know, these are data science PhDs who are, who, who can pre-train models. I, I don't think we need that kind of data science expertise in every development team, uh, which is actually a good thing, right?
This is how, uh, AI would scale, uh, because I think if you go back again before the Gen AI revolution just from a couple, three years ago, uh, you know, I think, uh, you know, data science, uh, you would need that kind of data size, uh, and, um, you know, because every model was kind of bespoke. Uh, but I think in this new era, era, it is possible, uh, to, uh, actually, um, have some data science, but, um, be able to get away with, uh, especially with tools. Like, like I just, again, I don't wanna talk about Red Hat and what are you doing with the Instruct lab, but that's the intent of in Instruct lab, is that you really don't need, uh, you know, data science, right?
Like you, all you really need is, okay, I have this bunch of PDF documents, which I want to infuse or, or some other format documents. I want to infuse this, uh, open source model that comes as part of Red Hat's rail AI product. Uh, you know, I want to infuse it.
There's a simple three or four step process to get it done. Uh, it's just basically ACL I, uh, you know, you see I lab generate with generate synthetic data. You see, I lab trained, and that does the training of the Granite model and outcomes, a customized granite model.
And that, again, you don't need to know any data science to do this, right? Like, you could do, everybody could do it. I could do it, you know, we all could do it.
You know? So I think, you know, that's the intent and I think you'll see a proliferation of most such tools. So, uh, yeah, I think that hopefully answers the question, right?
Like, you know, I, I think gen AI has unleashed their possibilities that were previously Yeah. It's just like that television commercial, even a caman can do it, right? Um, the question I would ask you though is, are AI models gonna be infused in every application from here on out?
I mean, is it just gonna be a standard component of every app, or do you think it'll still be a relatively narrow set of use cases? I think there is a, there's a lot of, um, uh, applications, uh, that can benefit. Uh, I mean, from ai, I mean, if you think about AI in general, right?
We kind of talked about gene ai and specifically within gene ai, we talked about large language models, um, here, but if you open that up a little bit, right? Like you have the whole predictive all the way from what is used to be still called, you know, predictive AI to generative ai, right? Like, and predictive AI is everything from, you know, uh, linear regression.
You know, if you think about that way, right? Like, you know, uh, and correlation and you know, much more of a, uh, you know, objective answer versus lot of time model is more about, uh, probabilistic statistical, uh, as well as, uh, you know, um, you know, uh, language based, but also like we have multimodal, um, uh, you know, um, models now, which can, uh, do more than just, um, uh, natural language. They can understand images, they can understand voice, they can understand, um, you know, other such modes.
And so, uh, long story short, what I'm trying to say is that if you open up the aperture to all these models, like traditional, um, you know, uh, ml, uh, machine learning models to all the way to, you know, what are modern foundational models, uh, using, including multi, uh, modal languages, I think, I mean, think about, another way to answer your question is which application doesn't depend on some amount of interaction with natural language of some kind, some amount of interaction with some kind of a image, some kind of interaction with some kind of a voice, some kind of a interaction with some kind of a, uh, you know, data, right? Like, you know, I mean, every application needs this. And so AI applies to all of these.
And therefore I would say, uh, that, you know, AI is pretty much, uh, you could imagine now not all of them will be large language models, but I could, I would say that, uh, AI will be in almost every application. All right, folks, sharing it here. AI is gonna be pervasive, but if you wanna succeed, you might wanna start small.
Hey, thar, thanks for being on the show. Thank you, Mike. Appreciate it.
All right. And back to you guys in the studio.