OPEA: Reducing the Barriers to Enterprise AI Adoption at Cloud Native Now 2024
Generative AI is on everyone’s mind since the arrival of large language models. How can enterprises adopt these technologies to increase their productity? Further, deploying solutions requires AI, cloud and operations expertise. We will cover OPEA: Open Platform For Enterprise AI, a newly launched Linux Foundation AI and Data open source project, that brings together recipes for common use cases, leveraging the best in open source technologies spanning embedders, vector databases, large language and vision models, guardrails and more, all securely, that can run in the cloud, on-prem or at the edge. Come and learn how to accelerate your GenAI journey and influence the future of GenAI.
Transcript
Hello everybody. I'm delighted to be here today and talk to you about oia, reducing the barriers to enterprise AI adoption. I'm Ali BGA and an open source cloud architect from Intel.
And I have had the help of several contributors from Intel for this deck and some of them are Arun, Rama Ker. Hi, how Iris? And there of course.
So many, many more. So with that, let's get started. Okay, so a little bit about me.
I'm a cloud native architect focused on ai, confidential compute, security and performance. I have been working in open source for over a decade and my first foray in this space was with OpenStack when I realized our servers were underutilized and it had to be shared for better resource utilization, total cost of ownership reduction, et cetera, et cetera. And I have worked on applications that span autonomous driving, healthcare, remote monitoring and management, uh, telcos along long journey, a long career.
But I first joined Intel as a power performance architect. I have worked on this yon service. What else about me?
I have been a program committee member and a co-chair and things like that for CubeCon open source summits. And I have a PhD in machine learning and over 25 patents and I love to garden. So with that, let's get started.
So what and why about upia? Where are we today and what's next? So that's our agenda for today.
So we really have to be hiding under a stone if we haven't heard about AI today. So Jenna AI is top of mind, but realizing value from it is still a mystery. It's a challenge and it's not yet then in production at many places, but we do want it to, you know, release the potential of all the productivity savings enhancements we can get.
So we wanna try and make that happen. So what are the top challenges for, you know, enterprise AI adoption? What's stopping us?
What's hindering us? What are those roadblocks? So it's not technically easy.
So that's one of the barriers. Technical implementation cost, I mean there's cost associated with training, large language models, vision models, whatever. There's also cost associated with offering them as influence services.
And if you're thinking it costs a lot for just building the model, millions of dollars, uh, you know, billions of pieces of data, even trillions of pieces of data for training. Um, the cost for inference is a recurring daily expense and you know, they can be as much as um, the amount of energy and resources we use for a transcontinental flight. So that's where we are going with cost and that's where it's important to also be sustainable and efficient and optimize our code.
And last but not least, there's talent scarcity. I mean, AI is been around for more than a 50 years, um, but it's not like we have lots of people who are experts in that area. And if AI is hard, it's also equally true that cloud is not easy.
So the people who know cloud don't know ai, the folks who know AI don't know cloud. So there's this dichotomy of uh, knowledge expertise. So these are our barriers today and this is from a Gartner report.
So what will it take to tame this ecosystem? Because we mentioned, you know, technically it's complex, uh, if you're a cloud expert, you know, you have to be able to design your apps, find places to deploy them, deal with things like resiliency, CICD, so on so forth. And a lot of those problems have been solved, especially with managed services.
And of course they're in-house experts who deal with on-prem infrastructure and keeping it on up. We've come a long way thanks to Kubernetes and all the features and bells and whistles in there. And then we also have dealt with things like security.
I mean, you cannot have an enterprise application without some amount of authentication, authorization and so on so forth. But if you're a business, you also have to say, what's my return on investment? Is it worth my time to, you know, leverage generative AI solutions, whether it's for document summarization and why do we even want document summarization?
We're in this information age where there's a clut of information, everybody makes lots and lots of white papers, news articles, there's also GitHub re reports that you need to study. There's product documentation. Can we make that easier to consume?
And that's where document summarization comes in. It's also not easy to procure some specialized hardware like GPUs are hard to get. Just yesterday I was trying to resource some GPUs and it wasn't easy.
All my instance requests were dying, failing, et cetera. So there are a lot of issues in this space and we wanna see how we can ease some of these barriers to adoption. And if all that gets solved, I mean, where do we wanna apply ai?
Just about everywhere. Healthcare, drug discovery, anomaly detection for networks and anomaly detection for even your credit card expenses. We wanted to also support us in research and innovation in just about every space, whether it's fuels, whether it's chemistry, whether it's, you know, economics.
How do we make things better for all of us? So with that, let's take a look at what a typical generative AI pipeline looks. Why do I use the word generative ai?
Because if large language models and large vision models and you know, all the stable diffusion, et cetera, was the talk of the town and it's bringing in all excitement, there's one problem with that. That's just one piece and that's just this piece over here, okay? It's this like little box called LLMA real application needs a whole lot more.
And that's not even talking about some of the things we deal on a day-to-day basis in an any solution. And in your enterprise there's of course all your authorization, authentication, firewalls to detect, you know, denial, a service type of attacks and so on so forth. So all the regular security issues are there.
But talk on top of that, what AI brings to this picture and what is that? If you have a large language model, it's something static, it was trained maybe with millions of dollars cost with billions of, you know, samples like all your Wikipedia, all your news articles and so on so forth. Those were still static.
It, it could be at the level of a fifth grade English speaker, it could be a graduate school student, you know, who has an awesome vocabulary, but it's still static. Asking a large language model, something about what's happening in the war in, you know, Gaza, it's not gonna be able to answer 'cause it was trained maybe more than two years back. Likewise for what's happening in the Ukraine and you know, Russia and that war.
So we need to bring along with these models, proprietary data. That's what makes it relevant for you in your enterprise. That's what makes it relevant for a domain like healthcare or credit card fraud detection or network anomaly detection and you know, security anywhere, everywhere and that sort of stuff.
So we need to bring in proprietary data and that's expensive, that's sensitive, that's might be guarded by regulatory services. So that's where you wanna bring in your regulated data, your high value proprietary data in the house into these LLMs and how do you reference them? How do you bring them online quickly?
That's where vector databases come in. So why am I sharing all this data? Because I'm, you know, sure that not everybody is aware of all these aspects of retrieval augmented generation.
So I'm pointing out each of these little boxes in some detail. So not only do you bring in all your proprietary data, that could be in any format. It could be word documents, PDFs, uh, bug databases, name it, even research papers, you know, like archive.
These have to be ingested, they have to be chunked into little pieces that are digestible. You can't have like, you know, a a thousand page bible just dropped in because it's so much information. How do you bring down the attention into small chunks?
And that's where embedding models come into place. That's also where you take that embedded model, convert it into vectors. So you can kind of do similarity metrics and say this is relevant to my question and I can use that to better answer the question from the user, which comes in from this user query angle.
So you have a vector database, you might retrieve multiple pieces of information and you need to be able to order them. Like sometimes you wanna have things that are more relevant in time. Was it recent?
Is that something more related to today's, you know, whatever it is in the news. Last but not least, you combine all this, pass it into your LLM and have it provide a response for you. But is that response good enough?
Is it redundant? Is it repetitive? Is it using profanity?
And that's where guardrails come into place. So you use guardrails, do some post-processing and finally lo and behold you give a response to your user. So this is your typical rack pipeline, okay?
And then things can be made even more better and interesting by taking user feedback because they could be train the trainer type of stuff saying no, the response was bad, let's improve it. Let's give feedback, let's do some fine tuning of these large language models and so on so forth. You might also want to have some kind of guardrail.
So if you get some uh, bad user requests like you know, how do I kill somebody? How do I get away with murder type of things, you stop them right there and not waste your resources trying to answer them. So with that, after having taken a deeper dive into what a rag pipeline looks, let's see what this OPM vision is first and foremost, what does OPS stand for?
It's open platform for enterprise ai. What's our vision? We wanna reduce this barrier to adoption.
How can we therefore help you? And that's where customizable reference solutions come into place. The first and foremost one is chat Q and a.
You have your database of whatever propriety data. Like let's say I wanna go to Home Depot and order a refrigerator. They have a whole bunch of appliances and they differ in height weight with energy efficiency.
And I might just have a little corner in my kitchen where this fridge has to fit. So the first thing that can reduce my aggravation and you know, speed my selection is giving it those dimensions. So the chat bot will say, please send me your dimensions.
What are you looking for? These are the options. And we go from there.
Say, oh, okay, so you do wanna steal one and blah blah blah and so on and so forth. So there are some definite quick low barrier type of reference implementations we want to offer like document summary, chat q and a, and of course to help our software engineers, you know, code translation, code generation and so on so forth. Further, we wanna be able to make these pipelines, we call them rack pipelines or gen AI solutions using composable building blocks.
So think Lego building blocks and then you can build dinosaurs out of them. Buildings, castles, whatever. So that's your rack pipeline.
And then we wanna also be able to support multiple deployment options. Sometimes people wanna run them on prem, sometimes they wanna run it in a cloud nearby. Sometimes they wanna run it in an edge somewhere for low latency.
So support all these deployment options. And last but not least, like we mentioned with the guardrails, we want responsible ai, we want security, we want performance. 'cause we do wanna have a planet that's sustainable.
And as part of that there will be things like evaluations, frameworks, benchmarks and so on, so forth. So where are we today? We're having some code, we have some technical steering committee for our project.
We have some partners. And let's take a quick look at these. So our code, if you were to do just Google search as a GitHub, opi, and lo and behold you'll land on our main page.
We have, as I mentioned, those reference implementations. So gen AI examples, those building blocks that I mentioned, gen AI components, we'll rename them as components. Uh, our docs are in flux, they're scattered, they are engineering oriented.
And I just had a lot of feedback from our documentation gurus yesterday on how we can improve this. So these will definitely get better. And then we have gen AI infra.
And why is that important? Because no matter what pipeline you build, you eventually have to land it in some cloud on-prem, public heterogeneous, you know, hybrid type of cloud. And these have to be efficient and you know, can we chain things faster?
Can we drop in observability telemetry? Can we load balance better faster? So that's all gonna fit under this junk AI infra bucket.
And this is still evolving the evaluation part, like we will have benchmarks, uh, contributed by our partners for each of those little Lego block components. Then we'll have end-to-end pipeline tests. And then of course there's all the guardrails, uh, responsibility, quality of answers, hallucination and things like that.
And this is still evolving and we hope you'll be part of this journey and contribute there. So with that, let's take another look at what these RAC pipelines are. As I mentioned, they've made of building blocks like the embeds, the L lms, the guardrails.
The typical user can come and reach this through an HTTP or GRPC request. It'll go through a gateway. There'll be all the mutual TLS and TLS and load balancing and firewalls, all the usual goodies that you have.
But we'll also have some AI specific things like what's a gateway for an AI kind of application. There might be something like rate limiting on number of requests, number of tokens you can use over a period of a month and things like that. And maybe even like which LLMs you can access that could be part of your authentication authorization aspects like my Melany and I am not allowed to use maybe Misra or Chacha PT four because it's expensive or whatever type of stuff.
And then of course all the components, they fall into place as a pipeline and are using reuse and agents make it more interesting and more complex, but it makes it also more powerful. So we haven't yet addressed agents, but that's coming down the road. So with that, where are we today in terms of code?
We took a look at that, but where are we in terms of some demos? So you can take a look at them and then try them, see them say, oh, you know, these guys are early, but they have something cool, let's try and get involved. Let's try and use it.
So we have a few demos and they're up on our wiki page. So that's a link for it. It'll be part of the slide deck and you can use it.
Then uh, we have something at Intel called the type developer cloud, uh, that provides you early access to our hardware. We have also tested our solutions on a public cloud offering, in this case Amazon. We've also worked with our partner Red Hat to provide OIA on OpenShift ai.
And last but not least, why not do more on our laptops? They're getting more and more powerful today. You have um, discreet GPUs, you have integrated GPUs, you have so many more cores so we can start doing more and more.
And that's the whole AI PC kind of bucket that you're seeing. And they're coming more and more on the news. It's not just Intel, but we have other competitors who are making offerings.
So a little bit more about the Intel developer cloud. It provides you early access to our hardware. Hardware that hasn't even yet come out into the market.
It's not broadly available, but if you can use it, you can try it. And it also comes with some intel optimized software stack. So it runs soup soup the best it can on Intel hardware because of some software modifications.
It's also a place where we are offering Intel's GPUs. They're very cost efficient and they called gaudi. And at, at this point in time, Gaudi three is, you know, under development close to release.
But Gaudi two is in our Intel developer cloud. And that's an alternative to difficult to obtain our expensive GPUs from our competitors. So you can train over here, you can even just offer it as an inference service.
You can deploy at scale and uh, you know, here share with you like what you might want to use under different circumstances. We have a bunch of offerings you can use just our CPUs for medium type of inference. Smaller models, if you're looking to fine tune models, you'd be suggest using our larger VMs.
Again, it's on our eons. Um, if you wanna do some big time training with lots of data, but still a smaller model, we suggest using our Gaudi to our Mac series and so on so forth. So we give you some menus and some recommendations over here.
Uh, red Hat OpenShift, as I mentioned earlier, they're also working with us on opi, making sure that their platform, the underlying layer is the foundation and then we can drop OP pieces on it. And just yesterday we had a community, Dave, Chris, our Red Hat partner gave some demos over there. So back to looking at that pipeline, another view of it.
Uh, when I mentioned that we are running upia in public clouds and on Amazon, basically we took that whole pipeline. Every little orange box here was just running on our Amazon EKS cluster. As we're going out further and reaching out to more partners and collaborators, you know, we are looking at, you know, the cloud from Oracle and uh, we are gonna have a demo with them.
Demonstrate not only using an open source off the shelf, uh, model, but their, uh, internal deployment of the model, their internal databases. So the options will increase in this modular plug and play multi-vendor kind of solution that's inviting, inclusive and so on so forth. Okay, what do we use in the Kubernetes space?
Uh, we have something called the Gen AI Microservice connector. Typically what we've heard is people use LAMA index, lang chain haystack, and that's made it possible to create these services, connect these services, and lo and behold, you start off with a prototype and you get your hands dirty and it's all working very fast. But a comment we've also heard is that people typically drop these, these crutches or this uh, this framework when they get into production because they want finer control or they don't even want the overhead of the framework.
So with that in mind, we designed our gen AI microservice connector that does nothing really other than pure Kubernetes. It leverages the Kubernetes operator pattern, the custom resource definitions, and then it can help you integrate with other technologies like service measures that provide you authentication, authorization, load balancing, AB testing. You know, all that's more start of options.
And one of the other things that it does apart from the AB testing is also lets you do some kind of, um, balancing of how much, um, you know, traffic goes here and there. So the traffic engineering aspects of it. So we use the gen AI microservices, uh, connector, we deployed it and then after that we just deployed the YAML file using your usual tube couple deploy or apply type of space and it all works.
So it was very thrilling. The I PC that I mentioned, this is just a quick look at what these offerings typically give you. Uh, you have already GPU cards on your machines, like my Zoom call right now is supported by that.
There are also neural processing units on some of these and there's of course your CPUs and typically you have more than one of these. Even your little phones nowadays are multiple CPUs. So AI can come to your PC and it can do things like, um, instant near real time translation of your meeting, uh, you know, from one language to the other.
Of course, we are all seeing transcription these days and things like blurring effects of your video and you know, maybe even adding makeup to you, but that doesn't stop there. It could even be an audio assist device. It could even help you with writing your documents and so on so forth.
So it's really us, it's our imagination where we take this technology and how we increase our productivity. So a little bit about the technical steering committee of opi. OPI is about four months old.
It was launched as a vision of what we would like to do in terms of opening this whole ecosystem of ai. Uh, not locking anybody in a wall garden, whether it is one vendor's GPU or another's or one cloud versus another or one, uh, framework versus another. We wanna bring all the best together and we have a awesome team of, uh, members for our technical steering committee.
We have Logan from LAMA Index. We have Nathan who is a system integrated CDW. We have Kerr who's our AI expert from Intel, my colleague.
We have Armor fidelity and we have Robert from Comcast who also comes from, uh, Terraform expertise. We have Steve from Red Hat who comes in with also security expertise and he is gonna apply it for ai. So he, I like the way he says, let's write a mind map of what it is in AI and what do we need to secure.
It means models it's data. And we have awesome Melissa who ran our community day yesterday to say, Hey, where are we now? What do we wanna see next?
And it was just awesome the way she ran that show yesterday. And we have Justin from DACA expertise, the original people who help popularize containerized solutions. So with that awesome TSC, we have meetings, we have responsibilities, and our responsibilities are to like guide this project, take input from you, the community, um, make sure that we have work groups to address different aspects of it, like an evaluation work group, uh, an end user workload type of work group, security work group, and so on so forth.
So, and all these meetings are open, everybody can come, they're recorded. And, um, help us guide this project further. And our partners today, there's several of them as I mentioned.
You know, all our TSE members also are partners and I'm super happy to have LAMA Index folks From there we have Neo for J who brings in graph databases because that's a better way of pulling out information. We have Jfr, Melissa, uh, VMware is looking at it, and Zills who brings in a vector database of their own and of course Red Hat. Um, we are also approaching others.
Uh, one of the things we want in opi, and as you see we have some here, but where we want to expand is also have representation from different hardware vendors. We also want to have representation from different CSPs and end user community. So we meet our needs much better together.
So what's next? Our near term goals. So there's much more details on the roadmap, but I'm just sharing what our near term goals are.
Uh, over this month we will be bringing in authorization and authentication using TIO service mesh. We'll be polishing our documentation, we'll be running it on different hardware platforms, and we started looking at what we need to support agents, um, more midterm over the next few months. And I'm really saying months here is there'll be more work on optimization, evaluation, observability, and so on, so forth.
Longer term, as we are looking for more partners, we are also gonna want to deal with fine tuning models because the word on the street is you don't really need large language models and all the expense and resources they need, the energy they consume, uh, fine tune models, which get us a further much faster. Um, but they take time to build of course. So how do you get involved?
So we need to refine and expand our project plans. Just scan this code and you can get into links that'll help us please get involved and more links for you. So where's our GitHub?
Where's our wiki and how you can reach out to us. We are also setting up something called AI explore, and we'll be showing you, you know, basically play with these apps, those, you know, six odd generative AI example pipelines that we have set up. So you get a feel for it, what AI can help you with.
And then of course the rest is up to your imagination. And then you can use the building blocks to get there. So please join the project, please look at what we have, please try it.
Please file issues and say, Hey, you have bugs, or your documentation sucks, or whatever you need to, or please include this new model. Or I'd like to, you know, contribute maybe a whole dashboard, a whole studio. So we were recently talking to some folks from, um, Malaysia and they have this awesome dashboard and they're called Embedded LLMs.
So that could be one of those next things we include into the project. And bring your enterprise use cases because understanding your use cases will help us better develop features. Uh, right here at Intel, we have a group that's offering generative AI solutions to improve productivity.
And they brought some awesome requirements to us. And sometimes you can just be an open source engineer who's designing without enough grounding and they're really grounding us. And sometimes it can be low hanging fruit, but it's very important to have those things.
Uh, so please also help us popularize this project. If we wanna really democratize ai, we need the democracy, the people to help make it happen. So it's really you, it's you contributing, it's you using, it's you bringing in your use cases and it's you growing this community.
So with that, thank you.