Andreea Munteanu, Canonical | KubeCon + CloudNativeCon Europe 2023
MLOps is used in various organizations that operate on very sensitive datasets. With Kubernetes and its ecosystem – Kubeflow, strict confinement for K8s using AppArmor profiles, confidential computing and blockchain-based tokenization – you can achieve a very safe and compliant setup.
Transcript
This is texturing TV. Hello and welcome back to kubecon cloudnativecon Europe. We're here with Andrea Montu from canonical and we're talking about Ai and kubernetes and how all these things are going to come together because we have ml Ops devops, but it's not clear to me where these will ultimately converge.
So undramp, I'm glad to be here. I'm AIML product manager, we think canonical having a data science background and I'm quite happy to tell you more about ml Ops. I always Define it shortly as devops from machine learning and I think your confusion is right because envelops in general is quite a new practice AIML in general is new it has been around for quite some time, but it was not so popular before generative Ai and a few other projects that made everyone quite interesting made organizations.
We think there's strategy and their budgets you see a lot of reports saying that if you don't start with AI you might be out of business which can be scary. Right but it's I think a new Revolution that comes in place and but mvelops brings to the table is reliability repeatability importability such that it connects different dots and different Technologies and what you're flow is because you already mentioned it. It's a it's a developed platform.
It's an open project that at the moment. It's incubating and CNC. Is right here.
And in canonical we have our own distribution. We are working closely and it's the glue of a bunch of other machine learning tools. So we are integrated with spark mlflow and if you wonder why because that might be another question right is you cannot I don't think you can do everything within one tool, but you can connect all the dots such that data scientists don't have to struggle anymore with all the drivers compatibility is all the how do you move it from here to there?
How do you run it at scale if you want to if you just experiment how do you monitor it? So you have one platform that has it all for you but also because it's integrated with a bunch of other tools. I feel like a lot of organizations are hiring data scientists who wind up doing a lot of Plumbing.
So can we kind of take away a lot of that lifting and make it a little more seamless for them and actually have them I don't know work on data science the truth is that There there at the moment. There is a huge skill gap on the market with data science and also The data scientist that are hired often they spend way more time on a lot of other things not on writing code and building machine learning models or data science models. They need to spend a lot of time.
I'm at the data site cleaning it to processing it finding where the data sources are because organization. They don't have a huge focus on data until a couple of years ago. And then once they have something that our model that they're happy with they struggle understanding how to run it at scale because there is Need for more compute power.
There is need to monitor. There is Need for serving and if there is not a reliable way to do it, it's very complicated so shortly. Yes MLPs actually makes lives of data scientists machine learning Engineers a lot seamless and a lot more.
Collaborative because at the moment there is quite a big Trend in the industry to have data scientists working in silos. So you have to data scientists in marketing, they do data science for marketing and then you have two others for sales operations and they do it for that but they don't collaborate even if they use the same data sets and there are a lot of things that could be. Done.
Jointly, for example cleaning the data you do it once and then you enjoy it. But now they do it twice and it's a lot of waste of time at the end of the day which often Falls as AI does not bring enough in Africa and on investment, but it's not because of data scientists, but rather because there is a lot of manual work that could be automated at the end of the day. Is there some duplication of effort because the data scientists is creating an artifact and a lot of times they're storing it in something that feels like a feature store what they call a feature store.
And yeah, we have all these git repositories. So does that need to converge? Well, they have different roles.
Right but you capture it correctly because all those at the end of the day machine learning models are deployed using python code. That's the truth. They write python code that runs into jupyter notebooks initially, but if they don't put it in in a place that can be visible for everyone else such as gate repositories.
It's almost impossible to know what you're working on for example, so you're right here and then there are different types of stores that need to be developed feature stores for understanding each other relevant features for different data sets model store as well because it's you don't do much learning ones you develop a model and you never look at it again, you do it constantly and repeatedly and some of the models might have use cases in in various areas of the business because they might need have the same needs if you think of automations for example, they can easily be replicated based on models that have been developed by other teams. Are the models being built primarily in the cloud or we've seen a lot of people doing stuff on premises these days because the data is kind of sensitive what you where are we building these models? I think you're touching on a very interesting topic now because We our keynote at kubecon is really about the qmlops on the public cloud and how you can use highly sensitive data on the public cloud with confidential Computing but going back to your topic on.
Where do we develop AIML models? I think it's everywhere often Enterprises start experimenting on the public Cloud. So they use AWS.
They use Azure, they use Google and but there are topics or projects that I would like to do on Prem or they would like to scale on Prem due to usually pricing constraints and that's it. This is where also mlops is important because how do you do that transition from the public Cloud to the private cloud or how do you do you enable actually hybrid clothes scenarios? Because you might have part of the data on the public Cloud part of it on the private cloud and you don't want to move it because it's very expensive.
So this is where MLPs comes in place ensuring and offering ml hybrid Cloud scenario is a multi-clouds and I use to be doable for data scientists everybody and his brother is talking about large language models generative AI The ones that we see are built by Microsoft and they're massive and they required years to build. But are we going to see smaller large language models that people are going to build for more narrow use cases. And is this going to become commonplace?
I think what should I repeat? And this great is to to bring AI in front of everyone and shook is that hey, it's here. It's ready to be used and it's easy to be used.
It's not this unicorn out there that no one can touch on us. They're extremely Technical and I think this is great because this is what gave a lot of confidence to people to professionals but also to organizations to try it out but to go back to your question, if you if you go a bit behind you'll see a bunch of a startups doing them having used kids having use cases with llm models actually which are much smaller and I think it's a trend that is going to evolve quite quickly from now on. It's They tried to put the open the box that is going to to solve a lot of problems in a bunch of Industries and it takes of course it takes time to develop them, but it is going to happen.
We see a lot of people building AI models on top of kubernetes. What's that? One thing you see people doing that maybe is a mistake or something that you wish they knew before they got started so that they wouldn't run into some type of issue.
I mean, what's your best advice to folks? Well when it comes to machine learning, I think people often get scared. I got scared when I started with with AI it was 10 years ago.
I was doing my thesis and I was very confused on is it safe can I do it but the most important thing that anyone can do is to understand a problem that they want to solve right? And once you have the the problem clear in your mind you want to generate sound that sounds better. I don't know.
It's just an example then, you know what you're trying to solve and the next step is to ensure that you have enough data to train models. There is no machine learning project without data, even if we want to do it, but it's not and these are the two most important things for me have a problem and have data the next step is of course to to try to build it reliably. Don't try to work in your own Silo because at some point is going to be overwhelming.
I think a lot of data scientists struggle with that at the end of the day. They're often Engineers who shifted or statisticians who shifted and in both of these? Areas of work we are used to to work by ourselves.
So try to collaborate and try to find joints. Efforts left but not least always bear in mind that the end audience of your machine learning model is not an engineer. It's very often business people people who don't have any technical background and how do you showcase your work in front of them can be your successful your failure.
If if you're doing something that can be shown to dashboards. Tell us all the story to those dashboards such that they can understand the problem quite quickly. They can understand what's what are the outliers and also they could easily see trans for example, so it's easy for them.
You don't need to spend more than 30 seconds on a dashboard to understand it if you want to dig deeper. Yes, you should have the tools to do it. But if you just want to get a glance, I think that should be easy and I think that's where data scientists.
Still have to learn on how do you tell a story from your work? How do you how do you put it in front of people who benefit out of it? All right, folks if you're scared of AI take comfort in the fact that you're not alone.
Hey, Andrea. Thanks for being on the show. Thank you very much for having me.
All right, folks. We'll be back in a couple of minutes.





