AI Showdown: The Good, The Glitchy, and The Risky | RSAC Virtual 2024
Nowadays, with bad actors trying to compromise open-source projects “en mass”, it’s important to put aside the excitement about AI just long enough to make the right decisions. Introducing AI into your CI/CD system might sound fun but there are reasons why making the wrong choices will bite you and your organization.
In this session, we’ll look at each of the top AI frameworks and assess them under the cold light of day. We’ll look at their background, how (and on what) they are trained etc. and generally try to assess accuracy and usefulness today and tomorrow. We’ll also look at which ones have good security, which ones follow great coding practices and even which ones pass the emerging open-source standards and legislation being developed for components and AI models by organizations like OWASP, OpenSSF and governments worldwide. In AI, the best framework isn’t just smart, it’s the one that doesn’t outsmart your safety.
Transcript
Hello everybody. Nice to see a packed room for this presentation. It's an important topic, so nice to have everybody here.
AI showdown, the good, the glitch, and the risky. During this couple of slides, I will try to bring a presentation about what are the unknown parts of the AI and how the LMS that stepped in our life, not more than one year ago, changed the way how we should look also in the software supply chain. And without note further ado, let's get started.
I'm Opi Pop. I'm in the technology space for almost 20 years, and together with my fellow Steve, that we had a conversation ongoing to see how the things are happening in the AI space. He liked to play with pictures, so all the glitchy visuals are courtesy of Steve Pool that used to work for Sonotype for until not long ago.
And through this presentation, we look at four parts, a brief history and everything happened that brought us here, a couple of examples and see how the things are moving towards integration in the software space. Then we look at the software supply chain and try to see how generative ai, and not only generative AI is changing the landscape. And last but not least, to see how the Java landscape more particularly embraced it and tried to put things together.
So let's move forward. A couple of years, um, almost two decades actually ago, the famous Mark Anderson from Anderson Horowitz, a company that is doing a lot of, uh, VC in the software space, had a, had an epiphany why software is eating the world. And that was proved to Bero to be right.
Looking at the landscape in the last 15 years, almost four 16 years, different types of software came together and reached 100 million users in different periods of time. In 2008, Spotify was the first to start on this journey, and it ended 11 years to get to 100 million users. After that Dropbox in the middle of the period needed only four years to get to 100 million users.
And lately this period seemed to shrink more and more. TikTok, probably one of those applications that you cannot live with without on your phones, needed only nine months exactly the amount of period to give birth to a new baby. And in 2 20 22, Chad GPT made history needing only two months to get to 100 million users.
That's impressive. But last year, almost to the point, Anderson Horowitz had another epiphany. They just considered that AI will save mankind, AI will save world.
And let's see if they, they were right according to the Green Software Foundation report from 2023, we are quite far from it because currently the software industry requires as much energy as the transportation airlines and, uh, ship transportation all combined. And it's all more than that because we need also a bunch of water to build everything. Probably just generating one of those pictures that I'll show you in a bit needed half a liter of water with each image generator kind of, I'll doubt that I, we will get to the point where AI will, will save humanity in the next couple of years because if we don't have a plan B, we are here to stay.
But disregarding the tendency lately where we ask the AI about everything, I just went to human intelligence to find out what is actually AI intelligence and JA and jra. And if we don't have a planet B to move to, I don't think AI will be in the position to help us save humanity in the next period. But disregarding the tendency in the last period, I just went to one of the codes of dextra regarding the AI intelligence and he states the following, the question of whether a computer can think is no more interesting than the question of whether a submarine can swim.
And I couldn't resist to not ask JGBT to do something with this code. I just prompted it and asked it to transform it into a picture. Because in the end, a picture is worth a thousand words, right?
And in the starting from the left hand side, you can just see how the prompt changed. I asked it to generate a picture in support of, uh, this quote. It transformed into a prompt that helps Dali to transform the image.
And you have to admit it, it has some creativity and it looks quite nice. I just put it on my living room wall to just have some art there. But I didn't just get rid of human behavior.
And I just went to ask the virgin internet, the virgin internet is the part where humans still wrote content, not ai. And what they are saying is that generative AI is the part where machine learning is generating counting mimicking content already written by humans that it used for training purposes. For me, it just sounds like, you know, the copycats that you can see in the, in the windows, those that are waving your, their hand like that.
But yeah, probably a little bit more impressive than that. And I just wanted to see how far can I go with questioning the artificial intelligence. I took a picture from, um, it, as you can recognize it maybe, and I asked the AI whether it can describe it, and it was quite good at it.
He just took it and just said, okay, it's from it. And it's a kid that is just crossing across the moon on its bicycle together with its with his alien friend. But I wanted more than that.
I just ask you to get into more detail and just drop, you know, the long prompts that it usually has with 200 and something words. I didn't put it here because it's more, more or not nonsense. And the the idea is pretty much clear.
But then I asked it to create the picture and that's what I came out with. But let's see how, how far we got here. Initially, it just showed me that it has ethics.
It just prompted me to tell me that it's impossible to generate it because it's from a movie and it, it's protected by copyright, but it just gives strives to be as close as possible to what they wanted to do. And you can see here, this is probably the, the planet earth. If we continue generating pictures like this daily, we'll remain without water and we'll just be on a dry earth looking at the moon.
And let's see, something that is clear, I took a picture of myself, of, of a much younger self and asked it to just transform me in a other persona, a personalities ready for the space odyssey to move from planet A to planet B as we are probably killing this one. And it just transformed it. Uh, and it just transformed it into, um, astronaut as you can see.
But then I wanted to, to adapt a bit the picture I asked it to move a bit to to the right. And this is what I got. I got an extra helmet when I asked to remove the helmet, it just brought me to my real, real self.
I'm not at high, you know, but in the end, it removed the helmet. And when I asked to just slightly, slightly turn me again, it just bounced me and turned it turned me exactly as I asked him to do. And pretty much I got to the, my normal self.
As you can see, I am exactly the, is the one describing the picture. So I just wanted to do more than that. I just said, okay, take the, take me and put me on a spaceship.
I'm ready to move to planet B and that's it. I can wave my friends goodbye. And we can just look at the future, right?
As you can see, the way how the things are happening are quite glitchy. And imagine that this is only a picture and it's self describing how the things are are changing and how much input from the human. It's still needed to get to the point where you are.
Imagine now that you're talking about software or configurations for something on your pipeline or anything else that needs attention, it still needs a lot of tweaking on your side. And at this point probably will be much better suited if you wrote those things yourself. Or maybe there are things that you don't fully understand and you'll be unable to check whether it's possible or not, uh, to be correct or not.
And in the end, the choice of AI will lock you in. And that's the problem. How will you move from one point to another?
How will you change tooling between themself? And let's see, a couple of examples of what's happening in reality, because everybody was striving to be the first one to catch the AI train, not to miss it. So we can see the benefits in how to healthcare, where to help help into finding quicker solutions.
It'll help transporting multiple things from the videos or audio recordings to text. Or you can imagine a smart bot that will just cut the time that we need when discussing on the phone with multiple representatives from different parts of the world. And it'll be a lot fo a lot faster.
It'll be closer to our context. And it can be, it can be done in parallel. It'll be 24 7.
We don't have to be from nine to five to discuss with somebody. Or maybe finally the public sector will have a friendly phase to what uh, we need and to get it to our best interest. But is it actually so, and a couple of the companies in the private space try to do it.
A company from UK tried to bring LMS in the space where they discuss in terms of support for parcel parcel delivery. It got really angry at one of the customers and it started cursing it. Air Canada tried to do the same and the bot convinced the human passenger that he can, he can cancel his ticket and nothing will happen.
Surprisingly enough, Canada Air Canada believed that they can just blame the AI and then they, the AI will go to jail. Unfortunately, they lost a huge amount of a pile of money because of that. So they had to pay the fine for their bot.
At this point, we didn't get to argue. So I find it only normal. And if you just want to turn yourself in a hacker, just reach to your phone in your pocket.
And probably if you have Amazon, you have the right tool in your pocket because the LLM that they have behind the scenes, it was able to generate Python code for malware transformations only based on plain questions. And if the guardrails around the L LMS are better, oard will just go around it even if you want to believe it or not. So if you look at it according to experts in the field, nine out of 10 companies are open to risk of cyber attacks, you to hackers that are using ai.
And there are a lot of things, just think about the black mamba that used, um, charge G PT to generate keying attacks that are changing with every, every second passing. Wow, that's exaggerated, but it just transforms to evade modern EDR security. And as we have elections pretty much all around the globe, deep fakes and deep fake fishing, it's something that grows exponentially by day passing.
So be careful. Who do you believe that you heard on the phone or you, who do you believe you saw in a, in a video? Just make sure that everything is as you believe it.
But let's look how the AI is changing the developer space. We speak a lot about co-generation. We speak about, um, pictures, we speak about summarization, but actually the whole ecosystem, the whole SDS, the whole SDLC chain is getting changed for the part where we can improve our architectural design.
The part where we just make better, better roadmap together with the AI to the point where we test it, deploy it, build it, understand, monitor or operate it. Just think about situations where you have those huge piles of data stored into logs where it just can have a better understanding of what's going to happen. So AI can be distributed, everything as you can see in the diagram, but what's actually in the middle of a couple of names of tools that are being around the market.
The GitHub copilot is promising to generate code for you to augment the way how you build it. The star code is a model that try to do that as well. And there are a couple of other tools like Code Rabbit and the tools from code that are helping you generate unit tests that are generating you, uh, that are helping you generate code reviews and pretty much everything in between.
And there are a lot of them and each of them has their own minor glitch to look into. But how does the landscape really look like? And here let's see how the landscape changed in terms of tooling for the developer.
And the picture is taken from Holly Commons who zoomed in in the way how the software development tooling changed in the last 80 PL plus years. We started with, uh, the SAMR and then we moved to libraries frameworks and so on and so forth. And now we are scared that the AI coding system will take our jobs.
We are not at here. Um, if you think about it, for me it looks like more like a modern id. Of course it'll just help out in multiple drawers, but IDs didn't scare us, right?
AI shouldn't either. Let's try to put it to the best we can to our help and try to be build on top of it. And the same people from, uh, Anderson Horowitz just came back with, um, emerging LLM stack architecture showing us how the things will change.
And you can see that it's a huge difference between what we believe. It's not only the model here in the, in a corner, it's a huge pile of things that are all interacting with things coming from the user with output, going to the, to the same user. But in the hand there are also ways of interacting between the, the moving parts of the, this very complex supply chain.
And your supply chain will definitely become huge. The attack surface will be huge as well. Just imagine that you now have a virtual assistant that knows pretty much everything about yourself.
Probably you'll just ask it to buy flowers to your wife or coughings to your partner and now knows the color, knows the name, your credit card and everything else in between. And because we don't have that many carlay, it'll be probably a mix between cybersecurity practices today and social engineering skills to make sure that you're not being convinced to do something and to allow it to move to, to the enterprise like in its own garden. And one of the, the edges and probably the sharpest one is AI poisoning.
How can we protect ourself from models that we don't know that they're altered? And me, I'm coming from behind the iron curtain. I'm coming from a former communist country and probably I would be the bad guy in the movie with, um, spies.
I'll be the sleep arrangement that, uh, lives across the street from you and you'll be the one that, uh, just believes that I am the the best neighbor you you ever had until the point where something triggers a different way for me to be, to behave. And that's something that might happen to models as well. We have models that are behaving as normal in staging environment during the development and everything else in between.
But at a given moment of time when it's in production, it'll just change sites and it'll become a very ill willed friend, a very ill willed robot that will just try to do everything that you might be surprised to do it, but it might happen. And that's the problem because currently this is how the supply chain looks like. Or it used to look before we had ai, you develop the code, you selected the dependencies, you tested it, integrated it, deployed it, you through security and compliance and then maintenance.
And in the end, sunset, those, those components. But now the things got more complicated if they weren't already. Now you have to control the ai, understand its provenance, how it was trained, evaluate and continuously retrain it.
And last but not least, legal governance. And that's pretty much different all over the globe. But hey, is it, are we there yet?
Yes, we are. According to the state of open source report, you can see that there are already, and this is a picture taken from uh, last year state of the supply chain, uh, from Sonotype, it's already being visible that there are dependencies in the software that we use. The blue lines are the transformers, the long chain is growing as well.
And on the other side is OpenAI. So people start using it. And now we have to just think about it.
DevSecOps is a, is a thing. Just think about how devs, ssec AI ops would look like. Yes, that's complicated, right?
And you really believe that that's complicated. Think again, dec devs sec AI ops is the new next thing. And that's important because now you can just have a lot, a lot of things happening in front of your door and not even knowing.
So let's see a bit how we can just dig into this black box and try to make the best out of it. And obviously the best way of choosing a tool is finger in the air to see where the wind is blowing. But is it that the best way of doing it?
And let's see, how many models do we actually have there? This is a picture of HuggingFace, the repository for open source models. In March this year, it had half a million models.
Now 200,000 more were added to the space almost. So how do you actually make sure that you don't get into remote code execution problems, vulnerabilities that were injected, for instance in the Lama CCP Python library that is used for inferring LAMA models. Unfortunately I don't have a clear response.
There is no recipe for picking it. There are ways of looking at it. For instance, the Mitric Corporation put together the Atlas matrix that is providing different mechanisms to which you have to understand the problems that you're looking into.
And there other way of looking at it into is the, the LLM top 10 vulnerabilities from a osp. There are no clear recipes to do it, so you have to do it yourself or you can use frameworks that are already existing. For instance, sife, the safe AI framework from Google that provides a guiding into the way how you can improve the way how we use AI into your company.
And if you believe that, uh, nobody's trying to help think again in March, 2024. So this year at CubeCon Europe, CNCF promised to help to bring AI closer to you to make, um, MLOps something that is a lot easier and how better to do it than providing a cloud native reference architecture. In the lower part you see the harder and moving towards the closer part to the software development where you have your workload and here you have all the types of personas that you can imagine.
They just believe that they can join forces. And they did that together with, um, a couple of the partisans of open source software, Google, Mistral and alama. And they need to do, to do it quickly because as you see, the software supply chain is getting more and more hits, um, in terms of, uh, vulnerabilities.
You can see it from the, the same state of supply chain development that it's increasing. A quarter of a million malicious packages are discovered yearly, or at least that was the number for 2023. The unfortunate Pro, the unfortunate case is that it's the double amount that it happened in the last previous three years.
And how can you actually protect against it? I personally tried to use open SSF scorecards because they're looking at malicious maintainers the way how build systems are working, source code compromises, but also malicious packages. And in order to prove that, I looked at the for four libraries that can be used in the Java space to infer AI models.
And unfortunately because they're moving at such a fast space, all of them are in the lower part in terms of scoring long chain for four J, which is the most, uh, famous one. And I like it because it's built on top of, um, vanilla Java. So it's easy to integrate with other stuff.
It has advanced rag practices and it's, um, native counterpart. The lung for the lung chain is quite fast developing. 4 and this actually improved in the last couple of months.
It was under six not long ago. 7 and looking at spring ai, the part of the spring ecosystem is just half the scale. And that's unfortunate because spring is a huge ecosystem that has a lot of users.
And last but not least, J Lama, which is taking advantage of the vector API and um, as a fallback on the Native Panama, it has a quite decent score giving them that it's sort maintained by a very small group of people and doesn't have any kind of company behind it. So be careful of what you are, you're trying. And before bringing anything in your project, maybe you would like to look into the way how the models behave by using Alama port mam, J Lama or even DevOps gi because they will allow you to see how the pro the model behaves in different ecosystem and and responds to different types of questions.
I particularly like the the DevOps Genie one because it just tries with multiple models and then it allows you to compare the way how the models behaves in different types of questions by providing any, some graphs, but you should try them yourself. And what's the best way of looking at it? Is there a right, is there a wrong?
Well, actually it depends because each project is different. And even if, um, there is this famous code going around the internet that stay that states that it shouldn't learn any new programming language, we should take this with a huge grain of salt because actually this is quite complicated. And popularity usually doesn't mean safety on the internet.
Just looking at some of the, the tools they had north of 100,000 stars, but actually the scoring was under five. And of course we need to keep the humans in the loop. And more than that, just make sure that you test the, the outcome that you'd expect.
Make sure that you have the gray, make sure that you have the guardrails in place for everything that you need. And also make sure that you get the proper model for your use case. And of course, popularity doesn't mean safety.
For instance, there are tools out there that have north of 100,000 stars on GitHub, but actually the score is under five. And even if we are tempted to just automate everything, we need to keep the humans in the loop now tomorrow and probably next year as well. And in order to make sure that you guard yourself from everything that can happen, test the outcomes, continuously test them because there are so many problems that might occur.
Make sure that everything is on the side, on the safe side. And with that, thank you.