Dear CIO: Navigating the Shadows – GenAI’s Promise, Peril, and the Path Forward with John Willis at AIE 2024
Transcript
Hello. There it is, uh, John Willis. Uh, the presentation today is called Dear I-O-C-I-O.
And, uh, and most of you know me, or if you don't know me, I could primarily go by Bot Loop or John Willis DevOps on LinkedIn if you're looking for me. Anyway. So, uh, one of the things that, um, has prompted this, which was, I think we went back 10 years ago, maybe a little less, you know, uh, it was probably maybe seven years ago, if I'm being a little more accurate.
We, uh, we were concerned about, like, DevOps had been sort of in process. We were doing a lot of great things, and then we sort of felt like we didn't, like explain what we were doing very well to auditors in large banks. And so we wrote a paper called Dear Auditor, and it was, it was very much like an apology letter and like, man, I wish we would've thought about this, but now we're gonna think about this.
And, and so we were a lot of us in DevOps and, you know, for those of you, so we ran out the first dev DevOps, generative AI hackathon, a tech strong Van Book Raton last year. It was great. We brought in a lot of DevOps thinkers and, um, and we started thinking about like, not only what does it mean for DevOps, what is it, what does it mean?
Like how can we create opportunity in some of this new generative AI stuff, you know, but like what, you know, I think some of the conversations like, what is it gonna mean for life of people who support? And then that conversation has continued to the point of like, we are realizing that the technology is very brittle. It's expansive, and it, like, I I kind of joke that there's a technical debt tsunami coming down to the people who are basically in charge of protecting, we're protectors, you know, DevOps, dev SecOps, infrastructure and operations, SRE.
And so we started recently thinking about, uh, part of Gene Kim's organization creating this paper called sort of Dear CIO. And then what this one is, is, dear CO beware that if you thought shadow it was bad, wait till you see what shadow AI might look like. So that's basically what this presentation is about.
And so I wanted to walk you through, um, some things that to sort of be the, get to the, the technical debt tsunami, talk about the threats and maybe some suggestions of how, you know, not answers, but like the questions that you should be thinking about. So first we'll go through an introduction. If you don't know who I am, I'll try to go through that reasonably quick, but, but I'll keep it germane to like why I am telling this story now.
Um, and then we'll go, I think there's a, an opportunity for us to get level set on, on what, what are the sort of the componentry of these things that we're gonna have to support. So we just say chat, GPT, or we talk about lang chain or, uh, rags. So I, I, I, I little section on that.
And then I wanna talk about like, the potential technical debt. What does this mean? What have we learned in the past?
What should we be thinking about? And then threats and, and scope. com right?
Know that I've written about, I dunno somewhere in there, but I think I, I'm working on my 13th book right now, which actually is the history of ai, which is gonna be a blast, hopefully that's stunning. In the year, I've written a number of books over the years, probably most notable is the DevOps handbook was co-authored. We also created some papers about sort of devs stack s or specifically what we call DevOps, automated governance.
And then I've had a passion on Dr. Deming. So I've recently, uh, in early this year, the, the, the paperback version of that book came out.
And I've worked for a number of companies. I've had like 10 plus startups and, and I do work very closely with, uh, the great people at Techstrong ai, and, uh, and on a couple of vendors. Now I've really, oh, for almost two years now, I've been very focused on sort of generat AI as it applies to DevOps.
What are the solutions that we can solve in our domain? And more recently, what are the problems spaces that we may encounter? IE shadow ai?
And so I've actually gone out and solicited myself to a couple of, like, I pick my clients, my clients don't pick me. Um, so I tried to find a nice compliment of clients, um, starting with MongoDB. They have a Vector database, we'll talk about Atlas Vector database.
So I've been really focused on that. I feel really comfortable with their product suite, their management and their leadership in the enterprise on the right hand side, uh, an open source project called Phoenix Arise. This is observability.
And we'll talk a little bit about the difference between observability and Genive AI versus sort of the way we think about, you know, cloud native computing. You know, so it's a different, it's different. It doesn't measure latency and performance at latency, correctness and hallucinations.
And so we'll talk more about that. Uh, uh, also a hedgehog, which is interesting where I didn't see myself working with GPUs, but as I learned more about training models, um, I found that, uh, this hedgehog is, you know, very much, um, what they call, um, composable infrastructure. And it really is the way, if we're going to have to run on-prem, we're gonna have to have network configurations, configuring GPUs and all the sort of software defined constructs that we have to do for networking.
Very complex in this world. And, um, the, the, uh, Mike Dekin, who was the creative of a CI for Cisco, he's a CTO over here. And like a joke, when Mike Dekin asks you to help him out, you say, yes, he's a brilliant man.
And then last but not least, is another tool that really covers the operational side of this stuff. The people who wrote, um, rancher, uh, created a new company called Acorn, and they focus in on something called GPT Script. So this has been a great mix for me to learn, produce more material, and help you to the extent that you'll enjoy this presentation, you can thank them for helping me learn from a lot of their technologies and their brilliant people.
So look, why me? You know, like, and I, I and I, I've said this, I, this is a joke that I used to bring out in the early days of, uh, sort of cloud. And I said that we're sort of like operations people or people who worry about infrastructure operations or protectors, if you will.
Uh, we're like cicadas. We literally, um, the world gets really messed up. We sort of either wake up or maybe now they listen to us and we have to clean everything up.
And I, I realized that I've actually been doing this for five decades. I actually started in an, I mean, I actually started in the eighties with mainframe stuff, but my first sort of transitional next gen big problem that I tried to get involved is where, you know, back in the day you had data centers with mostly mainframes, but some distributing computing. And you just had this consoles and you literally, as many humans as you could put, the consoles always were like a four to five to one, 10 to one ratio ratio.
And so a lot of us created this idea of like, how can we automate the things you see on the console? It's called automated operations. And the point I wanna make here is in every one of these next gens, you kind of, there's sort of like, if I'd love to create a bar chart where there's the promise of the next gen delivery, there's the actual fulfillment, which is somewhere between 50 and 60% best case.
And then there's a level of technical debt that like gets left behind. And as you get to the next generation, which is the two thousands, which probably most of you're familiar with, it'll just start a configuration management. I used to call it configuration ation first generation, but then you had the, the CF engine puppet chef, then to Ansible really sort of infrastructures code, right?
Which, which sort of took a lot of the conversation, really created commerce. Like maybe the cloud couldn't have been as successful without it. Again, think about that stack chart of the promise, the delivery, the technical debt, and meanwhile, the prior generation, the technical debt is compounding.
Then we get 2010, the, the, the, uh, that whale character that I'll just call it proposal infrastructure, but containers, platforms, Kubernetes, all those things right here, again, the promise, the delivery. And at each level, you're having compounded technical debts. We never actually complete the debt.
The gap between the promise and the, uh, delivery is always significant enough. And then we're always compounding technical debt. Now we're in a another one where I can guarantee you, as great as these things are, and I'm big fan of generative ai, I mean, I've, like, it has changed my life.
I've redirected my career again in my fifth decade of doing these things. Um, but the point is, I can guarantee you there's gonna be a promise, a delivery, and a compounded technical debt, including all the past generations compounding, right? So I wanted to stick that in the back of your head as we go through this, you know, dear, CIO, you, you know, this is a repeatable pattern.
And you know, and I think you, this, I I, I was doing some research and in fact, the, the, uh, Einstein quote is actually misquoted. It actually is from a, um, environmentalist, um, feminist called r May Brown is insanity, is doing the same thing over and over again, expecting different results. And it actually goes back to like the 18th century.
But, but she's the sort of the canonical, uh, citation quote. Alright? So, so that's the intro.
And then I want to talk about like, okay, so now there's this, all this stuff, and most of us, you know, I've had the advantage. Maybe I've got a year and a half Ted start on most of you, maybe all you are experts, but like, if you're, most people I'm meeting at DevOps days, and the, the people I do a lot of work with we're, I've got like a year, maybe a little more year head start on you on the terminology, what these things look like, what they're gonna do, what their impact is. And I thought a lot about in the early days of being an operations or a protector during, you know, infrastructure operations sort of pre, before we coined the word, um, a DevOps, but like, we were sort of doing those things, right?
What was the, something that helped us ground it was the lamp stack, right? It just, you know, like, okay, there was all this stuff going on. It was Apache, and then people were doing PHP and some people were still doing Pearl and, and, and, and there was MySQL.
And it was all like, what was going on here, right? Like, and like it was all getting thrown at us really fast. And what do we have to do?
We had to support this stuff. And to me, I think what's fundamentally, um, worked really well is the grounding in at least a stack, an acronym that meant a stack that allows us to sort of ask questions about, okay, what are you using for relational database? What are you using for, you know, sort of the web orchestration, right?
And I've been thinking a lot about this, and I didn't, like, I invented this taxonomy, and not that I want it to be a plaque that John Willis created the LMA stack, but I, I've been noticing that vendors have been usurping some form of a stack. Like they'll put their letters, uh, as a, and I thought, well, wait a minute, what if we step back and said, what are the things that seem to be showing up in most conversations around general AI that are actually gonna cost support opportunities, right? And so there is what we call language model orchestration, um, in either a large language model, LMS or small language models.
And we'll, we'll go through these, but think about the Lang Change LAMA index observability I talked about, like this is how do we sort of monitor or evaluate the answers and the questions we'll go into that. You've probably heard about rag more specifically vector databases, retrieval, augmentation generation. And then there's a whole set of principles around how are we gonna maintain the models that we use, the foundational models, the embeddings that we use from other models, and how are we gonna create all of the sort of things that make work internally for our enterprise.
And then sort of an emerging conversation is, you know, sort of, you'll hear it rephrased as an army of LLMs or a mixture of experts, or, uh, another friend of mine calls an army of bots that solve lots of problems like this autonomous agent, um, or, um, agent personalization is really interesting. So I think this is a good way to at least I can have a conversation with people. So we're at least grounded in like, what's your, um, what's your language model orchestration choice?
What rag are you using? And so the fear then back to D-F-C-I-O is what I don't want to see happen is us to ignore these conversations. And in a year and a half from now, you are now supporting 30 vector databases from 20 different vendors.
You're supporting like five different orchestration engines from two or three vendors. You're supporting, like what is antithetical to everything we learned on how to manage the constraints to create flow, right? Same thing with model providers.
And so if we walk through quickly some of the orchestration tools, I mean, you probably have heard of Lang Chain Llama Index. Um, the reason I sort of highlight DSPY is I think what we're finding is the Lang Chains llama index are great getting started tools. And you know, I I, I'm okay with eating my words in this world because no, anybody who tells you they know exactly what's going on is full baloney because it's changing so fast.
But I think what we're finding is these tools do a lot of, under the covers work for you. So you ask a question, it converts your question into an embedding very technical, that embedding has to match the foundational model or the rag the vector database. And there's, there's even sort of some re-ranking and re questioning depending on the, the parameters that you chose.
Um, there's a lot of stuff that's going on that's being done for you. And I think if the question is that you're writing some chat bot for your corporate headquarters to describe or given it a recommendation of where to go to lunch, yeah, Lang chain, pretty easy, straightforward, use of modern VE database. But if the answers are how to customize, uh, from a buyer perspective, a, a vehicle, or how to answer a question for somebody climbs up on telephone poles that fit very complex, um, things that are going on with electrical wires and, and all the complexities of the dangers, the answers could actually kill somebody or cost a lot of money.
And there, we hear, heard the Canada story, right? Imagine you're buying a car and you have a dialogue, and like when you get your car, it's like, yeah, no, no, I said it was supposed to be purple. This is blue, right?
Like, you know, again, I'm exaggerating. But the point is, we're finding that the more programmable or immutable, that orchestration you take away all those abstractions and things that make it easier is where people are winding up for high consequence answers. Uh, this is an architecture that we've actually came and develop developed at the Boca Raton, the first DevOps stays.
There's a great video ad on that, um, of what we did, um, in Boca Rat at Techstrong, and this is one of the, um, from Joseph Enox, but ev uh, enterprise Vision technology. But we, we really worked around like, what is this gonna look like? And, you know, if you have more questions, I'd certainly reach out to us about this.
And I talked about observability, and the thing I'll tell you about observability, it is not really the performance, the latency, the CPU, you know, which is still mandatory. Like, so if somebody says, we do LLM perform like a honey chrome and not to pick on them or Dynatrace, we do LLM observability, what they're telling you is they're monitoring to the components that possibly are a mon manageable either on-prem or that they can see some insight to. But what these tools are actually doing is telling you the percentage of correctness of the answer or the percentage of hallucinations.
It's a whole different area. And it's, I I tend to work with a rise. I think they're open source tool, but Lang Chain has a built in Lang Smith.
And, uh, and so what are we doing here? We're managing like hallucination management. We're, we're setting sort of thresholds or, um, or, um, setting the bar at like, I want 98, 90 3% correctness in my answers.
I, and I'm actually using, like you think about TDD, instead of creating sort of mocks, I'm actually creating questions. And the questions are creating the output of the relevance, the bias, the toxicity, and the drift. And that's what we're monitoring.
And so this is a tool, this is one of those, so tools where it actually is giving me, um, explanations of hallucinations, right? Really, really fascinating stuff. Also tell me latency as well, but also tell me percentage of correctness and like very powerful tools.
Um, and then you have the Vector database. And I think vector database right now are the sort of lifeblood of the conversation for enterprises, right? Like, uh, I mean, the truth of the matter is, most people in large enterprises who are building generative ai, their own internal chat bots or internal copilots are not training their own models.
Most people are putting their data and fine tuning their data into a vector database and then using the evaluation software and a process of data, data ops engineering to get the right answers. And so, again, I I, you know, I could go on and on away. I think MongoDB is well fitted for the enterprise.
You know, you might find some of these little tools look great for green fields and, but like, this is a company that is, that actually supports Vector A is in the document object format and has been doing this stuff at scale for many years. So, but these are some of the participants, and there are more, there's a hundred now. And so when I'm thinking about dear CIO concerns, I'm, I'm worrying about like, uh, if I'm doing this, like, yeah, it sounds great, and everybody's gonna want to build a rag.
Everybody wants to take their PDFs and get some glorious answers, which you do. I mean, like, you can take a PDF of a bunch of information, throw it into Claude, and it'll give you incredible answers. In fact, uh, one of the things I did with it Revolution on my Deming book is we put in my book and then created a study guide just from loading a book without it doing any data engineering anything.
And the, the study guide was so amazing that I didn't know they had done that for me. I asked who wrote the study guide because it was so good, and it turns out it was literally no engineering and just outta the box. But again, if the answers are life or death or can destroy your brand, then you don't wanna rely on just enough.
Uh, but oversharing, um, making sure that the right answers are going the right people, right? This is gonna be, these are hard problems. You know, we've had these problems in, in just, you know, relational databases and managing, you know, sort of our back for all things like Oracle over the years.
It's, it's gonna be even more complicated now. Uh, simple data, discovery leaks, log leaks, um, and in there, you know, the adversaries, you know, the either intentional or unintentional or intentional, um, you know, the sort of the, the internal adversaries, like there's just a lot of room to do some like terrible things by, you know, just poisoning the data. You know, like almost like, uh, you know, creating Easter eggs in data, right?
That it could give insight to something once you leave the car. I mean, it's scary stuff, right? Um, I mean, I'm not saying don't do this stuff.
I'm just saying, dear, CEO beware. And then I just wanted to conflate a little bit like the, the models and the providers are not necessarily linked, but I think there is sort of like some synergy, like OpenAI, probably Microsoft Azure, Azure Open as probably for the enterprise, the better place to be Google and Gemini. But even like CLO and clo, like you can run that anywhere.
Amazon's gotten really good at it with Bedrock. And then of course LA of Meta. And then there's the small language models, um, which are really interesting.
These are really good at the edge, very specific purpose, high fidelity, low cost, um, you know, so again, a lot of trade-offs in how you build this stuff. But minstrel is a darling child. Microsoft FI three is like really promising and Google Gemma.
And so, all right, next section I wanna talk about, well, you know, sort of the technical debt that supposedly, uh, that is a Chet GBT generated, uh, tsunami of a computer farm, right? So there you go. That's supposedly a, a tsunami coming in.
And, and so I, I thought it'd start with this, uh, interesting thing that just came, I just read recently. Um, it's a, um, LinkedIn work index trend and report. And, and what's really interesting here is, I won't bore you the whole thing, but 75% of knowledge workers around the world are generative using generative ai.
Fine, okay? That's what I expect. 78% are doing the BYO ai, right?
Does this sound a little different than what we went through with, uh, cloud and shadow it? But here's the thing that I, you need to sort of comprehend here, dear CIO or you, you should be shouting at your dear, your CEO about this, is that if you think about shadow it, if I had a hundred thousand person organization, I mean Amazon, you know, AWS, um, you know, EC2 S3, the p the, the, the target audience who are really gonna use that, and I'm being generous, was three to 5,000 out of a hundred thousand people. And look at a mess it created with Shadow it.
If you believe this, and I do that, if you've got a hundred thousand person person and 78% of 75% people are bringing their own ai, this is gonna be of epic proportion technical debt comparative, like shadow AI will be at a, a proportionally, I don't know how many orders of magnitude, but more than two or three of what we saw with, or at least two or three of what we saw with Shadow it, right? So, um, so that's sort of my first data point. And I, you know, I think there's the cautionary tale, like don't be these guys.
They're Canada group, you know, these people. Um, and, and, and the sort of, the, the interesting and behind the scenes thing was that, um, they probably weren't even using GPT uh, two, which is like pre to 2019, right? Because if you look back when the last, when the actual incident happened, it was probably some intelligent bot that they were created.
But that wasn't the point that Air Canada's brand was damaged because they allowed a bot to answer a question that it shouldn't have answered. And, and that would've been okay, except the New York Times article about it was the poor guy who didn't have a whole lot of money was promised $1,200 in a refund to go to his mom's funeral. And the mean old Air Canada wouldn't pay him because they said it was, uh, the chatbot gave the answer and they didn't, right?
Like, don't let this happen to you. Um, and so we wrote a letter and we're gonna sort of like, we got a paper coming out later, probably in the summer, and this'll be, it'll be sort refined and we're doing a lot of research, but like, dear, CIO, these are some things, and I'll have the slides available for you that you can get, but like, be aware. 'cause the other problem that I think I'm seeing, I'm not think I see it, is like the CIO, the CEO for certain, and the CIO to a certain extent is starting to believe the myth of all this is gonna reduce head count across the board.
And in some cases, you know, if you have like your a hundred thousand person organization, you got 500 people, you know, correcting colons and commas and semicolons, then they're probably gonna job is gonna change. Um, if you had 20,000 Java developers at a bank, you're probably only gonna have a couple of thousand. But what isn't going away is protecting the brand support and all that, that's gonna increase.
So there should be, there might be this blind spot of thinking that it's a pure reduction play across the board. And my belief it is not. And know Chris Brown, OEC two, I worked with a chef, said, you know, you squeeze the blue balloon balloon and the oxygen just goes somewhere else.
And that's the point of this CIO letter. And what's interesting is Google learned about these kind of problems that we're gonna see now, 10 years ago, they wrote a paper in, in, in 2013 called Hidden Technical Debt Machine Learning. And so much of it's a very technical paper, but it's readable of what they learned 10 years ago is exactly as you go through it.
And I won't go into gory detail and the paper we're gonna produce this summer, we will spend a lot more time, but it's just amazing. You know, they're always 10 years ahead of us in terms of scale problems, right? And so they document incredibly well what they experienced on the things that like, this is going to happen to us, you know?
So one of the principles, like change anything, changes everything. Like this is a non-deterministic world, right? Where everything is based on probability.
There isn't no, like, this is the CI/CD run and it should run this way. And here's what the door metrics are like, you know, like some of that. But like, we live in a probabilistic world.
Generative AI is a probabilistic world. Um, um, you know, correction cascades, this is another interesting, like, these have these, you know, when you build these models, you know, without a lot of like technical debt or un or cleansing, what it learns, it learns forever. And, you know, and some of these normal networks are gonna be hard to undeclared consumers.
This comes, we see in this all the time, you know, I created a model, like there's a famous story recently where, um, a large email provider, cloud-based email provider created, um, a, a a copilot version of it within the first day. Um, internal adversaries, or just internal seekers, right? We can call 'em, you know, corrupt or not corrupt, figured out it was a Wall Street firm that did this.
They were able to figure out what all the bonus bonuses were for all the, the, uh, executives, right? So the there, you know, again, it gets harder and harder for these things to create RA and who gets to see what, you know, the idea is we're putting all this data together. Anyway, I'm going on.
There's data dependency debt. There's, um, I talked about evaluations. Um, all right, so let's get into the threats.
Again, I'm going through these quickly, you know, 'cause I, I want to fit this within a, a timeframe that sort of works within the agenda today. And I do, do, um, you know, so this, um, uh, two shameless shoutouts my book on Deming, certainly please buy it. Or, um, I do do a workshop and I'm doing, uh, the half day and one day workshops for people on, you know, where we actually write code, we learn more about the, the alarm stack from a code delivery.
We solve problems, and we go through a lot of these sort of organizational design and technical debt. But, um, so, but then the next section is general, uh, gen AI threats. And so that, um, I I will say this, I'll come out and say this.
I think what everything that I've read from NIST on AI is just nonsense. Just nonsense. I, I would dare anybody to challenge me.
I'd love to have that debate. Tell me I'm wrong, or convince me I'm wrong. But I mean, the, the, the way they're describing generally, it just sounds like all the threats you could have made about Google or general search.
But I will say, oh, I is doing incredible job, in my opinion, of really rolling up their sleeves and trying to address, you know, nobody has the right answers right now. And I like, if I'm implying that I have all the answers here, like, please do not accept that or go down that path. I am trying to learn this journey.
I might just be a year ahead of you. Maybe I'm a year behind you. I don't know.
But the one thing I liked about the recent OAS paper is on ai. And is, uh, is that, like, it talks about the, the first off, let's be clear, there are threats. You know, I've heard a large entertainment, um, executive tell me company tell me this is do or die.
Like this isn't something we can say no to. Um, so, so what are the threats of not using generally competitive disadvantage market perception? Like, why?
Like, again, I think in the not too distant future, I'm going to expect that I don't have to pull out that thick book in my glove compartment on why this blinking yellow light is happening on my dashboard, right? I'm gonna want to be able to ask either my phone, probably my phone of what is that yellow light light, you know, that has this weird thing on it, or even take a picture of it, right? Um, like, so I'm, there's gonna be a perception of like, why are you not doing this?
'cause all your competitors are innovation technician operational inefficiencies. Like I said, I don't know, there's a world where we should have 20,000 Java developers. I'm not saying we shouldn't, I'm just not sure.
That may make sense given the ability, uh, with things like copilot and co-generation and stuff like that. Um, inefficient allocation of human resources. Yes.
But, okay, so what are the, the threats? So the threats are, we're, we're really, this is a whole new world. You know, one of the things I talked about, shadow it, which was if you, if you go back in sort of the history, and I've been doing this five decades.
So I, I, like, I, I have a good perspective of all the things I've done wrong, things I've done right? Whatever industry has done wrong, what our industry has done, right? You know, when cloud first came out, and I was actually considered one of the early clouderas, and that, it's a silly term, but there was a group of us that were considered people that you should listen to in the cloud.
And I was one of those early hundred or whatever. And, um, and the thing was, it was confusing to a lot of us who were classic sys admins. This maybe predates a little bit of DevOps.
Maybe at the same time, DevOps was being created. Um, you know, for the, those who don't know, I, you know, I was the only American, the first DevOps stays. Um, I created me and Damon Edwards and a few other of us, Andrew Cliche from Mark Hink, who created the first, uh, dev stays in the US at LinkedIn.
Um, you know, so like, I've been around. So the point being that, um, what cloud seemed really confusing at first, and even like the lamp stack. So to seem confusing from a classic CIS admins perspective.
But then we realized at the end of the day, cloud was just virtual. Say it was at the end of the day, it was network compute storage. And like, oh, okay, well, that, that abstraction to start an instance was different than VMware, but like, it was still a virtual.
In other words, we got over it pretty quick and we're able to normalize our knowledge. I would argue that in this world, it's gonna be a longer, um, slope, uh, or tail to get over it because it's a whole, everything's different. It's non-deterministic.
It's probabilistic. It's, it's mostly comes from academia, right? So, so the immaturity of like what you would expect is what Val and sticks and CVEs and the, like, it's getting good.
And we're starting to get some good research on bug bounties for regenerative ai. But, but the point being, like these adversarial attacks, and I'll, I'll give you some examples a little bit, are just far different than anything we're used to. Or they, they meet some criteria of an old pattern, but they're done, delivered in a way, like, oh, wow.
Right? And you'll see some of those. So there, in fact, you'll see it here in a second.
New malware opportunities, right? Um, there's a number of really interesting, um, that I've been tracking, uh, you know, sort of like how the adversaries are taking advantage. One is, you know, when you use like code generation, like copilot and stuff, right?
Uh, just like, like when you ask chat GPTA question, it can sometimes hallucinate and there's a science behind that. It's probable answers. And, and, but then also code can hallucinate.
So what you'll see sometimes is you'll ask for some code and it'll give back a library to install, and the library doesn't exist. And then, you know, if, if you didn't know about this hack, you basically run in and says library doesn't exist, then you realize, oh, that was not really a, that was a different name. Or maybe it's been renamed or, but what the adversaries are doing is they're going out and finding all those hallucinations in code, especially in libraries, and they're actually in installing those libraries so that when you basically run that the library runs and maybe it sort of does what you think it's supposed to do.
It's send the covers, they're ping you, um, the, uh, like the HuggingFace stuff like Jfr has done incredible job like documenting, and there's others. But I've been, I, you know, I'm a good fan of j Jfr. Jfr is a good fit, a good, um, community.
Like they're part of the Textron community. I'm part of their community. So like, uh, like, we like them, we like Sona type too.
So we like 'em all. But, um, but I will say that, um, they're, um, the, the, they acquired a company called Voodoo from in, in, uh, uh, from Israel a couple years ago. And they're just these, they are really figuring out some cool stuff.
And so HuggingFace, so like, sometimes everything new is old or whatever. Like, you know, we ran into this with like, don't just install a puppet manifest without testing it. Don't just install a chef recipe without testing.
Don't install a docker image without testing it or sandboxing it. Well, it's same thing. Hugging base, hugging based.
The adversaries are, you know, one of those things that's happening is you, is you're embeddings tends to save or not, right? Because you can have, uh, executable bike code in embedding, right? Like, and so like, again, the adversaries are really out to hunt right now, and they're, they're keeping up with this stuff faster than, so there's these phishing teams.
There's like, you know, you know, I talked about, you know, like the, the idea of, um, you know, ations on certain libraries or certain, like the, the belief that is all the sort of things that you get from a a, a copilot is authoritative, and therefore, like very much like you sort of, you like not to pick on like your aunt or uncle who doesn't know it. When they get this question from some company that says, Hey, you know, we, you, we know more about your account because it's been compromised and you hit the link, right? Well, that because you, if you went without some education, you believe there's an authority.
Like, oh, it came in an email. It looks like a very authoritative email. Let me hit that link.
Like, most of us have gotten really good at that, but like, what? It's a redo. Now.
It like, what if Chey B tells us, Hey, by authority, I am telling you that this is the answer, you should go here, right? Like, that's happening. Uh, reverse engineering is an interesting problem space.
Um, the, I told you about the, you know, be able to reverse engineering the, uh, the copilot for email, right? Or a lot of these embeddings are in the wild. So if you've got enough CPU resources, if somebody has how your data's been vectorized, it's not incredibly hard to reverse engineering from how you would search to how you get the data to actually just give me the data, the original data.
So, um, we're working on some, like, how do you encrypt, uh, embeddings and you don't really encrypt embeddings, but you could do matrix multiplication. So it's really cool. So stay tuned to some of the stuff I'll be writing about innovating acne.
And then, you know, so the defects, I have it last because it's not that it's not important, but it's, that's the one that's like very glaringly obvious. The ones that aren't obvious is like the hallucination on code. Uh, you know, the, the embeddings in, in, in, um, you know, the, the, the, you know, the sort of the bike code in embeddings or the, um, reverse engineering.
So, and here again, again, um, you know, I had to read when I first saw the OS top 10 for LM applications, ta, here we go again. But it wasn't until their paper came out recently, um, which is like an owas for LA Generat AI or whatever. And I, I, I'll, I'll try to get the link and I reread these and I'm like, you know what?
This is all happening and this is like, good job, oasp really good. I mean, Joes has great, right? Like, we like the work they've done over the years for security is, is, you know, incredible and honestly, um, not always a hundred percent accurate, but like incredible, right?
But prompt injection, this is a real problem. Like sort of saying who I, you know, I'm this, and therefore you should give me these answers or insecure output handling, right? Or, um, you know, poisoning the data model, denial service supply chain.
I mean, like these s SLMs are interesting, but like, we're gonna hear stories about somebody basically taking, uh, a data brick, copying a small s lm and walking out the building with it, which might have like your algorithm algorithmic trading a copilot for algorithmic trading. Like, like, trust me, this is gonna happen. Sensitive disclosure, um, you know, excessive agency, overreliance, you know, uh, model theft.
That's the one, right? Where like, I think with the LLMs or the large language models, the ones you train, they're gonna be a little more difficult to get those. But if we're throwing soms at the edge, um, you know, I mean, you know, I remember when, uh, you know, I heard the first heard the story about like running Kubernetes in all the, uh, Chick-fil-A retail stores.
Like yeah. Is the person who cooks the fries gonna have to then go reboot the Kubernetes cluster or update the crud? Like, like, uh, like, but like, it's not like, like we probably will see some of the weird edge cases of like, um, maybe GPUs running on the edge.
And, but the point is it brings a whole nother set of threats, right? And so, like I promised you, um, in the beginning, I don't have all the answers. I am incredibly interested in, in creating community, uh, quote, dear CIO double quote as a paper, as a, let's all get together, let's get the protectors on the same page.
The people who DevOps DevSecOps, SRE infrastructure and operations, like the, like, what we do. And let's not let you know, one of the things that like, that concerned us early on is we're seeing a lot, a new, couple of new positions being, creating a chief data ai, a chief CDAO, uh, A-C-D-A-I-O, a Chief Data AI officer. Um, yeah, it was, you know, I mean, like, there, there charter might be inconsistent with the CIO's charter, right?
Because they're gonna be fast moving. Let's get something. Maybe they're, maybe your new chief AI officer is some brilliant, uh, Stanford AI professor.
I don't know how much that professor's gonna care about GDPR or uh, GRC risk control audit. Um, how do you just roll, roll up your sleeves and protect the brand? How do you not let a CA Canada happen?
I mean, again, I'm not saying they won't, but that's sort of what we're gonna try to discover in the paper. Um, you know, I, I heard a large insurance company recently where now, uh, they don't know all the answers, but they're creating required training for everybody in the company to take this sort of like, checkbox, did you check off that? You did take this training?
You can't get, you know, you have to get through the, we've all been through this. You have to go through the training so it knows that you actually did. At least you've been told what the corporation thinks you should do and shouldn't do.
So you can't do the do now ask forgiveness later. Oh, well I didn't know you weren't supposed to use, uh, OpenAI's chat GPT No, no. You were told if you're going to use this, you had to get a request and you had to go through a formal process and you had to use Azure OpenAI, right?
Like, you know that now, now you can't say, well, I didn't know, right? So I think that's a really interesting first step, um, platform engineering, right? Like, like, like we know platform's gonna play an important role here.
Patrick Abar, you know, the godfather of DevOps is doing a lot of explanation here. I think SRE is gonna have to get involved way earlier. Let's not just wait to say, oh, you know what, maybe SR like, remember my, my comment earlier, I don't wanna see an organization like a bank that has 30 vector databases, you know, 15 variations of orchestration, a hundred different models, which were then none of 'em are sort of curated in any software supply chain.
You know, what if SRE started becoming really good at like, say, um, you know, um, vis and, um, and, and Mongo to be at all Vector search, like, and we asked everybody who wanted to be SRE supported that you have to have one of those two vector names. 'cause those are the two we're really good at those ones we run at scale and like, so, or all the other lama, like, like, and then you have to have an exception or like, we won't manage it, right? I think this is gonna be incredibly important for, for SRE to get involved and be educated and learn this and ask these questions.
Secure supply chain, right? All this development of our rags, our model embeddings, all like, we have to treat that like a software supply chain. And then everything we did with automated governance, you know, if you read the Investments unlimited book that was co-author on like, we should be creating digitally signed at the stations of the decisions we're making to create these checkbox and co-pilots so that when we get audited or we in fact have to explain a breach, we at least have evidence of decisions we made to get you that answer.
And even though in the middle there, there's a really complicated neural network, at least from a human perspective, we've explained our rationale. Um, yeah. So I mean, that's, uh, there's a couple more things too.
I think as I got a minute maybe left. Um, one of the things I think might remember, I, early on in this presentation I talked about always sweeping under the rug, you know, the compounded technical debt at each generation, you know, maybe now, and you know, me and Josephine Knox of enterprise vision technology working a lot on this is maybe now's the time to look at all that generational technical debt. And since the tools to do these conversions, you know, uh, I think Google is doing something really interesting.
They have the mainframe assessments program where they're looking at COBOL and JCL and, and giving you tools to convert that. But what if, like, and that's cool, but like, that's a small segment of the real complexity of all the sort of the stuff, the scaffolding we've created in large banks and insurance companies and retail. What if now is the time to step back and say, you know, and, and one of the things we're writing on the paper is like, is there, can we create this idea of innovation tax?
Like for every dollar you save on generational technical debt, you can now use $2 for innovation. And I know people are like, well, you know, and this is so like, like against the grain, right? Because every CEO is like, we've gotta be, first we gotta do this.
But it, I, there's no question that like I think this technical debt tsunami I talk about could bury us. This might be the time that we literally, um, you know, don't get outta the hole because the breaches are gonna be so hard. The technical debt that we have will be so complicated.
We've got so much in our large infrastructures that we've just been ignoring, ignoring, ignoring. And I understand why, 'cause I, you know, things that work historically for 50 years, but I think if, if there was ever a time to do it, maybe now is the time to do this, right? Because the technology of allowing these tools, and in fact that you can send your cobalt programs in JCL to Google and it gets it running on GCP is pretty phenomenal, and it actually works this time.
But I'm saying that's just a small segment of all the things that connect all the things that kicked an iPhone from a large bank all the way back to a mainframe system of record right Now. Maybe it's a time to think about, and these are some charts that Joseph Enox and I mostly Joseph, have put together some interesting data points of like how we might wanna think about innovative tax, how do we get performance? Uh, we'll be writing a lot more about this.
In fact, this is a lot of these charts are gonna go into the paper we're writing. Um, yeah. You know, and I'll end with, um, you know, I like, I think, um, my, my sort of more prolific work prior to this gen of AI was this book, it took 10 years to write.
Um, in fact, if you do any of my workshops, I use my book as the source to learn how to create gen of ai. You know, so like you use the book to ask questions about my book and I teach you how to curate the data, how to do the data ops, how to do the junking, if anything makes sense, how to create the high accuracy, the observability, componentries, the rags. And so all my workshops include this.
So I include this as my final slide. So anyway, thank you so much. Um, and uh, you know, again, I'm mostly known as boop, most places on LinkedIn, if you go to John Willis DevOps or John Willis Atlanta, pretty easy of fine.
I pretty much hang out on LinkedIn pretty much all the time now. So thank you so much.