Techstrong Gang – July 22, 2024
Mike, Mitch, Jon and special guests Tracy Bannon, John Willis and Camberley Bates, a chief technology advisor for data infrastructure at The Futurum Group, dive into the root causes and downstream impact of a catastrophic Windows outage caused by a CrowdStrike update. Then, the gang turns its attention to closing the divide that exists between developers and senior managers before delving into a growing uneasiness about the current state of artificial intelligence (AI) model security.
Transcript
Hello, I'm Mike Vard, and welcome to the Textron Gang. Alan Shimmel is on vacation for one more day, but he'll be back next episode. In the meantime, we're gonna be talking about this wow, mind blowing windows outers that occurred last week.
And then following that up, we're gonna have a chat about, there's a little bit of a disconnect between DevOps teams and their managers. And then finally, we're gonna take a look at some interesting new research that's come down the pike as it relates to vulnerabilities in AI environments. You're watching Textron Gang, and we'll be back in a minute.
All right, let's introduce our gang today. 'cause we got quite the lineup here. We're gonna start with Kimberly Bates from Futurum.
Kimberly, how you doing? Pretty good this morning. Just got back myself for a little bit of holiday.
All right, boy, people are taking vacations. This is a good sign, I think. All right.
We also have joining us once again from the Bay Area, John Schwartz. How you doing? I'm doing well.
Thanks Mike. Good to see y'all. And then a fellow we haven't seen in a couple of weeks, John Willis, who's one of the, uh, early godfathers, I guess, or concierge for the development of, uh, DevOps workflows.
John, how you doing? Yeah, I would say I'm the consigliere where Patrick Debar is the Godfather. So, there you go.
There you go. And then of course we have Mitch Ashley outta Denver, as always. Mitch, good to see you.
Good morning. Good, good to be with everybody. And then finally we have Tracy Bannon, who's one of the thought leaders in the realm of DevSecOps especially, and well, we're gonna have a lot to talk about that subject today.
So Tracy, good to see you. Good morning. It's good to see everybody.
Right. Kimberly, let's start with you for a little bit. Um, most, uh, folks spent the last weekend helping to clean up a lot of this, uh, outage and the complexity of all that is still unfolding.
But what's your sense of, what was the financial impact, the downtime, I mean, as, as, as events go? How big was this? Well, the list of, uh, kind and kinds of companies that were impacted is pretty widespread.
Um, the, the clearest one, I think probably people, especially since we're just talking about getting back from holiday, um, who's the airlines? So across the board we saw United Delta, American, Allegiant, frontier Spirit, just, and I'm sure there's some others that I, i I missed that were impacted. And they were either constantly flights to laying flights, all that has a huge impact to them.
And the amount of work that has to go into getting things restored. Um, we've also seen nine one ones impacted, and that goes beyond the financial, that's people that's, that's hurting people. Where we've had, I think it was Alaska that said they had a 9 1 1 down.
Um, there's a couple others that have had 9 1 1 down. Um, that's, that's a big deal. Um, hospitals have been hit.
Um, bring 'em, yeah, bring 'em up in, um, mass, Massachusetts, um, got hit. There was a couple of other ones that were, shouldn't say hit. They were down.
Um, it, what it appeared to be is patient records weren't available. Um, and what that impacts is the ability to serve the patients as well as they were saying they were going to, um, delay, um, operations. Dunno if that was critical operations or how they were doing.
We just critical operations are gonna go forward, but maybe, uh, Uh, elective operations wouldn't go forward. So we have that. We also talked about, saw that the news was hit or the, not the news, but the channels communication challenge.
So we saw Australia got hit, but I also, on Friday, I was listening to, um, to today's show. And, and they were hit, they came into, into the blue screen to death. So you've got, um, overtime, et cetera going on.
So you, uh, we saw Cloud strike, you know, took a, took a hit on their stock. Um, and we'll see where that ends up coming back. But then, then you have the trust factor that you have to reestablish.
So, do I trust Cloud CrowdStrike? Um, do I continue with it as ubiquitous as they are with people using it into the, into their, um, computers and their, their, the PCs, servers, et cetera. Kimberly, I was gonna add two more, uh, areas that I'm, that I observed.
One is actually with consumer. My daughter-in-Law is a manager of a Macy's store. And their systems were offline, no registers.
They were not able to, to contact their people. All of their corporate communications were down. Um, as well as some of the state level agencies, um, that are caught, that are dependent upon that are Windows dependent and, uh, head outages with their customer service areas.
So, uh, a lot of impacts that are trickling across, uh, or have been trickling across quite a few different areas. Yeah. There's Mitch, What, Mitch, what went wrong?
Let's, let's get into the details here of, of exactly what went wrong, Mitch, to as much as we know about it at this moment here. 'cause I don't think everybody understands what CrowdStrike was doing and how that kind of connected back to Microsoft. So give us a little bit of the technical details there.
Yeah, it, um, actually, it, it isn't too technically deep, at least what we know so far. It was primarily an update that would, was issued to some anti-malware security software, uh, in cloud Stripe, uh, claimed that it was content, not code, which implies things like signatures or whatever drives the software. Regardless, it's the same issue.
The, you know, the impact is, you know, just as big. Whether it was that or something else. Um, the, the, the, frankly, the issue is when you cause this blue screen of death, um, to a, to a Windows server, there's few other options to recover from it.
And it's tough to apply a patch to a piece of software that won't boot an operating system that won't boot. Now, Microsoft has initially provided the procedures as to reboot multiple times as many as 15 times, and it may recover. I'm not quite sure why that that would help you get there, but apparently it worked in some cases.
The issue is, you know, this is what Kimberly and and Trace were describing as a societal, uh, level impact, right? You may not be running a Windows server on your phone that you're doing a square transaction with for your cab service, but it impacts you. And, you know, a lot of our technologies, yes, they're virtualized and in the cloud.
So if it's a matter of applying a patch or an update and then rebooting, that can be done by operation staff that know how to do that, whether it's in your own data center as well. But we also have a lot of computers who are sitting out on top of an oil well sitting over a blue screen, maybe don't have a power power cycle capable, um, AC plug, uh, that it's, uh, that it's plugged into to power it. 'cause that's one of the ways you can rebo computers is not just go up and reboot it, but actually turn the power off to it, turn it back on remotely.
So this is one of those, the ripples of it, or the layers of it to recover, you know, are certainly not small, especially given the widespread impact on so many different industries. And can you imagine, if you didn't think to in install a, a remotely manageable power strip up in the, you know, the, uh, monitors that are showing what the latest arrivals are, wherever those computers are, that you need to re review reboot, if you can't get to those very easily, you know, you may be cl up on ladders and going into closets you haven't visited in a long time, And go, John, you've been doing this DevOps thing forever. What's your assessment of, uh, how did we wind up with a single point of failure like this?
And, and what, what should we learn from this? Well, one comment I wanna make on the, if you're running Windows servers on the edge, you're outta your mind. It happens though.
Dust John. Um, but, uh, no, I, I think it is really, you know, it was like we, we've seen this movie a couple of times now, right? It, it's, uh, you know, I mean, the solo ones was adversaries, but it's the same problem.
I, in my travels, DevOps, I do get invited into a lot of software companies and, and it is amazing to me, you know, how they don't eat their own dog food. Um, you know, I've seen like prominent, well-known software that, that's running in, you know, almost every Fortune 500 company. Uh, yeah, we saw the solo ones for sure, right?
But with the, the, the half, half of their, uh, supply chains, they're not even running ci, they're still doing manual, um, let alone CICD. And then, like, what we've learned over the last, you know, four or five years is there's a lot more. There's salsa, there's SBOs, there's, you know, um, you know, just the whole, um, uh, you know, so many different tools.
Automated governance, which I've spent a lot of time on, right? You know, wrote a couple of papers and books about it, right? So like, I think there's the need for these vendors to be more transparent than just SBOs.
You know, the whole industry thinks if you have an SBO m then everything's baloney. It's a spar is just like what you have. It doesn't tell how you do your operations.
But the second part is, I think there's an authority bias here in that, and this is why I think I call out the DevSecOps and DevOps vendors more than others, is that I think in the case of now this is, uh, counterfactual, but I think in the case of, um, CrowdStrike, I suspect the, the vendors did, I mean, the customers didn't seem to test any of the software before they deployed it. Um, don't know that for a fact. Um, but, um, but the, you know, but the point is, I think they think because it's cloud CrowdStrike, and they're just, and they are, that's a, they're a great vendor.
I mean, in all research and everything. So, uh, and, and stuff happens. So, but, but the point being, I think people probably accept their software easier than others.
Maybe because they think, well, there's security. They're probably gonna do all these things that either on dog food. So I think that those are the two things, my observation of like what, you know, the, what the problems are here, Tracy is failing the test.
Part of the issue here is that I, I know you've always been an advocate for that particular thing, but it seems like maybe we just do take things granted. Well, It depends on who's responsibility it was to test, to John's point, um, um, is it a managed service that I'm using? And so I'm depending on the managed service to do the testing as opposed to my humans doing that testing.
So we do have to figure out the roles and responsibilities. I know that sounds trite, but we actually do need to make sure that we understand that SLA, um, there's a lot to be said about something rolling out at this level, uh, this broadly and having this big an impact. Uh, I'd like to see what the forensics show, how did they test it before they rolled it out, right?
Root cause analysis is massively important. It is not, uh, intended to be a witch hunt or scapegoating, but how are we going to learn and how are we going to make sure that we don't do this exact thing again, because we are seeing it again and again, and again, and again. To John's point, we need to help people get over some of these bad behaviors.
You know, trace, one of the things that, that this brings to me is kind of like security incidents will happen, right? Mm-Hmm. Software bugs will happen too.
And I'm not saying an instant like this widespread is common. It's not, but it, it can and will happen again. Mm-Hmm.
What it, what it makes me think of is, yes, we need to examine the software procedures, testing all of those things to see how it slipped out. It's also the deployment. If you're a, if you are a major software vendor, you're running in, you know, cloud, you know, hyperscalers environments, you're running in customer environments, do you really need to, uh, do every update cascade it, you know, within a matter of hours or have half a day?
Or can you really kind of ab test it and slow roll your initial updates to verify that things are still working? Well, I think if a little bit of patience, and, you know, it's easy for me to say Monday, Monday, Tuesday morning, quarterback in whatever day of the week we will with an event happens. Mm-Hmm.
But to me, it's also process. I mean, unless it's a serious, uh, security flaw, I think you're much wiser to take cautious, cautious, be cautious about how you roll things out, the pace of it, verify that things are working well, you know, customers are reporting problems. Continue that.
So maybe slow down a little bit. Well, Mitch, I'd like to just to shim in the, I think there's a too big to fail problem too here. Mm-Hmm.
Like solar winds, who knew? I mean, I think you, anybody who knew SolarWinds knew, but who knew, right? And CrowdStrike, like, I think these under these radars too big to fail.
And again, it goes back to what I said earlier. It's, it's about software, supply chain hygiene and transparency. And I think if, if we could identify the two big to fails, and we could basically force them to be more transparent than just providing SBOs, um, I think, you know, maybe we take it a little further to, to prevent these type cascade.
And they're gonna get worse, you know, with, with, with the technology exploding the way it is. So, John Schwartz, do you, John, go ahead. Lemme ask you, lemme bring John in here for a minute, guys.
Um, John, is this a, will we see more transparency? 'cause it is kind of a black eye, but, you know, no, Jesus. I mean, given, given what's been going on in the news cycle and on with the panel discussions we've had over the last couple weeks about security breaches, we're talking about open AI among, among other things.
Uh, and, uh, Google wants to buy Wiz, which shows it, what it thinks about cybersecurity is cybersecurity's their biggest acquisition ever when it happens. What, what's really unfortunate, and John mentioned, well, CrowdStrike is this, well-respected company that was fairly not well known. Well, now it is for all the wrong reasons, right?
This is a kind of a cautionary tale, uh, about technology when it goes wrong. And the one thing that I've always come across when reporting about cybersecurity, and I used to report a lot about it about 10 years ago, was people only care about outages and security and lost IDs when they are directly impacted. And I can't think of an incident that has rippled across North America and Europe, Asia, Australia, South Africa, what have you, over multiple industries.
We're talking about hospitals, airlines, banks, media companies, payment platforms, uh, online shopping websites. I mean, this, this, this was not only created a lot of collateral damage financially, but it could physically to people, uh, in hospitals. I mean, this could end up not only costing CrowdStrike, but where do we see what is happening to its stock?
But I'm thinking about class action since lost customers damage to its reputation, which kind of brings me to this point. And something I want to ask of the panel. What, what can it experts do?
Or what, what, what, what can it service leaders learn from this experience as, as well as how to protect from something like this happening in the future? I'm not sure if they can, but, and how do you handle a situ situ situation like this when it unfolding? It's a great question.
When you're, when you're running an i it organization, you have so many suppliers that you're working with. Even if you're closely aligned with the one like Microsoft, or maybe you're using a, you know, a lot of their technology or you're an open source or whatever it might be. It, it's still just like we have diversity of suppliers into buildings for broadband, uh, network access.
You still have to have some level of diversity in your suppliers. And, you know, we make decisions of saying, I'm gonna move all my, all my mail to Office 365, or whatever service you choose to use, you're now entirely dependent upon that service. And if something happens, it goes down.
And what are your alternate communications? Um, do you, do you suddenly fall back on, sorry, Mike Fazar Slack or other, other forms of communication to do that? Especially when it comes to, there's transacting business, and then there's just basic communications.
And when your servers are down, that can, all, that can affect both very easily. And so you're almost in a disaster recovery kind of scenario, or at least in, let's say a business resilience recovery mode of what do we do? What's happening?
Um, so you need to have clear communications with your suppliers so that, and we want them to be transparent with us Mm-Hmm. Um, and not over communicate and tell us things that don't matter or point us in the wrong direction, but what we need to be doing in terms of how we recover. And, and Mitch also point out, go ahead, please.
I was just gonna say quickly, I would, you know, I think there needs to be transparency in like a postmortem that said, you know, what was your, uh, software delivery? What was your supply chain architecture, you know, is the saucer supply chain levels for software architectures? What was, did you have any saucer level?
Like, I think these are things that nobody's asking for these two big fails. Like, we're not demanding postmortems. We're, you know, and we just keep, you know, these software companies just run, you know, sort of amuck when it comes to hygiene of things that, you know, other companies, like really big companies do really, really well.
Banks do it really well. Well, I'm also wondering if there's a, a big issue here on the blast radius. It, it's taken, how did it become this big and this broad when we are very dependent on a company that is, it sounds like, I mean, I'm not in the security side of the business side, the infrastructure piece of it that has to do with the data, the availability, reliability, protection area.
But one of the things that we always talk about, about is, you know, keep your blast radius at a point that you can, can make. I mean, what you're going through right now, they're going through right now is, you know, what are the critical apps that I gotta bring back up? You prioritization of all that.
Um, and how do you bring those applications up? And a, if you have to bring up the application, or is it I just re bringing up the endpoint and does it mean that I actually have somebody physically in some, some of the stuff that's out there saying you have to have somebody physically there to bring that endpoint device back up. I can't imagine if I'd walk into, you know, DIAI live here in Denver, um, and, uh, that, you know, you would have to walk through all of Denver and bring up every one of those streets, um, because they're linked to some sort of, you know, PC server or whatever.
That's, that controls 'em. I mean, it's a nightmare. In terms of the recoverability.
This goes back to continuity planning. Like there are lots of hygiene things. John mentioned this earlier about from a, a DevOps perspective.
And, but things like business continuity planning, which includes debt, disaster and recovery. Making sure that you have your organization has a commitment to understanding what your resiliency posture is, right? We table pop these things.
If this goes down, what will happen? We need to understand those. And having an SOM is important, but it ultimately, if we have an ecosystem that we are in charge of understanding where we have dependencies on other organizations and their SLAs, understanding our continuity plan and having a very specific understanding of when this electric socket fails, what are the other things that are going to, to fail downstream from it is, is good hygiene.
Right? We're, we're back to talking common sense again. Maybe common sense isn't so common, right?
All guys, guys, I'm gonna have to leave it here 'cause we're full up on 20 minutes already, and we could go on the subject probably for the entire day. And I'm sure we'll revisit it again during the gang sessions that are coming up. But I would just point out that, hey, maybe as part of your resiliency panel playing paper and pencil, check it out.
It kind of works from time to time. We'll be back in a minute. All folks, we're back and we're talking about the relationship between DevOps teams and their managers.
com about a pair of surveys that went out about this particular subject. And it was just surprising to see how much of a disconnect there was between what the developers said they were doing versus what the managers thought they were doing. And, um, for instance, there was a significant gap between the number of folks who were managers who thought scanning was happening on a regular basis versus what developers were saying.
Tracy, let's start with you. It's not generally a surprise, especially in bigger organizations, that there's a disconnect sometimes between managers and developers. But is it too wide right now?
And is it because of the complexity or what's going on in the rank and file? I think a lot of it has to track back to where we are in the growth cycle organizations continuing to automate sharing the data about their automation, sharing the data about what developers are doing. There's always this balance, and John, I'll probably talk to this, there's a balance between understanding all of the data about an individual developer and exactly what he, she they are doing versus the team versus the organization.
So those managers don't always get a clear view because they're just asking word of mouth. They may be not even walking down an aisle. They may be looking at a Slack channel, but not looking at the data that's being shared.
There's also a responsibility, though, for the developers, for anybody related to the SDLC, to be raising their hand and letting people know what they're doing. 'cause the ultimate value is when you get working software in production. So if everybody's not rallied around that these kind of things will continue to happen.
The the impedance right, is about human transparency right now. And that honesty that has to happen. John, what are your thoughts?
Yeah, no, I agree a hundred percent Puy. Um, you know, I think like you, this subject could go on forever, ever, ever, because the, you know, the technology mapping is a mess, right? It, it kind of, it, it even bleeds into the last discussion about hygiene.
I mean, developers are frustrated because, um, just the whole dependency map of technology is not well understood, right? Like, you know, you write 10 lines of code and it generates over a million, you know, a hundred thousand lines of code under the covers, you know, and all those libraries, you know, we would talk about outages and, you know, I mean, what, what happened about five or six years ago? Uh, disgruntle open source developer removed, uh, a blank and a third, the internet went down, right?
Literally, I mean, um, uh, in the JavaScript object code, right? But, but I think that the, one of the most interesting things to narrow it in is, I think the work, you know, I, I don't think Dora helps us. I don't think space helps us.
These are the sort of developer productivity metrics, but I, I, um, you know, I've been pretty vocal about, like, Dora helped us get to a great place in DevOps. Mm-Hmm. But, but I, I don't think it's useful, especially in these kind of discussions, you know, about manager, developer, you know, gaps.
I think the Dev X stuff though is really, really fascinating. So if you look into this stuff, it's, it's about monitoring feedback loops, cognitive load flow, state interruptions. And so, like, again, it, it sort of, when a survey just says that 97% of, of, you know, developers are dis crow, or they're wasting 70%, or say eight hours a week are lost time without really trying to understand how do you get to find out why.
And I, I think that I, I think out of all the metrics I've seen, I think most metrics that, that people talk about and show on slides and present about developer productivity are enable gazing. They don't get to the bottom line. I do believe what, what the Dev X folk had done, and, and Nicole, um, and, and, and a group of people is, I, I think this is the right path.
Finally, there's some metrics to me that, you know, I don't want to hear how many story points you had, right? But what I, you know, this, this stuff and the Dev X stuff really seems like, you know, that's where I put a, a focus on if I was concerned about the, these surveys and these gaps. Well, isn't it really, it, it's getting back to becoming human-centric.
That's what this is all about, is that we have, right, we know about taylorism and trying to replace people as though they're cogs and having no sense of what they're going through. Um, developers, It's not gonna bring up Deming, are you? No, no, I'm not.
But I've learned, I've learned from those who are the, are the, are the students of Deming. Um, but focusing on the humans becomes such an important part of this. We've learned it over the course of our careers, DevX and Rel, right?
This, this new role that's emerging that is somebody who is championing on behalf of the developer is a very, very important part of this mix. We've got to take care of the humans. Um, and we have to be consistently focused on that.
I mean, dev, we think about DevOps. DevOps and agility have been about making life better for humans. Um, but the metrics don't always help us understand the humans, uh, in this, uh, one final thing is that, um, I think we need to stop talking about individual developer productivity.
Yes. I, I think we need to stop talking about individual developer productivity. 'cause a single developer doesn't deliver team metrics.
That's okay. But let's understand the promoter score for that developer. Let's understand their satisfaction.
Let's understand their contribution. It might not be a story point that might not be burning down any kind of debt that might be actually educating the guy beside them. So there rant over.
Mitch, Mitch, jump in here for a second. Are, do we, Mitch, do we not want to admit that maybe we have a problem? And we're always kind of like on the IT side, we're trying to present the fact that we are the people who bring order to the chaos, and yet our own systems and our own processes are often chaotic.
So maybe, you know, is this kind of a moment where we just need to all stand up and look at each other and say, you know, it's okay not to be perfect. Hi, I'm Mitch Ashley and I make developers unproductive. That's the meeting we need to have.
I'll tell you a whole story. One of those humility you, uh, to give you some humility moments. One of when I first became a manager, um, at our, at our holiday party with our family that year, one of my, my cousins asked my daughter, who I think was about four years old, what does your dad do?
She said, well, he talks on the phone and he goes to meetings. And that was her definition, you know, outta the mouth of babes, right? And I think a lot of times what managers do is forget that if you're gonna have somebody paint your house, I'm not selling saying a developer's a house painter, but you're gonna have somebody do a job, paint your house, whatever it might be.
You're not gonna sit there and talk to 'em for six hours and then expect 'em to get a full days of work done in. Mm-Hmm. But we do that to our development teams.
'cause we think we need their expertise. We think they need to be in every meeting, um, that they may possibly need to be into. And they hate that.
They hate going, they hate wasting their time. 'cause all the time they're sitting there thinking about what they've really gotta do. They've gotta go back and do their real job.
So I think a, a, a source of this, and I agree with what Mike said. What, um, sorry, what, uh, Tracy said and John said, um, but in addition, I think leaders, you know, look at yourself for a moment. This, if, if these folks are really hard to find their valuable resource, we need them obviously to do great things with software.
Treat their time as as extremely valuable and use it wisely. You know, if you can do, do a 15 minute meeting instead of an hour meeting, or you can do with half the number of meetings, cut it back, just cut it back. Let these people go do their work and make sure that people are communicating either in other methods or when you do get together, you're using their time really wisely.
That's, that's my advice. I went that, that after Thanksgiving, I went back to work. I'm like, okay, what meetings do we need to cancel?
My daughter says, that's all I do. So guess what? I had lots of, lots of volunteered ideas, All kinds of things.
Asynchronous communication is a big part of this. How do I get out of a meeting and how do we move some of these things? Like, you should be able to, if you're doing your work in a way that's transparent, sometimes the manager doesn't need to come by, doesn't need to get an update.
They can go pull the information that they want. So there's some enablement and some transparency that we need in all this. And let's let people have conversations in Slack in whatever other async channel that they have, because that enables them to get their work done.
Right. Meetings, Meetings, yeah. Meetings are always, I mean, they're just counterproductive for the most part.
Not just for developers, but for anybody in general. Hey, Mike, uh, what, you wrote this fine story. What, what is, I think, what can I ask you what the impact of AI has on all of this?
Is it complicated or alleviate matters? I think that's still unknown. I theoretically it should reduce some of the toil, but, um, my sense of things, and John, you've been an early proponent, but I don't feel like, you know, AI's gonna magically transform everything about software development overnight.
I think it's gonna take care of some, you know, small, it, it scratches. I think it's two, two sides to that. One is there, it it is creating a tension point, but there's a tension point in that there's a belief at, at levels and a lot of executive levels, C levels that they don't need developers or they don't need as many developers.
Uh, and that's trickling down to sort of, I think in one of those surveys they even talked about sort of the angst of ai, right? So there's all sorts of angst and then there's sort of the part of the old school developers who are just going to resist these tools. But the truth is, um, in the end, the way we do everything we do is changing under foot right now.
That, that, that, um, you know, I'm not saying that, you know, I mean, if you looked at like, um, day traders 10 years ago, like Goldman Sachs would've had, um, you know, probably thousands of day traders. I think now they have like less than a hundred. Right?
That's 'cause of algos, right? Um, and it's the same thing, right? Like the, we're gonna see this change and reduction in the way things work.
0 is significantly better than G PT four, uh, the, you know, the, the, this NewLaw three five that's gonna come out, right? And GPT five is coming out like, like you can't ignore that these tools are changing the way we develop or create things. And, uh, but, but I think, again, there's two parts of it.
I think there's the create people who are resisting it. There's sort of the tension from above of like, we're gonna, we're gonna get rid of all develops. And we're not there yet.
No. But that's a meme that's happening at the, uh, seed level and a lot of large corporations. I Wanna, I wanna get Karen Building's opinion on that particular topic, if you don't mind.
Um, Go Ahead. We, businesses are more dependent upon software than ever, but I don't really feel like business executives have any understanding of this conversation we just had and what the implications of it are. And, um, so there's a gap between the managers and the development teams, but the, the gap between the business and the rest of the software development teams seems to be even wider.
I think what you're talking about is that we've changed where the factories are. We've started talking about this AI factory, but we've had this code factory for a very long time. And for the longest time it's been part of it.
It's not been ubiquitous to the, um, and what I was as, as I was listening to you guys, because you, the, it's a different world that I've spent in. I was kind of thinking about my dad, who was a captain of the ship, and one of the principles that he put in me was walk the ship floor, and he would always go down to the engine room and walk the engine floor. Um, and part of that was beloved by the people of the ship.
And I know that many of the people that have served under him in the, in the military as well as in the commercial world, but because of that, he was able to get more out of them. And you have in manufacturing people that walk shop floors, you have in mining people that walk the mining. And I'm wondering because we're so digital and because we're so, as you were saying, async, not synchronous communications, but async communications, and we sit behind these screens and we sit along and many, many times that ability to work together as a team is lessened than what it was.
Maybe if you're working on a construction field, if you're working in a mining operation, and maybe those disciplines need to come back, not only just for the, the team, but also for the management layers that come in and say, walk the code floor. You live with those people. So I've had remote teams globally dispersed teams since 2007, and I'd argue with that, that we've had phenomenal results.
I agree. But it was because we took protocol that we could emulate What happens if you're co-located in a building. So we have our virtual coffees where we get together, we have some rules of the road for where we pop up and just have conversations.
And those are allowable. We're not looking at your mouse clips to make sure that you're doing something. We encourage there to be that interaction, that synchronous interaction.
So I don't know that it necessarily is a force back to office, physical co-location as much as it is an intentionality to connect humans to humans. Focus on the humans, right? Focus on the humans.
We'll just keep saying that. Absolutely. And I think there's skills that we have developed or, you know, been forced upon us or developed as non coders in the world when we go, went through Covid that we didn't know how to operate in that way.
Mm-Hmm. And now all of a sudden, you know, you guys can teach us, but that doesn't mean, as you were talking about, the management layer has, knows what is going on. So it's embracing the management layer, needing to embrace these new capabilities and understand how that works through how an organization, you know, flows in All.
Well, to John would probably add in on this, that that transparency of the information, the manager being able to get the data, that's flow. Because, uh, in software engineering, what we're doing is, is there's information, there's metadata about your daily activities. It's flowing through the system.
There flow information that's available. Let's teach folks how to get after that and how to interpret it and how to understand it. So let's make it that when we're talking, we're focused on something that matters, not on nitty gritty because you didn't know how to go and pull or didn't have that information available to you, which, we'll, probably John, I'll take it back to, we still have a lot of organizations that have poor hygiene.
And I, and I think we've done it. I mean, you know, I mean, I think we've done a great job, you know, to the point, like DevOps, for example, it really is borrowed from Lean from Agile and all that's borrowed from Toyota Production Systems. So, you know, we, you know, uh, there are, there is sort of a virtual and, and and parallel nature of how we deliver software.
But I think if you look at the history of DevOps movement, it has been about trying to emulate or posture industrial economies into knowledge economies. And I think, you know, I think, I mean, there's a lot that still haven't read the books and haven't gotten the memo, but I, I think, um, that we've, we've, I, I think we, that's not a bigger problem is, I mean, the, the idea of sort of walking the floor, you know, Toyota would've called that code gba. So I, but I, I, I don't disagree that we are different in software delivery than creating cars or managing or, you know, running ships.
But, but I think we have done a fairly good job in, in, in, in, in really trying to emulate some of those industrial knowledge. All right, guys, I gotta cut this off here, but I would remind everybody in those books that John's talking about, there's a whole section called empathy. It's kind of how the whole DevOps thing works.
And it's all about trust. So, you know, practice what we preach. We'll be back in a minute.
I'm Bonnie Schneider, sustainability contributor to the Techstrong Group. I'm excited to introduce you to a groundbreaking new initiative from Techstrong Research, the sustainability pulse meter. The pulse meter offers valuable insights into how environmental responsibility factors into tech purchasing decisions for key players in the industry.
Position your company as a leader in the industry and differentiate from your competitors with a sustainability pulse meter offered exclusively from Techstrong research. All right, and we're back, and we're gonna be talking about some research into vulnerabilities as it relates to AI platforms. And specifically there's been some, um, analysis of what's going on on an SAP platform specifically.
But I'm gonna let John, um, walk us through this because he's the one looking at this. And I do mean John Willis here, but John, um, you seem pretty excited about this. There seems to be some progress being made here.
What's your take? Yeah, no, you know, I, I think I've mentioned this before. We're, we're working on a paper, uh, through a Gene Kims organization called Dear CIO, which is sort of, uh, off, you know, sort of modeled after the Dear Auditor letter we wrote years ago.
And, and we wanna sort of make the CIO aware of all these potential unseen dangers in generative ai. 'cause there's a lot of stuff here that is not obvious. There's a, a lot of these C-level people think, oh, I can get rid of all these people.
I don't need that anymore. I, same thing we found in Shadow it. And, and so NIST, I think is not doing a great job.
Uh, you know, uh, no disrespect to anybody who's a big panist. They're really, they, they're treating AI as, as if it was Google, you know, Google search. Like if you read their sort of, now O osp, I think is doing, you know, a much better job.
But they're very general and, and to everybody's defense, we're so early here Mm-Hmm. But these wiz guys is Wiz, I, oh, they're, they're, um, kind of a full service cloud security vendor, right? And, and I've been tracking them for now a couple of months now.
And, um, and they just posted recently, I think it was last week, about, uh, a vulnerability in SAP's, um, ai, um, service. Um, and, uh, but they, they, they posted a couple, they're all the same way. The, the, the same attack surface, which is these, um, AI as a service vendors, you know, hugging, face, replicate now sort of, uh, SAPs, uh, SSAP AI core.
Um, what they find is, and this goes to the dear, say, oh, there's so much unvetted code in this ai, the, the Python libraries, you know, we're used to vetting, you know, Java libraries for years, right? And, and some of this code has really never been run at scale in large enterprises or in the government. And I'd love to hear, soon as I finish, I'd like to hear Tracy's opinion on this as well.
But, um, what, what they're finding is, like all this stuff in like the models, there are a lot of times like, um, so a lot of these models, which, you know, people don't understand, have executable code in them. They'll have, like, for example, in Python, there's what they call pickle libraries. They, they allow you to serialize objects.
Well, what they do is they have to write to the, the file system. So they, they, so what they're able to do is do these lateral attacks, you know, run malicious models or update models. Um, and, uh, and in those malicious models, they get, they, they literally create lateral movement.
And what they found, like in the hugging phase one, they were able to escape out of the running the model, like a training model or running sort of some inference against it. And then they were able to escape into, um, a Kubernetes cluster. And a lot of these AI service people are not, don't have the data like 'cause back to hygiene.
They don't have the security and DevSecOps and DevOp hygiene. So they're, they, they don't have their Kubernetes clusters patched properly. So then they find vulnerabilities.
And the, the hugging face one is the scariest one, because what happened was they were able to show by running a malicious model that, that not only could they escape out of the model escape out of, uh, Kubernetes cluster, they literally got on. Uh, and because how you face is running on Amazon, they actually got on a multi-tenant host, uh, on, and, and they did a similar thing replicate, but the SAP P one is SAP's infrastructure. But what they were able to do there, same model.
It was like escaping out of the, the running the model and then, you know, finding vulnerabilities in a Kubernetes cluster. In this case, they actually created their own pod. And then, you know, Bob's your uncle at that point, because they were able to sort of implement an Argo work file.
They were able to, um, uh, attack, um, Grafana. They were able to get, um, AWS uh, actually, they were using, um, some AWS services. They found credentials.
They were, you know, they literally, the list, if you read the article, the list of compromises that they just, and, and here in this case, you know, hugging face is one thing, except for them having a multi-tenant access on a, on a physical, uh, Amazon host or, or a virtual vm. But, but this is like SAP's customers. So they were, they were getting customers prompts, they were getting customers credentials, they were getting cut, you know, all sorts of, um, you know, like really scary stuff, right?
Um, being able to read docile images, actually read, well, this is the scariest one is, uh, it's called false prediction. Uh, even US doesn't have this as a top 10, LLM, that where people can go in and put false predictions in models. This is like, you know, squatting, domain squatting, right?
Like, if I know that I can put false predictions into a model, I can maybe downstream have people do actions based on the response. It, it's the, the, the, uh, the breadth of the complexity here is, um, is just mind boggling. And I don't know, you know, and these, but, you know, I, I've gotta give these wiz uh, people incredible credit because I think, you know, they're pointing out stuff that should be scaring the heck outta everybody.
Tracy, call us for pause. Oh, gosh. There's, there's all kinds of things to be concerned about here.
I will say that while NIST has not yet stepped up, and my hope is that they will, but cisa, uh, is filling in some good gaps there. So I'm happy to see what's happening there. And at the end of the day, models are code.
And if a software engineer has not created the model, because they're not data scientists, come back to that in a second. If the person that's creating that model doesn't have the necessary understanding of the good hygiene, we're gonna continue to see this, and we're going to continue to see this. So there does have to be some intersection.
We do have to help to educate the data scientists, the data engineers. We have to think about working together, uh, much more closely. Uh, and what John is talking about, you know, these, the types of, of cyber impacts, uh, it's, it goes beyond the models themselves.
The concept of data poisoning has been around for a while. Uh, we're not really bringing it up and talking about it a whole lot. But what John brought up about being able to take somebody and provide them with false output, that's model poisoning.
And if I wanna poison a model, I don't necessarily have to make it something massive that you see right away, right? It can be very tiny, and I can build on that drip, drip, drip over a long time. Uh, so think about the impacts of that.
I can impact the medical field that way. Think about it with software engineering, all I need to, if, if we're training all of these models to help us to write software, and I'm doing it by looking at different open source repositories that are now have been lightly poisoned, small little bits, little drop of cyanide here, just a little bit, not enough to take you out, but enough to be uncomfortable. We're going to continue to, to see that.
Um, I have a lot of hope for what the future holds, but we are right now in the embryonic state where people are, it's, it's, it's the wild west, and people are trying to figure this stuff out. Um, gosh, I could talk about this for hours. I don't want people to suddenly hang their head and go, oh, this is, you know, this is terrible.
This is horrible. We've got a lot of really incredible experts that are focused on model assurance, on AI assurance. How do you make sure that the outcomes are valid?
How do you know and trust the corpus that's going in? There are a lot of really incredible people that are focused on that part of it, not on what the model does itself, not on who's create, but all of that AI assurance. My, I'm betting on those people.
Uh, I'm, I'm absolutely betting on those people. Hey, Kimberly, let me ask you real quick, how's the feeling in the, in the pit of your stomach right now as you listen to all this? You know, what's buzzing through my head is that for very long HPC market Never did anything to Do with, never wanted anything to do with data protection.
Um, and they never did anything. And then when we got into the ai the last two years, there still was nothing going on in there. Now, there were the gen ai, all of a sudden, it's kind of a topic, but as typical, you know, data protection or anything that has to do with data and that kinda stuff, it tends to be the tail of the dog.
And it's the last thing anybody thinks about. So I think that what you're talking about here, data poisoning or AI assurance, which is, I love that term, that's awesome term. Not just governance, but assurance that you know, what you're getting is real has, and I would love to see NIST address it.
If CISA is addressing it, fabulous. Because that's maybe how we're gonna get there is part of this. It's gotta be secure.
So before we even get to the other side of the data, it's gotta be secure. And that's the security guys. Um, and that gets unwieldy because now we're working with a lot of open source technology.
Um, and you're right, SBO m doesn't work. Just SBO m doesn't work. You know, I don't know what I don't know.
And so much of when I read those articles so much, it was going on with Wiz was not just an open source problem, but it was like somebody left a door open, didn't configure it. Right? You know, one of the things I'm thinking about ca as you remind me of Kimberly, is it's coding and it's also data, right?
That's one of the unique things about AI in whatever form of models, whether there's generative or, uh, machine learning. Uh, I'm thinking about, uh, back at the, um, reinforce event. AWS you know, close to the beginning of the year, the CISO got up and talked about the trillion chip and the enclaves that it lives in, and the separation of data that's used by ai, um, applications and software, uh, versus other parts of the system.
And what it, what it reminds me of is, in some way, I'm not saying that AI is a contagion that we have to control or anything like that, but John pointed out so many examples that the folks at Wiz are finding of sort of what are some of the, the blast radius term that you used earlier, what the blast radius can be, um, when something goes awry. And I'm not talking about AI mageddon taking over mankind. I'm just talking about we have a new kind of technology that we're using in a different kind of way.
What new types of data protection in terms of architecture, the application, and how it operates, the systems that it operates in, and do we need, uh, better protections to, you know, think of your Tracy back to our, our security, um, you know, microsegmentation and things like that, that we didn't use to, to limit the blast radius of new technologies. One, One of the points that's clear, and it's the reason why we're writing the dear CIO letter, because what we're seeing is CEOs are hiring chief AI officers. Chief AI officers are going off on their own and building infrastructure.
Mm-Hmm. You know, early on with open ai, one of the biggest bigger breach internal, they, they created it themselves. They were using a bunch of Python and parallel sync libraries with Redis, and they just didn't know there were vulnerabilities, right?
In its code. Like this is what's happening in hugging face. Hugging face is running Kubernetes, but they don't have, they don't have the expertise.
They don't even know what they don't know. So what you're seeing is all these AI experts that are absolute experts like a Chief ai or for a large bank's probably coming from Stanford, but they don't know, they've never been through an audit, they don't know what it takes. And this is the letter to the dear CIO is CIO, don't let them proxy what's gonna come back to you anyway, because it is, it's, Hey folks, it's, it's, uh, compute network and storage, and it's a bunch of middleware technologies.
All these, when you run these infrastructures, you run the Lang chains, you know, all under the covers. We're running stuff that we've been running for infrastructure all along, you know, um, you know, Kafka, Kubernetes, um, you know, things like Redis, MongoDB, right? Like all these tools are basically part, um, you know, uh, elastic and, and like the banks know how to do a really good job because people like Tracy learn from Tracy what to lock down, how to do it.
Unfortunately, these chief a IO officers or these hugging face infrastructure people, they don't know what they don't know. And that's the scariest part. That's the scariest part of the wiz story.
John Schwartz, anybody out in the valley talking about this among the tech bros? Or are they just kind of quietly ignoring all this? Well, um, you know, John brought something up that I found interesting.
There has been a debate out here about the CAIO and the necessity of having one, that the, the feedback I always hear here, here, here is that in a sense it's conflicting with the CTO, the CDO, the CSO, they're, they're kind of overlapping with one another, working against one another and creating situations that probably shouldn't exist. And I was just gonna throw this out maybe to John, to others about is a company, I mean, what type of company would need A-C-A-I-O-W? I mean, there are just a few.
Or would every one of those, well, Not just companies, Yeah. This is this crazy Idea. They're not just companies.
I'll hop in there before John says, and say, think about the government and think about every one of these agencies has a CIO, right? That's correct. Right?
Right. And the head of CDO, what I'm seeing is not instead of A-C-A-I-O, that they're actually bringing that underneath. And so it's the CDAO, so data and analytics, so they're bringing that together.
Um, but it's, I, I agree with everything you've said that we've got this dramatic overlap, all these roles, this is important. We need a c-suite to deal with it. This is important.
We need another c-suite to deal with it. Well, the problem to me also is that the CEO, and it just from all indications, everything I've, I've heard from you all or read or studies, that the CEO really doesn't care. I mean, they're, they're just, they're trying to just push these initiatives, you know, whatever the consequences may be.
I, I'm being very cynical about it, but I, I kind of get that impression and as well, there're creates this lettings ripple of, of chaos. Yeah. I mean, they're, they're in this, I mean, I've heard this quote from CIOs and, and from their CEO, which is, it's do or die.
AI is do or die, and they don't know what this technology is. So what's the sort of obvious thing is go find somebody who is an expert in this technology and help run your banking or financial or financial analyst business unit, which is a great idea. You know, Mark Schwartz wrote a book, um, A seat at the table, right?
He, he, he ran Homeland Secur, uh, ci the, uh, border protection or whatever, the border of, uh, Homeland Security. And he wrote a book in that book he talked about, he even said, when you basically create a parallel CDO to A CIO, you are basically, and I'm paraphrasing, admitting that you failed at, at infrastructure operations and applications. And so my only message is, is don't, I mean, if you wanna hire a Chief AI officer to Tracy's point, I, I sort of disagree.
I think more companies are putting A-C-I-I-O or A-C-D-A-I-O on, on a parallel level with a CIO. And I think that's a huge mistake. And I think that our letter to CIOs you know, we have, you know, part of Gene Kim's organization, which is Starbucks and Disney and a lot of big organizations.
You know, we're gonna see if we can get a letter to sign it and say, that is a mistake, CIO you're going to have to own the responsibility of this ai. Yeah. The, the CIOs that I've talked to, I did a story about this Dow Jones and the, the, the consistent answer I got from several CIOs is, this is just a mistake.
Um, this is just gonna create. Yeah, No, it is. And, but they're in a rock and a hard place, right?
Their CCEO is basically saying, I, you know, I can't wait on this is cloud all over again. That's why I wrote article called Shadow ai, right? It's cloud all over again.
I can't wait on you guys. I have to do this now. Even, you know, the imperatives here are so much more.
Um, and, uh, and so, you know, the, the CIOs are like, okay, uh, yep. And, and some It is just basically like, move as fast as you can. If it breaks, we'll, we'll, we'll worry about that later, but consequences can be, they're gonna be bad.
Guess what It come back to? This is what your point, John. It's gonna come back to the ccio shoulders.
Mm-Hmm. Yeah. Well, it's gonna, it's gonna tarnish the brand and who's gonna get in trouble when the brand gets re reputation problem.
And when you lose $5 billion of market cap in a day like Equifax did, like, who are they gonna call first? They're not gonna call the chief AI officer, they're gonna call cio. No, I, I suggest that we keep the CIO, uh, in their role and that they have their, I call, you know, the Knights of the round table.
It should be, should be instead of their chief officers, it should be their chief engineer. So if you have a Chief AI engineer, let's get them the experts that they need. I agree.
I agree. But there needs to be, um, somebody, it's just like anything else. You ultimately, somebody has to own the responsibility has to own the final decision making, even if it's informed and it should be informed by the experts.
But putting another, you another, how many drivers do you want for this car? Only one. Only one.
Yeah. Alright guys, I, we've seen this. Go ahead, Cameron.
The last word, Just, just last, you know, kind of, we've seen a bit of this before, um, with the Kubernetes crowd, the cube con crowd, it used to be all the dev people there in the last two years, or last year and a half actually, we saw the rise of the platform engineer power, um, because DevOps guys didn't wanna bother with that storage crap and that network crap and that server crap. So I think that trending, but not until we have a couple of, um, visible failure kind of things that happen and, and then people start addressing those requirements for hardening the environment. All right, folks, we gotta close it off here.
Hey, this was a great episode. I enjoyed every one of these conversations. I would just point out there's a fine line between Do or die and Do and die, right?
So think twice. That's a way to end a show. Oh, On that positive note.
Uplifting as always. Yeah. We'll just go do We have more uplift?
We have more uplifting coverage coming up on Text Textron tv. We invite you to stay tuned and check it all out. But as always, it's a joy to be with you guys.
And thank you all for being on the show, and we'll see you next time. com is the leading resource for news analysis and education on challenges facing the cybersecurity industry. com covers all aspects of cybersecurity, including data security, DevSecOps, cloud security, application security, network security, security threats and more.
com has the largest selection of security content featuring breaking news, blog posts, podcasts, and more. com to learn more. com.
Home of security bloggers network.