AI’s Impact on DevOps, Unified Observability, and the Cloudflare Outage | TSG Ep. 972
Mike, Jon, Fred Wilmot, Gina Rosenthal and Barbara Roos dive into the future of DevOps and software engineering in the age of artificial intelligence (AI) before assessing the degree to which observability might be unified following a move by Palo Alto Networks to acquire Chronosphere for $3.35 billion.
Then the gang looks into the impact the Cloudflare outage had this week as disruptions to crucial IT services appear to keep on coming.
Transcript
Hey, everybody. We've seen the future of DevOps in the age of ai. Maybe you're watching Textron Gang.
We'll be back in a minute. Welcome back everybody. We've got our usual assemblage of smart folks on the panels today.
And starting off with Gina Rosenthal, Fred Wilmot, John Schwartz, and we have a new member, Barbara Russ. And I guess I kind of wanna introduce Barbara A. Little bit 'cause everybody else has been on the show multiple times.
But Barbara, tell us a little bit about yourself. Sure. Uh, I'll start with, my name is Barbara Rose.
Uh, a lot of people get that wrong. No worries. Uh, I, I run Trailhead Communications, a consultancy that helps companies navigate the human side of AI adoption.
And my background is in, uh, the tech industry. I've been in communications change management culture work for the last 25 years. So, super excited about this AI revolution.
All right. Speaking of which, Alan and I were at a show up in Brooklyn this week. It was hosted by an outfit called Tesla.
And they were talking about, well, specifications for AI agents. And let me do my best here to kind of explain what's going on. But part of the issue with AI agents is, well, they give you superpowers, but they're also notoriously unreliable.
So people are creating specification files, which are essentially, you know, documents that tell the AI agent very narrowly what it's supposed to be doing. And then ultimately, once you get that AI more focused on a particular task, you can start daisy chaining these things to automate processes. And folks are talking about doing that within the context of DevOps workflows, because we need these things to be more reliable, right?
We can't have a bunch of AI agents running around just randomly generating some code that may be, I don't know, it has a bunch of vulnerabilities in it, or is just frankly, too verbose to run and gets kicked back by the software engineering team. Fred, you've been floating in around on DevOps for a while. Is this the right approach?
I mean, on the one hand, I kinda like the idea. On the other hand, I'm like a little bit concerned about, well, how are we gonna manage all these files? Tesla says they're gonna create a platform for this, but if you've been around Kubernetes, you are familiar with the phrase wall of YAML files.
So are we just gonna get more walls? It's a good question. I, I'm kind of thinking about it like it's the next, uh, it's like the US bump for, for agen, uh, workflows, the philosophy that, you know, we should probably think about having a persistent record of intent.
We kind of have this, so to make a standard for it though is an interesting concept, I think given the number of, of different variants in files. And so when you want to have, you know, a agent communication across using, uh, a to a or, or what have you, uh, MCP servers need to communicate the same, uh, actual effects as, as all of the, uh, agents start to collaborate outside of your sort of wall of trust or your wall of yams, right? As you put it.
The, the philosophy is really about how to understand whether or not that's going to improve things. So on the one hand, I would say, look, w we already have some solutions to this type of a thing, but on the other hand, I think they're, uh, it, it's a funded company and there's a, you know, there's a large amount of funding behind it. So my, my argument against that would be, look, uh, we had an ai, uh, uh, cyber, uh, cyber challenge at, at, uh, DEFCON this last year.
And those winners open sourced all of their frameworks. Those frameworks included the opportunity in a, uh, to, to find, disclose, uh, patch and deploy, uh, vulnerabilities in software, uh, and those types of things. The way that works best is when that's open sourced, how a standardized process works from a, a private company is a question mark for me.
So, but there's a need for it. Uh, is that the greatest need of all? No, I don't think so.
Gina, you, you have some experience in the land of operations among other skills and expertise, but as you kind of look at this, what's, what's your initial reaction? Well, my initial reaction, um, was isn't it just sounds like it's agent driven infrastructure as code. 'cause you're looking to put, and it sounds a lot like what we used to do with finish files for, um, for Jumpstart and Kickstart, right?
So you wanted to do a certain thing. If an agent is just a bundle of, it's just a bot that's assigned a specific task to go do, but you wanna make sure that tasks stay, you wanna be able to give that bot, um, uh, uh, a, a space. We want you to, we are gonna declare what you're gonna go do and all the other bots you've gotta go talk to.
How is this not agent driven infrastructure as code? And then my second thought was just like Fred was saying, we're already doing this as my mantra. This is just an extension of computer science.
We should be getting better not trying to reinvent the wheel. So I don't know how we have, how we get the communities to come together, right? Like, yes, you're thinking along the right track.
You're further along than we were when we had no tools. So yes, how do we accelerate it by showing you how we figured out how to do it already? Mm-hmm.
I think when I looked at it, it seemed to me we were coming up with a way to use code to make up for the limitations of the AI agents. There just got some fundamental problems, and we need to figure out how to manage that. But what they're saying is that we need to share these specification files among developers so that we don't have all create the same ones over and over again.
And then that leads to this wall that I was talking about. Um, Barbara, welcome to the show. But I guess, you know, it's pretty clear that we're gonna have some sort of leadership issue here in terms of how we manage this process because there seems to be a disconnect emerging between the developers that are in love with AI coding tools and the software engineers that are responsible for actually deploying this stuff, and maybe we need some more adult supervision.
What do you think? I think adult supervision is a great idea. Um, and I, I think collaboration is at the heart of this, um, you know, a across all kinds of industries and use cases, I think everyone's experimenting with AI and they're doing it individually and in their own way, and they're finding the things that work for them.
Um, but what needs to happen is we need to create communities of practice, uh, who are coming together and sharing what they're experimenting with, sharing what's working, what's not, creating a culture of experimentation. Um, and then from there, creating guardrails and systems and consistent use of, of these tools. Uh, and I, I think part of why this is happening in isolation is because we have that leadership gap where, um, leaders aren't, aren't talking about what the real vision for AI is and how, um, how the culture needs to shift in this, this new reality.
Mm-hmm. You know, Gina, to Barbara's point, most of the IT leaders that I have met usually have spent some time in the trenches and have some experience in this space. And yet, once they get promoted, something seems to happen.
They seem to get removed, they're divorced from the actual workflow and the things that are being done, and suddenly, you know, they're kind reading the latest report in the Wall Street Journal and making policy decisions and what happens and how do we kind of prevent that from happening? Yeah, that's a very interesting question, right? Because we all know what happens.
You get busy doing what the big bo what you're supposed to be doing, going between the big bosses and the actual technologist and, and making the company a profit and keeping everybody out of trouble. I, I love the term community of practice. I think that's a big part of it because I think the technical leaders who go on to, you know, these, um, more important roles, managing people and managing processes, um, need to be part of that community of practice.
And I think there's a great, just in tech in general, there's a huge, um, avoid of that anymore because everybody is on the hype monster. And so it's very hard to find real information about here's how you go from A to z. I would love to be able to see Tesla tell me, yeah, this is just, um, agent driven infrastructure as code, it's the next generation.
And then all of a sudden maybe you can tie, start teasing those communities together and building a community of practice that gives people on the top level that don't have the time to look into things and maybe get their hands dirty anymore. It gives them something that sounds reasonable versus we've figured out a brand new thing. It's a brand new thing.
So now we're gonna have to do a brand new thing to manage it when we all know if we've got the experience behind us. That's not necessarily true. We need to build on the foundations that we've all climbed through and in, in those trenches.
So I think it's getting that information to, uh, to the leaders who are technical in a way that is technical and it's not hype driven, or it's not, um, analyst defied, you know, that it's actually tied to reality. And that then they can direct, you know, they can encourage their, um, their teams to go join community of practices that are truly technical to dig in into the details. And the technology leaders can, um, find ways to that to direct them that is tied back to the business.
I think we're just missing a, a community of practice to help us get through the hype. You know, that's interesting that you say that, Gina, 'cause I, when you said A to ZI was thinking of how the way this narrative is unfolding, this hype narrative for now. And I think about Silicon Valley, they present the pie in the sky utopian vision in these announcements.
And there are, there are announcements every day, as Mike and I can attest sadly. And then they, they, so they give us, on one hand they tell us, this is what we're gonna announce. It's bigger, better or faster.
It's gonna do everything for you, imaginably automated, and, uh, we'll, we'll give you the end result as well. But there's no transition. There's, there's nothing in the middle on how do we get there.
And I think that keeps occurring. It keeps rearing its ugly head in this AI agent year of 2025, like the devil in the details. And almost every segment we talk about, and most of the segments we talk about involve that gap or some sort of problem, whether it's security observability, um, what the agents do, how they're coordinated, and the impact on the people in the middle.
So I think we're gonna be hearing a lot more from, like, from Barbara about how we're gonna navigate this A to C journey. There's a big gap. There's like a Grand Canyon gap between the two sides.
I think actually AI has a lot to do with it, right? So coming from the product marketing side, if you have given all of your marketing budget to ai, so you lay off all the marketers that had any experience in the industry, and then you give it all to brand new people and say, use AI to help fill in the gaps, it fills in the gaps. But there's no, they, they don't have, the AI doesn't have the, the, uh, the expertise in the domain and neither do the new word marketers.
And so that's part of what we're seeing. We're seeing these great looking articles come out, but there's no depth to it at all. Yeah, there's no good stuff.
I think that's, that's critical. I mean, we're, we're seeing that across All, all kinds of spaces that are adopting ai. You need to pair AI with human expertise and wisdom.
And, um, you know, a AI fundamentally, at least for now, is still derivative. You still need innovation, and that comes from people. And so I think pairing AI with human beings who are kind of giving it the right guidance and instruction, and then can also judge the output of its work and decide is this flawed?
Is this good? Um, you know, continue to guide. I mean, in a way it's like the, the practitioner on the front lines using the AI becomes like a leader or a manager themselves of the ai, and they need to develop a lot of those same leadership skills and coaching and mentoring skills.
Can I ask, can I Ask you, oh, I'm sorry, Barbara, can I ask you Oh, go for it. A quick question. So when you, me, it's interesting.
So the human in the loop equation, or this, this concept are most companies, and I don't wanna put you on the spot, but how many companies out there are doing a good job of, of the integrating the human in ai? Because I'm not hearing a lot of those. Maybe they're in this initial process of doing this.
Uh, I think there's some who are doing a great job. Um, you know, uh, I would point to, like companies I've talked to recently, Cornell is Networks Marsh. Um, I, I see examples where they're, they're really embracing that mindset, um, building their business as, you know, an an AI first kind of company.
But I think most companies are struggling with it for sure. They, they don't know how to lead in this new reality. And, um, you know, they, they're leading with tools, uh, instead of leading with people and mindsets and, and leadership skills.
I think that we're on a spectrum, right? So we started out with these copilots and now we're moving up to smarter AI agents, and hopefully things will get a little bit better. But I have noticed this trend where, you know, execs show up and they're like, oh, this is gonna be great.
We're gonna increase productivity and we're gonna have all these wonderful outcomes. And then when it doesn't happen and the rank and file starts telling 'em, well, this stuff doesn't really work as well as you think, then what happens next? I think at least the good leaders is they roll up their sleeves, they get in there and they start actually working with this stuff, and then they discover what the limitations are, and they're suddenly a lot more cognizant of just what's real and what's not real.
But Barbara, is that kind of the cycle of things? Oh, a hundred percent. I mean, we've seen it, uh, over and over with technology revolutions in the past.
You know, you, you have the hype cycle, and then you have the trough of disillusionment. And I think, you know, we're, we're going into that, um, that trough. And what I think it's gonna take is leaders getting past the hype and the excitement about all the potential, which is totally valid.
Uh, we're all excited about the potential. Um, but I, I think if leaders are living in their imagination of what's possible, instead of rolling up their sleeves and, and getting their hands dirty, like they need to walk the walk and they need to experience for themselves what this technology can do and what it can't. And then they need to be role modeling with their teams, you know, showing them, this is how I use it myself and my work.
This is how it's transforming what I do. And then, you know, their, their teams can then in turn figure out what that means for their work. Mm-hmm.
Fred, coming back to software development, I was at this conference, and a fellow made me laugh. He said, this stuff is great. I'm running into the same walls 10 times faster.
Um, Yeah, I, I, I, I, I couldn't be further from that opinion, to be honest with you, right? My, my, uh, my scale of writing code is five x, you know, and the ability to, uh, multitask while doing that with numbers of agents doing multiple workloads is incredible. Um, not to say that it doesn't require the same level of persistence of understanding and testing and rigor that other things do, but it's in essence, like hyperthreading a person doing the work with agents, doing the work that you have to supervise.
So, I'm with you a little bit on that, Barbara. I think one of the biggest concerns is really about who can do that most effectively. And that's really where you're getting, you know, folks that have been doing this a long time, have a lot of experience, can, can really ize their experience with a number of ways of parallelism there.
Um, I I, there's a lot of speculation about what the outcomes of this is, right? If you take somebody that's relatively okay at writing software, and you multiply that with somebody that's relatively okay at writing software, you're gonna get, you know, some effect, right? Whether it's a ripple effect of massive code with lots of other things to be concerned about, or, you know, you, you have a much improved way to, you know, make a force of 10, fight like a hundred.
Um, the real issue is like some of the standardization. I think if we, you know, if we think about, you know, coming back to Tesla, like if you wanna make something like an industry standard, okay, well, SBOs were created by the, you know, Linux Foundation, right? And they were also, you know, fundamentally supported by the Oasp community.
And so when you want to get industry adoption on things, you typically have to open source it. So, you know, I'm a little bit man on, you know, a a a company sort of building an industry standard, uh, you know, that's funded, uh, because inherently there's, you know, uh, cause for question about whether or not there's, uh, you know, some concern for their, uh, intent. But, you know, ultimately there, there's a set of requirements here.
It's a natural e evolution, like, like Gina said, and I think, uh, this situation is here, right? The question is more like, how do you handle, you know, things like authentication and authorization. How do you manage an identity of 10,000 agents, right?
Uh, a hundred thousand agents, a million agents, uh, as opposed to sort of like, is it going to happen? It's happening, uh, for sure. And I think evidenced by, you know, this example of, Hey, we need these kinds of things to make sure we have a, you know, a persistent record of intent so that, so that when these agents start over again, there's so many agents that we have to have some reasonability that they'll come back to where they're, where they're centered to do the work.
'cause there's just too many to manage, really. So, Fred, to your point, I think I have noticed this trend, and I saw it at the event, but there does seem to be something of an AI divide emerging in the software development community, and there's folks like you that know how to make these AI agents dance, and then there's the mere mortals that are kind of struggling to figure out how to organize and manage all this stuff. So we'll that gap get wider, or can we close it?
That's a good question. Um, I, I'd like to think that that gap will get closed, but I think it'll, it won't get closed by people, right? So the challenge we have here is, and I think we're seeing this with, you know, sort of entry level jobs being, uh, waylaid, um, some of the largest companies we know, massive layoffs for folks.
Also, middle management getting sort of let go in the sense why, because you've got folks that have been doing this job for 15 years and, you know, they can, they can command a fleet of agents doing, you know, relatively good work and, and, uh, staying with the same context, right? So some of the lossiness that happens when you have, you know, humans talking to humans, then talking to more humans, um, much less so when you have a human talking to, you know, a hundred agents, a thousand agents. Uh, I, I'm not saying from a, uh, from a civilization perspective, that's great.
But, you know, from an efficiency and a work stream perspective, I think there's an awful lot of optimization truth in it. And I think, you know, as the scale continues to grow, that's where, you know, people I think are going to dig in to see what is that economy and scale for optimization and efficiency. That's really useful, but you're gonna get these other problems that we've already solved in other ways now with this set of problems and this set of infrastructure.
Just like, just like Gina suggested. Yeah. The thing, uh, the thing I was thinking when you asked that question, Mike, was you're talking about the two types of developers, but you didn't say anything about ops, and this has a huge impact on ops and whether it works at all, you know, with that kind of ops mindset.
And so I think there has to be, to me, it's pretty exciting that you could manage a fleet of bots. And, but, but like Fred said, you have to have all of that kind of data center hygiene with it. And you also have to have the observability and the reportability, especially with ai, um, for what's going on and how people's information is being used.
So, um, I know a lot of people are getting let go. My my gut feeling is it's because the budgets have been deflected away from other things to just invest in ai. Um, and my gut feeling is that we will need just as many people managing.
It's just that the, that the, the output's gonna be much greater. But I think the mistakes are gonna be multiplied as well, and can be more catastrophic than we've ever seen, which since I'm not involved will be pretty exciting to watch too. All right, folks.
Well, I'm gonna leave this conversation here and just note that, you know, the philosophers are right. Once again, the future is here and just unevenly distributed, we'll be back in a minute. You've earned it.
This, the spotlight, the responsibility, the weight of teams, companies, and entire industries fall on your shoulders, lives depend on your decisions, your home life included that work. You are protected physically and digitally. Nothing gets through your team without a fight.
But in a globally connected world, everyone sees you, including those who mean to cause you and your organization harm. And now home your sanctuary attackers see an opportunity. Your digital front door is wide open.
And what compromises your home can breach your boardroom. Because the devil's greatest trick isn't targeting your workplace firewall. It's convincing you that your personal life isn't at risk.
Black cloak, digital executive protection, defending the new attack surface your personal life. Well, it wouldn't be a weak in it without another significant merger and acquisition. And this time, Palo Alto Networks is acquiring a company called chronosphere.
They're a provider of an observability platform that is generally used in IT ops and for application development. But Palo Alto Networks apparently sees an opportunity here to apply this more broadly. 35 billion to prove that point.
John, um, what's your take on this? 35 billion a lot of money these days? Or is it just lots Of bargain?
That's a bargain? Uh, met, met, met is spending $600 billion on their infrastructure, aren't of course, right, they're gonna spend all that money. Um, yes, I, I'll, I I won't be not facetious after that.
Um, but Palo Alto Networks announced its earnings, and as part of the earnings, it announced this acquisition of this observability platform, which you wrote about recently, I think a couple of days ago, Mike. And, which was interesting because earlier this week, Chronosphere previewed these AI capabilities in its observability platform to help identify root causes of issues and provide remediation suggestions, um, among other things. So in a sense, Palo Alto Networks is getting into absorbability.
I ha have a hard time saying that word, by the way. Um, so in a, in a sense they're getting into it. And I think that the idea behind this is this push into the market observability market at time when AI applications are creating a lot of demand for system monitoring and performance management.
Um, and it's, it's, it's trying to, I believe trans transform observability from passive monitoring into autonomous remediation. So that's, that's the own, that's the, the end goal. Um, I think we're gonna see a lot more acquisitions, uh, especially they're gonna, it is gonna pick up, given the kind of the political climate where they're gonna be, we're gonna be rubber stamping acquisitions now.
Plus there's gonna be a lot, there're gonna be a lot of smaller companies that are gonna look for an Nexus strategy because there's still some sort of lingering fear about what's gonna happen in the markets. There's a debate, but I think we're gonna see a lot more of this, and the big companies are gonna get bigger. But I think it was a good strategic move by Palo Alto Networks.
Um, and, um, you, we'll see, they, they're probably not done. They'll probably continue to do these types of deals. That's true.
You know, you made me laugh because, you know, observability is like one of those words, like jocularity, everybody knows it, but nobody wants to say it 'cause they kick on. Yeah. Yes.
Um, but yeah, no, what is you, I was gonna ask you, Mike, though, what did you think? I mean the, uh, the, the interesting timing and of, of this acquisition, I know that it's been in the works for a while, obviously, but the, the timing uses really well for Palo Alto. I think this bodes well for the future because the thing about observability is the, the core idea is we're moving beyond a modern area of predefined s of metrics that we're gonna be able to collect all this telemetry data and analyze it so we can get to the root cause of an issue faster.
Well, that has been driven mainly out of the app dev and DevOps world because they're trying to improve performance and reliability. But these issues apply to security and IT ops and all across the landscape. And as that occurs, I think what we need to see is maybe some unification.
'cause you know, the security people are collecting telemetry data too, and so is it ops and so is the DevOps teams. And now we got more telemetry data that we're collecting that nobody knows what to do with and it costs a fortune. But Fred, is there an opportunity here to unify all this stuff?
Sure. My agents don't know what to do with all that information. Right?
That's the theory. Uh, you know, we, I think there's an opportunity here for, not just for Palo Alto who sees a bunch of writing on the wall, but you know, if they have a massive fleet of agents, right, they need to manage and get the telemetry and the optics around this. So observability is, you know, it's, it's rudimentary for everybody that's not a pure cyber company.
But now we understand, okay, cyber also includes, you know, agent behaviors and all these other things. AI is the du jour. So if you're going to wander in with software and hardware and do these things, man, it'd be awful impressive to have something that allows you to navigate, manage, and disseminate that information.
'cause the key is resiliency and agility, you know, not, uh, not whether or not we can stop cyber attacks, right? That's the business problem is not cyber attacks. The business problem is resiliency and agility.
Mm-hmm. Gina, what's your take on all this? Can we all maybe link arms and have a kumbaya moment, it, ops, DevOps, security people we're all gonna like, talk about the same thing at the same time for the first time ever?
Well, sure. I mean, that's the, the dream of it. And this is a perfect, perfect use case for so-called ai.
It's, it's really machine learning, probably a little deep learning, but it's the perfect use case for it. There's too, too many alerts, there's too much going on. And if you had only known this server was getting a little too hot 30 minutes ago, you could have done some action to make sure nothing went down.
So I, us as an ops person's dream, are you kidding me? You love it, Barbara. It's no secret that there's not a lot of love loss between application developers and security people.
Security people tend to view developers as kind of, well, the root cause of all evil. 'cause they created the software that led to the vulnerability that led to the breach developers. They, the security people are, well, they're just in the way they need, we need to build software faster.
And I got features to do and deadlines to meet. And, well, I can't be bothered with all the security stuff that generates alerts, most of which turn out to be nothing. How do we kind of bridge this?
'cause this is a cultural issue as much as it is a technical issue. Maybe now they can be, uh, united in, uh, casting their blame on AI instead of each other. There you go.
But, um, I, I, I think, you know, one of the big questions here is, you know, as you have, um, agentic remediation happening, um, who audits the decisions that are being made? Um, what if the remediation goes wrong? Uh, what kind of governance do you have in place?
And I think all of these teams are gonna have to work together to define that. Um, otherwise they're still gonna be pointing fingers at each other. All right, Fred, you laughed, but can I take all these people and just maybe throw 'em in a room and lock the door until somebody sees sense or what?
I love it. Uh, the trope is terrific. Uh, I, I don't really see that problem as much as maybe other people do, uh, from that standpoint.
Um, but I agree, uh, Barbara's got a great assessment of both what the risks are, and I think you absolutely should lock people in a room. And maybe it's a little bit of a, you know, two men enter one man leave, you know, from, uh, from Mad Max. But ultimately, uh, I think this will help drive better specifications, better standardization in companies.
And instead of talking past each other about, I think this vulnerability has a high level of probability, uh, and that's very low on my development priority list. It's, this is a real thing. We have a lot more telemetry that helps share that.
And we have common data to look at rather than, you know, my tools say these things and your tools to those things, and our bosses have to argue about prioritization. So I think it's a really, it, it could be a really big step up to diffuse some of the, you know, I'm a people person problem of how do you navigate both the requirements to the operational effects of it. So I think it's, uh, the, the future's bright.
Alright, Barbara, does that work? Throwing people in a room and locking the door? Is this a management technique?
It, it actually kind of does to a certain extent. Um, and a a lot of it is, you know, what are the conversations you have in that room? Um, I, I think fundamentally, no matter what your role is, uh, people can benefit from putting themselves in each other's shoes and understanding where the other person is coming from.
And, um, when you humanize each other and, and understand each other's motivations, that's where you start to find shared value and, and shared understanding of, of the problem and, and get to solutions. So yeah, lock 'em up. Lock 'em up.
Heard it. Wait, that's a political campaign. That's a political comment.
Yeah, you guys, I'm glad though, Barbara, I'm glad you're mentioning the, the, the importance of humans. God, what a concept. I mean, given all that we've heard about Agen AI and how it's gonna eviscerate middle management or replace people, or, you know, take all these jobs and become part of like the customer service or the workforce, I'm, I'm glad that they're, these companies are starting, it's starting to dawn on them that people actually are kind of important in the whole process.
I, I, I think that's gonna be the big differentiator honestly, in, in who wins and comes out ahead in, in this period of change, is the companies that see, um, the importance of, of humans, um, for their future growth and future opportunities, um, and who are investing in upskilling their employees. Um, yeah, it's great when AI frees you up so that you don't have to do this tedious repetitive work anymore. Don't let those people go, keep the expertise and the wisdom and experience they have and grow and build it so they can do higher level functions and continue to be the advantage for you and your future growth as a company.
Mm-hmm. I guess part of my soul here is that maybe I'm just kind of not really using it at, at the level of scale that other folks may be talking about, like with Fred, but I often find that I'm like frustrated because by the time I validate everything the AI generated, I I've done it the right way the first time myself. So, Well, that's, I mean, that's a lot like having an intern, right?
Um, you know, you, you only get in as much as you, or you only get back as much as you put into the intern, but then in time the intern grows in its capabilities, its judgment, um, and it can work more autonomously and, and grow into mature, seasoned, um, expert. And, and I think AI is the same way. You, you know, going back to what Gina said earlier about going from A to Z, we're not gonna get there by going from a straight to Z.
We're gonna go A to B, B2C, C to D, and down the path. And so you have to start small and build that trust in the AI's capabilities that it's gonna do what you wanted it to do, that you trust what it, what it did, the decisions it made. And then when you have that base level of trust, you got to B, then you can start working on C.
But if you try to go straight to Z, you're gonna have a really bad experience. You're gonna give up and walk away and say, this doesn't work. And, you know, and then we failed.
Mm-hmm. I guess maybe, uh, I, I just wanna hire somebody else's AI intern after they trained them and see how that goes. I feel that way about Claude actually.
Yeah. Okay. So Gina, though, let's bring this full circle.
It seems to me that this AI stuff will get better as we expose more telemetry data to it. And we don't have that data as, as widely as we'd like to think. So do we put the cart before the horse and we created all these AI agents and then expose them to, you know, random bits of data, but we really need them to focus on the telemetry data?
Well, I think that's the rub, right? Like, I, I think you can put the agents on whatever system that you have or whatever information you have, but you, it has to be focused on the right things for your business and for what, for the job that you're trying to do. There's so much promise to ai, let's get rid of the hype.
But if you have a, a business need and you have that much data, uh, especially log data and, and real time monitoring data that's just perfect to, to set agents on and, and help resolve a lot of issues before they start. So, you know, you don't get caught off guard by something that you didn't even see coming. So, um, I, I think the agents are, it's, I think it's great.
Everybody are starting to play with them. I hate that we call them agents because I still think they're bots. I don't see the difference.
Someone can change my mind about that terminology, but, you know, why not let let the computers run with the things they know how to do with supervision? I think that's a, a great use of it. So, friend, coming back for a second to what we were talking about last segment there, is there an opportunity to kind of maybe democratize observability?
And I'm asking the question because I, I've explained this to more people, and I care to admit, and almost universally I get the same response, which is, yeah, wow, that sounds great. Followed by three seconds of pause and then it goes, yeah. Well, but I have no idea what questions to ask in the first place.
Yeah, I think there's gonna have to be some consensus driven behavior around it, some standardization around it. Um, especially as more and more, uh, agent interactions happen across, you know, the vast quantities of MCP servers that every company is standing up to communicate with their, what used to be APIs. It was now, as, you know, a natural language processing exercise.
Uh, the, the challenge is the same. So, you know, today we would say, and we'll we'll get to this in a minute, but let's imagine that we have several companies that are critical partners for us. And, you know, something happens with those critical partners.
You know, how do we understand, uh, how do we inform and how do we account for that from a resiliency perspective as a, you know, partner said company. Historically, we haven't really had a way to do that, but this offers a number of ways for us to think about, you know, if we have agents that are monitoring things about telemetry, that historically we would have, you know, maybe an model doing this, maybe we're doing statistics and aggregations and all these other things which are trivial tasks, uh, for, you know, a set of ag agentic flows. And so if you have a fleet of folks, uh, agents in this particular case, or, uh, there's a little difference, I think maybe in, in, in agents and bots, you know, we, we should certainly have a rock paper scissors, uh, engagement on that one.
Um, but the benefit you get is, you know, I can task a fleet of these guys to go do this specific thing, which is observe, you know, my, you know, organization's telemetry, compare it to others in this sense and get a, you know, a good bellwether as to whether or not something is happening appropriately and make a decision about something like, do we need to find another more resilient route for a thing, uh, in the future? And I think it'll also help hold companies more accountable, right, to their actual metrics that they say they uphold. Mm-hmm.
All right. Well, John, last question on this one though, but, uh, you're out in the valley. Is this the beginning of, you know, mergers and acquisitions across Yeah, I think it is space.
Yeah. I, I really, I really do. I mean, and I, and I, it's not, it's not related, but I mean, I think what opens really opened the floodgates was this decision in the meta case involving the FTC, trying to fight the WhatsApp and Instagram acquisitions that it had approved a decade ago.
Um, that's another story in itself. But yes, I think it's gonna be an acceleration. There's a lot of money, and I think there are a lot of companies that are really nervous about what we're, where we're headed, um, with all this debate between bubble and, and boom, I think there will be some companies that cash out and, uh, the large companies are gonna pick 'em off.
All right, here we go folks. It's gonna be cleanup in the observability aisle, rub it back, Discover Techron Group, the epicenter of tech innovation. We are your go-to for reaching IT leaders and practitioners worldwide.
Our secret impactful content that sparks awareness, engagement, and top quality leads with us. You'll access editorial websites, streaming videos, virtual events, custom content analyst research, and more. Join our satisfied clients.
Let's revolutionize your tech journey. Contact us today and tell your story to the world in the most powerful way with Textron Group. Hey folks, we're back on this Friday with our last topic, which is this CloudFlare outage, and it occurred earlier this week, and I think it only lasted maybe three or four hours, but then it cascaded for a while for people to recover.
And it's very similar to what we saw with the AWS outage and a Microsoft outage. And we seem to have these large scale outages these days. Gina, is this just like the new cost of doing business and it is the way it is?
Or is there something to be done about this? And are we too dependent upon a couple of things out there that have so many dependencies that they can take down? Well, everybody, Well, that's a couple of questions, right?
So, yeah, yeah. To me, I think this is the first thing me and my friends talked about was like, that's a lot of single point of failure that you may not even know is your single point of failure. So what happened was with CloudFare Cloud, uh, can't even talk today with CloudFlare, uh, was their CDN, their content development network went down and it was internal, it was a bug.
And I wanna kind of read from their outage report. It was a change to a database system's permission. It caused the database to output multiple entries into a feature file that's used by their bot management system.
And then the feature file doubled in size and that propagated to all the machines in their network. And that, uh, ba basically was what caused the problems. So the, the feature that that file helped the bot management system keep up to date with all the threats on the content management system.
So it's kind of like a, a, a, a story of warning based on everything that we've talked about today, which is kind of interesting, right? So we have a bot management system responsible for taking care of, uh, however it worked, taking care of any kind of security threats. Uh, a mistake was made by somebody or maybe by a bot, I don't know, in the permissions that were assigned to the server.
And it kind of cascaded this cascaded failure. They definitely needed some observability so they could catch this as it was happening so they could get rid of it. There's definitely some questions about, okay, um, was it a bot?
Was it a human? Um, was it what, you know, like what does this spot management system, you know, how, why does it need this file? Like, all of the questions that come up into my mind was, how did it actually work?
Um, so this is gonna happen, I think as we're getting used to automatically, uh, or having, having agents not saying that this is what happened, but having agents go out and, and make changes. It should have been as simple change as what it sounds like and it wasn't, and something ha it was just a bug in the system, which that also happens. Um, which caused this effect to have their clients go down.
So like for me, the thing it affected like x it affected open ai, which now Im impacts everybody's working life because everyone's using it. For me, it impacted my radio stations that I listened to online. I was annoyed and I had to listen to YouTube 'cause I was too lazy to get up and set up my record player.
So, but I have to, you know, I need the music to go on. So it kind of was like a disruption in my work life. Um, so there's so much we depend on that we have no idea that it's depending on a content management system that is, um, being protected by CloudFare flare and could go down because CloudFlare had an issue with data changing database permissions that kept caused a problem.
Um, so, so yeah, there's a couple of things. There's number one, providing the, my providers of my radio station, the providers of X, everything, everybody else, they're depending on CloudFlare and nobody else. So can they switch over to that other provider if something goes on?
Or is there even another provider that says good as CloudFlare? And then, um, just the consumers, you're not knowing what goes on and, you know, we're technical, we can figure it out, but other people trying to get on X or just trying to run their, whatever they do for work with OpenAI just kind of stuck like it's not working. Um, and then calling all of us.
So, So full disclosure, full disclosure, tech strong is also a customer of CloudFlare. We experienced some of that outages ourselves, but, um, here's what I'm trying to get at here. Fred, let's kick this to you.
Is it the fault of somebody who created the config file that pushed this button? Who probably feels awful right now? Or is it just that the system itself is flawed in a way that's create, is gonna create a problem and it could have been anybody in any time and maybe, you know, we shouldn't be beating up on the four little engineer in the config file when the architecture may be the issue?
Uh, so I, I think, so it's a different scale of problems when you look at 20% of the internet in general, right? And this system was built for hyperscale DDoS attacks, right? It's a handful of, of services that sort of take all this information from click house cluster that says, look, what are the latest and greatest?
So it does this every five minutes, right? And so that same way you think about routes propagating with a network device like a router, uh, or a switch, and, and the same sort of thing happens and the construct isn't, you know, whether or not as we move into this, this was a, this was an ML process that generates this thing, and it's a bunch of features for a model to make decisions, uh, that, you know, then get informed by this, right? And the question is really about there's a file limit size, right?
Ultimately like a very human problem, uh, that wasn't dynamically adjusted. Um, okay. Or database permissions that might have changed during the course of, uh, this process.
But, you know, I think to your question, Mike, it's, um, you know, these things, maybe this is the worst outage since 2019 for CloudFlare. We've seen a number of these different things. We're going to see some of these types of interruptions, but they're all, you know, methodologies that are, I would say, much further, uh, and more impactfully designed for resiliency than what most people deal with in their companies.
They're gonna run into these types of things from time to time. I think that, you know, the questions would be like, what's the bellwether for whether or not your file replication looks like, you know, there's a lot of after action, I'm sure these guys are working through and like, yep, we're gonna automate that thing. We're gonna put this back in the process.
I need observability on this particular element here. Right? All of that that'll all get after action, I'm sure super hardcore.
And the question is just sort of when you turn the keys over, would that have made a difference if a human did it right? Is probably the, you knows, prolific argument, uh, versus, you know, whether or not, uh, a machine learning algorithm did it, or, uh, an agent did it. And at that scale, uh, when we think about what the problem is, maybe with hyperscale DDoS, like a human's not gonna solve that problem anyway.
It doesn't matter, uh, at that economy of scale and that magnitude, right? That's gotta be, uh, an automated workflow. And that process has to be, you know, pretty instantaneous five minutes, you know, is a, is a substantive time period in that type of, uh, in that type of threat landscape.
And I think, um, you know, I don't wanna let, uh, kler off the hook, but I mean, it's just a different level of problem, uh, as broadly as it affected everything. Uh, yesterday, us everybody, uh, from interacting with customers or their out their outputs as well. Uh, same challenges, Fred, how do you, they, they, they were quick to, uh, to specify that this was not, uh, outside threat.
But that seems like a pretty obvious place to introduce a threat when you, you know, hin hindsight quarterback, Monday morning quarterback kind of thing. How can they be sure it wasn't orchestrated from outside? That's a great question.
I think that's why the, you know, first, uh, you know, the CEO would say that the first thing they did was evaluate whether or not that was an outside attack. Uh, 'cause that's a presumption, right? Murder attack all the time.
The, uh, the ASU botnet, which is, uh, sort of like they're, you know, contending with this thing right now, which is probably their first, the first jump was this might be actually that, that problem. And so they work backwards, I think, uh, to get there. And, and at that economy of scale, uh, it makes reasonable sense because that's like a persistent threat for them.
So I'm with you. I, I think the challenge is, you know, once you get into the diagnostics of that, uh, the time that it takes to make an observability decision, and then again, internet scale, the time that it takes to unwind, that takes time. And so rate of propagation across the globe, you know, and all those things, what was the fix revert to a file, you know, that worked well, right?
Like, like we know so well, that's, that's always the answer, revert whatever that change was, like, put it back, fix it. But I wanna get To two points here with Barbara though. So one is I think we can give the cloud player people props because they own this pretty quickly and they got up on social media and they basically said, you know, are bad and, you know, we apologize.
And, and so that is a good thing on one hand, correct, Barbara. I mean, that's the way to kind of handle these things. HA hundred percent, uh, owning it, acknowledging it, being transparent about, uh, what's happening is, is critical.
Okay? Second part of that question is that may be cold comfort to the IT people that contracted them in the first place. 'cause I'm sure they're getting a call from their boss going, how come the website's down?
How much revenue are we losing? And then the third question is invariably, well, who picked CloudFlare? So I, I mean, I, I think the scale, um, uh, at which, uh, these systems are operating, um, it, you know, going back to what Fred was saying about whether or not it's, uh, the fault of a human or a bot, I, I honestly don't think it matters.
Um, it, in this day and age, it's how do you, what do you do when you have an issue that comes up, you know, have, how have you prepared your people and your systems to navigate that issue and respond as quickly as possible? Um, uh, figure out the right solution if that's reverting back to the last version, whatever. Um, uh, it's, you know, what's your, what's your fail safe plan?
Um, what's your, um, you know, how are you preparing your people and your systems to respond to these outages? Um, that's the, the focus now as much as uptime is, John, you and I are probably one of the few people out there outside of Wall Street that actually read, you know, 10 Ks and SEC statements and are they now gonna include things like we're over? Oh, they do.
Yeah, they do. All Providers. Yeah.
So in every, in the, all these documents, they have the risk assessment. So they, they point out their outages, or this is actually, that's this's a really good way to find stories, by the way. So you look for, you look for you, you do a search of under risks, and they do, they, they mention everything that went wrong, but they bury it deep within the, the document.
You know, I actually think, and I'll, and as an Xfinity customer, I'm used to outages. And I'm wondering if, given what's happened with AWS and what happened with CloudFlare, and I think CloudFlare correct me if I'm wrong, didn't, wasn't there another incident several months ago? Um, I think, but, and regardless, it's something we're gonna become accustomed to.
I, unfortunately, I think it's part of this whole kind of dynamic that we're living under and living with. Um, so I think we're gonna see more of these outages, um, as these companies make their transit transitions and a lot of 'em are making major transitions, trans transformations. Um, I think this goes with the territory.
Alright, Fred, last question on this whole thing. Is this an argument for chaos engineering? Because theoretically you should just be ripping things out just to see what breaks anyway.
Oh, man. Uh, On the spot, Fred. I, I think these guys regularly practice chaos engineering.
Uh, I, I think there's always an n plus one system that doesn't have the rigor and resiliency you expect when some magnitude occurrence happens you didn't plan on. And yeah, uh, regularly implanting chaos telemetry data in your regular operating procedures, uh, is a great thing to do. It also creates change.
It, it also creates a set of, uh, unknown variables. Uh, and in certain systems you'd absolutely wanna reduce all of those things so it doesn't have a place everywhere. But, uh, I'm sure that there's a handful of folks sitting in a room right now gaming out every other possibility for this system and the 10 that are adjacent to it, that it impacts upstream and downstream.
Um, and so, you know, what'll be great is the outshot, right? For folks that have similar sort of telemetry requirements down the road, whether it's a CloudFlare or an AWS or even, you know, your Netflix, right? From that perspective, right?
To get back to your chaos, the theory. So, All right folks, well, I think what we've established here is what I'm gonna call the new Monty Python School of IT Management. Expect the unexpected.
Hey, thanks everybody for being on the show and sharing your thoughts and your insights, and please stay tuned for the rest of the text on TV lineup. It's gonna be awesome. And we'll see you all again early next week.