Techstrong TV November 21, 2025
Watch our live stream Monday through Friday, featuring exclusive news, announcements and conversations with IT leaders and experts on topics ranging from digital transformation to #DevOps, #Cybersecurity, #CloudNative, #Containers and deep-dives into specific technologies and best practices. http://techstrong.tv/
Transcript
Hey, everybody. We've seen the future of DevOps in the age of ai. Maybe you're watching Textron Gang.
We'll be back in a minute. Welcome back everybody. We've got our usual assemblage of smart folks on the panels today.
And starting off with Gina Rosenthal, Fred Wilmot, John Schwartz, and we have a new member, Barbara Russ. And I guess I kind of wanna introduce Barbara A. Little bit 'cause everybody else has been on the show multiple times.
But Barbara, tell us a little bit about yourself. Sure. Uh, I'll start with, my name is Barbara Rose.
Uh, a lot of people get that wrong. No worries. Uh, I, I run Trailhead Communications, a consultancy that helps companies navigate the human side of AI adoption.
And my background is in, uh, the tech industry. I've been in communications change management culture work for the last 25 years. So, super excited about this AI revolution.
All right. Speaking of which, Alan and I were at a show up in Brooklyn this week. It was hosted by an outfit called Tesla.
And they were talking about, well, specifications for AI agents. And let me do my best here to kind of explain what's going on. But part of the issue with AI agents is, well, they give you superpowers, but they're also notoriously unreliable.
So people are creating specification files, which are essentially, you know, documents that tell the AI agent very narrowly what it's supposed to be doing. And then ultimately, once you get that AI more focused on a particular task, you can start daisy chaining these things to automate processes. And folks are talking about doing that within the context of DevOps workflows, because we need these things to be more reliable, right?
We can't have a bunch of AI agents running around just randomly generating some code that may be, I don't know, it has a bunch of vulnerabilities in it, or is just frankly, too verbose to run and gets kicked back by the software engineering team. Fred, you've been floating in around on DevOps for a while. Is this the right approach?
I mean, on the one hand, I kinda like the idea. On the other hand, I'm like a little bit concerned about, well, how are we gonna manage all these files? Tesla says they're gonna create a platform for this, but if you've been around Kubernetes, you are familiar with the phrase wall of YAML files.
So are we just gonna get more walls? It's a good question. I, I'm kind of thinking about it like it's the next, uh, it's like the US bump for, for argentic, uh, workflows.
The philosophy that, you know, we should probably think about having a persistent record of intent. We kind of have this, uh, to make a standard for it though, is an interesting concept, I think given the number of, of different variants in files. And so when you want to have, you know, a agenta communication across using, uh, A to a or, or what have you, uh, MCP servers need to communicate the same, uh, actual effects as, as all of the, uh, agents start to collaborate outside of your sort of wall of trust or your wall of ya mls, right?
As you put it. The, the philosophy is really about how to understand whether or not that's going to improve things. So on the one hand, I would say, look, w we already have some solutions to this type of a thing, but on the other hand, I think they're, uh, it, it's a funded company and there's a, you know, there's a large amount of funding behind it.
So my, my argument against that would be, look, uh, we had an ai, uh, uh, cyber, uh, cyber challenge at, at, uh, DEFCON this last year. And those winners open sourced all of their frameworks. Those frameworks included the opportunity in a, uh, to, to find, disclose, uh, patch and deploy, uh, vulnerabilities in software, uh, and those types of things.
The way that works best is when that's open sourced, how a standardized process works from a, a private company is a question mark for me. So, but there's a need for it. Uh, is that the greatest need of all?
No, I don't think so. Gina, you, you have some experience in the land of operations among other skills and expertise, but as you kind of look at this, what's, what's your initial reaction? Well, my initial reaction, um, was isn't it just sounds like it's agent driven infrastructure as code.
'cause you're looking to put, and it sounds a lot like what we used to do with finish files for, um, for Jumpstart and Kickstart, right? So you wanted to do a certain thing. If an agent is just a bundle of, it's just a bot that's assigned a specific task to go do, but you wanna make sure that tasks stay, you wanna be able to give that bot, um, uh, uh, a, a a space.
We want you to, we are gonna declare what you're gonna go do and all the other bots you've gotta go talk to. How is this not agent driven infrastructure is code. And then my second thought was just like Fred was saying, we're already doing this as my mantra.
This is just an extension of computer science. We should be getting better not trying to reinvent the wheel. So I don't know how we have, how we get the communities to come together, right?
Like, yes, you're thinking along the right track. You're further along than we were when we had no tools. So yes, how do we accelerate it by showing you how we figured out how to do it already?
Mm-hmm. I think when I looked at it, it seemed to me we were coming up with a way to use code to make up for the limitations of the AI agents. There just got some fundamental problems, and we need to figure out how to manage that.
But what they're saying is that we need to share these specification files among developers so that we don't have all create the same ones over and over again. And then that leads to this wall that I was talking about. Um, Barbara, welcome to the show.
But I guess, you know, it's pretty clear that we're gonna have some sort of leadership issue here in terms of how we manage this process because there seems to be a disconnect emerging between the developers that are in love with AI coding tools and the software engineers that are responsible for actually deploying this stuff, and maybe we need some more adult supervision. What do you think? I think adult supervision is a great idea.
Um, and I, I think collaboration is at the heart of this, um, you know, a across all kinds of industries and use cases, I think everyone's experimenting with AI and they're doing it individually and in their own way, and they're finding the things that work for them. Um, but what needs to happen is we need to create communities of practice, uh, who are coming together and sharing what they're experimenting with, sharing what's working, what's not, creating a culture of experimentation. Um, and then from there, creating guardrails and systems and consistent use of, of these tools.
Uh, and I, I think part of why this is happening in isolation is because we have that leadership gap where, um, leaders aren't, aren't talking about what the real vision for AI is and how, um, how the culture needs to shift in this, this new reality. Mm-hmm. You know, Gina, to Barbara's point, most of the IT leaders that I have met usually have spent some time in the trenches and have some experience in this space.
And yet, once they get promoted, something seems to happen. They seem to get removed, they're divorced from the actual workflow and the things that are being done. And suddenly, you know, they're kind of reading the latest report in the Wall Street Journal and making policy decisions and what happens and how do we kind of prevent that from happening?
Yeah, that's a very interesting question, right? Because we all know what happens. You get busy doing what the big bo what you're supposed to be doing, going between the big bosses and the actual technologist and, and making the company a profit and keeping everybody out of trouble.
I, I love the term community of practice. I think that's a big part of it because I think the technical leaders who go on to, you know, these, um, more important roles, managing people and managing processes, um, need to be part of that community of practice. And I think there's a great, just in tech in general, there's a huge, um, avoid of that anymore because everybody is on the hype monster.
And so it's very hard to find real information about here's how you go from A to z. I would love to be able to see Tesla tell me, yeah, this is just, um, agent driven infrastructure as code, it's the next generation. And then all of a sudden maybe you can tie, start teasing those communities together and building a community of practice that gives people on the top level that don't have the time to look into things and maybe get their hands dirty anymore.
It gives them something that sounds reasonable versus we've figured out a brand new thing. It's a brand new thing. So now we're gonna have to do a brand new thing to manage it when we all know if we've got the experience behind us.
That's not necessarily true. We need to build on the foundations that we've all climbed through and in, in those trenches. So I think it's getting that information to, uh, to the leaders who are technical in a way that is technical and it's not hype driven, or it's not, um, analyst defied, you know, that it's actually tied to reality.
And that then they can direct, you know, they can encourage their, um, their teams to go join community of practices that are truly technical to dig in into the details. And the technology leaders can, um, find ways to that to direct them that is tied back to the business. I think we're just missing a, a community of practice to help us get through the hype.
You know, that's interesting that you say that, Gina, 'cause I, when you said A to ZI was thinking of how the way this narrative is unfolding, this hype narrative for now. And I think about Silicon Valley, they present the pie in the sky utopian vision in these announcements. And there are, there are announcements every day, as Mike and I can attest sadly.
And then they, they, so they give us, on one hand they tell us, this is what we're gonna announce. It's bigger, better, faster, it's gonna do everything for you imaginably automated, and, uh, we'll, we'll give you the end result as well. But there's no transition.
There's, there's nothing in the middle on how do we get there. And I think that keeps occurring. It keeps rearing its ugly head in this AI agent year of 2025, like the devil in the details.
And almost every segment we talk about, and most of the segments we talk about involve that gap or some sort of problem, whether it's security observability, um, what the agents do, how they're coordinated, and the impact on the people in the middle. So I think we're gonna be hearing a lot more from, like, from Barbara about how we're gonna navigate this A to C journey. There's a big gap.
There's like a Grand Canyon gap between the two sides. I think actually AI has a lot to do with it, right? So coming from the product marketing side, if you have given all of your marketing budget to ai, so you lay off all the marketers that had any experience in the industry, and then you give it all to brand new people and say, use AI to help fill in the gaps.
It fills in the gaps. But there's no, they, they don't have, the AI doesn't have the, the, uh, the expertise in the domain and neither do the new word marketers. And so that's part of what we're seeing.
We're seeing these great looking articles come out, but there's no depth to it at all. There's no good stuff. I think that's, that's critical.
I mean, we're, we're seeing that across all, all kinds of spaces that are adopting ai. You need to pair AI with human expertise and wisdom. And, um, you know, a AI fundamentally, at least for now, is still derivative.
You still need innovation, and that comes from people. And so I think pairing AI with human beings who are kind of giving it the right guidance and instruction, and then can also judge the output of its work and decide is this flawed? Is this good?
Um, you know, continue to guide. I mean, in a way it's like the, the practitioner on the front lines using the AI becomes like a leader or a manager themselves of the ai, and they need to develop a lot of those same leadership skills and coaching and mentoring skills. Can I ask, can I Ask you, oh, I'm sorry, Barbara, can I ask you a quick question?
So when you meant, it's interesting. So the human in the loop equation or this, this concept are most companies, and I don't wanna put you on spot, but how many companies out there are doing a good job of, of that integrating the human in ai? Because I'm not hearing a lot of those.
Maybe they're in this initial process of doing this. Uh, I think there's some who are doing a great job. Um, you know, I I would point to, like companies I've talked to recently, Cornell is Networks Marsh.
Um, I, I see examples where they're, they're really embracing that mindset, um, building their business as, you know, an an AI first kind of company. But I think most companies are struggling with it for sure. They, they don't know how to lead in this new reality.
And, um, you know, they, they're leading with tools, uh, instead of leading with people and mindsets and, and leadership skills. I think that we're on a spectrum, right? So we started out with these co-pilots and now we're moving up to smarter AI agents, and hopefully things will get a little bit better.
But I have noticed this trend where, you know, execs show up and they're like, oh, this is gonna be great. We're gonna increase productivity and we're gonna have all these wonderful outcomes. And then when it doesn't happen and the rank and file starts telling 'em, well, this stuff doesn't really work as well as you think, then what happens next?
I think at least the good leaders is they roll up their sleeves, they get in there and they start actually working with this stuff, and then they discover what the limitations are, and they're suddenly a lot more cognizant of just what's real and what's not real. But Barbara, is that kind of the cycle of things? Oh, a hundred percent.
I mean, we've seen it, uh, over and over with technology revolutions in the past. You know, you, you have the hype cycle, and then you have the trough of disillusionment. And I think, you know, we're, we're going into that, um, that trough.
And what I think it's gonna take is leaders getting past the hype and the excitement about all the potential, which is totally valid. Uh, we're all excited about the potential. Um, but I, I think if leaders are living in their imagination of what's possible, instead of rolling up their sleeves and, and getting their hands dirty, like they need to walk the walk and they need to experience for themselves what this technology can do and what it can't.
And then they need to be role modeling with their teams, you know, showing them, this is how I use it myself and my work. This is how it's transforming what I do. And then, you know, their, their teams can then in turn figure out what that means for their work.
Mm-hmm. Fred, coming back to software development, I was at this conference, and a fellow made me laugh. He said, this stuff is great.
I'm running into the same walls 10 times faster. Um, Yeah, I, I, I, I, I couldn't be further from that opinion, to be honest with you, right? My, my, uh, my scale of writing code is five x, you know, and the ability to, uh, multitask while doing that with numbers of agents doing multiple workloads is incredible.
Um, not to say that it doesn't require the same level of persistence of understanding and testing and rigor that other things do, but it's in essence, like hyperthreading a person doing the work with agents, doing the work that you have to supervise. So, I'm with you a little bit on that, Barbara. I think one of the biggest concerns is really about who can do that most effectively.
And that's really where you're getting, you know, folks that have been doing this a long time, have a lot of experience, can, can really ize their experience with a number of ways of parallelism there. Um, I I, there's a lot of speculation about what the outcomes of this is, right? If you take somebody that's relatively okay at writing software, and you multiply that with somebody that's relatively okay at writing software, you're gonna get, you know, some effect, right?
Whether it's a ripple effect of massive code with lots of other things to be concerned about, or, you know, you, you have a much improved way to, you know, make a force of 10, fight like a hundred. Um, the real issue is like some of the standardization. I think if we, you know, if we think about, you know, coming back to Tesla, like if you wanna make something like an industry standard, okay, well, SBOs were created by the, you know, Linux Foundation, right?
And they were also, you know, fundamentally supported by the Oasp community. And so when you want to get industry adoption on things, you typically have to open source it. So, you know, I'm a little bit now on, you know, a a a company sort of building an industry standard, uh, you know, that's funded, uh, because inherently there's, you know, uh, cause for question about whether or not there's, uh, you know, some concern for their, uh, intent.
But, you know, ultimately there, there's a set of requirements here. It's a natural e evolution, like, like Gina said, and I think, uh, this situation is here, right? The question is more like, how do you handle, you know, things like authentication and authorization.
How do you manage an identity of 10,000 agents, right? Uh, a hundred thousand agents, a million agents, uh, as opposed to sort of like, is it going to happen? It's happening, uh, for sure.
And I think evidenced by, you know, this example of, Hey, we need these kinds of things to make sure we have a, you know, a persistent record of intent so that, so that when these agents start over again, there's so many agents that we have to have some reasonability that they'll come back to where they're, where they're centered to do the work. 'cause there's just too many to manage, really. So, Fred, to your point, I think I have noticed this trend, and I saw it at the event, but there does seem to be something of an AI divide emerging in the software development community, and there's folks like you that know how to make these AI agents dance, and then there's the mere mortals that are kind of struggling to figure out how to organize and manage all this stuff.
So we'll that gap get wider, or can we close it? That's a good question. Um, I, I'd like to think that that gap will get closed, but I think it'll, it won't get closed by people, right?
So the challenge we have here is, and I think we're seeing this with, you know, sort of entry level jobs being, uh, waylaid, um, some of the largest companies we know, massive layoffs for folks. Also, middle management getting sort of let go in the sense why, because you've got folks who have been doing this job for 15 years and, you know, they can, they can command a fleet of agents doing, you know, relatively good work and, and, uh, staying with the same context, right? So some of the lossiness that happens when you have, you know, humans talking to humans, then talking to more humans, um, much less so when you have a human talking to, you know, a hundred agents, a thousand agents.
Uh, I, I'm not saying from a, uh, from a civilization perspective, that's great. But, you know, from an efficiency and a work stream perspective, I think there's an awful lot of optimization truth in it. And I think, you know, as the scale continues to grow, that's where, you know, people I think are going to dig in to see what is that economy and scale for optimization and efficiency.
That's really useful, but you're gonna get these other problems that we've already solved in other ways now with this set of problems and this set of infrastructure. Just like, just like Gina suggested. Yeah.
The thing, uh, the thing I was thinking when you asked that question, Mike, was you were talking about the two types of developers, but you didn't say anything about ops. And this has a huge impact on ops and whether it works at all, you know, with that kind of ops mindset. And so I think there has to be, to me, it's pretty exciting that you could manage a fleet of bots.
And, but, but like Fred said, you have to have all of that kind of data center hygiene with it. And you also have to have the observability and the reportability, especially with ai, um, for what's going on and how people's information is being used. So, um, I know a lot of people are getting let go.
My my gut feeling is it's because the budgets have been deflected away from other things to just invest in ai. Um, and my gut feeling is that we will need just as many people managing. It's just that the, that the, the output's gonna be much greater.
But I think the mistakes are gonna be multiplied as well, and can be more catastrophic than we've ever seen, which since I'm not involved will be pretty exciting to watch too. All right, folks. Well, I'm gonna leave this conversation here and just note that, you know, the philosophers are right.
Once again, the future is here and just unevenly distributed, we'll be back in a minute. You've earned it. The spotlight, the responsibility, the weight of teams, companies, and entire industries fall on your shoulders, lives depend on your decisions, your home life included that work.
You are protected physically and digitally. Nothing gets through your team without a fight. But in a globally connected world, everyone sees you, including those who mean to cause you and your organization harm.
And now home your sanctuary attackers see an opportunity. Your digital front door is wide open. And what compromises your home can breach your boardroom.
Because the devil's greatest trick isn't targeting your workplace firewall. It's convincing you that your personal life isn't at risk. Black cloak, digital executive protection, defending the new attack surface your personal life.
Well, it wouldn't be a weak in it without another significant merger and acquisition. And this time, Palo Alto Networks is acquiring a company called chronosphere. They're a provider of an observability platform that is generally used in IT ops and for application development.
But Palo Alto Networks apparently sees an opportunity here to apply this more broadly. 35 billion to prove that point. John, um, what's your take on this?
35 billion a lot of money these days? Or is it just lots Of bargain? That's a bargain.
Uh, met, met is Met is spending $600 billion on their infrastructure, aren't of course, right, they're gonna spend all that money. Um, yes, I, I'll, I I won't be not facetious after that. Um, but Palo Alto Networks announced its earnings, and as part of the earnings, it announced this acquisition of this observability platform, which you wrote about recently, I think a couple of days ago, Mike.
And, which was interesting because earlier this week, Chronosphere previewed these AI capabilities in its observability platform to help identify root causes of issues and provide remediation suggestions, um, among other things. So in a sense, Palo Alto Networks is getting into absorbability. I ha have a hard time saying that word, by the way.
Um, so in a, in a sense they're getting into it. And I think that the idea behind this is this push into the market observability market at time when AI applications are creating a lot of demand for system monitoring and performance management. Um, and it's, it's, it's trying to, I believe trans transform observability from passive monitoring into autonomous remediation.
So that's, that's the own, that's the, the end goal. Um, I think we're gonna see a lot more acquisitions, uh, especially they're gonna, it is gonna pick up, given the kind of the political climate where they're gonna be, we're gonna be rubber stamping acquisitions now. Plus there's gonna be a lot, there're gonna be a lot of smaller companies that are gonna look for an exit strategy because there's still some sort of lingering fear about what's gonna happen in the markets.
There's a debate, but I think we're gonna see a lot more of this, and the big companies are gonna get bigger. But I think it was a good strategic move by Palo Alto Networks. Um, and, um, you know, we'll see, they, they're probably not done.
They'll probably continue to do these types of deals. That's true. You know, you made me laugh because, you know, observability is like one of those words, like jocularity, everybody knows it, but nobody wants to say it 'cause they football.
Yeah. Yes. Um, but yeah, no.
What is you think I was gonna ask you, Mike, though, what did you think? I mean the, uh, the, the interesting timing and of, of this acquisition, I know that it's been in the works for a while, obviously, but the, the timing used really well for Palo Alto. I think this bodes well for the future because the thing about observability is the, the core idea is we're moving beyond a modern area of predefined s metrics that we're gonna be able to collect all this telemetry data and analyze it so we can get to the root cause of an issue faster.
Well, that has been driven mainly out of the app dev and DevOps world because they're trying to improve performance and reliability. But these issues apply to security and IT ops and all across the landscape. And as that occurs, I think what we need to see is maybe some unification.
'cause you know, the security people are collecting telemetry data too, and so is it ops and so is the DevOps teams. And now we got more telemetry data that we're collecting that nobody knows what to do with and it costs a fortune. But Fred, is there an opportunity here to unify all this stuff?
Sure. My agents don't know what to do with all that information. Right?
That's the theory. Uh, you know, we, I think there's an opportunity here for, not just for Palo Alto who sees a bunch of writing on the wall, but you know, if they have a massive fleet of agents, right, they need to manage and get the telemetry and the optics around this. So observability is, you know, it's, it's rudimentary for everybody that's not a pure cyber company.
But now we understand, okay, cyber also includes, you know, agentic behaviors and all these other things. AI is the du jour. So if you're going to wander in with software and hardware and do these things, man, it'd be awful impressive to have something that allows you to navigate, manage, and disseminate that information.
'cause the key is resiliency and agility, you know, not, uh, not whether or not we can stop cyber attacks, right? That's the business problem is not cyber attacks. The business problem is resiliency and agility.
Mm-hmm. Gina, what's your take on all this? Can we all maybe link arms and have a kumbaya moment, it, ops, DevOps, security people we're all gonna like, talk about the same thing at the same time for the first time ever?
Well, sure. I mean, that's the, the dream of it. And this is a perfect, perfect use case for so-called ai.
It's, it's really machine learning, probably a little deep learning, but it's the perfect use case for it. There's too, too many alerts, there's too much going on. And if you had only known this server was getting a little too hot 30 minutes ago, you could have done some action to make sure nothing went down.
So I, US is an ops person's dream. Are you kidding me? You love it.
No, Barbara, it's no secret that there's not a lot of love loss between application developers and security people. Security people tend to view developers as kind of, well, the root cause of all evil. 'cause they created the software that led to the vulnerability that led to the breach developers.
They, the security people are, well, they're just in the way they need, we need to build software faster. And I got features to do and deadlines to meet. And, well, I can't be bothered with all the security stuff that generates alerts, most of which turn out to be nothing.
How do we kind of bridge this? 'cause this is a cultural issue as much as it is a technical issue. Maybe now they can be, uh, united in, uh, casting their blame on AI instead of each other.
There you go. But, um, I, I, I think, you know, one of the big questions here is, you know, as you have, um, agentic remediation happening, um, who audits the decisions that are being made? Um, what if the remediation goes wrong?
Uh, what kind of governance do you have in place? And I think all of these teams are gonna have to work together to define that. Um, otherwise they're still gonna be pointing fingers at each other.
All right, Fred, you laughed, but can I take all these people and just maybe throw 'em in a room and lock the door until somebody sees sense or what? I love it. Uh, the trope is terrific.
Uh, I, I don't really see that problem as much as maybe other people do, uh, from that standpoint. Um, but I agree, uh, Barbara's got a great assessment of both what the risks are, and I think you absolutely should lock people in a room. And maybe it's a little bit of a, you know, two men enter one man leave, you know, from, uh, from Mad Max.
But ultimately, uh, I think this will help drive better specifications, better standardization in companies. And instead of talking past each other about, I think this vulnerability has a high level of probability, uh, and that's very low on my development priority list. It's, this is a real thing.
We have a lot more telemetry that helps share that. And we have common data to look at rather than, you know, my tools say these things and your tools to those things. And our bosses have to argue about prioritization.
So I think it's a really, it, it could be a really big step up to diffuse some of the, you know, I'm a people person problem of how do you navigate both the requirements to the operational effects of it. So I think it's, uh, the, the future's bright. Alright, Barbara, does that work?
Throwing people in a room and locking the door? Is this a management technique? It, it actually kind of does to a certain extent.
Um, and a lot of it is, you know, what are the conversations you have in that room? Um, I, I think fundamentally, no matter what your role is, uh, people can benefit from putting themselves in each other's shoes and understanding where the other person is coming from. And, um, when you humanize each other and, and understand each other's motivations, that's where you start to find shared value and, and shared understanding of, of the problem and, and get to solutions.
So yeah, lock 'em up. Lock 'em up. Heard it.
Wait, that's a political campaign. That's a political comment. Yeah, you guys, I'm glad though, Barbara, I'm glad you're mentioning the, the, the importance of humans.
God, what a concept. I mean, given all that we've heard about ex agentic AI and how it's gonna eviscerate middle management or replace people, or, you know, take all these jobs and become part of like the customer service or the workforce, I'm, I'm glad that they're, these companies are starting, it's starting to dawn on them that people actually are kind of important in the whole process. I, I, I think that's gonna be the big differentiator honestly, in, in who wins and comes out ahead in, in this period of change, is the companies that see, um, the importance of, of humans, um, for their future growth and future opportunities, um, and who are investing in upskilling their employees.
Um, yeah, it's great when AI frees you up so that you don't have to do this tedious repetitive work anymore. Don't let those people go, keep the expertise and the wisdom and experience they have and grow and build it so they can do higher level functions and continue to be the advantage for you and your future growth as a company. Mm-hmm.
I guess part of my soul here is that maybe I'm just kind of not really using it at, at the level of scale that other folks may be talking about, like with Fred, but I often find that I'm like frustrated because by the time I validate everything the AI generated, I I've done it the right way the first time myself. So, Well, that's, I mean, that's a lot like having an intern, right? Um, you know, you, you only get in as much as you, or you only get back as much as you put into the intern, but then in time the intern grows in its capabilities, its judgment, um, and it can work more autonomously and, and grow into mature, seasoned, um, expert.
And, and I think AI is the same way. You, you know, going back to what Gina said earlier about going from A to Z, we're not gonna get there by going from a straight to Z. We're gonna go A to B, B2C, C to D, and down the path.
And so you have to start small and build that trust in the AI's capabilities that it's gonna do what you wanted it to do, that you trust what it, what it did, the decisions it made. And then when you have that base level of trust, you got to B, then you can start working on C. But if you try to go straight to Z, you're gonna have a really bad experience.
You're gonna give up and walk away and say, this doesn't work. And, you know, and then we failed. Mm-hmm.
I guess maybe, uh, I, I just wanna hire somebody else's AI intern after they train them and see how that goes. I feel that way about Claude actually. Yeah.
Okay. So Gina, though, let's bring this full circle. It seems to me that this AI stuff will get better as we expose more telemetry data to it.
And we don't have that data as, as widely as we'd like to think. So do we put the cart before the horse and we created all these AI agents and then expose them to, you know, random bits of data, but we really need them to focus on the telemetry data? Well, I think that's the rub, right?
Like, I, I think you can put the agents on whatever system that you have or whatever information you have, but you, it has to be focused on the right things for your business and for what, for the job that you're trying to do. There's so much promise to ai, let's get rid of the hype. But if you have a, a business need and you have that much data, uh, especially log data and in real time monitoring data that's just perfect to, to set agents on and, and help resolve a lot of issues before they start.
So, you know, you don't get caught off guard by something that you didn't even see coming. So, um, I, I think the agents are, it's, I think it's great. Everybody are starting to play with them.
I hate that we call them agents because I still think they're bots. I don't see the difference. Someone can change my mind about that terminology, but, you know, why not let let the computers run with the things they know how to do with supervision?
I think that's a, a great use of it. So, friend, coming back for a second to what we were talking about last segment there, is there an opportunity to kind of maybe democratize observability? And I'm asking the question because I, I've explained this to more people, and I care to admit, and almost universally I get the same response, which is, yeah, wow, that sounds great.
Followed by three seconds of pause and then it goes, yeah. Well, but I have no idea what questions to ask in the first place. Yeah, I think there's gonna have to be some consensus driven behavior around it, some standardization around it.
Um, especially as more and more, uh, agent interactions happen across, you know, the vast quantities of MCP servers that every company is standing up to communicate with their, what used to be APIs. It was now, as, you know, a natural language processing exercise. Uh, the, the challenge is the same.
So, you know, today we would say, and we'll we'll get to this in a minute, but let's imagine that we have several companies that are critical partners for us. And, you know, something happens with those critical partners. You know, how do we understand, uh, how do we inform and how do we account for that from a resiliency perspective as a, you know, partner said company.
Historically, we haven't really had a way to do that, but this offers a number of ways for us to think about, you know, if we have agents that are monitoring things about telemetry, that historically we would have, you know, maybe an model doing this, maybe we're doing statistics and aggregations and all these other things which are trivial tasks, uh, for, you know, a set of ag agentic flows. And so if you have a fleet of folks, uh, agents in this particular case, or, uh, there's a little difference, I think maybe in, in, in agents and bots, you know, we, we should certainly have a rock paper scissors, uh, engagement on that one. Um, but the benefit you get is, you know, I can task a fleet of these guys to go do this specific thing, which is observe, you know, my, you know, organization's telemetry, compare it to others in this sense and get a, you know, a good bellwether as to whether or not something is happening appropriately and make a decision about something like, do we need to find another more resilient route for a thing, uh, in the future?
And I think it'll also help hold companies more accountable, right, to their actual metrics that they say they uphold. Mm-hmm. All right.
Well, John, last question on this one though, but, uh, you're out in the valley. Is this the beginning of, you know, mergers and acquisitions across Yeah, I think it is space. Yeah.
I, I really, I really do. I mean, and I, and I, it's not, it's not related, but I mean, I think what opens really opened the floodgates was this decision in the meta case involving the FTC, trying to fight the WhatsApp and Instagram acquisitions that it had approved a decade ago. Um, that's another story in itself.
But yes, I think it's gonna be an acceleration. There's a lot of money, and I think there are a lot of companies that are really nervous about what we're, where we're headed, um, with all this debate between bubble and, and boom, I think there will be some companies that cash out and, uh, the large companies are gonna pick 'em off. All right, here we go folks.
It's gonna be cleanup in the observability aisle, rub it back, Discover Techron Group, the epicenter of tech innovation. We are your go-to for reaching IT leaders and practitioners worldwide. Our secret impactful content that sparks awareness, engagement, and top quality leads with us.
You'll access editorial websites, streaming videos, virtual events, custom content analyst research, and more. Join our satisfied clients. Let's revolutionize your tech journey.
Contact us today and tell your story to the world in the most powerful way with Textron Group. Hey folks, we're back on this Friday with our last topic, which is this CloudFlare outage, and it occurred earlier this week, and I think it only lasted maybe three or four hours, but then it cascaded for a while for people to recover. And it's very similar to what we saw with the AWS outage and a Microsoft outage.
And we seem to have these large scale outages these days. Gina, is this just like the new cost of doing business and it is the way it is? Or is there something to be done about this?
And are we too dependent upon a couple of things out there that have so many dependencies that they can take down? Well, everybody, Well, that's a couple of questions, right? So, yeah, yeah.
To me, I think this is the first thing me and my friends talked about was like, that's a lot of single point of failure that you may not even know is your single point of failure. So what happened was with CloudFare Cloud, uh, can't even talk today with CloudFlare, uh, was their CDN, their content development network went down and it was internal, it was a bug. And I wanna kind of read from their outage report.
It was a change to a database system's permission. It caused the database to output multiple entries into a feature file that's used by their bot management system. And then the feature file doubled in size and that propagated to all the machines in their network.
And that, uh, ba basically was what caused the problems. So the, the feature that that file helped the bot management system keep up to date with all the threats on the content management system. So it's kind of like a, a, a, a story of warning based on everything that we've talked about today, which is kind of interesting, right?
So we have a bot management system responsible for taking care of, uh, however it worked, taking care of any kind of security threats. Uh, a mistake was made by somebody, or maybe by a bot, I don't know, in the permissions that were assigned to the server. And it kind of cascaded this cascaded failure.
They definitely needed some observability so they could catch this as it was happening so they could get rid of it. There's definitely some questions about, okay, um, was it a bot? Was it a human?
Um, was it what, you know, like what does this bot management system, you know, how, why does it need this file? Like, all of the questions that come up into my mind was, how did it actually work? Um, so this is gonna happen, I think as we're getting used to automatically, uh, or having, having agents not saying that this is what happened, but having agents go out and, and make changes.
It should have been as simple change is what it sounds like. And it wasn't, and something ha it was just a bug in the system, which that also happens. Um, which caused this effect to have their clients go down.
So like for me, the thing it affected like x it affected OpenAI, which now Im impacts everybody's working life because everyone's using it. For me, it impacted my radio stations that I listened to online. I was annoyed and I had to listen to YouTube 'cause I was too lazy to get up and set up my record player.
So, but I have to, you know, I need the music to go on. So it kind of was like a disruption in my work life. Um, so there's so much we depend on that we have no idea that it's depending on a content management system that is, um, being protected by CloudFare flare and could go down because CloudFlare had an issue with data changing database permissions that kept, caused a problem.
Um, so, so yeah, there's a couple of things. There's number one, providing the, my providers of my radio station, the providers of X everything, anybody else, they're depending on CloudFlare and nobody else. So can they switch over to that other provider if something goes on?
Or is there even another provider that says good as CloudFlare? And then, um, just the consumers, you're not knowing what goes on and, you know, we're technical, we can figure it out, but other people trying to get on X or just trying to run their, the whatever they do for work with OpenAI just kind of stuck like it's not working. Um, and then calling all of us to, So, so full disclosure, full disclosure, tech strong is also a customer of CloudFlare.
We experienced some of that outages ourselves, but, um, here's what I'm trying to get at here. Fred, let's kick this to you. Is it the fault of somebody who created the config file that pushed this button?
Who probably feels awful right now? Or is it just that the system itself is flawed in a way that's create, is gonna create a problem and it could have been anybody in any time and maybe, you know, we shouldn't be beaten up on the four little engineer in the config file when the architecture may be the issue? Uh, so I, I think, so it's a different scale of problems when you look at 20% of the internet in general, right?
And this system was built for hyperscale DDoS attacks, right? It's a handful of, of services that sort of take all this information from click house cluster that says, look, what are the latest and greatest? So it does this every five minutes, right?
And so that same way you think about routes propagating with a network device like a router, uh, or a switch, and, and the same sort of thing happens and the construct isn't, you know, whether or not as we move into this, this was a, this was an ML process that generates this thing, and it's a bunch of features for a model to make decisions, uh, that, you know, then get informed by this, right? And the question is really about, there's a file limit size, right? Ultimately like a very human problem, uh, that wasn't dynamically adjusted.
Um, okay. Or database permissions that might have changed during the course of, uh, this process. But, you know, I think to your question, Mike, it's, um, you know, these things, maybe this is the worst outage since 2019 for CloudFlare.
We've seen a number of these different things. We're going to see some of these types of interruptions, but they're all, you know, methodologies that are, I would say, much further, uh, and more impactfully designed for resiliency than what most people deal with in their companies. They're gonna run into these types of things from time to time.
I think that, you know, the questions would be like, what's the bellwether for whether or not your file replication looks like, you know, there's a lot of after action, I'm sure these guys are working through and like, yep, we're gonna automate that thing. We're gonna put this back into the process. I need observability on this particular element here.
Right? All of that that'll all get after action, I'm sure super hardcore. And the question is just sort of when you turn the keys over, would that have made a difference if a human did it right?
Is probably the, you know, it's prolific argument, uh, versus, you know, whether or not, uh, a machine learning algorithm did it, or a, an agent did it. And at that scale, uh, when we think about what the problem is, maybe with hyperscale DDoS, like a human's not gonna solve that problem anyway. It doesn't matter, uh, at that economy of scale and that magnitude, right?
That's gotta be, uh, an automated workflow. And that process has to be, you know, pretty instantaneous five minutes, you know, is a, is a substantive time period in that type of, uh, in that type of threat landscape. And I think, um, you know, I don't wanna let, uh, kler off the hook, but I mean, it's just a different level of problem, uh, as broadly as it affected everything.
Uh, yesterday, us everybody, uh, from interacting with customers or their out their outputs as well. Uh, same challenges, Fred, how do you, they, they, they were quick to, uh, to specify that this was not, uh, outside threat. But that seems like a pretty obvious place to introduce a threat when you, you know, hin hindsight quarterback, Monday morning quarterback kind of thing.
How can they be sure it wasn't orchestrated from outside? That's a great question. I think that's why the, you know, first, uh, you know, the CEO would say that the first thing they did was evaluate whether or not that was an outside attack.
Uh, 'cause that's a presumption, right? Threat attack all the time. The, uh, the ASU botnet, which is, uh, sort of like they're, you know, contending with this thing right now, which is probably their first, the first jump was this might be actually that, that problem.
And so they work backwards, I think, uh, to get there. And, and at that economy of scale, uh, it makes reasonable sense because that's like a persistent threat for them. So I'm with you.
I, I think the challenge is, you know, once you get into the diagnostics of that, uh, the time that it takes to make an observability decision, and then again, internet scale, the time that it takes to unwind, that takes time. And so rate of propagation across the globe, you know, and all those things, what was the fix revert to a file, you know, that worked well, right? Like, like we know so well, that's, that's always the answer, revert whatever that change was, like, put it back, fix it.
But I Wanna get to two points here with Barbara though. So one is I think we can give the cloud player people props because they own this pretty quickly and they got up on social media and they basically said, you know, are bad and, you know, we apologize. And, and so that is a good thing on one hand, correct, Barbara.
I mean, that's the way to kind of handle these things. HA hundred percent, uh, owning it, acknowledging it, being transparent about, uh, what's happening is, is critical. Okay?
Second part of that question is that may be cold comfort to the IT people that contracted them in the first place. 'cause I'm sure they're getting a call from their boss going, how come the website's down? How much revenue are we losing?
And then the third question is invariably, well, who picked CloudFlare? So I, I mean, I, I think the scale, um, uh, at which, uh, these systems are operating, um, it, you know, going back to what Fred was saying about whether or not it's, uh, the fault of a human or a bot, I, I honestly don't think it matters. Um, it, in this day and age, it's how do you, what do you do when you have an issue that comes up?
You know, how, how have you prepared your people and your systems to navigate that issue and respond as quickly as possible? Um, uh, figure out the right solution if that's reverting back to the last version, whatever. Um, uh, it's, you know, what's your, what's your fail safe plan?
Um, what's your, um, you know, how are you preparing your people and your systems to respond to these outages? Um, that's the, the focus now as much as uptime is, John, you and I are probably one of the few people out there outside of Wall Street that actually read, you know, 10 Ks and SEC statements and are they now gonna include things like we are over? Oh, they do.
Yeah, they do. All Providers. Yeah.
So in every, in the, all these documents, they have the risk assessment. So they, they point out their outages, or this is actually, that's, that's a really good way to find stories, by the way. So you look for, you look for you, you do a search of under risks, and they do, they, they mention everything that went wrong, but they bury it deep within the, the document.
You know, I actually think, and I'll, and as an Xfinity customer, I'm used to outages and I'm wondering if, given what's happened with AWS and what happened with CloudFlare, and I think CloudFlare correct me if I'm wrong, didn't, wasn't there another incident several months ago? Um, I think, but, and regardless, it's something we're gonna become accustomed to. I, unfortunately, I think it's part of this whole kind of dynamic that we're living under and living with.
Um, so I think we're gonna see more of these outages, um, as these companies make their transit transitions and a lot of 'em are making major transitions, trans transformations. Um, I think this goes with the territory. Alright, Fred, last question on this whole thing.
Is this an argument for chaos engineering? Because theoretically you should just be ripping things out just to see what breaks anyway. Oh, man.
Uh, On the spot, Fred. I, I think these guys regularly practice chaos engineering. Uh, I, I think there's always an n plus one system that doesn't have the rigor and resiliency you expect when some magnitude occurrence happens you didn't plan on and yeah, uh, regularly implanting chaos telemetry data in your regular operating procedures, uh, is a great thing to do.
It also creates change. It, it also creates a set of, uh, unknown variables. Uh, and in certain systems you'd absolutely wanna reduce all of those things so it doesn't have a place everywhere.
But, uh, I'm sure that there's a handful of folks sitting in a room right now gaming out every other possibility for this system and the 10 that are adjacent to it, that it impacts upstream and downstream. Um, and so, you know, what'll be great is the outshot, right? For folks that have similar sort of telemetry requirements down the road, whether it's a CloudFlare or an AWS or even, you know, your Netflix, right?
From that perspective, right? To get back to your chaos, the theory. So, All right folks, well, I think what we've established here is what I'm gonna call the new Monty Python School of IT Management.
Expect the unexpected. Hey, thanks everybody for being on the show and sharing your thoughts and your insights, and please stay tuned for the rest of the text on that TV lineup. It's gonna be awesome, and we'll see you all again early next week.
Hey everyone, welcome back to our day three coverage of Kubernetes of KubeCon here in Atlanta. It's been an exciting couple of days. We've had a lot of guests, we've talked about a lot of things, but like any good conference, in person conference, the best conversations take place in the hallways.
And I've had a lot of those too, and I'll be talking and writing about that in the days and weeks to come. Let me introduce you though to our next guest. His name is Scott Rosenberg.
Hey, Scott is with a company called Terrace Guy, is that right? Yeah. Excellent.
Scott, welcome to Text Drunk tv. It's great to have you, Ben. Thanks.
It's great to be here. Absolutely. So we'll talk about Terra Sky, we'll talk about Cube Con, but let's take a moment talking about Scott, give people a sense of who you are, where you've been, like your journey.
Yeah, yeah. So I grew up in Chicago, moved out to Israel, um, you know, back in 2005, and really started my tech career around 20 14, 20 15. Started in the mainframe world.
Really. You know, they still exist today, apparently. They, They know they still, they, and they will 20 years from now.
Exactly. Believe me, they're not going anywhere. It started there, moved into more the VMware, uh, you know, virtualization private cloud area.
And then all of a sudden around 20 18, 20 19, just this bug of Kubernetes hit me. Um, that's when I joined Terrace Guy as well. Um, okay.
And ever since I've been just working up in the Kubernetes space, I've been privileged to be a contributor to Kubernetes for six plus years. Very cool. Working on a bunch of the other ecosystem tooling, cross plane backstage, um, and like what, what, what I'm known for, which is awesome, is I have never been into a single session at CubeCon that I'm not doing because as you were saying, the hallway track are the most interesting conversations Absolutely.
That there can possibly be. And you get to meet all these amazing people here. I can watch the sessions live or the recorded sessions later on.
Absolutely. And you, I just wanna get to meet the people because You only have that window. Right.
Exactly. And then everyone goes back to the world And all these people I'm writing with on Slack every day, it's like you find Well once you meet. Oh, that.
And that's a big thing too. You know what I find people I meet on Zoom, they're always taller in real life on Zoom, but Exactly. That's have a Body.
Well, because it's, because you never, you know, it's hard to judge how tall someone is from here. Exactly. Right.
But anyway, you know, you mentioned Backstage and of course Backstage Con this year Yeah. Was a big success platform con here was also, you presented a platform con Here. Yeah.
And a backstage Con. Oh, you did both. I, I had four sessions this year at, uh, cube Cut.
So, exactly. But good For you. Yeah.
Well, why don't you share a little bit about what you, uh, presented on? Yeah. No.
So I, backstage Con was really about building out a platform with cross plane backstage, um, and trying to bring together the consumption layer and the operation side from cross plane and trying to fix that. But like, my favorite conversation, my favorite talk I gave this year was actually a platform con, um, which was with a amazing, uh, guy Shimmer m uh, from a company called Play Tika Uhhuh. Um, and they're a unicorn gaming company in Israel.
I, amazing company. And we've been working with them for probably like 10 years or so. Really?
Okay. 5. Oh really?
Okay. Um, so back in 2016, and really it was, we talked about the journey that, you know, I've been going on with him together and how we have moved from a do it yourself chaos that happens with early adopters, um, into a real full CNCF bank like platform now. And all the benefits that it's brought them and what the challenges were.
Um, and it was an amazing conversation because it was the interweaving of the technology with the cultural elements and how it was able to be done, um, with CNCF projects. Love it. I love it.
Um, Scott, so at some point, it sounds like to me, he didn't say it, but you made the transition then to really platform engineering. Exactly. Right.
I, that's basically what we did, right? It's so much of this was a cultural change. Right.
Um, one of the challenges like that we always see is DevOps. One of the reasons that there's that line, DevOps is dead. Right.
Which is so, okay. It's a cute line. It's marketing.
Exactly. But why is that kind of true? Because DevOps never reached its full potential.
DevOps was supposed to be a culture and became the synonym for Jenkins, GitHub actions, CI, whatever it is. Right. And Kubernetes and as no DevOps was a culture.
Those tools fit well into that DevOps culture. What I love about platform engineering is that platform engineering really is an implementation way to make DevOps actually succeed. And the two are actually together.
Platform engineering is a mechanism that allows us to actually bring DevOps a platform to fruition Platform. Right. Right.
And that's what a platform helps with. And really what we did with them was, I, this was a over a year project of assessments, of interviews, of gathering data, everything, data driven, and really coming and understanding the different pains of each team, and then going and building out this platform together with the teams to get their buy-in to really restructure things around and build that full platform out from the ground up. Got it.
Right. With the challenges of Brownfield and large scale and all of that Real life. Exactly.
Yeah. You know, it, it's funny, Andrew Clay Schafer is one of the, not founders, but a big, a big name in the DevOps space. He always used to say, he still says, I guess the DevOps you get is the DevOps you deserve.
And Yeah. Part of that is, is because you're right. A good, the, the heart of DevOps was culture, right.
It was how we interact in a team. And for too many people, they used to pay lip service to culture and they'd, what I call bag dive into the tech. Exactly.
Yeah. Yeah. Culture, culture, culture.
Jenkins. Exactly. Culture, culture, gi ups.
But they never really, they never really invested in the culture. Exactly. And I think the other thing is, if you look at DevOps that way, it's not meant, it's not meant to replace Agile.
No. It wasn't supposed to be from the beginning of time to the end of time. Right.
Right. I I think it needed platform engineering because it had to sit on something. Right.
Now you've got, so now the way I look at the world Yeah. You've got the platform, right? So it's like on the first day God created the earth.
Right. Right. You got a platform.
Exactly. It's one way of looking. Probably a little sacrilegious.
Exactly. But all right. All's good.
You, but you know, you got the earth now, now we can build on. Right. We could build it on the second day we created developers in code.
Exactly. And, and, and, and we do it in Scrum and, and Agile. And that's the best way or monolith.
Right? Right. But we, we had this code or coders, and then we had ops and, and you know, it was like, can and evil, almost very different.
DevOps tries to bring that kain enable, if you will. Right. Not possible.
Kain enabled together where we, we, um, you know, it, but it it sits on top of the platform. Exactly. And Agile plays in there and DevOps plays in there and SRE is in there Oh, yeah.
Of security and observability and all these things that we build on this beautiful Exactly. Earth, this beautiful platform. Yeah.
It's just, it's almost like we did it best backwards though, in that we Right. We brought the platform in after all these things were running around. And, and I think the truth of matter is, and this is like one of the things I've been saying in like a bunch of the platform engineering, uh, different communities, is that, you know, having a platform is not a new thing.
No. We have had platforms for decades. What we haven't had is a unified platform.
Right. And that is the difference here. Platform engineering is the idea that we need to focus on the platform level instead of just having it as a given.
And we have to unify that platform because what we had was a bunch of platforms in the past, and we did have platforms, but they were all disparate platforms. Yes. And we had chaos.
What we want is now just let's bring those platforms together into one unified one and build all of those tools on top of it. Right. Which is so many people say like, oh yeah, we wanna build our first platform.
I said, you have platforms, you wanna evolve into a unified platform. Yep. Right.
And that's the key, I think. Well, once we get a build, then we could rest for a day. But anyway.
Exactly. Let, let's go. Let, I wanna move to IDP.
Yeah, right, because you were back backstage con. Right. You know, just on that in IDP though, it's the, I will say the tech world can't figure out acronyms for the life of them.
'cause what I always say is you connect with your IDP to your IDP to visualize your IDP. You have an identity provider, right. To your internal, Talking To your internal developer platform.
Right. Which IDP are we talking about? You talking about I met the internal developer platform portal and both.
Exactly. But this, you know, they used to say when platform, when I first became aware of this whole platform engineering movement, you know, basically if you got a handle on Kubernetes, you, you had a platform. Right, right.
And the engineer was doing that. I think that we've seen that focus change from, you know, handling your Kubernetes cloud native sort of architecture to, to IDPs. Yeah.
And, you know, and platform portal there. Um, and it's interesting because it's, it's, some may say, well, it defo us from working, you know, perfecting Kubernetes. We got enough people working here Exactly.
To perfect Kubernetes. It's okay. It's, it's not going to exactly be an orphan.
Um, but, but we are seeing that. But here's another thing. You know, they say two people, three opinions.
Yeah. And we're starting now to see, uh, differences of opinion different communities. Yeah.
org. They have a, a booth down the hall from me here. Yeah.
A couple rows back. I work very closely with Luca and the community. Yeah.
We saw a platform come here, put on by CNCF. They offer certifications. Those people offer certifications.
Is this town big enough? Uh, I think it is. I think that, you know, obviously there's always gonna be some challenges, right.
In differing opinions. But I think it's very similar to what we saw in the GI UPS world happen when you had weaveworks that had Flux. Right.
Right. Flux. Then you have Intuit that go and do Argo ccb, and then you come out and Flux V two comes out.
Right. And you have these two tools. You even have the other ones, like on the side.
Yeah. No, there are patch Fleets and the carve cap controller, and you have some others. We see it in service mesh too.
Right. But what ended up happening, they all come, right. They go in their own directions.
They're figuring out what GI Ops is, and then what happens, the flux guys and the Argo guys come together and create the open GI ops initiative, which comes as great. The, this is what GI ops means. We're different implementations of the same thing.
And that's fine. Choice is Good. Exactly.
And I think that this is healthy because having two organizations, having two groups trying to figure out and really distill down what platform engineering really is and what its core tenets are, is so critical. Because once that happens, the two can come together, let everyone go off on their own approach. Similar to how people with agents, they run twice an agent, whether they're doing a no, absolutely.
Let's see what we get and try it out. But variety's the spice of life. Exactly.
Right. And so, and I think, I think that's what open source is about too. It's about choice.
It is, it's about freedom to, to decide, you know, and, and if, and if, and if Backstage, the code doesn't work exactly the way I want it to, I'm free to change that too. E Exactly. It's right.
And I mean, like Backstage is one of the best, I think proofs of like what a real, like open source community is looking like. Because what we've seen is like Spotify, when they built it initially and they contributed top stream, they work in a very specific way. And then you had companies like VMware that came in, right.
And they wanted to do something a bit different, and they changed. Absolutely. Red Hat came in and you see these companies have different, A opinion, but it didn't stop Spotify from saying, you know what, we wanna sell a commercial version of it, which is awesome, our vision, and it's all good.
It is big enough for everyone And, and everyone can push it in their Direction. That brings up another, another question though I want to ask you. Yeah.
I think one of the questions for the platform engineering community is, do you want commercial IDPs, like the Spotify backstage? Which kind of, it gives it to you in a nice package. It's all wrapped up nice.
It looks good. It smells nice. Here you are packaged, just add water, right?
Or do you want me to give you sort of that old time open source, here's all your packages, you pick what you want, put it together and it's a DYI kind of thing. Right? So, you know, it's a really great question.
And it's like, you know, it's something that we've been grappling with for a while with a bunch of our customers and the people that I'm talking to. And I think it really just comes down to there's no one size fits all. And it depends on the customer.
Yeah. You can't be a jack of all trades. No.
And therefore, in a company that's large enough, right? A big enough unicorn style company that has the ability to really go and fine tune that internal developer platform to be exactly how they want and can base it off of open source and can have a large enough DevOps team to maintain something like that, that's a very viable solution that's probably gonna give them a better platform long term. Because there's a reason it's called an internal developer platform, because typically it's built internally.
Internally, right? Now that doesn't mean that on specific components that you don't have the expertise on, you shouldn't buy an enterprise version of it. Right.
I think that, on the other hand, platform engineering can help so much in the smaller companies as well. Right. And there, I think that you have to go with these commercial Offerings.
They're prepackaged Because a, you just, They don't have the resources. They don't have the resources, and they don't hit the challenges No. That these platforms can't cover and Scale and stuff like That.
Yeah, I agree. If you Have 50,000 nodes in Kubernetes, I promise you that most of the paid solutions out there ain't managing it. Well, no, you are gonna need to figure it out yourself.
Agreed. But if you have 50 nodes in your environment, I promise you, any of those IDPs out there are gonna to help you. So Agreed.
Hey, man, I appreciate it. Thanks for coming on here. Awesome.
This was great. Scott, if people wanna maybe follow you or Yeah, stay on top it. What's the best web?
cloud, uh, you can check me on weekly, uh, podcast together with Victor Farik on, uh, the DevOps School kit. Victor's great friend. Tell Victor I said hi.
Hi, will. I'm a good friend, a good old friend of Victor. Victor's great.
He is a great Guy. Uh, so we do an ask me anything, ask us anything on every Thursday, basically. Very cool.
Um, and reach me out on Kubernetes Slack, CCF Slack, cross plain Slack, LinkedIn, any of 'em. Exactly. All over.
Love the community and great being here with you. Thank You, Scott. We'll come back again.
Hey, we're live. We got more coming to you. And as we wrap up our, uh, day three coverage here of Q Con Con, native Con, this is Alan Shimel, standby.
Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. We're here with Charlie Cartwright, who's director of the Quick, uh, AI agents to help you all workflow. And we're gonna a little chat about, well, just how do we get all the AI agents to start working for us?
You know, explain to us how you guys envision all this coming together, and how do I actually like, tie something together across different products and services to actually drive the outcome as intended? Yeah, absolutely. I think with, with Amazon Quick Suite, you know, we're, we're going to reimagine the way that, that people do work in, in the workplace.
And what you're touching on is exactly what we've been focused on. I think if we rewind the clock clock, like in the last two years, I think we've all all had those amazing moments of delight where we were using consumer AI applications and it was able to do wonderful things for us. But doing that in an enterprise setting, like there were point in time, uh, benefits, but you didn't necessarily have those applications, those AI applications connected to an information ecosystem.
And, and that's really the key here to be connected to your data. Um, like data stores, for example, right? To bring in all that structured data.
And the same thing for document repositories, to have access to all your strategic documents, your internal wikis and web search that can come in. And then equally important is to be connected to all these applications for which users often, uh, operate with on a daily basis. And that creates this like, layer of context that's also relevant for humans to be able to get access to relevant information quickly to make decisions.
But that is the same context that agents need to be able to execute on my behalf to interoperate with one another and to actually operate with autonomy. And if they're connected to those applications to take actions, and they have the enterprise context that is needed to properly perform those actions, then you start to get the real utility that we've, that we haven't seen in some of the enterprise applications From the uninitiated. Give me an example of how that plays out or something that maybe you're personally doing that the rest of us would go, wow, I didn't know I could do that.
Yeah, yeah. No, thank you. Um, so this is the, this is one of the amazing things about this product is, you know, me and my team, this entire organization, we get to utilize it and work with it and experience the benefits.
Um, just recently I was working on, you know, something to look at, like a competitive landscape, for example. And previously, this is something that I would go to various different sources. One, like a competitive landscape document maybe that I, I wrote, uh, a few months ago, right?
And then I would look across information, my, my email, various document repositories and, uh, dashboards and data that I use to understand maybe some things internal and external related to adoption, uh, product capability performance. And then couple that with a lot of time searching on the web to understand what's out there, how these capabilities compare. And this is something I was able to do with our research agent in Quick Suite, um, which previously had taken days coordinated across several people, and I was able to do this in, in less than two hours.
And, and so it kind of like this remarkable transformation where something had taken me so much time just to go across, you know, tens of different sources to collect that information. And then I finally get to the point where I can start to analyze it and think about it. And I was able to short track that and, you know, within two hours, think about some of the opportunities that we have as we build our product, as we go toward general availability in comparison to, to what's out there, for example.
So like a remarkable change in the amount of time that me and my team spend on a daily and weekly basis just because of how well Quick Suite is integrated into and has access to all the enterprise information that I utilize to do my job on a day-to-day basis. And the same is true for people at various different, um, roles in the organization and different levels in the organization. How much faith can people have in the output of these tools?
'cause everybody's been a little bit concerned about hallucinations. So how do I validate what the AI agents are coming back with is what is actually in that data? Yeah, Mike, great question.
And, and of course, this, this is something like, there's no bigger trust buster than when you are asking a question and you know that information, you can see that information and, and here it is this like AI agent or, uh, assistant can't represent that information correctly. And so of course we're doing, um, benchmarks where we look at different question and answer pairs that we can ground truth against so that we can do things to make sure that we answer with the right relevance. Also, um, when you have structured data, right, like actual data sets and things like that, and all this context across the, the various information ecosystem within quick the document repositories, uh, access to email information in Word documents, various drives, you're able to give the right context to the underlying model.
So the, the need to hallucinate and the likelihood of a hallucination also decrease. And, you know, that can often be the case when that context isn't there. And that model is essentially trying to answer that question the best it can that the user's posing.
And in this case, we're able to equip these agents and the underlying models with that context. So not only can they coordinate, communicate, take action, but they're able to do that with, uh, users as well. Hmm.
Is there per chance any way that there's a setting for the AI agents, so I can kind of determine how aggressive I want it to be to answer a question so that maybe some of those wrong answers get tempered a little bit because as far as I can tell, the AI agents are just trying to please us and they go to any l to do it. Yeah. One, one of the features that we have, the capabilities that we have in the quick suite is the ability to create your own chat agent, like create your own chat assistant, essentially.
And, and with that you can, um, change and dial aspects of the, the persona, the instructions for that chat agent, just how creative that chat agent goes, how much it needs to stick to a script, so to speak, right? To to go in the bounds of a particular goal, to only look at context, um, for a certain set of underlying, uh, data stores or document repositories or a subset of other applications and actions. So you can scope that down, but then also fine tune that persona and kind of dial that creativity.
And what you can do is get an agent to respond very much to the objective that you're trying to achieve. So I can fine tune that. How customizable is the overall experience?
You mentioned the chat, but um, every company has slightly different workflows, so how does the platform actually know what my current workflow is? And for that matter, will it suggest ways to improve my workflows? Yeah, so I think to the, to the first part of that, um, Amazon Quick Suite I is, is highly flexible in terms of how this can be utilized specific to a user's workflow.
So if that's certain daily routine task that I, you know, a user's taking information from one application or one source and moving that to another, those are things that can be automated with a, a wide spectrum of automation capabilities that we have from low code to more advanced enterprise workflow automation. And so really what it is, it just kind of depends on, um, the different applications and information that are available. And that's what really drives that customization and the opportunity to automate some of those workflows that exist within an organization.
But it's not that the capabilities themselves, um, are restricted or, or so specific that they can't generalize across different organizations workflows. It's again, back to that context and those procedure documents that can be given as context. So the enterprise context and maybe a procedure document or an SOP that can be used then by that genic plan agentic planner to create an automation workflow.
Mm-hmm. How do I manage the level of permissions that I give to each of these AI agents? 'cause from what I've seen so far is in their quest to please us, they can be overly aggressive in accessing any and all data they can find.
So do I have to think through what it is exactly? I'm gonna let that AI agent see. Yeah, I think this is, this is both important for users and agents.
So as, as with Amazon and Quick Suite users now have access to information and so much information that, that maybe they didn't realize, right? Or it wasn't used on a daily basis. And so the things that, that we are very mindful of and, and take with a high degree of responsibility is making sure that users only access the information content that they have access to already.
And that through quick, we're not giving access to anything they shouldn't. And the same is true with agents, and specifically with agents. When I, you know, when I create a custom chat agent, for example, I can actually equip that chat agent, I can narrow its context, it can start with the same context that I have access to, for example.
Or I can further narrow that and I can give it access to a subset of actions, um, or just a few applications even. So the flexibility is really there to, to take, to take that agent and focus it on a very specific goal so that my interactions, I can limit the context and I can limit the applications and the actions that it can take within those applications. And also on the reverse, the other side of that spectrum is also capable where I can take and, and customize that agent or make sure that it has similar access to many of the applications, document repositories, internal wikis, and data stores that I have access to.
Mm-hmm. Um, a lot of organizations are experimenting with AI and they, some of them have even gotten so far is to build their own AI agent. But I kind of feel like we're at that point again where we're trying to figure out when do I build versus buy an AI agent.
And I think a lot of the use cases that people are coming up with are just gonna be things that are gonna be standard features of a platform like yours. So where's that line between build and buy in your mind? I think the, the build and buy discussion always comes down to a speci, a specific set of use cases, objectives, and goals.
And we look at Quick Suite, um, some of the, the one P agents that come with Quick Suite, for example, a PhD level researcher, right? A business analyst that can not only look at information that's in a dashboard, but that can understand and extract and generate insights on the underlying data set, which is very powerful. And then a team of automation experts.
So we have this broad set of capabilities that come with Quick as part of our One P agents. And the other thing that we offer there, Mike, so let's say that a, a customer or a user, and maybe it's a power user, they wanna introduce their own agent, right? And, and they wanna build that on like AWS infrastructure or somewhere else.
That agent can then be brought into and can be coordinated with, um, among the Quick Suite agents to maybe offer some specialized capability that is not there in Quick Suite today. And so what we're trying not to do is take, take this position that, you know, everything has to be here in Quick Suite. And just like the applications, we want users to be able to connect applications and data stores, whether they're part of AWS or not part of AWS, we wanna make sure that we can meet users where they work and how they work.
And that quick is extensible into those applications, whether it's a web browser or it's word for example. And so that's a very, very important, uh, capability within Quick. And so we wanna make sure that for organizations that need to build a specialized agent, we need to enable that, and we also need to support that agent to have access to essentially this enterprise comprehensive information so it can operate with the autonomy and the agency that the first party agents can within quick as well.
What LLMs or platforms are the agents using to accomplish these tasks? And are they permanently tied to those LLMs or will they change and swap them out as LLMs Advance and different ones are available at different price points? That's exactly it.
We, we evaluate and we change models based on the different capability to achieve the best outcomes, um, for users for that given capability, right? And so, and so that's really important because as models continue to evolve, we wanna make sure that our users are getting the best experience possible. And so nor do we wanna be, uh, held to one particular model.
And sometimes one particular model isn't the best for all the broad set of capabilities that are available in quick. Hmm. So what is the pricing for quick look like?
I mean, a lot of folks are, um, concerned that over time that these will just get more expensive and cost prohibitive, and other folks are saying, well, a lot of these agents maybe are low cost today, but over time the cost will add up. How do I think about the total cost of doing some sort of agentic AI workflow? Yeah, so, um, we, we basically have have two, two SKUs or two tiers, um, a professional enterprise.
And in, in the, the professional or the $20 tier, um, users get access to all the capabilities that we talked about today. So they can chat across their data, they can use, um, you know, PhD level research, uh, business analyst, and they can also automate workflows like no code workflows with a capability that we call flows. And so they have access to all those capabilities and the, the things that we do there, um, for research and automation where there's these agentic hours that can essentially be accumulated.
Um, we offer agentic hours for those capabilities and for users that need to consume even more of those agentic hours for that research agent, for example, there's a $40 skew where we extend those hours and then also offer consumption so that users can go into, you know, as many hours as they need for those power users. And the same thing with automation workflows. I think I had mentioned briefly earlier that we support the spectrum of automation.
And so one is this no code, natural language automation builder, where you can think about like routine tasks being automated. Like maybe, maybe I wanna create like a weekly business review. I wanna automate aspects of that or, or fully automate weekly business review or, or pull information to prepare for customer meetings and send that out to the respective, um, owners for those meetings to resolve some open customer issues.
Like those types of things can be automated and flows, but if we're looking at more complex automation in the creation of that complex enterprise automation for that, we have quick automate. And in the $40 skew, um, folks can author and create that enterprise level automation. And that, that automation comes with Angen planner.
Uh, it also has like the build, observe and deploy capability versioning that you would expect in more development like tools and also supports human in the loop so that users, um, so, so that like when the workflow maybe gets stuck or the designer of the workflow, the author of the workflow just wants to put in a, like a criteria for a human to be the decision maker if a policy, um, reaches above a certain threshold or something like that and quick automate that can be supported. So ba mike, so basically to, to reiterate, we have two SKUs, a $20 SKU that offers access to all the capabilities we just discussed. And then for power users that want to create these complex automate enterprise automations that's available in the $40 skew, along with higher limits on the ENT hours for research and automation.
So how will workflow automation evolve in the office going forward? Because I think, you know, there's a lot of requests that individual departments have put into it to help them build something, and then there's a low code tool and it went back and forth and usually dies in the vine. And then there's also the whole notion that, um, I, I wanna be able to create a workflow that's kind of disposable.
I might only need it for a little while. So is that kind of how this is gonna evolve? We'll see more disposable workflows because individual end users will be able to do things without necessarily requiring somebody who knows how to code something.
They do everything for them. Yeah, that, that, that's exactly it. And, and so we expect like all of these types of workflows that add add weight and complexity and und differentiated work that, that users have to go through today just to complete their job, but they don't necessarily add that differentiated value to their goals and objectives.
Um, these are things that previously would probably be stacked up on like a center of excellent department, right? That may be looking at robotic process automation or even, uh, teams of engineers that were looking to automate certain capabilities, but maybe it just couldn't be justified with some of the other things that were on the plate there. And now users in a no-code, natural language prompt can automate those workflows.
And so to your point, even if that workflow doesn't maybe have weeks or months of longevity, or it serves for a a point in time, um, that can still be highly beneficial and a no-brainer for a user to go in there and automate that workflow. And of course, the same is definitely true on the other end of the spectrum for the more complex automation use cases where you can use agents and natural language and maybe a process document to get started to create what is a much more complex comprehensive workflow. And then the user can go in there and a adjust or insert workflows as they need at any level and debug as necessary.
Of course, that comes with test runs so that validation can be there, but it's absolutely the case, Mike, that we will see a level of automation opportunity that was just not attainable before. And so, and I'm really excited, you know, even in, in internally before we, uh, went with general availability with Amazon Quick Suite, you know, we had tens of thousands of Amazonians that were basically battle testing this. And so it was, it was energizing to see like the excitement and the love that they had and the utility they were getting across these capabilities.
And we were seeing exactly what you're talking about where we have folks that had probably never au not probably, that had never automated a workflow before that were coming to us and saying, Hey, look what I've done. I used to have this, I used to collect this survey across all these departments across the breadth of Amazon, and that took weeks and weeks. I think they described it as months even, and they were able to do that in, in an hour with Amazon quick.
And so, you know, these success stories are, are remarkable and it's so great to see them. So we are seeing that exact transformation that you described internally at Amazon and, and we're also hearing of these great success, success stories, you know, with our hundreds of external beta customers as well that entered the, the Quick Sweep beta program before launch. Alright, Hey folks, you heard in here you can go after some big giant AI project all you want and you might swing for the fences and miss, or you can play small ball and just go after all these smaller little workflows that at the end of the day collectively are probably gonna make a much bigger difference.
Hey Charlie, thanks for being on the show. Thanks, Mike. Thanks for having me.
All right. And thank you all for watching the latest episode of the Techstrong AI Leadership Insight series. You can find this episode and others on our website.
We invite you to check those out. And until then, we'll see you next time. Hey everyone, we're back here live at Cube Con, cloud Native con wrapping up our day three coverage.
I want to introduce you to our, our next guest, and if I mispronounce his name, I apologize. Goov Sena. Yes.
Gorav is Gov's a real life platform engineer for a major automotive company. Um, he gave two talks here at Q Con this year. One was on behalf of the open SSF and one was on behalf of the, uh, PE Open Telemetry, oh, Excuse me, the hotel.
Yes. Uh, conference Now, Monday of this week was satellite conference day. So we had hotel, uh, we had the hotel conference, we had the open SSF conference.
We, they had others. We had the PE conference, we had a backstage conference. I go call, they were Argo, Argo, c, DK, and the rest, you know, it's the day for all of them.
But gov let, let's hear about your two, uh, your two presentations. Sure. But before we do, yes, without, I don't want to, you know, we're not mentioning companies or anything like this, but give an I people an idea of, you know, how long have you been a platform engineer?
What do you, what's your journey been like? Yeah, so I've been a platform engineer for about 10 years now, and, uh, I today work for a major automotive company. So my work is basically, I, I lead a platform engineer team to build an internal developer platform.
So think like many verticals, observability, Kubernetes, CICD, database as a service, messaging as service. So that's the work we do today for our internal teams to improve their developer productivity and efficiency. Very good.
And uh, my talks that you mentioned about one was for the open telemetry. So basically the use cases are how do we do the over the air deployments to the vehicles? How do we monitor them through our open telemetry stack?
That was my talk basis on the generalized hotel concepts. The second talk was around the open security. So how do we do the secure software supply chain?
Like how do we make sure that software that we're deploying to the vehicles are secured by nature? So that was a talk around the, around the open security for Excellent. Yes.
You know, that, that's one of the interesting things about platform engineering, talking about observability, talking about security, but it all kind of fits in right to this idea of a platform, an underlying platform that we build upon, whether it's DevOps or, or cloud native or you know, we we're building on that platform. It, it, it touches all of those things. Got it.
Um, what, what was some feedback from the audience from, you know, peers and other platform engineers there? What did they think? Yeah, So I got a huge interest on the way we do open telemetry today.
And the reason being is that, uh, in my current role, we are using hotel stack for all the three signals, logs, metrics, traces, and how do we use this in a vendor agnostic way, right? It's an open, it's an open source standard. How do you scale millions of vehicles that are support on the platform?
How do you monitor them? How do you monitor the, the, the progress of the software delivery? How do you enhance the cloud software development through open telemetry stack?
Not only in terms of the observing, but also the, from the deployment perspective. So I got a huge interest, uh, in terms of like, uh, people coming in and asking like, how do we use open source software? How you able to like manage the fleet of collectors?
How you able to upgrade them? How you able to get the maximum performance out of those collectors, right? So that, that's a, that's the area of interest among the, among the audience that I got basically asking about like the, the, what's the blueprint, right?
Not essentially sausage making, but the blueprint, like how the high level concepts work, right? The other, uh, other theme around the open security con was basically the way we do our software deployments in terms of CICD pipelines. How do you make sure that those images that we are deploying are free of critical vulnerabilities?
How we are making sure that those images that we are actually taking from third party dependency softwares, they are trusted. So what are, what's the process? How we are doing the, the trust in the zero sec in the zero security network architecture so that there is a, not only the software that we are running, but it's dependence also, we make sure that they are in compliance with our security compliance point of view.
Right? So those are the love it two manual things. Yes.
If I appreciate that. And you know, the nice thing is probably both of those were recorded, right? Yes.
So they'll be available to download, correct. On, on the, uh, c fcom website, CNCF website. Yes.
Let me ask another issue or another question for folks out there who are saying, you know, I, this platform engineering sounds interesting, I'd like to be a platform engineer. Yes. Or, you know, I considered that, give people a sense of your day to day as a platform engineer.
What do you work on? What do you Yes. What are some of the problems you're solving?
Yes. Like what, what is it like to be a platform engineer? Oh, I, so I came into platform engineering field when, right when Kubernetes was getting shaped up in 2015 year era.
So it's been 10 years. And, uh, before then I was a software engineer. So coming from the engineering background, when I was an application developer, my knowledge, my skillset were only confined to that particular business domain level knowledge.
When I transition to platform engineer, my breadth of knowledge skill sets, not only in terms of the infrastructure, but how do we operate the services as a service operator, I was able to find out the business value because I today host many different applications that powers different critical user experience workflows on a, as a, as a operations engineer, I am able to make sense of why this services exist in the platform, how can we make it more performance in our infrastructure, develop development, and how can we power the, those developers to do their job efficiently? So it gives me a pleasure being a platform engineer, not only knowing the breadth of the, all the operations for those services that we operate on, but also powering them, enabling them in the, through the infrastructure in the cloud native way. So yes, I love it.
I enjoy my role at the platform engineering. Do you? Yes, very much Before platform engineering was a thing.
Yes. What did you call yourself? I was a, a generalist software engineer.
So I was working on, uh, things like, uh, video and coding. So I was, before my platform engineer, I had a background on writing the software that actually was in your mobile hand hand cells, for example, I'm talking about pre Android, pre iOS era, pre Apple era, where you used to have like a Nokia hand, hand cells, blackberry hand, handheld sets, and you had a very low power devices in terms of the capabilities back then because the compute, how you are able to render your media playback on your, on your handheld sets. I was working on those media encryption and audio compression algorithms back then and my contributions then actually got open sourced to the Oh really?
To the Android. Oh, that's great. So the android, the first Android version 2008 that came out was running the software that I worked on.
That you are talking. Yes. So good for you man.
Yes. Thank you for that. Yes.
Alright. Hey, we gotta wrap up. Yes.
Gov. Thank you. It pleasure for coming on.
Thank you for what you do for the community. Thank you. Hey, that's gonna wrap up CubeCon day three here, man.
I hope you've enjoyed it. Atlanta will miss you. Uh, we'll be back tomorrow with our regular tech strong tv, but for now, this is Alan Shimmel, we're out.
Hey guys, thanks for the throwaway here with Chris McHenry, his Chief Product Officer for Aviatrix. And we're talking about, well, the degree to which maybe nation states are stealing encrypted data that they plan to decrypt someday and using quantum computers and how prevalent this all is. Chris, welcome to show.
Thank you very much, Mike. Good to be here. Some people my friend, are a little skeptical of this is even happening.
So what evidence do we have so far that nation states are doing this? Is this a theory or do we have some actual examples thereof? Yeah, it, you know, I think it's, it's difficult to ever say exactly what a nation state is doing, uh, but it is really interesting.
I think there's several pieces of evidence, especially with the investment that nation states are making in quantum computing. Now, uh, particularly we know that China has made some very big investments, uh, in, in terms of de developing the technology. And one of the primary use cases, uh, that are has been most attractive to nation states and for potential military applications has been the fact that it has the ability to, to break most of the fundamental standards of modern encryption.
So I think, I think the, you know, the, the things we know about quantum breakthroughs and we know who's making the investments in them, definitely lean towards the fact that, you know, there is some serious interest in this space and, and we need to be aware of it. Now, I will say one more really, really interesting piece of evidence that we have is one of my favorite, um, examples of a hack probably happened about 10 years ago, and it was a manipulation of the global internet routing tables that temporarily, temporarily rerouted about 40% of the traffic through, uh, through China. Uh, and so, you know, you gotta ask the question like, why is that even interesting to some extent, especially when most of the traffic on the internet is encrypted.
And then we did see last year, uh, a, uh, a really high profile attack on US service providers by an organization known as Salt typhoon, which is generally associated, um, with the Chinese nation state. Uh, and they, uh, managed to penetrate a lot of the core US service providers. And, and I think the use cases there were primarily about snooping on, um, very targeted communications.
Uh, it it, you know, relative to national security, um, maybe not as relevant for enterprise data, but we do also know that there's a long history of stealing trade secrets. So, I mean, it's, it is, uh, it, it's hard to say ever if something specific is, is happening, but we have a lot of evidence that points towards this being a potential potentially very real issue. Yeah.
Some people, of course, are waiting for what's known as Q day, which is when the quantum computers get smart enough to break some of those encryptions. Um, the question I have is how close are we to that and might we even know when Manay arrives? Yeah, I, I it's a really, really, it, it's incredibly difficult to predict.
Um, I think one of the things that's really interesting though is we have seen very meaningful progress in the development of quantum computers over the last couple years. Uh, not just from nation states, but also from private companies with announcements from Microsoft and Google and the investments that they're making there. Uh, you know, we also saw the federal government ratify post quantum encryption standards last fall, and we're gonna start to see a lot of those come into regulations in the next couple years.
But my best prediction is that we're looking at like early 2030s, but I think the preparations that organizations, the technology to help be preventative about this is, is gonna be, it's here now in many cases, and, uh, we're gonna start to see this be a, a best practice. Uh, really next year. I, at my, my prediction is the next fall is really gonna be the year where everybody starts to kick off their, their quantum encryption upgrade programs.
Now, one thing that's important is this is not the first time many enterprises have gone through this. 3. That was maybe 2016 ish timeframe.
So, uh, we are practiced in this, but I don't know that anybody really has a process. And so whether it's quantum or not, we are going to need to upgrade our encryption algorithms over time as both traditional and quantum computers get better. And I think next year is gonna be the year where enterprises really put this on their whiteboards.
How do I have this conversation with business execs who are gonna kind of look at you and say, well, let me get this straight. They're stealing data today that there might decrypt three or four years from now. And the business exec goes, I don't think any of that data is gonna be valuable three or four years from now.
I think that's a really real question. Right? And, uh, and so, you know, it, it, it, a lot of it depends on, on when Q day ultimately happens, but if you don't start preparing for how you can keep your encryption is incredibly important.
Like, we want to best practice as all of your data in transit and in rest needs to be encrypted, and organizations need to invest in technologies to allow them to keep their encryption up to date. This is really just another threshold there. So I think the harvest now decrypt later schemas are obviously, um, you know, very interesting and very pressing and will have an incredibly, you know, when Q day comes, we'll have a huge, huge, huge impact on how we think about cybersecurity.
Uh, but it's, uh, you don't know when that's gonna happen, right? And, and, and in most cases we probably won't even know because a lot of it is funded by nation states when it actually does happen. And so that's, you know, it, it, it, we are, we're getting close, we've seen the signs and, and now really is the time to prepare.
To your point, it may have already happened and we just don't know about it, but how big a lift is it to swap out my encryption scheme is on different products and what kind of effort is gonna be required and, you know, what kind of funding for that matter? Yeah, it's a great question, and I think that is one of the big challenges, right? 3, that we put it systems in place to make that easier down the road.
But I was, I've had multiple conversations with enterprises even in the last couple weeks where they know it's gonna be an, a huge challenge for them to upgrade. One of the, one of the big challenges is that, you know, not all of your applications running in your environment are modern applications. And so we need to think about multiple different layers of security here.
And, uh, and that's really where tools in the network I think will become incredibly powerful, uh, to be able to, uh, do network level encryption even when the applications can't. So it's like you are solving the problem at the trunk of the tree rather than at the leaves. You wanna do both.
Um, but there will be a lot of, there will be a lot of techniques and I think technology to, to ultimately improve that process. We're definitely investing in that space. You will see from us next year a one click easy button to do PQE and, and all of your, you know, for all of your application traffic in the cloud.
And, uh, and I think a lot of other vendors will be investing in this, in this space too. Some folks I talked to are counting on forklift upgrades. Their assumption is that sometime in the next four years, they're going to upgrade their servers, storage infrastructure and applications and whatever it may be.
And whatever that new thing is, it will come with the appropriate level of encryption. Is that a decent strategy? Uh, no.
Is the, is the short answer to that que short answer to that question, right? I mean, all you have to do is look at, uh, typical enterprises and, and, and see, especially ones that have been, uh, around longer than 10 years if you don't forklift your entire application environment. I mean, up until recently when you booked an airline ticket, it still went back to a mainframe.
I was talking to a health insurance company, uh, the other day where a lot of their claims are still processed in cobol. I mean, there's, it, it, it, it, it, it's a, there's no such thing as a forklift. Every organization does rolling upgrades, right?
And, uh, and so that's where, you know, where, where you can make big impacts is looking at the, the roots of the problem, right? Going down to first principles. And I do think many organizations are gonna look at services at the infrastructure layer that will help them to, um, to get larger swaths of their environment ready, both from a data storage encryption perspective.
That's gonna be a big one. Um, you'll already see that in some of the cloud providers. And then obviously also at the network level.
One of the big challenges at the network level is that, um, it's, it, it in many cases may be hardware dependent. So I think what we're gonna see is a lot of software solutions and some innovation in software that will help you upgrade your encryption without necessarily doing a, a hardware forklift. So what is your best advice to folks about how to go about putting together the plan and executing it and getting everybody on board?
Because there's just a lot of moving parts? Yeah, I mean, uh, my advice almost always, and, and you'll hear me say this across a variety of cybersecurity topics, is look at first principles. We have a tendency as organizations to play whack-a-mole.
'cause we're looking at, you know, the, the things that are at the higher levels that's like, Hey, I want to go patch my servers to eliminate the vulnerabilities. Well, you know, if your servers weren't available on the internet, maybe, maybe, maybe they wouldn't be as exploitable, right? Same thing, I think with encryption.
I talked to a lot of companies who say, Hey, we have encryption at the application layer, right? Great. Do you have it on every single application in your entire environment?
And the answer is, I've literally never seen a customer that that can say yes to that, right? So, uh, what's really powerful is when you think about it at lower layers of the stack, when you think about at the network layer, at the storage infrastructure layer, those are places where we can make a really big and broad impact without having to play whack-a-mole. What do you think the government's gonna say about all this?
Will they come up with some mandates? I know that NIST has been involved with creating some of the specifications, but, uh, at what point might someone show up and say thou shout? Yeah, I, I think it's gonna be next year, honestly.
Um, but I don't think it's just gonna be government, right? I mean, typically the way that the government mandates work is that they'll, they'll make rec strong recommendations to industries that they contract with, right? I'd say from an enforcement mechanism perspective, what is really interesting is when the industry specific, um, regulations or, or, uh, you know, uh, standards start to come out, like PCI as an example, uh, I, I strongly suspect in the next 18 months we'll see a modification to PCI that recommends particular encryption standards, uh, really focused on, on, on, on post quantum encryption.
And, uh, and so it'll be the government first and then we'll see the industry specific, uh, industry specific recommendations and standards follow. Do you think the insurance companies are gonna show up and say, Hey, we're not gonna give you that cybersecurity insurance if you don't have fix this encryption issue? Yeah, it's a great question.
I mean, I think, uh, what you typically see with insurance from a cyber insurance perspective is that you have to have your first principles covered. Encryption is always one of those, and as soon as the standards change, it definitely will follow that from a, from an insurance perspective. We've seen that multiple times.
All right. Hey folks, you heard it here. We're upgrading our encryption.
It's just a question of when and how painful and how long. And the sooner you start, the easier it'll be. Hey, Chris, thanks for being on the show.
Thank you, Mike. All right, and back to you guys in the studio. Hello everyone, and welcome to DevOps experience, uh, advancing DevOps for AI Native.
Before I start my talk, I would like to thank Techstrong group for putting together this experience with the industry leaders, practitioners, uh, coming together and advancing the DevOps and AI native concepts for software delivery. I'll start this talk with some news. Uh, for not so good reasons, AI native software made some news lately.
One of the AI native software startup companies promised a no code AI generated software, but reportedly relied heavily upon underpaid human programmers to compensate for AI shortcomings. There was another incident closer to where I live. Uh, a major Canadian airline company lost a court case after a chat bot provided a false information about travel discounts and the court hold accountable the company for that AI generated device.
And lastly, I would also point out one of the catastrophic failures where an AI agent from a reputable, uh, AI native software organization wipes the production database and then lies about it. So, apparently it seems that, um, you know, production grade AI native software is still in making, and companies are having trouble operationalizing AI native software. So, what I would be doing today and discussing about how DevOps practitioners could help in this journey, you know, creation of organization, structure, ai, and I first mindset, the culture curation, which needs to happen, AI native companies, uh, holding them accountable and responsible for, uh, ethics, privacy, uh, responsible ai, and bringing it at, at a core of the software organization alongside the rapid growth, which is happening now.
If you are still, uh, listening to this talk, I would like to introduce myself. I'm the founder for the Canada DevOps Community of Practice here in Canada, which has several chapters. I'm also the producer for, uh, summits Canada, which does micro and macro summits.
One of our greatest event, uh, the DevOps hackathon is coming along in Toronto in November. My call to action is leadership Practices and Communities, and I will further discuss a it in little more detail, but I wanna kind of project, uh, the possible future, uh, with DevOps, uh, when ai, AI native is generating so much of news buzz and hype. So how DevOps practitioner with data machine learning and AI can survive the strong.
So, history reminds us DevOps is a collective journey towards evolution of software practices. And AI enhances predictive and adaptive decision making, which means that we are finding new ways to use data sources that matters for flow, feedback and experimentation. We also are up for collaborating with 19 million developers and coding assistants.
We all know that AI native vulnerability storm is on the cards, so we have to approach it with a lot of consciousness where human and machine interactions matter. We have also seen and advocated for open source projects in the past. The DevOps practitioners have seen the highest performers, uh, for us has been open source projects.
And not to say that it's not only faster but valuable, which matters, right? So we can learn a lot of things from our DevOps journey when we are going to this AI native space. Estimates suggests that leading tech giants will invest approximately $1 trillion on the development of AI in the next five years, which entails that rising demand of for reducing software development cycle cost, as well as accelerating deliveries.
New methods are prospected to be considered in the software delivery process, like new and easy and human friendly or machine friendly way where soft software abstraction layers will be built up. And DevOps market is projected to grow, but it also needs a strategy lift, right? And this facelift will happen through different facets.
Uh, we will talk about it in the next, uh, few slides. So projecting the possible future with DevOps, and I call it often DevOps plus because there will be more practices, more processes, more projects, which will contribute to this journey. I, in the first phase of this talk, I will talk about addressing the operational challenges of AI adoption and how DevOps practitioners could be of help and support here.
We also would talk about all operational design of organizations and future organizations and what DevOps practitioners have to do to align to this change curve. And lastly, I would talk about a RS management policy and governance. So getting started, there are foundational challenges with AI native projects.
Uh, Gartner has predicted 40% of agent a KI pro projects will be canceled by end of 2027, due to cost overrun, unclear business value and inadequate risk controls. More broadly, GATA has also forecasted 85% of AI projects will fail to deliver value to CIOs. So from here we can recognize that the challenge is evolving and the complex expectations we have on the leaders, and how leaders practitioners in this space could take this opportunity and dive deeper into how they could support this transformation of the technology curve.
When we talk about operational AI native, uh, operationalizing AI native software, we are talking about a lot of pressure, which is developing on mid layer and senior level leaders, which who must advocate for AI first culture. It also entails that large scale upskilling, 50 to 60% of jobs should be transformed, and it is pivoting and to be replaced by ai adopting new technology. Experimenting with novel approaches is not a choice anymore.
Authenticity, transparency, ethical AI use will be essential for ma maintaining that mainstream trust. And it is needless to say that rapidly evolving cybersecurity space is also putting us into a lot of challenges. Now, when you see this urgency and you spot this urgency, the leaders and the practitioners have started to take charge of the situation, and they've started to act upon, right?
So individuals with let's say, limited AI knowledge, you see that they are overestimating the ability to, to prompt and interpret AI systems leading to poor decision making. And how, you know, ai, uh, outputs should be assessed or, you know, put in, uh, practice underestimation of AI risk, underestimation of AI limitations. Uh, users are often overlooking that ai, especially large language models, lack true understanding, judgment, common sense, and IO is also amplifying this effect by readily making, uh, available the AI generated answers, for example, AI generated code, for example, to create an inclusion of expertise and knowledge.
So you will see a lot of sprawl in the AI space as tsunami of shadow AI applications initiatives, which also brings a lot of challenge on the table. And it leads us to this thought process that behind every effective AI deployment lies leadership that plans innovation with disciplined execution. It is important to understand the strategic alignment of, you know, the initiatives, the key initiatives which we are putting in place through the executive management.
But we also need practitioners like DevOps practitioners. We need a strategy for cloud platforms. We need strategy to architect these solutions at a very, very fundamental level.
We also would like to have, um, intelligence feedback loop, uh, associated with all this and the data pipeline to ensure that we need to be more conscious about the full stack approach of execution, not only, uh, curating strategic key initiatives and just leaving people alone. Another aspect is that what we have seen from the past is that technology alone is never enough for two productivity. So if as a leader, as a practitioner, you believe that AI is savior for everything, I think you got to think through a little bit, because I'll give you some points to ponder upon.
When electricity came in, it required factory redesign, not just replacing steam engines. When computers came in, software, digital infrastructure, new workflows, everything was needed. Upskilling of people was an essential element in AI world.
It demands data pipeline, governance, new skill, cultural rapture. These all things will complement to the productivity curve, which you will be ensuring in your organization. I will also advocate that Darwin's story of evolution needs not to be limited to the world of organic beings, but can be extended to the world of innovation, which means survival of the fittest, which means that, you know, understanding the AI Esker is crucial strategic decision making.
An innovation flies in the fact that how well you forecast as an organization forecasting future progress, performance, market saturation, all depends on you, how strong feedback loops you have created. Right? And there where DevOp practitioners could be of help.
Also, when you are making investment decisions, identify which part of the AI esker specific AI technology can help your business, right? So experimentation in a streamlined way through, you know, curating a list of strategic initiatives. Navigating multiple s-curves means that, you know, how do you keep your flow intact?
Why you are actually having an intersection of s-curves such as scur for large language models and technology obstacles have to be overcome to ensure that you reach that productivity of, uh, from through the technologies life cycle. And this is, again, um, complex mix of, uh, things which are happening as we speak. Setting clear goals and objectives is the call of the hour.
The key components of an AI native s disruptor is how well your project objectives, product objectives, and financial objectives intersect. So how do you link your AI product OKR to your financial KRI will also ensure, ensure that we have brought in DevOps practitioners in play here. Evolution of practices, for example, operational practices like SLO management, AI assistance, law monitoring analysis, continuous resilience automation.
All this is needed to ensure that we curate this list from a strategic point of view, and we should have feedback loops, like, how do we optimize delivery, including managing progressive delivery or observatory driven development, or managing AI generated, right? So all this all in all will make an organization more resilient towards the pitfalls of, uh, AI native organization and AI native software, uh, you know, processes. And how do you also take the next step is again, trying to understand the organization structure.
Allocate your 20% of, for example, your time, r and d time to experimental projects and achieving that faster deployment through AI would only happen when organization design facilitates that, uh, part. So I will also kind of, uh, articulate this in a way that it makes sense to upskilling and reskilling. Uh, you know, 85 billion jobs, uh, might go unfilled fulfilled by 2030 because there aren't enough skilled people to take them according to contrary.
5 trillion in unre unrealized revenue. So skill shortage is on the card. There are many aspects of learning, many people are getting behind due to lack of time.
What self development in reducing a major skill shortage cost is another factor. More and more services and applications. And, you know, there is a plethora and of tools, which, you know, it makes us hard to believe that we only started with this evolution of practices a decade ago or something, and complexity.
It is inevitable that the complexity is increasing with rapid adoption. There are new risk, new operational overhead and new cognitive load on practitioners, which takes us to the next slide. That what, as a practitioner, as a leader, as an individual, as a team player, what could I do to dify this?
And there are four components of this, um, operational design for future organization. And the four components is what it means to me as an individual practitioner or leader. What it means to my team, what it means to my enterprise, and how should I take this in a holistic way.
So think about building your own AI native skill radar. How, what kind of new skills are needed in for this new era? What new teammates you'll have, like AI code companions, how do you build an async first culture to AI collaboration?
What kind of AI native workflows would be needed? And sustainability with ethics and responsible AI and getting together all this, how DevOps could help you in this process or journey in the next five years. As a software developer, what we will see is that employees who are generalist with broad skills, but know contemporary areas like cyber or cloud or subject matter expertise, are likely to struggle for those who excel in programming.
Individual programming languages, for example, currently in use may also feel a sense, uh, sense of insecurity because there is a need to understand multiple languages, which will grow quickly, right? And at the moment, it skills which are nearing end of lifecycle, like manual testing, for example, ui, ux, uh, designing, cross-functional team members can automate these tests, testing skills. And cloud-based engineers would also be able to do a lot of these ui, ux, and ui, you know, onboarding and capability.
So these are things which people should kind of look at from upskilling or re-skilling perspective. Now, how should I start this journey? How should I take the first step?
What I would recommend is that ask yourself three questions. What have your roles achieved in the past? How well your role or yourself, your team have collaborated with others, and which new skill is needed for upskilling?
And if you think about this, you will come across organizations which aim to integrate AI into your business models, creating new values, and giving rise to next generation of software companies. AI is with is a speed and a scale multiplier for software features. So what we need is to learn how to collaborate with generative ai, develop new AI applications, design, and also curate, and, uh, do content for that.
Empathy and ethics plays an important role. Onboarding of responsible AI is essential. And then your left hand side of the brain, which is more problem solving, critical thinking, mathematical and computational knowledge system thinking, communicating with intent, and then so on and so forth.
The roles which we probably would see in the upcoming years is data scientists, machine learning engineers, data engineers, AI specialists, AI researchers, prompt engineers, and AI programmers. And also you will see a rise of code companions in your team. So imagine if you te treat each of your AI agent or a companion as a new high potential employee, how would you you treat them?
You would provide comprehensive onboarding material, keep them with necessary tools, create a safe space for trial and error, and establish effective communication channels. Similarly, if you see from a software engineering practice perspective, um, there are more and more AI engineers and AI programmers in the loop. So if you see on this slide, you'll have a lot of perspective on AI engineer, full stack engineer, ML researchers, ML engineers, data scientists, research engineers.
These are new functions roles, and probably these are areas of upskilling, which can give you some ideas of how can you take that journey and take that steps in the right direction. Now, when we talk about AI ready enterprises, so what could happen is that we would be relying on a lot of AI companions and these AI companions, uh, um, need AI friendly interfaces for team collaboration. We also need fast and slow thinking agents for realtime, re realtime reasons, and for research interactions, comprehensive knowledge base, documentation, testing, quality checks, data pipelines.
All of this would be needed to ensure that we have foundational capabilities to design the next generation future organization, which are more AI centric. We also need, uh, new skills for new era, as we have mentioned earlier, recognizing patterns of bias, risk and failure, understanding how knowledge is shaped by who it creates and what it ends up excising, ethical and aesthetic judgment, not just functional adequacy. When we talk about organizations, uh, and individuals, every organization and every individual will have unique journey, right?
So there is no one size fits all, uh, you know, organization design. At this point in time, at least, we could have some AI development team, which can help, uh, develop, customize, maintain AI models, especially n lms for various software development tasks. We could have data management team, which can manage the data required for creating and refining AI models.
We, uh, will see a rise of AI integration testing team, which integrates AI solutions into existing processes and systems, and ensuring their functionality and reliability. Also, we probably will see a rise in AI ethics and compliance team ensuring the ethical, ethical development and deployment of AI systems. And moving forward, we would also see, um, a rise of AI agents, an agent which is backed by generic or specialized lang, a along large language model like software development agents, product management agents, operation support agents, or research agents.
So, you know, the initial investment in developing training and integrating these AI systems for software development organizations can be substantial. And also maintaining updating or training these AI systems to keep them effective can also incur significant cost. Organizations should be conscious of the fact that all this is coming, uh, along with the AI hype, and we have to be prepared to deal with these AI led micro change initiatives, uh, at the core.
Now, I will also talk about scaling, um, ai, risk management policy and governance in a nutshell. If I could do a polling question at this point in time that, what are some of the perceived risks for AI native organizations? Is it inadequate or misleading output due to limited training data, legal or ethical implement, uh, implications in copyright, infringements, cybersecurity, vulnerabilities in managing large data sets, impact of job roles, or workforce displacement, or all of the above?
I'll give you a second to guess, but my answer to this would be e all of the above. With the rise of, you know, organized hackers with the rise of, uh, AI enabled, uh, you know, malicious applications with the rise of kiosk gpt and black hat AI tools and wool gpt, I think it is evident that we got to be more serious to implement a culture where, you know, we have more, uh, capacity to deal with security posture. We get to get to have more, uh, people as well as more investment into bias audits.
For example, uh, AI model explainability scores are important, and if organizations don't lead with governance, they will fall short. They will adopt AI in silos with limited scale or return of investment. Creating prioritization, uh, for, uh, you know, not creating prioritization for these results would result into fragmented investment, uh, being reactive on addressing risk and damaging both reputation and finances.
And lastly, repeat mistakes across views throughout the AI adoption lifecycle, which can cost multi, uh, afford, right? I have written a paper on this, uh, aspect, responsible ai, which is, uh, to be published, uh, through IEEE conference. So if you want to get hold of more information around responsible AI and how to instrument responsible ai, that's, uh, that is the paper for you.
But operationalizing, uh, responsible AI through processes and policies in an organization, governance structure is not a choice anymore. Uh, to create, uh, standards for risk steering or monitoring throughout the project lifecycle is of utmost important. Ensuring that checks and controls are in place to validate that all use cases are compliant to an organization or as legislative regulations, which are also changing as we speak, driving security reviews for gen AI platform technical architecture, for example, uh, audits for those architectures and recommending organization definitions of ai, current risk tier and AI security and privacy, uh, racing metrics and publishing it and making, uh, advocacy for it is also at Atmos importance.
Now, I come to the last section of this presentation where we also have to plan for next 18 to 24 months. So we have touched base upon a lot of aspects like, um, upskilling operating model change management. How do we handle AI led change management?
How do you build, uh, RACI metrics where, you know, AI is your core companion? Identifying specific use cases or strategic use cases as your first pivot points, right? Tools and technology.
What kind of investment is needed in partners in ecosystem building in tech stack? How do we advocate, uh, the strategy for AI native, uh, tool stack and technology onboarding from an executive to, uh, you know, uh, cloud platform to DevOps practitioners is of utmost important. And lastly, I would say risk management.
Uh, I mean, we all know that vulnerabilities are getting introduced as we speak. So how do we navigate this vulnerable AI vulnerability strong? And how do we put in place, uh, frameworks, best practices, industry standards, uh, advocacy groups, open source initiatives where we bring ahead and control on AI led, you know, um, or AI integrated, uh, vulnerabilities.
With that, I would come to an end of this session. I would like, uh, to also introduce you to a few of the other initiatives where you can follow me. Of course, I write for, uh, ML Khan Magazine.
I also have published my blogs on Dev meo. Um, you can join me for my next workshop, uh, on the DevOps Con Conference. And one of the biggest events, which we are, uh, doing in November is DevOps for Gene AI Hackathon, which is coming in Toronto, November 3rd.
Stay tuned. Our good old friend John Willis, will be with us along with industry practitioners, teams and communities. With that, uh, I hand it over back to the DevOps Experience Conference, and if you have questions, uh, please do connect with me off online or, uh, through my LinkedIn profile.
Thank you very much. Network automation is the future of network operations, but what standards are you adhering to? Are you even aware of what the standards are in this episode of the tech field?
A podcast automation Needs standards. Welcome to the Tech Field Day podcast, where we bring together a group of influential IT experts from across enterprise IT to discuss hot topics in the technology industry. This podcast is brought to you by Tech Field Day, which is a part of the Futureum Group, and is often recorded in association with one of our events.
We're here today at Networking Field Day, and we'll be talking about automation and the need for standards. But before we do that, I wanna have our guests introduce themselves so you know who they are. Starting with Denise.
Hi, Denise Donahue, uh, network architect, technical author, all around Network Geek. I am, uh, Steve Plucka, network architect with, uh, DQE communications out of Pittsburgh, and also General Geek. I'm Kevin Myers, also network architect, um, and, uh, general Network Geek, IPV six, geek, uh, service provider geek, um, routing and switching geek.
Alright, and of course, I am Tom Hollingsworth event Lead for all things related to networking here at Tech Field Day. Let's jump into this episode. No doubt, you have probably heard about the importance of network automation, whether you are starting that journey yourself, you're finding yourself in the middle of a long project, or in some cases even wrapping up this, uh, amazing thing that you now have done to make your life so much easier.
But were you following the standards the, the whole time? Did you know that there were standards? Did you lie to me and tell me that you thought there were standards?
Because there aren't in this episode of the Tech Field Day podcast, the premise is automation needs standards. So let's talk about this for a minute, because this was something, as soon as you guys said that this was the topic we wanted to talk about, I was immediately on it because the last time that I checked, there is no written standard for automation. There are a lot of standard protocols and ideas that we use, but I find it funny that when you start talking about, well, automation, usually you get barrage with a whole bunch of questions of, you know, which scripting languages are you gonna use and which platforms are you gonna use?
And is everything gonna be written in Camel case or Pascal case? Like, like you have to figure things out here. What is it about network automation that lends us to being, for lack of a better term, creative with the way we implement it?
Well, I think a, a big part of what it is, is this development all along. We, we've been CLI jockeys forever. Mm-hmm.
Uh, and now that we're getting into the move of automation, the first thing we did was create a couple of particular protocols, whether you're using YAML or whatever standard in that sense that you're having to communicate with the devices. So now we've got enough of those going that we're now doing scripting and, uh, each vendor is individually creating a platform to do their equipment on. And then there's a handful of, uh, multi-vendor companies out there that are picking and choosing which platforms they're gonna support.
So I think we finally reached the critical mass where enough of this automation is happening and enough of the bits and pieces and tools are there that we need to get together as an overall community and create that standard on this is, this is how it should be abstracted, this is how it should, should work, and figure out who the hell's responsible for that. Yeah, it reminds me of the early days of the, of networking in the internet where there was just like, my protocol and your protocol and this, you know, and just everybody had their own idea about how things were gonna go, how data was gonna be structured and everything. But then there was, at that point there was a small enough community 'cause small enough group doing this that they could get together and make and create standards and fight over who was gonna win.
Mm-hmm. I think the challenge that I see is that it, where we're at right now with automation is the vendors really drive automation. Whatever vendor it is that you're choosing really drives the automation.
And yes, there are open tools and there are open frameworks, but to the point of the podcast with the person sitting down and figuring out how they're going to assemble those is doing it in whatever way they think is best. And a lot of it comes from the coding community, which not all of those practices translate over to network engineering. And so I think we do have, there are some standards that might provide some guidance, industrial control world, um, world of energy.
They've had some automation standards around for a while, but it does take a bit to adapt those into networking things like the Purdue model. I think those are places to start. If we think about, let's take automation that's specific to networking and standards, and what does that look like for an enterprise network?
What does that look like for a service provider network and figuring out those best practices? Well, I think one of the challenges that you mentioned there is that the two industries that kind of have those are also very highly regulated. And so a lot of the standards that kind of come out of this are not so much, this is the best way to do it, as much as if you don't do it and prevent these outcomes, we are going to sue you out of existence or fine you until the pain stops.
Like, you know, for example, um, something like P-C-I-D-S-S drives a lot of the way that we design certain kinds of networks because we can't have certain things interacting with, with each other, or we have to have certain retention policies and things like that. And as far as I know, there's no kind of guidance in the industry right now of let's just say automating healthcare facility. Like, you know, if, if HIPAA applies or if, if this kind of patient protection thing applies, that's gonna direct the way that you do things and create kind of a defacto standard, if not a Deger standard.
But again, that's vendor specific. You look at HIPAA and you've got the different, um, medical records vendors, you know, the, the, um, and they're, they're doing their own automation within that, which then they've extended on into the hospitals. Um, so then does that gonna be the same?
How's that gonna carry over to all the other companies, types of enterprises? And there's a difference between policy and standard. Mm-hmm.
HIPAA is the overall umbrella policy that's saying what you can and cannot do. What we need as engineers and as vendors is, uh, a way to implement that policy and create the workflows necessary. So far we've been concentrating on the individual task level things, uh, and we need to step up to, to workflows and, and processes that Yeah.
They have to be driven by some overall arching policy Or business logic, Uh mm-hmm. Yeah. On how, how this is gonna work and the categories you're talking about, right.
As well. But, but that needs to be a team coming together and we need a, a group that does that, that doesn't bi solely benefit, it can't be the vendor, right? Mm-hmm.
It has to be the, the engineering teams in general as a, as a group. Yeah. Manners kind of comes to mind.
Ma NRS mutually assured norms for routing security, which is basically a lot of carriers are involved in it, CDNs, and it's a neutral organization that says, these are the policies that we feel that we need to implement for routing security on the internet and in the DFZ. And then they take that a step further and have technical recommendations that if you're Cisco, if you're Juniper, if you're Nokia, here's how you implement that policy. Here are recommended guidelines.
And that's not quite a standard to Steve's point. But that to me is kind of along the lines of where you might need to go with automation is recognize the problem, define the problem, and then get down to those, those technical details of how do you implement this? And in my opinion, the challenge is automation is coming outta the world of coding.
And what does coding love to do? They love to jump to the newest programming language whenever there's a hot new language. I mean, how many have we been through?
And I think that's the biggest challenge we'll have as network engineers, is we look at standards and protocols over a very long term arch. The world of coding is always jumping to the next new thing. So how do you, how do you, how do you bridge that divide of how do we pick something, how do we pick a winner in the world of automation to create standards around that isn't going to get left in the dust of the world of coding?
You don't Wanna run my pearl Script? I sure Pascal, I wanna go. Yeah, I want do it in Pascal.
Yeah. Alright. So who's gonna do this?
What stand, what standards body are we going? Well, before we get to that point, I think that that something that Kevin brought up that, that is very germane to your point, is there are groups that are kind of creating these informal agreements, manners, for example, it's in the name right? Mutually, um, Assured norms for routing security, Mutually issued norms, not requirements, not restrictions, norms.
And we see this a lot in our society, right? Like something like holding the door open for people. That's a norm in certain spots in, in the us but it's not a rule, it's not a requirement.
And the difference is when you cross that line of a group of people getting together as an industry organization saying, this is how we're gonna do things, doesn't have any weight, because I can choose to leave that organization whenever I want and not abide by those rules. But to your point, there are standards bodies out there Yeah. That provide guidance on how things are gonna be implemented in the us The two biggest ones are IEE in the IETF.
I would also add ISO as a global standards body. Um, I, I would recite chapter and verse with iso, but nobody does that. These groups are formed from people who analyze the problem, decide on a, a method for recommended implementation.
It is voted on, it is debated, it is argued, it is voted on some more and then it is released. Why have we not seen anything from any of these standards bodies about this yet? You'd have to ask them.
I would, I would say there Hasn't been a demand. The only one that I've worked with is IETF and I, I've done, you know, a little bit of work in the ITF and so that's the one I'm probably the most familiar with and it's also the most open. Mm-hmm.
I think mm-hmm. Yeah. ISO and IEE are, you know, they're solving real problems, very complicated problems.
But that is usually, you know, there's someone is funding that to go to go do and solve that problem. And again, that's a long arc, I think programming and coding and automation and are all there together and they iterate very quickly. So I think you have to start somewhere, whether it's in the IETF and we talk about the IETF is the right place, or whether it's a new organization.
I think you've gotta look at that and say, how do you create an organization that can help to define those, you know, first maybe norms, then best practices, and then ultimately standards. And with the IETF and the ITF is very, it's very protocol and operation specific. So I mean, I could see it going there, but I could also see challenges in, you know, in getting it through there.
Which working group does it go into? How does that, you know, I see pros and cons in it going into ITF and, and. Mm-hmm.
You bring up a really good point here, and I'll reference everyone's favorite XKCD comic from Randall Monroe. There are seven ways to solve this problem. I know I'll create a standard that encompasses them all.
There are now eight ways to solve this problem. Yeah, exactly. Like one of the, and one of the things that we've seen over the years, like IIEE is very focused on protocols, right?
That's where we came up with the ethernet. That's where we came up with wifi and a lot of those other things. But with IETF, I feel like there is a lot of debating things until they're dead.
And that's the way they eventually come to a consensus is, okay, are we done arguing about this? Does anybody still care? Perfect.
Now it's a standard because these are the, the remaining people. And we saw that with TRILL and SPB that kind of became pseudo standards, but didn't really, and TRILL is actually a really good example of this because the, the definition of what was going to be ITF standard TRILL never developed because the vendors were so focused on making sure that their version was the standard. And I'll go out on a limb here and say that one of the reasons why I think that that happened was because there was no bad guy.
What do you mean Power over ethernet and trunking on switches? 3 af? Because Cisco had a competing standard and everybody in the industry mm-hmm.
Lined up and said, we wanna do the exact opposite of what they're doing. Oh, you want us to wrap the switch in a, in a, a trunk tag? No, we want to put it in the header.
'cause it's what you're not doing. Oh, you want to carry power over 5, 6, 7, and eight on the ethernet wires? No, no, no.
We want to use 1, 2, 3, and six. Because you're not doing that in a way. Standards sometimes are driven as a big middle finger to a competitor.
So maybe what we need is Red Hat to come out and say, oh, well we're gonna use Ansible standard network automation, and the rest of the industry might wake up and go, you know what f you we're gonna go do things our way. You know, I I think that you two have got, have, have some excellent points that kind of could be drawn together. Like the why part, what's driving it.
You talked about funding, um, you talked about, um, you know, vendor pushing. Um, you talked, you talked about, okay, you know, the, the reason for doing it just as a, a reaction, also a need for it. And obviously there's a need for anybody who does much work in this and is dealing with multiple vendors of equipment, multiple vendors of software.
There's definitely a need. Um, how are you going to get someone to do it? Is the thing That's the big challenge.
'cause I think you look at, um, you know, does it, we've taken the existing, you know, protocols, if you will, and we've take, we have a p we have APIs, we have mm-hmm restconf, NETCONF, all of different things. And, and the companies that are out there have taken the existing available protocols, created a few new, and that's what we use. And then how you cobble that together is up to you.
So I think the real question is, is does automation, you know, whether that's, you know, automation control, that's, you know, pushing configuration, um, you know, telemetry related to automation. Anything that's in the realm of automation, does that become an 800 series protocol or is it gonna be like an RFC? Is it a standard?
I think that's probably the first problem you've gotta sit down and solve is, are we happy with what exists? Or do we want to build something new that encompasses the entire world of automation? And the question then becomes, if you're gonna build something new, how do you get people to sign onto it?
Exactly. Mm-hmm. 'cause one of the things about a standards body is, is that when the standard is defined, everybody has to play by that standard, or it's not a standard.
Mm-hmm. And you run into these situations where it's like, okay, there are four people that all believe that this is the way to do it. We have to pick one of these solutions, especially if they're not able to be merged.
Right? Like, you know, if the, if these two people have one that's pretty similar, we could probably merge them together. And then 50% of the use cases are kind of done like that.
But it goes back to, well, if I don't like your solution to that, I'm not gonna be in part involved in this working group anymore because I'd rather do what I do because it's easier on my developers. My customers prefer it this way. And in some cases, if I can get enough people to sign on onto what I'm doing, I become the defacto standard anyway.
Hmm. Do your customers prefer it that way though? Or is it just that they have to take it?
Well, that, that's part of it is if you don't know there's any other options out there, my way is the highway. Yeah. Well, the, the, the customers are drowning right now.
We're, we're just dealing with the day-to-day problem on how to keep hundreds or thousands of devices doing what they need to do and changing the way they need to change in a, in a reliable way with what we have in front of us. So, uh, it's, it's time to take that step back and, and to, and to try to figure out how, um, I think the big difference between this, what, what we, the problem we have here and, and the normal standards is that it involves a, a workflow, um, type operation. It's, I, I guess the closest thing we've had previous to this is MPLS, where there's a whole lot of steps that have to happen to create this end to end, right?
And everybody on the path has to play nice with each other for it to, to work in, um, in a vendor, in a, in a mixed vendor environment. But even that is not as complicated as what we have now because we need to, you know, you know, we need to do, you know, pre-check change, post check, you know, rollback if necessary. Uh, you know, in the post check this, the, the whole workflow access this that we sort of inherited from the developer world is the, uh, is the part that is alien to, you know, the average network engineer.
Yeah. And that's the biggest challenge I see is I don't, you know, I, I haven't written code and anger since the nineties 'cause I've decided to be a network engineer and I didn't want to do coding. But, you know, now that we're in that world where code is very much a part of, of network engineering, the challenge that I see is I either go to a vendor and I have to go buy something and they're gonna tell me how to do it.
Or I've gotta sit there on their gear only exactly right on their gear only in most cases. Or I've gotta go out and look at what the development community is doing and figure out, okay, which of the thousand different ways that somebody has figured out how to build this framework, do I go, it's the build or buy. That's the challenge that you always have mm-hmm.
As an organization is do I buy it or do I build it? And if I build it, what is my am I doing? Because as a network engineer, let's do a protocol.
Okay, well what they're like, you know, half a dozen routing protocols, there's layer two, there's a, there's a sit limit, uh, number of protocols. You know, you're dealing within boundaries that are pretty well defined when you look at the protocol stacks of network engineering. But you look in the world of automation and coding and it's almost limitless in what you can do and what you can build.
And that's the sea of uncertainty that I think network engineers or people, especially people at the network architecture level, you know, where we came up in the old world of networking. And the newer engineers are, you know, nothing against them at all. It's just they're coming up in a different world of learning coding and network engineering.
We're looking back at a long career of here's why we did things this way, here's the choices we make. You've been through enough networks to realize that the choices that you made maybe weren't the right ones, and you do it differently mm-hmm. Later.
So how do you take that knowledge and work with the network engineers that are coming into the field now to help them to also help inform the automation? And I think you need standards and frameworks for that too, as as to how to make and evaluate those decisions. The other thing that I think we need to be aware of is the trap that we can fall into by turning this over to a standards body.
This is something that's happened in the wifi world recently where we're going to introduce a standard and we're gonna make sure that everybody follows it, but we need everybody to buy in on the standard. So what we're gonna do is we're gonna take some of the hairier pieces of it and we're gonna make them optional. And that is something that has happened quite a bit in the last couple of releases of wifi, wifi succeed, and wifi seven, where some of the things that make it good are optional and don't have to be, uh, implemented by you if you don't have the technology or the desire to make it happen.
And so what we're left with is a slightly better than the last one standard where we hope that everybody plays nice on these other pieces. How can we prevent a standards body from coming in to restore order and ultimately making things worse? Because the things that we need to restore order are left for optional, because, well, if you don't make this optional, I won't sign on to the standard.
Ouch. Maybe that's why there hasn't been a standards body. And that kind of comes back to that whole how do you get everybody in the room to compromise?
And for the purposes of this podcast, everybody knows that the definition of compromise is when nobody gets what they want. Mm-hmm. Well, the other challenge you have is if you were, let's say you were to take this outside of something like the IET after, you know, ISO or IEE, how long does that take to, to build that organization, to get people that want to contribute?
You know, when are you gonna, is it gonna be years before you see a work product that's, you know, that is relevant and, and going to help you? Because I think that's a whole other challenge of if you do create another standards body, what's the, you know, the ramp up of that is, is gonna be a lot. I'm sure they're working on finalizing the standards for building wooden sailing ships in the ISO like this year.
So, you know, there are only like nine centuries behind at this point. You only need 90 bucks to access the standard. Yeah, Exactly.
And, and that's the other problem too, that a lot of times you need a lot of resources. You need to create working groups, you need to create, uh, people, you need chairs, you need folks who have disposable time. And then for people to be able to access the standard, you have to have resources to, to do that.
Either they're people who have access or that you pay for it. And, and then you create this rolling problem of, well, if I don't, if I can't see the standard, I'm not gonna follow the standard 'cause I can't follow a document that doesn't exist. And how do you prevent that?
I mean, we, we bag on the I-E-T-F-A lot because it feels like there, it's just basically an excuse for people to fly around and argue with each other. But they do get things done pretty quickly because somehow all of that arguing eventually does lead to some kind of an eventual consensus. 'cause I think they realize they're not gonna get any more out of it than just the, the heated discussion.
They're not like an a giant organization like ISO or IEE that kind of direct things, but use that as a way to fund the organization. Yeah. They were, they, they're definitely the right structure.
Um, for this type of thing. I'm just not sure if the, there's enough, there may need more participants. The half of the half or maybe even more than half of the right people and companies are involved in the IETF.
But adding in this additional layer of, of, of, um, of the workflow aspect of it and understanding the history of it too. You know, saving those configurations, which we've never really historically done. You know, it's a, to compare to, you know, the, the pre and post check access of the, of the workflow.
These are not traditionally things that network equipment vendors are, are good at or know They're, they're good at getting you a box that does a thing. Yeah. Yeah.
Yeah. And, and that's the, that's sort of the missing piece of the IATF, uh, the, the fundamental section I think has worked so far. You know, we have, you know, the, the neck off, the yaml, the, all the options available for the communications and the, uh, and the saving.
What we don't have is the, uh, is the discipline of the, of the programming community in the workflow as aspects of this and the history aspects of it so that we can see what changed and, Um, yeah, And, and why I think it's also the, um, it's getting the, there's very few people that have a very, very long deep network engineering background and also have the coding background that understand how those worlds come together. So I, I worked on a project one time where a coder was off trying to solve a problem, and he spent like a month or two writing code. And what he had we'd written was Radius, he had built Radius.
And I said, you know, there's a protocol on the router that already does what you're, what what you just spent two months writing. He is like, what is it? It's called Radius.
You just go turn it on. And that's where you can have a brilliant programmer that understands coding and even understands trying to solve a problem. But if they don't have the domain knowledge of what is possible in network engineering, then you may spend, you know, you're sitting there beating your head against the wall solving a problem that's already solved.
And I think that's just as important as getting a standards body and defining the standards is what already exists that we want to leverage so that we're not reinventing the wheel. Mm-hmm. Mm-hmm.
And, and ultimately I think that that is where we're at is that as networking folks, we have worked very hard to create these small areas where we can standardize on certain things. Whether they're things like Radius or they're manners type, uh, assured norms. What we need is a neutral third party to step in and effectively get everybody's ego out of the conversation.
This is how things are gonna be done. I get that you're not the way, that's not the way you do it, but this is how we are going to do it going forward. And if you can have that happen with someone who has the personality to pull it off in an industry where people are going to have to agree to disagree about certain things and put their competitive advantage aside for the betterment of society, I think what you'll ultimately find is that we can standardize automation if we're willing to put in the effort.
That will just about do it for this episode of the Tech Field Day podcast. We want to thank everyone out there for following along. com/podcast.
You can also find show notes and bios of all of our guests. com for more information on our upcoming events. com for more information about the future on group.
We'll see you soon. Hey everybody, we've seen the future of DevOps in the age of ai. Maybe you're watching Textron Gang, we'll be back in a minute.
Welcome back everybody. We've got our usual assemblage of smart folks on the panels today. And starting off with Gina Rosenthal, Fred Wilmot, John Schwartz, and we have a new member, Barbara Russ.
And I guess I kind of wanna introduce Barbara A. Little bit 'cause everybody else has been on the show multiple times, but Barbara, tell us a little bit about yourself. Sure.
Uh, I'll start with, my name is Fennell's Barbara Rose. Uh, a lot of people get that wrong. No worries.
Uh, I, I run Trailhead Communications, a consultancy that helps companies navigate the human side of AI adoption. And my background is in, uh, the tech industry. I've been in communications change management culture work for the last 25 years.
So super excited about this AI revolution. All right. Speaking of which, Alan and I were at a show up in Brooklyn this week.
It was hosted by an outfit called Tesla. And they were talking about well, specifications for AI agents, and let me do my best here to kind of explain what's going on. But part of the issue with AI agents is, well, they give you superpowers, but they're also notoriously unreliable.
So people are creating specification files, which are essentially documents that tell the AI agent very narrowly what it's supposed to be doing. And then ultimately, once you get that AI more focused on a particular task, you can start daisy chaining these things to automate processes. And folks are talking about doing that within the context of DevOps workflows, because we need these things to be more reliable, right?
We can't have a bunch of AI agents running around just randomly generating some code that may be, I don't know, has a bunch of vulnerabilities in it, or is just frankly, too verbose to run and gets kicked back by the software engineering team. Fred, you've been floating around on DevOps for a while. Is this the right approach?
I mean, on the one hand, I kinda like the idea. On the other hand, I'm like a little bit concerned about, well, how are we gonna manage all these files? Tesla says they're gonna create a platform for this, but if you've been around Kubernetes, you are familiar with the phrase wall of YAML files.
So are we just gonna get more walls? That's a good question. I, I'm kind of thinking about it like it's the next, uh, it's like the US bomb for, for argentic, uh, workflows.
The philosophy that, you know, we should probably think about having a persistent record of intent. We kind of have this, uh, to make a standard for it though is an interesting concept, I think given the number of, of different variants and files. And so when you want to have, you know, a agenta communication across using, uh, a to a or, or what have you, uh, MCP servers need to communicate the same, uh, actual effects as, as all of the, uh, agents start to collaborate outside of your sort of wall of trust or your wall of yaml, right?
As you put it. The, the philosophy is really about how to understand whether or not that's going to improve things. So on the one hand, I would say, look, w we already have some solutions to this type of a thing, but on the other hand, I think they're, uh, it, it's a funded company and there's a, you know, there's a large amount of funding behind it.
So my, my argument against that would be, look, uh, we had an ai, uh, uh, cyber, uh, cyber challenge that at, uh, DEFCON this last year. And those winners open sourced all of their frameworks. Those frameworks included the opportunity in a, uh, to, to find, disclose, uh, patch and deploy, uh, vulnerabilities in software, uh, at those types of things.
The way that works best is when that's open sourced, how a standardized process works from a a private company is a question mark for me. So, but there's a need for it. Uh, is that the greatest need of all?
No, I don't think so. Gina, you, you have some experience in the land of operations among other skills and expertise, but as you kinda look at this, what's, what's your initial reaction? Well, my initial reaction, um, was isn't it just sounds like it's agent driven infrastructure as code.
'cause you're looking to put, and it sounds a lot like what we used to do with finish files for, um, for Jumpstart and Kickstart, right? So you wanted to do a certain thing. If an agent is just a bundle of, it's just a bot that's assigned a specific task to go do, but you wanna make sure that tasks stay, you wanna be able to give that bot, um, uh, uh, a, a a space.
We want you to, we are gonna declare what you're gonna go do and all the other bots you've gotta go talk to. How is this not agent driven infrastructure as code? And then my second thought was just like Fred was saying, we're already doing, this is my mantra, this is just an extension of computer science.
We should be getting better not trying to reinvent the wheel. So I don't know how we have, how we get the communities to come together, right? Like, yes, you're thinking along the right track, you're further along than we were when we had no tools.
So when, how do we accelerate it by showing you how we figured out how to do it already? Mm-hmm. I think when I looked at it, it seemed to me we were coming up with a way to use code to make up for the limitations of the AI agents.
They're just got some fundamental problems and we need to figure out how to manage that. But what they're saying is that we need to share these specification files among developers so that we don't have all create the same ones over and over again. And then that leads to this wall that I was talking about.
Um, Barbara, welcome to the show. But I guess, you know, it's pretty clear that we're gonna have some sort of leadership issue here in terms of how we manage this process because there seems to be a disconnect emerging between the developers that are in love with AI coding tools and the software engineers that are responsible for actually deploying this stuff, and maybe we need some more adult supervision. What do you think?
I think adult supervision is a great idea. Um, and I, I think collaboration is at the heart of this, um, you know, a across all kinds of industries and use cases. I think everyone's experimenting with AI and they're doing it individually and in their own way, and they're finding the things that work for them.
Um, but what needs to happen is we need to create communities of practice, uh, who are coming together and sharing what they're experimenting with, sharing what's working, what's not, creating a culture of experimentation. Um, and then from there, creating guardrails and systems and consistent use of, of these tools. Um, and I, I think part of why this is happening in isolation is because we have that leadership gap where, um, leaders aren't, aren't talking about what the real vision for AI is and how, um, how the culture needs to shift in this, this new reality.
Mm-hmm. You know, Gina, to Barbara's point, most of the IT leaders that I have met usually have spent some time in the trenches and have some experience in this space. And yet, once they get promoted, something seems to happen.
They seem to get removed, they're divorced from the actual workflow and the things they're being done, and suddenly, you know, they're kinda reading the latest report in the Wall Street Journal and making policy decisions and what happens and how do we kind of prevent that from happening? Yeah, that's a very interesting question, right? Because we all know what happens.
You get busy doing what the big bo what you're supposed to be doing, going between the big bosses and the actual technologist and, and making the company a profit and keeping everybody out of trouble. I, I love the term community of practice. I think think that's a big part of it because I think the technical leaders who go on to, you know, these, um, more important roles, managing people and managing processes, um, need to be part of that community of practice.
And I think there's a great, just in tech in general, there's a huge, um, a void of that anymore because everybody is on the hype monster. And so it's very hard to find real information about here's how you go from A to z. I would love to be able to see Tesla tell me, yeah, this is just, um, agent driven infrastructures code, it's the next generation, and then all of a sudden maybe you can tie, start teasing those communities together and building a community of practice that gives people on the top level that don't have the time to look into things and maybe get their hands dirty anymore.
It gives them something that sounds reasonable versus we've figured out a brand new thing. It's a brand new thing. So now we're gonna have to do a brand new thing to manage it when we all know if we've got the experience behind us.
That's not necessarily true. We need to build on the foundations that we've all climbed through and in, in those trenches. So I think it's getting that information to, uh, to the leaders who are technical in a way that is technical and it's not hype driven, or it's not, um, analyst defied, you know, it's actually tied to reality.
And that then they can direct, you know, they can encourage their, um, their teams to go join community of practices that are truly technical to dig in to the details. And the technology leaders can, um, find ways to that to direct them that is tied back to the business. I think we're just missing a, a community of practice to help us get through the hype.
You know, that's interesting that you say that, Gina, 'cause I, when you said A to ZI was thinking of how the way this narrative is unfolding, this hype narrative for now. And I think about Silicon Valley, they present the pie in the sky utopian vision in these announcements. And there are, there are announcements every day, as Mike and I can attest sadly.
And then they, they, so they give us, on one hand they tell us, this is what we're gonna announce. It's bigger, better or faster. It's gonna do everything for you, imaginably automated, and, uh, we'll, we'll give you the end result as well.
But there's no transition. There's, there's nothing in the middle on how do we get there. And I think that keeps occurring.
It keeps rearing its ugly head in this AI agent year of 2025, like the devil in the details. And almost every segment we talk about, and most of the segments we talk about involve that gap or some sort of problem, whether it's security, observability, um, what the agents do, how they're coordinated, and the impact on the people in the middle. So I think we're gonna be hearing a lot more from, like, from Barbara about how we're gonna navigate this A to C journey.
There's a big gap. There's like a Grand Canyon gap between the two sides. I I think actually AI has a lot to do with it, right?
So coming from the product marketing side, if you have given all of your marketing budget to ai, so you lay off all the marketers that have any experience in the industry, and then you give it all to brand new people and say, use AI to help fill in the gaps. It fills in the gaps. But there's no, they, they don't have, the AI doesn't have the, the, uh, the expertise in the domain and neither do the new word marketers.
And so that's part of what we're seeing. We're seeing these great looking articles come out, but there's no depth to it at all. There's no, yeah, good stuff.
I think that's, that's critical. I mean, we're, we're seeing that across all, all kinds of spaces that are adopting ai. You need to pair AI with human expertise and wisdom.
And, um, you know, a AI fundamentally, at least for now, is still derivative. You still need innovation, and that comes from people. And so I think pairing AI with human beings who are kind of giving it the right guidance and instruction, and then can also judge the output of its work and decide, is this flawed?
Is this good? Um, you know, continue to guide. I mean, in a way it's like the, the practitioner on the front lines using the AI becomes like a leader or a manager themselves of the ai, and they need to develop a lot of those same leadership skills and coaching, mentoring skills.
Can I ask, Can I ask you, oh, I'm sorry, Barbara, can I ask you a quick question? So when you meant, it's interesting. So the human in the loop equation or this, this concept are most companies, and I don't wanna put you on the spot, but how many companies out there are doing a good job of, of the integrating the human in ai?
Because I'm not hearing a lot of those. Maybe they're in this initial process of doing this. I I think there's some who are doing a great job.
Um, you know, I I would point to, like companies I've talked to recently, Cornell's Networks, marsh, um, I, I see examples where they're, they're really embracing that mindset, um, building their business as, you know, an an AI first kind of company. But I think most companies are struggling with it for sure. They, they don't know how to lead in this new reality.
And, um, you know, they, they're leading with tools and instead of leading with people and mindsets and, and leadership skills, I think that we're on the spectrum, right? So we started out with these copilots and now we're moving up to smarter AI agents, and hopefully things will get a little bit better. But I have noticed this trend where, you know, execs show up and they're like, oh, this is gonna be great.
We're gonna increase productivity and we're gonna have all these wonderful outcomes. And then when it doesn't happen, and the rank and file starts telling 'em, well, this stuff doesn't really work as well as you think, then what happens next? I think at least the good leaders is they roll up their sleeves, they get in there and they start actually working with this stuff, and then they discover what the limitations are, and they're suddenly a lot more cognizant of just what's real and what's not real.
But Barbara, is that kind of the cycle of things? Oh, a hundred percent. I mean, we've seen it, uh, over and over with technology revolutions in the past.
You know, you, you have the hype cycle, and then you have the trough of disillusionment. And I think, you know, we're, we're going into that, um, that trough. And what I think it's gonna take is leaders getting past the hype and the excitement about all the potential, which is totally valid.
Uh, we're all excited about the potential. Um, but I, I think if leaders are living in their imagination of what's possible, instead of rolling up their sleeves and, and getting their hands dirty, like they need to walk the walk and they need to experience for themselves what this technology can do and what it can't. And then they need to be role modeling with their teams, you know, showing them, this is how I use it myself and my work.
This is how it's transforming what I do. And then, you know, their, their teams can then in turn figure out what that means for their work. Mm-hmm.
Fred, coming back to software development, I was at this conference, and a fellow made me laugh. He said, this stuff is great. I'm running into the same walls 10 times faster.
I'm, Yeah, I, I, I, I, I couldn't be further from that opinion, to be honest with you, right? My, my, uh, my scale of writing code is five x, you know, and the ability to, uh, multitask while doing that with numbers of agents doing multiple workloads is incredible. Um, not to say that it doesn't require the same level of persistence of understanding and testing and rigor that other things do, but it's in essence, like hyperthreading a person doing the work with agents, doing the work that you have to supervise.
So, I'm with you a little bit on that, Barbara. I think one of the biggest concerns is really about who can do that most effectively. And that's really where you're getting, you know, folks that have been doing this a long time, have a lot of experience, can, can really ize their experience with a number of ways of parallelism there.
Um, I I, there's a lot of speculation about what the outcomes of this is, right? If you take somebody that's relatively okay at writing software, and you multiply that with somebody that's relatively okay at writing software, you're gonna get, you know, some effect, right? Whether it's a ripple effect of massive code with lots of other things to be concerned about, or, you know, you, you have a much improved way to, you know, make a force of 10, fight like a hundred.
Um, the real issue is like some of the standardization. I think if we, you know, if we think about, you know, coming back to Tesla, like if you wanna make something like an industry standard, okay, well, SBOs were created by the, you know, Linux Foundation, right? And they were also, you know, fundamentally supported by the Oasp community.
And so when you want to get industry adoption on things, you typically have to open source it. So, you know, I'm a little bit ma on, you know, a a a company sort of building an industry standard, uh, you know, that's funded, uh, because inherently there's, you know, uh, cause for question about whether or not there's, uh, you know, some concern for their, uh, intent. But, you know, ultimately there, there's a set of requirements here.
It's a natural e evolution, like, like Gina said, and I think, uh, this situation is here, right? The question is more like, how do you handle, you know, things like authentication and authorization. How do you manage an identity of 10,000 agents, right?
Uh, a hundred thousand agents, a million agents, uh, as opposed to sort of like, is it going to happen? It's happening, uh, for sure. And I think evidence by, you know, this example of, Hey, we need these kinds of things to make sure we have a, you know, a persistent record of intent so that, so that when these agents start over again, there's so many agents that we have to have some reasonability that they'll come back to where they're, where they're centered to do the work.
'cause there's just too many to manage, really. So, Fred, to your point, I think I have noticed this trend, and I saw it at the event, but there does seem to be something of an AI divide emerging in the software development community, and there's folks like you that know how to make these AI agents dance, and then there's the mere mortals that are kind of struggling to figure out how to organize and manage all this stuff. So will that gap get wider, or can we close it?
That's a good question. Um, I, I'd like to think that that gap will get closed, but I think it'll, it won't get closed by people, right? So the challenge we have here is, and I think we're seeing this with, you know, sort of entry level jobs being, uh, waylaid, um, some of the largest companies we know, massive layoffs for folks.
Also, middle management getting sort of let go in the sense why, because you've got folks that have been doing this job for 15 years and, you know, they can, they can command a fleet of agents doing, you know, relatively good work and, and, uh, staying with the same context, right? So some of the lossiness that happens when you have, you know, humans talking to humans than talking to more humans, um, much less so when you have a human talking to, you know, a hundred agents, a thousand agents. Uh, I, I'm not saying from a, uh, from a civilization perspective, that's great.
But, you know, from an efficiency and a work stream perspective, I think there's an awful lot of optimization truth in it. And I think, you know, as the scale continues to grow, that's where, you know, people I think are going to dig in to see what is that economy and scale for optimization efficiency, that's really useful, but you're gonna get these other problems that we've already solved in other ways now with this set of problems and this set of infrastructure. Just like, just like Gina suggested.
Yeah. The thing, uh, the thing I was thinking when you asked that question, Mike, was you were talking about the two types of developers, but you didn't say anything about ops. And this has a huge impact on ops and whether it works at all, you know, with that kind of ops mindset.
And so I think there has to be, to me, it's pretty exciting that you could manage a fleet of bots. And, but, but like Fred said, you have to have all of that kind of data center hygiene with it. And you also have to have the observability and the reportability, especially with ai, um, for what's going on and how people's information is being used.
So, um, I know a lot of people are getting let go. My my gut feeling is it's because the budgets have been deflected away from other things to just invest in ai. Um, and my gut feeling is that we will need just as many people managing.
It's just that the, the, the, the output's gonna be much greater, but I think the mistakes are gonna be multiplied as well, and can be more catastrophic than we've ever seen, which since I'm not involved will be pretty exciting to watch too. All right, folks. Well, I'm gonna leave this conversation here and just note that, you know, the philosophers are right.
Once again, the future is here and just unevenly distributed, we'll be back in a minute. You've earned it. The spotlight, the responsibility, the weight of teams, companies, and entire industries fall on your shoulders.
Lives depend on your decisions, your home life included, that work your protected physically and digitally. Nothing gets through your team without a fight. But in a globally connected world, everyone sees you, including those who mean to cause you and your organization harm.
And now home your sanctuary attackers see an opportunity. Your digital front door is wide open. And what compromises your home can breach your boardroom.
Because the devil's greatest trick isn't targeting your workplace firewall. It's convincing you that your personal life isn't at risk. Black clerk, digital executive protection, defending the new attack surface your personal life.
Well, it wouldn't be a weak in it without another significant merger and acquisition. And this time, Palo Alto Networks is acquiring a company called chronosphere. They're a provider of an observability platform that is generally used in IT ops and for application development.
But Palo Alto Networks currently sees an opportunity here to apply this more broadly. 35 billion to prove that point. John, um, what's your take on this?
35 billion a lot of money these days? Or is it just lot's A bargain? That's a bargain.
Uh, me met is Met is spending $600 billion on their infrastructure, aren't of course, right, they're gonna spend all that money. Um, yes. I, I'll, I I won't be not facetious after that.
Um, but Palo Alto Networks announced its earnings, and as part of the earnings, it announced this acquisition of this observability platform, which you wrote about recently, I think a couple of days ago, Mike. And, which was interesting because earlier this week, Chronosphere previewed these, these AI capabilities in its observability platform to help identify root causes of issues and provide remediation suggestions, um, among other things. So in a sense, Palo Alto Networks is getting into absorbability, I ha have a hard time saying that we're Bo way.
Um, so in a, in a sense they're getting into it. And I think that the idea behind this is this push into the market observability market at time when AI applications are creating a lot of demand for system monitoring and performance management. Um, and it's, it's, it's trying to, I believe trans transform observability from passive monitoring into autonomous remediation.
So that's, that's the own, that's the, the end goal. Um, I think we're gonna see a lot more acquisitions, uh, especially they're gonna, it is gonna pick up, given the kind of the political climate where they're gonna be, we're gonna be rubber stamping acquisitions now. Plus there's gonna be a lot, there're gonna be a lot of smaller companies that are gonna look for an exit strategy because there's still some sort of lingering fear about what's gonna happen in the markets.
There's a debate, but I think we're gonna see a lot more of this, and the big companies gonna get bigger. But I think it was a good strategic move by Palo Alto Networks. Um, and, um, yeah, we'll see, they're, they're probably not done.
They'll probably continue to do these types of deals. That's true. You know, you made me laugh because, you know, observability is like one of those words, like jocularity, everybody knows it, but nobody wants to say it 'cause they trip on.
Yeah. Yes. Um, but yeah.
Now what did you, I was gonna ask you, Mike, though, what did you think? I mean the, uh, the, the interesting timing and of, of this acquisition, I know that it's been in the works for a while, obviously, but the, the timing used really well for Palo Alto. I think this bodes well for the future because the thing about observability is the, the core ideas we're moving beyond a modern era of predefined ed of metrics that we're gonna be able to collect all this telemetry data and analyze it so we can get to the root cause of an issue faster.
Well, that's been driven mainly out of the app dev and DevOps world because they're trying to improve performance and reliability. But these issues apply to security and IT ops and all across the landscape. And as that occurs, I think what we need to see is maybe some unification.
'cause you know, the security people are collecting telemetry data too, and so is it ops and so is the DevOps teams. And now we got more telemetry data that we're collecting that nobody knows what to do with and it costs a fortune. But Fred, is there an opportunity here to unify all this stuff?
Sure. My agents don't know what to do with all that information, right? That's the theory.
Uh, you know, we, I think there's an opportunity here for, not just for Palo Alto who sees a bunch of writing on the wall, but you know, if they have a massive fleet of agents, right, they need to manage and get the telemetry and the optics around this. So observability is, you know, it's, it's rudimentary for everybody that's not a pure cyber company. But now we understand, okay, cyber also includes, you know, agentic behaviors and all these other things.
AI is the du jour. So if you're going to wander in with software and hardware and do these things, man, it'd be awful impressive to have something that allows you to navigate, manage, and disseminate that information. 'cause the key is resiliency and agility, you know, not, uh, not whether or not we can stop cyber attacks, right?
That's the business problem is not cyber attacks. The business problem is resiliency and agility. Mm-hmm.
Gina, what's your take on all this? Can we all maybe link arms and have a kumbaya moment, it, ops, DevOps, security people we're all gonna like, talk about the same thing at the same time for the first time ever? Well, sure.
I mean, that's the, the dream of it. And this is a perfect, perfect use case for so-called ai. It's, it's really machine learning, probably a little deep learning, but it's the perfect use case for it.
There's too, too many alerts, there's too much going on. And if you had only known this server was getting a little too hot 30 minutes ago, you could have done some action to make sure nothing went down. So I, us as an ops person's dream, are you kidding me?
You love it. No, Barbara, it's no secret that there's not a lot of love loss between application developers and security people. Security people tend to view developers as kind of, well, the root cause of all evil.
'cause they created the software that led to the vulnerability that led to the breach developers state to security. People are, well, they're just in the way they need, we need to build software faster. And I got features to do and deadlines to meet.
And while I can't be bothered with all this security stuff that generates alerts, most of which turn out to be nothing, how do we kind of bridge this? 'cause this is a cultural issue as much as it is a technical issue. Maybe now they can be, uh, united in, uh, casting their blame on AI instead of each other.
There you go. But, um, I, I think, you know, one of the big questions here is, you know, as you have, um, agentic remediation happening, um, who audits the decisions that are being made? Uh, what if the remediation goes wrong?
Uh, what kind of governance do you have in place? And I think all of these teams are gonna have to work together to define that. Um, otherwise they're still gonna be pointing fingers at each other.
All right, Fred, you laughed, but can I take all these people and just maybe throw 'em in a room and lock the door until somebody sees sense or what? I love it. Uh, the trope is terrific.
Uh, I, I don't really see that problem as much as maybe other people do, uh, from that standpoint. Um, but I agree, uh, Barbara's got a great assessment of both what the risks are, and I think you absolutely should lock people in a room. And maybe it's a little bit of a, you know, two man enter, one man leave, you know, from, uh, from Mad Max.
But ultimately, uh, I think this will help drive better specifications, better standardization in companies. And instead of talking past each other about, I think this vulnerability is a high level of probability. Uh, and that's very low on my development priority list.
It's, this is a real thing. We have a lot more telemetry that helps share that. And we have common data to look at rather than, you know, my tools say these things and your tools say those things, and our bosses have to argue about prioritization.
So I think it's a really, it, it could be a really big step up to diffuse some of the, you know, I'm a people person problem of how do you navigate both the requirements to the operational effects of it. So I think it's, uh, the, the future's bright. Alright, Barbara, does that work?
Throwing people in a room and locking the door? Is this a management technique? It, it actually kind of does to a certain extent.
Um, and a a lot of it is, you know, what are the conversations you have in that room? Um, I, I think fundamentally, no matter what your role is, uh, people can benefit from putting themselves in each other's shoes and understanding where the other person is coming from. And, um, when you humanize each other and, and understand each other's motivations, that's where you start to find shared value and, and shared understanding of, of the problem and, and get to solutions.
So yeah, lock 'em up. Lock 'em up. Heard it.
Wait, that's a political campaign. That's a political comment. Yeah, You guys, I'm glad though, Barbara, I'm glad you're mentioning the, the, the importance of humans.
God, what a concept. I mean, given all that we've heard about Agentic AI and how it's gonna eviscerate middle management or replace people, or, you know, take all these jobs and become part of like the customer service or the workforce, I'm, I'm glad that they're, these companies are starting, it's starting to dawn on them that people actually are kind of important in the whole process. I, I, I think that's gonna be the big differentiator honestly, in, in who wins and comes out ahead in, in this period of change, is the companies that see, um, the importance of, of humans, um, for their future growth and future opportunities, um, and who are investing in upskilling their employees.
Um, yeah, it's great when AI frees you up so that you don't have to do this tedious repetitive work anymore. Don't let those people go, keep the expertise and the wisdom and experience they have and grow and build it so they can do higher level functions and continue to be the advantage for you and your future growth as a company. Mm-hmm.
I guess part of my soul here is that maybe I'm just kind of not really using it at, at the level of scale that other folks may be talking about, like with Fred, but I often find that I'm like frustrated because by the time I validate everything the AI generated, I I've done it the right way the first time myself. So, Well, that's, I mean, that's a lot like having an intern, right? Um, you know, you, you only get in as much as you, or you only get back as much as you put into the intern, but then in time the intern grows in its capabilities, its judgment, um, and it can work more autonomously and, and grow into a mature, seasoned, um, expert.
And, and I think AI is the same way. You, you know, going back to what Gina said earlier about going from A to Z, we're not gonna get there by going from a straight to Z. We're gonna go A to B, B2C, C to D, and down the path.
And so you have to start small and build that trust in the AI's capabilities that it's gonna do what you wanted it to do, that you trust what it, what it did, the decisions it made. And then when you have that base level of trust, you got to B, then you can start working on C. But if you try to go straight to Z, you're gonna have a really bad experience.
You're gonna give up and walk away and say, this doesn't work. And, you know, and then we failed. Mm-hmm.
I guess maybe, uh, I, I just wanna hire somebody else's AI intern after they train them and see how that goes. I feel that way about Claude actually. Yeah.
Okay. So Gina, though, let's bring this full circle. It seems to me that this AI stuff will get better as we expose more telemetry data to it.
And we don't have that data as, as widely as we'd like to think. So do we put the cart before the horse and we created all these AI agents and then expose them to, you know, random bits of data, but we really need them to focus on the telemetry data? Well, I think that's the rub, right?
Like, I, I think you can put the agents on whatever system that you have or whatever information you have, but you, it has to be focused on the right things for your business and for what, for the job that you're trying to do. There's so much promise to ai, let's get rid of the hype. But if you have a, a business need and you have that much data, uh, especially log data and, and real time monitoring data that's just perfect to, to set agents on and, and help resolve a lot of issues before they start.
So, you know, you don't get caught off guard by something that you didn't even see coming. So, um, I, I think the agents are, it's, I think it's great. Everybody are starting to play with them.
I hate that we call them agents because I still think they're bots. I don't see the difference. Someone can change my mind about that terminology, but, you know, why not let let the computers run with the things they know how to do with supervision?
I think that's a, a great use of it. So, Fred, coming back for a second to what we were talking about last segment there, is there an opportunity to kind of maybe democratize observability? And I'm asking the question because I, I've explained this to more people than I care to admit, and almost universally I get the same response, which is, yeah, wow, that sounds great.
Followed by three seconds of pause and then it goes, yeah. Well, but I have no idea what questions to ask in the first place. Yeah, I think there's gonna have to be some consensus driven behavior around it, some standardization around it.
Um, especially as more and more, uh, agent interactions happen across, you know, the vast quantities of NCP servers that every company is standing up to communicate with their, what used to be APIs that was now as, you know, a natural language processing exercise. Uh, the, the challenge is the same. So, you know, today we would say, and we'll we'll get to this in a minute, but let's imagine that we have several companies that are critical partners for us.
And, you know, something happens with those critical partners. You know, how do we understand, uh, how do we inform and how do we account for that from a resiliency perspective as a, you know, partner said company. Historically, we haven't really had a way to do that, but this offers a number of ways for us to think about, you know, if we have agents that are monitoring things about telemetry, that historically we would have, you know, maybe an ML model doing this, maybe we're doing statistics and aggregations and all these other things, which are trivial tasks, uh, for, you know, a set of ag agentic flows.
And so if you have a fleet of folks, uh, agents in this particular case, or, uh, there's a little difference, I think maybe in, in, in agents and bots, you know, we, we should certainly have a rock paper scissors, uh, engagement on that one. Um, but the benefit you get is, you know, I can task a fleet of these guys to go do this specific thing, which is observe, you know, my, you know, organization's telemetry, compare it to others in the sense and get a, you know, a good bellwether as whether or not something is happening appropriately and make a decision about something like, do we need to find another more resilient route for a thing, uh, in the future? And I think it'll also help hold companies more accountable, right, to their actual metrics that they say they uphold.
Mm-hmm. All right. Well, John, last question on this one though, but, uh, you're out in the valley.
Is this the beginning of, you know, mergers and acquisitions across Yeah, I think it is space. Yeah. I, I really, I really do.
I mean, and I, and I, it's not, it's not related, but I mean, I think what opens really opened the floodgates was this decision in the meta case involving the FTC, trying to fight the WhatsApp and Instagram acquisitions that it approved a decade ago. Um, that's another story in itself. But yes, I think it's gonna be an acceleration.
There's a lot of money, and I think there are a lot of companies that are really nervous about what we're, where we're headed, um, with all this debate between bubble and, and boom, I think there will be some companies that cash out and, uh, the large companies are gonna pick 'em off. All right, here we go folks. It's gonna be cleanup in the observability aisle, rub bit back, Discover Techron Group, the epicenter of tech innovation.
We are your go-to for reaching IT leaders and practitioners worldwide. Our secret impactful content that sparks awareness, engagement, and top quality leads with us. You'll access editorial websites, streaming videos, virtual events, custom content analyst research, and more.
Join our satisfied clients. Let's revolutionize your tech journey. Contact us today and tell your story to the world in the most powerful way with Techron Group.
Hey folks, we're back on this Friday with our last topic, which is this CloudFlare outage, and it occurred earlier this week, and I think it only lasted maybe three or four hours, but then it cascaded for a while for people to recover. And it's very similar to what we saw with the A AWS outage and a Microsoft outage. And we seem to have these large scale outages these days.
Gina, is this just like the new cost of doing business and it is the way it is? Or is there something to be done about this? And are we too dependent upon a couple of things out there that have so many dependencies that they can take down?
Well, everybody, Well, that's a couple of questions, right? So, yeah, yeah. To me, I think this is the first thing me and my friends talked about was like, that's a lot of single point of failure that you may not even know is your single point of failure.
So what happened was with CloudFare Cloud, uh, can't even talk today with CloudFlare, uh, was their CDN, their content development network went down and it was internal, it was a bug. And I wanna kind of read from their outage report. It was a change to a database system's permission.
It caused the database to output multiple entries into a feature file that's used by their bot management system. And then the feature file doubled in size and that propagated to all the machines in their network. And that, uh, ba basically was what caused the problems.
So the, the feature that that file helped the bot management system keep up to date with all the threats on the content management system. So it's kind of like a, a, a, a, a story of warning based on everything that we've talked about today, which is kind of interesting, right? So we have a bot management system responsible for taking care of, uh, however it worked, taking care of any kind of security threats.
Uh, a mistake was made by somebody, or maybe by a bot, I don't know, in the permissions that were assigned to the server. And it kind of cascaded this cascaded failure. They definitely needed some observability so they could catch this as it was happening so they could get rid of it.
There's definitely some questions about, okay, um, was it a bot? Was it a human? Um, was it what, you know, like what does this bot management system, you know, how, why does it need this file?
Like, all of the questions that come up into my mind was, how did it actually work? Um, so this is gonna happen, I think as we're getting used to automatically, uh, or having, having agents not saying that this is what happened, but having agents go out and, and make changes. It should have been as simple change is what it sounds like.
And it wasn't, and something ha it was just a bug in the system, which that also happens. Um, which caused this effect to have their clients go down. So, like for me, the thing it affected like x it affected open ai, which now impacts everybody's working life because everyone's using it.
For me, it impacted my radio stations that I listened to online. I was annoyed and I had to listen to YouTube 'cause I was too lazy to get up and set up my record player. So, but I have to, you know, I need the music to go on.
So it kind of was like a disruption in my work life. Um, so there's so much we depend on that we have no idea that it's depending on a content management system that is, um, being protected by CloudFare flare and could go down because CloudFlare had an issue with data changing database permissions that kept caused a problem. Um, so, so, you know, there's a couple of things.
There's number one, providing the, my providers of my radio station, the providers of X, everything, anybody else, they're depending on CloudFlare and nobody else. So can they switch over to that other provider if something goes on? Or is there even another provider that as good as CloudFlare?
And then, um, just the consumers, you're not knowing what goes on and, you know, we're technical, we can figure it out that other people trying to get on X or just trying to run their, the whatever they do for work with open ai just kind of stuck like it's not working. Um, and then calling all of us related to, so, so Full disclosure, full disclosure, techstrong is also a customer cloud flare. We experienced some of that outages ourselves, but, um, here's what I'm trying to get at here.
Fred, let's kick this to you. Is it the fault of somebody who created the config file that pushed this button? Who probably feels awful right now?
Or is it just that the system itself is flawed in a way that's create, is gonna create a problem and it could have been anybody at any time and maybe, you know, we shouldn't be beating up on the four little engineer in the config file when the architecture may be the issue. So I, I think, so it's a different scale of problems when you look at 20% of the internet in general, right? And this system was built for hyperscale DDoS attacks.
It's a handful of, of services that sort of take all this information from click house cluster that says, look, what are the latest and greatest? It does this every five minutes. And so that same way you think about routes propagating with a network device like a router, uh, or a switch, and, and the same sort of thing happens and the construct isn't, you know, whether or not as we move into this, this was a, this was an ML process that generates this thing, and it's a bunch of features for a model to make decisions, uh, that, you know, then get informed by this, right?
And the question is really about, there's a file limit size, right? Ultimately like a very human problem, uh, that wasn't dynamically adjusted. Um, okay.
Or database permissions that might have changed during the course of, uh, this process. But, you know, I think to your question, Mike, it's, um, you know, these things, maybe this is the worst outage since 2019 for CloudFlare. We've seen a number of these different things.
We're going to see some of these types of interruptions, but they're all, you know, methodologies that are, I would say, much further, uh, and more impactfully designed for resiliency than what most people deal with in their companies. They're gonna run into these types of things from time to time. I think that, you know, the questions would be like, what's the bellwether for whether or not your file replication looks like, you know, there's a lot of after action, I'm sure these guys are working through and like, yep, we're gonna automate that thing.
We're gonna put this back in the process. I need observability on this particular element here. Right?
All of that that'll all get after action, I'm sure super hardcore. And the question is this sort of, when you turn the keys over, would that have made a difference if a human did it right? Is probably the, you know, its prolific argument, uh, versus, you know, whether or not, uh, a machine learning algorithm did it, or, uh, an agent did it.
And at that scale, uh, when we think about what the problem is, maybe with hyperscale DDoS, like a human's not gonna solve that problem anyway. It doesn't matter, uh, at that economy of scale and that magnitude, right? That's gotta be, uh, an automated workflow.
And that process has to be, you know, pretty instantaneous five minutes, you know, is a, is a substantive time period in that type of, uh, in that type of threat landscape. And I think, um, you know, I don't wanna let, uh, kler off the hook, but I mean, it's just a different level of problem, uh, as broadly as it affected everything. Uh, yesterday, us everybody, uh, from interacting with customers or their out their outputs as well.
Uh, same challenges, Fred, how do you, they, they, they were quick to, uh, to specify that this was not, uh, outside threat. But that seems like a pretty obvious place to introduce a threat when you, you know, hin hindsight quarterback, Monday morning quarterback kind of thing. How can they be sure it wasn't orchestrated from outside?
That's a great question. I think that's why the, you know, first, uh, you know, the CEO would say that the first thing they did was evaluate whether or not that was an outside attack. Uh, 'cause that's a presumption, right?
Or attack all the time. The, uh, the ASU botnet, which is, uh, sort of like they're, you know, contending with this thing right now, which is probably their first, the first jump was this might be actually that, that problem. So they worked backwards, I think, uh, to get there.
And, and at that economy of scale, uh, it makes reasonable sense because that's like a persistent threat for them. So, I'm with you. I, I think the challenge is, you know, once you get into the diagnostics of that, uh, the time that it takes to make an observability decision, and then again, internet scale, the time that it takes to unwind, that takes time.
And so rate of propagation across the globe, you know, and all those things, what was the fix revert to a file, you know, that worked well, right? Like, like we know so well, that's, that's always the answer, revert whatever that change was, like, put it back, fix it. But it, I wanna get to two points here with Barbara though.
So one is I think we can give the cloud player people props because they own this pretty quickly and they got up on social media and they basically said, you know, are bad and, you know, we apologize. And, and so that is a good thing on one hand, correct, Barbara. I mean, that's the way to kind of handle these things, A hundred percent.
Uh, owning it, acknowledging it, being transparent about, uh, what's happening is, is critical. Okay? Second part of that question is that may be cold comfort to the IT people that contracted them in the first place.
'cause I'm sure they're getting a call from their boss going, how come the website's down? How much revenue are we losing? And then the third question is invariably, well, who picked CloudFlare?
So I, I mean, I, I think the scale, um, uh, at which, uh, these systems are operating, um, it, you know, going back to what Fred was saying about whether or not it's, uh, the fault of a human or a bot, I, I honestly don't think it matters. Um, it, in this day and age, it's how do you, what do you do when you have an issue that comes up? You know, how, how have you prepared your people and your systems to navigate that issue and respond as quickly as possible?
Um, uh, figure out the right solution if that's reverting back to the last version, whatever. Um, uh, it's, you know, what's your, what's your fail safe plan? Um, what's your, um, you know, how are you preparing your people and your systems to respond to these outages?
Um, that's the, the focus now as much as uptime is, John, you and I are probably one of the few people out there outside of Wall Street that actually read, you know, 10 Ks and SEC statements and are they now gonna include things like we're over? Oh, they do. Yeah, they do.
All providers. Yeah. So in every, in the, all these documents, they have the risk assessment.
So they, they point out their outages, or this is actually, that's, that's a really good way to find stories, by the way. So you look for, you look for you, you do a search of under risks, and they do, they, they mention everything that went wrong, but they bury it deep within the, the document. You know what I actually think, and I'll, and as an Xfinity customer, I'm used to outages.
And I'm wondering if, given what's happened with AWS and what happened with CloudFlare, and I think CloudFlare correct me if I'm wrong, didn't, wasn't there another incident several months ago? Um, I think, but, and regardless, it's something we're gonna become accustomed to. I, unfortunately, I think it's part of this whole kind of dynamic that we're living under and living with.
Um, so I think we're gonna see more of these outages, um, as these companies make their transit transitions and a lot of 'em are making major transitions, trans Transformations. Um, I think this goes with the territory. Alright, Fred, last question on this whole thing.
Is this an argument for chaos engineering? Because theoretically you should just be ripping things out just to see what breaks anyway. Oh, man.
Uh, On the spot, Fred. I, I think these guys regularly practice chaos engineering. Uh, I, I think there's always an n plus one system that doesn't have the rigor and resiliency you expect when some magnitude occurrence happens you didn't plan on.
And yeah. Uh, regularly implanting chaos telemetry data in your regular operating procedures, uh, is a great thing to do. It also creates change.
It, it also creates a set of, uh, unknown variables. Uh, and in certain systems you'd absolutely wanna reduce all of those things so it doesn't have a place everywhere. But, uh, I'm sure that there's a handful of folks sitting in a room right now gaming out every other possibility for this system and the 10 that are adjacent to it, that it impacts upstream and downstream.
Um, and so, you know, what'll be great is the Outshot, right? For folks that have similar sort of telemetry requirements down the road, whether it's a CloudFlare or an AWS or even, you know, your Netflix, right? From that perspective, right?
To get back to your chaos, the theory. So, All right folks, well, I think what we've established here is what I'm gonna call the new Monty Python School of IT Management. Expect the unexpected.
Hey, thanks everybody for being in the show and sharing your thoughts and your insights, and please stay tuned for the rest of the text on TV lineup. It's gonna be awesome. And we'll see you all again early next week.
Hey everyone, welcome back to our day three coverage of Kubernete of Kon here in Atlanta. It's been an exciting couple of days. We've had a lot of guests, we've talked about a lot of things, but like any good conference, in person conference, the best conversations take place in the hallways.
And I've had a lot of those too. And I'll be talking and writing about that in the days and weeks to come. Let me introduce you though to our next guest.
His name is Scott Rosenberg. Hey, Scott is with a company called Terra Sky. Is that right?
Yeah. Excellent. Scott, welcome to Text Drunk tv.
It's great to have you here, Ben. Thank You. It's great to be here.
Absolutely. So we'll talk about Terra Sky, we'll talk about Cube Khan, but let's take a moment talking about Scott, give people a sense of who you are, where you've been, like your journey. Yeah.
Yeah. So I grew up in Chicago, moved out to Israel, um, you know, back in 2005, and really started my tech career around 20 14, 20 15, started in the mainframe world. Really.
You know, they still exist today, apparently. They, they know they still, they, and they will 20 years from now. Exactly.
Believe me, they're not going, it Started, they're moved into more the VMware, uh, you know, virtualization private cloud area. And then all of a sudden around 20 18, 20 19, just this bug of Kubernetes hit me. Um, that's when I joined Terra Guy as well.
Okay. Um, and ever since I've been just working up in the Kubernetes space, I've been privileged to be a contributor to Kubernetes for six plus years. Uh, very cool.
Working on a bunch of the other ecosystem tooling, cross plane backstage, um, and like what, what, what I'm known for, which is awesome, is I have never been into a single session at CubeCon that I'm not doing because as you were saying, the hallway track are the most interesting conversations Absolutely. That there can possibly be. And you get to meet all these amazing people here.
I can watch the sessions live or the recorded sessions later on. Absolutely. And I just wanna get to meet the people because You only have that window.
Right, exactly. And then everyone goes back into the and the world, All these people I'm writing with on Slack every day, it's like you find the Person well once you meet, meet them, right. Oh, that, and that's a big thing too.
You know what I, five people I meet on Zoom, they're always taller in real life on Zoom, but Exactly. That's have a Body. Well, because it's, because you never, you know, you, it's hard to judge how tall someone is from here.
Exactly. Right. But Anyway, you know, you mentioned Backstage and of course Backstage Con this year Yeah.
Was a big success platform Con here was also you presented a platform con Here and a backstage con Oh, you did both. I, I had four sessions this year at ah, cube cut. So, exactly.
But Oh, why don't you share a little bit about what you, uh, presented on? Yeah, no. So I, backstage Con was really about building out a platform with cross plane backstage, um, and trying to bring together the consumption layer and the operation side from cross plane and trying to fix that.
But like, my favorite conversation, my favorite talk I gave this year was actually a platform con, um, which was with a amazing, uh, guy er ma, uh, from a company called Play Tika Uhhuh. Um, and they're a unicorn gaming company in Israel. I, amazing company.
And we've been working with them for probably like 10 years or so. Really? Okay.
5. Oh Really? Okay.
So back in 2016. And really it was, we talked about the journey that, you know, I've been going on with him together and how we have moved from a do it yourself chaos that happens with early adopters, um, into a real full CNCF bank like platform now, and all the benefits that it's brought them and what the challenges were. Um, and it was an amazing conversation because it was the interweaving of the technology with the cultural elements and how it was able to be done, um, with CNCF projects.
Love it. I love it. Um, Scott, so at some point, it sounds like to me, you didn't say it, but you made the transition then to really platform engineering.
Exactly. Right. I, that's basically what we did, right?
It's so much of this was a cultural change, right? Um, one of the challenges like that we always see is DevOps. One of the reasons that there's that line, DevOps is dead.
Right. Which is so, okay. It's a cute line.
It's marketing. Exactly. But why is that kind of true?
Because DevOps never reached its full potential. DevOps was supposed to be a culture and became the synonym for Jenkins, GitHub actions, CCI, whatever it is, right? And Kubernetes, and that's, no, DevOps was a culture.
Those tools fit well into that DevOps culture. What I love about platform engineering is that platform engineering really is an implementation way to make DevOps actually succeed. And the two are actually together.
Platform engineering is a mechanism that allows us to actually bring DevOps a platform to Fruition. Right? Right.
And that's what a platform helps with. And really what we did with them was, I, this was a over a year project of assessments, of interviews, of gathering data, everything, data driven, and really coming and understanding the different pains of each team, and then going and building out this platform together with the teams to get their buy-in to really restructure things around and build that full platform out from the ground up. Got it.
Right. With the challenges of Brownfield and large scale and all of that Real life. Exactly.
Yeah. You know, it, it's funny, Andrew Clay Schafer is one of the, not founders, but a big, a big name in the DevOps space. He always used to say, he still says, I guess the DevOps you get is the DevOps you deserve.
And part of that is, is because you're right. A good, the, the heart of DevOps was culture, right? It was how we interact in a team.
And for too many people, they used to pay lip service to culture and they'd, what I call bag dive into the tech. Exactly. Yeah.
Yeah. Culture, culture, culture. Jenkins.
Exactly. Culture, culture, gi ups. But they never really, they never really invested in the culture.
Exactly. And I think the other thing is, if you look at DevOps that way, it's not meant, it's not meant to replace Agile. No.
It wasn't supposed to be from the beginning of time to the end of time. Right. Right.
I I think it needed platform engineering because it had to sit on something. Right. Now you've got, so now the way I look at the world Yeah, you've got the platform, right?
So it's like on the first day God created the earth. Right? Right.
You got a platform. Exactly. That's one way of looking.
Probably a little sacrilegious, but all right, well it's good you, but you know, you got the earth now, now we can build on, right. We could build on, on the second day we created developers in code. Exactly.
And, and, and, and we do it in Scrum and, and Agile. And that's the best way or monolith. Right?
Right. But we, we had this code or coders and then we had ops and, and you know, it was like Cain and enable almost very different DevOps tries to bring that Cain and enable, if you will. Right.
Not possible. Kain enable together where we, we, um, you know it, but it it sits on top of the platform. Exactly.
And Agile plays in there and DevOps plays in there and SRE is in there. Oh yeah. And security and observability and all these things that we build on this beautiful Exactly.
Earth, this beautiful platform. Yeah. It's just, it's almost like we did it best backwards though, in that we Right.
We brought the platform in after all these things were running around. And, and I think the truth of matter is, and this is like one of the things I've been saying in like a bunch of the platform engineering, uh, different communities, is that, you know, having a platform is not a new thing. No.
We have had platforms for decades. What we haven't had is a unified platform. Right.
And that is the difference here. Platform engineering is the idea that we need to focus on the platform level instead of just having it as a given. And we have to unify that platform because what we had was a bunch of platforms in the past, and we did have platforms, but they were all disparate platforms.
Yes. And we had chaos. Well, one is now just let's bring those platforms together into one unified one and build all of those tools on top of it.
Right. Which is so many people saying like, oh yeah, we wanna build our first platform. I said, you have platforms, you wanna evolve into a unified platform.
Yeah. Right. And that's the key, I think.
Well, once we get it built, then we could rest for a day. But anyway. Exactly.
Let, let's go. Let, I wanna move to IDP. Yeah.
Right. Because you were back backstage con. Right.
You know, just on that in IDP though, it's the, I will say the tech world can't figure out acronyms for the life of them. 'cause what I always say is, you connect with your IDP to your IDP to visualize your IDP. You have an identity provider, right.
To your Internal, talking To your internal developer platform. Right. Which IDP are we talking About?
You're talking about I meant the internal developer platform portal and both. Exactly. But this, you know, they used to say when platform, when I first became aware of this whole platform engineering movement, you know, basically if you got a handle on Kubernetes, you, you had a platform.
Right. Right. And the engineer was doing that.
I think that we've seen that focus change from, you know, handling your Kubernetes cloud native sort of architecture to, to IDPs. Yeah. And, you know, and platform portal there.
Um, and it's interesting because it's, it's, some may say, well, it defo us from working, you know, perfecting Kubernetes. We got enough people working here Exactly. To perfect Kubernetes.
It's okay. It's, it's not going to exactly be an orphan. Um, but, but what we are seeing that, but here's another thing.
You know, they say two people, three opinions. Yeah. And we're starting now to see, uh, differences of opinion different communities.
Yeah. org. They have a, a booth down the hall from me here.
Yeah. A couple rows back. I work very closely with Luca and the community.
Yeah. We saw a platform come here, put on by CNCF. They offer certifications.
Those people offer certifications. Is this town big enough? Uh, I think it is.
I think that, you know, obviously there's always gonna be some challenges, right. And differing opinions. But I think it's very similar to what we saw in the GI UPS world happen when you had weaveworks that had Flux.
Right. Right. Flux.
Then you have Intuit that go and do Argo ccb, and then you come out and Flux V two comes out. Right. And you have these two tools.
You even have the other ones, like on the side. Yeah. There are fleas And the carve cap controller, and you have some others.
We see it in service mesh too. Right. But what ended up happening, they all come, right.
They go in their own directions. They're figuring out what GI Ops is, and then what happens, the flux guys and the Argo guys come together and create the open GI ops initiative, which comes as great. The, this is what GI Ops means.
We're different implementations of the same thing. And that's fine. Choice is good.
Exactly. And I think that this is healthy because having two organizations, having two groups trying to figure out and really distill down what platform engineering really is and what its core tenets are, is so critical. Because once that happens, the two can come together, let everyone go off on their own approach.
Similar to how people with agents, they run twice an agent when they're doing a No. Absolutely. Let's see what we get and try it out.
But variety's the spice of life. Exactly. Right.
And so, and I think, I think that's what Open Source is about too. It's about choice. It is, it's about freedom to, to decide, you know?
And, and if, and if, and if Backstage, the code doesn't work exactly the way I want it to, I'm free to change that too. E Exactly. It's right.
And I like Backstage is one of the best, I think proves of like what a real, like open source community is looking like. Because what we've seen is like Spotify, when they built it initially and they contributed top stream, they work in a very specific way. And then you had companies like VMware that came in, right.
And they wanted to do something a bit different. And absolutely Red Hat came in and you see these companies have Different opinions, but it didn't stop Spotify from saying, you know what? We wanna sell a commercial version of, which is our vision.
And it's all good. It is big enough for everyone And everyone can push it in their Direction. But that brings up another, another question though I want to ask you.
Yeah. I think one of the questions for the platform engineering community is, do you want commercial IDPs, like the Spotify backstage? Which kind of, it gives it to you in a nice package.
It's all wrapped up nice. It looks good, it smells nice. Here you are package, just add water.
Or do you want me to give you sort of that old time open source, here's all your packages, you pick what you want, put it together and it's a DYI kind of thing. Right? So, you know, it's a really great question.
And it's like, you know, it's something that we've been grappling with for a while with a bunch of our customers and the people that I'm talking to. And I think it really just comes down to there's no one size fits all and depends on the customer. Yeah.
You can't be a jack of all trades. No. And therefore, in a company that's large enough, right?
A big enough unicorn style company that has the ability to really go and fine tune that internal developer platform to be exactly how they want and can base it off of open source and can have a large enough DevOps team to maintain something like that, that's a very viable solution that's probably gonna give them a better platform long term. Because there's a reason it's called an internal developer platform, because typically it's built internally. Internally, right?
Now that doesn't mean that on specific components that you don't have the expertise on, you shouldn't buy an enterprise version of it. Right. I think that, on the other hand, platform engineering can help so much in the smaller companies as well.
Right. And there, I think that you have to go with these commercial Offerings. They're prepackaged Because a, you just, they Don't have the resources.
They Don't have the resources, and they don't hit the challenges No. That these platforms can't cover and Scale and stuff like That. Yeah, I agree.
If you Have 50,000 nodes in Kubernetes, I promise you that most of the paid solutions out there ain't managing it. Well, no, you're gonna need to figure it out yourself. Agreed.
But If you have 50 nodes in your environment, I promise any of those IDPs out there are, are gonna be able, able to help you. So, agreed. Hey man, I appreciate it.
Yeah. Thanks for coming on here. Awesome.
This was great. Great. God, if people wanna maybe follow you or Yeah.
Stay on top. What's the best web? cloud, uh, you can check me on weekly, uh, podcast together with Victor Farik on, uh, the DevOps school kit.
Victor's great friend. Tell Victor I said hi. I will.
I'm a good friend, a good old friend of Victor. Victor's great. Exactly.
He's a Great guy. Uh, so we do an ask me anything, ask us anything on every Thursday, basically. Very cool.
Um, and reach me out on Kubernetes Slack, CCF Slack, cross plain Slack, LinkedIn, any of them. Exactly. All over.
Love the community and great being here with you. Thank You, Scott. Thank you.
We'll come back again. Hey, we're live. We got more coming to you.
And as we wrap up our, uh, day three coverage here of Cube Con Con Native Con, this is Alan Shimmel, standby. Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. We're here with Charlie Cartwright, who's director of the quick, uh, AI agents to help you all of the workflow.
And we're gonna a little chat about, well, just how do we get on AI agents to start working for us? You know, explain to us how you guys envision all this coming together and how do I actually like tie something together across different products and services to actually drive the outcome as intended? Yeah, absolutely.
I think with, with Amazon Quick Suite, you know, we're, we're going to reimagine the way that, that people do work in, in the workplace. And what you're touching on is exactly what we've been focused on. I think if we rewind the clock clock, like in the last two years, I think we've all all had those amazing moments of delight where we were using consumer AI applications and it was able to do wonderful things for us.
But doing that in an enterprise setting, like there were point in time, uh, benefits, but you didn't necessarily have those applications, those AI applications connected to an information ecosystem. And, and that's really the key here to be connected to your data. Um, like data stores, for example, right?
To bring in all that structured data. And the same thing for document repositories, to have access to all your strategic documents, your internal wikis and web search that can come in. And then equally important is to be connected to all these applications for which users often operate with on a daily basis.
And that creates this like, layer of context that's also relevant for humans to be able to get access to relevant information quickly to make decisions. But that is the same context that agents need to be able to execute on my behalf to interoperate with one another and to actually operate with autonomy. And if they're connected to those applications to take actions and they have the enterprise context that is needed to properly perform those actions, then you start to get the real utility that we've, that we haven't seen in some of the enterprise applications For the uninitiated.
Give me an example of how that plays out or something that maybe you're personally doing that the rest of us would go, wow, I didn't know I could do that. Yeah, yeah. No, thank you.
Um, so this is the, this is one of the amazing things about this product is, you know, me and my team, this entire organization, we get to utilize it and work with it and experience the benefits. Um, just recently I was working on, you know, something to look at, like a competitive landscape, for example. And previously, this is something that I would go to various different sources.
One, like a competitive landscape document maybe that I, I wrote, uh, a few months ago, right? And then I would look across information, my, my email, various document repositories and, uh, dashboards and data that I use to understand maybe some things internal and external related to adoption, uh, product capability performance. And then couple that with a lot of time searching on the web to understand what's out there, how these capabilities compare.
And this is something I was able to do with our research agent in Quick Suite, um, which previously had taken days coordinated across several people, and I was able to do this in, in less than two hours. And, and so it kind of like this remarkable transformation where something had taken me so much time just to go across, you know, tens of different sources to collect that information. And then I finally get to the point where I can start to analyze it and think about it.
And I was able to short track that and, you know, within two hours, think about some of the opportunities that we have as we build our product, as we go toward general availability in comparison to, to what's out there, for example. So like a remarkable change in the amount of time that me and my team spend on a daily and weekly basis just because of how well Quick Suite is integrated into and has access to all the enterprise information that I utilize to do my job on a day-to-day basis. And the same is true for people at various different, um, roles in the organization and different levels in the organization.
How much faith can people have in the output of these tools? 'cause everybody's been a little bit concerned about hallucinations. So how do I validate what the AI agents are coming back with is what is actually in that data?
Yeah, Mike, great question. And, and of course this, this is something like, there's no bigger trust buster than when you are asking a question and you know that information, you can see that information and, and here it is this like AI agent or, uh, assistant can't represent that information correctly. And so of course we're doing, um, benchmarks where we look at different question and answer pairs that we can ground truth against so that we can do things to make sure that we answer with the right relevance.
Also, um, when you have structured data, right, like actual data sets and things like that, and all this context across the, the various information ecosystem within quick the document repositories, uh, access to email information and Word documents, various drives, you're able to give the right context to the underlying model. So the, the need to hallucinate and the likelihood of a hallucination also decrease. And, you know, that can often be the case when that context isn't there.
And that model is essentially trying to answer that question the best it can that the user's posing. And in this case, we're able to equip these agents and the underlying models with that context. So not only can they coordinate, communicate, take action, but they're able to do that with, uh, users as well.
Hmm. Is there per chance any way that there's a setting for the AI agents, so I can kind of determine how aggressive I want it to be to answer a question so that maybe some of those wrong answers get tempered a little bit because as far as I can tell, the AI agents are just trying to please us and they go to any l to do it. Yeah.
One, one of the features that we have, the capabilities that we have in the quick suite is the ability to create your own chat agent, like create your own chat assistant essentially. And, and with that you can, um, change and dial aspects of the, the persona, the instructions for that chat agent, just how creative that chat agent goes, how much it needs to stick to a script, so to speak, alright, to, to go in the bounds of a particular goal to only look at context, um, for a certain set of underlying, uh, data stores or document repositories or a subset of other applications and actions. So you can scope that down, but then also fine tune that persona and kind of dial that creativity.
And what you can do is get an agent to respond very much to the objective that you're trying to achieve. So I can fine tune that. How customizable is the overall experience?
You mentioned the chat, but um, every company has slightly different workflows, so how does the platform actually know what my current workflow is? And for that matter, well it suggest ways to improve my workflows. Yeah, so I think to the, to the first part of that, um, Amazon Quick Suite I is, is highly flexible in terms of how this can be utilized specific to a user's workflow.
So if that's certain daily routine task that I, you know, a user's taking information from one application or one source and moving that to another, those are things that can be automated with a, a wide spectrum of automation capabilities that we have from low code to more advanced enterprise workflow automation. And so really what it is, it just kind of depends on, um, the different applications and information that are available. And that's what really drives that customization and the opportunity to automate some of those workflows that exist within an organization.
But it's not that the capabilities themselves, um, are restricted or, or so specific that they can't generalize across different organization's workflows. It's again, back to that context and those procedure documents that can be given as context, so that enterprise context and maybe a procedure document or an SOP that can be used then by that agent planner to create an automation workflow. Mm-hmm.
How do I manage the level of permissions that I give to each of these AI agents? 'cause from what I've seen so far is in their quest to please us, they can be overly aggressive in accessing any you all data they can find. So do I have to think through what it is exactly?
I'm gonna let that AI agent see. Yeah, I think this is, this is both important for users and agents. So as, as Amazon Quick Suite users now have access to information and so much information that, that maybe they didn't realize, right?
Or it wasn't used on a daily basis. And so the things that, that we are very mindful of and, and take with a high degree of responsibility is making sure that users only access the information content that they have access to already. And that through quick, we're not giving access to anything they shouldn't.
And the same is true with agents and specifically with agents. When I, you know, when I create a custom chat agent for example, I can actually equip that chat agent, I can narrow its context, it can start with the same context that I have access to, for example. Or I can further narrow that and I can give it access to a subset of actions, um, or just a few applications even.
So the flexibility is really there to, to take, to take that agent and focus it on a very specific goal so that my interactions, I can limit the context and I can limit the applications and the actions that it can take within those applications. And also on the reverse, the other side of that spectrum is also capable where I can take and and customize that agent or make sure that it has similar access to many of the applications, document repositories, internal wikis and data stores that I have access to. Mm-hmm.
Um, a lot of organizations are experimenting with AI and they, some of them have even gotten so far as to build their own AI agent, but I kind of feel like we're at that point again where we're trying to figure out when do I build versus buy an AI agent. And I think a lot of the use cases that people are coming up with are just gonna be things that are gonna be standard features of a platform like yours. So where's that line between build and buy in your mind?
I think the, the build and buy discussion always comes down to a speci, a specific set of use cases, objectives, and goals. And we look at Quick Suite, um, some of the, the one P agents that come with Quick Suite, for example, a PhD level researcher, right? A business analyst that can not only look at information that's in a dashboard, but that can understand and extract and generate insights on the underlying data set, which is very powerful.
And then a team of automation experts. So we have this broad set of capabilities that come with quick as part of our One P agents. And the other thing that we offer there, Mike, so let's say that a, a customer or a user, and maybe it's a power user, they want to introduce their own agent, right?
And, and they wanna build that on like AWS infrastructure or somewhere else. That agent can then be brought into and can be coordinated with, um, among the Quick Suite agents to maybe offer some specialized capability that is not there in Quick Suite today. And so what we're trying not to do is take, take this position that, you know, everything has to be here in Quick Suite and just like the applications, we want users to be able to connect applications and data stores, whether they're part of AWS or not part of AWS we wanna make sure that we can meet users where they work and how they work.
And that quick is extensible into those applications, whether it's a web browser or it's word for example. And so that's a very, very important, uh, capability within Quick. And so we wanna make sure that for organizations that need to build a specialized agent, we need to enable that and we also need to support that agent to have access to essentially this enterprise comprehensive information so it can operate with the autonomy and the agency that the first party agents can within quick as well.
What LLMs or platforms or the agents using to accomplish these tasks, and are they permanently tied to those LLMs or will they change and swap them out as LLMs Advance and different ones are available at different price points? That's exactly it. We, we evaluate and we change models based on the different capability to achieve the best outcomes, um, or users for that given capability, right?
And so, and so that's really important because as models continue to evolve, we wanna make sure that our users are getting the best experience possible. And so nor do we wanna be, uh, held to one particular model. And sometimes one particular model isn't the best for all the broad set of capabilities that are, are available in quick.
Mm-hmm. So what is the pricing for quick look like? I mean, a lot of folks are, um, concerned that over time the these will just get more expensive and cost prohibitive, and other folks are saying, well, a lot of these agents maybe are low cost today, but over time the cost will add up.
How do I think about the total cost of doing some sort of agentic AI workflow? Yeah, so, um, we, we basically have have two, two SKUs or two tiers, um, a professional enterprise and in, in the, the professional or the $20 tier, um, users get access to all the capabilities that we talked about today. So they can chat across their data, they can use, um, you know, PhD level research, uh, business analyst, and they can also automate workflows like no code workflows with a capability that we call flows.
And so they have access to all those capabilities and the, the things that we do there, um, for research and automation where there's these agentic hours that can essentially be accumulated. Um, we offer agentic hours for those capabilities and for users that need to consume even more of those agentic hours for that research agent, for example, there's a $40 skew where we extend those hours and then also offer consumption so that users can go into, you know, as many hours as they need for those power users. And the same thing with automation workflows.
I think I had mentioned briefly earlier that we support the spectrum of automation. And so one is this no code, natural language automation builder where you can think about like routine tasks being automated. Like maybe, maybe I wanna create like a weekly business review.
I wanna automate aspects of that or, or fully automate weekly business review or, or pull information to prepare for customer meetings and send that out to the respective, um, owners for those meetings to resolve some open customer issues. Like those types of things can be automated and flows, but if we're looking at more complex automation in the creation of that complex enterprise automation for that, we have quick automate. And in the $40 skew, um, folks can author and create that enterprise level automation.
And that that automation comes with Angen planner. Uh, it also has like the build, observe and deploy capabil ability versioning that you would expect in more development like tools and also supports human in the loop so that users, um, so, so that like when the workflow maybe gets stuck or the designer of the workflow, the author of the workflow just wants to put in a, like a criteria for a human to be the decision maker if a policy, um, reaches above a certain threshold or something like that and quick automate that can be supported. So ba mike, so basically to, to reiterate, we have two SKUs, a $20 SKU that offers access to all the capabilities we just discussed.
And then for power users that want to create these complex automate enterprise automations that's available in the $40 SKU along with higher limits on the agent hours for research and automation. So how will workflow automation evolve in the office going forward? Because I think, you know, there's a lot of requests that individual departments have put into it to help them build something and then there's a low code tool and it went back and forth and usually dies in the vine.
And then there's also the whole notion that, um, I, I wanna be able to create a workflow that's kind of disposable. I might only need it for a little while. So is that kind of how this is gonna evolve?
We'll see more disposable workflows because individual end users will be able to do things without necessarily requiring somebody who knows how to code something. They do everything for them. Yeah, that, that, that's exactly it.
And, and so we expect like all of these types of workflows that add add weight and complexity and undifferentiated work that, that users have to go through today just to complete their job, but they don't necessarily add that differentiated value to their goals and objectives. Um, these are things that previously would probably be stacked up on like a center of excellent department, right? That may be looking at robotic process automate or even, uh, teams of engineers that were looking to automate certain capabilities, but maybe it just couldn't be justified with some of the other things that were on the plate there.
And now users in a no-code, natural language prompt can automate those workflows. And so to your point, even if that workflow doesn't maybe have weeks or months of longevity or it serves for a a point in time, um, that can still be highly beneficial and a no brainer for a user to go in there and automate that workflow. And of course, the same is definitely true on the other end of the spectrum for the more complex automation use cases where you can use agents and natural language and maybe a process document to get started to create what is a much more complex comprehensive workflow.
And then the user can go in there and a adjust or insert workflows as they need at any level and debug as necessary. Of course that comes with test runs so that validation can be there, but it's absolutely the case, Mike, that we will see a level of automation opportunity that was just not attainable before. And so, and I'm really excited, you know, even in, in internally before we, uh, went with general availability with Amazon Quick Suite, you know, we had tens of thousands of Amazonians that were basically battle testing this.
And so it was, it was energizing to see like the excitement and the love that they had and the utility they were getting across these capabilities. And we were seeing exactly what you're talking about where we have folks that had probably never, not probably, that had never automated a workflow before that were coming to us and saying, Hey, look what I've done. I used to have this, I used to collect this survey across all these departments across the breadth of Amazon, and that took weeks and weeks.
I think they described it as months even, and they were able to do that in, in an hour with Amazon quick. And so, you know, these success stories are, are remarkable and it's so great to see them. So we are seeing that exact transformation that you described internally at Amazon and, and we're also hearing of these great success success stories, you know, with our hundreds of external beta customers as well that entered the, the Quick Sweep beta program before launch.
Alright, Hey folks, you heard in here you can go after some big giant AI project all you want and you might swing for the fences and miss, or you can play small ball and just go after all these smaller little workflows that at the end of the day, collectively, I'm probably gonna make a much bigger difference. Hey Charlie, thanks for being on the show. Thanks Mike.
Thanks for having me. All right. And thank you all for watching the latest episode of the Techstrong AI Leadership Insight series.
You can find this episode and others on our website. We invite you to check those out. And until then, we'll see you next time.
Hey everyone, we're back here live at Cube Con, cloud Native con wrapping up our day three coverage. I want to introduce you to our, our next guest, and if I mispronounce his name, I apologize. Gorav Sena.
Yes. Gorav is Gov's a real life platform engineer for a major automotive company. Um, he gave two talks here at Q Con this year.
One was on behalf of the open SSF and one was on behalf of the, uh, PE Open Telemetry, oh, excuse me, the hotel. Yes. Uh, conference.
Now, Monday of this week was satellite conference day. So we had o hotel, uh, we had the hotel conference, we had the open SSF conference. We, they had others.
We had the PE conference, we had a backstage Con Argo, they were Argo Argo, CD Con. And the rest, you know, it's the day for all of them. But gov let, let's hear about your two, uh, your two presentations.
Sure. But before we do, yes, without, I don't want to, you know, we're not mentioning companies or anything like this, but give an like, people an idea of, you know, how long have you been a platform engineer? What do you, what's your journey been like?
Yeah, so I've Been a platform engineer for about 10 years now. And, uh, I today work for a major automotive company. So my work is basically, I, I lead a platform engine team to build an internal developer platform.
So think like many verticals, observability, Kubernetes, CICD database as a service, messaging as a service. So that's the work we do today for our internal teams to improve their developer productivity and efficiency. Very good.
And, uh, my talks that you mentioned about what was for the open telemetry, so basically the use cases are how do we do the over the air deployments to the vehicles? How do we monitor them through our open telemetry stack? That was my talk basis on the generalized hotel concepts.
The second talk was around the open security. So how do we do the secure software supply chain? Like how do we make sure that software that we are deploying to the vehicles, uh, secured by nature?
So that was a talk around the, around open security part. Excellent. Yes.
You know, that, that's one of the interesting things about platform engineering, talking about observability, talking about security, but it all kind of fits in right to this idea of a platform, right? An underlying platform that we build upon, whether it's DevOps or, or cloud native or you know, we we're building on that platform. It, it touches all of those things.
Good. Um, what, what was some feedback from the audience from, you know, peers and other platform engineers there? What did they think?
Yeah, so I got a huge interest on the way we do open telemetry today. And the reason being is that, uh, in my current role, we are using hotel stack for all the three signals, logs, metrics, traces, and how do we use this in a vendor agnostic way, right? It's an open, it's an open source standard.
How do you scale millions of vehicles that are support on the platform? How do you monitor them? How do you monitor the, the, the progress of the software delivery?
How do you enhance the cloud software development to open telemetry stack? Not only in terms of the observing, but also the, from the deployment perspective. So I got a huge interest, uh, in terms of like, uh, people coming in and asking like, how do we use open source software?
How you able to like manage the fleet of collectors? How you able to upgrade them? How you able to get the maximum performance out of those collectors, right?
So that, that's a, that's the area of interest among the, among the audience that I got basically asking about like the, the, what's the blueprint, right? Not essentially sausage making, but the blueprint, like how the high level concepts work, right? Other, uh, other theme around the open security con was basically the way we do our software deployments in terms of CICD pipelines.
How do you make sure that those images that we are deploying are free of critical mal liberties? How we are making sure that those images that we are actually taking from third party dependency softwares, they are trusted. So what are, what's the process?
How we are doing the, the trust in the zero sec in the zero security network architecture so that there is a, not only the software that we are running, but it's dependency also, we make sure that they are in compliance with our security compliance point of view. Right? So those are the Love it.
2 million things. Yes. If I appreciate that.
And you know, the nice thing is probably both of those were recorded, right? Yes. So they'll be available to download, correct?
On, on the, uh, CNC website CNCF website. Yes. Let me ask another issue or another question for folks out there who are saying, you know, I, this platform engineering sounds interesting, I'd like to be a platform engineer.
Yes. Or, you know, I considered that, give people a sense of your day to day as a platform engineer. What do you work on?
What do you Yes. What are some of the problems you're solving? Yes.
Like what, what is it like to be a platform engineer? Oh, I, so I came into platform engineering field when, right when Kubernetes was getting shaped up in 2015 era. It's been 10 years.
And, uh, before then I was a software engineer. So coming from the engineering background, when I was an application developer, my knowledge, my skillset were only confined to that particular business domain level knowledge. When I transitioned to platform engineer, my breadth of knowledge skill sets, not only in terms of the infrastructure, but how do we operate the services as a service operator, I was able to find out the business value because I today host many different applications that powers different critical user experience workflows on a, as a, as a operations engineer, I am able to make sense of why the services exist in the platform, how can we make it more performance in our infrastructure, develop development, and how can we power the, those developers to do their job efficiently?
So it gives me a pleasure being a platform engineer, not only knowing the breadth of the, all the operations for those services that we operate on, but also powering them, enabling them in the, through the infrastructure in the cloud native way. So yes, I enjoy my role at the platform engineering. Do you?
Yes, very much. Before platform engineering was a thing, what did you call yourself? I was a, a generalist software engineer.
So I was working on, uh, things like, um, video and coding. So I was, before my platform engineer, I had a background on writing the software that actually was in your mobile hand headset, for example, I'm talking about pre Android, pre iOS era, pre Apple era, where you used to have like a Nokia, hand handsets, blackberry hand handsets, and you had a very low power devices in terms of the capabilities back then because the compute, how you are able to render your media playback on your, on your handheld sets. I was working on those media encryption and audio compression algorithms back then and my contributions then actually got open sourced to the Android.
Android to the Android. Oh, that's great. So The Android, the first Android version 2008 that came out was running the software that I worked on That you are talking.
Yes. So good for you man. Yes.
Thank you for that. Yes. Alright.
Hey, we gotta wrap up. Yes. Gov.
Thank you. It a pleasure you for coming on. Thank you for what you do for the community.
Thank you. Hey, that's gonna wrap up CubeCon day three here, man. I hope you've enjoyed it, Atlanta.
We'll miss you. Uh, we'll be back tomorrow with our regular tech strong tv, but for now, this is Alan Hummel, we're out. Hey guys, thanks for the throwaway here with Chris McHenry as Chief Product Officer for aviatrix.
And we're talking about, well, the degree to which maybe nation states are steal and encrypted data that they plan to decrypt someday and using quantum computers and how prevalent this all is. Chris, welcome to the show. Thank you very much, Mike.
Good to be here. Some people my friend, are a little skeptical of this is even happening. So what evidence do we have so far that nation states are doing this?
Is this a theory or do we have some actual examples thereof? Yeah, it, you know, I think it's, it's difficult to ever say exactly what a nation state is doing, uh, but it is really interesting. I think there's several pieces of evidence, especially with the investment that nation states are making in quantum computing.
Now, uh, particularly we know that China has made some very big investments, uh, in, in terms of de developing the technology. And one of the primary use cases, uh, that are has been most attractive to nation states and for potential military applications has been the fact that it has the ability to pro to break most of the fundamental standards of modern encryption. So I think, I think the, you know, the, the things we know about quantum breakthroughs and we know who's making the investments in them, definitely lean towards the fact that, you know, there is some serious interest in this space and, and we need to be aware of it.
Now, I will say one more really, really interesting piece of evidence that we have is one of my favorite, um, examples of a hack probably happened about 10 years ago, and it was a manipulation of the global internet routing tables that temporarily, temporarily rerouted about 40% of the traffic through, uh, through China. Uh, and so, you know, you've gotta ask the question like, why is that even interesting to some extent, especially when most of the traffic on the internet is encrypted. And then we did see last year, uh, a, uh, a really high profile attack on US service providers by an organization known as SA Typhoon, which is generally associated, um, with the Chinese nation state.
Uh, and they, uh, managed to penetrate a lot of the core US service providers. And, and I think the use cases there were primarily about snooping on, um, very targeted communications. Uh, it it, you know, relative to national security, uh, maybe not as relevant for enterprise data, but we do also know that there's a long history of stealing trade secrets.
So, I mean, it's, it is, uh, it, it's hard to say ever if something specific is, is happening, but we have a lot of evidence that points towards this being a potential potentially very real issue. Yeah. Some people, of course, are waiting for what's known as Q day, which is when the quantum computers get smart enough to break some of those encryptions.
Um, the question I have is how close are we to that and might we even know when that day arrives? Yeah, I, I, it's a really, really, it, it's incredibly difficult to predict. Um, I think one of the things that's really interesting though is we have seen very meaningful progress in the development of quantum computers over the last couple years.
Uh, not just from nation states, but also from private companies with announcements from Microsoft and Google and the investments that they're making there. Uh, you know, we also saw the federal government ratify post quantum encryption standards last fall, and we're gonna start to see a lot of those come into regulations in the next couple years. But my best prediction is that we're looking at like early 2030s, but I think the preparations that organizations, the technology to help be preventative about this is, is gonna be, it's here now in many cases, and, uh, we're gonna start to see this be a, a best practice, uh, really next year.
I, I, my my prediction is the next fall is really gonna be the year where everybody starts to kick off their, their quantum encryption upgrade programs. Now, one thing that's important is this is not the first time many enterprises have gone through this. 3.
That was maybe 2016 ish timeframe. So, uh, we are practiced in this, but I don't know that anybody really has a process. And so it, whether it's quantum or not, we are going to need to upgrade our encryption algorithms over time as both traditional and quantum computers get better.
And I think next year is gonna be the year where enterprises really put this on their whiteboards. How do I have this conversation with business execs who are gonna kind of look at you and say, well, let me get this straight. They're stealing data today that they might decrypt three or four years from now.
And the business exec goes, I don't think any of that data's gonna be valuable three or four years from now. I think that's a really real question, right? And, uh, and so, you know, it, it, it, a lot of it depends on, on when Q day ultimately happens, but if you don't start preparing for how you can keep your encryption is incredibly important.
Like, we want to best practice as all of your data in transit and in rest needs to be encrypted, and organizations need to invest in technologies to allow them to keep their encryption up to date. This is really just another threshold there. So I think the harvest now decrypt later schemas are obviously, um, you know, very interesting and very pressing, and we'll have an incredibly, you know, when Q day comes, we'll have a huge, huge, huge impact on how we think about cybersecurity.
Uh, but it's, uh, you don't know when that's gonna happen, right? And, and, and in most cases we probably won't even know because a lot of it is funded by nation states when it actually does happen. And so that's, you know, it, it, it, we are, we're getting close, we've seen the signs and, and now really is the time to prepare.
To your point, it may have already happened and we just don't know about it, but how big a lift is it to swap out my encryption scheme is on different products and what kind of effort is gonna be required and, you know, what kind of funding for that matter? Yeah, it's a great question, and I think that is one of the big challenges, right? 3, that we put it systems in place to make that easier down the road.
But I was, I've had multiple conversations with enterprises even in the last couple weeks where they know it's gonna be an, a huge challenge for them to upgrade. One of the, one of the big challenges is that, you know, not all of your applications running in your environment are modern applications. And so we need to think about multiple different layers of security here.
And, uh, and that's really where tools in the network I think will become incredibly powerful, uh, to be able to, uh, do network level encryption even when the applications can't. So it's like you are solving the problem at the trunk of the tree rather than at the leaves. You want to do both.
Um, but there will be a lot of, there will be a lot of techniques and I think technology to, to ultimately improve that process. We're definitely investing in that space. You will see from us next year a one click easy button to do PQE and, and all of your, you know, for all of your application traffic in the cloud.
And, uh, and I think a lot of other vendors will be investing in this, in this space too. Some folks I talked to are counting on forklift upgrades. Their assumption is that sometime in the next four years, they're going to upgrade their server storage infrastructure and applications and whatever it may be.
And whatever that new thing is, it will come with the appropriate level of encryption. Is that a decent strategy? Uh, no.
Is the, is the short answer to that que short answer to that question, right? I mean, all you have to do is look at, uh, typical enterprises and, and, and see, especially ones that have been, uh, around longer than 10 years. Like, you don't forklift your entire application environment.
I mean, up until recently when you booked an airline ticket, it still went back to a mainframe. I was talking to a health insurance company, uh, the other day where a lot of their claims are still processed in cobol. I mean, there's, it, it, it, it, it, it's a, there's no such thing as a forklift.
Every organization does rolling upgrades, right? And, uh, and so that's where, you know, where, where you can make big impacts is looking at the, the roots of the problem, right? Going down to first principles.
And I do think many organizations are gonna look at services at the infrastructure layer that will help them to, um, to get larger swaths of their environment ready, both from a data storage encryption perspective. That's gonna be a big one. Um, you'll already see that in some of the cloud providers.
And then obviously also at the network level. One of the big challenges at the network level is that, um, it's, it, it in many cases may be hardware dependent. So I think what we're gonna see is a lot of software solutions and some innovation in software that will help you upgrade your encryption without necessarily doing a, a hardware forklift.
So what is your best advice to folks about how to go about putting together the plan and executing and getting everybody on board? Because there's just a lot of moving parts? Yeah, I mean, uh, my advice almost always, and, and you'll hear me say this across a variety of cybersecurity topics, is look at first principles.
We have a tendency as organizations to play whack-a-mole. 'cause we're looking at, you know, the, the things that are at the higher levels that's like, Hey, I want to go patch my servers to eliminate the vulnerabilities. Well, you know, if your servers weren't available on the internet, maybe, maybe, maybe they wouldn't be as exploitable, right?
Same thing, I think with encryption. I talked to a lot of companies who say, Hey, we have encryption at the application layer, right? Great.
Do you have it on every single application in your entire environment? And the answer is, I've literally never seen a customer that that can say yes to that, right? So, uh, what's really powerful is when you think about it at lower layers of the stack, when you think about at the network layer, at the storage infrastructure layer, those are places where we can make a really big and broad impact without having to play whack-a-mole.
What do you think the government's gonna say about all this? Will they come up with some mandates? I know that NIST has been involved with creating some of the specifications, but, uh, at what point might someone show up and say thou shout?
Yeah, I, I think it's gonna be next year, honestly. Um, but I don't think it's just gonna be government, right? I mean, typically the way that the government mandates work is that they'll, they'll make rec strong recommendations to industries that they contract with, right?
I'd say from an enforcement mechanism perspective, what is really interesting is when the industry specific, um, regulations or, or, um, you know, uh, standards start to come out, like PCI as an example, uh, I, I strongly suspect in the next 18 months we'll see a modification to PCI that recommends particular encryption standards, uh, really focused on, on, on, on post quantum encryption. And, uh, and so it'll be the government first and then we'll see the industry specific, uh, industry specific recommendations and standards follow. Do you think the insurance companies are gonna show up and say, Hey, we're not gonna give you that cybersecurity insurance if you don't fix this encryption issue?
Yeah, It's a great question. I mean, I think, uh, what you typically see with insurance from a cyber insurance perspective is that you have to have your first principles covered. Encryption is always one of those, and as soon as the standards change, it definitely will follow that from a, from an insurance perspective.
We've seen that multiple times. All right. Hey folks, you're heard it here.
We're upgrading our encryption. It's just a question of when and how painful and how long. And the sooner you start, the easier it'll be.
Hey, Chris, thanks for being on the show. Thank you, Mike. All right, and back to you guys in the studio.
Hello everyone, and welcome to DevOps experience, uh, advancing DevOps for AI native. Before I start my talk, I would like to thank Tech strong group for putting together this experience with the industry leaders, practitioners, uh, coming together and advancing the DevOps and AI native concepts for software delivery. I'll start this talk with some news.
Uh, for not so good reasons, AI native software made some news lately, one of the AI native software startup companies promised a no code AI generated software, but reportedly relied heavily upon underpaid human programmers to compensate for AI shortcomings. There was another incident closer to where I live. Uh, a major Canadian airline company lost a court case after a chat bot provided a false information about travel discounts and the court hold accountable the company for that AI generated advice.
And lastly, I would also point out one of the catastrophic failures where an AI agent from a reputable, uh, AI native software organization wipes a production database and then lies about it. So apparently it seems that, um, you know, production grade AI native software is still in making, and companies are having trouble operationalizing AI native software. So what I would be doing today and discussing about how DevOps practitioners could help in this journey, you know, creation of organization structure, ai, my first mindset, the culture curation, which needs to happen, AI native companies, uh, holding them accountable and responsible for, uh, ethics, privacy, uh, responsible ai, and bringing it at, at a core of the software organization alongside the rapid growth, which is happening now.
If you are still, uh, listening to this talk, I would like to introduce myself. I'm the founder for the Canada DevOps Community of Practice here in Canada, which has several chapters. I'm also the producer for, uh, summits Canada, which does micro and macro summits.
One of our greatest event, uh, the DevOps hackathon is coming along in Toronto in November. My call to action is leadership practices and Communities, and I will further discuss a it in little more detail, but I wanna kind of project, uh, the possible future, uh, with DevOps, uh, when ai, AI native is generating so much of news buzz and hype. So how DevOps practitioner with data machine learning and AI can survive the strong.
So history reminds us DevOps is a collective journey towards evolution of software practices. And AI enhances predictive and adaptive decision making, which means that we are finding new ways to use data sources that matters for flow, feedback and experimentation. We also are up for collaborating with 19 million developers and coding assistance.
We all know that AI native vulnerability storm is on the cards, so we have to approach it with a lot of consciousness where human and machine interactions matter. We have also seen and advocated for open source projects in the past. The DevOps practitioners have seen the highest performers, uh, for us has been open source projects.
And not to say that it's not only faster but valuable, which matters, right? So we can learn a lot of things from our DevOps journey when we are going to this AI native space Estimates suggests that leading tech giants will invest approximately $1 trillion on the development of AI in the next five years, which entails that rising demand of for reducing software development cycle cost, as well as accelerating deliveries. New methods are prospected to be considered in the software delivery process, like new and easy and human friendly or machine friendly way where soft software abstraction layers will be built up.
And DevOps market is projected to grow, but it also needs a strategy lift, right? And this facelift will happen through different facets. Uh, we will talk about it in the next, uh, few slides.
So projecting the possible future with DevOps, and I call it often DevOps plus because there will be more practices, more processes, more projects, which will contribute to this journey. I, in the first phase of this talk, I will talk about addressing the operational challenges of AI adoption and how DevOps practitioners could be of help and support. Here.
We also would talk about of operational design of organizations and future organizations and what DevOps practitioners have to do to align to this change curve. And lastly, I would talk about a RS management policy and governance. So getting started, there are foundational challenges with AI native projects.
Uh, Gartner has predicted 40% of agent AI pro projects will be canceled by end of 2027, due to cost overrun, unclear business value and inadequate risk controls more broadly, Gartner has also forecasted 85% of AI projects will fail to deliver value to CIOs. So from here we can recognize that the challenge is evolving and the complex expectations we have on the leaders, and how leaders practitioners in this space could take this opportunity and dive deeper into how they could support this transformation of the technology curve. When we talk about operational AI native, uh, operationalizing AI native software, we are talking about a lot of pressure, which is developing on mid layer and senior layer level leaders, which who must advocate for AI first culture.
It also entails that large scale upskilling, 50 to 60% of jobs should be transformed, and it is pivoting and to be replaced by ai adopting new technology. Experimenting with novel approaches is not a choice anymore. Authenticity, transparency, ethical AI use will be essential for ma maintaining that mainstream trust.
And it is needless to say that rapidly evolving cyber security space is also putting us into a lot of challenges. Now, when you see this urgency and you support this urgency, the leaders and the practitioners have started to take charge of the situation and they've started to act upon, right? So individuals with let's say, limited AI knowledge, you see that they're overestimating the ability to, to prompt and interpret AI systems leading to poor decision making and how, you know, ai, uh, outputs should be assessed or, you know, put in, uh, practice underestimation of AI risk, underestimation of AI limitations.
Uh, users are often overlooking that ai, especially large language models, lack true understanding, judgment, common sense. And AI is also amplifying this e uh, effect by readily making, uh, available the AI generated answers, for example, AI generated code, for example, to create an inclusion of expertise and knowledge. So you will see a lot of sprawl in the AI space as tsunami of shadow AI applications initiatives, which also brings a lot of challenge on the table.
And it leads us to this thought process that behind every effective AI deployment lies leadership that plans innovation with disciplined execution. It is important to understand the strategic alignment of, you know, the initiatives, the key initiatives which we are putting in place through the executive management. But we also need practitioners like DevOps practitioners.
We need a strategy for cloud platforms. We need strategy to architect these solutions at a very, very fundamental level. We also would like to have, um, intelligence feedback loop, uh, associated with all this and the data pipeline to ensure that we need to be more conscious about the full stack approach of execution, not only, uh, curating strategic key initiatives and just leaving people alone.
Another aspect is that what we have seen from the past is that technology alone is never enough for two productivity. So if as a leader, as a practitioner, you believe that AI is savior for everything, I think you got to think through a little bit, because I'll give you some points to ponder upon. When electricity came in, it required factory redesign, not just replacing steam engines.
When computers came in, software, digital infrastructure, new workflows, everything was needed. Upskilling of people was an essential element in AI world. It demands data pipeline, governance, new skill, cultural.
These all things will complement to the productivity curve, which you will be ensuring in your organization. I will also advocate that Darwin's story of evolution needs not to be limited to the world of organic things, but can be extended to the world of innovation, which means survival of the fittest, which means that, you know, understanding the AI ES curve is crucial strategic decision making. And innovation flies in the fact that how value forecast as an organization forecasting future progress, performance, market saturation, all depends on you.
How strong feedback loops you have created. Right? And there where DevOp practitioners could be of help.
Also, when you are making investment decisions, identify which part of the AI esker a specific AI technology can help your business, right? So experimentation in a streamlined way through, you know, curating a list of strategic initiatives. Navigating multiple s-curves means that, you know, how do you in keep your flow intact?
Why you are actually having an intersection of s-curves such as scur for large language models and technology obstacles have to be overcome to ensure that you reach that productivity of, uh, from through the technologies lifecycle. And this is again, um, complex mix of, uh, things which are happening as we speak. Setting clear goals and objectives is, uh, the call of the hour.
The key components of an AI native tructure is how well your project objectives, product objectives, and financial objectives intersect. So how do you link your AI product OKR to your financial OKRI will also ensure, ensure that we have brought in DevOps practitioners in play here. Evolution of practices, for example, operational practices like SLO management, AI assistance, law monitoring analysis, continuous resilience automation.
All this, uh, is needed to ensure that we curate this list from a strategic point of view. And we should have feedback loops, like, how do we optimize delivery, including managing progressive delivery or observability driven development, or managing AI and related risks extra, right? So all this all in all will make an organization more resilient towards the pitfalls of, uh, AI native organization and AI native software, uh, you know, processes.
And how do you also take the next step is again, trying to understand the organization structure. Allocate your 20% of, for example, your time, r and d time to experimental projects and achieving that faster deployment through AI would only happen when organization design facilitates that, uh, part. So I will also kind of, uh, articulate this in a way that it makes sense to upskilling and reskilling.
Uh, you know, 85 billion jobs, uh, might go unfilled fulfilled by 2030 because there aren't enough skilled people to take them according to contrary. 5 trillion in unre unrealized revenue. So skill shortage is on the card.
There are many aspects of learning, many people are getting behind due to lack of time. What self development in reducing a major skill shortage cost is another factor. More and more services and applications.
And, you know, there is a plethora and of tools, which, you know, it makes us hard to believe that we only started with this evolution of practices a decade ago or something, and complexity. It is inevitable that the complexity is increasing with rapid adoption. There are new risk, new operational overhead and new cognitive load on practitioners, which takes us to the next slide.
That what, as a practitioner, as a leader, as an individual, as a team player, what could I do to dify this? And there are four components of this, um, operational design for future organization. And the four components is what it means to me as an individual practitioner or a leader, what it means to my team, what it means to my enterprise, and how should I take this in a holistic way.
So think about building your own AI native skill radar. How, what kind of new skills are needed in for this new era? What new teammates you'll have, like AI code companions, how do you build an async first culture to AI collaboration?
What kind of AI native workflows would be needed? And sustainability with ethics and responsible AI and getting together all this, how DevOps could help you in this process or journey in the next five years. As a software developer, what we will see is that employees who are generalists with broad skills, but no contemporary areas like cyber or cloud or subject matter expertise, are likely to struggle for those who excel in programming.
Individual programming languages, for example, currently in use may also feel a sense, uh, sense of insecurity because there is a need to understand multiple languages, which will grow quickly, right? And at the moment, the IT skills which are nearing end of lifecycle, if like manual testing, for example, ui, ux, uh, designing, cross-functional team members can automate these test testing skills. And cloud-based engineers would also be able to do a lot of these ui, ux, and ui, you know, onboarding and capability.
So these are things which people should kind of look at from upskilling or reskilling perspective. Now, how should I start this journey? How should I take the first step?
What I would recommend is that ask yourself three questions. What have your roles achieved in the past? How well your role or yourself, your team have collaborated with others, and which new skill is needed for upskilling?
And if you think about this, you will come across organizations which aim to integrate AI into your business models, creating new values and giving rise to next generation of software companies. AI is with is a speed and a scale multiplier for software features. So what we need is to learn how to collaborate with generative ai, develop new AI applications, design, and, um, also curate and, uh, do content for that empathy.
A ethics plays an important role. Onboarding of responsible AI is essential. And then your left hand side of the brain, which is more problem solving, critical thinking, mathematical and computational knowledge system thinking, communicating with intent, and then so on and so forth.
The roles which we probably would see in the upcoming years as data scientists, machine learning engineers, data engineers, AI specialists, AI researchers, prompt engineers, and AI programmers. And also you will see a rise of code companions in your team. So imagine if you te treat each of your AI agent or a companion as a new high potential employee, how would you you treat them?
We would provide comprehensive onboarding material, equip them with necessary tools, create a safe space for trial and error, and establish effective communication channels. Similarly, if you see from a software engineering practice perspective, um, there are more and more AI engineers and AI programmers in the loop. So if you see on this slide, you'll have a lot of perspective on AI engineer, full stack engineer, ML researchers, ML engineers, data scientists, research engineers.
These are new functions roles, and probably these are areas of upskilling, which can give you some ideas of how can you take that journey and take that steps in the right direction. Now, when we talk about AI ready enterprises, so what could happen is that we would be relying on a lot of AI companions and these AI companions, uh, um, need AI friendly interfaces for team collaboration. We also need fast and slow thinking agents for real-time, re real-time reasons, and for research interactions, comprehensive knowledge base, documentation, testing, quality checks, data pipelines.
All of this would be needed to ensure that we have foundational capabilities to design the next generation future organization, which are more AI centric. We also need, uh, new skills for new era, as we have mentioned earlier, recognizing patterns of bias, risk and failure, understanding how knowledge is shaped by who it creates and what it ends up to. Excising ethical and aesthetic judgment, not just functional adequacy.
When we talk about organizations, uh, and individuals, every organization in every individual will have unique journey, right? So there is no one size fits all, uh, you know, organization design. At this point in time, at least, we could have some AI development team, which can help, uh, develop customized, maintain AI models, especially l LMS for various software development task.
We could have data management team, which can manage the data required for creating and refining AI models. We, uh, will see a rise of AI integration testing team, which integrates AI solutions into existing processes and systems, and ensuring their functionality and reliability. Also, we probably will see a rise in AI ethics and compliance team ensuring the ethical, ethical development and deployment of AI systems.
And moving forward, we would also see, um, a rise of AI agents, an agent which is backed by generic or specialized lang, a long, large language model like software development agents, product management agents, operation support agents or research agents. So, you know, the initial investment in developing training and integrating these AI systems for software development organizations can be substantial. And also maintaining updating or training these AI systems to keep them effective can also incur significant cost.
Organizations should be conscious of the fact that all this is coming, uh, along with the AI hype, and we have to be prepared to deal with these AI led micro change initiatives, uh, at the core. Now, I will also talk about scaling, um, ai, risk management policy and governance in a nutshell. If I could do a polling question at this point in time that what are some of the perceived risks for AI native organizations?
Is it inadequate or misleading output due to limited training data, legal or ethical implement, uh, implications in copyright infringement, cybersecurity vulnerabilities in managing large data sets, impact of job roles, or workforce displacement, or all of the above? I'll give you a second to guess, but my answer to this would be e all of the above. With the rise of, you know, organized hackers with the rise of, uh, AI enabled, uh, you know, malicious applications with the rise of kiosk gpt and black hat AI tools and wool gpt, I think it is evident that we got to be more serious to implement a culture where, you know, we have more, uh, capacity to deal with security posture.
We get to get to have more, uh, people as well as more investment into bias audits. For example, uh, AI model explainability scores are important, and if organizations don't lead with governance, they will fall short. They will adopt AI in silos with limited scale or return of investment.
Creating prioritization, uh, for, uh, you know, not creating prioritization for these results would result into fragmented investment, uh, being reactive on addressing risk and damaging both reputation and finances. And lastly, repeat mistakes across views throughout the AI adoption lifecycle, which can cost multi, uh, afford, right? I have written a paper on this, uh, aspect, responsible ai, which is, uh, to be published, uh, through IEE conference.
So if you want to get hold of more information around responsible AI and how to instrument responsible ai, that's, uh, that is the paper for you. But operationalizing a responsible AI through processes and policies in an organization, governance structure is not a choice anymore. To create, uh, standards for risk steering or monitoring throughout the project lifecycle is of utmost important.
Ensuring that checks and controls are in place to validate that all use cases are compliant to an organization or legislative regulations, which are also changing as we speak. Driving security reviews for gen AI platform technical architecture, for example, uh, audits for those architectures and recommending organization definitions of ai, current risk tier and AI security and privacy, uh, racing metrics and publishing it and making, uh, advocacy for it is also of at most importance. Now, I come to the last section of this presentation where we also have to plan for next 18 to 24 months.
So we have touched base upon a lot of aspects like, um, upskilling operating model change management. How do we handle AI led change management? How do you build, uh, racing metrics where, you know, AI is your quote companion identifying specific use cases or strategic use cases as your first pivot points, right?
Tools and technology. What kind of investment is needed in partners in ecosystem building in tech stack? How do we advocate, uh, the strategy for AI native, uh, tool stack and technology onboarding from an executive to, uh, you know, uh, cloud platform to DevOps practitioners is of utmost important.
And lastly, I would say risk management. Uh, I mean, we all know that vulnerabilities are getting introduced as we speak. So how do we navigate this vulnerable AI vulnerability strong?
And how do we put in place, uh, frameworks, best practices, industry standards, uh, advocacy groups, open source initiatives where we bring a hand and control on AI led, you know, um, or, or AI integrated, uh, vulnerabilities. With that, I would come to an end of this session. I would like, uh, to also introduce you to a few of the other initiatives where you can follow me.
Of course, I write for, uh, ML Khan Magazine. I also have published my blogs on Dev meo. Um, you can join me for my next workshop, uh, on the DevOps Con Conference.
And one of the biggest events, which we are, uh, doing in November is DevOps for Gene AI Hackathon, which is coming in Toronto, November 3rd. Stay tuned. Our good old friend John Willis, will be with us along with industry practitioners, teams and communities.
With that, uh, I hand it over back to the DevOps Experience Conference, and if you have questions, uh, please to connect with me off online or, uh, through my LinkedIn profile. Thank you very much. Network automation is the future of network operations, but what standards are you adhering to?
Are you even aware of what the standards are in this episode of the Tech Field Day Podcast? Automation Needs Standards? Welcome to the Tech Field Day podcast, where we bring together a group of influential IT experts from across enterprise IT to discuss hot topics in the technology industry.
This podcast is brought to you by Tech Field Day, which is a part of the Futurum Group, and is often recorded in association with one of our events. We're here today at Networking Field Day, and we'll be talking about automation and the need for standards. But before we do that, I wanna have our guests introduce themselves so you know who they are.
Starting with Denise. Hi, Denise Donahue, uh, network architect, technical author, all around Network Geek. I am, uh, Steve Plucka, network architect with, uh, DQE communications out of Pittsburgh, and also General Geek.
I'm Kevin Myers, also network architect, um, and, uh, general Network Geek, IPV six, geek, uh, service provider geek, um, routing and switching geek. Alright, and of course, I am Tom Hollingsworth event Lead for all things related to networking here at Tech Field Day. Let's jump into this episode.
No doubt, you have probably heard about the importance of network automation, whether you are starting that journey yourself, you're finding yourself in the middle of a long project or in some cases even wrapping up this, uh, amazing thing that you now have done to make your life so much easier. But were you following the standards the, the whole time? Did you know that there were standards?
Did you lie to me and tell me that you thought there were standards? Because there aren't in this episode of the Tech Field Day podcast, the premise is automation needs standards. So let's talk about this for a minute, because this was something, as soon as you guys said that this was the topic we wanted to talk about, I was immediately on it because the last time that I checked, there is no written standard for automation.
There are a lot of standard protocols and ideas that we use, but I find it funny that when you start talking about, well, automation, usually you get barrage with a whole bunch of questions of, you know, which scripting languages are you gonna use and which platforms are you gonna use? And is everything gonna be written in Camel case or Pascal case? Like, like you have to figure things out here.
What is it about network automation that lends us to being, for lack of a better term, creative with the way we implement it? Well, I think a, a big part of what it is is this development all along. We, we've been CLI jockeys forever.
Mm-hmm. Uh, and now that we're getting into the move of automation, the first thing we did was create a couple of particular protocols, whether you're using YAML or whatever standard in that sense that you're having to communicate with the devices. So now we've got enough of those going that we're now doing scripting and, uh, each vendor is individually creating a platform to do their equipment on.
And then there's a handful of, uh, multi-vendor companies out there that are picking and choosing which platforms they're gonna support. So I think we finally reached the critical mass where enough of this automation is happening and enough of the bits and pieces and tools are there that we need to get together as an overall community and create that standard on this is, this is how it should be abstracted, this is how it should, should work, and figure out who the hell's responsible for that. Yeah, it reminds me of the early days of the, of networking in the internet where there was just like, my protocol and your protocol and this, you know, and just everybody had their own idea about how things were gonna go, how data was gonna be structured and everything.
But then there was, at that point there was a small enough community, small enough group doing this that they could get together and make and create standards and fight over who was gonna win. Mm-hmm. I think the challenge that I see is that it, where we're at right now with automation is the vendors really drive automation.
Whatever vendor it is that you're choosing really drives the automation. And yes, there are open tools and they're open frameworks, but to the point of the podcast with the person sitting down and figuring out how they're going to assemble those is doing it in whatever way they think is best. And a lot of it comes from the coding community, which not all of those practices translate over to network engineering.
And so I think we do have, there are some standards that might provide some guidance, industrial control world, um, world of energy. They've had some automation standards around for a while, but it does take a bit to adapt those into networking things like the Purdue model. I think those are places to start.
If we think about, let's take automation that's specific to networking and standards, and what does that look like for an enterprise network? What does that look like for a service provider network and figuring out those best practices? Well, I think one of the challenges that you mentioned there is that two industries that kind of have, those are also very highly regulated.
And so a lot of the standards that kind of come out of this are not so much, this is the best way to do it, as much as if you don't do it and prevent these outcomes, we are going to sue you out of existence or fine you until the pain stops. Like, you know, for example, um, something like P-C-I-D-S-S drives a lot of the way that we design certain kinds of networks because we can't have certain things interacting with, with each other, or we have to have certain retention policies and things like that. And as far as I know, there's no kind of guidance in the industry right now of let's just say automating healthcare facility.
Like, you know, if, if HIPAA applies or if, if this kind of patient protection thing applies, that's gonna direct the way that you do things and create kind of a defacto standard, if not a deger standard. But again, that's vendor specific. You look at HIPAA and you've got the different, um, medical records vendors, you know, the, the, um, and they're, they're doing their own automation within that, which then they've extended on into the hospitals.
Um, so then does that gonna be the same? How's that gonna carry over to all the other companies, types of enterprises? And there's a difference between policy and standard.
Mm-hmm. HIPAA is the overall umbrella policy that's saying what you can canning cannot do. What we need as engineers and as vendors is, uh, a way to implement that policy and create the workflows necessary.
So far we've been concentrating on the individual task level things, uh, and we need to step up to, to workflows and, and processes that Yeah. They have to be driven by some overall arching policy Or business logic, Uh mm-hmm. Yeah.
On how, how this is gonna work and the categories you're talking about, right. As well. But, but that needs to be a team coming together and we need a, a group that does that, that doesn't be solely benefit.
It can't be the vendor, right? Mm-hmm. It has to be the, the engineering teams in general as a, as a group.
Yep. Manners kind of comes to mind. Ma NRS mutually assured norms for routing security, which is basically a lot of carriers are involved in it, CDNs, and it's a neutral organization that says, these are the policies that we feel that we need to implement for routing security on the internet and in the DFZ.
And then they take that a step further and have technical recommendations that if you're Cisco, if you're Juniper, if you're Nokia, here's how you implement that policy. Here are recommended guidelines. And that's not quite a standard to Steve's point.
But that to me is kind of along the lines of where you might need to go with automation is recognize the problem, define the problem, and then get down to those, those technical details of how do you implement this? And in my opinion, the challenge is automation is coming outta the world of coding. And what does coding love to do?
They love to jump to the newest programming language whenever there's a hot new language. I mean, how many have we been through? And I think that's the biggest challenge we'll have as network engineers, is we look at standards and protocols over a very long term arch.
The world of coding is always jumping to the next new thing. So how do you, how do you, how do you bridge that divide of how do we pick something, how do we pick a winner in the world of automation to create standards around that isn't going to get left in the dust of the world of coding? You don't wanna run my pearl script?
I sure Pascal, I want to go. Yeah, I want to do it in Pascal. Yeah.
Alright. So who's gonna do this? What stand, what standards body are we going?
Well, Before we get to that point, I think that that something that Kevin brought up that, that is very germane to your point, is there are groups that are kind of creating these informal agreements, manners, for example, it's in the name right? Mutually, um, Assured norms for routing security, Mutually assured norms, not requirements, not restrictions, norms. And we see this a lot in our society, right?
Like something like holding the door open for people. That's a norm in certain spots in, in the us but it's not a rule, it's not a requirement. And the difference is when you cross that line of a group of people getting together as an industry organization saying, this is how we're gonna do things, doesn't have any weight, because I can choose to leave that organization whenever I want and not abide by those rules.
But to your point, there are standards bodies out there Yeah. That provide guidance on how things are gonna be implemented in the us The two biggest ones are IEEE in the IETF. I would also add ISO as a global standards body.
Um, I, I would recite chapter and verse with iso, but nobody does that. These groups are formed from people who analyze the problem, decide on a, a method for recommended implementation. It is voted on, it is debated, it is argued, it is voted on some more and then it is released.
Why have we not seen anything from any of these standards bodies about this yet? You'd have to ask them. I, I would, I would say there hasn't been a demand.
The only one that I've worked with is IETF and I, I've done, you know, a little bit of work in the ITF and so that's the one I'm probably the most familiar with and it's also the most open. Mm-hmm. I think mm-hmm.
Yeah. ISO and IEE are, you know, they're solving real problems, very complicated problems. But that is usually, you know, there's someone is funding that to go to go do and solve that problem.
And again, that's a long arc, I think programming and coding and automation and are all there together and they iterate very quickly. So I think you have to start somewhere, whether it's in the IETF and we talk about the ITF is the right place, or whether it's a new organization. I think you've gotta look at that and say, how do you create an organization that can help to define those, you know, first maybe norms, then best practices, and then ultimately standards.
And with the IETF and the ITF is very, it's very protocol and operation specific. So I mean, I could see it going there, but I could also see challenges in, you know, in getting it through there. Which working group does it go into?
How does that, you know, I see pros and cons in it going into ITF and, And. Mm-hmm. You bring up a really good point here, and I'll reference everyone's favorite XKCD comic from Randall Monroe.
There are seven ways to solve this problem. I know I'll create a standard that encompasses them all. There are now eight ways to solve this problem.
Yeah, exactly. Like one of the, and one of the things that we've seen over the years, like IIEE is very focused on protocols, right? That's where we came up with ethernet.
That's where we came up with wifi and a lot of those other things. But with IETF, I feel like there is a lot of debating things until they're dead. And that's the way they eventually come to a consensus is, okay, are we done arguing about this?
Does anybody still care? Perfect. Now it's a standard because these are the, the remaining people.
And we saw that with TRILL and SPB that kind of became pseudo standards, but didn't really, and TRILL is actually a really good example of this because the, the definition of what was going to be IETF standard trill never developed because the vendors were so focused on making sure that their version was the standard. And I'll go out on a limb here and say that one of the reasons why I think that that happened was because there was no bad guy. What do you mean Power over ethernet and trunking on switches?
3 af? Because Cisco had a competing standard and everybody in the industry mm-hmm. Lined up and said, we want to do the exact opposite of what they're doing.
Oh, you want us to wrap the switch in a, in a, a trunk tag? No, we wanna put it in the header. 'cause it's what you're not doing.
Oh, you want to carry power over 5, 6, 7, and eight on the ethernet wires? No, no, no. We want to use 1, 2, 3, and six.
Because you're not doing that in a way. Standards sometimes are driven as a big middle finger to a competitor. So maybe what we need is Red Hat to come out and say, oh, well we're gonna use Ansible standard network automation, and the rest of the industry might wake up and go, you know what fu we're gonna go do things our way.
You know, I I think that you two have got, have, have some excellent points that kind of could be drawn together. Like the why part, what's driving it. You talked about funding, um, you talked about, um, you know, vendor pushing.
Um, you talked, you talked about, okay, you know, the, the reason for doing it just 'cause of a reaction, also a need for it. And it's obvious there's a need for anybody who does much work in this and is dealing with multiple vendors of equipment, multiple vendors of software. There's definitely a need.
Um, how are you going to get someone to do it? Is the thing That's the big challenge. 'cause I think you look at, um, you know, does it, we've taken the existing, you know, protocols, if you will, and we've take, we have a p we have APIs, we have mm-hmm restconf, NETCONF, all of different things.
And, and the companies that are out there have taken the existing available protocols, created a few new, and that's what we use. And then how you cobble that together is up to use. I think the real question is, is does automation, you know, whether that's, you know, automation control, that's, you know, pushing configuration, um, you know, telemetry related to automation.
Anything that's in the realm of automation, does that become an 800 series protocol or is it gonna be like an RFC? Is it a standard? And I think that's probably the first problem you've gotta sit down and solve is are we happy with what exists?
Or do we want to build something new that encompasses the entire world of automation? And the question then becomes, if you're gonna build something new, how do you get people to sign onto it? Exactly.
Mm-hmm. Because one of the things about a standards body is, is that when the standard is defined, everybody has to play by that standard or it's not a standard. Mm-hmm.
And you run into these situations where it's like, okay, there are four people that all believe that this is the way to do it. We have to pick one of these solutions, especially if they're not able to be merged. Right?
Like, you know, if the, if these two people have one that's pretty similar, we could probably merge them together. And then 50% of the use cases are kind of done like that. But it goes back to, well, if I don't like your solution to that, I'm not gonna be in par involved in this working group anymore because I'd rather do what I do because it's easier on my developers.
My customers prefer it this way. And in some cases, if I can get enough people to sign onto what I'm doing, I become the defacto standard anyway. Hmm.
Do your customers prefer it that way though? Or is it just that they have to take it? Well, that, that's part of it is if you don't know there's any other options out there, my way is the highway.
Yeah. Well, the, the, the customers are drowning right now. We're, we're just dealing with the day-to-day problem on how to keep hundreds or thousands of devices doing what they need to do and changing the way they need to change in a, in a reliable way with what we have in front of us.
So, uh, it's, it's time to take that step back and, and to, and to try to figure out how, um, I think the big difference between this, what, what we, the problem we have here and, and the normal standards is that it involves a, a workflow, um, type operation. It's, I, I guess the closest thing we've had previous to this is MPLS, where there's a whole lot of steps that have to happen to create this end to end, right? And everybody on the path has to play nice with each other for it to, to work in, um, in a vendor, in a, in a mixed vendor environment.
But even that is not as complicated as what we have now because we need to, you, you know, we need to do, you know, precheck change, post check, you know, rollback if necessary. Uh, you know, in the post check this, the, the whole workflow aspect of this that we sort of inherited from the developer world is the, uh, is the part that is alien to, you know, the average network engineer. Yeah.
And that's the biggest challenge I see is I don't, you know, I, I haven't written code and anger since the nineties 'cause I'm gonna decide to be a network engineer and I didn't wanna do coding. But, you know, now that we're in that world where code is very much a part of, of network engineering, the challenge that I see is I either go to a vendor and I have to go buy something and they're gonna tell me how to do it. Or I've gotta sit there on their gear only exactly right on their gear only in most cases.
Or I've gotta go out and look at what the development community is doing and figure out, okay, which of the thousand different ways that somebody has figured out how to build this framework, do I go, it's the build or buy. That's the challenge that you always have mm-hmm. As an organization is do I buy it or do I build it?
And if I build it, what is my am I doing? Because as a network engineer, let's do a protocol. Okay, well what they're like, you know, half a dozen routing protocols, there's layer two, there's a, there's a sit limit, uh, number of protocols.
You know, you're dealing within boundaries that are pretty well defined when you look at the protocol stacks of network engineering. But you look in the world of automation and coding and it's almost limitless in what you can do and what you can build. And that's the sea of uncertainty that I think network engineers are people, especially people at the network architecture level.
You know, where we came up in the old world of networking. And the newer engineers are, you know, nothing against them at all. It's just they're coming up in a different world of learning coding and network engineering.
We're looking back at a long career of here's why we did things this way, here's the choices we make. You've been through enough networks to realize that the choices that you made maybe weren't the right ones and you do it differently mm-hmm. Later.
So how do you take that knowledge and work with the network engineers that are coming into the field now to help them to also help inform the automation? And I think you need standards and frameworks for that too, as as to how to make and evaluate those decisions. The other thing that I think we need to be aware of is the trap that we can fall into by turning this over to a standards body.
This is something that's happened in the wifi world recently where we're going to introduce a standard and we're gonna make sure that everybody follows it, but we need everybody to buy in on the standards. So what we're gonna do is we're gonna take some of the hairier pieces of it and we're gonna make them optional. And that is something that has happened quite a bit in the last couple of releases of wifi, wifi succeed, and wifi seven where some of the things that make it good are optional and don't have to be, uh, implemented by you if you don't have the technology or the desire to make it happen.
And so what we're left with is a slightly better than the last one standard where we hope that everybody plays nice on these other pieces. How can we prevent a standards body from coming in to restore order and ultimately making things worse? Because the things that we need to restore order are left for optional because, well, if you don't make this optional, I won't sign on to the standard.
Ouch. Maybe that's why there hasn't been a standards body. And that kind of comes back to that whole how do you get everybody in the room to compromise?
And for the purposes of this podcast, everybody knows that the definition of compromise is when nobody gets what they want. Mm-hmm. Well, the other challenge you have is if you were, let's say you were to take this outside of something like the IET after, you know, ISO or IEE, how long does that take to, to build that organization to get people that want to contribute?
You know, when are you gonna, is it gonna be years before you see a work product that's, you know, that is relevant and, and going to help you? 'cause I think that's a whole other challenge of if you do create another standards body, what's the, you know, the ramp up of that is, is gonna be a lot. I'm sure they're working on finalizing the standards for building wooden sailing ships in the ISO like this year.
So, you know, there are only like nine centuries behind at this point. Yeah. Only need 90 bucks to access the standard.
Yeah, exactly. And, and that's the other problem too, is that a lot of times you need a lot of resources. You need to create working groups, you need to create, uh, people, you need chairs, you need folks who have disposable time.
And then for people to be able to access the standard, you have to have resources to, to do that. Either they're people who have access or that you pay for it. And, and then you create this rolling problem of, well, if I don't, if I can't see the standard, I'm not gonna follow the standard 'cause I can't follow a document that doesn't exist.
And how do you prevent that? I mean, we, we bag on the I-E-T-F-A lot because it feels like they're, it's just basically an excuse for people to fly around and argue with each other, but they do get things done pretty quickly because somehow all of that arguing eventually does lead to some kind of an eventual consensus. 'cause I think they realize they're not gonna get any more out of it than just the, the heated discussion.
They're not like an a giant organization like ISO or IEE that kind of direct things, but use that as a way to fund the organization. Yeah. They're, they're, they're definitely the right structure.
Um, for this type of thing. I'm just not sure if the, there's enough, there may need more participants that half of the half or maybe even more than half of the right people and companies are involved in the IETF. But adding in this additional layer of, of, of, um, of the workflow aspect of it and understanding the history of it too, you know, saving those configurations, which we've never really historically done, you know, to, to compare to, you know, the, the pre and post check access of the, of the workflow.
These are not traditionally things that network equipment vendors are, are good at or know They're, they're good at getting you a box that does a thing. Yeah. Yeah.
Yeah. And, and that's the, that's sort of the missing piece of the IETF, uh, the, the fundamental section I think has worked so far. You know, we have, you know, the, the net off, the yaml, the all the options available for the communications and the, uh, and the saving.
What we don't have is the, uh, is the discipline of the, of the programming community in the workflow as aspects of this and the history aspects of it so that we can see what changed and, Um, Yeah, and, and why I think it's also the, um, it's getting the, there's very few people that have a very, very long deep network engineering background and also have the coding background that understand how those worlds come together. So I, I worked on a project one time where a coder was off trying to solve a problem, and he spent like a month or two writing code. And what he had we'd written was Radius, he had built Radius.
And I said, you know, there's a protocol on the router that already does what you're, what what you just spent two months writing. He's like, what is it? It's called Radius.
You just go turn it on. And that's where you can have a brilliant programmer that understands coding and even understands trying to solve a problem. But if they don't have the domain knowledge of what is possible in network engineering, then you may spend, you know, you're sitting there beating your head against the wall solving a problem that's already solved.
And I think that's just as important as getting a standards body and defining the standards is what already exists that we want to leverage so that we're not reinventing the wheel. Mm-hmm. Mm-hmm.
And, and ultimately I think that that is where we're at is that as networking folks, we have worked very hard to create these small areas where we can standardize on certain things. Whether they're things like Radius or they're manners type, uh, assured norms. What we need is a neutral third party to step in and effectively get everybody's ego out of the conversation.
This is how things are gonna be done. I get that you're not the way, that's not the way you do it, but this is how we are going to do it going forward. And if you can have that happen with someone who has the personality to pull it off in an industry where people are going to have to agree to disagree about certain things and put their competitive advantage aside for the betterment of society, I think what you'll ultimately find is that we can standardize automation if we're willing to put in the effort.
That will just about do it for this episode of the Tech Field Day podcast. We want to thank everyone out there for following along. com/podcast.
You can also find show notes and bios of all of our guests. com for more information on our upcoming events. com for more information about the RUM Group.
We'll see you soon.