Understanding the Impact of the AWS Outage & Cloud Infrastructure Challenges | TSG Ep. 949
Mike, Mitch, JP Morgenthal and Stephen Foskett, president of the Tech Field Day arm of the Futurum Group, dive into the lessons to be learned from this week’s global Amazon Web Services (AWS) cloud outage.
Then the gang takes a look at NetApp’s evolving data management strategy in the age of artificial intelligence (AI) and how Kubernetes could be better managed using Talos Linux in the wake of two conferences held last week.
Transcript
AWS outage, outrage, or just the cost of doing business. You're watching Textron Gang. We'll be back in a minute.
Hey, folks, we're back. Everybody woke up Monday morning to another AWS outage. This one may be bigger than most involved, a DNS issue that was trying to access a particular database in an East coast region.
So there was a single point of failure, but everybody's kind of running around with their hair on fire for a couple hours in the morning. And, and then finally, AWS said they fixed the issue. And it took a little while for that to come all back online.
But JP, let's start with you. Is this just the new norm? Is this what we're gonna have to expect?
And sometimes these things are bigger than we anticipated, but for all the talk on these clouds, there are single points of failure. Well, I don't know that it was a single point of failure. It sounded more like it was a, uh, uh, cascading effect.
Um, and it, it, it, it didn't just take out, um, that it one entry point. It affected other entry points. It's more like a virus in many regards and how it responded, right?
Being able to affect multiple services, not just itself. Uh, I, new norm. It depends, right?
I mean, you have a choice. When you leverage, uh, hyperscalers, one of your choices is that you are gonna own all your own infrastructure and, uh, just leverage their hardware, uh, much like, uh, data, you know, or outsourcing your dataset. And, uh, in that particular case, if you had done that and you had cloud flare backup as your DNS, and you were at all your own database servers running in virtual machines, the likelihood that you would've gotten impacted by this would be very low.
However, the money would be flowing outta your pocket, um, faster than you would like. Uh, the reason you go to the cloud is to leverage pay for consumption pricing models, uh, which means that you're now dependent on the cloud service provider delivering a certain level of availability. And, uh, I happen to know that AWS is really, really good at architecture, uh, and design, uh, of scalable, highly available clouds.
Now, that doesn't mitigate the fact that occasionally, um, there is human error or, uh, or a mechanical failure, uh, that was unaccounted for just you, you couldn't see that this, if, if, you know, if this occurred, you would see these cascading effects. And I think that's what we're seeing. These are edge conditions.
It's happening less and less. I mean, I, I saw reports of like, this is the le you know, this is happening. It's not happening more and more the last major out of 2024.
And yeah, we're gonna hit edge conditions because we can't know everything and equipment ages out, right? So we're learning how to run clouds, and at the same time, we're painting the bus. This is going down this, this, the road at 80 miles an hour, because more and more people are coming onto that cloud.
So, uh, I, I think what you're just seeing is, you know, learning, you're seeing culmination of effects, and you're seeing a organization that is designed to have mean time to failure or, uh, mean time to repair at a very, very high level. At the same time, it's interesting how these effects can cascade throughout the internet. I mean, I, I disagree with you.
I I wouldn't characterize it as a virus, simply because I don't wanna scare mom and pop at home. But that being said, it is interesting that so many services are dependent on this. Um, you know, my wife came, uh, woke up this morning and found that she couldn't successfully log into Wordle.
Um, maybe that's an AWS problem, probably. Uh, another one I saw on Mastodon, uh, somebody shared an account of, uh, being woken up by their cat this morning because their wifi connected cat feeder didn't feed them because AWS was down. And, um, I, that sounds plausible to me, and that sounds like 2025 to me when, um, I, I could totally see a situation where people are reporting issues with, um, electronic doorbells, um, cameras, logins, and, you know, not, not like, oh my gosh, I can't get into the bank, but oh my gosh, I can't get into the house because everything runs through AWS uh, probably the most telling was, uh, the old Cismen joke, which is, none of our systems should be affected because we don't use US East one and one of the other cismen says, oh, except for that service, you know, and then nothing works.
Welcome to the world. And I just dropped Something. It, it was, uh, so it was around DynamoDB and I think a lot of backend, uh, app, uh, backends for applications, leveraged DynamoDB as a data store for, uh, you know, the state.
So your Wordle app, right? And somebody quickly, that's an app that probably doesn't have any, you know, it's not transactional. So, hey, yeah, let's just use DynamoDB.
It's quick, it's easy, it's cheap. Um, so when you take out DynamoDB, I think you're taking out a lot of state management for a lot of apps, which is why half of it might work, and half of it may not work, because when it goes to get the state, it's like that's when it hits the wall. Daniel DB is an extremely popular service in AWS, it's one of the most frequently used databases, but I believe it's also part of used and manage their infrastructure.
To your point, JP, about state, because this outage affected EC2 and S3 and some of the operations of the infrastructure. Um, I, I think the analogy that works for me the best, just thinking about how you guys were kind of how to, how to describe this. It's like, it's like, uh, cascading, uh, brownouts in the electrical grid, right?
One, one over capacity, or something happens in one part of the grid grid that can actually have a whole cascading effect all the way up to, you know, larger parts of the grid. And in a sense, that's kind of what DNS is, is apparently this was a DNS issue inside the Dynamo DB service, uh, for this API endpoint. Um, it wasn't part of their Route 53 db, uh, DNS service.
Um, but DNS takes a while to propagate. And so when errors happen, they go out and they propagate, and then they have a problem, and they get bigger and bigger because they're propagating. Meanwhile, back at the ranch, you're like, okay, we figure out what it is here, push this back out.
Well, that's all controlled by a time to live parameter of how long does the cash live of those DNS settings. So it is one of these rolling both the ca the outage and the bringing it back up takes a while because even though they may have fixed the DNS with the servers at Amazon, there's a lot of caches out here that have the old information in it that are gonna wait till that, uh, time to live expires and then refresh. So lemme ask Steven this.
There's some folks who are saying that we need to fix DNS and that some of this stuff is too antiquated to scale, and is that something we can actually do? Or is this, we're stuck with it 'cause it's been around so long? Well, it's funny, you know, I wouldn't put it on, you know, I wouldn't say that we need to fix DNS writ large.
Uh, I do think there's a need to, um, figure out a system to make it a little bit more resilient to situations like this. Um, but, you know, I mean, D-N-S-D-N-S isn't broken. Something broke in DNS and, and I think that, um, you know, as JP started off with, I think the thing that's most incredible and remarkable is that failures like this are not increasingly common failures like this are surprisingly rare.
And, and especially when you give, when given the situation where literally anyone in the anyone in the world with the technical skills can participate in DNS and push updates to DNS, uh, that would lead us to assume that this thing would be incredibly fragile. But it's not. I think that, um, ultimately what's gonna come down here, we'll find out that the, the root cause of this was somebody pushing something out that they shouldn't have pushed out.
Um, And frankly, I, I don't think we're gonna fix that. No, and, and it's funny. D-D-D-N-S is, uh, you know, pretty resilient.
I, i, it, it actually, my understanding is that the whole idea of moving, uh, from IP four to IP six was that issues like this would be easier to resolve. Uh, there's a but that never occurred global scale. We are still a world that is heavily dependent upon IP four, and until we get to IP six, we're not gonna have all those nice advanced features, which is like being able to basically move your network, you know, in seconds.
Entire networks can just transform themselves dynamically through the IP infrastructure, right? We are left with old laggard arc net created infrastructure. That's like, I, we really can't do much.
We, we, we've pushed it as far as we can push. DNS now has DNS sec. You have to have, you know, you know, trusted certified pushes in order to affect.
So we've taken away the ability for anybody to just screw it, the entire infrastructure of the internet and shut it down. Um, but there's only so far we're gonna go, because IP four has its limitations. Mitch, I think I heard fighting words from Steven.
He was blaming the DevOps people for pushing something out. If it's not the DevOps people, it's the developers. You know, so we, we know where to go to, to, to blame next.
If we the DevOps, Let's blame the DBAs instead. Can we blame the DBAs? Please?
Uh, you, you know, you think about it, it's, it D-N-D-N-S, it's amazing that DNS works as well as it is because it is a very, it's a very cascading kind of hierarchical system about how zone pushes, zone transfers happen, uh, to push updates. And, but it's also not in real time, meaning you can't just say, oh, quick changes and push this out, and it hits all the servers and it hits every server out there. It doesn't have sort of that emergency oh, oh, crap setting.
It's the, it's in the header of the packet, the oh crap packet, um, to say, update this. So it's, it's, you know, yes, it's, um, long in the tooth. It sure runs our network pretty well today.
And yes, there's an advantages in IPV six, but it shows you how hard it is to change from one generation of technology to the other. So I don't, I don't know who will get the blame for this. I'm sure some, you know, someone will be held up as they, whoops, that was not the process that we're supposed to follow, probably, you know, something like that.
Skip to skip or, you know, usually you have ways of prevent, you know, like you said, Mike, this doesn't happen. Or I guess Steven, this doesn't happen very often, so it isn't like we're seeing lots of DNS errors. It happens, but not very frequently.
There you go. Jp, one other thought here. Is it feasible to use another cloud to back up AWS or is that just too costly and too complicated?
And most people aren't gonna do it? I think some people do it for mission critical stuff. I don't think it's plausible, um, to do for your app infrastructure, for example, um, we talked, uh, when I mentioned earlier, you, if you're running your own database, fine, you could probably run that in AWS and Azure simultaneously, and you, you know, uh, hoping that you get replication in the DNS quickly enough, 25 minutes, you could switch your, uh, you know, globally switch your pointer from AWS to the Azure instance.
But you're paying a lot for that, a advantage. And, and for a very limited subset of issues that might occur. Now, if you have SLAs that are gonna be costly to you, if you are down, if you don't meet them, then it's worth it.
Because the fees you could be paying to your customers for being down are high and, and outweigh the not doing that. 98, let's say, okay, which is still extremely high. But let's say you had that, I forget exactly how many days a year that equip allows you to be down, but it's like 12 or something, right?
Then, then take the one cloud live with the outage, it's gonna get fixed quickly, right? And, and deal with the issue. So it, it really, I think it's a, it's an economic decision for the business to want to do that.
But yes, it is plausible to design your application to leverage two clouds and with one cloud being the backup of the other cloud. You know, Mike, to that, to your point of your question, I think one of the cascading effects of this will be customers reevaluating, not necessarily leaving AWS if this was a chronic issue, they might consider changing, but, um, reevaluating their own application architecture and infrastructure that you're using, they're using in the services. And, you know, maybe they wanna switch from a dyno DB to a Oracle or, or SQL server database that they can replicate and, and coordinate, have fail over across different cloud services.
Or maybe it's, there's things about their applications that are, they can shore up to make it more resilient. They'll, people will at least look at that and say, okay, this could happen again. Is there something we could do on our side, uh, either less than the pain?
I wouldn't bet that they're gonna do that. I would, I would bet that they're gonna leave systems exactly as they are, and we'll have another outage in a year. Oh, I don't, I'm not saying they're gonna change.
Everybody's gonna go whole scale change things. I think the, and the organizations I I work in, when an incident happens, you, you step back and say, okay, right now, what about our part of this? Is there anything on our side?
We need to, are we dependent too dependent on Dynamo DB or whatever? And you may or may not decide to do anything, but it causes you, if you don't ask the question, you're at least, I think you're derelict in your duty, your responsibilities. All right, folks, I think we've chatted about this long enough, but I would just remind everybody of one thing.
They push these updates usually out sometime in the middle of the night. So you, if you're in it, you're the one who's gonna get the call in the middle of the night. We'll be back in a minute.
You've Earned it. The spotlight, the responsibility, the weight of teams, companies, and entire industries fall on your shoulders, lives depend on your decisions, your home life included that work. You are protected physically and digitally.
Nothing gets through your team without a fight. But in a globally connected world, everyone sees you, including those who mean to cause you and your organization harm. And now home your sanctuary attackers see an opportunity.
Your digital front door is wide open. And what compromises your home can breach your boardroom. Because the devil's greatest trick isn't targeting your workplace firewall.
It's convincing you that your personal life isn't at risk. Black cloak, digital executive protection, defending the new attack surface your personal life. Hey, folks, we're gonna talk now about what NetApp is up to, because the folks at Tech Field Day were at their conference last week, and there's a lot going on with NetApp in terms of not just storage, but they're managing data and expanding their ambitions.
And we, of course, have a story about all that on Techstrong it, you should check that out. But Steven, you were there. Give us the download.
Well, yeah. This is, uh, my, in, uh, NetApp Insight Conference, uh, in, in as many years. Um, I've been there a lot.
Uh, I was actually one of NetApp's first customers, and it's been interesting to see what they've done in three decades of development with this platform. Um, you may not know NetApp, but essentially they are the, um, probably the biggest, most popular storage platform in the world at this point. Um, dedicated storage.
Now, that doesn't mean that they're the biggest, uh, probably S3 is bigger. Uh, it doesn't mean that, uh, they're the biggest in enterprise tech either. Uh, companies like Dell and HPE sell an awful lot of storage.
But, um, you know, NetApp is a dedicated storage company, and it's always interesting to get their angle on it, because, because they're not part of Dell or HPE or Lenovo or whoever, uh, because they're not part of AWS or Google, they're in a unique position because they have to play nice with others, and they always have had to play nice with others in terms of protocols and open source and partners. And that was, uh, really one of the big takeaways for me, uh, that they are continuing to, um, to work, uh, really, really well with partners. Um, for example, NetApp has native storage integrations in AWS, uh, Google Cloud and Azure.
And each of those is different. Uh, it's not just that they're running their thing in all three clouds, it's that they made a thing that meets the needs of the customers of each of those clouds, and they're all compatible, but entirely different products with actually confusingly different names. Um, and, and in many cases, uh, oh, they would admit that.
I think in many cases. It's interesting too because, um, their stuff is so integrated that like, if you're in Azure and you want like a, you know, a, an NFS server, you, you, you might end up being a NetApp customer without even knowing it because it's just part of the dropdown box. So it, it's a really clever platform play for them.
And my question going into this is how does this extend into ai? Now, my fear was that we were gonna sit there and NetApp was gonna tell us, we're not a storage company, we're a data company. We understand data, we're an AI platform.
Thankfully, that did not happen. Um, and, and I gave the everybody from the CEO on down a big pat on the back in private conversations about that, because frankly, uh, they ain an AI company or a data company. They're a storage company, but they're using AI in a clever way.
Um, one of the biggest challenges, as we've talked about here, is the challenge of unstructured data. Multimodal data. You know, essentially, companies have a lot of database data, which is structured, but they also have way more data that's images and files and whatever, wherever, wherever it may be.
Um, how do they leverage that in this AI world? And that's what NetApp is using AI for. So they launched this new platform that's the A FX, that's all singing, all dancing bells and whistles.
It does the thing, but more importantly is that they have this AI data engine that runs in front of their storage. And its job isn't to run AI applications, its job is to classify and categorize and bring structure to all that data. So essentially they're using AI to scan and categorize and tag data, and then they present that to AI applications and get this, the, one of the net one of the engineers at NetApp actually said magic words in my ears, which is, there's no way to secure LLMs.
The only way to keep your data from leaking is to keep your data from getting into the LLM in the first place. And I just wanted to hug him because yes, yes. That's the only way it works.
And that's essentially what NetApp's focus is. They're trying to figure out a way to basically present this data in a safe way to AI applications instead of just saying, okay, here's all the data. Better not do anything bad with it because you know what's gonna happen, right?
I think it's interesting 'cause we're seeing, uh, Steven, the, the data platforms, I too was an early NetApp customer, and they have some of the best data management software, I think of any, any of the platforms. But you're seeing, uh, data, you're seeing storage companies and database companies come into their own in the AI era, because everyone's talked about, Hey, how do we get to this data, but also how do we secure it and protect it, and the privacy and governance and all of that. Well, that's all built into these platforms.
It's built into NetApp. It's Oracle had a similar announcement, a little bit different, but they had a AI did a platform announcement that also categorizes data and creates, uh, vectorized information to be able to use that. But it's designed to contain the data within that environment and not into the models.
And I think that's what large enterprises are, are gonna enforce religiously, is we don't want our data escaping into models. And those platforms are critical for doing that. I'm confused, Steven.
So do we want storage to be managed separately or do we want storage and data management to converge in the age of ai? Well, I, I think we gotta be smart about it. Um, I think that there's this, there's a role for, um, object storage in the cloud.
I think there's a role for, you know, local object storage. Uh, I think there's definitely gonna be forever a need for structured data for databases, uh, whether they're SQL or no sql. But I also think that there will forever be a need for unstructured data stores.
And frankly, those are really hard to build. I think that one of the things that has become incredibly apparent as we've seen companies, you know, cloud companies as well as enterprise companies and just sort of new startups try to enter the storage space. Storage is deceptively hard to do.
And I think that, again, it's one of those things where it's hard enough that you have to put incredible resources into it. Another, uh, private conversation I had at NetApp, um, that I'll share with everybody because I guess I don't respect privacy, is, um, about Google and AWS and Azure. Um, they have put the same kind of developer resources into their storage platforms that companies like NetApp and Pure Storage and Vast have put into their platforms.
In other words, uh, AWS is as much a storage company as NetApp is simply because they hired enough people and put enough manpower and put enough effort into developing a credible storage solution. But even them, even they are looking and saying, you know what? We don't do it all.
We don't want to do it all. We want to work with companies like NetApp and Pure and Vast and CCA and so on, uh, to integrate their storage with our thing. Because ultimately, in answer to Mike's question, ultimately we will have dedicated storage even as we're developing our own really good storage Products.
All right, jp, are you buying that argument? What do you think we should do with storage in, in the age of the cloud? Because when I go up on the cloud, there is no storage admin.
So, Uh, it's, it's been a, for me, storage is, as Steven said, it, it is a deceptively complex area. There was a time before I worked at EMC where I was like, dude, you hook up a disc. How hard is this?
Right? It's like, and then you start to get into, you know, all, all, you know, everything from like how much data can be transferred to, um, you know, is the data, you know, is it networked or is it not? And how is the data stored?
How do you, you know, how do you retrieve the data? How do you accelerate and optimize how, you know, storage of large volumes of data? And then on top of that, we have volume sizes, you know, increasing exponentially, and people are just dumping stuff.
I mean, we're up to petabytes of data just being dumped. And, uh, frankly, I, I would go on to say that it, it's, now we do, I'm sorry, we have zetabytes, you have zetabytes of data. And I would say it, you know, if you look, took a step back and you looked at that zettabyte of data, it reminds me of the garbage hills that I used to see when I would go to Staten Island where they dumped all the garbage.
And one day, one day they'll flatten that and build something on it. But for now, it's just a huge smelly chunk of grass, right? And, and, and so it's, you know, I don't think people cultivate that, those, you know, their data well, I don't think that there's a lot of organization that goes behind that.
We certainly haven't done a very good job of expiring data, putting, you know, doing content, real content management about what we put out there. So stuff will save for years and Lord only knows what's in there. Uh, we haven't done a very good job at looking what's in, being put out.
There could be keys, could be passwords, could be anything. So I think in that mound of zettabytes, you know, there's a subset of data that has been protected, shielded because of its value. Its value has been identified and cultivated, and then the rest of it's all just sitting there and waiting, waiting for something to occur.
If the platform, that's a good use for Ai. You know, I mean, that's the thing. It's, it's, it's u useful to use an LLM to Do that task.
Yes. Yeah. And now it is a great opportunity to just stick something on that, that an agent to just, your job is to go through zettabytes of data and determine what's out here, what has value, and help organize and, and classify it.
Um, that would help, that would be a good first step. And I don't think, think a human was ever gonna be able to do that. So, Oh, it's, it's entirely the amount of data.
It, it data is like the wiring closet for your network. It looks great the day you did it, and you shut the door. As soon as you shut the door, it's gone.
Chaos, right? It automatically will start changing on its own. And data has that sort the same evolution.
We have really nice models that we use to create databases. We have unstructured data, but those things also evolve 'cause all kinds of exceptions and things happen over time. But now what we're lacking is today the seman, the semantic meaning of the, the data is encapsulated in the software that accesses it and does things with it that doesn't help AI models use that data.
We might have, uh, the same field or column called account in 20 different databases, but they mean something very, very different. Maybe that's formatting of the structure, maybe it's account of what all these kind of things. So there's companies are going both SaaS companies and, uh, we have to do this with our own enterprise apps, is go back through and say, let's create the kind of data dictionary used to call these, the semantic meaning of what is that, wouldn't say it's the summary or it's the sum, or it's the average, well average what?
And putting some definition around that so that when you are trying to use it from an LLM through into the data platform, it knows, okay, I'm asking for this information. Okay. That I, that's what I'm looking for.
I can use it. Yeah. I'm actually gonna take a, uh, a second to, to do a gratuitous, uh, call out to myself.
Uh, I, I wrote a book in 2005 called Enterprise Information Integration, A Pragmatic Approach, which really looked at leveraging semantics to do integration way too early. For most people, it was ontology, it, it's root was building ontologies, which is like, you know, telling frogs to build castles. Um, and now the technology has arrived that can actually help.
Uh, and, uh, so I am updating that book, um, for the AI era into talking about how, uh, how this, you know, you can now leverage AI to drive integration through semantics. Uh, and it, and it, and it's a really interesting story. I agree that every, the roots of everything, the value of that AI is bringing to this picture is the ability to, to enrich all that data that's out there with additional metadata and understand it All right?
I'm sure there's one copy in the, in the, in congress somewhere that everybody has that book, but, you know, they're looking forward to the the next signing. Steven, I do have a question for you though. Um, so I see all the database companies to Mitch's point suddenly, or in the data management space, and I see all the storage companies are pretending to be data management companies.
At some point, is there gonna be like some mega mergers involving all these databases and storage companies and we're gonna see something interesting happen and, uh, all in the name of ai, Um, maybe, uh, but I think the Omega mergers will have a lot more to do with businesses and, uh, wall Street and venture capital and all those kind of things than it will have to do with the need to integrate these things. I mean, frankly, um, we've already, uh, mega consolidated it. Lemme tell you, going to a tech conference in 2025 is a lot different than going to it in 2015.
Remember JP, when we were, you know, we would go to these conferences and there would be, you know, 80 companies all doing roughly the same thing. There ain't 80 companies in the whole industry anymore, it feels like sometimes. Um, so I do worry a little bit about kind of where that, where that is heading, because frankly, um, you know, again, as a storage nerd, I'll tell you, storage is hard to do.
Uh, security is hard to do. Databases are hard to do. All these things are hard to do.
And it's one of those interesting things that you, you know, you watch companies come, you watch them announce what they're doing, uh, and you watch them make the same mistakes that have been made so many times before in terms of availability and, and how to build products and how to deliver services. And, uh, you know, you kind of wish that, uh, that they could maybe learn the lessons of the past, but it seems like we humans are, are pretty bad at that. Um, one, one last thing I'll give a plug for though, uh, we did livestream, uh, the, uh, tech Field Day presentations on, uh, techron tv and we're gonna be posting the video recordings, uh, by the time you see this, they'll be posted on YouTube as well as, uh, running on the techron app.
Now, now, I haven't kept up with the details of, of the acquisition and where pieces have gone, but when Oracle acquired Sun, uh, interestingly enough, uh, they internally were working on the o Oracle cloud, which was built because they wanted to be the cloud provider to deliver, uh, bare metal opportunity as much as, uh, shared, you know, uh, virtual machines. But they also, when they acquired Sun Acquired Storage, uh, and now, uh, and we know that they're heavily involved in leveraging their, uh, OCI architecture, which is closer to bare metal with GPU architecture and giving some very, very high performance. And now, and of course the, the Oracle database company wing it, you could envision that Oracle could be looking at creating a consolidated, um, you know, entity that, that brings AI storage and the data all into one appliance at some point.
Sure. And you could think of somebody like HPE or Dell or IBM doing the same thing, right? Right.
It, well, let's leave it there guys, but it's an interesting theory. 'cause if JP is right, then basically Sun Reverse acquired Oracle and we'll see where it from there. Hey, hey, Mike.
Just a quick shout out. Be sure folks should check out futurum signal on data intelligence platforms. Uh, Brad Shiman put together, uh, we're using ai, so kind of a little farther up the stack in the data management data intelligence part of it, but it's an excellent report.
You can go search for rum and signal and find the page to download it. All right, and with that, we'll be back in a minute. Discover Textron Group, the epicenter of tech innovation.
We are your go-to for reaching IT leaders and practitioners worldwide. Our secret impactful content that sparks awareness, engagement, and top quality leads with us. You'll access editorial websites, streaming videos, virtual events, custom content analyst research, and more.
Join our satisfied clients. Let's revolutionize your tech journey. Contact us today and tell your story to the world in the most powerful way with Textron Group.
Hey folks, last week I was in Amsterdam at an event called tuscon. It was hosted by Sedera Labs, and they were talking about something called Tali Linux, which is a distribution of Kubernetes that comes with a lightweight distribution of Linux embedded in it. And they stripped out all the stuff in Lennox that has nothing to do with Kubernetes.
So basically they reduce the size of the attack surface, and they make the whole thing a lot more efficient. Um, Mitch, I was really surprised when I got there because I thought this would be a small event. There was about 500 people there, and they had companies like Roach was in there and some of the biggest retailers in, uh, Europe, and it was very much a European driven event.
But, um, I think they're onto something. And I was wondering here, is this something that Red Hat and everybody else missed as an opportunity? Because, you know, today they make you get Kubernetes and they make you get a lightweight distribution of Kubernetes and, and Lennox, sorry.
And you're supposed to figure it out. So these people were all saying, we're sick and tired of figuring it out. We just want one thing.
What do you say? Well, the, the options have just kind of exploded of what you can use. And some of it is, you know, caused by the alternatives to VMware kind of problem of folks that might want to consider modernizing or doing something else.
But there are so many options, um, because you just take a company like Red Hat, you know, they created a slim down version of OpenShift. OpenShift is pretty complicated to set up and run, but they created a lightweight version of it. Um, SUSE has a K three version of Kubernetes.
It's a smaller footprint, light, light footprint, and good for the edge. You have other alternatives. You know, in addition to KVM for virtualization, you have linkerd for containers.
Um, you have of course, open to is an alternative to Terraform. So this whole infrastructure software layer of managing it, but also programmable or infrastructure's code is really a big, big area for platform engineering teams. People that are managing the software part of the infrastructure below the application stack, kinda between the operating system and, and that stack.
So I think a number of players are stepping into this to say, yeah, we need the heavyweight solution, but we gotta have the lightweight solution or the, let's say, less friction solution to implement. Ah, jp does this speak to anything that you've experienced in your career, or do you think, you know, this is just something we're kind of looking at as a momentary blip and everybody will do the same thing soon? No, I, I, you know, I, I see it moving more towards, uh, appliance, which I would, I wondered what took so long.
Uh, I, for years I've wondered why people wanna manage things at a YAMA level, uh, manually, right? This is, uh, you know, well, at least with AI and, and now you have, uh, CICD pipelines being built by ai, A lot of that, uh, overhead was, uh, moved to, uh, or out of the need to have a human actually build those files. But they're still being used.
Uh, they still have errors in them when they're done by LLMs, they are not perfect. Um, and, you know, so this, you know, the, this is the opportunity to do, to say that a cluster is a box. You just turn it on, plug it in, which is kind of where Kubernetes really needs to go, uh, because it's just too, it's too, you know, expansive to manage at, at a scale when you're doing something at scale, it just takes too much effort, too much management overhead.
Um, and it's, uh, it's unclear, you know, the, there isn't like one way to do it. So I, I, I think this is an opinionated approach to Kubernetes, which is fine. It may not be, you know, it may not fit all applications the same way, but it will work for a large number of organizations and reduce the overhead and labor overhead for managing a Kubernetes cluster.
Yeah, Mitch, they were also saluting GI ops and they were alluding the fact that they were gonna put some sort of application delivery platform on top of this thing, and, uh, would be more probably some type of CD capability. And I'm guessing it might be some open source thing that they're just gonna implement. But, um, is this another example of where the CD part of CI is starting to peel away from each other, and we're starting to see CI on one side for developers and CD on the other side for operators?
I think it is, and it's also just representative of the complexity of the environment. It's partly what JP was talking about, you know, all the configuration files. And in a GI UPS approach, basically everything is not only, you know, version controlled of what the infrastructure is with the code to build it, or the terraform configurations to build it, or the Kubernetes configurations, all of the above.
And then, yes, we automate those, you know, with tools, Ansible, and other tools today. But I think with ai, we have a much, a greater opportunity to not only automate, but automate the deployment of IT with human control, with human intervention of that environment. And I think we're gonna move away from a day of where it is stick-built by humans, you know, sort of the stick-built house and using not just automation, but more intelligence systems to help us configure, manage, and update these kinds of things.
Of course, they need to stop hallucinating a little bit before we can do that too much. But that's, that's the beauty of a good ops approach. It's not an easy thing to shift to.
It's, it's definitely a paradigm change. Yeah, it, it's interesting that, uh, there are so many different, uh, attempts to make Kubernetes more enterprise friendly, more developer friendly. Uh, you know, you talked about SUSE and K three s.
Um, I'm a big fan of Rancher before it became, you know, uh, what it is now at suse, I still like the, the product. I still use the product. Um, you know, Docker went in that direction.
OpenShift went in that direction. Um, all of these are, I think, attempts to sort of break that boundary down and give platform engineers a consistent and, uh, uniform, uh, approach, you know, a a platform to work with. And on the, on the, a AI front, like Mitch, you were just mentioning, um, one of the aspects that that is intriguing here is that, uh, there have been examples of proper Kubernetes, uh, Docker containerized, you know, NextGen, uh, web application configuration out there for long enough that most of the modern, um, LLMs were actually trained with such a library, uh, you know, corpus of knowledge about the proper way to configure this stuff.
I found that LLMs are surprisingly good at platform engineering, simply in my opinion, because there's so much good examples out there of platform engineering. They've, they've, they're able to do it. So it's one of those rare instances where I've actually had a very good, uh, inter, you know, use of LLMs in SIS administration.
All right. I talked to folks about using LLMs in Kubernetes, and they said it was just too damn complicated. And that the lms, you know, they could explain a few things, but they had serious doubts about using it on the operation side.
But we'll see how that goes. Um, I also wanna point out, I, I've u I've U I've used it for against, I used it to d develop and, uh, deploy on, on Azure. Um, on AWS I've had it build biceps.
I've had it build cloud formations. I've had it build Terraform scripts. And, uh, for the most part, not only is it really good at getting the baseline architecture and deploying that, um, if you tell it that the app that was deployed isn't working, uh, I've seen it go through and it knows how to look for problems that you typically have in a, uh, multi-tiered application between the front end and the back end and, and deliver and test all of those things looking for the problem.
So, uh, I, I can't, I don't know if that was older data, but I, I've, I would even agree that it, it is one of the things that it does really well is manage infrastructure and deployment and deployment of applications in that infrastructure. Funny enough, that is gonna be one of the major topics at CubeCon in Atlanta and Sedera Labs is gonna be there. And so are we, and we're gonna have a bunch of interviews and all kinds of fun stuff going on there.
So please come to Atlanta and stop by and say hello to us if you see us on the show floor just for grins. But I want to thank our guests for sharing their knowledge and insights as usual. And of course, we'll be back again tomorrow.
Now. Stay tuned for the rest of the Techstrong TV lineup.