The 2024 Kubernetes Benchmark Report | Predict 2024
The 2024 Kubernetes Benchmark Report is out! Join us as we review the results of over 150,000 scanned workloads to learn what’s working and what needs to be improved. Join to see how you compare and get advice on what you do and don’t need to do.
Key Takeaways:
Benchmark yourself against others, see what you are doing well and what you need to improve and get past issues that may be limiting your ability to gain the full value of Kubernetes.
Transcript
All right, I'm, uh, really excited today to share a little bit about Fairwinds Kubernetes Benchmark Report for 2024. Uh, we're really, uh, kind of proud of some of the work that our team has put together, uh, over the past, uh, few weeks getting this report ready for the new year. And, um, excited to kinda share some of the findings that we observed, uh, by studying over 330,000 workloads.
Um, so just as a quick introduction, my name is Joe Pier. I'm a product manager here at Fair Winds. Uh, been in the Kubernetes and container space for, um, a little over five years, um, and, uh, been working on the Fairwinds Insights product.
Um, like I mentioned, we actually, uh, have done this report now for, uh, this is the third year. And, uh, it's the, we've analyzed over 330,000 workloads, uh, to inform the data behind this report. Um, it's the largest number of workloads we've ever analyzed for the report, and it includes, um, more than a dozen different types of policies covering kind of reliability, uh, security, as well as cost efficiency.
Uh, and when it comes to Kubernetes, we find that organizations really have to, you know, consider all three types of, uh, checks here in order to make sure that they are running workloads that align with best practices. This gives you a little bit of an example of the types of policies we've evaluated. Uh, I won't go through every single one, but, uh, you'll, we'll be covering some of these in today's presentation.
Um, the final report actually, uh, covers a number of different cate, uh, categories of policies and provides a lot more depth than what you'll, we'll be able to cover in today's session. So I do recommend that you download the report. I think what you'll see in today's presentation is sort of a, a summary and a bridge version, and hopefully kind of, uh, uh, see a little bit of where the industry's going in terms of, uh, both kind of Kubernetes best practices and sort of how well organizations are aligning to those best practices as well.
Okay. So a common graph that you're gonna see in this report is, uh, will look like this. And, uh, what, what I'll do is spend a little bit of time just explaining how to read the report and how to read the different charts and graphs that you're seeing.
Uh, so on the left, uh, on the Y axis here, you'll see sort of the percentage of workloads impacted within an organization. And on the X axis, you'll see the percentage of organizations, uh, that were evaluated, uh, in terms of how many of their workloads were actually impacted, um, you know, by the, uh, uh, as a percentage. And so, a great way to kind of read this report or this example here is, um, if you take this example, the number of organizations with less than 10% of workloads impacted has fallen, uh, from 46% in 2022, uh, to 21% in 2023.
And so when you see something like that where, uh, the number of organizations that have such a small percentage of workloads impacted decrease, it actually demonstrates that the problem might be getting harder to control. And so that's some of the ways that we've been able to kind of highlight information in this report. And, um, you'll see that in just kind of various different, uh, aspects of today's presentation.
And, and, um, when you read the, the final piece, so let's kick it off with sort of the highlight from this year's, uh, analysis, which is really this, this really kind of, uh, interesting fact that, um, over one third of organizations today and specifically 37% of organizations, uh, need to actually right-size their containers, uh, to improve efficiency. And when we dig into some of that, uh, those details, we'll notice that actually, uh, this 37% of organizations have 50% or more of their containers that are over-provisioned. And it's an interesting finding because, uh, what we notice is that a lot of times developers have to guess their resource requests and limits, uh, when they go to deploy it, because they don't really have the tooling or the feedback loops to tell them what, uh, those resource requests and limits should be.
And a lot of times, developers will guess high, they will over-provision by, uh, giving their application too much memory or too much compute. And when left unchecked or unmonitored organizations end up incurring lots of additional, uh, uh, compute spend as a result. And so, uh, this is the first year that we actually started looking at this data.
Um, and again, you'll see that the 37% of organizations that need to right-size, you know, 50% or more of their containers really represents, um, the bottom, uh, part of this gra uh, this chart. Um, but it's also interesting to see that there's actually a large cohort of organizations, in this case, 57% of organizations that have, uh, less than 10%, less than or equal to 10% of their workloads impacted. Uh, so some organizations do seem to get this right.
And, and what we're really excited to, uh, do is monitor this progress, um, over, over the years. So this being our first year measuring this, we wanna see sort of how well has this, um, uh, trend, uh, improved or not improved going into next year as well. So, um, again, first year we're kind of baselining this, but, uh, next year should help us understand where that trend is going.
Another aspect of kind of Kubernetes efficiency that's important to solve is really making sure that, uh, both memory and CPU requests are set on deployments before, uh, they're actually set to running Kubernetes. Now, Kubernetes technically makes these settings optional, uh, but when you don't set memory or CPU requests, it can actually make it difficult for Kubernetes to properly schedule that workload. And so, uh, what we're seeing is that, um, you know, this is becoming more and more of a systemic issue.
Uh, this year we're noticing that 78% of organizations have at least 10% of their workloads missing CPU requests, and this is up from about 50% last year. Um, so again, I think the, the numbers are kind of are all over the map here. You'll see that, you know, um, pretty much every organization has some amount of this problem.
Uh, but you know, interestingly enough, this is actually, uh, uh, fairly easy thing to solve with guardrails and, and policy enforcement mechanisms for Kubernetes, where you can kind of give developers feedback at the time of their pull request or at the time of their deployment when they are missing these settings. And, and, you know, use that as an opportunity to educate developers as well. So we, we think that while a lot of organizations may be having missing CPU requests, we also think this could be a potentially easily solved problem with, you know, policy enforcement and guardrail tools as well.
Let's shift a little bit to reliability, um, and kind of, uh, look at a few different trends that we've, uh, observed in this year's report. So just at a very high level, about a quarter of organizations today are relying on a cash version for 90% of their images. And, and what does that mean?
Well, it really means that the pull, uh, policy is not set to always, which is a general best practice. Uh, you, you kind of want to make sure that your containers are pulling, uh, the latest image so that you don't have inconsistency. And, uh, it also can help you from a security perspective as well to make sure that when you do push an image, um, with, with updates, that it's actually pulling in that latest image as well, not just the cash version.
So, uh, this is just a, a general best practice that, you know, we're seeing that, uh, right now, about 24% of organizations are relying on, uh, this poll policy not being set to always, uh, for 90% of their images. Another pattern that we're seeing is that container health checks seem to be missing or ignored in some, uh, deployments as well. And so, uh, right now about, uh, 66% of liveness and 69% of readiness probes, uh, are missing in, um, in Kubernetes deployments.
And it's important to set these because it helps Kubernetes automatically restart containers and ensure that the applications are available to receive traffic and then ultimately serve users. So this is actually considered one of the more basic, uh, ways to ensure application reliability in cobe. And we're still seeing that organization struggle to various degrees here.
Um, I think part of it is because the, the configuration does require a little bit of application specific input. Uh, so development teams need to, you know, consider, you know, what are the changes they need to make to their application in order to make sure that the health checks works for Kubernetes. Uh, so we're hoping that this trend kind of improves over the years as well.
Another trend that we identified is that deployments are missing replicas. Uh, this is another general best practice is to make sure that there's a few different, uh, you know, there's a couple replicas available for, for pods and that, uh, right now we're noticing that 30% of organizations actually have less than 10% of their deployments missing replicas. Uh, so this is a, uh, an, an improvement over 2023, uh, but still kind of, uh, you know, highlights, uh, that if you look at the graphic here, that some organizations have, you know, much more than just 10% impacted.
There might be, you know, um, lots of applications missing replicas. And, and I think sometimes this is because of the, what we call the copy and paste problem, sometimes a deployment from one team that's missing this best practice gets copied from another team who's looking to get their application deployed. And so they may be, you know, propagating these misconfigurations where replicas aren't set on the previous team, and now the new team is using that same configuration without replicas.
And so you can see that this becomes sort of a wider problem. Uh, again, this is usually a very quick fix, uh, uh, like a one line change to your infrastructure as code. And, um, we're hoping to see, you know, even though we're on the, on an improvement path here, that even more and more organizations, um, have fewer and fewer deployments missing these replicas.
Shifting gears a little bit, we'll also take a look at security. Now, security in Kubernetes kind of means can mean a lot of different things. Uh, we look at security, uh, from two lenses in this report.
One is from image vulnerabilities as well as, uh, from the kind of the configuration, uh, itself. So the YAML or the helm chart that's being deployed. Um, at a very high level, we're noticing that about 28% of organizations are running about 90% of their workloads with insecure capabilities.
So that means that they're adding some, some sort of insecure capability like, like net admin. And, uh, a lot of times it actually might be necessary for some applications or workloads to have these additional capabilities, but sometimes, uh, it may not be. And, and it could be, you know, accidentally added to apps going back to that original copy and paste problem where one team copies the configuration from another team as a, as a starting point and, you know, inadvertently propagate some of these misconfigurations going forward.
Uh, so we always look to make sure that applications start with not having these dangerous or insecure capabilities added, and, uh, that, you know, helps ensure kind of a good baseline from a security perspective. One positive though trend coming outta this year's report is that we're actually seeing fewer containers set to run as root. Um, so 30% of organizations today are running 70% or more of their containers as root, which is actually a drop from 44% in, in last year's, uh, report.
And part of me thinks that this is an ex another example of sort of a low hanging opportunity to, uh, fix, uh, make a quick win, a quick fix to containers by, you know, essentially turning off the ability to run as root, uh, which again, is a one line change. And, uh, I think we also see that this example, this type of misconfiguration example is, is sort of very popular, uh, when talking about the, the issues of, uh, misconfigurations of Kubernetes. A lot of organizations talk about, as an example, running as root being a, a common example of that.
So, um, it's great to see that this trend is going in the right direction in that fewer and fewer organizations, um, have, uh, a vast majority of their containers running as root, and that, that seems to be going in decline, which is awesome. Um, and I think it's important to note that, you know, running a container as root, just overall, it increases the risk of a malicious user taking advantage of that root privilege, uh, as part of a larger attack. So you want to kind of, from a defense in depth perspective, um, you know, by default have your container not run as route unless it absolutely needs to because of some special, uh, need or use case, uh, for that app.
So again, this is going in the right direction, and we hope, we hope it, uh, uh, stays that way going forward as well. Switching gears a little bit away from kind of misconfigurations, we'll, we'll talk about image vulnerabilities. And so this is, you know, the, uh, image vulnerabilities that, uh, may exist in running containers or as part of, you know, scanning, uh, container images as part of your, you know, CICD process or your shift left process.
Um, and I think this is an ongoing challenge for many organizations. It's, it's an ongoing problem, but we do see some signs of progress in this year's report. Um, so if we actually dig into the, the first section where we show the percentage of workloads impacted, uh, 26% of organizations have less than 10% of their workloads affected, which is an improvement from, you know, 12% in 2023.
So we're seeing essentially a greater percentage of organizations with fewer workloads impacted due to image vulnerabilities. And I think that, um, is a signal of both kind of organizations upgrading their third party containers to newer, less vulnerable versions, but also integrating and scanning more of their containers so that they have a, a process in place for this. Um, in the report, you're also gonna see a section where we talked about uns scanned images.
So, uh, Fairwinds, uh, is able to kind of help companies identify if there's images running their cluster that they have not scanned. And, um, this has greatly improved, uh, over the year. We're actually seeing almost 84% of organizations getting almost complete scan coverage of, of containers in their runtime that's up from 64% last year.
Um, so I, I think that's a great sign that organizations are, are kind of doing the first step, which is scanning as many of their images as possible so that they understand their risk, and then, and then taking remediation after, after that. Uh, so, you know, we hope that next year we even see a higher percentage of, of organizations with fewer workloads affected. One of the enhancements that we made to Fair Wind's Insights last year was we added, uh, some specific checks related to, uh, uh, the NSA hardening guide.
So, uh, the NSA actually released, uh, Kubernetes hardening guidance, I think back in 2021. Um, and, uh, there was a number of great recommendations there. And we actually expanded the number of checks that Fairwinds Insights offers to, um, match, you know, what the recommendations were in the NSA hardening guide.
Um, so a lot of new security checks kind of made its way into the fairwinds Insights platform this year. Uh, one of those checks is actually verifying if there's a network policy, uh, configured for, for workloads and network policies are, are increasingly important because it helps you kind of segment workload traffic and, and, you know, uh, ensure that you've got controls around which pods can speak to which pods. And so we wanted to get a sense of how is the industry, you know, doing on this particular, uh, policy.
Um, and so I think we see kind of, you know, two, two types of organizations. Uh, 37% or about a third of organizations today have less than 10% of their workloads without a network policy. And that's actually a great sign that, um, there's a lot of network policy adoption happening in some organizations where they're making sure that their workloads have a network policy set.
Um, but on the other hand, there's still a majority of organizations that have, you know, way more than 50% of their workloads without a network policy. So it means that they're deploying the Kubernetes, their, their workload's running fine, but you know, that workload can, can speak to any other workload in the cluster. And so I think it shows that the industry still has a little bit of ways to go to make sure that network policy adoption is even more widespread and more adopted.
Um, and so just to kind of give a little bit of an example of why we think this is important, um, you know, network policies help you limit that egress and ingress traffic. And so, um, when you have that ability to control the traffic, it allows you to kinda, again, from a defense in depth perspective, you know, prevent any undesirable access to, um, uh, to those pods. So those are some of the, the summaries and the highlights from the report.
Again, I think it's probably only, we're only covering about a quarter of the information that the report has this year. Um, but I wanted to kind of also help organizations understand what is a path forward? Like if you're running lots of Kubernetes today, how do you ensure that your teams are following reliable security and cost efficient best practices?
And I think that's really where Fairwinds Insights can, you know, provide a lot of value. It can provide, you know, guardrails to help you solve your business problems. Uh, whether it's, uh, ensuring that your images are free of vulnerabilities or that your, uh, your, your workloads are aligned to standards like the NSA hardening guide or aligned to standards like SOC two or ISO 27,001.
There's a big security reason, uh, to provide developers with, with guardrails and feedback around their, their configuration hygiene. I think increasingly in 2023, we, we did notice that a lot ofop, uh, organizations were very cost conscious. So they wanted to make sure that they had a, a way to measure their container usage, but also right size containers to properly make sure that it's using the correct memory and C uh, CPU, and they're not overspending in ways that, you know, incurs additional cost or just wastes compute resources.
So, uh, Fairwinds does provide sort of both Kubernetes cost allocation as well as container rightsizing recommendations, uh, and that that's helped organizations in some cases save over 25% on their container costs. And then finally, this notion of guardrails is sort of core to everything that we do. So, you know, in order to make sure that engineers have the tools to take action on this feedback, you wanna be able to provide guardrails at different steps, steps in the process, whether it's at time of pull request, when they're making their infrastructure as code changes or at the time of deployment, uh, also known as the time of admission when applications are being deployed into the Kubernetes environment.
You wanna give that feedback to developers and have, you know, both like a way for them to remediate things easily, but also ensure consistency so that you're not introducing risk or, or overprovision to applications along the way. And these are kind of the core capabilities that, that Fairwinds Insights provides and how our customers are getting value. So I, I do encourage you to kind of take a look at, um, the Kubernetes configuration benchmark report for this year.
Uh, like I said, we only really covered about a quarter of, of what's in that report, and there's a lot more, uh, broken out by security cost and reliability, so you can kind of see the different patterns. com. Uh, reach out to me on LinkedIn.
I'm happy to, uh, point you in the right direction. And, um, I think that's really kind of what we're hoping to, uh, cover today. And so, uh, thanks again for the time and looking forward to hearing your thoughts out there in the community.





