Observability/DevOps 2.0 at SKILup Days 2024
With the evolution of observability, DevOps and beyond, driven by data, AI and computation, we will see new observability pillars driving DevOps practices. Predictive visibility will allow deep and broad forecasts of cloud and business systems. Workload data and telemetry allow for model creation, validation and tuning. Autonomous remediation for rapid, targeted response to cloud and business issues will be automatic, even when teams are not available.
Alignment of cloud and business metrics spanning across the production floor and virtual cloud will merge to expose valuable insights. These new pillars will improve observability ROI by forcing a shift from consumption-based pricing models allocating funding for DevOps 2.0.
Key learnings:
– Why the new ways to observe, identify and repair infrastructure are different from the old technology will change DevOps practices.
– Why accessibility, low cost of quality, organized and large quantities of data is essential for DevOps 2.0.
– Why these change will allow the redistribution of resources to more higher value activities.
Transcript
Hello, uh, my name's Mark Callahan, and I am the CEO and founder of Cloud Canaries. 0 or two point x. It kind of depends, but also what happens beyond that.
And I think that's pretty important to bring up right now as we go through this, uh, evolution. Uh, also we're gonna touch upon what is driving obs observability as well. So stay tuned.
We're gonna go through, uh, some slides that I hope you find, um, very interesting and helpful and, and a little bit entertaining. So here we go. 0, predictive visibility, and that includes a, a list of, of, of items.
It's continuous monitoring, SLA, uh, compliance, actual and forecast data. So you can, you can work with, you can alarm notify on actual data or forecast data from API to workflow. So observability, uh, it has to be more than just monitoring APIs.
It really has to be monitoring the, the user experience, and that really includes workflows. Um, the second pillar is, uh, autonomous mitigation. Now, what that really means is, you know, prevent fix, but do no harm without interventions of the op ops team.
It also kind of means that you wanna be able to do this really fast if something bad is happening. And so that's within 15 seconds, but it could be weeks or even months that, um, the system will basically self-heal. So you have to keep that in mind.
Um, when you have this type of mitigation, you want to be predictable in the outcome. When it, it, it has to, the way it has mitigated an issue, uh, or resolved it, uh, has to be kind of within the, I think, a set of, uh, domains that are, uh, are reasonable. Um, also kind of aligning actions with intent.
And this kind of goes back with predictable outcomes. Um, you can fix the problem with a hammer or, um, a sledgehammer and, and you wanna be able to, uh, do it. That kind of aligns with the intent.
0, we believe the cloud canaries is convergence of metrics, connecting the dots. The dots are from the cloud to your business. And once you've done that, you've opened the door for a set of very interesting opportunities.
Predict anything, forecast anything, add business insights. So connecting the dots between the business metrics, cloud metrics, convergence of those two, this is, again, a really important thing. All these three pillars are, are, um, very important and will change the way DevOps functions and operates, and the requirements that are gonna be built upon, uh, um, that for, for the DevOps teams, and I think all really, really good.
0, we have our three pillars. 3, but the idea is this will be a set of increments that, um, that will, um, happen in observability. Um, I just didn't wanna confuse everyone.
0, as we mentioned, three pillars. Uh, predictive visibility, autonomous mitigation, uh, convergence of the metrics, cloud metrics, business metrics. This will have an incredible impact on the business, very positive.
And the cloud too and beyond. At Cloud Canaries, we believe that AI models, and these are models built by machine, um, uh, learning tools that are being supplied by many different vendors that create, uh, models. Basically, for the most part, neural networks will become the solution.
In other words, you're gonna have one platform and many different models is the models are gonna contain all the intelligence that you need to, to, um, solve that a particular problem. And because, and because it's not a, you don't have as many platforms and, and models provide the solution, you're gonna have many solutions because they're gonna be much easier to, um, to deploy, to manage, and you're gonna be able to roll it out to your entire, uh, all of your infrastructure, which again, is going to be amazing. It's gonna be incredibly helpful for the cloud and the business.
So once again, beyond, it's gonna be, you're gonna be buying or renting or subscribing to models, not platforms. Think about that. It is gonna happen.
It's happening right now in some areas. 0 and beyond? So let's take a look at that.
Um, developer leaning and people say, moving to the left, I, I don't like to use that term, but it's gonna be, um, developer leaning. DevOps is gonna be, um, leaning to developers to do a couple things. Building observability in code.
There's some things that you just have to do. You have to do it in the code, and you can't just add on a little library to do it. And, and the, the, the, the clearest example is workloads, um, and workflows.
So, um, how does a, how does a workflow, um, work? It, it's a, you know, it's a, a process of going through, uh, several different stages within, uh, kind of a user experience. And you create this, this workflow.
And, um, you might, it might be a series of workloads, don't really care, but, uh, observability is gonna have to be built in to really capture that, um, that information. Also leveraging CI and cd, continuous integration and continuous delivery with ai. So again, I think that's gonna work with observability to actually make a better, more, um, uh, better applications, but also better models.
And we'll talk about that a little bit later on. Second pillar, cloud gatekeepers. It's not gonna be just an ops, it's going to expand, and DevOps are going to be the gatekeepers.
If there's something that's going to happen in the cloud, you're gonna be talking to DevOps, the DevOps organization. I call 'em the guardians of the cloud. And that's really what it is.
What, you know, DevOps will be focused on predictive outcomes. Your reaction time is going to, you're not gonna have that much re uh, the re reaction time that you're gonna be investing is gonna be fairly small because it's working. Everything's working.
So predictive outcomes and DevOps is gonna own p and l, you know, profits and losses. It's gonna own the performance. And SLA targets, again, it's more of a gatekeeper, guardians of the cloud kind of role versus an operations or developer security.
It's all kind of all the above within an organization. The third PO pillar is, is cloud analyst. And, and maybe there's a better term for it, but it's really kind of making sure that, um, the cloud is aligning, uh, to the business intent and that the RS, uh, is, uh, is understood.
So aligning how the cloud is being operated, what you're doing to the end in mind, which is making customers happy. Uh, measuring success, predicting success. It's important also on the downside is, you know, measuring, uh, when things go wrong are when there's a failure and learning from that.
Um, and those failures might have, might have nothing to do with, uh, the cloud. It's really driven by the business. 0 and what's gonna come beyond that.
And I think what's coming beyond that is gonna be very, very exciting too. So let's take a quick look at that. 0, developer leaning, it's in the code ban, uh, cloud gatekeeper, um, guardians of the cloud.
Uh, the scope of, uh, DevOps grows roles, grow cloud analysts, making sure that the business and and activities in the cloud are aligned and using the, the massive data through, um, understanding the flows of, of, uh, solutions or workloads of solutions, uh, uh, or creating models that can help the business, um, you know, take care of customers and have a fair, you know, profit. So beyond is, um, kind of extension of that more diverse specialized roles for DevOps. And we're seeing this to some extent today.
We see this in, uh, security. We see this in infrastructure, uh, solutions. This is gonna grow.
Uh, the key will be to always make sure that the DevOps continues as a community, because that's really what it has been. Um, and, and that's gonna be the key to its long-term success. 2.
You can, you can decide when we, when we accomplish that. Let's take a look at, uh, the next slide. So I tried to map this into like, what's the average time that a DevOps individual will be spending working, and what will they be doing?
Um, and, and I hope you, the, the big takeaway, I hope you see here, and, and it's, you know, pretty obvious is rea the time that you have to spend on reaction, there's a disaster. There's two disasters, there's three disasters will be reduced. It has to get reduced.
That's really the only way to allocate the time to do the other tasks that will have immense value for the business and the customer. 0, uh, the time that you have to spend on reacting to some emergency is gonna go down. And the reason for that are observability.
Being building the code, um, we're able to mitigate, um, the problems before they even happen through forecasting. We're able to gen through our forecast due prevention, understand what the problems are, and, and deal with them in a very reasonable time. And that will impact the amount of time that DevOps development, operations security will actually have to be reacting to something versus planning to, you know, solve problems that are, that are six months or two years, uh, away.
Um, beyond that, i, I, I think it will, um, it will continue. There might be a little bit less development aspect because observability has been built into the code, and that's now a standard process. There should be more economic to actually do.
The big takeaway in the beyond is the addition of business insight. As DevOps teams continue to build their mod forecast models, um, you're gonna be able to generate business insight that, you know, the business side will, well, we'll be very excited. It'll be, it'll be, it'll be their secret sauce for the business, the how to take care of the take, how to take care of their customers.
So you're gonna see a, a massive amount of that because DevOps are gonna be managing the models through AIOps, or however you wanna define it to me, is still part of DevOps. Um, so that's the big takeaway with beyond the ability to really communicate value and insight to the business. 0, but on a smaller scale.
So then the question is, well, what's driving OB observability? 0, right? Yeah.
It boils down to limited ROI talk to anyone, well, at least, and at least I have in the DevOps world, and you, you get the end, the end in mind that businesses aren't getting their, the ROI that they expect it. Some of it has to do with solutions being overpriced, consumption based. So you can't, you know, do a full rollout as I mentioned here.
Um, these, the, the company can't afford it. Cloud costs, they really dropped when you moved from the, the, your data center to, to the cloud, but now they're going back up cost to support overlapping vendors. You don't necessarily need that.
If, if the, uh, if the model that you're generating is the solution, you don't need as many platforms. The number of vendors and, and consult consultants will go down at, you know, um, because you're just not gonna need them. Uh, regulatory rules that's continued, that will continue, uh, to, to increase.
Um, especially if you're, uh, a global company, whether you're in EU or you know, the US or, you know, wherever you, you, you're, you're positioned, um, the rules are gonna increase. We mentioned, um, organizations can afford to do a full rollout. And again, I, you know, it's, it's just amazing when, when I go and speak with, uh, DevOps, uh, you know, VPs of DevOps, uh, operations and, and they say, well, we only actually rolled our observability solution out to like 20% of our infrastructure because we really can't afford it.
Um, and the turbo insights not delivered fast enough. Ultimately, that is the, the big driver for, uh, a limited ROI. 0, AMV odd.
So let's go to the next, next slide here. So what are the drivers? Um, um, so there's a, there's software drivers.
Biggest one is Open Telemetry tool is tool slash vendor agnostic. High adoption, it's amazing how many of software vendors have adopted open, uh, telemetry. It's all standard standardized.
It might not be necessarily be perfect, but vendors are all supporting it. And again, if there is an issue, it can be resolved pretty, pretty quickly because it's standardized. Second pillar code first.
This is having a dramatic impact on the quality of software that's being generated. Um, you know, both as a feature, as features, but also as within a steady state, uh, where software is released and, and running. Um, you'll be able to provide better monitoring, observability, um, security, uh, in, in, you know, all the time, uh, with, uh, continuous, um, observability.
So building code build in code, not a add in. And it's standardized code first. The last one, the last pillar is, uh, can you continuous integration and continuous, um, deployment in code.
That's pretty, pretty typical right now. CICD, um, but also in the forecast models that you're creating. Um, there's no reason why we shouldn't be doing that.
We should always be making our models, um, smarter. Um, and we can do that now. We'll talk about that in a few, few, few moments.
Better code and models, better user experience. Ultimately that is our goal. Observability technology drivers model ready data at Cloud Canaries, since we collect the data, we know how to easily, uh, use that data in our models that we, we create no matter who, whose tools the customer is actually using.
Pretty cool, um, machine learning AI tools, again, there's a lot of different flavors or shapes, I guess shapes. Um, all of them have certain pros and cons. It's should be up to the, the user to actually determine what they want to use.
And the last piece here is compute, because our theory of cloud canaries is that you don't necessarily need, uh, you know, a thousand data scientists. What you need is a lot of compute and generalized, um, uh, uh, tools that, that generate, generate models. You'll get your best performance bang for the buck that way that's happening right now.
Now, with that, we can, we can take, um, we can allow continuous model, uh, validation. And this is the key to, to, uh, one part of the puzzle because we're constantly collecting data, workload data telemetry. We're using that data to build a model.
We're using the models to generate forecasts. Eventually in time, we generate actual data that can be compared to forecast data and we can say, Hey, it's right on, uh, uh, uh, or No, it's not. And then tune that model, uh, and, and take the error out or as much as you can until you get, you know, more actual data or you get more forecasts.
So it's a continuous loop, and this is gonna have a massive impact on the whole idea of models. Become the solution, not the platform. Think about that one platform.
Many solutions 'cause of many models, what happens? Solutions get smarter, better, faster. It's pretty simple.
You know, collect, validate, build your model with, uh, models, with, uh, tools that you want to use, forecast, predict, add, insight, notify and act. And act is obviously the, a critical piece here for our customers. Beyond solutions, as I, as I said, you're gonna have more models, fewer platforms.
You know, our claim at Cloud Canaries is that the, the, the number of solutions are going to become very, uh, they're very numerous because you can be able to build models and be able to use the same infrastructure, same, uh, platforms to use different models for different problems. And from digital experience, contract negotiation, you name it, we can have a model for that. Costs are gonna come down, support for those.
That platform is gonna cut, is, is going to be low. And you're gonna build a set of consultants who are focused on a couple platforms and, um, validating models. So it's, it's gonna be, it's, I think it's, it's uh, it's a paradigm shift and, uh, it's going to, uh, make, uh, better products, happier customers and, and, and more time for us to, uh, think of new, new ways to innovate.
That concludes our little presentation. I wanna thank you very, very much for listening. Uh, if you wanna learn more, you know, just go to Cloud Canaries.
Thank you very much. Have a great day.