A New Paradigm for Platform Engineering with Hamish Watson at Techstrong Con 2024
The integration of Artificial Intelligence (AI) into DevOps practices presents a groundbreaking opportunity to revolutionise platform engineering, fundamentally enhancing operational efficiency, reliability and innovation in computing environments.
This session will share the role of AI in transforming the role of engineers by automating complex workflows, optimizing resource management, and ensuring high levels of system performance and security. The session will start by looking at the challenges faced by engineers and developers in the dynamic and complex ecosystem of platform computing (on-premises and in the cloud), including the management of scalable resources, continuous integration and deployment (CI/CD) pipelines, and real-time monitoring and incident response.
Transcript
Hi everyone, and welcome to my session, A new paradigm for platform engineering, empowering DevOps processes with ai. My name is Hamish Watson. Uh, I'm gonna be presenting to you today.
Um, I am a DevOps consultant here in Christchurch, New Zealand. Um, I've been doing DevOps for basically since it's been around, and it's been really exciting over the, over the years to see it mature within, uh, companies. But also we are now on a, um, a new paradigm where we are now having AI come along.
And, um, I feel like it's a step change, uh, within our industry. Um, doing sessions like this, bringing DevOps to the masses is a personal passion of mine, um, and helping people understand AI data. And the cloud is, uh, a company driver.
I have my own company, Morphe, uh, based out of, uh, New Zealand working with, uh, clients all around the world. And basically at, at the core of what I do is I am a, a technologist who understands business value, um, allowing companies to deliver value quicker, um, but also helping, um, you know, engineers know how to incorporate some of the latest, um, DevOps processes, um, using tooling, um, to also be a part of that delivery of business value. Um, and if I could sum up who I am in one word, it's, uh, make stuff go.
Uh, if you follow me on Twitter or X as it is these days, um, often I will be, um, uh, using the hashtag make stuff go anytime that I'm talking around. And, uh, DevOps. Alrighty.
So in this session, we're gonna do a very, uh, brief introduction of DevOps, an introduction of where AI can help, and also some examples of tooling that, um, I've been using out in industry. And then we'll wrap it up and, and see if there are any questions. When I thought about, um, what, uh, DevOps is, and also around cloud or platform engineering, I really liked, uh, this quote by, uh, Donovan Brown, uh, used to be, uh, principal DevOps project product manager at Microsoft, where DevOps is the union of people, process and products to enable the continuous delivery of value to our end users.
And when I was preparing for this, um, session, I've been in it over 25 years, and I was thinking about, you know, how did we used to engineer our platforms? And in the old days, you know, we had all these different work requests that were manually transferred from person to person. Um, you know, the developer would have to raise a request with the server team who would then deal with the networking team.
There might be a backup team. And so we had all these different silos within our, uh, processes, um, to actually get our platform stood up, uh, our platform monitored, and also our platform secured. Whereas, you know, these days I'm looking at codifying, um, my infrastructure, um, you know, uh, be it using, uh, Terraform, using some form of scripting or, uh, in the bottom right hand corner there, you know, using generative ai when I'm writing Ansible, because I don't know about you, but often, you know, I'll use a technology or, or a language.
Um, but, you know, I may not use it for, you know, a a wee while, might be a a month, couple of months, whatever. And it's quite good having, you know, some hints, um, from generative AI tools there, helping me along, um, and being able to, um, you know, just have that hint, uh, reminder and go from there. There are pain points associated with building platforms.
I've been building platforms, uh, since the late nineties. And, you know, historically, you know, managing our infrastructure or resources was a manual process, you know, would physically put servers in place, configure them, you know, only when they were configured to the right settings, et cetera, et cetera. And, you know, having that manual workload, um, is not one that can scale.
You know, uh, if I'm building, uh, physical servers or even virtualized servers manually, you know, you can't scale me. Whereas if I can codify it, you know, suddenly, you know, we can build a thousand servers all standard from, you know, a template and, you know, push it out. And of course, there's costs.
You know, we, back in the day, we had, you know, teams of infrastructure engineers, and all they did was literally just, you know, uh, put in massive amounts of hardware. And, you know, having our process siloed through the way meant that we had, you know, professionals at each step who would do one particular thing, they did it well. But, you know, having network engineers, hardware maintenance technicians, that kind of thing, um, you know, we've gotta pay all these people.
I talked about scalability before. And basically, you know, scalability and availability, um, are related to the speed in which we can actually push out our platform, our resources, whatever. And, um, you know, if we didn't, um, have, you know, enough, um, resources in terms of, you know, massive data centers, you know, this is one reason why we've embraced the cloud, because, you know, our SLAs associated with when, how things are available, well, now we can actually purchase, you know, that, uh, tier of availability and of course, monitoring, um, you know, we've got all the infrastructure in place.
How do we actually keep an eye on it to ensure it's performing optimally? More importantly, how do we do that in a standard, um, way? You know, often we'll have a production server, it's set up with diagnostic settings, um, for, you know, sending data off for observability, but do we have the same settings set up across every single server, you know?
And so, um, we need, um, the ability to have a consistent approach to how we, um, deploy our infrastructure or engineer our platforms. This here is an animation that I love doing because this is, uh, CICD in action, right? So we have continuous integration where we're making changes.
Um, we're writing code, we're hopefully doing unit tests or some form of scanning at the source. So when I write within my IDE of choice or integrated developed environment, I then, you know, I've got a range of tests that are happening. I then push up to build server, um, oh, well, sorry.
I push up to source control, which builds on our build server, and then we're slowly going through our various environments. And so we're pushing out two environments that have been built from source control. And so all through this process, you know, we've got AI at different points.
AI may be checking for vulnerabilities of my code as I'm writing it. It may be hint giving me hints for the code as I'm writing it, um, through our build test. We may be using, uh, AI in the background to scan for certain things, um, through our build process.
And as we go through our testing, we may be using generative AI to create some of our test data, um, and also, uh, some of our, um, our testing frameworks, you know, um, and then if we're happy that everything has gone well up until this point, we'll then push through to pre-production and then production all the way through here. We do have continuous monitoring, and that monitoring in itself will have an AI component, because think of the, you, the gigabytes and terabytes and possibly petabytes of data that is being generated through, um, our monitoring and observability tools. You know, getting, um, machine learning over that and being able to get information from this massive amount of data, um, is an absolute, um, help in terms of knowing what is actually going on in our, our platform.
Right? So here's an introduction. So the merging of AI and in platform engineering and DevOps basically represents an integration where AI's predictive analytics and automation capabilities actually complement.
You know, what we are doing in the cloud. And as platform engineers, we will be building, um, resources on-prem. And one of the reasons why I like certain languages, say like Terraform, um, is the fact that I can use the same paradigm when I am deploying infrastructure, say, on premises as well as what I'm doing in the cloud.
And so we get operational efficiency anytime where we're reducing our manual, uh, workloads, either via developing, deploying, or managing, um, our assets, uh, on our platform. And also automated problem solving. Um, you know, AI helps us, um, uh, to detect problems, um, and also, um, it can help in form the resolution.
And again, you know, having machine learning, um, algorithms going over, um, our resources based in the cloud, um, looking at what is happening at the moment, being able to detect the problems, and, you know, and that detection of problems is where AI tools can identify patterns, um, predict potential issues before they become critical, um, and suggest or even implement, uh, solutions without our intervention. And, you know, that sort of comes down to optimization. Um, it not only speeds up our DevOps cycle to again, deliver business value, um, but it also allows us to have more robust and fault tolerant systems.
And of course, you know, innovation is key to what we do, um, within our industry. Um, and being able to continuously improve means that the convergence of artificial intelligence, platform engineering and DevOps processes, it really gives us an environment of continuous improvement and innovation. Um, you know, we can rapidly prototype, we can rapidly test and deploy innovative features, where before we couldn't because we, we couldn't do things at scale.
We were doing things in consistently, um, but also we were weighed down by the sheer weight of the data associated with running our platforms. And so now we can actually get actionable insights that drive the continuous evolution of our services basic, and which means that we can not only meet, um, our business and user expectations, but more importantly exceed them. All righty, so cloud infrastructure.
So I talked a little bit about scaling. Um, and so we now move towards where we have intelligent resource allocation and scaling. And basically we can dynamically analyze application performance, user demand in real time.
And that means that we can automatically adjust, you know, our resource allocations, uh, to meet these needs efficiently. Um, we can streamline the deployment process by automating the configuration of our environments. We can select, you know, the optimal configuration settings based on our application requirements, um, but more importantly, look at, you know, historical performance data, and which reduces again, the effort for setup and minimizes, um, uh, human error, um, through, again, through analysis of logs and monitoring data, we can predict potential, um, system failures or performance bottlenecks before they even occur.
Um, and, you know, we can now move towards self-healing systems. I've worked with a lot of clients, you know, where we set up some of the, the, the health checks, uh, in the background, and then we'll set up, um, you know, algorithms in terms of, uh, if we get, you know, uh, a, if we breach a threshold of say, a 500, uh, HTT P status codes. So if we get a huge amount of errors, you know, we can use ai use, uh, cloud platform algorithms in the background to actually self-heal, to actually remediate, uh, those issues.
Alrighty, security, I, I wouldn't be able to talk about DevOps unless I actually talked about, uh, security. Uh, and again, uh, AI pay plays a very, very important part here. And so we are doing predictive threat detection.
So we're analyzing, uh, not only historical data, but we're actually looking at real data in real time from a whole heap of sources within our environment to identify patterns and behaviors, which may actually indicate that we have potential security threats. And, you know, again, by leveraging machine learning algorithms, AI can predict attacks before they occur, which again, allows us to do preemptive actions to basically safeguard our data, our applications, and at the end of the day, our users. And so this proactive, uh, approach to security significantly reduces the risk of data breaches, and it ensures that we are having continuous, um, protection against, you know, cyber threats.
Because cyber threats are evolving, our platform needs to evolve to mitigate against it when we detect a potential threat. Um, you know, if we have our security systems, which are being powered by AI in the background, um, you know, we can initiate response protocols to mitigate risks, uh, as they, before they happen or as they happen. And so we may isolate our effective systems, we may deploy patches, um, or adjust, you know, firewall rules to block, you know, malicious traffic as it's happening.
And by automating, um, these responses, artificial intelligence ensures that our security incidences are addressed immediately, often before they can cause significant damage. Um, and again, that gives us, uh, integrity and availability of our platforms. Um, you know, we need to be compliant.
Um, and so we need to embrace, um, a paradigm of continuous security. And so having AI systems continuously monitor cloud, uh, cloud or platform environments, um, for compliance, you know, think of when I think of compliance, um, you know, I think of a lot of, you know, paperwork, uh, a lot of regulations, well, we can use, you know, um, uh, any form of GPT, right? Um, so, uh, generative, pre-trained transformer, um, can be created to go over all the compliance materials, look at our platform, and we can create that ourselves.
We can engineer our own GPT that will be, uh, doing a whole heap of work in the background for us in terms of compliance. And, you know, we can, by using this approach, um, you know, we can ensure encryption is used when necessary access controls are, you know, properly implemented, um, and configurations are secure. And so, again, you know, this is a good strong, secure posture where embracing means that, you know, our deployments of platforms, you know, those platforms, be it web apps, databases, uh, microservices, you know, all through our platform, you know, all of those different things, right?
Can remain compliant with industry standards and regulatory requirements. And that reduces, you know, legal and operational risks. And of course, v you know, um, identifying vulnerabilities within our platform.
So again, we can harness the power of AI to, um, automatically scan our code dependencies and, you know, our infrastructure or platform for known vulnerabilities using machine learning algorithms to identify patterns indicative of potential security issues. Um, and then, you know, once we've detected it, you know, we can use AI and again, machine learning to pri prioritize our vulnerabilities based on the severity, uh, the exploitability, um, and the potential impact on our system. So again, this enables our teams to focus on fixing the most critical issues first.
And this not only reduces the window of exposure, but also optimizes, you know, the allocation of our security resources, ensuring that again, you know, our efforts are concentrated on where they should be and where they have the greatest impact on improving, uh, our security posture. Um, continuous integration and continuous delivery. And so AI is a game changer here because, you know, it enhances our pipelines by automating our routine tasks, um, optimizing our processes based on, again, you know, real time feedback within, um, our builds our tests and releases.
Um, and, you know, by utilizing machine learning algorithms, AI can actually predict the best time for our deployments, uh, identify bottlenecks in our deployment process. And, you know, a key theme here is to suggest improvements to it. And so this leads to, you know, faster and safer, um, releases of our software.
And as I said before, you know, business value that we are delivering without sacrificing quality or reliability. And again, what that means is that our engineering teams, well, any of our teams development operational teams, DevOps teams, platform engineering teams, it means that we can move up a level and actually focus on more complex and creative tasks. You know, when I think about when I first started building pipelines, you know, building them in Jenkins and using octopus deploy back in the day, you know, a there was a lot of work to bid in the pipelines, um, but there was a whole heap more work to maintain that level of automation that we wanted.
Well, these days, you know, a lot of that we can automate a lot of that. We can, you know, incorporate machine learning, generative AI to actually do a lot of that. And so we can then, you know, start doing more creative, um, pipelines.
Um, and yeah, before changes emerged and deployed, hopefully right on, uh, my laptop or um, computer, you know, AI can analyze the code and infrastructure changes to predict their impact on the system. And there's some great extensions out there that you can put into, um, whatever IDE you are using for coding, so that as I'm writing the code, not only is it suggesting some of that, um, good code, um, but it's also looking at potential failures or, um, vulnerabilities that I might be introducing in my code scanning for secrets that, all that kind of thing. They shouldn't be there.
We should be using a secret manager, um, you know, basically by preventing, um, you know, those vulnerabilities getting into, uh, our code or, you know, code that's less than optimal, um, you know, reduces our need for rolling back once we've deployed out to production. And also ergo it will prevent, uh, disruptions. And again, you know, a key theme here for platform engineering is the fact that, you know, we can dynamically allocate resources based on the needs of, you know, our development or deployment processes and, you know, analyzing the workload, um, and using performance data.
You know, AI can scale resources up and down in real time, ensuring that our builds and tests run efficiently without unnecessary expenditure. And what I mean by that is often I'll work with qa, uh, quality assurance engineering teams, and so we'll use pipelines to bring down, you know, uh, an anonymized version of production. But then, you know, we can dynamically, um, allocate our resources, um, uh, you know, scale it up and down as required, um, but do it in a cost efficient manner.
You know, uh, I want to be able to throw a workload at my platform, see it scale, but then once I've finished, you know, bring it all down, remove it, whatever, because I don't wanna pay for resources if I'm not using them. And that's true of whether I'm engineering a platform in the cloud or maybe using some hypervisor technology on premises, you know, everything costs if it's running. And so by incorporating these, we're not only speeding up the development and deployment cycles, but we're also optimizing our costs, making everything more efficient and, uh, cost effective.
So, um, advanced anomaly, excuse me, advanced anomaly detection and alerting. So again, you know, when analyzing metrics, we're analyzing logs, and these logs are from disparate sources. Some of it might be structured data, some of it might be unstructured data from various sources.
And so, you know, incorporating AI into this part of our platform allows for immediate alerting on potential, um, issues before they escalate, enabling quicker response times and minimizing the impact on our, um, on service performance. Um, you know, we can start upon detecting, you know, an issue. Um, you know, there are AI enhanced tools out there.
I've got two examples there. I'm not, um, saying that you have to use these, these are two examples of AI tools that I've used, um, during my time. You know, they can automatically perform root cause analysis, determine the source of our problem, uh, whether it's in the application code infrastructure or some specific service dependency.
And, you know, detecting and responding to incidents, uh, is great, but also being able to contribute to, you know, continuous optimization, uh, of cloud services and our infrastructure, be it on-prem. Um, you know, by analyzing patterns over time, you know, um, monitoring software can provide insights into resource utilization, you know, at performance, um, and user, uh, experience. So what's some best practices?
Start small. Um, incorporate ai, AI early in your development, um, lifecycle, you know, integrating AI capabilities early in the development process, you know, during the planning and design stages, ensures that, you know, the insights that AI can give us, the automation and being able to embed this into our, um, application infrastructure, basically, you know, it enables us to be able to do predictive analytics, intelligent automation, and real time monitoring to be, you know, these are core aspects of, of our system rather than, you know, afterthoughts, oh, we should do this. So my advice is to start small, start early, um, adopt, you know, a a continuous learning and improve mindset.
You know, this is, there are three ways of DevOps. Make your work visible, get feedback loops as early in your process as possible. And the third way is around risk.
Being able to experiment in a way that doesn't, you know, being able to take risks, um, because we have all, all of our platform is built in such a way that we know that we can spin up an environment it's production like, and go from there. Um, you know, when incorporating AI into our engineering and DevOps, uh, process, it's crucial to prioritize ethical considerations, data, prior privacy. And this includes transparently, you know, using AI algorithms, securing sense of data, and ensuring that AI driven decisions don't inadvertently introduce bias or, um, or discrimination.
So where are we going? So I've talked about, you know, predictive analytics and decision making or autonomous operations and self-healing systems and AI driven development and tooling. I think that we are on a great path at the moment, and, you know, this is a great time to be part of incorporating AI into your DevOps processes, um, for cloud engineers, right?
So I've just seen, uh, that we have one question, uh, in the chat, so I'll just go to that. Um, alrighty. So, um, the question is, is AI a potential threat to the job security of platform engineers?
No, it's not. Um, you know, it's, it's understandable to think that AI could be a threat to the job security of us, whereas in reality, AI acts more of as an enhancer rather than a replacement. And so I see AI as being complimentary to what we do.
Um, it's going to enhance our problem solving capabilities, but we still need human oversight, right? We still need, um, you know, they need to be trained on specific data sets and continuously tuned. Um, so yeah, I think that AI is, you know, it's our peer.
It's there to help us innovate, strategize quicker and faster than we ever have before. And again, being able to move up and use complex, um, you know, do complex work and everything. But, um, and of course, you know, um, being able to help us with regulatory, thank you so much for being part of this session.
I really do, uh, appreciate it and have a good rest of your day.

