Adam Robertson, Engage AI | DevOps World 2023
Adam Robertson, head of DevOps at Engage AI, discusses multi-cloud Kubernetes to accelerate DevOps practices. Multi-cloud Kubernetes offers a number of benefits for DevOps teams, including increased agility, scalability and resilience. In this interview, Mitch and Adam discuss how to use multi-cloud Kubernetes to accelerate DevOps practices, such as continuous integration and continuous delivery (CI/CD), as well as some of the challenges and best practices for managing multi-cloud Kubernetes environments.
Transcript
This is Techron tv. Hey everybody, welcome back. We are a DevOps world here in Santa Clara, Silicon Valley, 2023, talking with some of the great people.
There are lots of talks that are going on at the same time, uh, as we're doing these interviews conversation. So yeah, I kinda have the best of both worlds being part of this if you're online or here in person. So I'm joined by Adam Robertson.
Adam, welcome. Adam is, uh, head of DevOps with Engage ai. Mm-Hmm.
Yep. Happy to be here. Thank you.
Good to have you. Good to have you. It's, it's ironic 'cause we just did a prep panel, a prep call for a panel we hadn't met in person before.
Now we meet, and then we're gonna be on a panel and we're gonna be talking about cloud native and, uh, some of the more nuancey things about not just it's microservices and your good, right? Mm-Hmm. Well, what does it really mean?
What does it really entail? Mm-Hmm. And you've also, you're, you're in the last stages of a book you're coming out with right?
About similar for the same topic. That's correct. Yeah.
Yeah. Um, finishing up the final edits now, and it's set to be, um, published next month at the end of November. Okay.
Well, I know from, I have not wrote a book, but I, I have a lot of friends who have, and I know from their experiences, you gotta have a lot of passion about what you're writing about because it's, it's a lot of work to get it all the way to the finish line, so I can, you know, admire you for getting it. It. Thank you.
You know that close to the finish line. Mm-Hmm. So, so talk to us a little bit about, um, and I understand from your book, it's not just another how to on Kubernetes and how to set it up and, you know, niceties of, of that.
There's plenty of resources to do that. You're talking more the kind of ecosystem around cloud native and patterns of applications, architectural patterns, um, how you deploy it. What are the considerations for Mm-Hmm.
Improving resiliency. Not to give an intro to your book, but it seems like, it sounds like that's kinda what you're pursuing in this. Yeah.
Um, I think you, if you take a lot of the, you know, the cloud native, um, you know, principles and what is required of cloud native applications, it kind of leads you to what the book primarily talks about, which is multi-cloud. Um, in order to actually be like fully resilient, fault tolerant, um, and also just to satisfy your operational needs for a, a workload, but also your business needs things like, you know, cost-effectiveness. You can actually, you know, in a multi multi-cloud environment, you're able to, uh, mitigate cost differences between clouds and migrate workloads from one to the other in the event of, um, you know, price increases are, are also, another big thing is availability of resources, right?
Uh, GPUs are pretty hot items these days. And so, And, uh, limited supply. Absolutely.
And that's part of the, part of the problem is you almost, if, if you are AGPU intensive application, um, you need to have, uh, uh, especially if you are responsive to customer, uh, um, customer needs and you need to scale, you need to have automated systems set up to do so, and you may not be able to do so in one region of a particular cloud provider. Um, so the book primarily focuses on, um, uh, architectural patterns and, you know, it, it certainly talks about the pros and cons and everything, and it also talks about certain necessities you need to have to complete, um, multi-cloud Kubernetes, um, from a, you know, service mesh perspective and CICD and monitoring, security alerting and everything that's gonna be slightly different. So I was not trying to write another book about Kubernetes, 'cause there's so many, but I was really trying to be forward looking and saying, how do companies prepare, um, you know, like an ounce of preparation's worth, you know, 10 pounds a cure?
Um, how do they prepare for actually having, you know, five nines in their application? And for a global audience with the minimal amount of operational overhead, uh, with putting in a lot of just, uh, putting in a lot of initial thought and preparation into launching their applications to their infrastructure and a global scale. Well, let's pull this apart a little bit because I think a lot of us, you know, talk, hear about multi-cloud.
Sounds like they're great concepts, but don't have any idea really to, to avoid so you aren't locked into one vendor cloud, right? But think about workload portability across clouds, resilience of applications, um, in multi-cloud. What are the things you do to kinda set yourself up if you're starting on one CSP?
But you know, you, you're gonna have to do a global multi-cloud Yeah. Deployment. I think the, one of the biggest things you can do is just be very forward-thinking.
You know, when you talk about cloud native applications, IAC always comes up infrastructures code. And the only way that you're gonna be able to automate a lot of systems and a lot of, um, like scalability or even first time, um, um, let's say the first time you launch something, but if you're gonna do it within an individual developer portal, right? You need infrastructures code.
And things like Terraform, um, are gonna be having, you know, a, a different provider for each cloud, right? And if you design it correctly from the very beginning and you know that you're gonna go to multi-cloud down the line, you're not designing everything for one cloud, you're saying, okay, this is the provider I'm using for now, but when we expand down the line, we're just gonna be able to modularize, you know, plug this in, and then all of a sudden you're on GCP, then it's your then, you know, I, or lineup, whatever, whatever you're looking for. Does it lead you to you using more open source things like a terraform instead of a proprietary or a cloud specific service?
Or not necessarily. Yeah. The reason why is because cloud, you know, cloud specific services are very cloud specific.
So, you know, if you're looking for serverless, um, you know, you want to be able to design it and code to where it takes advantage of Amazon Lambda versus, you know, another serverless, you know, offering for another cloud. So when you start actually tailoring your solutions for a cloud, um, offering that only exists on that cloud, you're no longer cloud agnostic. So using IAC and using Terraform, it's been around a long time.
Um, but you know, you're able to allocate resources, you know, based in code, and then depending on where you're actually deploying to your target environment, it actually is able to fill in the gaps for you. Mm-Hmm. So that just provides, you know, minimal operational overhead and quite a bit of automation to where if you have an engineer who wants to go in and create an application, you know, if there's a self-service portal they can do, so it actually generates a terraform, runs it against, you know, multiple environments.
And, um, you're just set and good to go. If I could, if I could ask a chapter in your book to, for you to write, maybe you did was talking about, um, stateless and stateful Yes. And the impact of doing either one mm-Hmm.
Are you much better off being moving to stateless and what does that really mean for, from a development architectural pattern standpoint? And are there other implications if you've gone to state stateless? Absolutely.
There's a couple chapters on that. Okay. Just right.
It is actually a very important concept, not only just to have a fundamental understanding of, but to actually, uh, implement interior architecture and design of your software, um, stateless applications that require data. They can run anywhere. They're gonna be, you know, very lightweight, they're portable easy, so obviously if you're gonna deploy them to multiple clouds, it's quite simple.
And you can have them running, uh, in parallel without having to communicate to each other because they are indeed stateless. So obviously you want to be able to identify which of your applications are good candidates for that, and really push in that direction. You know, some things, obviously you can never go stateless database systems.
However, you can create, you know, multiple stateless systems that coordinate with an API endpoint that is stateful that provides that data. And then you have a centralized managed point for your data that makes things operationally a whole lot simpler. Um, as long as they meet your, you know, requirements for latency and throughput and everything like that.
Um, it ends up being a much more easy way to, uh, manage large scale, uh, global systems. I assume there's things like even just cash technology, cash, cash aside, cash ahead. Well, yeah, that kind of give you a bit of an appearance of stateless.
Absolutely. But, and think about it this way too. Like if you, primarily in the eu, people don't usually move that often.
Mm-hmm. Right? So if you log on Facebook and in France, you know, it might take you 30 seconds to load your newsfeed, but that's because it's getting all the data from the us.
999% of the time, that's not the case. You're localized. So having local caches for your stateless services works perfectly fine.
People don't notice it's rapid fire. Um, and in the case you need to, uh, go back and retrieve a certain amount of data, you're gonna get some latency. However, that's a, such a small percentage of your use cases that it's on the whole doesn't really affect your users.
They're properly applied. You can do YouTube can do this at home, right? YouTube does it.
Yes. Uh, how about, so, so again, shifting topics here. Resilience is, you know, everybody's talking about it.
Mm-Hmm. Organizational resilience to in the cloud. Um, well first of all, how do you, how do you define resilience in the context of cloud, cloud native, multi-cloud?
Yeah, I would actually take it a bit further and not define it in terms of cloud native, but turns it from the user's perspective. It's just, it's always there. It's always on.
And no matter what happens, act of God or you know, um, you know, uh, you know, you, you run outta GPUs like the end user doesn't notice, right? You, you're fault tolerant. You're able to, uh, move resources around as needed in response to downtimes.
Like if level three has network outage in the east coast, you're able to very easily switch things over to the west coast. And actually even better if you have it an automated system that does it for you to where you don't even know, you just get a report and your data dog saying, Hey, we transferred this over here. And you're like, cool, it worked.
So expect the unexpected, right? I always, I always say, I say this too much where my teams tend to say, stop saying it, but, you know, hope for the best plan for the worst. And when I plan for the worst, I plan for, uh, what I call the seven forty seven test, which is 7 47 lands on your data center.
So it just wiped. And you don't know why, but you don't really care. You need to get the system up.
So if you actually have these fault tolerant resistance systems that are self-healing and also, you know, able to quickly allocate resources or, sorry, uh, allocate, um, requests over to another resource, like another data center, you actually don't skip a beat. You may have a few seconds of outage while you update DNS, but in the, at the end of the day, you, you're prepared for it. And, um, you users end up, you know, not noticing, which means money keeps flowing and your business keeps operating.
I always use the analogy, we, we see all these, uh, robot, uh, uh, videos of people, you know, banging 'em with a, with a two by four, pushing 'em over when they're trying to jump up on a whatever and they recover. I mean, they still manage 'em, may falter me, whatever. It's, it's, do they recover, stand back up and kind of can proceed from there.
To me, that's what resilience is. Yeah. And if you think about the best way we can bang our, you know, our robots is like high traffic road, right?
Let's say for instance, black Friday is coming up, right? If everyone is not preparing for, you know, you know, NX traffic, like I never, I never say it's a thousand X or a hundred x. You don't say that.
You say it's n right? Because you don't know. Right.
You just need to make sure you have the scalability and Whatever you plan for it would be bigger. Yeah. And, and, and again, hope or plan for the worst.
Um, but that, you know, the kicking of the robot essentially is perhaps, you know, you've reached a certain limit. So you then need to scale or you need to divide, you know, and actually, you know, expand out to multi reed or multi-cloud as well. Um, and hopefully you've already had the system set up and the, um, the processes set up as well to automatically do it in response to your traffic.
So you don't actually have to get kicked. You, you see the guy coming and raising his leg and you're like, I'm ready. And even if he kicks you, you're like, okay, we're good.
'cause there's two, there's three of me. Now You, you know, I, I'm thinking about one of the things about moving to the cloud over time. You learn some assumptions that you made when you're own in your own data center that mm, you maybe have to kind of think about it a little bit differently.
Maybe it's radically different. What are some assumptions in architectural design if you're deploying even multi-cloud, you've kind of got that figured out. What are maybe some assumptions you should rethink or think about more deeply in terms of resilience?
I think there's probably two that I would just cite. One is people assume they've done all the due diligence and requirement gathering and they usually haven't. Mm-Hmm.
So, you know, I always tell my teams also like, if I have six weeks to build a boat, I'm gonna plan for five, right? 'cause a good amount of planning, research, you know, requirements gathering, making sure you understand your scope and everything like that, um, allows for the build should just be tick, tick, tick, right? Um, so spending the more time than you generally think you need to do requirements gathering, you know, from stakeholders, from business stakeholders, from your, you know, your engineering teams, from your own team about the operational requirements.
You know, things that people don't often think about. You're like, okay, what's, um, you know, what's gonna happen if this goes down, it's a third party we have no control of, right? And then all of a sudden you're like, oh, we need a backup plan for this.
And like, what's your backup plan to your backup plan in case that fails, right? So a lot of people don't spend the necessary time. They get excited, they dive in and then they build something and it works until it doesn't.
The only things I would put in that category too is resilience to change. Absolutely. The change is gonna, we bought a company, now all of a sudden we are 10 x traffic, but we have, we have to kind of rethink a re-architect add whatever, a new system to interface to that isn't resilient.
Absolutely. Whatever it might be, how you think it is today, it's not a static thing. The design of this A hundred percent.
So, and that's the second thing I was gonna say is assume that you're going to run into corner and edge cases that you had no idea. And how are you going to respond to that? So a good example is, um, really critical, you know, um, you know, security alerts that you get, like if you're hit with one, you know, and you're actually under compliance framework, you have to actually allocate, you know, you have to resolve it within 48 hours, right?
Can you do that? Like, people don't think about that when you're building a system, but you're like, Hey, if this system's regulated, even if it's not, it's a good idea, but you know, how are we gonna be able to patch this? Or just like you said, like we're buying a company.
Um, and good good example. Um, I was part of a gaming company that got bought by a large, um, Disney and um, uh, they basically told us we had to move to Disney status centers. We had not prepared for that.
And so when we actually looked at how long that was gonna take, eight months, and they were like, no, we need it in six weeks. And it's like, you know, we never once thought about the idea that we're gonna have to move complete data centers and a, you know, just and especially do so within six weeks. So basically had to drop everything and just stop all projects to just be able to allocate that.
'cause we just never thought of that as a possibility. So it is always the, so what are all the assumptions you make? And then assume that they're all gonna be whacked, Gonna go away and also be open to the fact that they're wrong.
'cause a lot of times people think this is a solved problem, it's not always a solved problem. It may have been two years ago when you designed the system, but it doesn't mean it is today. So there's that initial shock, you know, then that robot gets kicked, it stumbles, like, you know, if you can minimize that stumble, you get back on track a whole lot faster.
And so, great. Well, um, we're gonna be on a panel here in a little bit, so we'll get to explore this some more. Is, are you independently publishing your book or is it gonna be available when it does come out?
It's, uh, with ABPV publishing out of India. So it's, um, I have it on my LinkedIn. Um, it's available on their store and on Amazon.
And then of course I'll be selling it, um, on my own. Okay. Alright.
Look for Adam Robertson. Thank you very much for joining us and really enjoy talking to you. Yeah, I mean, we covered a pretty big swath there and, uh, of course Adam took it, took it pretty deep for that short of a conversation, which is great.
Stay tuned. Interview coming up next. We'll see you in a few minutes.





