Tim Hockin – Operationalizing Kubernetes: Interview with Tim Hockin
Operationalizing Kubernetes: Are we there yet? Yes, there are some Unicorns leading the pack, but is the Kubernetes ecosystem mature enough to truly operationalize at scale? What are the telltale signs that we have reached — or will reach — this inflection point? What is on the short-term and long-term horizon that will better enable this? We will explore these questions and more with Tom Hockin, Principal Software Engineer at Google, and one of the originators of Kubernetes.
Transcript
Hi, everyone, thanks for joining us today on Operationalizing Cloud Native with Kubernetes. Really glad you were able to join us for this virtual event. We have a full day of learning around Cloud Native and Kubernetes.
We have some amazing speakers, workshops, booths and expo halls and other activities. I want to introduce you to one of our keynote speakers today, and we're really actually thrilled to have him back. This is the second year keynoting our virtual event on Cloud Native and Kubernetes.
And it's our friend Tim Hockin of Google. Tim, welcome back and thanks for being here. Thank you.
Thanks for having me. It's great to be back. So, Tim, I didn't want to mess it up and I'm not quite sure what your title is these days with Google and Google Cloud.
So why don't why don't you tell our audience and that's where we will get it right from the horse's mouth. Sure. So the official title, I guess, is a principal software engineer, but I refer to myself as a troublemaker in chief.
Ok, you know what? So I'm from a security infosec background, and we used to wear the people who like to break things and then see why they broke. So I get I get the troublemaker in chief being a very productive role.
And of course, Tim, you were one of the original team members are the Kubernetes team that Google pioneered and is now part of Cloud Native Computing Foundation. And and, you know, the rest, as they say, is history. Tim, the theme this year at our event is operationalizing Cloud Native with Kubernetes.
You and I were talking off camera. 0. It's amazing.
I didn't even think it was five. I thought it was actually coming up on four. But you corrected me.
I also mentioned that I just saw a survey that seventy five percent of new applications being deployed now are deployed on a Kubernetes is based infrastructure, which is I mean, phenomenal. Right. I think I think in five years that kind of market, not only penetration, market domination, it boggles the mind.
Right. It's a little humbling to think about how many people are really trusting the stuff that we've been doing and building. Yeah, I guess in your place, it is humbling.
It's also got to be quite rewarding, I think, to know that you've had this kind of of influence, of impact on what we're doing. And, you know, and of course, 2020 been a year for the record books with everything we've had to deal with. You know, and I always tell my children and my friends and relatives, thank God for the Internet, thank God for the cloud.
Thank God for so much of the technology that keeps us going. Right. And I think where we'd be without it, I think where we'd be without Kubernetes in 2020.
Right. It's been a godsend in that way as well. But Tim, as we said off camera.
Just because of everything we've accomplished doesn't mean it's done or we should rest on our laurels or there isn't more to do. I wanted to ask you because I mean, you're in a unique catbird seat where what what are the kinds of things you think we need to be working on to make it better to further operationalize, you know, the cloud native stack and Kubernetes? So, you know, I interact with a lot of customers both through my role in open source and through my role at Google.
And I get a good picture of some of the things that they struggle with. And yeah, we are five years in which on the one hand feels like an eternity. On the other hand, feels like a blink of an eye.
People still struggle with a bunch of things. And that 75 percent number is is amazing, but it's specifically about new applications. And so, Some of the things that I see a lot of customers struggling with are not new applications.
The so-called Brownfields and integrating there and operationalizing is also part of how do you interface between the existing applications and the new applications? How do you bridge those universes? Communities is somewhat opinionated about how it operates.
And so people still struggle with aspects like networking, like storage, which is where those are areas I spend a lot of my time. But they also struggle with cluster upgrades. Still, five years later, cluster upgrades are still not as easy as they could be.
They still struggle with observability and visibility of what's going on and debugging the system and debugging their applications. Communities is a big catalyst for people moving to microsurgeons applications. And in so doing, they're making their they're changing their problems.
I don't see them making them harder, but they're changing the way that they operate. And they need to learn how to debug across distributed systems instead of across monolithic systems. A great, you know, the brownfield greenfield thing is an excellent point.
I think we should jump into a little bit because that seventy five percent number is deceiving because it is greenfields, right? My God, I don't doesn't every engineer out here watching this wish that they only worked in green fields? Right.
And they always start with a nice, clean slate. You build what they want, how they want it. But, you know, the world, of course, doesn't work that way.
And not to mention a lot of these brown fields. And we see with communities, they're not all going to cloud or on the public cloud. Some of them are on bare metal back at the data center.
Some of these applications, many more, as a matter of fact. I mean, if you want to comment on this, I'd love your opinion more than we than we are led to believe or may think. Do you have any idea of, like, what percentage are bare metal back at the data center?
I don't have a number for percentage and I don't want to hazard a guess, but it's more than I think anybody anticipates and it's not shifting as fast as people hoped it would shift. And interestingly, it's still growing. People are still choosing to deploy stuff in their own data centers for myriad reasons.
And they're looking to get some of the same cloud native experiences, but in their data centers and to specifically link up to the rest of their experience. They might move some of their workloads into the cloud and they want to interface with it from the rest of their their environments, their stack, their network, whatever you want to call it. Yeah, I mean, Tim, we see that and we also see people, quite frankly.
I mean, they're putting Kubernetes on their bare metal back at their data center as well. And, you know, this whole hybrid thing I've heard you talk of Kubernetes on mainframe. Right.
And being able to hybrid front and back end systems of record systems of engagement still within this cloud native stack almost, which is kind of not ironic, but kind of oxymoronic that you'd be talking about cloud data stack back in the in the data center, not on the cloud. So not really native. But anyway, you're right.
So there's that whole there's that whole piece of it. Tim, you mentioned a couple of other areas, though. Networking was one of them.
Storage is one of them. Another one I'll throw out there is security. So let's take each of those, if you don't mind, and give us where are your views, you know, what can we do better?
What do you think is coming down the pike on each of those, if you don't mind? Sure. So I'll start with the things that are nearest and dearest to my heart.
The networking is where I've spent a ton of my time over the last five years. But I'm not by background and networking person. Most of what I know about networking, I've learned through the Kubernetes project.
And so I'm watching a lot of customers struggle with integration's network models. How do how do I take a criminalities cluster which loves to eat IP addresses and integrate it with my network, which may already be fragmented and distributed, and have the IP space started out across different regions or data centers or offices or clouds? And how do I how do we integrate those things in a way that doesn't hamstring our ability to to get the things that we want out of cloud native and communities.
So this is an ongoing challenge and it's something we work with on individual bases for customers, but also something that we're working with sort of across the board in upstream, trying to find the core primitives that we need to make it easier for people to do these things. In fact, I did a talk just yesterday, a webinar first CNCF talking about network models and how clusters can integrate with existing networks. You know IPV6 is perpetually on the horizon.
And it's in some sense the saving grace for a model like Kubernetes, on the other hand, it's very difficult for a lot of enterprises to reach yet. So we're working with that sort of stuff. And I think that will be a big part of getting rid of some of the barriers on the networking site.
On the storage side, especially, big enterprises are very attached to their enterprise storage solutions and they're wonderful, whether there's software solutions or hardware solutions, they've put a lot of energy into making them robust and building trust around them, and they want to use those with their communities experience. And so that's been a challenge for a lot of customers, getting the model the way that they want it to operate in terms of safety and security and migrate ability. You mentioned hybrid, and there's been a lot of chatter lately around various forums, Twitter and blogs, about multi cloud and hybrid and how it doesn't really achieve what people are hoping to get out of it.
And there's some I think there's some merit to that argument. But also I think it's maybe going down some of the wrong path. But if a customer recently was asking, how can I take a snapshot of a volume running in a communities cluster on Google Cloud and restore it on their on prem data center so that they can move their application between cloud and on the right.
And in that sentence alone, there's probably for like, oh my gosh, big problems. And so we're trying to take the lessons away from these customers and find the core things that we can build into upstream. And then security man, what do you say about security?
It's it's a journey, not a destination. And it's it's interesting to watch. I'm not a security expert.
It's interesting for me to watch from a little bit of a distance the way that the security experts and the testers are finding gaps in Kubernetes, models and finding little gotchas that they can turn into big gotchas and how we can go about attacking those things. The good news is I haven't seen anything so far that makes me believe that we're fundamentally off target. The bad news is there's a constant stream of things that we need to pay attention to and think about.
So it's definitely been interesting there. And we're now seeing sort of I think of it as a second wave of integrations where customers are not looking for deeper integration with existing security solutions and being able to run deeper traffic analysis and introspection on their cluster nodes, or being able to do more of the virtual networking sort of problems where they can do deep packet inspection on the wire and finding ways to integrate commands with that is relatively new. No, I mean, so for instance, I know Google Cloud recently announced the beta program.
I don't think it's in GA yet as a certificate authority. Right. So if you're hosting infrastructure and Google Cloud, let's say Kubernetes infrastructure.
You mentioned the IP usage. One of the things the new the new normal is now the way we do things is every single one of those IP addresses now has to have a unique certificate or should, you know, best practices have a unique certificate that identifies that IP address is bonafied, as authentic. And so with the explosion of IP addresses, becomes an explosion of certificates and with an explosion of certificates comes an explosion of certificate management and the potential for four men in the middle and certificate fraud and stuff like that.
So, you know, problems, progress sometimes begets more problems that you didn't anticipate. It's kind of like the Phoenix project and the goal book. Before that, you clear one bottleneck only to discover it opens up another bottleneck.
But nevertheless, you know, progress marches on, Tim. And I think we have made progress in a lot of these things that you're talking about. And that's part of the operation.
Have a tough time with this word, operationalizing of of this cloud native in Kubernetes is. What is a novel thing that you may run into in two or three customers once once you normalize it and operationalize it, it becomes very much par for the course thing going forward. And as time goes on, we get more of that.
Tim, if you if you've got to look, you know, short term, I'm talking about not long term. Where do you see? Where do you see sort of the biggest gaps, the biggest potential, you know, priorities that we need to address to to really operationalize this?
The biggest the biggest gaps you shared with me, the topic beforehand, but I tried not to do my homework because I wanted this to feel off the cuff. Tonight you got me. Yeah, no, no.
Big gaps are relative. Right. But what do you know?
And I know networking is dear to you. Storage is dear to you. So when you're a hammer, everything looks like a nail a little bit.
But, you know, in speaking to customers, where do you, where is the pain? So I'm not going to hit the nail, that metaphor. I think the biggest pain that people are struggling with is not networking or storage, though.
Those are issues you probably not the biggest issues. The biggest issues are things like application management and the sort of workflows and integrations with their CI/CD and how to do those things in a sane and safe way that they can manage and oversee and get information and data about, and that they can do things like hitless upgrades in a way that they trust completely, that they don't think about it anymore, that it is not a constant source of worry. And that implies their applications, but it also applies to their customers.
They should be able to completely trust that their customers are going to upgrade and not have to worry about it. And I work near a managed service and we try to make that true for users. But I don't think the level of trust is there yet.
I don't think we have the track record across the board to make people not really worry about what's going to happen. I do not disagree with you there. What about you, Tim?
I've heard people say this and your take that we need to get to the point where it really, you know, Kubernetes is about cluster management and and we really we just focus in our cluster manager. We don't worry about Kubernetes, you know, x point x release kind of, you know, idiosyncracies. We're really let's just focus in on cluster management, how we do that to how to how we deploy these applications.
We're not there yet, but we're not there. I mean, if you remember the sort of early days of smartphones and you get a notification and said, hey, there's a new Gmail update, do you want to download it now? And you go, you know, I'm traveling next week.
Maybe I'll wait until after I'm done traveling down. But that's because I don't know if it's going to work or not. I can't afford for it to be broken.
Long gone. And after dozens and dozens of these, you finally start notifying me. That's fine.
Just do it right. And we need to get we're not there yet. Truth be told, I'm still there on my iOS upgrades to what you know, I always worry it's going to kill my battery and I'm not going to be OK.
Don't don't do it when I'm going to be on the road. But but you're right. We're not there yet.
And we're still we're still struggling there. Tim, there's been a lot made sense that there's a lot made. But, you know, managed Kubernetes is it's a very popular business model.
Managed Kubernetes as a service, if you will. And a lot of people say the reason for it is because Kubernetes still too hard for many organizations. And it's just it's one of those things where maybe you're better off outsourcing it.
You were there. I mean, I don't think that was the original intent behind this. What do you what do you think about that as we go forward, less managed Kubernetes as a service, or is that, you know, the way it's you know, the way it might go?
So it was never intended. I didn't intentionally make it hard, right? It was not a nefarious plan to instill managed services, but there is some amount of inherent complexity in the problem that is just trying to tackle.
And the more customer cases that we handle, the more diversity of workloads and problems we're attacking, the more surface area we add in, the more subtle it gets. And we add every every feature we add is for a reason. And we don't we don't try to make it difficult, but there is some amount of it that we just can't get past.
Now, I don't mean to throw our hands up and pretend like there's no solutions here. We're definitely working on it as a community, trying to find ways to make it better and safer and easier for people to manage. But at the same time, I think for a lot of people, they want to focus on the things that are really important to them.
And running their clusters isn't really their their core competency. It's a word. It's a phrase that sort of fallen out of the limelight.
But their core competencies are not running a cluster. The core competencies are doing whatever it is their business does. So anything that they have to do in money or engineering that they spend to do things to manage their clusters is kind of low ROI.
And I think a lot of people just would rather pay Google or Amazon or Microsoft or anybody who's running their services for them if they can find a way to do that. And I think that's the trend, not just with the trend, with everything big databases and those sorts of things is a reason that the cloud business is growing as quickly as it is. There are cases where that doesn't work and that intersection is unfortunate.
Right? When you're on Prem and you have to run your own infrastructure, you have to internalize some amount of expertize around Kubernetes. And that's I think, where we're trying to find ways to make it a little easier for people to operate stuff and a little more reliable.
And if we get to that place where you can just say, that's fine, upgrade my cluster, I don't care, then people will sleep better and they will not need as hefty of an SRE team to run their clusters. They can focus back on the things that really matter to them. But we're not there yet still.
So, no need to panic, you're not losing your jobs any time soon, you know, just when they finally thought they were sitting on top of it. So, Tim. As we sit here now, it's, well, October 2020, this will be October 1st, 2020.
And if I ask you to look ahead to October 1st, 2021, we still having the same conversations or do you think a lot of these are in the rearview mirror? 2021, I think we will be having a different conversation down the same avenue, the specific targets of our ire will will have shifted. We will have fixed some of the networking and storage and security problems and we will have introduced some new ones or discovered some new ones.
We will have broached a new topic or new workload that we never thought about before, or we'll have some customer who has some really interesting and legitimate requirements that Kubernetes doesn't handle yet. And we're going to figure out how do we tackle those things. I don't think honestly, I don't think we'll ever get to a place when we say getting better is not a topic that we're done, we're not going to get done with this.
Just like Linux is never done. There's always more to do. There's always new hardware, there's always new services.
There's always new attack vectors that we need to think about. I don't think anything in technology is ever done, Tim. And I'm older than you and I've never seen one thing ever finished.
It's always it's always better. There's always a next version. There's always more to do.
So I can't imagine that wouldn't be the case here. Tim, one of the things that I think people are beginning to realize is that there's more to cloud native than cool. Kubernetes, Right.
You know, the CNCF has done a tremendous job of incubating and graduating and shepherding, if you will, many great projects as part of this cloud native ecosystem, if you will. I know, you know, they put out that chart with all of the cloud native projects and companies and the folks on Twitter had a field day write about what a what a mess and how unwieldy it is and all of that. And I think it's also a testament to the success of of what they've been doing and that it is such a vast ecosystem if it is a little unwieldy.
It is, but it's nevertheless vast. How does but yet Kubernetes remains at the heart of so much of that cloud native ecosystem. Do you think that continues or you know, do you think something like service assurances that kind of sits on top and we have APIs coming in and allowing stuff to come in and out, micro services and so forth?
Does that become where the where the action is or discover that is still at the heart of it? I mean, I think Kubernetes is the sort of exemplar for embracing the ideas of cloud native. You know we could spend a solid half hour talking about what that really means.
But the if I can catch a cloud native in a single word, it's decoupling. It's unhinging everything from everything else as much as you can and finding ways to make sure that you're not conjoined in failure modes or conjoined in evolution. And I think what the roadmap represents is exactly that.
People are looking to disconnect things from each other. You see a lot of projects like, well, I'm going to focus on training and that's going to be our project. And if you can decouple via a standard API, then they can have interchangeable implementation.
You look at service mesh, and you say, OK, I want to decouple my services from each other so that I'm not as monolithic and not as linked to each other. I don't have to do these one binary rollouts. I don't have to pick the same language.
And you can do this for for everything in that cloud, need a roadmap. And so I think it's it's not a problem. I do think Kubernetes continues to be at the heart of it.
But I think there's an increasing focus over time of higher level abstractions, whether those are paths like abstractions for workloads or service. Mesh abstractions for connectivity or policy management. I think that hopefully the abstraction continues to raise up enough and maybe at some point in the future, people start seeing the Kubernetes underneath it all.
And that's fine. I think actually one of the interesting things about the Kubernetes project is it may be that the Kubernetes' API and declarative style and the mindset outlasts the actual project itself. Absolutely, Tim, I got one last question for you somewhere out here right now, I'm looking into the camera.
Somewhere out here, there's a young person, maybe not so young person, but a person who's new to Kubernetes who heard about this virtual event and it was free. So they signed on and they're looking to change their life and embark on a career in Cloud, Cloud Native and Kubernetes They look at you and say, my goodness, this was a man who helped, you know, humbly, I don't want to embarrass you, but this was a guy who helped put this whole thing together. What advice do you have for that person out there right now who is looking to embark on a career around this?
What what should they be doing? Great question. I think the hallmark of success in our industry right now is adaptability and the willingness to Tibet to pivot, to learn something new and to go and start over again on a different topic.
You mentioned earlier sort of jokingly, right? Like we're not putting anybody out of a job, but we might be encouraging a lot of people out there who are sort of traditional sysadmins or DevOps people to go in and learn a different way of approaching similar familiar problems and to grow and broaden that scope. And for those people who are just looking at careers going, oh, my God, it's so complicated.
I don't know what to do with it. The managed services all have the ability to get a little bit of free service. And for me personally, there's just nothing better than rolling up my sleeves and jumping in and doing something and just trying it out and using it.
And if it doesn't do what you expect it to do, that's why I go back to those docs to go go back maybe to the community. We have a wonderful Kubernetes community and and understand what's going on. Why was it that my comprehension of the problem that was off target or maybe the product itself actually just does the wrong thing?
Certainly Kubernetes is not above reproach in terms of writing books. We have more than our fair share. And so don't be disheartened by the perceived complexity of the thing.
Start at the beginning. Simple stuff. When simple stuff becomes simple, then you move on to the more complicated stuff.
Fantastic, Tim Hockin, principal engineer, Google. Yeah, thank you. I won't do I'm terrible with titles, I'm better with names, but thank you for joining us again once again here this year on our virtual event around Cloud Native this year Operationalizing Cloud Native with Kubernetes.
We hope maybe we'll have you back next year and we could look back at this and see where we went and we'll have this whole thing out there. But until then Tim, be safe and be well. And thanks for joining us.
Thanks to you to.