Eli Birger, PerfectScale | KubeCon + CloudNativeCon Europe 2023
PerfectScale CTO Eli Birger discusses the challenges associated with controlling and governing Kubernetes capacity at scale, which makes it complex for DevOps and platform engineering teams to balance the cost and resiliency of their environments. Additionally, he explains how the PerfectScale Platform gives teams the visibility and insights to ensure peak Kubernetes performance, at the lowest possible cost.
Transcript
This is texturing TV. Hey everyone. We're back here live at kubecon.
The floor is open. They're kicking soccer balls over there. Everybody's walking around and we're bringing you day three coverage here Friday from Amsterdam.
It's been a great show. I want to introduce you to our next guest. He's the co-founder and CTO of perfect scale.
His name is Ellie bierga. Yeah. Hello Hi, how are you doing today?
Hey Ellie. Nice to meet you. Nice to have you on the show.
So I mentioned that you are the CTO and co-founder of perfect scale. Before we talk about perfect scale though. Let's talk a little bit about you.
Tell us kind of your your journey. Yeah, so my journey, I managed devops for many years and I established in last few years. I established many multiple SAS platforms big product multiple environments mainly based on the kubernetes in recent years.
Yeah. Well, everything is based on kubernetes and in recent years. so Ellie I founded a few companies myself usually for to found a company you you got to see a problem and it's most often a problem that you will encounter in your work life.
But you say to yourself. I can't be the only one. This is not just my problem.
This is a problem other people have too. And there's got to be a better way. I got to build something that solves this problem.
Is that kind of what happened with you with perfect scale tell us how you came up with the idea. And what drove you in perfect scale. Yes, I'll an absolutely so that's exactly what happened to me.
Do you know when you establishing a new thing in this case? It was a it's fast product based on the kubernetes. There is a problem of the day one of operation when you building everything when you hooking up the ci/cd thinking how do you build everything new stuff?
And then you found yourself with the problem of the second day operation? Yeah where you actually build everything, but now you build it for many years forward and now you have to maintain and this poses the company for the favorite for a very new and different challenge of the second day operation and there we found it's extremely challenging. To maintain those clusters that we build over a period of time with many different touching hands with many chefs in the kitchens developers are keep developing their applications.
The customers are coming and going the part of the low patterns is always changing. How do you maintain and make sure That the Clusters are set up properly are capable to withstand the load of your customers the fish effectively without wasting too much resources on the other hand and I it was extremely hard for me to find those needles in the haystack of data of where I'm not perfectly tuned and perfectly scale. So this is this that was the point where I met Michael Founders, and I said guys we have a big problem here.
So we started by talking to other companies and asking how do you guys solve this problem? And we found that this pain is common and that was the point where we decided that. Yes.
We are going to solve this problem for everyone. That's great. I love those kinds of stories because a really gives you insight into it.
So You found a planet scale. Well a dumplatic scale. Excuse me.
Perfect. Scale. Perfect scale.
Yeah, there is a company planted scale too. I remember them but perfect scale tell us and share with the audience. What what how does perfect scale solve this tell?
You know, what what is the company do exactly? Yes. So we developed a product that helped devops and SRE teams to understand exactly how the systems how they clusters in multi-cluster multi-cloud environment behaves because normally you start with a single development cluster you develop your application, then you spin up the your staging and production cluster than the Clusters are growing and more and more teams are evolved in World in this So how do you make sure that the people are making right decisions in terms of resources?
Because right definition of resources actually affects your capabilities of scaling the horizontal photos of scalar the cluster Auto scaler every single days on those definitions and those definitions might be set correctly at the first moment, but they need to change later on or they might be setting correctly for the first moment and creating problems like Like waste of cloud resources on one hand or problems with resilience. If you are underestimated how much resources do you need? Then you have evictions you have out of memory.
If you trolling your customers are suffering. So we are running inside the cluster and making constantly sure that you have the exact right amount of resources for all of your Fosters. Love it Sally.
I've been coming to kubecon. for a long time a lot of cube counts here and the US it used to be kubecon was mostly for Developers. right it was about Multi multi-threaded apps and and you know developing applications and then using kubernetes in order to deploy them.
It was very developed percentric one of the trends. I see at the show this year in recently is cloud native is not just for developer certainly as a matter of fact, it's more for platform engineers and sres and the Ops people as much as it is. for Developers And and really, you know you when you look at the perfect scale Mission, this is an ops mission.
No doubt about it, right that's clearly Ops and it brings up something. You know, we've seen the platform engineering you're very so you get these platform engineers and they're sort of like architects in the beginning right? We're gonna and you hit it.
It's that day one. I'm gonna design the perfect perfect kubernetes platform for this app. And day one it is it works great.
But then real life happens, right and this gets pulled and pushed and that gets changed and that God said here and we found out about this and this and all of a sudden. Wow, we had like, you know, we scored scope creep, right? We had some creep here either.
We got to bring it back to that Christine day one. Well, we got to say no this is this is what it's like in real life. We got to make the system fit.
We can't make the app regress. And and that's really the mission. It sounds like from perfect scale right?
You are absolutely correct, and absolutely right that I can could not agree more. In general the industry of cloud native already passed the first stage where only developers mainly played with that and have this application or that application. This is now the reality we all live in this Reality main.
The main platforms that we all relay on starting from I don't know ordering a taxi ordering a food ordering whatever banks will everything is running to the in kubernetes and there is a lot of operation. because like someone have to maintain it has someone have to maintain it effectively. And this is the and now we're in the face where those tools that help operators to operate effectively are coming to the market.
Absolutely. So when I'd like you to do Allie, I want you to describe to the audience. So, what's the typical on-ramp?
Right. So a person maybe it's a platform engineer or an SRE or what have you to recognizes the problem. And then how did they get out to perfect scale?
And how do they how did they engage? And then, you know hopefully make their life better and easier by using this. Yeah, so basically there is there is kind of a paradigm that we are trying to change the Paradigm of constant fire fighting people are relaying on alerting and Monitoring Solutions to fix the problems once they occur however, this pattern is not scalable and in my personal opinion, this is the wrong yeah, okay that you burn out exactly so what we are trying to do we're we are observing everything that happens in your pastor we understand this is aality waves we understand the overall trend we understand that normally is that happens in your cluster and affects your scale and then we are capable to say Hey you are over provision here.
You just wasting a lot of resources. Hey, you are under provision here you have resilience problem. You most likely gonna fail or you already failed but your monitoring system just And no one too care.
So you keep flopping and flopping and no one takes care of that by that we achieving two things the first one like what's the mission in general for why we building those platforms in kubernetes. We're trying to get the best possible performance for our users. But if you ask also the sea level management, we're trying to achieve this best possible resilience and performance at the lowest possible cost.
So it's very hard to organize all this scattered data. Okay, because we have this the cats versus pets versus cattle Paradigm. We have many developers.
They have many microservices each microservice configure differently in each cluster. How person expected so someone who operates the system how he can analyze find this needle in a haystack of the proof of the particular problem, even before it happened and know how to address so perfect scale the does this analysis bring understand what's going on where the problems are again two types of problems the over provisioning and they're under provisioning sure and then it's also ordered them by the impact on your entire system. Bring all the needed evidence so you can immediately Act Perfect, man.
I love it. Give us an idea of like and if you're not comfortable don't but what does it cost for the like, how do you price this? Yeah.
So our pricing model is relatively simple. We are consumption based. It's based on how much resources we are.
We are monitoring for you. The most important thing we will be always on the positive side in terms of Roi the payment for for our platform will be much less than you will be able to save when you use waste and not only that you are not only reducing the Cloud waste in terms of money. You also reducing the carbon footprint of your systems, which is important today sustainable.
Yes. We are at the end of the day. We're all sharing the same Planet.
Yeah. No, let me put a quick plug in. We have a new video series coming up Barney Schneider who's written several books on climate change and so forth It's called Echo sustainability sustainable it you can watch it on Tech strong TV.
I think it started yesterday for Earth Day and it's a it's a big it's a big topic and it's gonna get bigger. So that's important. Yeah stuff Ellie Ellie.
We're about done. I want you to look into this camera though and tell people where did they go engage? Perfect Scout?
Yeah. So we are our boost located there in the area of startups. You will see how on that side.
io. This is our website that I know that's correct. Excellent Man Ellie.
Thank you very much you okay? Hey, perfect. Scale die.
Yo go check them out day 2K8 operations. You've got to be thinking about it. We're gonna take a break.
We're live here. Let's Dave kubecon. We'll be back in a minute.





