Patrick Bergstrom and Yasmin Rajabi, StormForge | KubeCon + CloudNativeCon NA 2022
Patrick Bergstrom, CTO of StormForge, and Yasmin Rajabi, VP of product management at StormForge, join Alan Shimel at KubeCon to discuss day two Kubernetes operations. Specifically, Yasmin and Patrick dive into optimization and how enterprises can leverage machine learning to their advantage.
Transcript
This is Textron TV. Hey everyone. We're back here live in Detroit for kubecon.
I think the short the show floor is not even officially open till 10:30 10 minutes ago. Yeah, so no, is it best friend? Oh, yes.
I love it. So it's open. You know, they did a nice job because you don't realize it the boots is so spread apart.
You don't have that density and people getting sick. Hopefully doesn't feel like you're walking around right you feel like it's open. Yeah good stuff.
Anyway, let me introduce you to our guests for this segment right here. We have Patrick Bergstrom. Yes.
Yep. Okay. Good morning.
Yeah, it's been rajabi. They are both with storm Forge. Welcome.
Thank you for having us out my pleasure. So you got to excuse. This is the Motor City.
Yeah, and they got this like race car noise running around here every like five minutes. Well, is that what it is the monorail? Okay.
I don't know if they actually hear it at home because these mics are actually pretty good it filtering. But yeah, it's crazy. Anyway guys, let's assume they don't know storm Forge.
We've never heard a storm Forge Patrick. What would you tell people? What's wrong Forge?
So stormforge we specialize in what we kind of consider day two operations for kubernetes specifically the optimization side of the house. So you think about anyone that's running a cluster of any size especially as it gets larger, when you get to the three four five thousand namespaces and Beyond it becomes very hard to dial in each individual namespace for the right resource consumption. And that's a huge problem that we've seen at the Enterprise now where we've got these big organizations that are spending millions of dollars on cloud.
Testing when really they don't need to be and so what we specialize in is leveraging machine learning to essentially predict the correct amount of resources that each namespace means and then you can actually turn us on and automatic mode and we'll go through make every adjustment for you save you a 40 to 50 percent on your Cloud hosting and more importantly it reduces the carbon footprint of your compute capacity which is which is huge that's really huge. I registered in the times today. Yeah of all the countries who promise to reduce their yeah missions and footprints.
I think 25 out of 185 are actually doing it. Yeah, that's miserable terrible. And you know, we're all gonna pay the price for that.
Yeah. So anything we could do to do that. It is really important.
You're right. So here you can't right obviously. Crazy show but you guys had some news recently around stormforge right?
I should be before we get into that. Hold on. I I didn't I gave them your names.
I know I gave like what you do, right? Yes. I'm the chief technology officer.
It's CTO. Okay guys had a product. Okay good.
That's who our people want to talk to so good stuff. All right. Let's jump into the news then.
Yeah, so we had a launch a couple weeks ago. Actually that enables our users to scale way more efficiently. So we're honestly right now the only software that can provide recommendations for both vertical and horizontal scaling and typically what we hear from our users is when they're they have the HPA enabled for kubernetes.
It typically scales on CPU and then if they want a vertically right size that if they change the request for CPU, then it can lead to CPU thrashing. If you're trying to enable the vpa and the HBO at the same time and so the challenge that we've heard from our users is I want to be able to scale horizontally in a resource efficient manner. So I want to do the right sizing vertically, but I don't want to then make that problem exponential across my environment and so with our latest launch the machine learning that Patrick talked about we look at your data and continuously provide you recommendations that will then write size your application, but also give you the recommendation for You should scale your HPA so that you're scaling the most resource efficient manner.
I love it. You know. You started off saying it's about day two kubernetes.
Right and a lot of people here names or phrases like that. They they think they understand. Yeah what it means.
But really you guys truly are day too. You're not about how to set up kubernetes how to do all that. These are real world real life issues that people run into once they've set up their kubernetes.
Infrastructure, and now they want to start optimizing they want to start scaling horizontally and vertically, but not that one at the expense of the other exactly. One of the things I hear when I talk to people out there is you know. You guys talk like the whole world's aren't kubernetes already and we're just looking we're just starting.
I wonder you know to a certain extent we all live in our own bubbles. Yeah, if you're if you're involved in day two Cooper Netties. Eureka everyone you talk to is you know on kubernetes are kubernetes aware that day too.
But real life. What are you seeing out there from companies who come to you? Yeah.
Are they premature? I think there's it really is that day two challenge right before I worked at at stormforge. I was a VP of operations at United Health Group and we use your kubernetes and the way that we used it was very like managing a massive cluster and we had we used automation for everybody poster.
Yep. We use automation for everything that we can it was segmented. We had multiple clusters depending on like the security space you're in and everything.
I got you. That sounds like good. Yes, right.
Yeah, but no but it boils down to you know, you have all these application teams are being told. Hey, you have to go to kubernetes and it's also now on the application team to determine what they need for compute memory all their requirements and they don't know how to use kubernetes and so in the end when I put on like my SRE hat and I'm using kubernetes as an SRE, I'm gonna dial the CPU all the way to the right dial the memory. The way to the right so I prevent throttling I prevent getting whom killed in my apps happy as a clan and you just pay a fortune and then the company pays a fortune but as an SRE like I don't care about that, right and but I Mega Millions know yeah and I've seen in the last couple of years.
Now this whole finops idea that's coming about and Yasmine. I know you've talked to some customers that are really concerned with that and oh, it's a huge problem. Oh you should yeah challenges that we're seeing is that Patrick's Point like the humans are the same.
So the people that we're working on building applications like a large e-commerce app. Those settings that they'd have to tune. It was like one or two things and now they're moving those applications to kubernetes and the challenges.
Okay. What do I set it to? I'm gonna set it to the max that I can because that's what I know and what we're seeing is that it's a lot of guesswork.
Teams are at developers are looking like I'll check out my gravada dashboard. Maybe I'll just double it or something. And that's what I'm gonna go with and then the sres I'd have to deal with their like what I'm steam is knocking at the door saying I see I have these reports.
I see that we're way over spending on cloud costs, but they don't have the tools to actually go fix it. There's a lot of tools out there that'll point out the problem but actually taking it a step further to give you the configuration that goes out and fixes the problem and does that continuously is a gap that we hear from our users. Absolutely.
I think that's a huge. Yeah, because you know, you're right. Same Ops is very good at saying hey this instance is costing us.
25,000 dollars a month, right and we get you know based upon resources and what we see being out we should be able to cut that in half. Yeah. Okay got about it, you know go do it and and it's one thing to even do it.
Like let's say in a convention conventional hypervisor environment where we have 15 20 years of data to play with the kubernetes. We don't yeah for the most part. I mean maybe Google does maybe but we don't have that kind of basically knowledge base.
Yeah. What what is the right dial setting? Well, it's interesting because every application is different too.
net, it's gonna be different from java from go from python from node. Everything's gonna require different configurations. I agree Engineers to make that trade off decision of like, hey, you need to cut costs.
They're not just gonna cut down their CPU because then you're trading off resiliency, and they don't want to be the person Made the change that resulted in some error some failure. And so that's where really knowing the balance and we're like I think machine learning is a the best solution for that because then it's based on data you're deciding what those configurations are and then making those decisions based on the data and what the machine learning comes up with rather than the guesswork of someone trying to figure out how is this going to impact me? There's way too many parameters and configurations that you have to take into account and not only that but it's one thing to do it when you have a you know, a dozen namespaces, but when you're on the tune of 7,000 10,000 12,000 name spaces, like sure you can throw humans at the problem, but that'll only feel so far right it's even if you spin up a project it's like maybe you'll look at it once and then you might revisit it again in a year, whereas with something like machine learning jesmond's point.
It's perfect for it because we're it's taking the guesswork out of it and we could literally do it every hour on the hour if we wanted to so, it's great. Yeah, let's talk about all right sounds Eight. We wanted to try it.
Where should people be on their kubernetes pass before they even so they don't waste their time and then once they're there, how did they engage with stormforge? I think that's really the power of our platform. So we have multiple products on our platform.
And the first one is more of a proactive. So as you're developing your kubernetes application, you want to deploy to production. We allow you to run load tests and experiment on what is that right configuration before you deploy it.
So on that sense you're being proactive and then more the day two story that we're like really strong about is once you're there at a large scale. How do you deal with it? So when you're fully deployed you have kubernetes up and running and you're looking at your cloud costume, you're like, oh not looking great.
That's right more of a like a set and forget it because you can trust that the software will come up with the right recommendation. And if you desire actually go and deploy those records. Yeah, that's that's it's early they want out here, but it's already a theme.
I'm hearing which is hey software is Getting smart enough trust it it will tell you and do the right thing to be the most efficient to get you the answer. You need to put it out there. I love that.
How do people engage with storm forgery today goes storm Forge? Not that I owe IO. Yeah, and when you can find us in both 91 here at Cuba if you happen to be right watching on your phone or something.
And then is there a free trial is there? Yeah, that's really interested in getting started. There's a button on our website.
We're happy to you know Reach Out do a trial with you walk you through kind of the platform and we have some kubernetes experts. They're all CK. I would imagine you do.
Yeah, we're happy to kind of diagnose and then work with you on the right path forward. I love it man. io.
They are live here at good Connor Cloud native kind, you know and Check them out or you could just check them out on their website get a free trial Patrick Yasmin. Thank you for joining us part of the rest of your week. It's only Wednesday.
We're here to Friday. Yeah, that'll be a great week. Thanks actually has been a pleasure meeting you we're gonna take a break.
We're in Detroit. We'll be back in a moment.
