Transforming App Performance Management – Shahar Azulay, Groundcover
Groundcover CEO Shahar Azulay explains how eBPF in the kernel of operating systems will be employed to transform application performance management (APM) in the Kubernetes era.
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with Shahara aslay. Who's the CEO for ground cover. They are pioneering advancements in application Performance Management for kubernetes.
And we're gonna be talking about what is the challenges with managing and monitoring kubernetes environments? Shahara welcome the show Hey Mike. Thanks for having.
We have had APM tools for every day. And a lot of people are also talking about observability these days what makes APM in the context of kubernetes different and why do we need a platform to address those issues? So I think a few major things of change like due to kubernetes cabernetes brings a lot of advantages.
Basically. It's it's an amazing abstraction system from you know, everything that is managing your deployment and managing your scale basically as a developer, you don't have to know how to do these things kubernetes knows how to do it for you, but it's created these obstruction also created a lot of difficulty in figuring out. what the heck is my application is doing basically so if we look like to the Past like the The back to the monolith era of you know a single machine running a single code developers used to you know, SSH into the machine to debug stuff kubernetes doesn't allow that anymore.
I mean poets just, you know, come to life and disappear and it's really hard to figure out what is going on without actually monitoring the kubernet stack and the application stack together. And the other thing is basically the scale. I mean, if before you run like a single monolith based on Java, whatever on a server and you had to monitor that process as an APM now, there's like 250 servers running all across these nodes Services.
I mean running all across season nodes and figuring out what's happening with all of this stuff more difficult. So the applications are inherently more distributed and we wind up with all these dependencies and the application may not fall over but you could spend weeks months trying to figure out why it's slow down because a particular microservice is not the way it's supposed. To and that becomes a challenge in its own right who is driving that conversation.
Do you think that that is within the purview of a devops team or is it the application owner or ultimately, who's the one who wakes up in the morning and says we have a problem using So I think that part of the reason that things became so API driven and so complex, you know, and we're seeing all these new professions emerge like SRE and you know production and Engineering basically, it's it's really hard to figure out the entire picture of what's going on in production just by knowing a single service and monitoring that or just by knowing the infrastructure and you know, the way my cicd is built and the way my my operations build to get my deployments into production. So there's a lot of like vacuum created between devops and them and I think there's a lot of different positions currently emerging to solve that. But basically we're seeing developers shifting into positions which are more either self-contained responsibility.
Basically, I'm responsible for my application from development to actual production and I'm responsible to get it monitored and working or I'm as a professional responsible for production, but I'm a Upper kind of blurring the bounds between what is devops and what is development which I think makes sense in complex environments. You can detach the two. Since becoming a little more of a team sport per se but it doesn't seem like we're gonna ship responsibility entirely left the developers.
There's just not enough quote unquote full stack people out there. However, as we go forward, do you think that it operations teams are a little intimidated by microservices and kubernetes in the whole environment. So they maybe not in engaging in as quickly as they could because they're basically how do I manage all this stuff?
Even if I do get it up and running? Yeah, I think is first it's intimidating because a lot of people don't really know the tech and I think we're at a stage of you know development and software engineering. Basically that development Stacks have become so complex that no one is expected to understand everything.
So it is intimidating. I mean, even you know, highly trained devops big organizations sometimes don't really know kubernetes all the way through they know, you know, the basics bits and bites that they had to learn to figure out how to use it. So yeah, I mean, there's a lot of intimidation there.
I think people know segments of what they're doing. For example, how should I what's like the basic best ways to separate a monolith into a microservice architecture with you know, a different business project running and how should I operate there, but And it's hard to foreign organization to actually control this entire understanding of what is kubernetes is it's right for the specific use case. How do I actually use it for my actual use case in to get the most out of it?
So yeah, I think I think it is intimidating. Basically. It's a really complex system and no one, you know, basically knows it all the way through although you can get get to use it really fast.
So it's kind of you know, it creates. Kind of a decent asset with what you can actually use and what you can actually understand. I think some developers are intimidated by that.
Yeah, you guys are using the phrase APM and other folks are touting observability platforms. Is there a difference in your mind or is one to get the other? I think an APM basically is trying to claim the full responsibility for the application there.
I think that as I mean we're used to the to either that the three pillars are observability, you know, the basic Beach of you know logs Matrix and traces and as time passed because the stack became so complex. We've we've seen a lot of kind of auxiliary tools emerging trying to cover part of the infrastructure that an APM like data dog or nearly couldn't get like, you know, getting more insights from kubernetes or getting more insights from different other deployment methods. So I think we're kind of diverged into solutions that try to help the apms get better and catch up with the technology where by saying that we're an APM.
We're trying to say that we will take full responsibility of the applications running in kubernetes for us. It means that we will monitor infrastructure all the way the obligation which basically an APM holds all these tears together using evpf. We kind of break these tears and allow all the value at once, but basically we're saying Well, we'll monitor the application from metrics to traces from logs of the application.
We're not an auxiliary tool just claiming we can get better insights from kubernetes as an infrastructure. We're going to we're going to monitor the applications as well. So I think that's that's us saying we're an APM.
What would be the relationship between an APM and Prometheus then going forward which a lot of people are using for monitoring? The cluster itself. Is that a data source for you, or do I just no longer need it?
I think for me is grafana. First of all, I mean the family of these open source Solutions first, they're amazing. We're a big fans, you know users in the past and definitely appreciative of these Solutions and I think a lot of the reason they they caught on so strongly is that apm's just can't provide the value in in high school in high throughput environments.
Basically you turn to build you turn to these Open Source Tax to provide infrastructure monitoring stuff. It's not like that data dog or neuralic or these other, you know, big players can provide what for medius monitoring like Prometheus exporters and grafana could give you as a stack but it people turn to that because it doesn't make sense to pay so much or store so much at the vendor side to get these value. So I think we're not a data source.
We're not seeing ourselves as Complement the stack but we we do believe that a lot of our customers are in this stack already already using from ideas for custom metrics which you know, basically these are fitted to the specific business logic you're running. So we're trying to bring our data which is completely self-contained ground cover provides Matrix out of the box all across the system from infrastructure to application. We're trying to bring them back as for media's compatible data source to your stack to where you're working because we're not trying to pull you away from the office or stack you're already working in.
If you have a dashboard, you know of other stuff you're already doing we're gonna complement that with you know, APM grade Matrix traces and logs that we just catch with EPF. So if that's the case as we go forward. We hear a lot about artificial intelligence.
Can we find a way to use? The data we're collecting to kind of save us from our cells. I imagine someday we might see I don't know danger Will Robinson button that says hey, you know, if you go do what you're about to do.
Yeah, you know your jobs in danger or something of that nature, but you know, how smart can smart get I think the first step there is basically first Distributing the way observability works. I mean if we look at the security domain in the past decade it clearly went through a similar process that observability is now going through only much earlier. I mean from, you know endpoint security to like edr's and these solutions that basically run rule engines at the edge crunching data at the edge making decisions at the edge close to where the data is observability is still stuck a long long way back in the position where your sheep all this data to the observability vendor and their things happen basically so you pay for this data you sheep all this data.
So but first we're Distributing the decision-making into observability stack so our agent basically can decide what to collect. Is it an interesting case to collect is this interesting fragment of data that can help you troubleshoot and there's a lot of Algorithms and machine learning and you know thoughtful thinking there on what to collect and how to collect in order to make observability first, very data efficient and very focused on the things you care about that's the first step on just making, you know, the passive part of that really intuitive smart integrated with the way you actually troubleshoot and work as an SRE as a devops as a Dev later on it can definitely go a step further into more the active part of you managing production and doing other stuff there and does this deployment make sense for for example, like, you know can already deployments is is this deployment actually working as expected right now in production should be alert on that, you know based on the on the different baselines at the agency on the previous versions of judicially deployed and so on so definitely a lot of the logic is gonna go there. I think the first part is Distributing it and creating agents or you know, front and parts running at the data side that can make these decisions.
in a way that you don't have to start tons of data just to get these insights. That's the you know the first part that's that we're trying to At least push forward. So the idea is rather than choking on a massive amount of data that I'm centralized somewhere.
I am doing the analytics more closer to the point where the data is being created in the first place and then I'm kind of sharing the aggregate results in a way that makes it possible for mere mortals to understand what's going on. Is that the core idea? Yeah, I mean basically first of all using evpf in a high level outcome of product.
That's the first thing grancover does EPF is a really raw technology. It's amazing. It's an amazing sensor.
You can basically get insights from anywhere from the kernel to do, you know, the the infrastructure of kubernetes all the way to the application. First of all, we're trying to you know wrap this up for the developers in a really high level input as you say that a person would figure out what's going on, but basically also, Um kind of digest the data on the Fly we create for example, a lot of Matrix from raw data as they fly by through the agent without writing them to disk or sending them anywhere without having you pay for you know, say like 10k of spans per second just to get the latency Matrix you wanna you know, put your SLO on as a business. So we're creating a lot of stuff inside the age of the fly letting it learn baselines and figure out if they're deviations for baselines and so on and capturing raw data when the rule-based engine basically decide as something is interesting like a high latency request or a broken status code or some kind of deviation from you know, normal behavior or anything a lot of different stuff that we're currently doing in the age.
It's like All right. So this is not your father's proverbial agent. This is taking advantage of EPF at the micro kernel level inside the operating system to capture and process that data that's required.
Do you guys think that ebpf is going to change the way we think about monitoring observability because of this capability. I'm not sure I think maybe you're one of the first platforms out there that's invoking that subsystems. So what's your ultimate expectations for me BPM?
And first of all, I think I mean ground cover is one of probably the best teams in the world only BPF, but it's not because we're amazing. It's because we have to be really very good at it because it's still erotic technology. It's still a very it's still very fast moving technology.
I mean the community currently there's kind of moving under our feet as we contribute back to it and and you know take value from from the community back to what we're doing. So for example, you know, even the protocols that EPF covers and the capability that it covers are changing on a monthly basis. So we have to be very good at and do as much as we can pushy VPS or big Believers and you have due to that but we believe that in a few years the EPF is gonna be a commodity all across the stack.
We see if the networking like service meshes and you know Network policies moving to EPF. We see it in security already a lot of security Stacks like endpoint security and workload security are already moving to being executed by EPF and Their ability I think is specifically interesting because it's also changes the organizational behavior of what observability is you're used to integrating an APM from the dev team here is to talking to the dev team, you know data team and you know back in team, please integrate this SDK into your code and that's deploy to production and and see the value of this APM. We just deployed suddenly one guy in devops for example can deploy you on a production at a thousand people company.
So it's also a really interesting organizational change of how hard it is how pre-planning is involved in integrating APM and I think ebm is definitely gonna create major major shifts there in how how fast you can deploy and deploy try out the new APM. I think it's really different than what it is today. It's much much harder to take to do these dogs.
Do you think apms and Cloud native? There will be more pervasive because in Legacy monolithic environments, it was people use it. Most Mission critical application because the agents were hard to deploy manage and in the whole thing was kind of expensive.
So are we getting to the point where apms will be a much more democratized thing in the land of kubernetes? Yeah, I think yeah for the first things we see teams currently doing is as you say, I mean, they're they're forced to be in a position where the they actually calculate the trade off of costs and and visibility of their application doesn't make sense. I mean, we see teams actually making decisions of you know, let's not deploy an APM on this application.
It's too high scale. We can't we can't keep up with the cost there and let's just monitor this one which you know is Mission critical and and stuff like that. This is a position where you don't want to put developers.
So first of all figuring out how to make APM scalable regardless of EPF, you know how to sample the data differently how to make things not volume-based price to not cause so much pain and so much, you know, kind of fear from the unpredictable build that's gonna come in the next month for teams around APM. That's the first step. And I think later on EPF also unravels a lot of different stuff about what is covered.
And what is not I mean you're used to choosing kind of cherry picking where you integrate the SDK and we also see it for example in Legacy and code. I mean, there's there's always this Legacy code no one no wants to touch in in a high end in a high scale Enterprise, you know one wants to change the code. No one want to integrate new stuff there.
So it's always left behind and ebf. Suddenly monitors everything we kind of not differentiating between anything so we can also see for example the control plane of kubernetes we can monitor. Sidecars if you have them in production, like nginx's running in the in the background if you if you're using them as proxies or whatever so we can actually monitor everything that you're running even though you didn't cherry pick where to to kind of integrate the APM and I think teams are surprised about you know, what you can cover with an API and there used to seeing it from the perspective of what they chose to monitor rather than what they're actually actually is going on in production.
All right, cool. Hey guys. I heard it here.
You don't have to tell anybody anymore that their application is just too big to manage or that matter not important enough. You can give everybody the equal attention. They deserve Shahar.
Thanks for being on the show. Thanks for having me. All right guys back to you in the studio.