Code Profiling and Observability – Ryan Perry, Grafana Labs
Ryan Perry, engineering director for Pyroscope at Grafana Labs, explains how code profiling extends observability all the way down to individual lines of code.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Ryan Perry, who's engineering director for Pyro Scope at Grafana Labs, and we're gonna be talking about continuous profiling of code, what that means and how it plays into observability.
Ryan, welcome the show. Thanks for having me. Excited to, uh, to be here talking with you again.
All right. So profiling isn't always considered a good thing in the regular world, but when it comes to code, we kind of need to understand what's going on, and I'm not sure everybody really knows what continuous profiling means and is all about. So walk us through it a little bit.
Sure, yeah, it is, uh, unfortunate that it shares a name with something, uh, you know, less exciting, but, uh, yeah, no, so continuous profiling in the code sense is basically this concept of attaching some sort of metric to, uh, to code that is running and being able to understand resource utilization for that code. So that can be CPU, it can be memory, it can be, um, stuff dealing with the network. And basically, uh, code profiling results in this output.
Typically, a flame graph, which gives you a breakdown by either like lines or something very granular that tells you how, you know, over some period of time, which is what the continuous nature is, uh, you know, how much resources were spent on line of code a, line of code B, line of code C, so on and so forth. Uh, whether that's again, CPU memory or some other kind of, uh, metric. And this has been fairly hard to track.
In fact, at least in my experience, applications and the code behaves erratically from time to time and doesn't always manifest itself in ways that we can see, but eventually it does come at a cost. Right. So are people looking at this to control costs, improve performance, or what are the benefits?
Yeah, no, there's, there's a lot of benefits and, uh, yeah, as, as you said, you know, there's, there's other signals out there, metrics, you know, standard, uh, graphs and such logs, um, traces is something that's become more popular. What profiles do is they have a much more kind of granular, uh, view on all that kind of stuff. And, uh, yeah, basically, you know, it can be used for cutting costs.
We see that really frequently, especially in today's environment where, uh, people are much more cost conscious than they may have been, you know, in, in better economic times. And then, um, for generating revenue, for a lot of companies, profiling is really good at directly translating into a decrease in latency. And as you can imagine for industries like, I mean, really in any industry, rideshare, e-commerce, uh, gaming ad tech, some places where latency FinTech where latency corresponds to either a loss in revenue or a gain in revenue.
Having these profiles to tell you, oh, this is how I can make my code faster, often correlates into, uh, more revenue there. And then obviously incident response. Uh, you know, often it's hard to put an exact number on it, but for many companies, the longer the incident is the, you know, the more money they're bleeding out.
And so having profiling is really good for, uh, being able to, uh, kind of get through any, any three of those use cases. We talk a lot about observability lately, and I feel like, um, maybe there's two flavors of observability. One is for the ops side, and the other one is for developers.
This seems like we're trying to give developers a little more insight into what's going on, but what's your sense of how is observability evolving and how do we get everybody to see the same thing at the same time? Yeah, no, I mean, yeah, that's, uh, that's the, the million dollar question for sure. I mean, I think it's, um, you know, the, the general promise of observability is being able to answer sort of, you know, new questions about your applications, about your code without having to write new code to answer those questions.
So it's being able to kind of dive in and say, you know, why did this request fail? Or why did this server run out of memory? Or whatever it might be.
And that's sort of the ultimate promise and, and compass of, excuse me, of observability. And so I would say as it's evolving, you know, it's, it's kind of changing from something where, um, you know, the, the developers write the code and then they throw it over the wall and then, you know, it ships and then somebody else is in charge of maintaining it and making sure, debugging it if it goes wrong. I think that's, that's definitely changing and, and has continued to change over time where everybody is more empowered with tools like Periscope, um, and other observability signals to be able to kind of be empowered themselves to make sure that they understand the impact that the code that's being written or the infrastructure changes that are being made have on the overall performance of the application.
And ultimately, the end users. A lot of times we will wrongfully accuse developers of being somewhat lazy and not conscious of the resources they're using. And I'm not sure that's a fair characterization.
It may be just we don't give 'em the insight, and if you give 'em the insight, they'll act differently. Yeah. Uh, I mean, a lot of times we wrongly accuse, a lot of times we rightly accuse, but either way.
Yeah, I, I mean, I definitely think that having the, having more tools available to you is, is, is ultimately what, what makes you more powerful and, and more, uh, equipped as a engineer or DevOps or whatever your position is, even executives and managers find ways to use these different tools to be better at their jobs, make teams more effective, more efficient, and that kind of thing. And so, um, yeah, I mean, it, it's really about getting all of these, these different tools and finding ways to use them together to, like I said, be able to answer a, answer a question, tell a story about, you know, what happened when somebody wants to understand that of these increasingly complex systems that have, you know, continued to grow over time. Where does this fit in the Grafana Labs portfolio?
'cause I know that's kind of increasing, but Mm-Hmm. A lot of folks, uh, you know, sometimes when you have too many tools, nobody knows what to use for what. Yeah.
Yeah. That is also, I would say on our side at Grafana, one of the things that we really focus on is finding a way to sort of bring all this together to tell sort of a story that, you know, that these, these different tools are fine individually. And, you know, there's lowkey for logs, Amir for metrics, tempo for traces, pyro scope for profiles, K six for load testing.
There's, there's many different tools that offer a different, you know, kind of angle on often the same situation or a different, uh, vantage point to view, uh, uh, incident or, or something from. But, uh, one of the things that we really focus on, and one of the reasons why we as pyro scope decided to join Grafana was because all of these tools really do work in conjunction at Grafana. And, uh, that's something that's just been in the DNA of of Grafana since the beginning of being able to bring all these tools into one dashboard, into one platform, so that you're able to, um, you know, not necessarily think of them all individually, but think of them collectively.
As, you know, you go into Grafana and you have this question, and you, you know, and sometimes profiling will give you the answer. Sometimes it's metrics, sometimes it's logs, sometimes it's some combination. But either way, you have one experience where these all work together to kind of help you solve whatever you're trying to solve or, or do whatever analysis you're trying to do.
Now, this is an open source project, so what are you looking for for help from folks? What would you like people to do? I mean, do they need to send coders or just, you know, lawyers and money?
Yeah, I mean, we'll, I mean, we'll take who whoever wants to. That's the, uh, the, the beauty of open source. Anybody who wants to play with it.
Uh, we, we highly encourage that. Um, obviously we love feedback. Um, one of the downsides of open source is that it is, you know, uh, harder to get feedback just because, you know, we, we rely on our community.
But, um, I think we've done a pretty good job, uh, both at P Scope and at Grafana as a whole of kind of cultivating that community, getting good feedback so that we can, um, kind of cycle that back into our, our roadmaps and make sure that what we're building is what people actually want to use. And so, um, I'd say just, you know, go to GitHub, uh, you know, follow the, follow the docs, uh, which hopefully will, will get you off onto a good start. And then, um, yeah, let us know how it works for you or how it could work better.
And we, um, definitely take that, uh, very heavily into account as we plan for, for what we work on, what people have explicitly asked us for from the community. We hear a lot about artificial intelligence these days and all the miraculous things that it will enable. And what's your sense of where with AI and how observability and continuous profiling might come together someday?
Yeah. Um, that's, you know, also a, uh, definitely a hot topic. I mean, we've done, uh, the last like hackathon we did, uh, internally at Grafana there, that was, uh, almost everything seemed to have a, a AI angle to it.
And so, um, I think there's still, you know, some time to figure out exactly how that ultimately manifests inside of the product. Um, uh, you know, we've seen cool things like being able to create dashboards automatically using AI or being able to, um, write queries. And, uh, you know, uh, one thing that people tend to, um, you know, that tends to be a barrier to entry for some is being able to, you know, query the data the way that they want to.
Especially if it's, you know, maybe someone who's not as familiar with code, someone on the business side, or maybe in executive who's not, you know, in the weeds as much anymore. And having AI for some of these tasks is actually really effective at, you know, kind of just parsing the docs and telling you how to write queries that will answer the questions, take plain English questions or, you know, whatever language you speak questions and turn it into, you know, something that will return a result for you in your dashboard or in Periscope or whatever it might be. That, uh, again, answers whatever you're trying to look for.
You've been doing this for a little while now. What do you have as kind of your pet peeve that you see people doing that you just shake your head and go, people were better than this. What can we be doing better?
Oh, that's a good one. Um, pet peeves. Um, yeah, I mean, I, I think, uh, one thing that, uh, that has emerged as a pet peeve as, um, like you said, we, we've, um, you know, we've been working on pyro scope for a while.
Um, we started it out of, uh, basically we were using it at our company, a, a previous company that I worked at, um, me and my co-founder. And, uh, we found a lot of value for it. A lot of times as, uh, people think about profiling, um, we're also really involved in like o hotel work on, uh, getting profiling accepted or open telemetry work, getting profiling accepted as a new signal there.
And so a lot of times people think about it as first you start with logs, you know, hello world, and then you start with, uh, you know, and then you move to metrics, and then you move to traces, and then you move to profiles. And they think of it as kind of like a linear progression of things that you add to your application when in reality, you know, it's more, you know, circular and you know, the lines all cross, uh, where, you know, sometimes you can start with profiling before you go to logs or you can start with, um, you know, metrics before you go to, uh, traces or whatever it might be. And so I would say my pet peeve is yeah, people who think that, you know, uh, there's some sequential order to, to the telemetry signals when in reality they all should just be, um, you know, if things will be much more efficient for you as a company and as a developer or DevOps person or whatever you are, if you use them together and use the right tool for whatever job you're trying to do.
All right, folks, well, you heard it here. We can actually see what's going on with lines of code and make some rational decisions, and it's not nearly as opaque as it used to be. And hopefully maybe we'll go back in and look at all that technical debt that's out there because, well, there's a lot of sloppy code when no one wants to admit it, but we can definitely do better.
Hey, Ryan, thanks for being on the show. Yeah, thanks for having me. Uh, as always.
All right, back to you guys in the studio.