KubeCon 2022 – SRE Show EP 8
Austin Parker of Lightstep, Andreas Grabner of Dynatrace and Ana Margarita Medina of Lightstep join Mitch Ashley for a KubeCon edition of the SRE Show. The panel discusses open telemetry, observability into developer environments and much more.
Transcript
it welcome back. We are here with a special edition the kubecon 2022 edition of the SRE show. Yeah, like the Muppets.
we saw and this is like the first time we've met you obviously, yeah. No, that's any normally we have we're sitting across the zoom from each other. But now we're sitting across the table.
You don't know if I'm you know, like three foot six or you know, seven foot tall. There's a whole state of that right where you see people in person. It's like oh wow.
I didn't know you were tall just and it's just it just takes the relationship that we build online tent to the next place right? It's cool. Oh really big advocate of So people so I have the pleasure Mitch Ashley and being joined by introduce yourself.
I know Margarita Medina. Well, I could never say that nice. Yeah, I love it.
When you say your name Andy Andy gravner. Yeah working from Dinah Trace from Diana Trace, but we presenting Captain today to cncf incubated project that and I'm agarito also helped kind of get off the ground. Very nice.
Yeah, and I was with the light step of course. Yeah and Austin Parker also the lights and co-host of the essay show Absolutely. So Austin what you start us out, we wanted to kind of get a SRE flavor of coon.
Yeah this year. So I I think the biggest thing that I'm seeing. Yeah, just some parent contrast right the last time I was a in-person cubecon was San Diego, you know 2019.
And I remember then open Telemetry was young SRE was people were adopting and we're accepting it. And now what you see is nobody's asking what is open Telemetry. People are saying we're all in on this right?
We're all in an SRE. We're all in on devout on whatever you want to say, but the big three things I sort of see out of this are one. Obviously if employment week two is a lot of interesting about service catalogs, right like backstage for example, which is an incoming project super popular and you're seeing a lot of stuff kind of like kept in the Sort of reliability engineering and resilience engineering, you know more than maybe you used to right.
But I'd be interested what you all think. Yeah. No, I think I can just agree with what you say people they don't do it.
You don't need to explain what the district with the traces anymore and the value of it, but right and I was fortunate to be at coupon in 2020. I think so. Oh my most the last we were just in Valencia and I actually met you with years today 2020.
Yeah now 2020 21 was a physical. Where in LA right? Well, yeah, and I think this is where you could already feel it open Telemetry is really about to take off but now it has really taken.
It's definitely just the question is now I think that's what I see in my discussions. What can we do with this data? How can we make it easy accessible and how can we get the best practices from site reliability engineering into kubernetes into the hands of everybody's developers so they can really Leverage The metrics and so keep the systems reliable because as we all know right you cannot just buy resiliency, you can apply reliability you need to build it in and the foundation is knowing if the system works as expected or not.
And if it's not as expected then you need to take actions and yeah like show up the work do the work every single day. And I mean I think for me one of the things that coming into this cute corn that we're starting to see is more people are understanding that their application is very complex and it's broken a lot for them already and they really need to drill down. So we are seeing that increase in salemometry.
I'm sorry type of topics, but I've also been a really big fan of what like just Cube and Trace us or doing that. We're starting to like Leverage some of like hey, there's a lot of complexity with Bernades, how is it that we can turn it backwards to really make it easier for folks to go again like understand. Seeing it in a way.
It's kind of what you're probably hope for a couple years ago, which is we'd love for things SRE to be part of the nomenclature. Not what is it? Yeah good or same with open Telemetry.
Anyways, we're really far ahead. I think it's it's interesting because you know in my mind, I think what I said was fantastic because what we realize is that systems are complex. And what's important is being able to understand your system in to understand it.
You need to model it and to model it you need to be able to do a lot of this sort of Base, you know, this foundational work around observability and then everything is kind of built. It's building blocks on top of that all the cool buzzwords of the past, you know, five or ten years even pay us engineering continuous delivery. We'll pick more buzzwords.
We're ready. you it's great that you have those but they're they're almost the sprinkles in the ice cream right like you need. Something to kind of put them on to hold them together and I think that's what observability is proving to be for sort of.
the SRE experience like I feel like we're just sci-fi movie or something. Yeah in between there between the booming voice and the the train. But this is just yeah, you're resiliency, right?
Because I obviously like my nails I'm having. Yeah. Did you get your service level agreement breached.
I think I'm in a crash loop back off right now after on day four kubecon for me. So it's actually a couple other things I've noticed one is talking about observability all the way into the developer environment. Right?
So that's moved into the normal culture and also productivity of everybody because you know playing is just type looking for how do we get the most out of the most valuable resources whether that's a real officer developers testers. So everybody's looking for what are the time wasters? How can we get the information into the right people's hands and the right time.
I think we're also starting to see where it's like, how can I have that as in a platform way? Like I'm tired to do it for every single application for every single service and then not my entire organization is actually implementing it like This way which I think is what some of the upcoming work with Captain that we're pushing is going to be. Yeah, it was interesting as you said right in 2019 when we started with Captain.
We tried to provide S3 best breakfasts from the outside on kubernetes. Meaning we we defined slos we implemented them with kept we shifted them left. But we did it from you can integrate this in your existing tools like Jenkins and all these felt like bolted on now and that's the cool thing.
We announced this week. We we have a kubernetes operator that basically brings S3 practices into the kubernetes cluster. Where is it in a declarative way developers can annotate their deployments and then Captain can automatically talk to you observability platform and say Hey, you try to deploy this but guess what the obstacility tells me environment current is not ready.
It's not healthy. You're dependencies are not there. You might be all available budget actually currently some maintenance is going on.
You should not deploy into a broken environment. And also with the same annotations Captain can then after the deployment is finished validates your ability data and says hey, you know Hello. Are you still meeting a rest solos?
It's a system still behaving as before. You see any new logs. You mentioned Trace test.
One of the things I talked with the founder of Twisters earlier, and he said Captain would be perfect to enforce the policy after the deployment to run a integration test functional test then get the trace that is collected and then validated with Trace test if the application still producing the right traces with the right amount of data, so in case the problem happens, we have all the data there. So we're really bringing a three practices into kubernetes. So it doesn't matter how you deploy.
You can enforce pre and post deployment policies as we call them and the obstac ability is just the heart of Captain. So we're pulling the data from open telemet the Prometheus and all the other vendors all the other data sources. Awesome.
Yeah. I mean I think is that we've seen the SRE space continue evolving through all these years from 2019 to 2022 of like coupons. We've all learned a lot and we also have to remember that kubernetes on its own has changed and like the way that folks are using it but I do think that the ability for folks to understand observability and open Telemetry as a project getting a lot more mature has really helped drives that forward and hopefully continues being at the Forefront like you really can't do reliability.
If you're not understanding how your systems are doing right now or where you want them to be like by setting up those best practices would like service level indicators and objectives and the kind of take it in a different angle. Even I think what we see here is you know, yes we have it's great that everyone's together right and we've we've been able to Kind of bring the cloud native Community back in person and it's very vibrant. It's very fun.
But we also have to recognize that, you know, the world is still is change and still changing right? We haven't cleaned the exited a pandemic or anything people are still in new types of work environments remote work is still very prevalent. And I think what you're actually seeing from talking to people is that we need more tooling more ways to sort of Not just enforce back practices, but to help bring people to speed that can't have that sort of really Hands-On like onboarding experience.
Right? If you're a remote organization or you are being brought on some of them play. Like you can't just sit back and have someone you know, go tap someone on the shoulder and be like, hey like what's up with this?
Right? How do I set an SLO and having this sort of Rich Tooling in this Richard's durability data to help codify SRE practices into code and and shift things left right earlier and dip cycle and more towards the developer is just super beneficial for creating not just like healthy teams, but also resilient teams, right and it's important part of SRE. It's not just the code.
It's not just the no it's about the people and I guess this is also why you I think it brought up next stage earlier. Yeah, there's tons of potential backstage. Like honestly, I think that is like not playing favorites here outside of Oakland.
But I I actually think that backstages like the most exciting cncf projects. I've seen in a while just because there's so much that that like there's so many ways that if you applied that right and you you had things built into it would be like really transformative to how like engineering works and as a discipline, you say somewhere around what backstage is for folks that may not yeah, sorry with it. So for people that don't know backstage is it's service cataloging effectively, right?
But it comes from Spotify. They did it internally open source, and I know they still use it a lot. But it's really about how do you you know, really common problem?
That is how do you start? What is the first thing you do when you use you sit down as a team, like alright, we need to create a new service. We need to create a new thing.
How do you start? How do you deploy it? How do you monitor?
How do you set alerts for it? How do you set us a loads for it? And a lot of time generally how that works.
I think in most organizations right now is you have wikis, right or you have you know, you have a knowledge base you have ticketing system and or you kind of have a Google Docs or various other things or you know, there's sort of knowledge. It's passed down from person to person and it changes and it's different over time. But it's really backstage.
You're able to have a Catalog that says like click here and you get a new service right and you can set it up to give you all the access you need, right? It's like okay your new service cool, click this and now you get your dashboards and click this and now you get alert set up and then and it all rolls up together. The second part of it is not just what happens when I want to start but what happens when I need to like find something out right?
I've got page. There's a fire. I need to know like what are the dashboards that I can go?
Look at? What are my service dependencies, right and tools like that that help sort of overlay this invisible world that code because when you think about it because any application any software, there's two parts. There's the part that you touch, you know, there's your apis.
There's your actual database there's the pods and whatever else but then all the stuff that you feel like the decisions and how stuff gets done like that second part is super important and that's what we need to pull in to. Like the way we understand systems through technology. That context yeah that context exactly.
But different than the other context of the limit. I'm also a huge fan of what backstage is doing. I've been like following them for a few years.
So to see them grow so much by this coupon. It kind of just validates that because with working in the SRE field we see that so many companies are like, oh man now I have a hundred microservices are like Oh, no, I got to a thousand. How did this happen?
How do I categorize this and then it's like you now have to create a system for it and then make every single engineering team put in the information and keep it up to date. So once you get to stand there, I did across companies. It is easier culture to Foster within it.
Um, and I I personally like what I'm excited for that project on its own is to continue bringing that cncf ecosystem together that we're actually able to integrate with all their tooling life same way that Captain is trying to do a lot more with open Telemetry. We're gonna be tying in more projects. We have a lot of like our goal Flex work.
I really want to see backstage be like wait. No like I'm just gonna be grabbing this information for you or completely Know if a really nice platform as you spin up your new class there and stuff like it. Yeah, it seems like also IT addresses something we've talked about is especially somebody new to an application or maybe new to the essay roll.
Where do we start when I have an issue when I wanted to get into something. I kind of got hives and you said wikis every team has a different hockey and if it's always and we all know we all know that wikis are the most up-to-date thing. Yes.
Yeah. All Wiki's just perfectly accurate and more things. Well maintained perfect right there Dynamic and they're constantly changing for you is just words get put in there.
Yeah. I mean there's a lot it's information sharing in general. It's such an important part of I mean if you think about like what is the job of a software developer right or a software games you're trying to we trying to share information and knowledge and trying to it's another racial safety.
Yeah you it's making me my train of thought but you're trying to share information knowledge, you know, not just because there's an incident and you need to go in and like say hey, here's what you should do. Here's what you should look like. You're sharing knowledge to help build people up, right you're sharing knowledge.
So you can bring in people that aren't familiar with your system or people that are more Junior their career and give them the context and the knowledge that they need to become better engineers and better srees, you know, and that is It's mostly invisible work. I think we all know people on our teams that do a good job at it and mostly they don't they don't really get rewarded or acknowledged a lot. But that active information sharing is so crucial to what we're trying to do in the space that I think that's a thing that makes a cloud native great just at a base level is this is build on information.
Sure, right like this is this community here's on sharing our knowledge and our skills and our abilities and letting everyone else kind of use the fruits of that. So that's why you have major companies dedicating millions of person hours of engineering a year to improving kubernetes and to improving open Telemetry and improving all these projects because all of us are better than any of us. And all of us together always gonna be stronger too, which is I think what makes the kubernetes community so strong or on its own the project to see a be successful release after release and I think it also shows that we are sitting here, right?
We presenting different observability vendors and we're on a table Yeah, and we are trying to improve you had to head we work together. Yes, and it's awesome as I was saying to someone the other day like the the pie is big enough for everyone right? There's it's not zero some I think the great thing about this is something I personally like would open Telemetry a lot right as to put on my my project cap.
the same camp historically one of the big problems in monitoring and observability is that you as a business you have to spend so much time and energy on building Integrations, right and building your own agent and the actual Telemetry you're getting from those is fairly into undifferentiated. It's it's commodity. So open times you recognizes that and I think that some people can look at the thing and say like, oh, well, of course, that's why the vendors are in, you know, why vendors are together for because you're taking something that everyone's having to kind of duplicate their work on and pushing it on the commons and saying like Okay, this shouldn't be on us but that's actually really good for end users because it means that we can now push the Telemetry itself further and further into like Hower, you know into the underlying Frameworks right into the language runtime into kubernetes into the service position and so on and so on and so forth so that it's less work for you.
So that in the future in five years or whatever open summary is just there it's everywhere, you know, the click a checkbox because it's always running you just have to go looking for it and there it'll be yeah and I agree with you, right? Obviously. I I've been working for dynasty for 15 years.
So we build agents for 15 years. Yeah, and we love open telemates because it takes a lot of effort to keep all of these agents having up to date and a lot of reverse engineering and yeah, I cannot leave with you. I hope you're predictions.
Correct that in five years. We are there right there that all systems will be interested. We know sometimes in software.
It takes a little longer. Yeah. Yeah, we're gonna say 2027 coupon 2027.
It was 30 Apple watch timer right a reminder for five years. He's not for graduation of open Telemetry or whatever. That should take less.
That's yeah, it did actually. Oh, yeah remind me that I want one remind me that new heads by then too. Yeah.
so we all have the benefit you're talking about being around for a while and see things change in a longevity with no cash engineering SRE and telemetry. Um, what is there anything that you were hoping? We'd be talking about at kubecon.
You're kind of looking for haven't seen yet. Or maybe I know there's some gnarly things we haven't talked about yet that we need to get to or the what are some of the things like that you might not see yet. We think we should be talking about any thoughts.
So the only thought that comes to mind as we were talking about backstage was the fact that like we still don't have a really good open source Cloud native incident response that is leveraging other tooling like backstage or open Telemetry where this information starts to get populated for you and I think that's something that might be coming as some of the other projects are gonna get more developed. And and I think is that I think it's just more of being better together. I think a lot of projects are really good about running on their own and if we were to work better together, we will have a better and user story and more companies might be adopting us or willing to check us out for some of their initial work in a technology.
Celebrities the model for that, right? together I would all gone I would say open feature is another cool project that was initiated in Valencia actually, and the reason I bring up open feature open features to standardized how we deal with future flagging because there's many different features of leggings vendors out there and within the organization you probably see different tools being used also homegrown tools. So we try to standardize it the reason why I'm bringing up in the context of a Serie because feature flakes are way to keep your resist your system resilient in a time where the system is kind of challenged either with a big announcement like we should but you can use feature Flags to actually mitigate a problem in the fast way until you find the real final solution.
I think feature flakes are a really great way to build more resilience systems also allows you to better experiment on topics that you're not rich sure if they're really resilient by default. I mean that's or we've I've also used like future Flags alongside we can also engineering where you just go away. Yeah.
Oh you just get the benefit but You do it in a small smaller glass. Yeah, and also what I heard from open feature in combination with open Telemetry is not not that we only just capture when a feature of leg is turned on in a trace but even using open feature Flags to turn on more open Telemetry in case there's a problem right? Let's say Austin is on the website.
He has a problem now, I know this because no ability platform tells me Austin is shell is has an issue right now. So let's turn on a feature flag for Austin to capture even more details for him so we can help him even better. I think there's there's like maybe two things to that latter point.
I think one of the things that remains a challenge maybe for cloud native in general is Is that sort of inner project coordination right because things can move a little slow things to be very deliberate in Cloud native. And there's there's certainly a when you're on the Cadence of a project you It may be as harder to go out and figure out like oh, let's talk to open feature. Right?
So in open Telemetry, we're working on something called op amp, which is agent management protocol for open agent management protocol and that would be for that case. You just specify right like we can dynamically adjust what's happening in terms of our television. But I think it's interesting actually because you both we all kind of mentioned this right.
There's not a lot in that sort of incident response. So that sort of Fire. Maybe in analysis and response side is lacking.
And I think a lot of that is because what we've identified over the past like year maybe is there's like a really big need for an open standard on query language. Because open Club dream is great. You have all this stuff you have metrics and traces and logs and in the future profiles and other types of signals.
but if you're gonna you know ship a piece of software, you know to an end user even you might say like, well, here's the alerts. I think you should have right or here's the dash work. I think you should use.
And there's not really a great way to define that in a truly open Cloud native way, right? There's prom ql obviously, but that's pretty optimized around sort of a metrics use case and we aren't just in a metrics world, right? We need a way to query our logs, you know way to query our Our traces anyway career profiles and our ebtf data and so on and so forth.
So I think it's almost a requirement that the community comes together and finds a way to create some sort of you know, Cloud native telemetric query language that can get broad adoptions support in order to see those other things get built. We just talked about this earlier today, right as well. Yeah, it's time for us after we standardized the way we collect the data.
Yeah, how we can access the data. Yeah, and it goes back into that like obserability is code making sure we're standardizing automating things right to just make those they want they too operations like a lot easier for Developers. And I would say this is a spoiler for the the audience at home.
Like this is something that we talk about we've talked about a lot at Open Country meetings this week. Right? It's something that we had people come up and say like this is a pressing issue for us and I think The community has heard it's where the governance it doesn't committees have heard it and contributors have heard it.
So there's you know, it's two where it's early days. So I don't want to make any promises but I will say like keep an eye out, you know, maybe there'll be some news in the future. But if an audience here that is listening to you.
It's okay. They can say, you know, here we go. That's the camera.
Oh, yeah. Yes, if you would like to see an open query language tweet at office. We need like Austin all Austin L Parker on Twitter tweet.
I want an open query language and I want it now text nine one one. Okay. Yeah, so part party question and you know light step.
It's been a great partner putting Austin co-hosting putting on this Surrey show. I'd love in appreciate you joining us too because I'd love to hear from you. What's what's it like to supporting open source food community activities.
Were you working for profit company? There's probably people the company say anything generating Revenue today. Well every day we all need to be helping that but there is a it's a different kind of company to be in an open source company, then, you know, you're you know, beat it out in the market CRM software, whatever maybe yeah.
It's a different world. I I feel like there's the the easy answer is you know Community is also What's important to a success of the business isn't it's at the end of the day? Yes, it's how much money is in your bank account, but There's a lot of things that factor into that, right.
So being a place that people want to work for is important being a place that like tries to live by some set of values in the community, you know in the world and We recognize even as competitors right that we have more to gain by working together. Then we have to lose by whatever. whatever right like But I've Loved about open Telemetry is that it's really led to this unfounded level of Comedy among a lot of vendors that you know, five years ago would have been very, you know.
right like you wouldn't probably have seen a lot of Datadog or dinotrace or happy or whoever people just sitting down on the panel together and Talking Shop like this I think working on open plunder is brought us all closer together. Right and I think otherwise very few would not invest in open sources an organization like we do it like you guys do. Yeah, you would be in your own little bubble that basically you get the feedback from the people that are already using a product and they you basically tell them how the world works but the world works the way the community sees the communities driving the world forward and so really investing as a company in open source being open.
It allows you to actually learn what's really happening at which problems need to be solved and yes, we're giving away a lot for free, right but in the end there's always something as always a way how it comes back right? We are all trusted advisors. So they see obviously you guys is somebody that knows about open Telemetry the same as me.
So maybe people are more inclined today and talk to us and also look at our commercial offerings because they trust us right because we help them in the past. That's the nature. I want to add one thing to it as well that I think.
When you are in your business like you do listen to your customers a lot, right? And one thing it's been very great to see with opening Telemetry is how much you know people that are participating companies are like going out and talking to their customers and saying well, what do you know? What do you need hotel to do for you?
Right or do you do for you? How can we and then taking that feedback from their end users and bringing it back to the opening Community? Right?
So we're at we're actually able to Really expand sort of the base of people we can talk to and listen to and get actual and get actionable feedback from because of all the different companies that are participating. Yeah, I think for me the end user one is one of the most key ones is like being able to engage with them is like what struggles are you having right now? What open source Technologies might be using or like what are you biggest issues?
Whether it's organizational or even within Cloud vendors themselves and how is it that there could be more education around more implementations within vendors themself within open source projects and I think one of the things that like these companies also gain is being able to say like, hey, we all know that technology is going to constantly be changing. We want to constantly be part of that Evolution. 30 takes us two years from now.
We're on this path together for sure. It seems like two that I've been in a lot of situations there's fingerpointing between vendors and when people are collaborating together, you can get to get the folks together. Let's talk about this and why isn't this working or how do we come up with a solution or enhance this open source project or whatever it might be.
So there's there is value to the End customer in addition to open source and openness in that collaboration amongst the technical and within use so well, we're I said, we're gonna go about 20 minutes we blue way back. So that's a great job, which is great. It is phenomenal conversation.
So thank you to both of you to lightstep for sponsoring and and being this and for coasting. You got it. You've got to be all big dog tired at the end of the day today.
I'm sure you've had plenty of time. You have two more days to go. We have all of today to go.
Oh he's oh, yeah, you've done really starts around 6PM everyone knows that I always twice in my voice. We got a fish. We still have a talk going on tomorrow.
Not over till it's over. I don't know it is a long week. It is a long week.
Well, thank you Andy. Thank you for stepping in Adriana. What's her last name?
Yellow. Thank you. You trying to head to step out in the Indy stepped in for us.
So we're missing you Adriana. We'll look forward to having you back too. Thank you all thanks for joining us for this special episode of the SRE show sponsor by light step looking forward to this will be available on Tech strong TV.
So you'll be able to check out all of our episodes there as well. And for those of you are here for the live stream. We got another interview coming up next so don't go away Dave Dave's not over yet.
The week's not over yet. We'll be back soon. Thanks everybody.
Hi everyone.
