Alois Reitbauer, Dynatrace | KubeCon + CloudNativeCon Europe 2023
Alois discusses what’s next for Kubernetes as it quickly grows in adoption and becomes a key foundation for cloud-native applications as well as OpenFeature’s approach to feature flagging.
Transcript
This is texturing TV. Hey everyone, we're back live here. We just had our lunch break on our final day of coverage from kubecon Amsterdam.
It's been a great week and we're gonna continue right up here till the end my next grit, my next guest is Alawar right bower. No, so you never me. I was Alice Alice right power.
He's with Dinah Trace. He's been with Dinah trace for 17 years. I know for most of you stay in your job a year and a half and it feels like forever 17 years is almost incomprehensible.
But this is how people work. This is how it used to be we all A lot of people stayed in their job and got a pension when they were done. Now.
Everybody knows that I started to work when I was eight years old. So well, that's a beautiful thing to start at eight years old it always one. Tell us your story though 17 years.
You got to be doing something right? Tell us about your journey. Yeah.
I think it's actually not nice Segway into the story and also the story of the whole of severability space overall because very often people ask me like, how can you do the same job for 17 years? Yep, and the truth is I didn't do the same job for 17 years obviously was always a demonitoring Performance Management space. When we started 17 years ago, we had a lot of very different problems to solve.
So when dinosaurs started a uniqueness when the product came out and then it was really a startup. We were the first ones who could create distributed traces in production. Without bringing the system down.
So the challenge was can you do it? Can you do it with little overhead? Can you get all the data?
We're able to memory dams profiling things today was open to limit. We think a table Stakes right? They were really challenging back then and believe it or not.
I mean, you know back then right system was 25 machines. That was a big that was a big environment 25 notes 25 notes of websphere. That was a lot of data tools just like them.
I so so this is 2005 1999. I'm I helped start what would 98 an asp application service provider and we are offering Lotus. No don't laugh Lotus Notes hosting on it wasn't even cold websphere.
Then. I think it was called Domino. Yeah.
Actually was even before Domino and then it became Domino. and there wasn't In '98 like VMware, was it really there yet? You didn't have that whole virtualization.
Yeah. It was just yeah, it wasn't easy. It was not not easy and and the idea of being able to monitor.
that was that was it was damn near impossible. To you know to really if you had that kind of setup and we were hosting. 10,000 feet Lotus Notes installs the monitoring of that was, you know, do it right and tracing them and when we tracing right we didn't even try Tracy this before there was Diner Trace, right?
How are you? I mean Yeah, maybe some of our people could have done it. But yeah, it was just this is why we didn't make our SLA.
But anyway, you know, it's an interesting thing. Otherwise. People may not understand what we mean by Dynamic tracing.
How would you explain it? So they be the idea behind it. I mean everybody's not more familiar was the distributed tracing is certainly Tracy.
Yes, excuse me, but back then it was really dynamic because you could trace through code, but it was really at the custom instrument that had to just run that application that instrumented mode. We were the first one so like instrumented so instrumentation actually meant like adding the tracing in there what we do with OpenTelemetry now, but doing it dynamically actually meant you did not touch the application you right the proper apis whether it was a jvm extra jvms for the first ones. I was a Java most of Enterprise code.
Yeah. It was still great extent is running on Java. You changed more or less the classes as they got loaded into the environment edit instrumentation based on rules.
And then build up your distributed phrasing and then like one of the key Innovations back then we had that you could change it at runtime. So you did not have to restart the jvm. You didn't have to recompile it but you could say trace this method now for me.
It's suddenly worked. It was like massive and even instrumentation would have worked right Dave. We see more and more.
People actually, yeah that this framework even if we walk around here. Yes, we support open Telemetry. Yes, we support Prometheus back.
Then you didn't have it and like your example with Lotus Notes a lot of software back then like these massive. Dags, like what it was a Reps for you this Jay say two eeve whatever call First application service you had to dynamical instrument because like 90% of the code had no instrumentation in them. So you wouldn't have a single distributed Trace.
No, and that's what a value proposition from dinosaurs in the very beginning. It's very early days came from oh you can do it without slowing down production and you can do it in environments. We can't even touch it.
We still see them to end. It's what I remember like helping people debug websphere and we were literally Instrumenting websphere during that startup phase of Webster if you're figuring out what's wrong and for people back then this was pure magic. Like, how can you even yeah.
No, it was it was it. You know, we called the I remember those days where it was. It was magic.
It was it was magic and a lot of look a lot of our audience out here. They they don't know. What world.
Working with it wasn't an OpenTelemetry or Prometheus, but for those of us of that age of you know. We really appreciate. What these open source the cloud native?
We're here at Cloud nativecon what it's meant. Right to people doing this job day in and day out and to me this this two views here. There's people who are doing this job who can use like an OpenTelemetry and Prometheus and then a diner trace on top of it.
Yeah. It's like stealing money. I mean compared to the old way you'd have to do this.
I mean, what up, what a difference but I'd like to focus for a little bit. What does it mean to dine a Trace? to have these great open source projects kind of underlying that you can then you see you don't have to get down at that plumbing level, but you could You know build on top of that and really deliver more value.
Yeah, so obviously we also a contribute that to OpenTelemetry and we were really from day one because we believed it and had a massive value proposition. I think it's good that we started this very early days. So without this becoming a history lesson, but it I'm sorry to make it that motivate.
I think it motivates a lot of what what I also like the from a financial development side what happened because what we had to do back then we reverse engineered all of these applications. So what do we have to instrument? Where where do we get the data out?
And that was a lot of this initial knowledge that we had like, how do we do this? How do we find out massive amount of the work that was reverse engineering applications? So that could be monitored and and everybody was doing it.
We were doing it then you're relics we're doing it. Data logs back then even Wiley was doing it. Yeah.
I remember that. There was a lot of companies doing reverse engineering applications so that they could provide monitoring and when OpenTelemetry came along I was actually way more excited that suddenly all this middleware providers. All these framework providers would do it for us.
Not just that we have less work. I mean that's that's great too, but they understand the replications and their Frameworks much better than we do. And everybody's like on par in the industry we can then take this data and focus more.
How do we analyze it? How do we value? Because the data we've talked about the 25 servers that were allowed today a thousand containers is nothing so they focus entirely shifted over the years.
Let's collect the data to Let's analyze process and make sense of the data. Couple of things on that right? So first of all you also when we talk about reverse engineering these applications the vendors who made those applications, we're not real happy not really about reverse engineering the application because the applications weren't like AppSec today where you know, they're the kind of frankensteined with 80% open source glued together with some custom code maybe 10 15% custom code and there's your application back.
Then the entire application was was closed was was you know opaque. So reverse engineering these applications was a was a big undertaking it took a lot of a lot of effort a lot of engineering and then the vendors themselves would really be upset about it. So they weren't like your friend.
They weren't cooperating necessarily, right? It wasn't interesting story. But a lot of them like when we came in back in the days, it's like making movie like that.
people move coming in because most of the times we were coming in when stuff was not working so back then you sold APM tools because something was not working. There was a proactive we have to observability from a strategic perspective. No.
Stuff is breaking in production help us that's actually was. The first couple of years of my career really was where your customer. Helped them fix those the application get it back working.
And remember you sometimes you walked in there that wanted to help that customers part of a proof of concept and that there was three render sitting in there. Looking at you. According to situation you actually give them insights into their replication that never heard before and they're very often.
Don't don't listen to these people what they tell us about our application. It's not entirely true, but that very often s***. It was usually the first one.
To two days of these engagements lasted longer than one day because you you really doing performance engineering with them? But then they started to see actually. If we had this visibility.
Our engagements would be sure our customers would be let our products really better. Yeah, I had situations were like very big three-letter companies were flying in people globally for these meetings. Really?
Yeah. I mean the flight cost today we talk about carbon footprint. You don't want to know what the carbon footprint of such an engagement was.
So in the beginning that didn't like it, but then they realize they could provide more customer value. And still not not having to deal with all of this because it was actually available. We even had very finally for One customer we had Like their special support group from one of those vendors during Black Friday and we were working together.
We were finding actually errors in their code that they were not aware of and at the end of the related like this t-shirt exchange like you do not a soccer game. Yeah. Yeah, you get our special you mind we give you mine.
Yeah, so crazy. It changed during the collaboration because they saw the value eventually. Absolutely.
You said something earlier too that I think is it's important though. We've gone from valuing the ability to collect data to Value the ability to analyze data and to me that's a huge a huge kind of to finding line of what it is because Collecting data today. Look we have the databases.
We have the infrastructure where we could collect. And we're willing to pay the fees for it collect as much data as we want but making that data into actionable intelligence or finding actionable intelligence is I think really what dynatraces businesses today it is and we started like 70 years ago when we released our back then you platform for the first time we used. AI approaches like causally, I having models to vet to find anomalies in applications and doing root cause analysis on top of it.
That was really where we're like our business started to really change. We're not long about data collection that that was still an issue in some environments and needed to be done properly well back then we've been doing it for so long for any of us kind of table Stakes. What's Europe that was there?
But then we started to see massive container environments running highly Dynamic. I remember like one of those early container environments effect and a very big customer. They were running tens of thousands of containers back then on measles.
And you saw a problem and some were disappearing others were appear in the system was trying to heal itself and even trying to understand how the system was changing underneath because these systems suddenly try to repair themselves on that level was was massive and it really helped people in the beginning. They didn't trust it like okay the eyes really smarter than I am. No smarter, but think about a GPS if in a city it might you might want more attention when you're automating.
Would that brings up a whole nother issue about AI though? Right, and I don't know do we want to go there can't go there we can go. So what do you think AI what effect is AI gonna have here on on dinotrades on on this kind of business on observability.
I think there's Ai and the eye like the like put it all together in a big buckets. So when we talk about AI dinosaurs, we have actually a couple of different layers and different actually methodologies we have Then one is more focused on the statistical side like the models we have for anomaly detection some people might claim. It's not AI but it still is it's using learning models underneath yesterday is yeah.
It's still a lot of statistics on a machine learning. Well, even like l&m's GPT analysis cast statistics, but it does amazing things. Then we have what we do with the causality analysis where we have like this more or less understanding how the environment works of serving it in real life.
Doing some causal and analytics like okay, if a service that's running on a host is that's running out of CPU is slow. This is kind of related. If this is service is consumed powders.
They're having errors. This is kind of related. That's the causality explainability.
That's really that step one. I see something in my environment like a lot of things like, how are it related how they do playing to each other? That's where we have to closely eyesight and obviously recently with that catch CPT and all this OpenAI approaches.
I think the fit into a whole another category. They are more inside. Okay.
What can we do once we know something like come up with a creative, but well grounded solution to a problem. Like my kubernetes configuration is not scaling. How would I have to modify it based on the load?
Like building like this next layer or even like you have an exception in your code. You want to get an explanation you Reaching out to to a GPT type L&M model that was trained on it. But and that heat is actually old playing together like level one is really need to understand your environment.
It's a lot of the statistical part. Then you make sense out of it. That's a lot of the causally I rule-based.
Like building up all the context and you're feeding this into this, llm so that they can really do something with it. But help you maybe change your combinate this configuration. So it's working or you just say, oh I have a problem here.
My jvms are going out of memory. Can you just create me an answerable script that automatically restarts them? So why should you have to type this right?
So I think that that's how like the different layers of starting to eventually play together, but the very Central Point is what we always like to say it down it's ways as well like context is King. Yeah. How much do I know and how much understanding do I have of this environment so that I can go in that?
Direction so I think eventually it's all kind of playing together. Absolutely. We are we're past time.
Yes. Yeah, it's okay, though. We had we had to break here.
We didn't have anyone waiting. For people want to get more information about Dinah Trace. What's the best website to go?
com or follow us on Twitter or on Twitter aren't Twitter. All right. I know you've got to play to catch back home to Austria.
Yeah. So now I'm in good shape. So good luck to it.
I hope you make it we're gonna take a break. We're gonna we've got about another hour hour and a half here. All right keep Calm before we wrap it up.
So stay tuned. We'll be right back in a minute.





