Internet Performance Monitoring Wins The Day – DevOps Dialogues EP17
This episode of DevOps Dialogues explores the gaps in application performance monitoring and how Internet Performance Monitoring, or IPM, proactively monitors and assures the complete user experience from where your users are located. Catchpoint CEO Mehdi Daoudi discusses their quick-start package for organizations wanting to start their resilience journey easily.
Transcript
Hey everybody, this is Mitch Ashley. Welcome to DevOps Dialogue. This is the podcast, the interview where we talk to the most interesting people about topics that are really top of mind.
So we're, I'm very pleased to have, uh, Medi Dowdy, who is co-founder and CEO of Catchpoint joining us. Good to be talking with you again. Thank you, Mitch.
Thank you. Happy New Year. Good to see you as well.
And thank you for having me. Absolutely. Happy, happy 2025.
Good to see you, me, Medi and team. We're at the tech field daily events. I think you're gonna be at some more if you're having tuned into that.
Be sure and do that. Some great content there. So it was interesting this morning I was having this conversation about, we're a wash with data, but we're not a wash with information.
We have a lot. And that's probably true. I'm guessing from a, from application performance monitoring standpoint.
I, I know there's a lot of lights, but who knows what that all means, right? When something's blinking or that's not blinking, which give us, give us, first of all, tell us about you and, and, uh, Catchpoint, and then we'll dive into how do we deal with this and a little more, uh, holistically. So Mitch, thank you again.
So my name is Mary, co-founder, CEO of Catchpoint. Been, uh, launched Catchpoint in 2008. So we've been at this by 16 years, almost, uh, worked at Double Click and Google where I was in charge of actually monitoring.
So I was on the buying, building, deploying and using the tools to, to keep, uh, uh, the ad technology system that, uh, double click was known for alive and performing super well. And, uh, I love monitoring, uh, whether we call it observability monitoring, et cetera. The reason why I love it is because when we do a good job, um, you deliver better services, you deliver, you have better outcomes, and then you, the monitoring becomes an enabler for running a better business.
And that's what I saw firsthand. Uh, and I was very proud of being part of that, of creating what often is called the culture of performance, which is like, Hey, how can we be the best at doing what we do with the most reliable, the most available, the most fa the fastest, et cetera. And so, and, uh, throughout the journey, of course, we learned a lot of stuff, uh, which is one of the, for example, you, you mentioned it is too much data, right?
The data overload, uh, because as humans, uh, we go through some kind of outage and we regret that we didn't have the right data. And so the immediate, uh, knee jerk reaction is like, okay, we're going to log everything now, right? And we're going to log, but nobody for, nobody thinks about how much going to cost.
Mm-hmm. So then there is usually A-A-C-F-O coming down on, on you and saying, okay, you just spent like x number of millions of dollars on storage just for the monitoring system. But the other thing is like, it's, it's, nobody knows how to interpret the data.
The correlation, the causations, the connecting the dots becomes even more, uh, difficult to make, right? So the more data you have, the more cardinality you have, the more like, I don't know, what am I looking forward? And, uh, and so I was just on the phone with the customer earlier and literally the, they were talking about how, you know, we went from looking for a needle in a haystack to looking for a needle in haystacks.
Mm-hmm. And, uh, and so more data, more silos, more people, et cetera, can, can lead to prop. Now the bad thing is like, it takes longer to detect a problem, identify a problem, and resolve it.
So I think, uh, I think we're too, for some recalibration of that, what I talk to customers is they're trying to figure out a solution to end that either through, uh, you know, of course, uh, wouldn't be 2025 without throwing ai, but those are the kind of tools and capabilities that are hopefully going to allow us to go through a lot of data faster, and then maybe helping us connect the dots better. I remember a day when we used to say, the best way to provide the highest quality is don't change anything. Well, that's not, that's not even possible.
It's all changing. It's like, you know, it's no longer a solid, it's a fluid, it's under constant change, different, and Even if you don't want to change Mitch mm-hmm. The internet is changing.
Your, your third party providers are changing. Amazon is AWS is making a change, GCP is making a change. Your SA applications, all of it.
Exactly. And so how do you get ahead of that? How do you, how do you keep up with the constant changes?
And oh, by the way, you can't go to the principal's office. I say, well, it's outside of my control. I'm not responsible for availability, performance, or reliability.
Mm-hmm. You're still in the hooks, right? I can't point the finger and have that me make any difference.
Right. Well, so where does a PM kind of end it's usefulness, it's useful life, and then how do you fill in that gap? I mean, I remember a PM was, oh, good, I can have some w servers out on the internet, load my webpage and measure the how fast it loads it.
Right? We're in a much different world now, but yes. I mean, a PM that even means a lot more than that.
Well, I, first, I, I think, uh, uh, what I usually tell folks I talk to, especially on the customer side, is, you know, terms, terminology sometimes can be limiting in the way we look at things, right? Mm-hmm. So if you think of your house as your application, you have valuable stuff inside.
You need to secure it. You need to make sure that you have your humidity monitors inside the house, et cetera. But then you also need an alarm system.
You need to, you, you need to make sure that, uh, can, can, hopefully nobody can get it. Um, so, so thinking about that from, from that perspective allows you to say, okay, what tool do I need to get the job done? And the job is very simple.
You need to be up, you need to be fast, you need to be, you need to be, uh, available to all your customers, right? So if you have users that are worldwide, you need to make sure that whatever tool you have or whatever perspective you have is, uh, is represents what the end user, where your end users are. Uh, and so, so I think it's first like walking backward.
What are we trying to accomplish? What are the metrics we want to do? Maybe say we need to improve avail availability, then what are the tools I need to do to, to do that?
Uh, but a PM is still the best. Uh, and tools like Dynatrace and, and, and, and New Relic, et cetera, do a fantastic job at, at making you understand what's going on in your house, right? Being on to understand when a, when a, somebody does a search, what database it's called, how long it took to query the, try to map the dependencies, et cetera.
But the challenge becomes for some companies that, that rely on many, many other third party services who's keeping an eye on the internet stack, right? That the same way you have an application stack, uh, what happens to CloudFlare? What happens if CloudFlare is having a problem?
What happens if Akamai is having an issue? What happens if the network in India is congested? Again, being able to understand all of that.
So it's not, I, I think it's understanding what each tool does, right? Uh, I don't use my toothbrush to comb my hair, obviously. Maybe I think it'll work for me, but, uh, bad example, Maddie, but you know what I means, right?
So it's like you need, you use the right tool for the right job. Interesting. Yeah.
It, it's a, it is a great point, and like your analogy, it's sort of like driving down the interstate. I know my inside of my car is all looking good, but the road conditions can change drastically, whether, correct. Yeah.
So any, there's things outside of your control really, essentially Correct, is what a lot of that is. Well, so talk about Catchpoint and some of the lessons you learned, you know, at, uh, at double click and that, uh, at Google that informed you of, okay, here's the next approach we have to take to answer the rest of the equation of what's going on in this picture. Right?
So very fortunate enough to have been part of, of the beginning of the internet kind, ofish, right? From the commercial standpoint in 97. And, uh, and so, and then Google, of course had a super, uh, focus on the end user, right?
So if you think of Google, you think of that, that search page that needed to load in, in Subic subsequent. And uh, and I think that set the stage for the, the whole concept of like, everything needs to be available. And Yahoo too.
I mean, Yahoo spent a lot of time inventing a lot of the tools and the concepts that, uh, exist today. So the end user is where it matters the most. It doesn't matter.
And this is what happened to me at DoubleClick one day. Uh, I walked into our knock, our network operation center, and I, I saw my team chilling and, uh, you know, as if nothing was happening, uh, and DoubleClick was broken. We were not serving ads, which a live a livelihood, but all the systems were green.
Like literally all of our internal monitoring was showing, okay, uh, no network issues, no database issues. The servers were fine, 15,000 servers were, we're up and running fine, et cetera, but we were not delivering ads. And, and so that's where the monitoring is, like, what are you monitoring?
Are you monitoring for an outcome or are you monitoring for CPU and memory kind of stuff. Mm-hmm. And so that was my big aha moment when it came to, you need to monitor what matters from where it matters, right?
Uh, uh, uh, it's, it's, it's so critical and it does what, this is the philosophy that still drive us today. So, um, and that was one of the biggest lesson is again, monitor the end user and monitor where the end user is. And then also, if you're an e-commerce, then can I buy something and add it to my cart and check out, right?
If I am a sneaker company, can I, if, if, if every time I go and pick size 10 you have an error, then something, you should do something about it, you should first know about it and then fix it. So I think driving the outcomes monitoring, should you be here to help businesses run better, right? So align the monitoring strategies to the business outcomes.
That's, I think, one of the biggest things I've learned. And, uh, and when I see customers, some of the customers and partners that we do that for, it brings a lot of joy to, to me just like to see that, that causation between better monitoring, better observability to direct impact to, to, to business outcomes. It, it reminds me of the metrics we always create for our technical organizations, as, you know, meantime between failure or whatever it might be, or know these kinda responsive things that they're all important.
Yeah. That doesn't mean the customer had a great experience though. Correct?
Because failure, we live in a world, failure's gonna happen. It isn't avoid failure at all costs. That's impossible.
It, it just so much is out of our control, right? Talk, talk about, so how do you, how do you do this from the end user's viewpoint? So you really are measuring as much as possible.
Are you really assessing, I should say, the experience, right, that you're delivering? So with the concept of we, we want to monitor from, from as many places as possible to simulate where the end users are. So that was one of the design philosophies of Catchpoint.
So we do what is in the industry called synthetic margin, which is a robotic process of monitoring. Um, it's like digital mystery shoppers, you know, mystery shoppers have existed for a hundred years, uh, where the digital version of it. So we have them, uh, located in the right cities, the right ISPs, the right carriers, the right telecom carriers, et cetera.
And those things, uh, those machines, they do very simple tasks. They basically simulate what an end user does, uh, and they do it across all the different stacks of the internet. So your DNS your network, your application, your APIs, your third party services, et cetera.
And our job is to really, from there, help customers triangulate the problem. So if you show up at your doctor, God forbid, tomorrow you're going to show up with a symptom, my head hurts. Great.
A good doctor is going to go through, okay, based on what I see, let me see if it's this, that, or whatnot. So monitoring and the data that we provide that needs to help the customer go through that triangulation as fast as possible so we can reduce the meantime to repair. And so what's also very important in our business is the data quality.
So we focus on the data quality, the signal, what we call this, the signal to noise ratio is very, very important because you don't want false positive, right? I mean, no hospital can deal with like people ev showing up at the hospital every time they cough, right? That you have to have fever, this, that whatnot.
So, so it's very important for us to deliver the right qu the right metrics and the right quality to be able to drive better triangulation. So that's one thing we do. The other one is we married, we enrich the data with other things.
So for example, synthetic and run. So real user monitoring, um, fantastic, uh, uh, vast way of, of answering the question. So what, right?
So the robots say there is a problem in Saudi Arabia, Ram should be able to say, oh yes, holy cow, it is a big problem. And oh, by the way, we dropped by 30% of the traffic. Uh, so again, it's like how do you put all these things together, uh, in one dashboard, et cetera, to answer the question, what's broken where?
And whose fault is it? Right? Is it us?
Is it the internet? Is it, is it a particular third party? And all of that needs to happen super, super fast.
You know, we do this SRE survey, we've been doing it for seven years now. Uh, and I'm very proud of the work that team does. And this year, something that, uh, was very interesting that came up and is performance is the new down.
Uh, so we went from like availability, and I've seen, we've seen that with some other customers where, you know, the, on the maturity side, they cared mostly about, am I up? Are we up? Is the stuff up and running to now performance, meaning that after three seconds, even though the site or the application is up, it's actually done because the person, the customer is not willing to tolerate that.
So the, the level of, of how much you're willing to tolerate slowness is going to be an indicator of, of availability. Uh, so again, how can we do all of this stuff as quickly as possible so customers can get to fix things as fast as possible themselves. Well, if we can wrap with the AI question.
Yes. On all of our minds, everybody's talking about agentic ai. It seems like we're not very far away from synthetic users that are AI agents and you know, a world of of, you know, I'm, I'm actually out there doing multiple things 'cause I've got agents Correct where My business does.
Is there anything that customers or organizations can do to kind of prepare for that unknown of what that future may look like? I, I would imagine the more you understand an instrument and understand the experience that you're delivering today, as you add a new factor into it now, now you could assess how to manage it better or understand it better versus I don't know what I'm doing now that just makes it worse, Right? So I, I think it's an excellent question.
So, uh, let's, let's answer it two ways. So the first one is, what are we doing to prepare for a world where now there's going to be a combination of humans using the internet and then synthetic agents, right? Um, uh, that are going to be also doing stuff like, uh, there, I was reading an article where Microsoft is, is allow you to create a robot to literally answer emails on Outlook without, without you doing anything.
So, so I think that doesn't change the way we look at things, which is like availability, reachability performance, reliability are, are, are pillars that exist in an AI or non-AI world, right? I would say, I would even argue that in an, in an AI world, the tolerance for speed, reliability, et cetera are going to go down and people, we, we need better, we need faster, et cetera. So I think, I think that is a fundamental thing.
Uh, the other part of your question, the way I look at it is when I talk to our customers and the SREs, the DevOps, et cetera, um, we're all trying to do our job better, faster, and be more productive and more efficient. Ultimately, that's what the, the promise and the revolution of ai. And so what we are seeing, uh, is we're seeing customers that have, uh, a more methodical approach to ai.
It's like, okay, pick three problems that we want to solve, rather than like peanut butter kind of thing. Like, let's put AI everywhere. So it's like, okay, what are the areas where we're having a hard time finding talent, we don't have enough manpower, uh, and let that drive, uh, uh, for example, either automation or whatnot.
Uh, but on the monitoring side, et cetera, there is definitely some incredible efforts that are being led to connect the dots faster, better, right? Being able to pull all the data and solutions. Like we're seeing a lot of that in Databricks where customers are pushing all kind of data into Databricks and then being able to connect the various dots at, at scale over there.
And people are seeing some really good benefits so far, Databricks, snowflake, et cetera. I think that's one area where we're going to see a lot of stuff. Now, the benefit of that, which is we're going to have less issues where people missed an alert, because I see that a lot with our customers.
Oh, we, we got too many alerts, or somebody took a large break and we missed something, that stuff is going to go away, or it's going to supplement or, or, or augment, however you want to look at it. But I think that's one of the benefit, I think that AI is going to drive better availability and reliability to, to, to a lot of companies. Uh, Well, very exciting.
It's interesting time. You live in interesting times up. This is one of the correct, funny funnest times in my career.
Funnest is a word. Yeah. Maybe, uh, tell folks where they can find out more about Catchpoint and learn more about what you all do and get engaged with you.
Sure. com. Obviously we're on LinkedIn, Twitter X, sorry.
Uh, our blog is fantastic. Highly encourage you to, to search that, uh, and, uh, read some of the content we produce, whether it's the SRE study that, again, 70 year in a row, uh, super impressive and or the reliability and resiliency report we published. org.
I'm sure some of your listeners, uh, know about WPT. Uh, and uh, so that's another free tool that, uh, that we have for the community to test your performance. So again, slow is the new down.
So start testing. Very good. There you go.
You heard it from the expert. Well, thank you, Medi, it's great to chat with you again. You thank you.
And we appreciate everybody tuning in to this episode of DevOps Dialogue. And look for me on another futurum event or a Textron tv. He's around.
We like having him on. Thank you so much. Take care.
Happy to everyone again.




