Mehdi Daoudi, CEO and co-founder of Catchpoint, joins Mitch Ashley to discuss how Catchpoint simplifies observability with Internet Performance Monitoring
Reliable and resilient digital experiences are critical to nearly all our businesses. Digital resilience requires full visibility into the entire Internet software and infrastructure stack. Mehdi explains how Catchpoint’s Internet Performance Monitoring Platform has added new capabilities that give ITOps and DevOps teams visibility from the user across the Internet Stack to the application code.
Transcript
This is Textron tv. Well, I have the great pleasure of being joined by Midi Dowdy, who is co-founder and CEO of, uh, Catchpoint. Welcome.
Good to be chatting with you. I'm still catching my breath from our week together. Absolutely.
Yeah, Mitch, it's, uh, it's, uh, two in a row, I guess in, in, uh, less than a week, so, but it's a pleasure to be with you again, Mitch. It is, it is. It's great to be back with you again.
Um, I had the pleasure of, uh, also being part of, uh, app dev Tech Field Day in Santa Clara, where Medi and, uh, members of the staff presented and talk with, you know, an audience asking a lot of tough questions, I think give very call. Yes. I really enjoyed, I really enjoyed the format, to be honest.
I think it's, uh, I think we should do more of these things. That was my message back when I came back. Well, that's good.
That's great to hear. That says that it, uh, you got a lot out of it. I know we had a lot of viewers.
I'll put a link to that video in this description when this goes up so well, for anybody who might not know what you all do, which Catchpoint is and, uh, else about that. And then let's get into kind of resiliency and how we need to do more than just kind of standard observability, maybe what some of the things that you all add to that picture. Sure.
Uh, thank you, Mitch. So, so we, we started the company in 2008. Before that, I used to run operations and monitoring for a company called Double Click, which was acquired by Google.
So I did that for 11 years. So I was battle tested on the other side of, of, uh, of, of, of thanks, which is on the receiving end, on the, on the, on the customer end. And we, we, my team and I wanted, uh, to basically create a better version of the end user monitoring tools that were available back then.
And so we embarked on this mission, uh, uh, to basically create an internet performance monitoring company. And we've gone through different categories, et cetera, observability, et cetera. But I think at the end of the day, the, the best way to describe what Catchpoint is, is we are a company that allows, uh, companies to basically have a good understanding of what's working and not working from an internet perspective, from an end user perspective, whether it's your infrastructure, your services, your APIs, your ecosystems and whatnot.
And, and really reduce that meantime to troubleshoot in a war room, right? So you go into a war room, there is a problem, and basically monitoring is like peeling an onion. So it's painful.
You want to eat the onion, but it's, it's a, it's a pain to get to the root cause. So how can we get as quickly as possible to where the problem is and, and eliminate all the things that, that could have caused the problem. The challenge now is just things have gotten so much more complex that finding that root cause.
It's like we, we don't look at the, at the needle in a haystack. It's really finding a needle in multiple haystacks. So now I'm really playing peek with the, with the needle across this cloud vendor, that cloud vendor.
And it's, it's very hard. So some days it feels like a needle and a stack of needles. Yeah.
They, or even better, I might, I might steal that. Uh, that from you. No, you're have at it, you know, um, I think just the human mind likes to compartmentalize problems, right?
So we understand what we understand and we either assume the rest is okay, you know, if we're running our own stack and applications and, or maybe we blame the things that we don't, don't control, don't have visibility into, but we don't have the ability to then pinpoint what that is. And if we don't consider the entire kind of system, if you will, of how what we're operating within, whether it's our code or our infrastructure, cloud infrastructure or DNS or, you know, security systems that are managing traffic and blocking bots, whatever it might be, it could be so many factors that are actually either singular, singularly, or in multiple ways contributing to a slowdown of performance or not delivering the kind of experience that you want. Right?
Absolutely. So I think the, the, the way you described it is really well, so think about the, for me to open a banking website or a Amazon or anything that we all deal with on a day-to-day basis, to either order a cab or order food, or it doesn't have to be always a technical thing, but at the end of the day, we're delivering customer experiences to people around the world. Uh, and, uh, so that delivery chain is almost like building a car.
It, like, all the things need to be assembled at the right time, and you cannot have a single millisecond except that it's not sequential. You have multiple things happening in parallel to make that browser page, uh, look good on your browser or on your mobile app. And so the, the complexity is just increasing there.
You know, I think when I started in this business, uh, at least the Catchpoint one, I think a single webpage had what, maybe a hundred requests, 60 requests, 60 different objects coming together today. It's not uncommon to see 600, 700, a thousand. I had the customer a few years ago ask us to basically increase our, what we call our waterfalls, to be able to display 2000 objects on a page.
So we, so you have 2000 things that need to happen at the speed of light, literally for the user to have a good user experience. Otherwise, there's, you know, if you're ad supported, you, you drop your ad revenues if you're, if you're selling software online or whatever it might be, uh, the, the impact can be huge. But this, this literally this complexity that is just mind boggling and then figuring out where the problem is so you can call on the right person.
Uh, uh, somebody on my team was on a crisis call with the customer of ours because sometimes we, we are part of the crisis calls and there were over 900 people on the call. Wow. So two, three o'clock in the morning, because usually somehow outages always pick two, three o'clock in the morning to happen, right?
While you're on vacation. While you're on vacation, or usually on Fridays even worse, right? So 900 people on the crisis bridge to say, okay, is it this, is it, this is it, this is it this.
And, and because there's so many systems that are involved in, in, in, uh, uh, uh, a bigger system. So, and, and that's the complexity is just growing exponentially, to be honest. And that's scary.
And there is no, there is no be, there is no universe where this is going to shrink to, oh, we're only going to have two calls on the webpage. Right? Things are just getting worse by the minute.
I have trouble getting, uh, three kids and, and dinner on the table to all show up at the same time. I can't imagine a call with 900 people, And you have no lead network latency in your house, so That's right. Right.
I shouldn't have a problem. You know, I'd, I'd love to hear a couple of things. One is, you know, you, you've been around since 2008.
Um, we all went through the, uh, covid crisis, the work from home. But one of the things that changed during that time period is of course, this rapid move to digital experiences. And we've, you know, accelerated so much of our work to become digital, uh, whether it's delivering to customers and partners, operations of our own business employees, et cetera.
How, how did that, that window, that emphasis and that rapid rise and, uh, really putting the emphasis on digital experiences shaped what you and, and Cashpoint does? Yeah, it's a great, literally, you, you nailed it when CO was the greatest transformation from a digital perspective that, uh, uh, that anybody could have asked for. I mean, I, I hate to use that because it's a lot of people died and it's a catastrophe at human scale.
But from a digital transformation perspective, that was the biggest boost that, uh, that has impacted every company. E everybody that had plans to transform or to improve their digital delivery, et cetera, had no choice. I mean, you had no choice.
You either ab embraced it or you died, or you literally just disappeared. And so, uh, I I, I, I think it was in New York where people had to learn how, you know, 70-year-old, 80-year-old people had to learn to use their cell phones. You get food delivery.
I mean, think about the digital transformation, um, at that level, right? Because there was, if you didn't eat, obviously that was a problem, right? So, um, so CIOs, CTOs had no choice but to embrace digital transformation at, uh, at the rapid scale.
And I, we saw huge, huge, huge transformation from our customer standpoint. I think the biggest thing that is still lagging in my opinion, is as much as most companies, obviously, uh, I don't think there, I don't think you talk about analog versus digital. I think now today is just like, it's digital, right?
So it's a, we assume as digital. Um, but I think from a, from a maturity model, a lot of companies are still thinking about, uh, from a binary standpoint, they're think they're talking about availability, for example. They're not talking about performance.
Where I think digital experience is the em embracing both availability and performance, because the customer expects a great user experience. And so I think when, when people don't think about user experience and they just, well, the website or the mobile app is up, you know, my job is done. I think that's a great disservice because the customer is expecting the journey to be concluded in a good, timely manner because their time is valuable.
And so that's the biggest, uh, in my opinion, still the biggest roadblock is companies and organizations still resisting the, the, the adoption of performance as another pillar of their digital transformation. I think one of the other examples that that's fascinated me the most, and, uh, we were involved in some, some, some work with some luxury brands, but when you think about it, one of the biggest, I mean, obviously people had to eat and people had to go to the hospitals and whatnot, so I think that was covered. But you still have people that had to buy Rolexes that still wanted to buy Rolex, and you couldn't go to a store, right?
Because they were closed. But the fact that even the luxury brands had to adapt to Covid, how, well, one of some of them adopted the ability to book a, a meeting online where you can go to the store and make a reservation that didn't exist before. Uh, we had customer, we had, we had some of these brands where you could basically have the, the, the piece of jewelry delivered to your home for you to try out.
So it just led to so many new innovations in terms of like, how do we deliver user experience to the customer in a different way? And I think that level of creativity was amazing. But again, performance is at the, is needs to be paramount in, in those, in those journeys.
When everything is in a digital marketplace, if something about the experience is less than satisfactory, changing is in many cases, this was moving to the next option, right? Absolutely. I, I, I, you, you know, so I, I use this example.
So in the good old days, you would get in your car, you would drive to Costco, right? Or to Sam's Club or whatever. Uh, and if, uh, the store tells you it's open at 9:00 AM then you expect the door to be at 9:00 AM If the door is not open at nine and you wait another 10 minutes, okay, nine 15, it's not that I have to go and move on with my life.
You have to get back in your car, drive another 50 miles to the next closest location or whatever. That is a very annoying thing. Now, online, I don't know about you, but I can open a tab on my browser in less than 200 milliseconds, and I am gone.
And not only am I gone, but the image, the brand image you left on me, on my brain is tarnished. And so how do you, all that money that you spend building, the trust is gone, and it's really about trust. At the end of the day, Mitch, I really believe that everything we do about it, in terms of delivering great user experiences is to establish the trust between a brand and their customers.
I, I heard that, I learned that from this great guy, Rob Marquee, who was the inventor of the NPS scoring at Bain Capital. It's really about trust. Like at the end of the day, this is about trust, period.
We're handing off our data, we're on the digital journey experience with Right. Uber website or application or, or whatever it might be. Talk, talk about that.
You mentioned, you know, 70, 80-year-old folks having to learn how to order food off their mobile app. What, what is such a move to mobile in addition to web? How does, how does that change the world for your customers?
It used to be kind of mobile was, yeah, I'd like fries with that. I'd like a mobile app too. Now it's mobile first, or equal mobile to web app, which you can, what you're delivering to your customers.
Sure. So mobile means a lot of things because you can, mobile is is it cellular mobile or is it wifi mobile or is it just, but bottom line, it's, what it means for me is the same, which is you need to deliver amazing user, user experiences anywhere, anytime 24 7, whether it's on a mobile device or on iPad, or a tablet or computer over a ER or wifi or 5G. So the complexity I was telling you about just keeps increasing for, for the customers is now they, they have to keep delivering, uh, the same levels of stuff through different channel, right?
So when we talk about omnichannel, it's not just store and, and online. It's, it's online desktop, online mobile, mobile app, this 5G wifi, six G six, wifi, six, et cetera, et cetera. So I feel for our customers, because they really have to keep delivering, and the bar keeps getting raised over and over and over again.
Um, I think it takes a great leader from on, on the IT side on, on our customer side to basically be that visionary person that's going to, uh, establish strong SLAs in terms of delivery, in terms of rallying the troops, in terms of saying, Hey, we're going to, we're going to deliver amazing user expectations. So user, user experiences, um, user expectations are Freudian slip here, right? So, but it's really about having the leaders set the bar very high and say, no matter where the users are going to meet us, we're going to meet them there and we're going to deliver awesome experiences.
And whether it's 5G, it becomes a non-issue. Some, some companies are resisted, right? Some companies are, well, we, we don't care about 5G because, or mobile cellular because we don't control.
Well, were you going to go and hand out, uh, at and t phones to people that can't get Verizon? I mean, it's insane, right? So are you willing to write off, I don't know how many subscribers they have, but let's say a hundred million on Verizon, a hundred million on at and t, a hundred million on T-Mobile, right?
So what you just block a hundred million users doesn't work Well. And like, like we tend to do across the board is the next thing, we'll solve all our problems. 5G was the miracle will change applications, and I can't even get 4G at my house.
So Something like, yeah, but then, but then you go to Europe and you can have like LTE on the subway and you can have a video called that a hundred feet on the ground without, uh, without any latency of some sort. The, the differences are amazing, are amazing. You know, one of the things is, I'm thinking about the a PM landscape of one of the measures of Yep.
Application performance and synthetic transactions, and is my application Sure. It, it, it is more than a monitoring. I mean, that was, that was kind of one era of monitoring, if you wanna think of it that way.
Yeah. How have things changed and how is it different? And how do you think about performance in the context of, uh, user experience KPIs or measurements of how can we really tell what the, the user's experience is, and are we delivering what we promised?
Yeah. So a long time ago, uh, when I was responsible for marching at double click, I started getting some alerts from this company called Gomez, which we used to use at the time that was giving us the, there you go. So that was, uh, my external signal, uh, system.
And, uh, I run into the knock, uh, double click, and I saw everybody really chilling, like, really cool. It's like, you guys, okay? It's like, yeah, we're good.
It's like, look, is everything okay? It's like, yeah, the network is up, the data centers are up, the databases are up. Everything was green except the end users were not getting ads.
And so for double click, I mean, it was ads, right? And so, so basically inside the firewall, inside 17 of our data centers, everything was perfect. Nothing was broken, everything was working as expected, except that from an end user perspective, it was broken.
So that, that first incident that happened for me in 99 was, if you want the, the biggest paradigm shift, and I literally, I turned off all the other monitoring systems. 'cause like, okay, if, if all these tools are telling us everything is okay, then they're giving us a false sense of security, then I don't need them, right? And, uh, so from there, I, I basically created this, this rule that said, okay, I'm going to think about performance.
I'm going to performance with a big P, not performance with speed, but I, I start looking at my, at my world in, in four categories. So reach, ability, availability, performance, and reliability. So can I get to you?
Can I get in the car and drive from my house to Costco, right? That's reachability. Can I get there?
Is the, uh, highway broken or not? Once I'm there, can I go inside the Costco store? That's availability, right?
So can, is the door open? Can I get in? That's availability.
There is performance, meaning that I am there. Did I find what I was looking for? I came for milk.
Did I find milk? Et cetera, et cetera. And was the checkout easy or did I wait two hours to check out because the cash resistor is broken, or we got, uh, the, I'm sorry, the, the, our computers are slow today.
So that's performance. And then the other one is reliability. And reliability was a, was a, was something I got to learn more when Google acquired double click and we got exposed to the whole SRE culture.
So reliability is like, are you able, uh, I remember this, it was amazing. I, I somebody from Google, uh, can you, can you share me with us? Your, your data?
And I shared the, the monitoring data of double click and said, this is a joke, right? It's like you, you're showing me a a a a time series, uh, that uses averages. And basically you're telling me that your performance is 200 millisecond.
That that's bs. He called, he called BS on me. He's like, I want to see your distribution.
I want to see how well you're performing at 1:00 AM, 2:00 AM 3:00 AM 4:00 AM and at 10:00 AM when there is peak, right? And so reliability was a new concept for me and for my team. It's like, okay, how, how, how can we deliver the consistent response availability, et cetera, 24 7.
So it's like, how well are you able to deliver the reliability? So those are my four. Uh, uh, if you want my cardinal rules, and I think from a monitoring a PM, whatever you want to call it, I think we need to be able to deliver that.
So, sorry for the long answer. So I just wanted to put that in context. So in today's world, I think it's extremely important to have, uh, different signals that are giving you the health of your business.
A PM is very passive, and it's, you still need that. I think it's extremely important to have also a more proactive approach to your monitoring strategy, to basically don't be caught your pants down where we are hearing about customers having a problem, right? That should not happen in 2024.
So you need to forgetting tools and names and brands and whatnot, you need different capabilities to basically paint that picture both passively, so a PM, open telemetry, observability, et cetera. But then you need a proactive approach that is constantly mystery shopping your environment, right? Can I resolve your DNS?
Can I, can I get your data centers? Can I, can I basically buy, uh, uh, uh, apples and oranges and, and add them to my cart and check out successfully? Because if you don't do that, and you're just waiting for, um, some baseline to tell you, Hey, we dropped the number of orders, that's too late, Mitch, you've already, you've already p****d off.
Excuse the French, you already p****d off a hundred thousand users and, uh, maybe lost a hundred million dollars in revenue. I mean, who wants that? And I had this conversation a few years ago with, uh, with this, uh, or former, uh, colleague, a former colleague of mine from DoubleClick who ended up at this company in Seattle.
And he said, you know, I don't need the proactive monitoring or synthetic monitoring. I have, I have an a PM tool. And he said, how is that working out for you?
He's like, well, good. We have a threshold. If you get a hundred thousand alerts, then, uh, a hundred, sorry, if, if our APIs has a hundred thousand um, errors, then we know it's a problem.
So I thought this was pre pandemic. So it's kind of interesting. It's like, so my analogy was like, if you're running a hospital, you're telling me you're going to wait for a hundred thousand people to die to say, um, Houston, maybe we have a problem.
Don't you want to find that out that the fifth person coming to your hospital, that there is a problem? And that's the proactive approach to monitoring. Um, so I wish people would be more, again, forgetting tools or this, that, but I think a proactive approach to monitoring is, is more important these days, especially because of the complexity of people accessing from all around the world, accessing on 5G, 4G, this, that, whatnot.
It's very important to, to proactively test things. It's a definitely a different world. You know, you mentioned, um, IT leaders being, being, having a vision capability, right?
Paying a vision of where we're going. Um, resilience is the topic now with many organizations. Yeah, probably multiple definitions, I think of resilience as, you know, the ability to, um, still be standing under unknown or unforeseen conditions as well as those we know we could handle.
And while you may not be able to handle everything, but you're able to weather much greater storms, uh, how, how do you, how do you advise, uh, your customers on thinking about resiliency as opposed to just uptime reliability, right? I say just uptime, reliability. Yeah, no, absolutely.
Yeah. So when you think about the world we all live in, is we live in an unknown, unknown. I mean, when you think about it, most IT organization know very well in 2024 to deal with the known situation, right?
They have, they've seen it before. They have a playbook, they have tools, they have all that stuff. The biggest challenge is the unknown.
Unknown, right? You don't know it's happening and you don't know how to respond to it. And that is the, the worst possible scenario, right?
That's what throw companies apart. And, um, and, and, and I don't think e even as humans, we, we don't deal very well with the unknown and no, right? That's the scary, uh, zone, right?
We, we all live. So I think from, uh, uh, what we tell folks when we get to have conversations with them about resiliency is like, listen, failures happen at Google. Failures happen at Amazon.
So you are not, I don't care what your name is, it's going to happen to you. So just admit it, just, it's not a question of what, uh, if it's a question of when and how bad if it'll happen, trust me, there is, uh, uh, it happens to everyone. So I think you need to build the resiliency, the redundancies to minimize that as much as possible.
I mean, for example, even a company as small as Catchpoint, we have three CDN com vendors, we have 3D NS providers. Why? Because we've seen the playbooks happen where one of them goes down, et cetera.
You don't want to be caught off guard. Um, I think resilience as well is making sure that you have the ability to speak the same language. And one of the biggest PBS I have in our industry is the fact that there is so much decent decentralization that has happened that you have, again, big companies, you have 5, 6, 7 teams working on products, infrastructure, et cetera.
They're all using different tools. And so they're all speaking the different languages. And when the crisis happened, what one thing I've noticed, and, you know, talking to other, uh, other companies in the, in the field, uh, there is, uh, the inability to agree that there is a problem.
I'm seeing a problem. No, the other tool is not showing any problem. And, and so there is this meantime to innocence or you people that buy tools just so they can absolve themselves.
This is the one that drives me crazy, to be honest, Mitch, right? It's just like, no, my tool is telling me everything is okay. You know, it's, it's not me.
It's the network guys, or it's, uh, it's the, the cloud guys or the whatever guys. And I think resiliency is not that, resiliency is not about blaming the other team in a company. Resiliency is about to able to recover as fast as possible and learn from it, right?
And I, that's why I love the whole SRE thing, but the true SRE, not just the name SRE for the sake of saying we, we do SRE, um, but, but resiliency is a state of mind in my opinion, right? I, it's, uh, and I don't think we do enough of it. And that's why I've been, you know, you've heard me at that, uh, at, uh, the event we were in last week.
I think it's time to have a chief resiliency officer in a company. The, the, the resiliency are somebody that can be that authority that can basically put teams together, unify some of the monitoring and observability data in one place, make sure that there is no more finger pointing. If that is one thing that is annoying me is the finger pointing that happens in companies, because the more people finger point, the less the, the less time you're spending on resolving and bringing the customer back online.
It's interesting, I think of the failure's not an option. Actually, failure's required. That's how we build resiliency, right?
Uh, you know, Mitch, the only reason, the only reason I'm sitting here in front of you today is like I took double click down in my career. Uh, I singlehandedly deleted the file that caused this chain reaction across all of our load balancers. And there was a file.
I said, I didn't know what it was. I just deleted it. Stupid rookies mistake, right?
And, uh, and, uh, I deleted that file and, uh, I was in my apartment in New York, and, uh, my boss calls me and said, did you do something? Like, yeah, I deleted that file. He's like, I don't know.
Did you do anything to the system? He's like, yeah, I purged a stupid file called one by one gift. I remember that.
And he said, well, you took us down. We've been chasing the problem for three hours. I said, maybe we'll call Maddy and see if he did something and that.
So I was asked to come the next day at the office. I said, that's it, Diego. You fire me.
I'm done. I, I am. That's my boss was a genius.
So that's why I said leadership matters, right? My boss was unbelievable. He sat me down, I was with my peers, and he said, what did you do?
Tell, tell us how, what went in your head? What did you do? Because we need to build resilience.
We need to build things around the next guy who's going to make the same mistake as you. And, um, so it starts with leadership. It starts by creating a, a safe space where mistakes are okay, failures are okay.
Uh, I've never fired anyone at Catchpoint on, in any department for making a mistake. But I will ask them, what have you learned? Like what, even in my interviews with, with the, uh, BDR and SDR R and SREs, like, tell me when was the last time you screwed up something and what did you learn from, like, what went through your learning process?
Because that's resilience, that's reliability. That's the ability to think about, okay, what am I going to make sure to never do again? And we don't do enough of that.
People are still scared. It has to be, uh, how has that error or that mistake helped us or helped you be better at what we do? Yeah.
Yeah. That's, that's the real value. Otherwise, you lost opportunity.
How do you, yeah. How do you get back on your horse, right? How do you get back on the horse?
How, what have you done? And I think we don't do enough of that. I've seen some amazing companies, LinkedIn is being one of them.
They talk about their, their failures and what they do about that. I think chaos engineering is important. I don't think we do enough of that anymore.
Uh, those, those kind of of events are super important to bring resiliency back into. In, in companies especially, again, you are as strong as your weakest link, right? So when you think about, I told you there are webpages or web applications that take six, 700 things to happen where you think all of them are always working.
And by the way, 90% of them, you're not in your control. I think about Apollo 13, failure's not an option. Yeah.
Actually, people dying in space is not an option. Failures, failure is required to solve. And Mitch, and the way I sometimes talk about this is forget for, for forget being able to deliver food or ordering your Uber, right?
That's when you think in the grand scheme of things, like you can survive, right? But imagine a hospital, imagine as you said, an airplane. Imagine things where lives are at risk.
Like a mistake is very costly, though. Look at the resiliency that is put in a hospital, right? There is not only one oxygen tank, there are 10, they're in a plane.
They're not only one, one system to guide the, the wheel or whatever they are. 10. So resiliency, resiliency.
So those guys think about it so much, and I think we need to bring that into companies that are delivering the customer experience because some customer experiences are, are as important because you fail as a company, you have to fire 10,000 engineers or 10,000 people because you didn't build the resiliency. That's a shame. Well, I think it's a resiliency is a great point to end on.
Thank you Maddie. It's great. Uh, pleasure talking with you.
Uh, folks who wanna know more about Catchpoint and some great resources that you all have, where can they go? I think go on our websites, read our blogs, uh, connect with us on LinkedIn. Uh, we're all, uh, love to share ideas and learn from you.
And, uh, maybe we can also add a few things to your arsenal. Oh, we have also great resources. The site reliability, uh, survey we do every year.
We've been doing this for six years. Some amazing content and some amazing, uh, feedback that we've gotten from the SRE community. So check that out.
We need to have you back on to talk about that Anytime. Pretty Important topic. Alright.
Thanks again. Thank you so much, Mitch. Thank you.