Causal Reasoning Software with Causely’s Francis Cordon
Francis talks about a new approach to incident response – causal reasoning software – which has the potential to transform the way DevOps teams operate, helping them filter signal from noise, automatically remediate issues and be more proactive.
Transcript
This is Textron tv. Hi everyone. Alan Chival here for Textron tv.
I've got a first time person on Textron with me today. His name is Francis Cordone. Francis is the chief customer officer for a company named Causley and don't believe we featured them before either.
So let's welcome him and, and find out more about, uh, Francis and Causley. Hi Francis, how are you? I am good, Alan.
Thank you for having me. How are you today? Very well, thank you.
So Francis, let's start with you. Tell us a little bit about you. Excellent.
Thank you. So, yes, um, it started in the nineties really with the pioneers of application performance management. Dating myself here a little bit, but I think that Okay.
Within a small company called Merger Interactive that ended up creating a lot of the concepts we use today and was bought many years later. By hp. By Hp, exactly, yes.
I remember I was, that's when I started in the, I was an engineer by, by training. And then that's when I started working in the performance management and monitoring space and then moved to United States with, I am originally from Spain. Uh mm-Hmm.
Moved to United States with a company called Compuware. And then I moved to Dynatrace where it was very, very small. That was there about seven years building this company that today is so well known.
And then I did something that changed everything in my career after so many years being a vendor, thinking I understood customers, I became a customer. I went on to run application reliability for the Bank of New York Mellon, a centralized team for 12 lines of business. And Alan, that was incredible.
O opened my mind. First week, two hour outage, 200 million penalty. Talk about PagerDuty and the cost of outages.
I think we're talking about a timeless problem that only keeps getting worse because of the complexity of these dependencies. And then I went on to become a vendor again because I had a passion in my heart to do it well, to take a lot those learnings that I had doing it in high stakes environment. So costly is, is from the ground up built to do this.
And I couldn't be happier to be in an environment that allows me to use both my vendor and my customer experience to bring that to the table. Excellent. Let's talk about Coley then.
Give us, give us the, uh, scoop there. Yeah, very good. So if you think about the history, I love to go, I don't know all maybe with age it happens, but it starts like everything is related to history.
I like explaining things by their history. Okay. And if you think about the phase we went through in the world of obligation, performance management and then observability, it, it was for so many years about further levels of visibility.
Give me more data, give me more data. Oh, we don't have the ability to instrument Java. Give me Java.
We don't have the ability to instrument that net, give it the net and of course network and server. And now containers constantly takes a step back and says, alright, with so much data, the problem we're having now is the fatigue with so much data, too many dashboards. But who is telling you when something happens and you have 300 things going on, who is telling you which one is the root cause and which are symptoms?
And if you think about it, the model, even with sophisticated solutions, and as the, uh, owner of resiliency and reliability at the Bank of New York, I had to spend millions in vendors. And they were good. They did their work, but they didn't do this single thing.
Differentiate root causes from symptoms, the cause and effect relationship wasn't there. So costly pills from the ground up to use data that more often than not is already there. And let's give that to the software to tell you root cause analysis and hopefully later we'll talk about that's not only reactive, that can also be proactive.
It has about three levels of maturity, but essentially that's costly and built from the ground up to do this with what we call our causal reasoning platform. Alright, I love that. Now, so let, let's dive in a, you know, let's dive in a little deeper there, Francis.
First of all, when we talk about the rising costs of incident response, there's really two pieces of it. Number one is the cost of incidents, period. As you mentioned when you were at Bank, New York, Maryland, $200 million, it's a fair amount of money, right?
It's A lot Of money for outages. So incidents, downtime, cost, businesses, money. So minimizing downtime is a huge priority.
But the second piece of it is a lot of companies say, okay, we need incident response software that allows us to minimize this. But it's not just buying software or signing up for a SaaS service. It's truly the implementation of that, of of using it, adopting it, living it, breathing with it that you then find out can cost a lot of money and not necessarily solve your problem.
Exactly right. So when we talk about causally helping with the rising cost of incident response while always causal reasoning software and you know, as defined by causley solving the rising cost of incident response, what costs are we talking about? Excellent.
I love how you framed that. So if you don't mind, I'll tackle maybe two or three key things you've mentioned there and developed a little bit in the context of costly and how it can help the cost. As the PagerDuty article refers to the cost, and to be honest, the, they mention, uh, a number of like 175 minutes on average.
We actually, in conversations know in many cases it's a lot more than that. Uh, I'm with multiple people involved. But the cost is something that even organizations struggle with because more often than not that cost is the operational expense of the people that will be involved.
But there is a lot more cost in the part that's difficult to calculate. What about the customers right there, right then that we're having a poor experience and they may jump out of your platform and go somewhere else. That cost is a lot more serious that the, the damage to the brand.
Uh, and then what we are miss in this too is engineers that now get involved systematically to do this because of the decentralization of the reliability practice and the complexity. So engineers jump on this, are also not doing their innovation work creating. And that's ultimately what puts a company ahead of the competitors that they're innovating.
Investors, especially in this time. The cost is a lot more than the number of hours or minutes, but what's going on with that? This is also why a costly, I like to partner with customers to establish the language of value, which is a lot related to what you mentioned of how is it implemented, right?
How do we talk about these things and how do we integrate in the processes? At the end of the day, what we're trying to do is let people do their work, innovate, create faster cycles of releases that put a comp, you can put a company ahead of the competition. That's the most important.
Not to mention attrition of employees. The engineers are called to be engineers 'cause they like to create not troubleshoot. And that's a whole shift.
However, the model today remains give data to experts to do troubleshooting. And maybe instead of three hours, they take an hour and a half. But why do we have to give it to experts that could be doing creativity or innovation?
So that's your first point. But then here's something very important to, to your second point, even innovation. 'cause if you read the PagerDuty, it's almost the impact of not having automation.
They keep referring to the cost of manual troubleshooting, manual, manual. But I can see people reading that and jumping too soon to oh yeah, we need innovation. But my question I ask almost as a challenge for everyone out there to think about, to learn from a best practices standpoint, how can you automate when you don't know, which is a root cause versus a symptom?
In other words, in any complex environment today, you're going to have hundreds of symptoms when something serious happens because of the ripple effect that the services dependency, the components and innovation, which is also something that has to be done for the standpoint of integrating with systems, uh, systems, et cetera, is actually not what we're looking for until we know which of these created the problem. Because if not, we're looking at like taking an aspirin when I have a headache without understanding if I have a headache every day, maybe there's a bigger reason and I should see a doctor about it. It may be a simple example, but when I was at the bank, this was a daily life, like just restarted.
We're so panicky about that because we couldn't determine without that tremendous spend of money and time, what was the root cost. So yes to the cost of manual, but to do it well, what we need first is to think costly, to understand what's the root cost. And everyone today they can open an LLM on their phone.
We have access to these incredible technologies. They're very good at correlation. They're not that good at causality.
This is something technology is barely now starting to do. Well, we are so good at having petabytes of data, but at best to establish correlation between them. And so I love the PagerDuty article because it, it gives a visibility into I must if I am running a, uh, IT operations and organization, I know I must automate instead of relying on manual troubleshooting.
But to do that, well you have to have causality because that's what's going to tell you when you automate this, then it goes ahead and prevents those issues and fixes the problem. And you talk about implementation. If I can throw one more then my, my passion is to work with organizations to say, when you've done done that, which is already ahead of the curve, who is doing causality in their current environment for current fires, to eliminate them at the root cause.
And then you put automation, but let's take it further Now if you can see that, or if you have a system like costly that can tell you that relationship of root concept and effect, even for issues that haven't happened yet, then couldn't you start building resiliently and that's the implementation you're referring to. After we work a little on those fires, let's move it left. Let's integrate into the workflows, into the processes, and let's start building things that don't break to begin with.
So we don't accept that broken model in it. Operations of things have to break and hopefully we're in a race to not take too long to fix them. That wouldn't work with my car.
I don't wanna accept a model which the, the KPI is that they, the mechanics take little time in fixing it. I don't want the car to break. And that's the journey we're on with costly and in partnership with our customers.
And I would argue an organization out there that's thinking about this maturely, let's learn to build resiliently so things don't break to begin with. And of course it's complex in in the current environment. So, And, and I don't, let me run this by and you tell me, you know, today we're all looking at how could ai, right, how can AI help us go faster, cheaper, better?
Um, and AI is very good at correlation, right? Whether we're talking about an ML ops, AI ops kind of thing, or even generator of AI and LLMs, it's, it's very good at correlation by taking, you know, big data sets and boiling that down is partial reasoning beyond today's ai. Yeah, excellent question.
It is a different take definitely because LLMs or gene AI isn't really, uh, it's forte isn't really thinking cost in, in a, in a cause and effect relationship cost. No. So I'll tell you a little bit about it and in general, and then also as it relate, relates to what approach costly had.
But also with ai, I have the feeling that we're leading this explosion of methodologies that in truth will be many, many, many, many, many tools AI's almost like many years ago, say agile. It's just an umbrella of many practices. Mm-Hmm.
At the end of the day, people care about the getting the job done and what technology we use only now because it's such a, a novelty. But in reality what we need is to get it done, get the job done e effectively. And for that we may have to combine multiple tools.
So thinking costly. And there is a scientific, uh, branch called, uh, causal ai. And it applies to medicine, it applies to it, it, it has and it thinks a little different than LLMs or gene ai.
What it is, is causal models and it has to capture knowledge that has existed. If you think about message brokers, if you think about, uh, HDP calls, RPC calls, Java obligations, all Kubernetes, all not These things have a footprint. These things have a way of behaving in relationship to each other and it gets beyond human scale because of the dependencies.
But this can be expressed in SAL language. And when we put that together, we call that a SAL model. And when we instantiate that with the topology of an organization and maintain it in real time, which is obviously very difficult to do, and that's the approach we take, then we have effectively a model that's constantly applying causality to your environment as it's alive.
And that would not work with LLM, but LLM would be really good to help you as a side, um, assistant in other things. For example, summarize for me what happened in the incident yesterday that would be perfect. Learn from my logs how many times that would perfect.
But not to tell your causality when it's happening. And preventively for that. We need a different technology in the greater sense of the world way of doing things.
And that's what we do with our causal reasoning is causal models and environments that are instantiated and maintained. Excellent. You know, we're almost outta time Francis.
I didn't even ask you for people who wanna find out more about causally. What's the website? io.
Dot io. Very good. Very good.
And yes, I encourage people to go there. We're actually really transparent there. People can see the description of essentially how we operate this description of our causal models and causal reasoning.
And also people can there see the platform in action through videos and a sound guided tour and then contact us to learn more. So I encourage people to go there and check it out. Excellent.
Francis, thanks for coming on and, and, uh, educating us here a little bit. Success. Wishing you good success with Causley and keep us posted.
Alan, I'm such a big fan of your work, so I was so honored to be here. Thank you so very much for having It's My pleasure. France is Gordon, chief Customer Officer at Causley here on Textron tv.
That's causley io. We're gonna take a break. We have a lot more here on Textron TV today.
Stay tuned.