Techstrong TV February 3, 2026
Watch our live stream Monday through Friday, featuring exclusive news, announcements and conversations with IT leaders and experts on topics ranging from digital transformation to #DevOps, #Cybersecurity, #CloudNative, #Containers and deep-dives into specific technologies and best practices. http://techstrong.tv/
Transcript
Hey, everyone. Welcome back here to text Trump tv. You know, I, I, if I don't have this gentleman on every two, three months, I start wondering what's going on.
So it's probably been about three months. Uh, I want to introduce you too. Well, you may know him already, but if not, meet Colton.
Andrews Colton is the founder and one time CEO. And now CEO Again, ed Gremlin. Hey, Colton, it's good to see you, my friend.
Happy New Year. Well, it's a little late, but I hope the New Year's off to a good start. How have you been?
Yeah, always a pleasure to chat with you, Alan. Thanks for having me on. Things have been going well.
Always excited to come share the latest and the greatest or dive into the details on the, you know, on the technical side. Very cool. Col, and I see you're the founder.
You were CEO then you weren't CEO for a while, then you came back as CEO. Um, but beyond that, prior to Grambling, you were at Netflix and, um, you know, that this is when Netflix was not that they don't innovate anymore. They buy things now, obviously, huh.
But, um, but you know, they were really an innovative technology company. They kind of pioneered the whole thing around like chaos, engineering and talking about scalability, right? They kinda wrote the book, a lot of cloud native kind of, uh, functionality and projects were spawned out of Netflix.
It, it was a great time to be there. Yeah, no, I'm, I'm really grateful. I mean, uh, for those that, dunno, I got to do a, a four year stint at Amazon focused on keeping the retail website up and available.
We built fault injection, chaos engineering tools there had a lot of success. But I remember being at a, a, a Velocity conference, I think in 20 12, 20 13, and, uh, hearing, uh, these companies talk about hearing Netflix talk about what they were doing on the resilience, reliability space, and Chaos Monkey had just come out and, uh, thought that was neat. I thought they could have done more there, but, but it was, it was a good first step.
But they did a great job promoting it and helping people understand why it was important and really driving that innovation. So I was excited to go there. I showed up, you know, there was room to pitch in and help out.
I helped us get another nine of availability, helped us build some grade tooling, and really that's what allowed me to have the opportunity to go found Gremlin. Um, I was giving a talk at a conference and I'm running into some VCs in the lobby, and we went back and forth and they said, and I was like, I'm gonna bootstrap, you know, I'll just wait it out. And they were like, hold on, you could start tomorrow.
Let's get going. Uh, and three months later we founded Gremlin. And coincidentally, uh, as of Sunday, it's been 10 years now, we've been out in market Another overnight success.
Yeah. Yeah. Not, not what I read on Twitter back in the day.
No. Well, that, you know, that's the deep dark underbelly of the whole thing. All these people who think, oh, yeah, startup guy, you know, overnight success.
No, this 10 plus years, it's not unusual to have that kind of effort in there. And you, you know, and you, but you're still a startup and still building and still learning and, and, and doing all that. Um, now Gremlin's Mission has, I don't know if you wanna say expanded, changed, evolved.
Chaos engineering is still obviously part of it, but chaos engineering now has, you know, it's, it's the lifecycle of a product to a feature in tech, right? Uh, today's product become tomorrow's features as you, you know, March in inextricably towards a platform and, and all of that good stuff. So what, what's rambling about today, Carlton?
Yeah, I mean, I think that's, it's such a great point. Chaos, engineering, cool idea, fun project, not really an enterprise discipline that drives reliability. And that's what we learned early days is a lot of folks wanted to like, you know, kind of have fun, we'll break some stuff, we'll see what happens, we'll order some pizzas, you know, we're doing good things.
And a lot of that didn't really result in the type of outcomes that they wanted because it was isolated. It was not repeatable. And to your point about things becoming commoditized, things becoming part of a platform, things really becoming just how we build software.
That's what we've seen a lot of growth in in the last few years, is people moving to, Hey, every team needs to do these basic tests. It's just, you know, it's unit and, and integration testing for our distributed systems. And when we do it, we find things on a regular basis.
We fix them before their issues, and we're just building better high quality software. So, absolutely. What, what we've got coming out right now is just along the same lines, uh, you know, having grown up in a lot of enterprises, a thing that almost every enterprise company does is some form of disaster recovery testing.
Hey, what happens if we lose a data center? Hey, what happens if a WSU East U US East one goes down, or GCP loses a region or a zone? How does my software, you know, behave?
Can I keep operating? And a lot of companies do this in a pretty manual process. Hundreds of engineers, a whole bunch of prep.
They're running it once or twice a year, and they, they, you know, they've gotta, they've gotta invest a lot of time. Then some companies, this is a weekend where everybody's on call and you've gotta just show up and be on that call just in case something goes wrong with your software while we execute this large event. So we thought, you know, this is a place that we could do better.
Uh, and what we did is we took Gremlin, which is already good. Our customers are already using Gremlin for this purpose. Some of the largest banks are using Gremlin to go do this kind of dedicated disaster recovery testing.
But we said, well, let's make it easy to do the right thing. So we built, uh, built into the product the way to model this large scale event this way to have the right safe preconditions. Let's make sure everyone's run it on their own first.
Let's make sure everything looks good. Do we wanna run it, uh, one big event, or do we wanna spread it out over a couple of weeks and let people, you know, prove their piece independently? Um, let's make sure we've got that halt button if things go wrong, you know, a way to clean it up and revert it so that we've got that safety net while we're running the experiment.
And I think part of how we've grown up is, the last piece is if you can't measure it and you can't turn back and present to the business or the auditors or the compliance folks, the evidence that you've successfully completed this, then it's just not as valuable. And so we spent a lot of time building really all of the, the right reporting and detail so that you can run this event, you can run it more often, uh, because you don't need as many people involved and you can streamline it. But then you also have everything you need on the back end to go hand to the business to say, we've done our due diligence.
You know, Colton, I think back to my days, you know, running companies that were, well even here, ra right? You know, disaster recovery plan. We're, we're totally sas.
We, we have none of our own infrastructure per se. So disaster recovery you plan is basically getting on the phone and begging, right? Um, um, I'm being facetious, of course, but you know, in other companies that I was involved in, we, we did have much more dr right?
We were running data centers and stuff like that. Um, I wish we had something like this back then. 'cause as you said, it was quote unquote a manual process.
And, you know, living here in South Florida, you would think everyone would have, what do you do in a hurricane as part of their DR testing? But you'd surprised, you'd be surprised, you know, oh, I didn't have that on my Bingo card. Well, you're in south Florida, you didn't think a hurricane might knock you out.
You might be flooded. You might not have power for a few days. What do, like, you know, where was the thought process there?
And it's, you know, so when we, Manuel is a, is a, is a good term, Colton, for covering up, just, you know, you can't cure stupid. And, uh, there was a lot of, I've seen a lot of stupid over the years down here with, with people, you know, who thought they had their DR plans in, in, in, uh, place until, until the stuff hit the fan. Yeah.
Um, Such a great analogy. But It is. But, you know, so having something like this, you know, this, this is like a, you know, sleeping under the blanket of security, right?
That, that kinda keeps you, you, you feel like you actually really do have it. It would also seem to me, Colton, that this might be something that maybe AI can help with in terms of, of, uh, you know, setting up, what, what are the test parameters? What is your test coverage here?
What should you be testing, what, you know, like, we probably don't test a lot down here for snow, right? I, well, who knows, but how are you using AI or anything here to, to help with that? Or is this really based on real world experience that you're now, you know, being able to scale up to multiple or infinite amount of, of customers?
Yeah, so part of what we have in just built into the Gremlin product is a set of recommendations where if things aren't going right, then we're gonna tell you how you can go fix them, how to go adjust them. Um, and one of the things we're always tuning and improving is just how do we make it easy for people to do the right thing? And that's really tell 'em what they should be doing.
I think this is one of the things I've had to learn in my career early as an engineer, come from Amazon and Netflix. Just assume everybody's done this a hundred times. They know what they're doing, just give 'em the tool and get out of the way.
But truthfully, there's a lot of folks that maybe haven't done this, or maybe they're, you know, they've got some junior folks on their team and they really need a little bit of guidance on the structure, on how to set it up, on how to model it, on how to run it, on how to interpret it. Uh, so we've got some of that. We're always improving that, uh, within the product.
So I'm not gonna say it's AI driven. I don't want to be sensationalists there, but we've built in a lot of recommendations that tell people how to fix what goes wrong and what they should be doing. Absolutely.
Colton, um, you know, thi this became available on February 3rd. It's in ga, February A as of the third then, or? Yeah.
Okay. Yeah. I told my team we've had it ready for a while and we've had it in beta with some of our large customers.
And one of my quality bars is have we run this ourselves in production? And the answer is yes, we have, we ran it and we learned some things. We passed the test, but we found a couple things we didn't like that we went and fixed, which is exactly why you run these exercises to uncover those things.
Absolutely. Go fix 'em. So that was why, that's why it's ready to launch.
'cause I said, engineering team, when you've run it and you feel good about it, that's when we're ready to go take it, then we're Ready to go. You know? And that, and that, right there was the beauty of of Chaos Monkey too, right?
Because it's like a live fire drill. Mm-hmm. Do you know what I mean?
Until you actually go through it, it's hard to anticipate everything. You, you've gotta kind of have been there, done that kind of thing. And that's what Gremlin brings to this.
So it's an excellent thing for people who wanna go check it out. Colton, what's their best kind of on-ramp for this? com.
We've got links to, you know, we'll have it, we have it up on the main page. We can tell you details about it. We've got a free trial.
You can go in and play with it yourself. And of course you can reach out to me or my team. We love to not just show people the product, but guide people in how to go model these exercises.
How to run them safely and effectively and partner with them to make sure they're building not just, uh, a chaos engineering experiment, but a reliability program for their company. Absolutely. You know, too many people, Colton, treat VR as a checkbox, you know, cyber insurance company asked, do you have a DR plan in place?
Yes. Have you tested it? Yes.
You know, and that's far As goes often is the tabletop exercise, or, you know, maybe once Yeah, we did a year ago Or five, you know, there, there is some of that too. 'cause not everyone is in Amazon or a Netflix. There's a lot of, you know, small, medium businesses who have a, a data center or some colo rack space.
Well, as you mentioned, if you're hosted in the cloud, as, as a lot of us are, uh, look, it was kind of a rough end of 2025 in the cloud. We saw Yeah. Some Major Outage with basically every cloud provider.
So yeah, the truth is, this isn't like wishful thinking. This is something that happens regularly and you can turn it Into a non event. It's not if it's when Yeah, it's not.
If it's when I get you, Colton, thanks for coming on here and telling us about this. What do you call it? Is it called Gremlin?
Disaster recovery testing? Yeah, gremlin disaster recovery testing. We, we went back forth on a name.
We had another internal name, and we said, look, this is what people run. This is what people call it. Let's descriptive, let's just call it what it is.
I love it. I love it, man. Colton, good luck with this.
I think it's a great product. You know, I've, as I said, I firsthand knowledge on this kind of stuff of where it comes back to bite you in the butt, man. It bites heart.
You know, it'd be a fool not to do it. Um, keep up the great work. Keep us posted.
You know, the clock's running now, right? It's end of January or beginning of February, actually. Um, we'll hopefully see you back here by April.
Yeah, That's the plan. I love it. All right.
Hey, Colton, be well. Colton Andrews founder Gremlin here on Tech Trunk tv. We're gonna take a break.
We'll be back. Hi everybody. Thank you for joining us on this series about AI in the mainframe environment.
My name is Mitch Ashley, and I lead the software lifecycle engineering practice with the Futurum Group. Now, this is part of a three part series, segment number two, where we're gonna be discussing infusing intelligence using AI as a partner in in mainframe environment with mainframe teams. Today I'm joined by Anthony Starro.
Anthony is senior director of, uh, architecture for AI at BMC software. Great to be talking with you, Anthony. Thanks.
Pitch. Thanks for having me. You know, when organizations feel they're ready, kind of take that next step.
How does generative AI start to make an impact in the mainframe environment? Yeah, so, you know, we start looking at AI as a partner. Yeah, I was a developer for 30 plus years.
Boy, I would've loved to have AI at some point in my career. Why? Because from a developer's perspective, generative AI would've taken a lot of the toil out of my day to day activities.
I could have used AI in a way where just the grunt work I had to do day in and day out during my developer developer journey. AI could have been a tremendous help in, in that regard, take the friction out outta my day. When it comes to the AI ops space, as an example, I spent a lot of years building, uh, data visualization solutions, dashboards, et cetera.
AI as a partner in that journey is how many times are we looking at operational dashboards and we're seeing blinking lights, or maybe we're not seeing certain things that may be in graphs or a data grid that's being shown. And, you know, we're, we're, we're looking at it. It would be nice if you had AI sitting there as your partner observing the same thing you are, and then pointing things out to you, uh, that you would otherwise miss.
So I really liked that portion of generative AI making an impact day in and day out. This is kind of tied to what we talked about previous, previously. If I'm a next generation mainframer and I'm new to the mainframe space, having AI as a partner in my daily journey, whether that's infusion and products that I'm using, or that's knowledge base access, as we discussed, all of that constitutes that generative AI bubble to help me start my journey on that mainframe space.
So that's really where I see it. Helping folks on the mainframe accelerate their day by taking things off their plate that they shouldn't be worried about, and focus on the more important high value items and innovation that they should be doing day in and day out. You know, going from AI as being ready as an advisor now into AI as a partner, it seems like we're using ai, we're working alongside us as we're performing work, and it's providing information and, you know, giving us insights.
Maybe that, as you mentioned that we didn't have. I'm curious if you have some kinda real world examples. Regenerative AI is helping us.
We do. I, I I I, I, I mentioned I talk to a lot of customers, and one of the things that comes up all the time is call ball, uh, code explain. They're looking at, you know, they got that next generation looking at code bases, that a 25, 30 year plus.
And there's a lot of complexity. There's code comments in there that are outta date. There are module comments that are there in there that are outta date.
Real world example, we've delivered through VMC. Any assistant that capability to do that code explain to do that. Document generation customers are using that.
They've given us really good feedback on the usage with that. That's a high value return. Another area that we focused on was our AI ops space, where BMC Amy assistant was able to go in there and, and, and when there was a, a problem in the environment, what's the root cause of this problem?
We have a lot of very sophisticated machine learning models and information that's presented to the user in that product experience. When generative AI came onto the scene, it was a great opportunity for us is we were able now to take something that was very complex, to explain and describe in a user experience, and have B-M-C-A-E assistant come in and then just explain what the problem is, what the root cause was in plain terms, not only that B-M-C-A-E assistant that was also able to give next step recommendations on how to resolve or prevent that problem from happening again. So these are real world problems that we've gotten feedback on, on how the AI truly helped elevate the business value that we're delivering out of our solutions.
You've talked about AI as a partner, giving us insights, helping us along, maybe triaging or diagnosing what issues might be. There's so much knowledge has gained in that process, but we're losing that knowledge with a lot of our workforce as they retire, move on, et cetera. Talk about that knowledge loss and how a AI can help us with that.
Yeah, so the knowledge loss is real as, as we all know. And, you know, an AI system is only as good as its knowledge. And what I mean by that is not the LLM knowledge, it's your institutional knowledge.
It's the knowledge that your staff carries day in and day out. How do we capture that? How do we infuse that into an AI system such that the AI system has more context and relevance and becomes smarter for the users using that?
That is critical. So how do we capture that institutional knowledge? We, we, we augment, or we have a facility within B-M-C-A-M-E assistant to capture that information and infuse it once it's infused, that's where the real power is.
And it's not just a chat conversation where you're gonna go to BMC AMY assistant and you're gonna ask questions and you get responses back based on that captured knowledge, right? We talked about that in the previous video that was around the advisor model. But, but imagine capturing knowledge and infusing in a, in a product experience, maybe it's in the AI ops space, capturing that knowledge so that when events happen and things happen within a dashboard or a report gets generated and there's alerts and exceptions and alarms that were generated, instead of just giving these generic type of, uh, insights out of the product, we can customize those responses and put them in terms that are relevant for that organization.
And the way we could do that is by capturing that tribal knowledge. I have to repeat that. I can't say tribal knowledge is capturing that enterprise knowledge that your senior staff has, and based on, uh, on processes and workflows they have when certain situations arise, we can infuse that in so the AI can respond in a way that, that, that is relevant to your organization, for your organization to take the right steps that's relevant for you.
That's what capturing institutional knowledge within an AI experience to give you back responses in that regard. So as we move along the journey that you've talked about, you know, adapting, adapting and using ai, going from the stage of giving advice as we're working to actually guiding work, um, I don't know that we want to turn it all over to the AI all at once, right? There's a step along the way, maybe a hybrid kind of AI talk about what that approach might look like.
Yeah, that, that's a, that's a really good question. So at this point, we've been talking about language models from the AI perspective and augmenting that with institutional knowledge or real time data as an example. But generative AI itself is not enough, right?
And when, when you, when you're looking at, um, infusing AI in, in, in your solutions and delivering that value to our customers, and a lot of times it's a, it, it's a, it's a, a, a collection of different AI techniques that we need to pull together to formulate the, the answer or the response that we want to give. And it's this combining of different AI techniques together to formulate that response. That's what we call hybrid ai.
So as an example, I mentioned the BMC, uh, a e ops insight product. There's a lot of sophisticated machine learning models driving that product, that experience. But we augmented that with generative a AI for the explainability and the next step recommendations.
So the, it's, it's the combination use of machine learning models with generative AI language models working them together. That's a great example of a hybrid AI solution. So you can combine things like, you know, classic rules-based AI with language models or rules-based AI with, with machine learning models.
What, what, whatever the application may be. It's this combination together. That's the way we talk about it, is in the terms of hybrid ai.
And that's a key element on our next video on how our AI agents think along their journey. Very good. That'll be a nice incentive to watch that third part, I'm sure.
Yes. Well, before we get there, any, any final advice you have for organizations as they're beginning to integrate generative ai? So when it comes to generative ai, um, again, you, you gotta be really practical on your use cases that you want to apply with, uh, generative ai.
Start with the explainability, no matter what your area is or, uh, your domain is that you wanna, uh, uh, apply ai, start with generative AI with explainability. Once you get that sorted out, move on to that next step. That next step is when you explain the situation, what's happening here?
How do we do the next step? What's my next step actions, whether that's to resolve a problem or the next step action could be very, could be simply as notifying someone that a certain situation is going on. But take those journey, that type of journey with your ai, do the, do the explain part, and then move to recommendations.
The next step with, with generative ai and how you're gonna do that is, especially with the recommend part, that's where you capture that institutional knowledge that you have with your senior staff that you've infused into your AI system. It could provide all those insights to the ai, ai, AI system on guiding it towards tho that output on those next steps. Well, thank you very much, Anthony.
It's great for you to share, uh, your experiences of your journey that you can help others along the way in that process. So you, you've joined us for our second segment in this three part series talking about infusing intelligence AI as a partner, generative AI in the main mainframe environment. We're sure happy that you've joined us, and we hope that you'll stick around and, uh, check out third segment.
We'll kind of give you a little hint there of some of the things that we're gonna be talking about. It's about acting with confidence with we have AI as an agent of change that's making things happen for us and with us in this environment. Thanks again for joining us with you on the next segment.
Hey everyone, it's Alan Shimel, founder, CEO here at Techstrong, and welcome to our continuing series on, uh, AI agentic AI in the future here with, uh, the Microsoft team and our RUM analyst team, as well as Techron. In this next episode, though, we're gonna be joined by Mitchell Ashley, uh, of Rum, who leads the software development lifecycle and building segment at futurum. And Mitchell is talking with Brian Good, whose official titles is corporate Vice President and agents marketing.
But Brian is really here talking about agent apps and chat, and it, it's, uh, you know, obviously very hot topic as we move to a ent ai workflow based basis. So let's join Mitchell and Brian here with me, and it's great to have them both. My name is Mitch Ashley, and I'm VP of practice, lead of the software lifecycle engineering practice at Futureum Research.
And, and Mitch, uh, my name is Brian. Good. Uh, I lead the business applications and agents team, uh, here at Microsoft.
Uh, let's start here. 2025 is described as kinda an inflection point for AI adoption. What do you think are the most significant changes that, uh, you've seen in organizations as they're using AI and agents in their businesses this year?
Well, I absolutely agree. 2025 is really an inflection point, and we'll sometimes describe it as the year that the Frontier Firm was born. And you might've heard us talk about the Frontier Firm before.
It's this idea of companies that are putting AI really at the heart of their business, and it's enabling them to do things like reinvent the way they engage with the customers or transform their business processes inside their company, um, and beyond. And it's really the year these frontier firms are sort of rising up and at time when I think we can learn a lot from these early adopters and understanding how they're deploying ai, how they're being successful, and then figure out how we can take those insights and bring 'em to, to our own businesses. That's where you can have outsized impact As AI transforms those functions.
What do you see as the new patterns of work that's enabled by copilot and agents as customers start leveraging the technology for productivity or innovation? I'll tell you what, I have studied these frontier firms as they, as they come up, and there's really three patterns that I see across these frontier firms. The first is really enabling, uh, employee productivity.
So they give every employee, uh, an AI assistant like Microsoft 365 copilot, and it helps them, those employees be more productive. The second pattern that I see is, uh, really these frontier firms deploying AI to automate existing business processes. So, for example, they already have a way of handling expense reports, but they can use AI to speed that up and reduce costs, and, and that certainly results in some benefits to the customer.
The third pattern is really where a customer, like starts from the beginning, let's say from first principles, and they reimagine a function altogether. They'll reimagine what it means to engage with a customer who has a, uh, an issue with their product, and they'll put agents at the heart of that. Most companies can only take on one or two of these functional transformation projects at, at any given time because it's a big lift.
Again, it's re-imagining a function from first principles. It's not just taking existing processes and and applying AI to them. How far along do you think most organizations are in their adoption cycle for generative ai?
Well, I think we're still in early innings, uh, that there's no doubt about that. Um, you know, customers are starting to deploy AI in different functions as we talked about, but certainly they haven't, in most cases, fully realized kind of the transformation that AI can have across their organization. So it's still very early innings, but I'd say we see very promising signs and, and green shoots, uh, of that, uh, of that adoption and success.
Uh, again, the key really here is to start function by function, think about a particular function, think about peeling it back to the business processes you need to go focus on. That's where you can have outsized impact. How do you define AgTech business applications, and why are they critical for organizations?
Yeah. Well, I do have a vision that there's, uh, the application of old is really transformed with ai, and we call that new type of application, the AG agentic business application. And the ag agentic biz app includes the assistant for the human to start to use.
It includes prebuilt business process agents that basically take the drudgery out of work, and then it's built on a data foundation. And so instead of being limited just to the data that, you know, maybe is in the CRM system or the ERP system, it joins that with data from other lines of business systems and even productivity data so that you have a rich set of data that you can build agents on top of, and you can empower your humans to get better decisions from. And so this idea that brings all those together, the assistant, the agent, the application, I call the Agentic Biz app, and I think it's a really big idea.
Let's talk about why you're optimistic about the future of agentic business applications. I have a lot of optimism here because number one, I see customers already deploying these applications, uh, and and really transforming their business. So I already see people starting to get great benefit from it, but at a more human level, the thing that gets me excited is like, if you think about every one of our jobs, like there's a lot of stuff that we end up having to do that really isn't adding joy to our lives.
Uh, you know, me doing an expense report or, you know, filling out a CRM uh, system, you know, with an update from a customer call, those aren't the things that, uh, that make us, uh, unique. Those aren't the things that bring joy. Those aren't really even things that add business value collectively.
But if we can delegate those things to AI and agents, I think we'll enable ourselves to actually go do far bigger things and really take our ambition to the next level. And so I am absolutely excited for what the future holds. Yeah, I'd love to hear about how you see the role of AI agents evolving just over the next few years, especially as organizations start to redesign core business processes drive outcomes for efficiency.
How, how do you see this taking shape? Yeah, great question. You know, odd, I I would say that we really believe there's a spectrum, uh, of, of agents.
You know, there will be very simple agents that someone will use that might just be grounded on a particular knowledge source and help you, you know, understand, you know, what, what might happen from that, uh, from that knowledge source. Then there'll be task-based agents, and then they'll also be more sophisticated, fully autonomous agents as well. And this more autonomous agents is where you can unlock a lot of business value.
These are agents that, you know, aren't called by a by a human. They're just running independently. They're triggered based on different actions, and they can really automate business processes in some cases from start to finish.
So we believe there's this spectrum of agents, and today, I'd say most ENT use cases start with those more simple kind of knowledge agents. Uh, but increasingly we're seeing customers bring in these task-based agents and autonomous agents, uh, to automate, uh, things, uh, end to end. How about the C-suite leaders?
How should they be thinking about investing in agent technologies for long-term value? Well, every C-suite leader I talk to today is already in on ai. Like every one of them recognizes that it's a competitive advantage if they can move quickly, and if they don't move quickly, they recognize it could be a disruptive force in their industry.
But let's face it, ai it's transforming businesses, it's transforming functions. It's, it's gonna reshape entire industries. And every C-suite leader I talk to recognizes that and once on board, what they're looking for though, is a partner.
They're looking for the tech, of course, they wanna make sure they've got the right tech, but they're also looking for a partner to help them shape this in their business. And that's where Microsoft, I think, comes into play. Uh, you know, we're not a large scale model maker, uh, but we do take the best of the models and bring 'em into the workplace, uh, so that companies can use them.
You know, we talk about AI being a disruptive technology, probably because it seems that it really can and will affect every part of our work, our lives, et cetera, certainly is affecting how we create software with new concept like agents, agent ai. Curious about your thoughts. Share with us, how does Microsoft view what the transformation is going to be like with AI really having an impact and a big benefit to businesses?
Yeah. Uh, many industry, uh, pundits will say that this move to AI agents is gonna, uh, lead to the rise of, uh, uh, more consumption like models or outcome-based pricing. And I think they're right.
I, that's definitely a direction that I see the world shifting as well. Um, the only thing I would, uh, balance that with is that many customers are still, uh, you know, most comfortable buying things on a, a per user or pre per seat basis. And so, uh, the approach I'm taking is, you know, how do I enable customers to buy offerings that they're comfortable with?
And that's typically like a per user or some type of per tenant type of type of license, while giving them the flexibility to, to grow and shift into more consumption models as they, uh, as their business changes. And so, uh, we are at an inflection point, just as we said, and I think one of the things that will change is the business model over time. Talk some more about AgTech.
We think about agents operating on their own or more autonomously. And why, why is that an important thing that organizations are looking to move to? Yeah, It really comes down to business priorities.
Like every business leader I talk to wants to find ways to grow their top line revenue, or they're looking for ways to automate, uh, things that, that they're doing so that they can save money or redirect folks to, uh, focus on more important activities. And, you know, autonomous agents really fit that bill. You know, imagine in the sales context, you can have an autonomous agent, you know, going through marketing leads and qualifying them before handing them off to a human seller.
That's work that wouldn't have gotten done in the past or would've been done by a human seller, uh, and have been relatively low value, not something they, they really enjoyed, uh, about their job. And so an AI agent can do that and add great value to the company, and again, help that company grow on the top line. Talk about Microsoft, uh, and the offering that you have in the context of an end-to-end tool toolkit, if you will.
Yeah. The functional transformation. You know, as far as the toolkit goes, we really believe that there's three essential parts to it.
The first is we think every employee should have an AI assistant in our case, uh, that's Microsoft 365 copilot. Uh, think of that as a productivity tool to help every employee work get through their workday and get more done. That can be quite transformative.
That next step up is where agents come into the picture, and I describe agents as really being for every business process or workflow in an organization. And we have, uh, a toolkit that enables customers to build their own agents. It starts with co-pilot studio, but even extends into our Azure capabilities with our Azure AI Foundry product.
So agents pair very nicely with that copilot that I described first, and then the final step is where the system of record becomes the system of action. And we'll sometimes call this the agentic business application, where you take a CRM system and you add agents and an assistant to it to really transform those three are the essential ingredients or building blocks for AI transformation, um, in a frontier firm or any business at this point. That's a great term.
Can you share some examples of how Microsoft customers are already seeing measurable impact from adopting agent business applications? Yeah, one example I I think I can give is lifetime. Uh, they're a great customer of ours, uh, based here in the United States.
They've used us as part of their finance and supply chain operations, and they've, uh, deployed agents to basically speed up how they handle and process e-commerce orders. In fact, I think it saved them 95%, uh, in terms of their order, e-commerce, order efficiency by deploying, uh, AI agents and agentic business applications to solve that problem. So I think that's a great example.
We also have, uh, another great example, uh, from Europe, um, uh, a large utility named Enco who deployed a multi-language, uh, AI agent, uh, for their customers to help them scale and address customer questions. Uh, it's a pretty cool solution, saves them time and again, helps them scale up. Now, it's a, it's a really exciting time in the world of agents.
You know, you talked about the C-suite and the kind of the three phases in this adoption curve as we adopt ai. What, what kinda recommendations do you have of like, how to get started? Well, maybe near some of the near term activities?
Well, you know, as I said earlier, I think AI is gonna transform every company, every function, every industry. And that opportunity is so vast, sometimes it's hard to know where to get started. And I've got two bits of advice really there to, to anybody that, uh, that is, is pondering that question.
Uh, the first thing is, again, start with a function. Pick a function that you want to go after and then peel it like an onion. You know, go look at the next layer, which are the business processes in that function that you can apply, uh, a copilot or agents to, to really transform.
The other thing though is like, don't get into analysis paralysis. Just pick a business process. Just go pick a business process to get started with.
And as you learn from applying AI to that business process, whatever it is, whatever your thorniest business process is, go apply AI to it. You are gonna learn, your organization's gonna learn, your culture will adapt, and then it will flow from there. So the opportunity is so vast.
Don't let that keep you from getting started. You gotta get started, start simple, pick one thing and go from there. So as we move to an agentic business environment, let's talk about the people.
How do you see the role of human creativity, judgment, leadership? How is that gonna evolve? Yeah.
Well, that's an excellent question, and one of the things that we'll often talk about is how AI and frontier firms is gonna transform the way we work. It's gonna change, uh, the org chart into more of a work chart. Uh, we sometimes talk about, uh, frontier firms taking on the Hollywood model where individuals will swarm around a problem, like focus on a problem and then disband, uh, you know, uh, when the, you know, once they've got a solution.
And so it's absolutely gonna change the way we work with others inside the workplace. It also, I think, is gonna give rise to a new idea that we call an agent boss. And you can imagine, just as a people manager today might take work and delegate it to different humans on their team, uh, you know, resolve conflicts and sort of manage performance.
Every one of us in the future is gonna do that, but not just with humans, but also with agents. So imagine taking a business problem, breaking it up into pieces, delegating it to agents or agents, uh, on your team, resolving conflicts that might come up, applying human judgment. Uh, it's gonna be an absolutely transformational moment, and I'm really excited about it.
I, I really believe that, you know, there is still a huge opportunity for humans with human ambition to go do great work amplified by ai. Talk a little about, about how you see the, the difference in, in the working together between applications, assistance agents, the different technologies. Well, that's an excellent question.
And you know, we really have this complete toolkit that spans everything from the assistant to the agent, to the application. And I think those three things work together in a symbiotic way. As an example, you can imagine that I might go to my assistant and, you know, ask a simple question, my AI assistant, like Microsoft 365 co-pilot and ask a question about, you know, uh, summarize the emails or help me respond to the emails I've got.
Or I might use an agent to update a CRM record after I meet with a customer, but I'm still gonna want to go to an application. Really, that's a purpose-built experience for me if I want a more specialized or fine tuned experience. So those three things really work together.
Talk a little bit about the industries or maybe the business functions that you expect to see reshaped first by agent AI transformation. Yeah, there are three or four, I'd say functions in particular that are, I'd say, um, you know, ground zero for, uh, functional, uh, transformation with AI and agents. Certainly customer service and customer experience is one that is, uh, absolutely being reshaped and say a very early adopter of ai, no doubt about it.
Another, where I'm seeing a lot of early AI adoption, uh, is in sales. Uh, and it's just because, just as we talked about, there's a big opportunity to use AI to basically increase capacity for that organization to grow top line revenue. So the business impact there is undeniable, but I also see it in places like finance and supply chain, where you can use AI to shorten the time it takes to, from an order to actually being able to ship it that order.
We're seeing people use AI to improve accounts, uh, payable and accounts receivable. Um, so we're seeing some pretty, uh, interesting impacts there. Let's talk a little bit more about the frontier firms.
You know, they're not just adopting technology, they're also changing the way they're do doing business around AI agents and co-pilot capabilities. You talk about, you know, what they're doing to invest and support the short-term ROI that they were looking to get for long-term innovation. The most successful frontier firms are doing actually is taking a functional approach.
And so certainly they'll think about their entire company, uh, but they'll really take an approach that's function by function. They'll think about their sales function as an example, and then they'll peel back the onion a little bit and understand which processes inside their sales department they can automate. Uh, using AI agents as an example.
They'll look at which, uh, things their salespeople need help with, where you could pair up an assistant like copilot to help them throughout their workday. And they'll even look at how they can bring agents in to really augment business capacity, maybe to grow top line revenue or take out costs. But starting function by function has really been the recipe that these, um, frontier firms, uh, are using to see success.
As an example, in our sales organization, uh, by deploying co-pilot and agents, we've been able to improve revenue per seller in some cases by almost 10% in some organizations. And, you know, if you think about that giving a seller 10% more capacity is like giving them an additional month, uh, in a year, uh, without actually having them spend any extra hours through the work week. And so there's some pretty remarkable results.
That's just sales. We've been able to do the same in customer service in our finance department, in our IT department. Our legal team has been able to reduce costs by 5%, uh, by deploying AI and agents, uh, within their, uh, within their functions.
How do we ensure trust and transparency as agent systems become more and more autonomous? Yeah. Well, there's two things that we want to do to help with trust and transparency.
The first is we have a rigorous set of principles on, uh, responsible ai. So for any AI that, uh, Microsoft deploys, we adhere to a an important set of guidelines. Uh, and that's really table stakes.
So that's the first piece, uh, responsible AI and our focus there. The second thing that I think is important is we now have an opportunity to start to quantify the value and the impact that many agents will have. And we often will call that in the industry, we'll call that evals or benchmarks.
And increasingly, I think we have an opportunity to help our customers understand how agents are being effective in their workplace using these evals and benchmarks on real world problems. So for example, uh, we, uh, uh, a typical, uh, workflow will be a sales leader doing sales research to try to understand like how they might want to organize accounts or territories or, you know, reshape planning as they think about the year ahead. We recently released a new benchmark, uh, that we call the sales research bench, and it shows how agents can actually help sales leaders in that very specific job, and it's quantified.
Uh, and so I think you'll start to see more of that. And between following responsible AI standards and then using quantitative measures like evals and benchmarks to understand efficacy, I think we're really on the cusp of helping customers know how they can deploy AI in a safe way for real business results. Well, thank you, Brian, for sharing with us your insights and can look into the future, what's happening with the Gentech business applications?
Thank you very much. You know, it's really amazing the pace at which AI has been adopted and continues to evolve. The technology evolves on a near daily basis, but at the same time, organizations have to figure out how they're gonna implement their AI strategies and what the business outcomes that they're most important to their business.
I think it's very fascinating how Microsoft has approached the market, both from the standpoint of addressing the individual and their productivity, but also thinking about new workflows, new models of business, but doing that with not a set of tools, but a set of capabilities that provide integration with data process, workflow, and agentic ai. No one knows for sure what that future holds and what an agentic business might really, really look like. But in situations like today, when we remove constraints of what we can do with technology thanks to ai, that's where the possibilities are created.
You know, using tools like copilot, using applications like Dynamics 365 and we'll, we'll see what end users as well as technologists bring to bear in their ideas and how they reshape businesses for day and tomorrow. And how about you? But I'm super excited about the future that we're creating together thanks to working with technology companies and with end users like yourself.
Hey guys, thanks for the throw. We're here with Sashi Koran, who's the CMO for Nile, and we're talking about how autonomous networks are, are starting to emerge in the age of ai. Sashi, welcome to the show.
Well, thanks for having me, Mike. So just how are networks going to evolve? We've been talking about automation and networking for as long as I can remember, and, and we had software defined networks, and that was supposed to lead to this nirvana, but we never quite seemed to get there.
And now we're kind of looking at it in the age of ai. So explain, if you would, what's changing here? Well, Mike, um, we see a lot of that, um, around us as consumers.
You know, you have, uh, autonomous vehicles now. We have a lot of, uh, things that are becoming, um, sort of functioning on their own. And, uh, when you apply it more into the business side of things, and particularly into networking, this has been a long sought after marijuana for many, you know.
So I think, uh, it's interesting to see the, the journey that's happened, uh, on the path to autonomy, uh, over the last, uh, couple of decades. But, um, you're right in terms of, um, really looking at it from an automation perspective. Um, and if you look at it in the context of the broader enterprise, uh, we've seen this journey with regards to a high degree of au autonomy and au automation happening in the data center where a few operators can actually manage, uh, a fairly, um, complex environment.
And, uh, we've started to see the journey happen in the wide area networks as well, um, with, uh, SD wan, as you might be familiar with. And that sort of gave the promise of, um, automation and management for number of, uh, devices. And, um, then it started to bleed into, um, you know, sassy and things like that.
And where we are at that journey is now into branch and campus environments. And if you start to look at, um, these environments, which is, um, you know, you could think of it as a bank branch, a university campus, a manufacturing entity, a healthcare institution, unlike the data centers, you know, this is where, um, you have users, you have, uh, devices, you have things. And so it's a far more complex environment to automate.
And, um, it also leads to, um, challenges with security challenges with, um, really making that network a lot more agile. And, uh, applying the notion of autonomous networking here is really a very big deal. And, uh, so that's sort of the journey we are on, which is to look at, um, incidents proactively look at patterns proactively, and make the network navigate to these ahead of anybody requiring to manually intervene.
And so, um, the more we can take away from manual intervention and manual configuration, manual provisioning, and go down this path of a hands off operations, then, uh, the network becomes more resilient, more secure, more agile. And so I think we're on the cusp of a lot of these things. And at Nile, we're, uh, at the forefront of driving these autonomous networks today.
Mm-hmm. And how does that manifest itself? Let's say that I have an end point and I am trying to access some service on the back end.
Does the network know that there's a down server somewhere and that it just automatically reroutes it or how smart is smart? Yeah, so it's a combination of a lot of these things, uh, whether you look at it from a network performance issues, whether you look at it from network resiliency or network security. I mean, all of these are things that happen in a real world environment.
And the more intelligent the network can be, the more it has these patterns recognized more, you take advantage of, uh, artificial intelligence and bake these into the foundation of the network itself, then the operations start to become a lot more autonomous without the need for manual intervention. And, uh, for a lot of, uh, you know, C-level, um, stakeholders, their mandate today is to drive transformation. But the mandate also is to do that while reducing risk and cost.
So it's sort of a balancing act and, uh, autonomous, uh, networking sort of helps bring that balance in a healthy way. 'cause uh, you know, you can move faster with agility, but you don't have compromises in terms of, you know, performance incidents, uh, reliability, resiliency incidents, or more importantly, security breaches. So a lot of these can be, um, uh, you know, navigated through this not notion of autonomous operations.
And, uh, icu, you've been covering the industry for a long time. Uh, you know, for every dollar that they spend on CapEx, there are six or $7 they spend in terms of operations over the lifecycle of the network. And if you can really get, uh, to tame that beast of operational cost and complexity, it becomes much easier for them to move forward with their transformation initiatives.
So that's really, uh, the crux of autonomous networks. It isn't just to bring, um, everything under automation. It's how do you do so without making it brittle?
How do you do so without, uh, increasing risk and complexity and how do you do so while reducing costs? So that's the, um, the leveler, you know, which we're, uh, achieving in, you know, customer environments day by day. There are, of course, a lot of types of AI these days.
So is this more of the predictive AI that's then being used to reroute the traffic? Or is it a combination of predictive, uh, generative and a little causal and all those things together kind of create this autonomous network? Yeah, it's, it's an evolution as with everything else.
And when you put things into real world production environments, um, it has to make sense, you know, so I think we look at the maturity of a lot of these technologies as they evolve. And, uh, at least with Nile's architecture, we're a bit different from, you know, box vendors that will put features in a box and then ship it, and somebody else is expected to own the service delivery. Uh, Nile actually, uh, owns that aspect of both the technology as well as the operations and the service delivery end to end.
So we have a greater purview of what's happening across the entire environment compared to any other vendor out there today. And so that gives us, in some ways an unfair advantage of being able to drive autonomy and to take these principles of, uh, AI and make sure that we're rolling these out with a great sense of responsibility. But the architecture is also designed in such a way that it is not a, a la carte approach with, you know, different versions for different people.
When we, you know, roll something out, uh, if we recognize something with one deployment, all the deployment globally gets to take benefit from it because we have stopped standardized on that architecture. So this in turn, you know, brings in a greater degree of resiliency. And, um, I think, um, month on month, year on, we are, we're becoming more mature, more autonomous as a result.
Mm-hmm. What becomes of the role of what we know as the network engineer these days? Uh, many of them have spent a lifetime using you know, CLI tools to troubleshoot networks.
But, um, how does their position kinda evolve into something maybe, I don't know, less stressful or more fun? I don't know. Yeah.
Look, I think, um, I have, uh, been in this industry for close to three decades, and I know you've been covering this industry for quite some time. And, uh, we've always talked about, uh, making the lives of people that are configuring networks and infrastructure in general a lot more easier, you know? So, um, I think, uh, this journey that we are on, on the road to autonomous networking is really meant to augment their skillset, make them much more productive, focus on more strategic initiatives.
'cause let's face it, um, almost every organization, no matter how big or small it is, including Fortune 500 organizations, are really being challenged to do more with less. Nobody's getting this flood of IT resources to grow their teams. And so the existing teams in many cases are, um, suffering from burnouts.
And, uh, many cases, especially when you look at the branch and campus environments that I talked about, you don't have the IT staff in 90% of, um, you know, these environments, even for large companies. So they end up outsourcing it, and that outsourcing is non-standardized, and they have to manage this, uh, operations and try to bring the consistency of the operational model. So these are challenges that hasn't necessarily been solved in the industry with the status quo that we see.
And so, if you can tame that beast, particularly on these edge environments where you don't have the IT staff or you don't have the skillset, or you've outsourced to multiple operators, if you're a global company with offices, you know, uh, worldwide, how do you kind of bring that and make sure that you have the degree of accountability for your organization? How do you bring consistency? So we think, uh, you know, it's something that, um, not just the network, but network and security engineers, these two things we view as two sides of the same coin are just going to welcome it.
And a lot of the response that we're seeing are from, um, network and security engineers, because we're helping make their lives a bit more easier and, um, you know, the environment more reliable and safer for their organizations. Mm-hmm. Do you think we'll get to the point where the network is much more secure than it is generally today in the age of autonomous networks?
Because I'll be able to have, maybe there may be more transparency in real time to see what the nature of the attack is and respond, or is it just gonna be, you know, a fundamentally hardened network and it's just by definition more secure? Yeah, look, I think, um, security is, again, a journey. And, uh, it's always a battle in terms of how do you outsmart your opponent.
Um, and it's a combination of, uh, a defensive strategy as well as something that's more offensive. And, um, you have to, uh, get lucky each time as somebody who's defending the network versus an attacker versus to sort of get lucky just once. Um, and so again, if you look at it, um, I, I look at holistically across the enterprise.
Uh, we have solved a lot of these problems as an industry in the data center where you moved away from perimeter security to looking at security for applications, looking at lateral security for workloads and things like that. And then we talked about SAS in the, in the van, which has been a topic of conversation for six, seven years now. And, you know, different vendors and providers are moving towards that model.
But if you look at the branch and campus environments, which is where the bulk of the users are, devices are, things are, it's still the, you know, wild West from the 1990s, and you cannot solve a security problem, thereby adding one more box or one more protocol, it needs a fundamental rethink. And that's really what we've done where we have, you know, brought in these, uh, data center class security constructs into these distributed branch and office, uh, campus environments, and, uh, espoused zero trust as a foundation. And, uh, you know, some of the things that we did recently, which we actually put out some news releases, were really, um, you know, getting into hacking environments with not a hardened model, but our standard offering that we, you know, give out of the box, so to say.
Um, and, uh, we had a million plus hacking attacks and zero breaches. We recently ran the entire network for Black Hat, which is a cybersecurity conference, 40,000 attendees, three plus days capture, world's largest capture the flag competition, zero incidents. So a lot of these are, uh, proof points of how we're able to build a zero trust fabric and layer in this networking and connectivity, um, on top of that foundation, which is radically different from every other approach that's out there in the industry today.
And we think, um, you know, it gives a huge leg up to organizations that can take advantage of this architecture advantage of this network as a service, you know, delivery. And this autonomous operations, the three of them sort of combining together to form a solution is a very, very powerful offering. Mm-hmm.
As you kinda look forward into this coming year, are we maybe approaching a point where the whole notion of, you know, downtime caused by a network outage might become obsolete? It's conceivable, um, you know, downtime can be caused by, you know, obviously, obviously software issues, hardware issues, but many times, you know, if you start to analyze these patterns, you see downtimes, um, being caused by manual errors, even security issues, if you look at it, a lot of them are due to manual errors, which are inadvertent. You know, they're not malicious by nature, they're inadvertent, they're accidental.
And so the more you can prevent some of these from occurring by going down this path by things that are built on zero trust foundation, things that don't require manual configuration, things that minimize manual intervention, we, we see, um, the resiliency as well as the security going up in order of magnitude to be higher, right? So, I frequently say that if you add more links into a chain, you're actually making the chain more weaker, not stronger, right? 'cause the chain is only as strong as the weakest link.
And so if we can minimize these kind of links that you artificially introduced because your architecture or your box configuration requires that then your, your, your, um, opening more doors for security breaches and interventions, and our approaches just been the opposite. And so I think, uh, as long as we can, you know, be intentional about when you need to manually intervene, um, and how do you, you know, bring in compliance? 'cause a lot of the complexity today in enforcing compliances, there are just too many variables and unknowns.
And so the more you can sort of bring those under your purview, then it becomes inherently much more secure. All right, folks, well, you heard it here. Autonomous networking is coming and it will certainly be automated alongside many other things that we're gonna use AI for.
The question now is how to get prepared for it. Hey, Sashi, thanks for being on the show. Thank You.
Having mine. All right. And back to you guys in the studio.
Hi everybody. So my name is Rashed. I'm at Fabrics ai.
I won't, uh, bore you with a long introduction. I, I help make things, I sometimes break them. And I have a, a few minutes to try to go a little bit deeper on what the, uh, uh, what we're doing with the middleware.
Um, uh, first of all, I just want to talk a little bit about motivation. So this, my, my piece is really about trying to convince you that we're very focused on, uh, one specific problem, which is how to take the agents from, um, from the prototype to production in a large enterprise environment. Um, I was, prior to joining Fabrics ai, I, I, I was an independent person building agents, and I, like everybody else, got extremely excited about what you could build really quickly.
Um, and then like everybody else, I started running into the wall of making this work in a real enterprise environment. So when I joined fabrics, I was actually surprised to find that there was already something in place, uh, to help deal with that. And so I'm going to talk to you about that.
Um, so ba basically, one of the, one of the, some of the, the key things that make this easy to, to get off the ground with an agent is you're, the main thing I'm gonna say is that you're working with it all the time. You're sitting there, you're talking to it, you're interacting with it. Every time something goes wrong, you pick up on it, you're able to, uh, iterate very, very quickly, and then you're eventually, without realizing it, kind of paint yourself into a corner where everything works great, but, but you've kind of guard railed yourself.
Um, and then you put it out in the wild, and then the variables change and everything, uh, falls apart. Usually in our experience, they fall apart for, uh, a number of reasons. So, what's on the screen right now are 16 examples of how things go wrong when you try to take a prototype agent and you try to move it into the enterprise.
Um, these are not the only things that go wrong. Uh, a lot of the things that go wrong go wrong with any agent, whether they're prototypes or deployed at scale. And so there's things you have to solve for if you wanna build any agent, you do have to solve for, how do I retrieve information in rag?
How do I, you know, do good prompt engineering? How do I, all of these things that we've kind of like as a community have been learning out loud and, and, you know, there's the buzz of the month about how to solve this problem and that problem, whether it's chain of thought and whatnot. I'm talking about things that happen when everything's working great, but then you try to scale it and it just breaks, and it, it basically breaks down into these three kinds of categories.
Most of them are gonna be in this context, um, uh, management bucket. And so when we're talking about context purity, which is something that was mentioned earlier, what we're talking about is we're talking about the fact that the LLM needs to get exactly what it needs to manipulate the statistical probability that it's gonna give you the output, you your desire, right? So you are actually playing with a statistical system, and you're trying to game those stats.
That's what you're trying to do. And there's a variety of ways that this breaks down. I'll just pick a couple here.
Um, one of them is gonna be large tool responses. When you're prototyping something, you've got mock data, and it's usually, you know, dozens, hundreds lines long. So anytime you call a tool, you get something, it, the LLM understands it and then perceives, um, in a production environment, sometimes the tools will, will return enormous data.
Sometimes they'll have to return data from thousands of remote systems, right? These are things that you typically don't play around with when you're prototyping, but you deploy, especially my background's in network automation. I'm, you know, I'm telling you, if you try to get an agent to work on your network, and you've got hundreds or thousands of devices out there, and you need to do something on that network, it's going to break unless you have a solution to handle that.
Um, the other ones is gonna be about operations, and how do you operationalize this? And a lot of these, uh, gotchas are gonna be around observability. As soon as you set it out there, you lose sight of it.
And when you lose sight of it, that's when it misbehaves. It's kinda like letting your, your dog go, you know? And they're gonna get into trouble if you're not watching 'em.
Um, and then the last one is gonna be, uh, you know, operational infrastructure. It's gonna be things like the security people, those pesky security people. They're gonna come after you because you're not being careful with the data.
You're not being mindful of what the LL m's doing with it. Um, you're, uh, not, uh, properly segregating different users. You don't properly handle the fact that the person using the agent may or may not have the same rights as the agent itself.
And how do you handle any kind of differences that arise from that? So these are not corner cases. They're going to be the types of things that everybody who tries to scale an agent is going to run into.
Um, so it just turns, it just, you know, coincidentally, completely coincidentally, turns out we've solved a lot of these things. And, um, we've solved it by basically taking three core principles, um, about how to properly, uh, build agents for scale. The first one is, don't resist the temptation to give the LLM the whole problem.
Move as much of the problem into the tooling layer as you possibly can. Okay? And fortunately, in a lot of use cases, there's ample tooling in place.
It's actually underutilized. I was at a recent conference where people were bemoaning the fact that network automation, for example, which is near and dear to me, um, is, has been slow to be adopted, but it's not for lack of tools. The tools are there, it's just that people are very resistant to using them.
Agents can be convinced quite easily to use those tools, right? But designing the tools, software practices, how to, where to put the, uh, the right levels of, of abstraction so that the tools are operating effectively. There's an art to it, and we'll talk about that.
The next one is curate the context feed. I try to tee up the problem for the LLM so that it always scores a home run. That means get any of the garbage that isn't gonna help it out of that context window.
Try to use context efficient formats, uh, token efficient formats when presenting, uh, the model of that information. Um, try to not make the LLM be the thing that carries the data from one tool to another, if at all possible. That's a waste, right?
And so then the last thing is gonna be all of the operational stuff. You have to not just observe the model. You have to observe the agent.
What does that mean? Well, the agent has a job responsibility like a person does. So manage that and, and look for outcomes.
Is it successful? Not just did it, was there a tooling error? Or, or did it take too long?
Or did it actually do the job successfully? You have to monitor it at that level, at the business outcome level. So the way do we do this between this, we, we, we call this the middle, uh, the middleware, right?
Um, and you've already heard the terms context engine and universal tooling and connectivity engines are the two functional pieces. This is a functional diagram. This is not an architecture diagram.
It's showing you the function that the fabrics AI platform provides that sits right between the age agent layer and the tooling, uh, uh, enterprise data and system layer. Okay? And it provides this three basic, uh, capabilities.
Um, a task relevant context, um, so that we're, we're making sure that the context is gonna help the LLM be successful coordination of tools and, and connectivity of tools. What that means is we let the tools talk to each other, uh, at least to pass data to each other, rather than bouncing through the LLM. The LLM can say, Hey, tool two, I just called tool one, the output's over here.
Go get it and then run, uh, run yourself, right? Um, and as well as the connectivity. So make it super easy and, and uniform.
Uh, some of the, the, the, the questions earlier, we're talking about normalization of data and about, um, making, uh, the, the data digestible, uh, at correlation time. So we want to be able to handle that, and that's what the tooling, uh, and connectivity engine is doing. And then the controls, which I've talked about, and I'm gonna show you a couple of examples.
So the, Yeah, I'm Gina Rosenthal, digital Sunshine Solutions. So the tools, so the middleware actually sits, is the first thing it sits on top of, are the tools. It, it collaborates with the tool, collaborates it, it communicates with the tools, and then that is what goes, that, that output of that is what goes into the infrastructure and the data.
But that be a true statement. Yes. I'm gonna double click on the, on each one of these boxes, and I'll show you exactly what goes on inside each one.
Okay? Cool. Thank you.
All right. So the first, uh, I, I should take a step back. Uh, the reason that fabrics AI is called fabrics AI is because we're, our platform is, is, is, is a trifa platform.
It's basically what we call a fabric, is gonna be anything that connects many to many. The data fabric is about connecting data to data. The automation fabric is about lacing together all of the automation processes and the policy management and all that stuff.
And the AI fabric is gonna be all about the stuff that happens at that level of abstraction, where you're talking about things like, um, like, uh, reasoning and, and prompts and, and, and all of that AI stuff, right? And so, two, the two components that I talked to you about, these are functional components that are, that, that can, that constitute, um, the, um, the middleware. They emerge out of that fabric, right?
That, that fabric is what they are made of. Um, the, the context engine is lives inside the AI fabric, and the tooling and connectivity engine lives inside the data fabric. So the context engine in a nutshell is this, it's this box that takes in all of the dirty, dirty context and cleans it up and provides pure signal to the LLM.
Um, there's four Main ways that it does this. Uh, one of them is, uh, intelligent caching. So what that means is that when a, when a, when a tool result comes in or when, uh, you know, any kind of data comes into the context, um, it is cached intelligently based on how much data it is, if it needs to be converted, it's converted, uh, in, in, and it, there's basically a, a virtual file system in which it goes, and then that becomes addressable.
It becomes addressable data rather than data that lives inside the context, okay? Mm-hmm. Um, the next one is con conversation compaction.
So we, we, we, we heard earlier, there's, there's a lot of, there's a lot of, uh, mind share around compacting the context. And a lot of our tools, like Claude Code, whatever, they, they've got that built in the, as you go, and as the context grows, it just compacts. It does this though, through summarization and summarization is very lossy.
Um, uh, especially for long operations, you end up losing something. And it's hard to tell at the time when you're making the summary what you need to keep and what you need to throw away. And if you make the wrong decision, you're kind of screwed.
It's not gonna work. You've lost the information that you need. So the, the conversation compaction, um, uh, technology that we have in the context, it's, we have a few clever tricks to make sure that the compaction maintains the relevant information for the current, uh, uh, conversation, um, dynamic.
Yeah. Is that easier to do? Because you're specifically working with tooling for enterprise applications, so you know what information you need to maintain when you do the compassion.
So what we do is we have a, we have a, a a clever trick that we pull off, which involves understanding. So usually agents have a multi turn interactions, right? Especially if you're talking to an interactive agent.
There's this multi turn things that's happening. And so you might say, for example, you know, what's the capital of France? And it might give you the capital, and you say, what's the weather in London?
And it gives you that. And you say, and what about Paris? Right?
And, and it needs to know, is it asking me about capitals or everybody's asking, right? So the trick that we pull is that we're always looking at what the current goal of the user is, and then we go back into the history and we create a summary that's tailored to the current actual goal rather than the past. And it gets a little tricky to do that, but that's basically what we do.
And so you end up with a series of little tailor made summaries, so when you chain those together, you actually preserve all of the information that's relevant. Thank you. Right?
So I'm not gonna have time to go into this too deeply, but I do wanna point out, you know, that the, that the context cache is, sits there in the middle of the context engine. It's this virtual file system where we put stuff, it lives there. And, and, and there's been a recent, um, paper that was, that was published about, um, recursive language models.
I dunno if anybody's read that, but recursive language model this is, is this idea that instead of getting an LLM to work with a gigantic document, it can spawn off, uh, a bunch of an, uh, an arbitrary number of subagents, and then assign bits and pieces of the, of the document to those subagents, um, through what, uh, in our EPL notebook and our, what that actually gives the, the whole basically benefit is that the, uh, the incredibly large document becomes addressable data rather than data that has to be completely ingested. You can seek it, you can profile it, you can histogram it, you can do all sorts of things to it to understand and feel around its edges, and then locate, pinpoint the bits and pieces of information that are going to be relevant to the thing that you're doing before you load it all inside your collect. What actually makes intelligent caching intelligent?
What, uh, it's, what are you, what are you calling the intelligence behind that caching? Um, so the, in the, the way that, well, there's different little bits and pieces. One of the pieces is gonna be that, um, we're not necessarily gonna cache something that is small.
That's a very basic decision that we're making. If it's small, just throw it in. Um, but if it's large, then we want to cache it.
And then, uh, what we want is, as we cache it, we glean certain bits of information from the data that we've cached, and we provide that to the LLM. So the LLM has a sense of what lives in the cache before it tries to work with it. It's almost like vectorization.
I mean, you're, you're extracting some dimensionality out of the data and, and providing some sort of a, a number stream to the, to the LLM so they can understand that it exists. And if he wants more data, go for it. Something like That.
Some sometimes like that. Yes. It depends on the, on the type of data.
Sometimes it's gonna be very numeric. So you want to, you want to maybe give it, uh, some sort of a basic profile, whether your mins and maxes and averages. What does the histogram of the data look like?
How many roads of data Do you have? Perspective? Yeah, yeah, okay.
Okay. With a pointer to the data, in case they Want to go back with a pointer to the data. And so you, you know, for example, that, you know, there's a, you know, 10,000 lines of data and you have a histogram that says that the first quintile ends at, uh, at, at row.
If Quin first quintile of this particular column, value ends at this row, and so you're able to request, you know, maybe the last three items and the first three items of that, uh, data frame, you know, so all of the machine learning type of techniques mm-hmm. We, we give those to the LLM through, uh, a set of fairly abstract tools. So I have a question based on this slide.
Yeah. Um, which is probably bigger than the question I wanna ask, but more to like, what is actually involved in your solution. So this context cache Yep.
Is an external source. Uh, so does, does that bring your own or do you guys provide a source, or how is that working? So we do, uh, our platform does have, you know, a, a data persister and data storage, but it could be external just as easily.
Yeah. Are you building a data graph DB underneath? We do have a grave grade, uh, uh, graph db.
Yeah. So we, our platform comes with graph. Okay.
Uh, as well as object storage and, you know, So your sources are becoming nodes and you're building all of what's building the edges around all of that, Depending on if the data, uh, is structured that way. Yes. Okay.
Gotcha. Yeah. Alright.
So, uh, Before you go on, I think, uh, Marianne had a question. Yes. Can you talk Marianne?
Can you guys hear me? Yeah. Now we can.
Yes. Thank you. Okay.
So I had a question. I see you have on your slides, uh, here, and purity. So my question was around what, what does contaminate your model?
And what are the enforcement controls? Is it isolation or something else? So usually what contaminates the context is going to be, um, uh, a bunch of extraneous information.
So, for example, take, take a, a, a tool, uh, that its output has a bunch of introductory text and a bunch of, you know, closing text or a lot of formatting characters. Let's say my tool returns HTML, right? I don't need the HTML tags, so why don't I just strip those out and just keep the text, right?
So that's because each one of those little characters in an HTML document is one token. Um, whereas words are, you know, oftentimes you can get like a long word where that's just two tokens long. Uh, but, but if you have an angle bracket on each side of it, that's another two tokens.
And so keeping the, the, the token count low, uh, not only lets you put more information that's relevant into the context without reaching that context rot threshold where the, the, uh, the effectiveness of the LLM drops off. Uh, but it also, you know, is reduces the cognitive load on the model as well. Does that answer your question?
You don't look like I answered your question. I guess my question was more around the controls. Uh, you talked about data context.
Is there some automated governance controls or security controls that prevent that? Uh, no. So we do have, as part of our platform, like as guardrails, right?
That guardrails is one of the operational features. So anything that goes, uh, prior to going to an LLM will pass through a guardrails model that will do safety checks and compliance checking that. But that is outside of the, uh, that is part of the security infrastructure of the middleware, but it's not specifically part of what we call the context engine.
And so is that automatic automatically deployed to do that in seconds, or is that, where does that fit in that process in your middleware? Yeah, it, it does introduce a little bit of latency between the prompt and the, and the LLM receiving it because it has to pass through. But those models are typ typically very fast, and they're, uh, specifically trained to, to, uh, you find, um, you know, uh, non-compliant, harmful type of instructions.
Okay. Thank you. Yeah.
I have a question, but I don't wanna subvert your presentation. I think what some of us are grappling with is, okay, you have a bunch of slides saying you've been there, done that. What's the customer get as a deliverable code?
Consulting support pilot projects? Yeah. We're a product company.
Yeah. Uh, so our, and, and so our product, uh, you know, we, we, we deliver that product in a variety of ways. But, uh, I would say that, uh, oftentimes what we're delivering is agents that are going to be, uh, operating inside of a, you know, enterprise and they're gonna be effective and scalable.
Uh, but for some customers, what we're uh, doing is we're saying, here's a platform in which you can develop your agents. And what we're saying is, as opposed to some other platform that's out there, um, that this, this is going to provide the facilities or all of the services, the platform based services that agents need in order to be scalable and reliable in, in an enterprise. And I would think, from what I've seen of other products, not in this space necessarily, but folks might want a pilot project where you demonstrate how to use the toolkit on some sample problem, maybe one or two.
Yeah. And then maybe some consulting support as they go along or the Yeah, Yeah. We always help.
We're always, you know, uh, we always get in there. Um, but our business model isn't primarily consultation. We want to give people, enable them to mm-hmm.
To do stuff. So rights to code, but not so much services, which that's doesn't Scale. That's, that's right.
Yeah. And we have a very extremely low code type of approach. I'll show you an example of that, uh, in, in a slide that's coming up here.
So just before I run outta time, I'm, I, I may not have time to do, uh, the demos here, but I'm, I'm going to, uh, talk to you about the connectivity engine a little bit because it's the other important piece of this. Um, this is the part that, uh, mentioned earlier that's connects to the tooling, right? So, uh, tools, uh, MCP tools, for example, they sit out there on a server.
And MCP was a great innovation because it creates this level of abstraction around APIs. Everybody loves them. And as we mentioned, however, they're in an enterprise environment, there's still plenty of tooling out there that isn't, um, wrapped by an MCP uh, server.
But even if it is, right, you have this problem, I, let's say I have an MCP server that's a MCP server that stands in front of my SQL database, and I have another MCP server that stands in front of my, uh, my, uh, my inventory system. And I somehow need to update, uh, do something that involves information from my inventory system and information from my SQL database. Well, great.
There's these two SEP tools. So the LLM can easily call the tools that are sitting there, but if anything needs to happen here based on something there, the LLM has to suck in all of the responses. Understand that, construct the data that it then pushes down to the other MCP server, and you just wasted a bunch of tokens, essentially performing a shuttling process from one thing to another.
Mm-hmm. Right? So the fact that we have this MCP, um, the, the, the, our particular, our, our engine of your, the, the, the tooling engine is, it creates kind of like a, a, uh, a net, uh, an MCP abstraction layer.
It's itself an MCP server that wraps other MCP servers and non MCP based tools. And because we do that, you get this joint and common execution environment for them. My tool results come back and they're in the tooling engine, and therefore I can move them to the next tool as an intermediary without ever touching the LLM.
I just let the LLM know that the tooling, the tooling results are there. Right? That's one thing.
Another thing is that, um, we have this YAML based way of defining tools because we provide some primitives inside of the, the tooling engine that pretty much lets you develop any tool. And through this, uh, YAML based abstraction, you can easily develop new tools and you can take a, a tool that might be, for example, SQL Query tool. And a lot of people, they love this when they're building a prototype.
They say, look, I hooked it up to an MCP server, it has a SQL Query tool, and now the LLM can write SQL queries. Awesome. Except it breaks on production because building SQL queries, if anything goes wrong, uh, it doesn't work.
And the LLM, the bigger the job you give it, the more likely that the SQL query it will build for you, will be wrong somehow. It's too complicated. It can't come up with it at inference time or whatever.
When you have, uh, this, um, this, uh, middleware layer, you're able to define a tool for specific types of queries that are going to be common in your enterprise workflows. And now it's just simply calling a tool that has a name and a bunch of parameters rather than constructing an expression that's brittle. Right.
And we can actually, you know, you can talk to the system and have it generate those for you. We have a, a thing that makes the tools for you, so you don't have to bother coding all that up. It'll do it because we have something called dynamic data discovery, where you can say, I'm interested in doing these things.
The system will connect to your MCP server or to your data source, and it will not just consume the entire API, it'll explore that API and figure out exactly where the data lives. Mm-hmm. And it'll skip the tables that don't have data.
It will focus on the tables that do, and, and, and, uh, it just generally means less responsibility for the LLM at inference time. It doesn't have to query the wrong table and say, oh, I've got zero results. Why is that?
Oh, there's another table with the same data every single time it runs. Right. You mentioned early on that you were doing sort of data science type analysis Yeah.
With the intelligent caching. Yep. Do you do similar type of data science work in this layer as well?
Yes. As you explore sources? Yes.
We, because data science is in our DNA, so we do, we, we do, uh, schema discovery and we, we code our agents when we write our prompts and all that stuff to do exploratory data analysis as just a matter that they work, always do EDA every time. And we give them the tools to do EDA because when you do your data analysis, you create shortcuts to good information that you otherwise would, uh, would miss. Great.
Thank you. Okay. Would you guys, um, like to see the demo, or actually we have a couple case studies.
Case studies would help me understand that. Yeah. All right.
So I'll, I'll just, I'll just, uh, I'll skip to a couple of case studies then. Um, so two, I have two, I have two case studies for you. This is a customer, they came to us.
They had, um, uh, a lot of documents. Uh, this is actually not an IT or a networking use case. It's straight up, we have an enormous cache of documents.
Each document is very, very large. We need to be able to, uh, to query those documents. We currently have a team of 12 people, and it takes them a long time to, uh, find the documents that we're looking for.
Can you do this with an agent? Somebody said, yes, sure. And they built a prototype and they did the demo, and they said, this is fantastic.
And so they tried to use it and it broke almost immediately, right? So then they came to us and they said, can you do better? Um, where we said, well, we're certainly gonna try.
And so, uh, we built, uh, this using our middleware with all of the features that I've been talking to you about. And so some of these questions that they had, like, you know, um, we're very basic questions. How come for one time I tell you to find the matching documents?
You gimme 12 results, and sometimes there's only three. Uh, this is the kind of stuff that drives people crazy when they use agents because it's probabilistic and it's, you know, not always the same. Well, we want to get that to the bottom of that.
So a lot of the techniques that I've been talking to you about, we developed while learning how to do this right. Um, and, uh, and, uh, so we were able to, you know, deliver consistent results. Um, and it convinced the customer that basically with the agent in a few minutes, they're able to get the results that are better than what their team of 12 people were usually able to accomplish in two or three hours.
Okay? So that's, that's one use case. And it's involves like, enormous amounts of large documents.
When you're prototyping, you would never use, you know, 50, 60,000 token documents with lists of thousands of them. Where when you try to find matches, you get 150 matches. Like, how do you, how do you handle 150, 60,000 token documents in an LLM unless you've got some sort of a context management strategy.
So that's what, that's what we're doing. The second case study is, um, is a case study where the, the customer had an enormous number of systems. This one came up a little earlier in a preview presentation.
So it's a very heterogeneous environment. Each one of these tools provides an amazing dashboard that lets you look at their, their information, right? And, um, and so they're like, don't replace, um, there was a question earlier about data lakes.
We can do a data lake, but they didn't want us to do a data lake. They said, no, no, we have, we're very happy with the job that each one of these tools is doing, gathering information and syn synthesizing it. It's just that in our business, sometimes issues come up that span these domains.
And right now I have to go and sit in front of one beautiful dashboard after another and correlate between them. And a person has to do this work. Can, can an agent do that?
So we took that on, uh, and, and we, we, we built, we built an agent. Um, and we basically, we, it's what I was just talking about was the, the, uh, dynamic data discovery is what led us do this a little bit better than just federating all of the data and ingesting it into our platform, which is what a lot of people would say is they say, oh, the data exists in all these places. Why don't you just pull it all in and, um, and into our system, and then we can do all the analysis.
Um, certainly there's times that we do that, but with our platform, you don't have to do that because the, uh, the, uh, middleware and the, the dynamic, uh, schema discovery is gonna be really good at, uh, understanding where the data lives and then going to get it when you need it. There's another thing that I really want to show you guys. I, I don't have time to, to to do the demo, but I do want to talk to you about the fact that we take a very different approach to, uh, evaluations as well.
I, I mentioned this at, at the top right. Um, a lot of times you'll see, uh, observability platforms that are simply, um, evaluating how many tokens, how, what the latency was, uh, whether there are any, uh, tool errors or anything like that. Um, but we have this really nifty system, um, that, uh, once a agentic session goes idle, we'll go back and review the entire session and perform a qualitative analysis of it, and it'll pull out various themes, what the topics that were covered were any lessons learned, any, uh, optimizations that are possible.
It creates these reports. It scores the agents dynamically. It considers any user feedback as well as any interactions.
Uh, and then, uh, it creates, um, an improvement strategy, kinda like a pip for each of the agents that the, uh, agent administrator is able to click through and apply on the go to update instructions, uh, over time and, and guide the agents to, uh, better performance Institutes. The first time I've heard putting your ais on a pip. Yeah, well done.
Well, you know, because we're giving them jobs that are human-like jobs. And this is kind of like, um, not only like the, the first point I mentioned where I said, push as much of your job into the tooling layer as possible. Mm-hmm.
But then there's still like, what are you asking the agent to do? And I, I really think what you want to ask it to do is something that you might ask a person to do mm-hmm. That they don't wanna do.
Like it's a person like activity that sits low on the person, like enjoyment scale. Right? Get the agent to do that, Right.
And then measure them, And then measure them like people, did you do a good job? Do I need to train you? Do I need to give you additional information that you didn't have before?
Right. Yeah. Okay.
And I think, uh, what, maybe since I have 15 seconds, we do have spend management and cost controls, uh, built into the platform, which is sometimes people forget about that. Agents end up costing you hundreds of dollars. It's kind of like, uh, back in the nineties when the kids used to call the party line, um, we, we have controls to help, uh, Hundreds.
Where are you getting your cheap AI from? Thank you very much for Welcome to the Security Boulevard, the cybersecurity podcast from the Future Room Group. Each episode explores a variety of topics within cybersecurity and the technologies that drive it.
com, security Boulevard, YouTube channel, Textron tv, and all of your favorite podcast platforms. Before we dive into today's episode, let's meet today's panel starting with Mitch. Mitch, it's good to see you.
Hey, great to see you, Mitch Ashley, I lead the software lifecycle engineering practice analyst practice at FU and kind dabble in security working with Fernando and team and software, supply chain, security, all that good kind of thing. It's good to have you and, uh, the other co-host this week. Uh, Mr.
Fernando Montenegro, Fernando, I, I, Montenegro, I lead cybersecurity research. And, uh, we had, we barely started, there's a correction to be made. Mitch doesn't dabble in security.
Mitch know security, right? So, uh, um, it's interesting because there is a, a, like, as the industry's changing and whatnot, the, the, the overlaps between, uh, security and observability, for example, right? Uh, his amazing at this.
And, and, and I rely on, on his judgment and knowledge of many, many times during the week. Well, let's end this episode right there before Fernando changes his mind. Thank you, Fernando.
And of course, I'm Tom Hollingsworth, event lead for security at Tech Field Day, a part of the Futureum Group. Let's dive into this episode now as we're recording this. The weekend was a little bit messy for most people.
There was a huge weather system that tracked through the United States. Uh, there were power outages, uh, schools were canceled. They even closed a couple of waffle houses.
And for those of you who are not in the US closing a waffle house as tantamount to shutting down the military, the government, and anything you've got, because it is the most reliable service that we have. But that made me think a little bit about the way that we plan for disasters, because a giant weather system may not be the kind of disaster that you think is, is going to cause a problem. But what about a security disaster?
What if, uh, you know, I don't know, some of your, uh, password files get leaked online? Or what if somebody manages to abscond with some of your data or crypto lock it? Uh, how do you respond to that?
Can you respond to that? And worse yet, is your response to that, oh, I better go check to see if my disaster recovery plan is up to date. Because we all know that sometimes people don't think about disaster until it's upon them.
So I'm gonna open this up to you, you gentlemen. Have we ever seen one of those situations in the past where it felt like the disaster recovery plan was making it up as we go? Oh, I can, I could name a few names and a few companies, but I think I'd get in trouble if I did that.
No, I think, I think, you know, it, it, it's funny. We call it a disaster recovery or, or business continuity plan or whatever it might be. It, it, it's sort of like any plan that you create, whether it's disaster recovery or a project plan, is outta date.
As soon as you hit save, 'cause the environment changes, something's different, something's new. You know, the, the CEO comes to you and says, we're gonna do this now. So it's, it's a little tough to create a fixed plan.
Uh, so I think you're almost really developing a system, and the plan is just a reference implementation of your system, of which you're disaster recovery process systems, people, et cetera, look like. And I think if you take that kind of an approach, you'll have a little more to borrow a term from, from Fernando's cyber resilience practice, you'll have more resilience in that system as opposed to, well, we did that in, uh, what is it, oh, 13. I think we're a little outta date.
We probably ought to update this. Yeah. A lot of things have changed since two th 2013, so, Yeah.
Well, so, uh, the, we also had, so I up in Canada, and we also had weather here, uh, I here in the, in the greater Toronto area. I haven't checked the news recently, but it, it may have been like record snow storm. Uh, it's no accumulation, thankfully.
It's, it's, uh, relatively light and fluffy snow. So I can, I can deal with it. Don't You call our, I'm sorry to interrupt, Fernan, don't you call our, our kind of storm that we're having, you just call that Monday up there, don't you?
Isn't that Sort of No, no, no. Listen, listen, I, and, and, and, and tying back, tying back to this topic, right? I think that one of the things about preparedness is that you always have to have a, um, uh, a threat model.
Like your threat model needs to have, uh, needs to be realistic to your environment, needs to meet your requirements, and needs to, to account for the things that, that you are, uh, willing to, to, to, to address. And in many places in the United States this week, uh, they're not, like, they're not statistically ready to, or it's not statistically, um, relevant to them to have preparations for this. So my heart goes out to the people who have been affected by this.
I know that, uh, um, I know that, uh, that, uh, it, it was very disruptive in many places where an inch of snow, an inch of ice is literally the community stopping kind of stuff. So, but yes, like the, the, the amount, these kinds of amounts up here in, in the greater Toronto area. Yeah, that's, that's Tuesday, right?
Uh, funny enough, uh, if you go further north in Canada, right? People who live further north think that about us here in Toronto, right? So, so like, it's, it's fine.
Like it's, uh, it's All relative. It's all relative. But, but, but Tom, to to, to your point on, on preparedness, right?
I, uh, uh, there is that, uh, quote that always attributed to Eisenhower, uh, sometimes variations with, uh, uh, with, uh, Mike Tyson, right? I mean, the, uh, plans are worthless, but planning is everything. And the other is everybody has a plan until they punch in the face, right?
Those are the, you can figure out which ones Eisenhower, which wants Tyson. It, it's a really good point. And, and to illustrate that, I actually wanna bring up something since, uh, Mitch mentioned 2013, uh, a lot of people lose track of the planning process because they never actually test their plan.
I worked with a customer many years ago at a previous job, and, and we don't get a lot of snowstorms in Oklahoma, but we get the other kind of weather that tends to level buildings, uh, especially in the springtime. Uh, we are in the middle of tornado alley, and, uh, I was working with the school and they had a solid backup plan. Uh, you know, we're gonna put things on tape, and if something happens, we've got the tapes.
And then one day they had to evacuate the main building because of a, uh, a tornado threat. And someone realized that their backup plan wouldn't work if the tapes were still in the building, if it was gonna get hit. And they had to literally run into the building to grab a box of backup tapes and throw it in their car as they're evacuating.
Fast forward a couple of years to 2013, and one of the largest F five tornadoes ever recorded, hit another school district administration building, and they had to sift through the rubble to pull out the drives of the servers, to insert them into different servers, to be able to run payroll for the teachers to be able to buy supplies, to start rebuilding their houses and things like that. Because they never thought what would happen if we got hit by a, a severe weather event so big? It literally leveled the building.
You don't think about those things until you're in the middle of a disaster. And that's one of the reasons why you have to test your disaster plan, because you need to know where the failure points are. Oh, the generators will immediately fall over if the power goes out.
Really? Did you test it? Will they automatically fall over?
Or does someone have to go hit the big switch to cause a failover? If that's the case, who's gonna do it if everybody's pinned at home in the middle of a snowstorm? And how does, you know, things like DNS records work and what's gonna happen if somebody's logged into the servers when the power cut happens?
Like if you don't test it and then jot down all the failure points, you might as well throw the binder out because it's useless to everybody. And I spot on, and I would argue there's one step before that, which is what goes into the plan in the first place, right? And I think that that is an area where, uh, bringing this back to cybersecurity, we've had, uh, multiple cases where incidents have happened where it's not like, it's not that they were completely novel, right?
It was the case where, yeah, you think this true a little bit, and oh, by the way, that can happen, right? And this is an area where I am, I'm, uh, uh, I'm really interested in the work that's being done in adjacent areas, right? So in aviation, for example, there's an entire field of study on near misses, right?
So why did this almost happen, right? And, and, uh, I know that in cybersecurity, I know that, uh, Adam Tack has written about near misses in the past. Mm-hmm.
And I know that, um, and I think that like, uh, Wendy Nater and Bob Lord, uh, uh, were talking about cisa last week. So Bob was at, at CISA not too long ago. Uh, and, uh, I think that they're doing some work on near misses as well, right?
Which is how do we as a community learn from the near misses, right? This bad thing, why did something horrible? I mean, it almost happened what got us there, right?
So, um, yes, very much about planning and very much about thinking through, sorry, where being creative about what can go wrong, right? It's risk management. My kids hate when I, when I, when I talk about risk manage, I use it to tease them, of course, right?
But, uh, it's risk management, right? And I think that, that, that our jobs in, in this industry is to elevate that risk management conversation. Tom, to your point, precisely to think through about these things.
By the way, we do have an official mascot now. Nate, the cyber cat has joined us. So this is Nate, everybody good buddy of mine.
You know, it, it's, it's interesting just kind of my own experiences with disaster recovery plans. I think one of the takeaways for me was, you know, even a plan can be a bad plan, right? Just because you have a plan doesn't mean it's a good plan.
And because you've tested it doesn't mean it's foolproof. That the things that worked when you tested it, tested it could still fail, right? You still could have some problem where the UPS doesn't kick in like it's supposed to.
The cooling doesn't happen like it's supposed to. But even more simple, more easy, easy kind of simple things you would think about, don't assume, like one of the, one of the plans that had to just completely revamp 'cause it was, um, really more than a decade old and how the technology was changed. Um, but when the assumption was that such and so employees live close to the building, and if this happened, if the computer data center overheated, they could come open the door and cool it down.
Well, what happens when you camp the ice storm, the fire, the whatever? And, and that there were things like that, that we addressed and, and, and fixed and made sure that that wasn't the, wasn't the the main way we were gonna solve it. And then as it happened, within five years, and I think it was then another three years after that, we first had floods that went right up to the edge of the building.
Nobody could get to the building. And a couple years later they had complete fire. A whole bunch of houses got built down, businesses got, but uh, burned down, um, right near our offices.
Again, can't get into it, right? Can't even get into the neighborhoods to do anything. So physical access is not an option.
So I think that's one of the things when you think about analysis, Fernando, is think about what assumptions we make that we shouldn't assume that's true, even if we've tested it. Does that ring true? I think that one of the things that rings very true is that the world is unpredictable.
The world. Like we can't guarantee anything, right? And, um, one of the, one of the concepts that I learned along the way is that I love, and, and, and it's very, very relevant to this, it's the notion of resulting, not sure if you ever heard of the concept of resulting.
Mm-hmm. So, uh, I read about it from, so Annie Duke, she's a, she was a poker player and now teaches decision science. And Annie Duke wrote a book called Thinking that, uh, basically how to make better decisions.
And the idea of resulting is when people mistake the quality of the outcome with the quality of the decision, right? And we cannot control the outcome. We can control the decision, right?
But we cannot, and and you may have made a horrible, a poor quality decision that still turned out okay, right? Oh, I'm gonna, I'm go on my tiptoes on this thing to pick up that one thing over there as opposed to the ladder, and I still fine, right? Um, I say that it's particularly, I don't mean to bring football into this, but uh, as we're recording this, the, they just defined the, the play, the, the teams for Super Bowl 60, right?
So Seahawks and Patriots, uh, 11 years ago, they had, that was the same, the same, uh, uh, matchup. And any, duke uses the example of what happened in Super Bowl four nine, uh, as an example of resulting because the, the, the Seahawks made the play at the one yard line that didn't work as people expected, and they were crazy about the result. But then when you look into the decision that was that the decision, that was a sound decision that led to that play.
Sure it didn't work because that's life. That's, that's football, right? That's the, but uh, sorry, I'm meandering as usual, but it's this, this notion of what is the quality of the, the decisions that you were making as a professional.
Are you, are you thinking through like the checklists that you need to do, Tom, to your point, to make the plans, like to think through like How good is your decision process to create those recovery plans, to create those, those plans that, uh, do they match reality? And I think the point that you guys are kind of bringing up that it's important to realize is that a lot of these, these plans are based on a series of assumptions that we need to make sure that we can validate. So here's a good one.
Let's say you have some kind of a data protection system in place that is creating regular backups that are stored offsite. What if you have to log into active directory to restore those? What happens if active directory is no longer available because it's been violated or it's been shut off or something?
'cause we're seeing that a lot now in cyber attacks, is that people are going after backups and people are going after active directory to halt any kind of, you know, cross system functionality. Well then how do you get your backups back? Can, can you get them back without being logged into active directory?
And those are the kinds of assumptions that you have to challenge. Kind of to Mitch's point, you know, someone will always be able to go over to the building and open the door, but what if they're not? How do we plan for that?
And I'm sure that you guys are probably sitting there thinking to yourselves in the audience, you know, oh, this is just a big tabletop exercise of how are all the crazy ways that I can, I can get around this? Well, in security, those crazy ways to get around things are exactly what your attackers are thinking of, right? They, they want to try to hit you from an angle that you're not expecting, that you haven't planned for.
I mean, I saw something today about, you know, people trying to get account, um, access by sending text messages or communicating with people, uh, through email going, Hey, this is, uh, at and t we just wanted to reset your password. 'cause we noticed somebody's trying to log into it. Can you give us that code that you just got texted?
It looks legit. All the links work, except, oh wait, that thing that you just got is actually gonna allow me to take over and then I'm gonna start hammering all of your other accounts for two-factor stuff and whatever. Like, those are the kinds of things that you have to be ready to adjust in your plan.
What happens if an admin gets compromised? What happens if your CEO's email starts spewing out, um, you know, spam or, or attacks or something like that? You have to think through that.
And I realize that a lot of this feels like exception handling, like, you know, programming in a auto, an autonomous car. Like what happens if a 7 47 lands on the highway? Like, you're right, the, the likelihood of it happening is low, but it's never zero.
So, like, you know, how do you Tom about this? Go ahead, Tom, about this is, it's not a linear process either. The, you know, they're, the, the, what you see as the attack may not actually be the attack.
There may be a secondary action is, which is really to get you to start the backups. And that's when they're gonna compromise them because the system that you're, you're restoring, uh, to is, is actually the one that's gonna compromise and is gonna steal the data. It's kinda like in, in, I guess in tank warfare, not to get too militaristic about it, but shells that penetrate tanks.
It isn't the first, uh, contact with the tank that the shell penetrates. It, it's really, that's just setting the stage to get through the explosives that happen that actually would then open it up to be penetrated by the follow-on, uh, secondary explosion from the, the shell that's attacking the tank. So it's, you have the added factor of what, not only what if what this happens, but what if we're not able to do what our action, the response is what if we're suspect about whether that environment's compromised or could be compromised if we take the restorative action.
So it, it's, it's a multidimensional kind of three dimensional chess, if you will, and tie back to Star Wars and or Star Trek and Spock. And this is what, this is an area where like, uh, one of the, the areas I work closely with is, is that, uh, part of the reason that the, the practice that that I run is called cybersecurity and resilience, is that I work very closely with the, with many of the, the, the backup and recovery vendors that are that, that work in this space, right? And one of the things that they do is precisely this notion of, okay, what does a clean room restore look like in a scenario where we need to restore ad before we do anything else?
As a matter of fact, a number of them are now tiptoeing their way into identity protection precisely on account of that. Mm-hmm. Not to throw the, the AI into it as well.
And, and they're using AI to help, uh, optimize that process. The, so it's, it's an evolution of, of, of, uh, it's an evolution of the technologies, an evolution of the capabilities that practitioners have to review their plans and say, okay, alright, what is it that, what assumptions are we making? Oh, we're assuming that we need to recover ad ourselves.
Oh wait, our vendor can now do that for us. Okay, can we trust that the vendor can do so? It, the, the, it changes the nature of what you need to do.
You're not recovering ad you are making sure that the vendor can recover ad for you, right? So, uh, the evolving nature of the, the, of the disaster recovery plan is not only that things drift for bad, like I, but sometimes they drift for good. Like, I mean, there are new capabilities into your environment that you didn't have before.
How can you make best use of Fernando question for you. You know, my thought is just like, you might use AI to analyze your business plan that you're, that you're creating, or maybe it's helping you create one. Certainly AI could be a good, uh, sounding board to run your disaster recovery plan through and test it, you know, put pressure on it.
Um, but also could be for scenario testing. Like what are the scenarios we're not thinking of? What are the latest kinds of scenarios we should start to consider how we adapt to, so AI could be your friend actually in helping you do a better job or keeping current with what's happening.
And you can rule out the, you know, far edge cases. 'cause you don't think that's practical for you to even validate or test or protect against. But certainly you're not relying on just your thinking.
It's like using an external consultant to help you. Uh, absolutely. The the challenge there, the, the him every, every episode at some point, AI comes in, right?
Uh, uh, the challenge, the thing I I challenge people there is that, look, yes, I agree wholeheartedly, but who is running the AI and whatcha are you, what are you using the ai? How are you using the ai? How much are you depending the ai, are you a backup and recovery specialist who is interacting with the AI to ask, hey, um, like precisely as I said, to do the, Hey, let's, let's think through this, this scenario.
What am I missing kind of thing. Or are you someone who is not as experienced in backup and recovery and you're coming to the AI for, oh, tell me what to do, because I dunno. Right?
And if it's the latter that is complicated because you dunno how to evaluate the, the quality of that outcome, right? As well as if you are an expert, uh, and you're just using it as a tool. It, we go back to the AI as a, as a tool for an experienced professional versus, uh, uh, a less experienced person using it.
And, and not catching the hallucinations, not catching the mistakes, not catching the, I say this as somebody who's using AI a lot, uh, for my, uh, uh, personal finance tracking, right? I, uh, it's great, but I know what I'm doing right? I, I know what to ask.
You know what I mean? I, I think it's important to realize that a lot of the pieces that get missed are kind of institutional knowledge. And I think where AI is gonna fall apart is that no one's ever documented those.
Mm. And and that's one of the reasons why I know for a fact, that's the reason why I originally started writing on my blog, was because a lot of the things that I was learning were things that maybe didn't exist anywhere else. Like we all know the XKCD, you know, the infamous, well, you know, who are you Denver coder seven?
And what do you know? Because so many questions get asked that never get answered. And when someone does answer it in their head and go, yeah, that's how this works.
They don't ever write down what that answer is. And in the old world that was, oh, that means that I've gotta solve that problem. But in the modern world, it's AI doesn't have a knowledge base to draw off of to create a solution to that problem.
And so we, we kind of find ourselves with that knowledge gap, right? 'cause this is what I have been helped with up to this point, and here's where I need to be and how do I jump that gap? Because I promise you, an algorithm is not gonna be able to jump that gap no matter how many GPUs you throw at it, because it really doesn't know how to think outside of the box.
That's where humans are still valuable in the loop. What happens if these conditions aren't met? What happens if this really random exception occurs?
And like, you know, it could be something as stupid as like a race condition. The power comes back on, but the network isn't up and it starts timing out because it will not connect. What do I do now?
We don't know what the answer is until we have tried it right now, before anybody else goes out there. Do not walk into your office tomorrow morning and just throw the master switch and go, let's see what happens. You need to kind of, you need to game through this.
This is what I want happen. This is what I expect to happen. Let's test this in small scales to see what occurs.
And then let's analyze what we tested to make sure that if something does go wrong, it's contained and fixable before we try it on a larger scale. Because if you do that, just yank, let's see what happens. Uh, in the industry, we call that an RGEA resume generating event.
You, you do not wanna be on the receiving end of one of those, especially nowadays. Funny, like I have so many things. First of all, uh, right?
Uh, Mitch, you're based in Denver, aren't you? Yes. Yeah.
Yep. And he knows, like, who knows, he might be that Denver cold Tom, who knows. Right?
Could be. Could be. Yeah.
Could. But, uh, funny, funny you mentioned that, that the throwing off the switch back when, when, uh, I was working on, on network security implementations, uh, I remember working with a, a healthcare customer where we built a, a multi-site, uh, uh, network with the firewall modules and switches and stuff like that and, and, and so on. And, and, and you, you can imagine what the architecture looks like.
And, um, one of the steps on the, the test plan was, okay, yank the power from the, from, from from NK Power Watch, okay, for 1, 2, 3, 4 seconds converge. Okay, we're good. Right?
Uh, but the, the firewall state crossed over like, okay, we're good. Right? But yeah, we had those things like, but to your point, that was not the test, right?
That was one step of a very, very well-designed test plan for, but it was a fun one to do. I still get nervous anytime anybody tells me to do anything destructive on purpose. And I'm like, are you sure I'm not gonna get in trouble if I do this right?
It's in the plan. You're witnessing me do the thing that's in the plan because it, it feels wrong to purposefully break something. Oh, yeah.
But that should be replaced with this, um, joyous feeling of, Hey, my backup plan worked as soon as it happened, whether it's logging into Azure active directory instead of the local copy if something blows up or Yeah. You know, we can get that data back and we can do a fractional restore so we don't lose like three months of database tracking. It's called an escalation of a resume generation event to career ending event.
You'll quickly become the story that gets told at every conference for the rest of your life. Exactly. Hey, remember that time that Fernando did XI, I've, I've made a few mistakes that what, thankfully, I don't think any of those rise to that level.
I'm sure not. Yeah, sure not. But, uh, but to go back to, to, to the, to the topic, right?
I think we're all circling around this notion of people needing to be, uh, needing to, to, to have the space and the knowledge and the, and the, the, the, the systemic thinking around creating the plans and the models and the, the, the countermeasures and so on for what they think is, is realistic. And, um, one of the areas that, uh, that this is, and, and, and we need to be pushing ourselves to keep doing that. So we had the, the whole incident with CrowdStrike a couple years ago, right?
Listen, it was, uh, uh, very, very impactful as we all know, right? But realistically, at which point, how far do you go in your testing, in your assumptions before making a decision of, you know what, I'm gonna trust this. And if this blows up, oops.
Right? I don't know. I, I've, I, as somebody who did endpoint security, I'm, I'm, I'm, I'm always nerve with around cradle mode stuff, so, Well, I think it, it, it's an important question to ask because past a certain point, there's not much that you can do.
CrowdStrike was actually a really good example of that. It's like, what happens if the colonel faults and everything goes offline? Well, it doesn't matter what my plan is to bring that thing back up if it will not come back up.
Or we run into the other problem of resource utilization, right? If I only have a limited amount of time or people to bring the business back online, what should they concentrate on? Because, yeah, I don't know.
Let's say we have another AWS outage. Like, I can't fix that. I, I can do everything I can on my side side to make it work as well as I can, but at a certain point, you have to throw your hands up in the air and realize, this is bigger than me.
And, and no amount of resources on my side, short of an infinite money glitch will allow me to fix this. And, and that's the other thing too, because we, we've seen this time and again, with disaster recovery plans, like, well, what, what's our option? Well, we can have a warm site, uh, secondary data center where we're doing continuous replication over a private link, and they're like, yeah, that sounds excellent.
And then you hand them the bill for what that's gonna cost per month, and they like all the color drains out of their face because they're like, well, how important is it? You're like, well, if you want that to be a cold standby site, that is a forced data replication. And we're not paying the license for two active active storage units like the cost to go down, but the R-T-O-R-P-O is gonna go up because we are going to have to, you know, physically transfer media over there or something like that.
And it's the back to that trade off. Like, like, I can't plan for everything. I also can't pay for everything too.
And here we're 30 minutes into the conversation talking about the economics again. Ah, See, we almost made it through the whole episode without saying economics. Thank you guys.
But, but it's Right. But, but, but here's the thing. Our role is to work with our principal.
So the, the, the, the principal in, in, in economics is the principal agent problem, right? Uh, when you hire somebody, right? You want to make sure that they are doing the things that you wanted them to do, right?
And that you have enough oversight over them that they're doing it. I mentioned this in the context that, you know what, creating a, a, a, a full tolerant active active is such a pain. I'm not gonna do that.
You know what? I'm just gonna go and, and, and, and put a, a couple of my, put the server under my desk here and, and, and, and, and call it a day, right? Uh, if the person who hired me to do that is expecting that active, active recovery, and I'm not, and I just created a, a little server under my desk, that is a, that is a market failure.
That is a, uh, and the person who hired me has to have enough oversight, has to have enough knowledge of what I'm doing to, Hey, hey, Fernando, why are you keeping a server? What's, where's our, where's our active active stack? Right?
Um, it, it's a, it's not a trivial problem. I don't, I don't mean to belabor the point too much, but it's the idea that, uh, you need to have incentives play a part. And the same person who, whose face drains, because they don't wanna pay for the active, active, what are they, what are they promising to the people above them, right?
If I'm an investor in the company and the company's supposed to be bulletproof, and then, uh, that that director doesn't pay for the active active, I'm not gonna blame the engineer who proposed the active, active and didn't get funded. I'm gonna blame the, the, the, the person who didn't approve it anyway. Well, you would hope.
You hope. That's how it works. You know, the, there's kind of taking that same scenario, Fernando, of, you know, someone decided, I don't think I'm not into it today.
I'm not gonna do, do an active, active, the, the, there's also the how far do you take things, right? Because this is, these are rabbit holes. We can all go down continuously how far and, and the you reach points where now this is out of our control.
Um, but is there a remediation or an action that we could take if that happened, whether it's a SaaS service or a physical plant thing, or whatever it might be. Um, you know, if enough things cascade together that that's a scenario where we would not be able to handle gracefully. Um, and I think that's, that's back to your risk analysis, um, of okay, how these are the kind, we're up good up to here.
We think we're good up to here. And we know that there would've to be several things that could happen together in order for this to escalate further. Can they?
Yeah. Well, they, they sure could, you know, could be a Murphy syndrome, but, you know, we're gonna call the question there and say, that's the level of investment we're, we're worth, we're, um, willing to make, right? Yeah.
And I think that's the, the economic conversation that's in engineer owns is, let me lay out the landscape for you of here's what we need to address, what the, what the cascading rings, how far we can take it, and what's the risk to the business and the value that we would assign us, how much money we spend to protect from that some point. Yep. We're gonna risks that's hitting the earth is not one we're gonna Maybe, but yes.
But, uh, I know we're, we're running away. And I go back, that's why I mentioned during the, during the, the chat today, this notion of the quality of your decisions, right? We, we have to, to have the good enough decisions.
We do exactly as the said, Mitch, here's what. And beyond that, that's the role of the life. Well, I think we're gonna have to wrap it here.
It was a good conversation. Hopefully we have, uh, given you some food for thought about your disaster plans, and if your disaster plans didn't work out the way you wanted to, maybe this is your opportunity to have a meeting and kind of discuss that. Uh, and when you do, you should go definitely check out some of the stuff that my, uh, co-hosts are working on.
Fernando, what's something that you've got coming up that people should be paying attention to? So, I, uh, I just just finished a report on, uh, cyber Physical Systems. The, the, the, the public version.
I think it's out, if you're a return subscriber, you have access to the full report. Uh, I'm starting to think about my, my next one. And, um, at the same time, we have other reports coming in.
And as a matter of fact, I think that Tom, you and I have one coming up in the not too distance future. Uh, that's another signal report. So that'll be great fun to do together.
But, and, uh, prepping up for, uh, the RSAC conference and, and, uh, I absolutely love the, the, the conversations I have there, the travel. Yeah, a little, I love a little less, but, but seeing friends in, in, in San Francisco is great. Well, speaking of r yeah, speaking of that, you reminded me, I have to get my slides turned in 'cause I have a speaking slot, uh, at R-S-A-C-I.
I think one of the things you'll see, of course, I spent a lot of time, um, researching and talking about ag development or AI assisted development. However, whatever term you wanna use, um, and I've kind of coined it as this year, is the developer's role evolves to engineering agents in the way that software is created. That's really what the development roles versus coding, working on code.
I think a lot of the same kind of changes are happening in the observability role and world. And one of the things we've done is elevated in my practice observability. You're gonna see a lot more reports coming out from that, which is a nice intersection that, that Orlando, uh, Orlando, Fernando, and I get to work together on.
Oh, boy, there's a, you know, a winner thought, let's go to Orlando that Fernando and I get to work on together. So, um, it's, it's changing from being an operational tool to, it's actually moving left. Like we kind of talked about security moving further up the chain, particularly way AI agents get developed.
But even more so, the interesting, the interest is change changing. And that's some of the things you're gonna see, uh, coming out from, uh, our respective practices and working together. So I'm excited about that too.
We wanna thank you for listening to this episode of the Security Boulevard podcast. If you enjoyed this conversation, please subscribe on YouTube or your favorite podcast application so you don't miss any of our episodes. We'd also love it if you'd leave us a rating and a review and a comment to help the show grow.
com in the RUM group. com, text strong tv website, or the Techstrong TV app, which is available on pretty much every device out there. Now, make sure you're following Security Boulevard on X, Twitter and lvd.
Uh, there's a lot more content for you out there. Thanks for tuning in. We'll see you all next week.
Hey, everyone, welcome back here to Text Trump tv. You know, I, I, if I don't have this gentleman on every two, three months, I start wondering what's going on. So it's probably been about three months.
Uh, I want to introduce you to, well, you may know him already, but if not, meet Colton. Andrews Colton is the founder and one time CEO. And now CEO, again, at Gremlin.
Hey, Colton, it's good to see you, my friend. Happy New Year. Well, it's a little late, but I hope the New Year's off to a good start.
How have you been? Yeah, Always a pleasure to chat with you, Alan. Thanks for having me on.
Things have been going well. Always excited to come share the latest and the greatest or dive into the details on the, you know, on the technical side. Very cool.
Colton, I said, you're the founder. You were CEO then you weren't CEO for a while, then you came back as CEO. Um, but beyond that, prior to Grambling, you were at Netflix and, um, you know, that this is when Netflix was not that they don't innovate anymore.
They buy things now, obviously, huh. But, um, but you know, they were really an innovative technology company. They kind of pioneered the whole thing around like chaos, engineering and talking about scalability, right?
They kinda wrote the book, a lot of cloud native kind of, uh, functionality and projects were spawned out of Netflix. It, it was a great time to be there. Yeah, no, I'm, I'm really grateful.
I mean, uh, for those that, dunno, I got to do a, a four year stint at Amazon focused on keeping the retail website up and available. We built fault injection, chaos engineering tools there had a lot of success. But I remember being at a, a, a Velocity conference, I think in 20 12, 20 13, and, uh, hearing, uh, these companies talk about hearing Netflix talk about what they were doing on the resilience, reliability space, and Chaos Monkey had just come out and, uh, thought that was neat.
I thought they could have done more there, but, but it was a, it was a good first step, but they did a great job promoting it and helping people understand why it was important and really driving that innovation. So I was excited to go there. I showed up, you know, there was room to pitch in and help out.
I helped us get another nine of availability, helped us build some grade tooling, and really that's what allowed me to have the opportunity to go found Gremlin. Um, I was giving a talk at a conference and I'm running into some VCs in the lobby, and we went back and forth and they said, and I was like, I'm gonna bootstrap, you know, I'll just wait it out. And they were like, hold on, you could start tomorrow.
Let's get going. Uh, and three months later we found Gremlin. And coincidentally, uh, as of Sunday, it's been 10 years now, we've been out in market Another overnight success.
Yeah, yeah. Not, not what I read on Twitter back in the day. Well, that, you know, that's the deep dark underbelly, the whole thing.
All these people who think, oh, yeah, startup guy, you know, overnight success. No, this 10 plus years, it's not unusual to have that kind of effort in there. And you, you know, and you, but you're still a startup and still building and still learning and, and, and doing all that.
Um, now Gremlin's Mission has, I don't know if you wanna say expanded, changed, evolved. Chaos engineering is still obviously part of it, but chaos engineering now has, you know, it's, it's the life cycle of a product to a feature in tech, right? Uh, today's product become tomorrow's features as you, you know, March inex inextricably towards a platform and, and all of that good stuff.
So what, what's Gremlin about today? Carlton? Yeah, I mean, I think that's, it's such a great point.
Chaos, engineering, cool idea, fun project, not really an enterprise discipline that drives reliability. And that's what we learned early days is a lot of folks wanted to like, you know, kind of have fun, we'll break some stuff, we'll see what happens, we'll order some pizzas, you know, we're doing good things. And a lot of that didn't really result in the type of outcomes that they wanted because it was isolated.
It was not repeatable. And to your point about things becoming commoditized, things becoming part of a platform, things really becoming just how we build software. That's what we've seen a lot of growth in in the last few years, is people moving to, Hey, every team needs to do these basic tests.
It's just, you know, it's unit and, and integration testing for our distributed systems. And when we do it, we find things on a regular basis. We fix them before their issues, and we're just building better high quality software.
Absolutely. So what, what we've got coming out right now is just along the same lines, uh, you know, having grown up in a lot of enterprises, a thing that almost every enterprise company does is some form of disaster recovery testing. Hey, what happens if we lose a data center?
Hey, what happens if a WSU East U US East one goes down, or GCP loses a region or a zone? How does my software, you know, behave? Can I keep operating?
And a lot of companies do this in a pretty manual process. Hundreds of engineers, a whole bunch of prep that run it once or twice a year. And the, the, you know, they've gotta, they've gotta invest a lot of time that some companies, this is a weekend where everybody's on call and you gotta just show up and be on that call just in case something goes wrong with your software while we execute this large event.
So we thought, you know, this is a place that we could do better. Uh, and what we did is we took Gremlin, which is already good. Our customers are already using Gremlin for this purpose.
Some of the largest banks are using Gremlin to go do this kind of dedicated disaster recovery testing. But we said, well, let's make it easy to do the right thing. So we built, uh, built into the product the way to model this large scale event this way to have the right safe preconditions.
Let's make sure everyone's run it on their own first. Let's make sure everything looks good. Do we wanna run it, uh, one big event, or do we wanna spread it out over a couple of weeks and let people, you know, prove their piece independently?
Um, let's make sure we've got that halt button if things go wrong, you know, a way to clean it up and revert it so that we've got that safety net while we're running the experiment. And I think part of how we've grown up is the last piece is if you can't measure it and you can't turn back and present to the business or the auditors or the compliance folks, the evidence that you've successfully completed this, then it's just not as valuable. And so we spent a lot of time building really all of the, the right reporting and detail so that you can run this event, you can run it more often, uh, because you don't need as many people involved and you can streamline it.
But then you also have everything you need on the back end to go hand to the business to say, we've done our due diligence. You know, Colton, I think back to my days, you know, running companies that were, well, even here at ra right? You know, disaster recovery plan.
We're, we're totally sass. We, we have none of our own infrastructure per se. So a disaster recovery you plan is basically getting on the phone and begging, right?
Um, I'm, I'm being facetious, of course, but you know, in other companies that I was involved and we, we did have much more dr right? We were running data centers and stuff like that. Um, I wish we had something like this back then.
'cause as you said, it was quote unquote a manual process. And, you know, living here in South Florida, you would think everyone would have, what do you do in a hurricane as part of their DR testing? But you'd be surprised.
You'd be surprised, you know, oh, I didn't have that on my bingo card. Well, you're in south Florida, you didn't think a hurricane might knock you out. You might be flooded.
You might not have power for a few days. What do, like, you know, where was the thought process there? And it's, you know, so when we, Manuel is a, is a, is a good term, Colton, for covering up, just, you know, you can't cure stupid.
And, uh, there was a lot of, I've seen a lot of stupid over the years down here with, with people, you know, who thought they had their DR plans in, in, in, uh, place until, until the stuff hit the fan. Yeah. Um, Such a great analogy.
But It is. But, you know, so having something like this, you know, this, this is like a, you know, sleeping under the blanket of security, right? That, that kinda keeps you, you, you feel like you actually really do have it.
It would also seem to me, Colton, that this might be something that maybe AI can help with in terms of, of, uh, you know, setting up, what, what are the test parameters? What is your test coverage here? What should you be testing, what, you know, like, we probably don't test a lot down here for snow, right?
I, well, who knows, but how are you using AI or anything here to, to help with that? Or is this really based on real world experience that you now, you know, being able to scale up to multiple or infinite amount of, of customers? Yeah, so part of what we have in just built into the Gremlin product is a set of recommendations where if things aren't going right, then we're gonna tell you how you can go fix them, how to go adjust them.
Um, and one of the things we're always tuning and improving is just how do we make it easy for people to do the right thing? And that's really tell 'em what they should be doing. I think this is one of the things I've had to learn in my career early as an engineer, come from Amazon and Netflix.
Just assume everybody's done this a hundred times. They know what they're doing, just give 'em the tool and get out of the way. But truthfully, there's a lot of folks that maybe haven't done this, or maybe they're, you know, if they've got some junior folks on their team and they really need a little bit of guidance on the structure, on how to set it up, on how to model it, on how to run it, on how to interpret it.
Uh, so we've got some of that. We're always improving that, uh, within the product. So I'm not gonna say it's AI driven.
I don't want to be sensationalists there, but we've built in a lot of recommendations that tell people how to fix what goes wrong and what they should be doing. Absolutely. Colton, um, you know, thi this became available on February 3rd.
It's in ga, February A as of the third then, or? Yeah. Okay.
Yeah. I told my team we've had it ready for a while and we've had it in beta with some of our large customers. And one of my quality bars is have we run this ourselves in production?
And the answer's yes, we have, we ran it and we learned some things. We passed the test, but we found a couple things we didn't like that we went and fixed, which is exactly why you run these exercises to uncover those things. Absolutely.
You go fix 'em. So that was why, that's why it's ready to launch, because I said, engineering team, when you've run it and you feel good about it, that's when we're ready to go take it, then We're ready to go. You know, and that, and that, right there was the beauty of of Chaos Monkey too, right?
Because it's like a live fire drill. Mm-hmm. You know what I mean?
Until you actually go through it, it's hard to anticipate everything. You, you've gotta kind of have been there, done that kind of thing. And that's what Gremlin brings to this.
So it's an excellent thing for people who wanna go check it out. Colton, what's their best kind of on-ramp for this? com.
We've got links to, you know, we'll have it, we have it up on the main page. We can tell you details about it. We've got a free trial.
You can go in and play with it yourself. And of course you can reach out to me or my team. We'd love to not just show people the product, but guide people in how to go model these exercises.
How to run them safely and effectively and partner with them to make sure they're building not just, uh, a chaos engineering experiment, but a reliability program for their company. Absolutely. You know, too many people, Colton, treat VR as a checkbox, you know, cyber insurance company asked, do you have a DR plan in place?
Yes. Have you tested it? Yes.
You know, And that's as far as it goes. Often as a tabletop exercise or, you know, yeah, maybe once. Yeah, we did a year ago Or five.
You know, there, there is some of that too. 'cause not everyone is an Amazon or a Netflix. There's a lot of, you know, small, medium businesses who have a, a data center or some colo rack space.
Well, as you mentioned, if you're hosted in the cloud, as, as a lot of us are, uh, look, it was kind of a rough end of 2025 in the cloud we saw Yeah. Major outage with basically every cloud provider. So, yep.
The truth is, this isn't like wishful thinking. This is something that happens regularly and you can it into Anonymous. It's not if it's when Yeah, it's not.
If it's when I get you, Colton, thanks for coming on here and telling us about this. What do you call it? Is it called Gremlin Disaster recovery testing?
Yeah, gremlin disaster recovery testing. We, we went back forth on a name, we had another internal name, and we said, look, this is what people run is what let people call it. Let's go descriptive.
Let's just call it what it is. I love it. I love it, man.
Colton, good luck with this. I think it's a great product. You know, I've, as I said, I've firsthand knowledge on this kind of stuff of where it comes back to bite you in the butt, man.
It bites heart. You know, it'd be a fool not to do it. Um, keep up the great work.
Keep us posted. You know, the clock's running now, right? It's end of January or beginning of February, actually.
Um, we'll hopefully see you back here by April. Yeah, that's the plan. I love it.
All right. Hey, Colton, be well. Colton Andrews founder Gremlin here on Textron tv.
We're gonna take a break. We'll be back. Hi everybody.
Thank you for joining us on this series about AI in the mainframe environment. My name is Mitch Ashley, and I lead the software lifecycle engineering practice with the Futurum Group. Now, this is part of a three part series, segment number two, where we're gonna be discussing infusing intelligence using AI as a partner in in mainframe environment with mainframe teams.
Today I am joined by Anthony Ro. Anthony is senior director of, uh, architecture for AI at BMC software. Great to be talking with you, Anthony.
Thanks. Pitch. Thanks for having me.
You know, when organizations feel they're ready, kind of take that next step. How does generative AI start to make an impact in the mainframe environment? Yeah, so, you know, we start looking at AI as a partner.
Yeah, I was a developer for 30 plus years. Boy, I would've loved to have AI at some point in my career. Why?
Because from a developer's perspective, generative AI would've taken a lot of the toil out of my day to day activities. I could have used AI in a way where just the grunt work I had to do day in and day out during my developer developer journey. AI could have been a tremendous help.
I i, in that regard, take the friction out outta my day. When it comes to the AI ops space, as an example, I spent a lot of years building, uh, data visualization solutions, dashboards, et cetera. AI as a partner in that journey is how many times are we looking at operational dashboards and we're seeing blinking lights, or maybe we're not seeing certain things that may be in graphs or a data grid that's being shown.
And, you know, we're, we're, we're looking at it. It would be nice if you had AI sitting there as your partner observing the same thing you are and then pointing things out to you, uh, that you would otherwise miss. So I really liked that portion of generative AI making an impact day in and day out.
This is kind of tied to what we talked about previous, previously. If I'm a next generation mainframer and I'm new to the mainframe space, having AI as a partner in my daily journey, whether that's infusion in products that I'm using, or that's knowledge base access as we discussed, all of that constitutes that generative AI bubble to help me start my journey on that mainframe space. So that's really where I see it.
Helping folks on the mainframe accelerate their day by taking things off their plate that they shouldn't be worried about, and focus on the more important high value items and innovation that they should be doing day in and day out. You know, going from AI as being ready as an advisor now into AI as a partner, it seems like we're using ai, we're working alongside us as we're performing work, and it's providing information and, you know, giving us insights. Maybe that, as you mentioned that we didn't have, I'm curious if you have some kinda real world examples.
Regenerative AI is helping us. We do. I, I I, I, I, I mentioned I talk to a lot of customers and one of the things that comes up all the time is call ball, uh, code explain.
They're looking at, you know, they got that next generation looking at code bases, that a 25, 30 year plus. And there's a lot of complexity. There's code comments in there that are outta date.
There are module comments that are there in there that are outta date. Real world example, we've delivered through VMC. Any assistant that capability to do that code explain to do that.
Document generation customers are using that. They've given us really good feedback on the usage with that. That's a high value return.
Another area that we focused on was our AI ops space, where BMC Amy assistant was able to go in there and, and, and when there was a, a problem in the environment, what's the root cause of this problem? We have a lot of very sophisticated machine learning models and information that's presented to the user in that product experience. When generative AI came onto the scene, it was a great opportunity for us is we were able now to take something that was very complex, to explain and describe in the user experience, and have B-M-C-M-E assistant come in and then just explain what the problem is, what the root cause was in plain terms, not only that B-M-C-M-E assistant that was also able to give next step recommendations on how to resolve or prevent that problem from happening again.
So these are real world problems that we've gotten feedback on, on how the AI truly helped elevate the business value that we're delivering out of our solutions. You've talked about AI as a partner, giving us insights, helping us along, maybe triaging or diagnosing what issues might be. There's so much knowledge has gained in that process, but we're losing that knowledge with a lot of our workforce as they retire, move on, et cetera.
Talk about that knowledge loss and how a AI can help us with that. Yeah, so the knowledge loss is real as, as we all know. And, you know, an AI system is only as good as its knowledge.
And what I mean by that is not the LLM knowledge, it's your institutional knowledge. It's the knowledge that your staff carries day in and day out. How do we capture that?
How do we infuse that into an AI system such that the AI system has more context and relevance and becomes smarter for the users using that? That is critical. So how do we capture that institutional knowledge?
We, we, we augment or we have a facility within B-M-C-A-M-E assistant to capture that information and infuse it once it's infused. That's where the real power is. And it's not just a chat conversation where you're gonna go to BMC Amy assistant and you're gonna ask questions and you get responses back based on that captured knowledge, right?
We talked about that in the previous video that was around the advisor model. But, but imagine capturing knowledge and infusing in a, in a product experience, maybe it's in the AI ops space, capturing that knowledge so that when events happen and things happen within a dashboard or a report gets generated and there's alerts and exceptions and alarms that were generated, instead of just giving these generic type of, uh, insights out of the product, we can customize those responses and put them in terms that are relevant for that organization. And the way we could do that is by capturing that tribal knowledge.
I have to repeat that. I can't say tribal knowledge is capturing that enterprise knowledge that your senior staff has, and based on, uh, on processes and workflows they have when certain situations arise, we can infuse that in so the AI can respond. And what in a way that, that, that is relevant to your organization, for your organization to take the right steps that's relevant for you.
That's what capturing institutional knowledge within an AI experience to give you back responses in that regard. So as we move along the journey that you've talked about, you know, adapting, adapting and using ai, going from the stage of giving advice as we're working to actually guiding work, um, I don't know that we want to turn it all over to the AI all at once, right? There's a step along the way, maybe a hybrid kind of AI talk about what that approach might look like.
Yeah, that, that's a, that's a really good question. So at this point, we've been talking about language models from the AI perspective and augmenting that with institutional knowledge or real time data as an example. But generative AI itself is not enough, right?
And when, when you, when you're looking at, um, infusing AI in, in, in your solutions and delivering that value to our customers, and a lot of times it's a, it, it it's an, it's a, uh, a collection of different AI techniques that we need to pull together to formulate the, the answer or the response that we want to give. And it's this combining of different AI techniques together to formulate that response. That's what we call hybrid ai.
So as an example, I mentioned the BMC, uh, AMY Ops insight product. There's a lot of sophisticated machine learning models driving that product, that experience, but we augmented that with generative a AI for the explainability and the next step recommendations. So the, it's, it's the combination use of machine learning models with generative AI language models working them together.
That's a great example of a hybrid AI solution. So you can combine things like, you know, classic rules-based AI with language models or rules-based AI with, with machine learning models. What, what, whatever the application may be, it's this combination together.
That's the way we talk about it, is in the terms of hybrid ai. And that's a key element on our next video on how our AI agents think along their journey. Very good.
That'll be a nice incentive to watch that third part, I'm sure. Yes. Well, before we get there, any, any final advice you have for organizations as they're beginning to integrate generative ai?
So when it comes to generative ai, um, I, again, you, you gotta be really practical on your use cases that you want to apply with, uh, generative ai. Start with the explainability no matter what your area is or, uh, your domain is that you wanna, uh, uh, apply ai, start with generative AI with explainability. Once you get that sorted out, move on to that next step.
That next step is when you explain a situation, what's happening here? How do we do the next step? What's my next step actions, whether that's to resolve a problem or the next step action could be very, could be simply as notifying someone that a certain situation is going on.
But take those journey, that type of journey with your ai, do the, do the explain part, and then move to recommendations. The next step with, with generative ai and how you're gonna do that is, especially with the recommend part, that's where you capture that institutional knowledge that you have with your senior staff that you infused into your AI system. It could provide all those insights to the ai, ai, AI system on guiding it towards tho that output on those next steps.
Well, thank you very much, Anthony. It's great for you to share, uh, your experiences of your journey that you can help others along the way in that process. So you, you've joined us for our second segment in this three part series talking about infusing intelligence AI as a partner, generative AI in the main mainframe environment.
We're sure happy that you've joined us and we hope that you'll stick around and, uh, check out third segment. We'll kind of give you a little hint there of some of the things that we're gonna be talking about. It's about acting with confidence with we have AI as an agent of change that's making things happen for us and with us in this environment.
Thanks again for joining us with you on the next segment. Hey everyone, it's Alan Shimmel, founder, CEO here at Techstrong, and welcome to our continuing series on, uh, AI agentic AI in the future here with, uh, the Microsoft team and our FU analyst team as well as Techron. In this next episode, though, we're gonna be joined by Mitchell Ashley, uh, of Futurum, who leads the software development lifecycle and building segment at futur.
And Mitchell is talking with Brian Good, whose official titles is corporate vice President and agents marketing. But Brian is really here talking about agent apps and chat, and it, it's, uh, you know, obviously a very hot topic as we move to a Newent AI workflow based basis. So let's join Mitchell and Brian here with me, and it's great to have them both.
My name is Mitch Ashley and I'm VP of practice, lead of the software lifecycle engineering practice at Futureum Research. And, and Mitch, uh, my name is Brian. Good.
Uh, I lead the business applications and agents team, uh, here at Microsoft. Let's start here. 2025 is described as kinda an inflection point for AI adoption.
What do you think are the most significant changes that, uh, you've seen in organizations as they're using AI and agents in their businesses this year? Well, I absolutely agree. 2025 is really an inflection point, and we'll sometimes describe it as the year that the Frontier Firm was born.
And you might've heard us talk about the Frontier Firm before. It's this idea of companies that are putting AI really at the heart of their business, and it's enabling them to do things like reinvent the way they engage with the customers or transform their business processes inside their company, um, and beyond. And it's really the year these frontier firms are sort of rising up and a time when I think we can learn a lot from these early adopters and understanding how they're deploying ai, how they're being successful, and then figure out how we can take those insights and bring 'em to, to our own businesses.
That's where you can have outsized impact As AI transforms those functions. What do you see as the new patterns of work that's enabled by copilot and agents as customers start leveraging the technology for productivity or innovation? I'll tell you what, I have studied these frontier firms as they, as they come up, and there's really three patterns that I see across these frontier firms.
The first is really enabling, uh, employee productivity. So they give every employee, uh, an AI assistant like Microsoft 365 copilot, and it helps them, those employees be more productive. The second pattern that I see is, uh, really these frontier firms deploying AI to automate existing business processes.
So, for example, they already have a way of handling expense reports, but they can use AI to speed that up and reduce costs and, and that certainly results in some benefits to the customer. The third pattern is really where a customer like starts from the beginning, let's say from first principles, and they reimagine a function altogether. They'll reimagine what it means to engage with a customer who has a, uh, an issue with their product, and they'll put agents at the heart of that.
Most companies can only take on one or two of these functional transformation projects at, at any given time because it's a big lift. Again, it's reimagining a function from first principles. It's not just taking existing processes and and applying AI to them.
How far along do you think most organizations are in their adoption cycle for generative ai? Well, I think we're still in early innings, uh, that there's no doubt about that. Um, you know, customers are starting to deploy AI in different functions as we talked about, but certainly they haven't, in most cases, fully realized kind of the transformation that AI can have across their organization.
So it's still very early innings, but I'd say we see very promising signs and, and green shoots, uh, of that, uh, of that adoption and success. Uh, again, the key really here is to start function by function, think about a particular function, think about peeling it back to the business processes you need to go focus on. That's where you can have outsized impact.
How do you define AgTech business applications and why are they critical for organizations? Yeah, well, I do have a vision that there's, uh, the application of old is really transformed with ai and we call that new type of application, the AG agentic business application. And the ag agentic biz app includes the assistant for the human to start to use.
It includes prebuilt business process agents that basically take the drudgery out of work, and then it's built on a data foundation. And so instead of being limited just to the data that, you know, maybe is in the CRM system or the ERP system, it joins that with data from other lines of business systems and even productivity data so that you have a rich set of data that you can build agents on top of, and you can empower your humans to get better decisions from. And so this idea that brings all those together, the assistant, the agent, the application, I call the AG Agentic biz app, and I think it's a really big idea.
Well, let's talk about why you're optimistic about the future of agentic business applications. I have a lot of optimism here because number one, I see customers already deploying these applications, uh, and, and really transforming their business. So I already see people starting to get great benefit from it, but at a more human level, the thing that gets me excited is like, if you think about every one of our jobs, like there's a lot of stuff that we end up having to do that really isn't adding joy to our lives.
Uh, you know, me doing an expense report or, you know, filling out a CRM uh, system, you know, with an update from a customer call, those aren't the things that, uh, that make us, uh, unique. Those aren't the things that bring joy. Those aren't really even things that add business value collectively.
But if we can delegate those things to AI and agents, I think we'll enable ourselves to actually go do far bigger things and really take our ambition to the next level. And so I am absolutely excited for what the future holds. Yeah, I'd love to hear about how you see the role of AI agents evolving just over the next few years, especially as organizations start to redesign core business processes drive outcomes for efficiency.
How, how do you see this taking shape? Yeah, great question. You know, odd I, I would say that we really believe there's a spectrum, uh, of, of agents.
You know, there will be very simple agents that someone will use that might just be grounded on a particular knowledge source and help you, you know, understand, you know, what might might happen from that, uh, from that knowledge source. Then there'll be task-based agents, and then there'll also be more sophisticated, fully autonomous agents as well. And this more autonomous agents is where, where you can unlock a lot of business value.
These are agents that, you know, aren't called by a by a human. They're just running independently. They're triggered based on different actions, and they can really automate business processes in some cases from start to finish.
So we believe there's this spectrum of agents, and today, I'd say most ENT use cases start with those more simple kind of knowledge agents. Uh, but increasingly we're seeing customers bring in these task-based agents and autonomous agents, uh, to automate, uh, things, uh, end to end. How about the C-suite leaders?
How should they be thinking about investing in agent technologies for long-term value? Well, every C-suite leader I talk to today is already in on ai. Like, every one of them recognizes that it's a competitive advantage if they can move quickly.
And if they don't move quickly, they recognize it could be a, a disruptive force in their industry. But let's face it, ai, it's transforming businesses, it's transforming functions. It's, it's gonna reshape entire industries.
And every C-suite leader I talk to recognizes that. And once on board. What they're looking for though, is a partner.
They're looking for the tech, of course, they wanna make sure they've got the right tech, but they're also looking for a partner to help them shape this in their business. And that's where Microsoft, I think, comes into play. Uh, you know, we're not a large scale model maker, uh, but we do take the best of the models and bring 'em into the workplace, uh, so that companies can use them.
You know, we talk about AI being a disruptive technology, probably because it seems that it really can and will affect every part of our work, our lives, et cetera, certainly is affecting how we create software with new concept like agents, agent ai. Curious about your thoughts. Share with us, how does Microsoft view what the transformation is going to be like with AI really having an impact and a big benefit to businesses?
Yeah. Uh, many industry, uh, pundits will say that this move to AI agents is gonna, uh, lead to the rise of, uh, uh, more consumption like models or outcome-based pricing. And I think they're right.
I, that's definitely a direction that I see the world shifting as well. Um, the only thing I would, uh, balance that with is that many customers are still, uh, you know, most comfortable buying things on a, a per user or pre per seat basis. And so, uh, the approach I'm taking is, you know, how do I enable customers to buy offerings that they're comfortable with?
And that's typically like a per user or some type of per tenant type of type of license, while giving them the flexibility to, to grow and shift into more consumption models as they, uh, as their business changes. And so, uh, we are at an inflection point, just as we said, and I think one of the things that will change is the business model over time. Talk some more about ag agentic when you think about agents operating on their own or more autonomously, and why, why is that an important thing that organizations are looking to move to?
Yeah, it really comes down to business priorities. Like every business leader I talk to wants to find ways to grow their top line revenue, or they're looking for ways to automate, uh, things that, that they're doing so that they can save money or redirect folks to, uh, focus on more important activities. And, you know, autonomous agents really fit that bill.
You know, imagine in the sales context, you can have an autonomous agent, you know, going through marketing leads and qualifying them before handing them off to a human seller. That's work that wouldn't have gotten done in the past or would've been done by a human seller, uh, and have been relatively low value, not something they, they really enjoyed, uh, about their job. And so an AI agent can do that and add great value to the company, and again, help that company grow on the top line.
Talk about Microsoft, uh, and the offering that you have in the context of an end-to-end tool toolkit, if you will. Yeah. For functional transformation, you know, As far as the toolkit goes, we really believe that there's three essential parts to it.
The first is we think every employee should have an AI assistant in our case, uh, that's Microsoft 365 copilot. Uh, think of that as a productivity tool to help every employee work get through their workday and get more done. That can be quite transformative.
That next step up is where agents come into the picture, and I describe agents as really being for every business process or workflow in an organization. And we have, uh, a toolkit that enables customers to build their own agents. It starts with co-pilot studio, but even extends into our Azure capabilities with our Azure AI Foundry product.
So agents pair very nicely with that co-pilot that I described first, and then the final step is where the system of record becomes the system of action. And we'll sometimes call this the ag agentic business application, where you take a CRM system and you add agents and an assistant to it to really transform those three are the essential ingredients or building blocks for AI transformation, um, in a frontier firm or any business at this point. That's a great term.
Can you share some examples of how Microsoft customers are already seeing measurable impact from adopting agent business applications? Yeah. One example I I think I can give is lifetime.
Uh, they're a great customer of ours, uh, based here in the United States. They've used us as part of their finance and supply chain operations, and they've, uh, deployed agents to basically speed up how they handle and process e-commerce orders. In fact, I think it saved them 95%, uh, in terms of their order, e-commerce, order efficiency by deploying, uh, AI agents and agentic business applications to solve that problem.
So I think that's a great example. We also have, uh, another great example, uh, from Europe, um, uh, a large utility named Enco who deployed a multi-language, uh, AI agent, uh, for their customers to help them scale and address customer questions. Uh, it's a pretty cool solution, saves them time and again, helps them scale up.
Now, it's a, it's a really exciting time in the world of agents. You know, you talked about the C-suite and the kind of the three phases in this adoption curve as we adopt ai. What, what kinda recommendations do you have of like, how to get started?
Well, maybe near some of the near term activities? Well, you know, as I said earlier, I think AI is gonna transform every company, every function, every industry. And that opportunity is so vast, sometimes it's hard to know where to get started.
And I've got two bits of advice really there to, to anybody that, uh, that is, is pondering that question. Uh, the first thing is, again, start with a function. Pick a function that you want to go after and then peel it like an onion.
You know, go look at the next layer, which are the business processes in that function that you can apply, uh, a copilot or agents to, to really transform. The other thing though, is like, don't get into analysis paralysis. Just pick a business process.
Just go pick a business process to get started with. And as you learn from applying AI to that business process, whatever it is, whatever your thorniest business process is, go apply AI to it. You are gonna learn, your organization's gonna learn, your culture will adapt, and then it will flow from there.
So the opportunity is so vast. Don't let that keep you from getting started. You gotta get started, start simple, pick one thing and go from there.
So as we move to an agentic business environment, let's talk about the people. How do you see the role of human creativity, judgment, leadership? How is that gonna evolve?
Yeah. Well, that's an excellent question, and one of the things that we'll often talk about is how AI and frontier firms is gonna transform the way we work. It's gonna change, uh, the org chart into more of a work chart.
Uh, we sometimes talk about, uh, frontier firms taking on the Hollywood model where individuals will swarm around a problem, like focus on a problem and then disband, uh, you know, uh, when the, you know, once they've got a solution. And so it's absolutely gonna change the way we work with others inside the workplace. It also, I think, is gonna give rise to a new idea that we call an agent boss.
And you can imagine, just as a people manager today might take work and delegate it to different humans on their team, uh, you know, resolve conflicts and sort of manage performance. Every one of us in the future is gonna do that, but not just with humans, but also with agents. So imagine taking a business problem, breaking it up into pieces, delegating it to agents or agents, uh, on your team, resolving conflicts that might come up, applying human judgment.
Uh, it's gonna be an absolutely transformational moment, and I'm really excited about it. I, I really believe that, you know, there is still a huge opportunity for humans with human ambition to go do great work amplified by ai. Talk a little about, about how you see the D difference in the working together between applications, assistance agents, the different technologies.
Well, that's an excellent question. And, you know, we really have this complete toolkit that spans everything from the assistant to the agent, to the application. And I think those three things work together in a symbiotic way.
As an example, you can imagine that I might go to my assistant and, you know, ask a simple question, my AI assistant, like Microsoft 365 co-pilot and ask a question about, you know, uh, summarize the emails or help me respond to the emails I've got. Or I might use an agent to update a CRM record after I meet with a customer, but I'm still gonna want to go to an application. Really, that's a purpose-built experience for me if I want a more specialized or fine tuned experience.
So those three things really work together. Talk a little bit about the industries or maybe the business functions that you expect to see reshaped first by agent AI transformation. Yeah, there are three or four, I'd say functions in particular that are, I'd say, um, you know, ground zero for, uh, functional, uh, transformation with AI and agents.
Certainly customer service and customer experience is one that is, uh, absolutely being reshaped and say a very early adopter of ai, no doubt about it. Another, where I'm seeing a lot of early AI adoption, uh, is in sales. Uh, and it's just because, just as we talked about, there's a big opportunity to use AI to basically increase capacity for that organization to grow top line revenue.
So the business impact there is undeniable, but I also see it in places like finance and supply chain, where you can use AI to shorten the time it takes to, from an order to actually being able to ship it that order. We're seeing people use AI to improve accounts, uh, payable and accounts receivable. Um, so we're seeing some pretty interesting impacts there.
Let's talk a little bit more about the frontier firms. You know, they're not just adopting technology, they're also changing the way they're do doing business around AI agents and co violet capabilities. Can you talk about, you know, what they're doing to invest and support the short term ROI that they were looking to get for long-term innovation?
The most successful frontier firms are doing, actually is taking a functional approach. And so certainly they'll think about their entire company, uh, but they'll really take an approach that's function by function. They'll think about their sales function as an example, and then they'll peel back the onion a little bit and understand which processes inside their sales department they can automate.
Uh, using AI agents as an example, they'll look at which, uh, things their salespeople need help with, where you could pair up an assistant like copilot to help them throughout their workday. And they'll even look at how they can bring agents in to really augment business capacity, maybe to grow top line revenue or take takeout costs. But starting function by function has really been the recipe that these, um, frontier firms, uh, are using to see success.
As an example, in our sales organization, uh, by deploying copilot and agents, we've been able to improve revenue per seller in some cases by almost 10% in some organizations. And, you know, if you think about that, giving a seller 10% more capacity is like giving them an additional month, uh, in a year, uh, without actually having them spend any extra hours through the work week. And so there's some pretty remarkable results.
That's just sales. We've been able to do the same in customer service in our finance department, in our IT department. Our legal team has been able to reduce costs by 5%, uh, by deploying AI and agents, uh, within their, uh, within their functions.
How do we ensure trust and transparency as agent systems become more and more autonomous? Yeah. Well, there's two things that we want to do to help with trust and transparency.
The first is we have a rigorous set of principles on, uh, responsible ai. So for any AI that, uh, Microsoft deploys, we adhere to a an important set of guidelines. Uh, and that's really table stakes.
So that's the first piece, uh, responsible AI and our focus there. The second thing that I think is important is we now have an opportunity to start to quantify the value and the impact that many agents will have. And we often will call that in the industry, we'll call that evals or benchmarks.
And increasingly, I think we have an opportunity to help our customers understand how agents are being effective in their workplace using these evals and benchmarks on real world problems. So, for example, uh, we, uh, uh, a typical, uh, workflow will be a sales leader doing sales research to try to understand like how they might want to organize accounts or territories or, you know, reshape planning as they think about the year ahead. We recently released a new benchmark, uh, that we call the sales research bench, and it shows how agents can actually help sales leaders in that very specific job, and it's quantified.
Uh, and so I think you'll start to see more of that. And between following responsible AI standards and then using quantitative measures like evals and benchmarks to understand efficacy, I think we're really on the cusp of helping customers know how they can deploy AI in a safe way for real business results. Well, thank you, Brian, for sharing with us your insights and kinda look into the future, what's happening with the Gentech business applications?
Thank you very much. You know, it's really amazing the pace at which AI has been adopted and continues to evolve. The technology evolves on a near daily basis, but at the same time, organizations have to figure out how they're gonna implement their AI strategies and what the business outcomes that they're most important to their business.
I think it's very fascinating how Microsoft has approached the market, both from the standpoint of addressing the individual and their productivity, but also thinking about new workflows, new models of business, but doing that with not a set of tools, but a set of capabilities that provide integration with data process, workflow, and agentic ai. No one knows for sure what that future holds and what an agentic business might really, really look like. But in situations like today, when we remove constraints of what we can do with technology thanks to ai, that's where the possibilities are created.
You know, using tools like copilot, using applications like Dynamics 365, and we'll, we'll see what end users as well as technologists bring to bear in their ideas and how they reshape businesses for day and tomorrow. And how about you? But I'm super excited about the future that we're creating together thanks to working with technology companies and with end users like yourself.
Hey guys, thanks for the throw. We're here with Sashi Koran, who's the CMO for Nile, and we're talking about how autonomous networks are starting to emerge in the age of ai. Sashi, welcome to the show.
Well, thanks for having me, Mike. So just how are networks going to evolve? We've been talking about automation and networking for as long as I can remember, and, and we had software defined networks, and that was supposed to lead to this nirvana, but we never quite seemed to get there.
And now we're kind of looking at it in the age of ai. So explain, if you would, what's changing here? Well, Mike, um, we see a lot of that, um, around us as consumers.
You know, you have, uh, autonomous vehicles now. We have a lot of, uh, things that are becoming, um, sort of functioning on their own. And, uh, when you apply it more into the business side of things, and particularly into networking, this has been a long sought after marijuana for many, you know.
So I think, uh, it's interesting to see the, the journey that's happened, uh, on the path to autonomy, uh, over the last, uh, couple of decades. But, uh, you're right in terms of, um, really looking at it from an automation perspective. Um, and if you look at it in the context of the broader enterprise, uh, we've seen this journey with regards to a high degree of au autonomy and au automation happening in the data center where a few operators can actually manage, uh, a fairly, um, complex environment.
And, uh, we've started to see the journey happen in the wide area networks as well, um, with, uh, SD wan, as you might be familiar with. And that sort of gave the promise of, um, automation and management for number of, uh, devices. And, um, then it started to bleed into, um, you know, sassy and things like that.
And where we are at that journey is now into branch and campus environments. And if you start to look at, um, these environments, which is, um, you know, you could think of it as a bank branch, a university campus, a manufacturing entity, a healthcare institution, unlike the data centers, you know, this is where, um, you have users, you have, uh, devices, you have things. And so it's a far more complex environment to automate.
And, um, it also leads to, um, challenges with security challenges with, um, really making that network a lot more agile. And, uh, applying the notion of autonomous networking here is really a very big deal. And, uh, so that's sort of the journey we are on, which is to look at, um, incidents proactively look at patterns proactively, and make the network navigate to these ahead of anybody requiring to manually intervene.
And so, um, the more we can take away from manual intervention and manual configuration, manual provisioning, and go down this path of a hands-off operations, then, uh, the network becomes more resilient, more secure, more agile. And so I think we're on the cusp of a lot of these things. And at Nile, we're, uh, at the forefront of driving these autonomous networks today.
Mm-hmm. And how does that manifest itself? Let's say that I have an endpoint and I am trying to access some service on the backend.
Does the network know that there's a down server somewhere and that it just automatically reroutes it? Or how smart or smart? Yeah, so it's a combination of a lot of these things, uh, whether you look at it from a network performance issues, whether you look at it from network resiliency or network security.
I mean, all of these are things that happen in a real world environment. And the more intelligent the network can be, the more it has these patterns recognized more, you take advantage of, uh, artificial intelligence and bake these into the foundation of the network itself, then the operations start to become a lot more autonomous without the need for manual intervention. And, uh, for a lot of, uh, you know, C-level, um, stakeholders, their mandate today is to drive transformation.
But the mandate also is to do that while reducing risk and cost. So it's sort of a balancing act and, uh, autonomous, uh, networking sort of helps bring that balance in a healthy way. 'cause, uh, you know, you can move faster with agility, but you don't have compromises in terms of, you know, performance incidents, uh, reliability resiliency incidents, or more importantly, security breaches.
So a lot of these can be, um, uh, you know, navigated through this not notion of autonomous operations. And, uh, Mike, you've been covering the industry for a long time. Uh, you know, for every dollar that they spend on CapEx, there are six or $7 they spend in terms of operations over the lifecycle of the network.
And if you can really get, uh, to tame that beast of operational cost and complexity, it becomes much easier for them to move forward with their transformation initiatives. So that's really, uh, the crux of autonomous networks. It isn't just to bring, um, everything under automation.
It's how do you do so without making it brittle? How do you do so without, uh, increasing risk and complexity? And how do you do so while reducing costs?
So that's the, um, the leveler, you know, which we're, uh, achieving in, you know, customer environments day by day. There are, of course, a lot of types of AI these days. So is this more of the predictive AI that's then being used to reroute the traffic?
Or is it a combination of predictive, uh, generative and a little causal and all those things together kind of create this autonomous network? Yeah, it's, it's an evolution as with everything else. And when you put things into real world production environments, um, it has to make sense, you know, so I think we look at the maturity of a lot of these technologies as they evolve.
And, uh, at least with NS architecture, we're a bit different from, you know, box vendors that will put features in a box and then ship it, and somebody else is expected to own the service delivery. Uh, Nile actually, uh, owns that aspect of both the technology as well as the operations and the service delivery end to end. So we have a greater purview of what's happening across that entire environment compared to any other vendor out there today.
And so that gives us, in some ways an unfair advantage of being able to drive autonomy and to take these principles of, uh, AI and make sure that we're rolling these out with a great sense of responsibility. But, uh, architecture is also designed in such a way that it is not a, a la carte approach with, you know, different versions for different people when we, you know, roll something out, uh, if we recognize something with one deployment, all the deployment globally gets to take benefit from it because we have stopped standardized on that architecture. So this in turn, you know, brings in a greater degree of resiliency.
And, um, I think, um, month on month, year on, we are, we're becoming more mature, more autonomous as a result. Mm-hmm. What becomes of the role of what we know as the network engineer these days, uh, many of them have spent a lifetime using, you know, CLI tools to troubleshoot networks.
But, um, how does their position kinda evolve into something maybe, I don't know, less stressful or more fun? I don't know. Yeah.
Look, I think, um, I have, uh, been in this industry for close to three decades, and I know you've been covering this industry for quite some time. And, uh, we've always talked about, uh, making the lives of people that are configuring networks and infrastructure in general a lot more easier, you know, so, um, I think, uh, this journey that we are on, on the road to autonomous networking is really meant to augment their skillset, make them much more productive, focus on more strategic initiatives. 'cause let's face it, um, almost every organization, no matter how big or small it is, including Fortune 500 organizations, are really being challenged to do more with less.
Nobody's getting this flood of IT resources to grow their teams. And so the existing teams in many cases are, um, suffering from burnouts. And, uh, many cases, especially when you look at the branch and campus environments that I talked about, you don't have the IT staff in 90% of, um, you know, these environments, even for large companies.
So they end up outsourcing it, and that outsourcing is non-standardized, and they have to manage this, uh, operations and try to bring the consistency of the operational model. So these are challenges that hasn't necessarily been solved in the industry with the status quo that we see. And so, if you can tame that beast, particularly on these edge environments where you don't have the IT staff or you don't have the skillset, or you outsourced to multiple operators, if you're a global company with offices, you know, uh, worldwide, how do you kind of bring that and make sure that you have the degree of accountability for your organization?
How do you bring consistency? So we think, uh, you know, it's something that, um, not just the network, but network and security engineers, these two things we view as two sides of the same coin are just going to welcome it. And a lot of the response that we're seeing are from, um, network and security engineers because we're helping make their lives a bit more easier and, um, you know, the environment more reliable and safer for their organizations.
Mm-hmm. Do you think we'll get to the point where the network is much more secure than it is generally today in the age of autonomous networks? Because I'll be able to have, maybe there may be more transparency in real time to see what the nature of the attack is and respond, or is it just gonna be, you know, a fundamentally hardened network and it's just by definition more secure?
Yeah, look, I think, um, security is, again, a journey. And, uh, it's always a battle in terms of how do you outsmart your opponent. Um, and it's a combination of, uh, a defensive strategy as well as something that's more offensive.
And, um, you have to, uh, get lucky each time as somebody who's defending the network versus an attacker versus to sort of get lucky just once. Um, and so again, if you look at it, um, I, I look at holistically across the enterprise. Uh, we have solved a lot of these problems as an industry in the data center where you moved away from perimeter security to looking at security for applications, looking at lateral security for workloads and things like that.
And then we talked about SAS e in the, in the van, which has been a topic of conversation for six, seven years now. And, you know, different vendors and providers are moving towards that model. But if you look at the branch and campus environments, which is where the bulk of the users are, devices are, things are, it's still the, you know, wild west from the 1990s, and you cannot solve the security problem, thereby adding one more box or one more protocol, it needs a fundamental rethink.
And that's really what we've done where we have, you know, brought in these, uh, data center class security constructs into these distributed branch and office, uh, campus environments and, uh, espoused zero trust as a foundation. And, uh, you know, some of the things that we did recently, which we actually put out some news releases, were really, um, you know, getting into hacking environments with not a hardened model, but our standard offering that we, you know, give out of the box, so to say. Um, and, uh, we had a million plus hacking attacks and zero breaches.
We recently ran the entire network for Black Hat, which is a cybersecurity conference, 40,000 attendees, three plus days cap, world's largest capture the flag competition, zero incidents. So a lot of these are, uh, proof points of how we're able to build a zero trust fabric and layer in this networking and connectivity, um, on top of that foundation, which is radically different from every other approach that's out there in the industry today. And we think, um, you know, it gives a huge leg up to organizations that can take advantage of this architecture advantage of this network as a service, you know, delivery.
And this autonomous operations, the three of them sort of combining together to form a solution is a very, very powerful offering. Mm-hmm. As you kinda look forward into this coming year, are we maybe approaching a point where the whole notion of, you know, downtime caused by a network outage might become obsolete?
It's conceivable, um, you know, downtime can be caused by, you know, obviously, obviously software issues, hardware issues, but many times, you know, if you start to analyze these patterns, you see downtimes, um, being caused by manual errors, even security issues, if you look at it, a lot of them are due to manual errors, which are inadvertent. You know, they're not malicious by nature, they're inadvertent, they're accidental. And so the more you can prevent some of these from occurring by going down this path by things that are built on zero trust foundation, things that don't require manual configuration, things that minimize manual intervention, we, we see, um, the resiliency as well as the security going up in order of magnitude to be higher, right?
So I frequently say that if you add more links into a chain, you're actually making the chain more weaker, not stronger, right? 'cause the chain is only as strong as the weakest link. And so if we can minimize these kind of links that you artificially introduced because your architecture or your box configuration requires that then your, your, your, um, opening more doors for security breaches and interventions, and our approaches just been the opposite.
And so I think, uh, as long as we can, you know, be intentional about when you need to manually intervene, um, and how do you, you know, bring in compliance? 'cause a lot of the complexity today in enforcing compliance is there are just too many variables and unknowns. And so the more you can sort of bring those under your purview, then it becomes inherently much more secure.
All right, folks, well, you heard it here. Autonomous networking is coming and it will certainly be automated alongside many other things that we're gonna use AI for. The question now is how to get prepared for it.
Hey, Sashi, thanks for being on the show. Thank you. Having me.
All right. And back to you guys in the studio. Welcome to the Security Boulevard, the cybersecurity podcast from the RUM Group.
Each episode explores a variety of topics within cybersecurity and the technologies that drive it. com, security Boulevard, YouTube channel, Textron tv, and all of your favorite podcast platforms. Before we dive into today's episode, let's meet today's panel starting with Mitch.
Mitch, it's good to see you. Hey, great to see you, Mitch Ashley, I lead the software lifecycle engineering practice analyst practice at Futurum, kind, dabble in security, working with Fernando and team and software, supply chain, security, all that good kind of thing. It's good to have you and, uh, the other co-host this week.
Uh, Mr. Fernando Montenegro, Fernando, I, I Montenegro, I lead cybersecurity research. And, uh, we had, we barely started, there's a correction to be made.
Mitch doesn't dabble in in security. Mitch knows security, right? So, uh, um, it's interesting because there is a, a, like, as the industry's changing and whatnot, the, the, the overlaps between, uh, security and observability, for example, right?
Uh, he is amazing at this. And, and, and I rely on, on his judgment and knowledge, uh, many, many times during the week. Well, let's end this episode right there before Fernando changes his mind.
Thank you, Fernando. And of course, I'm Tom Hollingsworth, event lead for security at Tech Field Day, a part of the Futurum Group. Let's dive into this episode now as we're recording this.
The weekend was a little bit messy for most people. There was a huge weather system that tracked through the United States. Uh, there were power outages, uh, schools were canceled.
They even closed a couple of waffle houses. And for those of you who are not in the us, closing a waffle house is tantamount to shutting down the military, the government, and anything you've got, because it is the most reliable service that we have. But that made me think a little bit about the way that we plan for disasters, because a giant weather system may not be the kind of disaster that you think is, is going to cause a problem.
But what about a security disaster? What if, uh, you know, I don't know, some of your, uh, password files get leaked online? Or what if somebody manages to abscond with some of your data or crypto lock it?
Uh, how do you respond to that? Can you respond to that? And worse yet, is your response to that, oh, I better go check to see if my disaster recovery plan is up to date.
Because we all know that sometimes people don't think about disaster until it's upon them. So I'm gonna open this up to you, you gentlemen. Have we ever seen one of those situations in the past where it felt like the disaster recovery plan was making it up as we go?
Oh, I can, I could name a few names and a few companies, but I think I'd get in trouble if I did that. No, I think, I think, you know, it it, it's funny. We call it a disaster recovery or, or business continuity plan, or whatever it might be.
And, and it's sort of like any plan that you create, whether it's disaster recovery or a project plan, is out of date. As soon as you hit save, 'cause the environment changes, something's different, something's new. You know, the, the CEO comes to you and says, we're gonna do this now.
So it's, it's a little tough to create a fixed plan. Uh, so I think you're almost really developing a system, and the plan is just a reference implementation of your system, of what your disaster recovery process systems, people, et cetera, look like. And I think if you take that kind of an approach, you'll have a little more to borrow a term from, from Fernando's cyber resilience practice, you'll have more resilience into that system as opposed to, well, we did that in, uh, what is it, oh, 13.
I think we're a little outta date. We probably ought to update this. Yeah, a lot of things have changed since 2 20 13, so, Yeah.
Well, so, uh, the, we also had, so I up in Canada, and we also had weather here, uh, I here in the, in the greater Toronto area. I haven't checked the news recently, but it, it may have been like record snow storm. Uh, it's no accumulation, thankfully.
It's, it's, uh, relatively light and fluffy snow. So I can, I can deal with it. Don't you call our, I'm sorry to interrupt, Fernan, don't you call our, our kind of storm that we're having, you just call that Monday up there, don't you?
Isn't that Sort No, no, no. Listen, listen, I, and, and, and Ty back tying back to this topic, right? I think that one of the things about preparedness is that you always have to have a, um, uh, a threat model.
Like your threat model needs to have, uh, needs to be realistic to your environment, needs to meet your requirements, and needs to, to account for the things that, that you are, uh, willing to, to, to, to address. And in many places in the United States this week, uh, they're not, not statistically ready. It's not statistically, um, relevant to them to have preparations for this.
So my heart goes out to the people who have been affected by this. I know that, uh, um, I know that, uh, that, uh, it, it was very disruptive in many places where an inch of snow, an inch of ice is literally the community stopping kind of stuff. So, but yes, like the, the, the amount, these kinds of amounts up here in, in the greater Toronto area.
Yeah, that's, that's Tuesday, right? Uh, funny enough, uh, if you go further north in Canada, right? People who live further north think that about us here in Toronto, right?
So, so like, it's, it's fine. Like it's, uh, it's relative. It's all relative.
But, but, but Tom, to to your point on, on preparedness, right? I, uh, um, there is that, uh, quote that's always attributed to Eisenhower, uh, sometimes variations with, uh, uh, with, uh, Mike Tyson, right? I mean, that, uh, plans are worthless, but planning is everything.
And the other is everybody has a plan until they're punch in the face, right? Those are the, you can figure out which one's Eisenhower, which one's twice. But, uh, It, it's a really good point.
And, and to illustrate that, I actually wanna bring up something since, uh, Mitch mentioned 2013, uh, a lot of people lose track of the planning process because they never actually test their plan. I worked with a customer many years ago at a previous job, and, and we don't get a lot of snow storms in Oklahoma, but we get the other kind of weather that tends to level buildings, uh, especially in the springtime. Uh, we are in the middle of tornado alley, and, uh, I was working with a school and they had a solid backup plan.
Uh, you know, we're gonna put things on tape, and if something happens, we've got the tapes. And then one day they had to evacuate the main building because of a, uh, a tornado threat. And someone realized that their backup plan wouldn't work if the tapes were still in the building, if it was gonna get hit.
And they had to literally run into the building to grab a box of backup tapes and throw it in their car as they're evacuating. Fast forward a couple of years to 2013, and one of the largest F five tornadoes ever recorded, hit another school district administration building. Wow.
And they had to sift through the rubble to pull out the drives of the servers, to insert them into different servers, to be able to run payroll for the teachers to be able to buy supplies, to start rebuilding their houses and things like that. Because they never thought what would happen if we got hit by a, a severe weather event so big, it literally leveled the building. You don't think about those things until you're in the middle of a disaster.
And that's one of the reasons why you have to test your disaster plan, because you need to know where the failure points are. Oh, the generators will immediately fall over if the power goes out. Really?
Did you test it? Will they automatically fall over? Or does someone have to go hit the big switch to cause a failover?
If that's the case, who's gonna do it if everybody's pinned at home in the middle of a snowstorm? And how does, you know, things like DNS records work and what's gonna happen if somebody's logged into the servers when the power cut happens? Like if you don't test it and then jot down all the failure points, you might as well throw the binder out because it's useless to everybody.
And I spot on. And I would argue there's one step before that, which is what goes into the plan in the first place, right? And I think that that is an area where, uh, bringing this back to cybersecurity, we've had, uh, multiple cases where incidents have happened where it's not like, it's not that they were completely novel, right?
It was the case where, yeah, if you think this true a little bit, and oh, by the way, that can happen, right? And this is an area where I am, I'm, uh, I'm really interested in the work that's being done in adjacent areas, right? So in aviation, for example, there's an entire field of study on near misses, right?
So why did this almost happen right then? And, uh, I know that in cybersecurity, I know that, uh, Adam Tack has written about near misses in the past. Mm-hmm.
And I know that, uh, and I think that like, uh, Wendy Nater and Bob Lord, uh, uh, we were talking about CISA last week. So Bob was at, at CISA not too long ago. Uh, and, uh, I think that they're doing some work on near misses as well, right?
Which is how do we as a community learn from the near misses, right? This bad thing. Why did something horrible?
I mean, it almost happened what got us there, right? So, um, yes, very much about planning and very much about thinking through, sorry, where being creative about what can go wrong, right? It's risk management.
My kids hate when I, when I, when I talk about risk, I use it to tease them, of course, right? But, uh, it is risk management, right? And I think that, that, that our jobs in, in this industry is to elevate that risk management conversation.
Tom, to your point, precisely too, think through about these things. By the way, we do have an official mascot now, Nate, the cyber cat has joined us. So this is Nate, everybody good buddy of mine.
You know, it, it's, it's interesting just kind of my own experiences with disaster recovery plans. I think one of the takeaways for me was, you know, even a plan can be a bad plan, right? Have a plan doesn't mean it's a, a good plan.
And because you've tested it doesn't mean it's foolproof. That the things that worked when you tested it, tested it could still fail, right? You still could have some problem where the UPS doesn't kick in like it's supposed to.
The cooling doesn't happen like it's supposed to. But even more simple, more easy, easy kind of simple things you would think about, don't assume, like one of the, one of the plans that had to just completely revamp 'cause it was, um, really more than a decade old and how the technology was changed. Um, but when assumption was that such, and so employees live close to the building, and if this happened, if the com computer data center overheated, they could come open the door and cool it down.
Well, what happens when you camp the ice storm, the fire, the whatever? And, and that there were things like that, that we addressed and, and, and fixed and made sure that that wasn't the, wasn't the the main way we were gonna solve it. And then as it happened, within five years, and I think it was then another three years after that, we first had floods that went right up to the edge of the building.
Nobody could get to the building. And a couple years later they had complete fire. A whole bunch of houses got built down, businesses got bought, uh, burned down, um, right near our offices.
Again, can't get into it, right? Can't even get into the neighborhoods to do anything. So physical access is not an option.
So I think that's one of the things when you think about that, the, uh, risk analysis, Fernando, is think about what assumptions we make that we shouldn't assume that's true, even if we've tested it. Does that ring true? I think that one of the things that rings very true is that the world is unpredictable.
The world. Like we can't guarantee anything, right? And, um, one of the, one of the concepts that I learned along the way is that I love, and, and, and it's very, very relevant to this.
It's the notion of resulting, not sure if you ever heard of the concept of resulting. Mm-hmm. So, uh, I read about it from, so Annie Duke, she's a, she was a poker player and now teaches decision science.
And Annie Duke wrote a book called Thinking that, uh, basically how to make better decisions. And the idea of resulting is when people mistake the quality of the outcome with the quality of the decision, right? And we cannot control the outcome.
We can control the decision, right? But we cannot, and and you may have made a poor quality decision, poor that still turned out okay, right? Oh, I'm gonna, I'm gonna go on my tiptoes on thing to pick up that one thing over there as opposed to getting the ladder.
And I still fine, right? Um, I say that it's particularly relevant. I don't mean to bring football into this, but, uh, uh, as we're recording this, the, they just defined the, the play, the, the teams for Super Bowl 60, right?
So Seahawks and Patriots, uh, 11 years ago, they had, that was the fame, the fame, uh, uh, matchup. And any, duke uses the example of what happened in Super Bowl 49, uh, as an example of resulting because the, the Seahawks made the play at the one yard line that didn't work as people expected, and they were crazy about the result. But then when you look into the decision that was that the decision, that was a sound decision that led to that play.
Sure it didn't work because that's life. That's, that's football, right? That's, but, uh, sorry, I'm meandering as usual.
But it's this, this notion of what is the quality of the, the decisions that you were making as a professional. Are you, are you thinking through like the checklist that you need to dom, to your point, to make the plans, like to think through? Like how good is your decision process to create those recovery plans, to create those, those plans that, uh, do they match reality?
And I think the point that you guys are kind of bringing up that it's important to realize is that a lot of these plans are based on a series of assumptions that we need to make sure that we can validate. So here's a good one. Let's say you have some kind of a data protection system in place that is creating regular backups that are stored offsite.
What if you have to log into active directory to restore those? What happens if active directory is no longer available because it's been violated or it's been shut off or something? 'cause we're seeing that a lot now in cyber attacks, is that people are going after backups and people are going after active directory to halt any kind of, you know, cross system functionality.
Well then how do you get your backups back? Can, can you get them back without being logged into active directory? And those are the kinds of assumptions that you have to challenge.
Kind of to Mitch's point, you know, someone will always be able to go over to the building and open the door, but what if they're not? How do we plan for that? And I'm sure that you guys are probably sitting there thinking to yourselves in the audience, you know, oh, this is just a big tabletop exercise of how are all the crazy ways that I can, I can get around this?
Well, in security, those crazy ways to get around things are exactly what your attackers are thinking of, right? They, they want to try to hit you from an angle that you're not expecting, that you haven't planned for. I mean, I saw something today about, you know, people trying to get account, um, access by sending text messages or communicating with people, uh, through email going, Hey, this is, uh, at t we just wanted to reset your password.
'cause we noticed somebody's trying to log into it. Can you give us that code that you just got texted? It's looks legit, all the links work except, oh wait, that thing that you just got is actually gonna allow me to take over and then I'm gonna start hammering all of your other accounts for two factor stuff and whatever.
Like, those are the kinds of things that you have to be ready to adjust in your plan. What happens if an admin gets compromised? What happens if your CEO's email starts spewing out, um, you know, spam or, or attacks or something like that?
You have to think through that. And I realize that a lot of this feels like exception handling, like, you know, programming in a auto, an autonomous car. Like what happens if a 7 47 lands on the highway?
Like, you're right, the, the odd likelihood of it happening is low, but it's never zero. So like, you know, how do you To about this? Go ahead, Tom, about this is, it's not a linear process either.
The, you know, the, the, the, what you see is the attack may not actually be the attack. There may be a secondary action is, which is really to get you to start the backups. And that's when they're gonna compromise them because the system that you're, you're restoring, uh, to is, is actually the one that's gonna compromise and is gonna steal the data.
It's kinda like in, in, I guess in tank warfare, not to get too militaristic about it, but, uh, shells that penetrate tanks, it isn't the first, uh, contact with the tank that the shell penetrates it, it's really, that's just setting the stage to get through the explosives that happen that actually would then open it up to be penetrated by the follow on a secondary explosion from the, its, uh, shell that's attacking the tank. So it's, you have the added factor of what, not only what if, what this happens, but what if we're not able to do what our action, the response is what if we're suspect about whether that environment's compromised or could be compromised if we take the restorative action. So it, it's, it's a multidimensional kind of three dimensional chess, if you'll, and tying back to Star Wars and or Star Trek and Spock, And this is what, this is an area where like, I, one of the, the areas I work closely with is, is that, uh, part of the reason that the, the practice that that I run is called cybersecurity and resilience, is that I work very closely with, uh, with many of the, the, the backup and recovery vendors that are that, that work in this space, right?
And one of the things that they do is precisely this notion of, okay, what does a clean room restore look like in a scenario where we need to restore ad before we do anything else? As a matter of fact, a number of them are now tiptoeing their way into identity protection precisely on account of that. Mm-hmm.
Not to throw the AI into it as well. And they're using AI to help, uh, optimize that process, process. The, so it's, it's an evolution of, of, of, um, it's an evolution of the technologies, an evolution of the capabilities that practitioners have to review their plans and say, okay, alright, what is it that, what assumptions are we making?
Oh, we're assuming that we need to recover ad ourselves. Oh wait, our vendor can now do that for us. Okay?
Can we trust that the vendor can do so? It, the, the, it changes the nature of what you need to do. You're not recovering ad you are making sure that the vendor can recover ad for you, right?
So, uh, the evolving nature of the, the, of the disaster recovery plan is not only that things drift for bad, like I, but sometimes they drift for good. Like, I mean, there are new capabilities into your environment that you didn't have before. How can you make best users?
Right? You know, so fernan this optimistic Fernando, Fernando, question for you. You know, my thought is just like, you might use AI to analyze your business plan that you're, that you're creating, or maybe it's helping you create one, certainly ai it could be a good, uh, sounding board to run your disaster recovery plan through and test it, you know, put pressure on it.
Um, but also could be for scenario testing. Like what are the scenarios we're not thinking of? What are the latest kinds of scenarios we should start to consider how we adapt to, so AI could be your friend actually in helping you do a better job or keeping current with what's happening.
And you can rule out the, you know, far edge cases 'cause you don't think that's practical for you to even validate or test or protect against. But certainly you're not relying on just your thinking. It's like using an external consultant to help you.
Uh, absolutely. The, the challenge there, the, the him every, every episode at some point, AI comes in, right? Uh, uh, the challenge, the thing I I challenge people there is that, look, yes, I agree wholeheartedly, but who is running the AI and what are you, what are you using the ai, how are you using the ai?
How much are you depending the ai, are you a backup and recovery specialist who is interacting with the AI to ask, hey, um, like precisely as I said, to do the, Hey, let's, let's think through this, this scenario. What am I missing kind of thing. Or are you someone who is not as experienced in backup and recovery and you're coming to the AI for, oh, tell me what to do because I know, right?
And if it's the latter that is complicated because you dunno how to evaluate the, the quality of that outcome, right? As well as if you are an expert, uh, and you're just using it as a tool. It, we go back to the AI as a, as a tool for an experienced professional versus, uh, uh, uh, less experienced person using it and, and not catching the hallucinations, not catching the mistakes, not catching the, I say this as somebody who's using AI a lot, uh, for my, uh, uh, personal finance tracking, right?
I, it's great, but I know what I'm doing, right? I know. I, I think it's important to realize that a lot of the pieces that get missed are kind of institutional knowledge.
And I think where AI is gonna fall apart is that no one's ever documented this. Mm. And and that's one of the reasons why I know for a fact, that's the reason why I originally started writing on my blog was because a lot of the things that I was learning were things that maybe didn't exist anywhere else.
Like we all know the XKCD, you know, the infamous one, you know, who are you Denver coder seven, and what do you know? Because so many questions get asked that never get answered. And when someone does answer it in their head and go, yeah, that's how this works.
They don't ever write down what that answer is. And in the old world that was, oh, that means that I've gotta solve that problem. But in the modern world, it's AI doesn't have a knowledge base to draw off of to create a solution to that problem.
And so we, we kind of find ourselves with that knowledge gap, right? 'cause this is what I have been helped with up to this point, and here's where I need to be and how do I jump that gap? Because I promise you, an algorithm is not gonna be able to jump that gap no matter how many GPUs you throw at it, because it really doesn't know how to think outside of the box.
That's where humans are still valuable in the loop. What happens if these conditions aren't met? What happens if this really random exception occurs?
And like, you know, it could be something as stupid as like a race condition. The power comes back on, but the network isn't up and it starts timing out because it'll not connect. What do I do now?
We don't know what the answer is until we have tried it right now, before anybody else goes out there. Do not walk into your office tomorrow morning and just throw the master switch and go, let's see what happens. You need to kind of, you need to game through this.
This is what I want to happen. This is what I expect to happen. Let's test this in small scales to see what occurs.
And then let's analyze what we tested to make sure that if something does go wrong, it's contained and fixable before we try it on a larger scale. Because if you do that, just yank Let's see what happens. Uh, in the industry, we call that an RGEA resume generating event.
You, you do not want to be on the receiving end of one of those, especially nowadays. Funny, like I have so many things. First of all, uh, point of order, right?
Uh, Mitch, you're based in Denver, aren't You? Yes. Yeah.
Yep. And he knows, like, who knows, he might be that Denver Tom, who knows, right? Could be, Could be.
Could. But, uh, funny, funny you me, that, that throwing off the switch back when, when, uh, I was working on, on network security implementations, uh, I remember working with a, a healthcare customer where we built a, a multi-site, uh, uh, network with the firewall modules and switches and stuff like that and, and, and so on. And, and you, you can imagine what the architecture looks like.
And, um, one of the steps on the, on the, the test plan was, okay, yank the power from the, the the, from, from, from one, from the, the rack yank the power watch, okay? For 1, 2, 3, 4 seconds pf converges. Okay, we're good, right?
Uh, but the, the state crossed over like, okay, we're good, right? But yeah, we had those things like, but to your point, that was not the test, right? That was one step of a very, very well-designed test plan for, but it was a fun one to do.
I still get nervous anytime anybody tells me to do anything destructive on purpose. And I'm like, are you sure I'm not gonna get in trouble if I do this right? It's in the plan.
You're witnessing me do the thing that's in the plan because it, it feels wrong to purposefully break something. Oh, yeah. But that should be replaced with this, um, joyous feeling of, Hey, my backup plan worked as soon as it happened, whether it's log me into Azure active directory instead of the local copy if something blows up or Yeah.
You know, we can get that data back and we can do a fractional restore so we don't lose like three months of database tracking. It's called an escalation of a resume generation event to career ending event. You'll Quickly become the story that gets told at every conference for the rest of your life.
Exactly. Hey, remember that time that Fernando did x I've, I've, I've made a few mistakes that what, thankfully, I don't think any of those rise to that level. I'm Sure not sure not, But, uh, but to go back to, to, to the, to the topic, right?
I think we're all circling around this notion of people needing to be, uh, needing to, to, to have the space and the knowledge and the, and the, the, the systemic thinking around creating the plans and the models and the, the, the countermeasures and so on for what they think is, is realistic. And, um, one of the areas that the, the, this is, and, and, and we need to be pushing ourselves to keep doing that. So we had the, the whole incident with CrowdStrike a couple years ago, right?
Listen, it was, uh, uh, very, very impactful as we all know, right? But realistically, at which point, how far do you go in your testing, in your assumptions before making a decision of, you know what, I'm gonna trust this, and if this blows up, oops. Right?
I dunno. I, I've, I, as somebody who did endpoint security, I'm, I'm, I'm, I'm always with around stuff, so, Well, I think it, it, it's an important question to ask because to ask past a certain point, there's not much that you can do. CrowdStrike was actually a really good example of that.
It's like, what happens if the colonel faults and everything goes offline? Well, it doesn't matter what my plan is to bring that thing back up if it will not come back up. Or we run into the other problem of resource utilization, right?
If I only have a limited amount of time or people to bring the business back online, what should they concentrate on? Because, yeah, I don't know. Let's say we have another AWS outage.
Like, I can't fix that. I, I can do everything I can on my side to make it work as well as I can, but at a certain point, you have to throw your hands up in the air and realize, this is bigger than me. And, and no amount of resources on my side, short of an infinite money glitch will allow me to fix this.
And, and that's the other thing too, because we, we've seen this time, and again, with disaster recovery plans, it's like, well, what, what's our option? Well, we can have a warm site, uh, secondary data center where we're doing continuous replication over a private link, and they're like, yeah, that sounds excellent. And then you hand them the bill for what that's gonna cost per month, and they like all the color drains out of their face because they're like, well, how important is it?
You're like, well, if you want that to be a cold standby site, that is a forced data replication. And we're not paying the license for two active active storage units like the cost to go down, but the R-T-O-R-P-O is gonna go up because we are going to have to, you know, physically transfer media over there or something like that. And it's the back to that trade off.
Like, like, I can't plan for everything. I also can't pay for everything too. And here we're 30 minutes into the conversation talking about economics again.
Ah, See, we almost made it through the whole episode without saying economics. Thank you guys. But, but it's Right.
But, but, but here's the thing. Our role is to work with our principal. So the, the, the, the principal in, in, in economics is the principal agent problem, right?
Uh, when you hire somebody, right? You want to make sure that they are doing the things that you wanted them to do, right? And that you have enough oversight over them that they are doing it.
I mentioned this in the context that, you know what, creating a, a, a, a full tolerant, active active is such a pain. I'm not gonna do that. You know what?
I'm just gonna go and, and, and, and put a couple of, put the server under my desk here and, and, and, and, and call it a day, right? Uh, if the person who hired me to do that is expecting that active, active recovery, and I'm not, and I just created a, a little server under my desk, that is a, that is a market failure. That is a, and the person who hired me has to have enough oversight, has to have enough knowledge of what I'm doing to, Hey, hey, Fernando, why are you keeping a server?
What's, where's our, where's our active active stack, right? Um, it's a, it's not a trivial problem. I don't, I don't mean to belabor the point too much, but it's the idea that, uh, you need to have incentives play a part.
And the same person who, who's face drains, because they don't wanna pay for the active, active, what are they, what are they promising to the people above them, right? If I'm an investor in the company and the company's supposed to be bulletproof, and then, uh, that that director doesn't pay for the active active, I'm not gonna blame the engineer who proposed the active, active and didn't get funded. I'm gonna blame the, the, the, the person who didn't approve it anyway.
Well, you would hope. You hope. That's how it works.
You know, there, there's kind of taking that same scenario, Fernando, of, you know, someone decided, I don't think I'm not into it today. I'm not gonna do, do an active, active. The, there's also the how far do you take things, right?
Because this is, these are rabbit holes. We can all go down continuously how far and, and the you reach points where now this is out of our control. Um, but is there a remediation or an action that we could take if that happened, whether it's a SaaS service or a physical plant thing, or whatever it might be.
Um, you know, if enough things cascade together that that's a scenario where we would not be able to handle gracefully. Um, and I think that's, that's back to your risk analysis, um, of Okay, how these are the kind, we're up good up to here. We think we're good up to here.
And we know that there would have to be several things that could happen together in order for this to escalate further. Can they? Yeah.
Well, they, they sure could, you know, it could be a Murphy syndrome, but, you know, we're gonna call the question there and say, that's the level of investment we're, we're worth, we're, um, willing to make, right? Yeah. And I think that's the, the economic conversation that's engineer owns is, let me lay out the landscape for you of here's what we need to address, what the, what the cascading rings of, how far we can take it, and what's the risk to the business and the value that we would scientists, how much money we spend to protect from that at some point.
Yep. We're gonna take risks. That, that's the, the meteor hitting the earth.
That's not one we're gonna protect for. Maybe, maybe not, but yes. But, uh, I know we're, we're running away.
And I go back, that's why I mentioned during the, during the, the chat today, this notion of the quality of your decisions, right? We have to, to have the good enough decisions. We do exactly as your, here's what and beyond that, that's the role that life.
Well, I think we're gonna have to wrap it here. It was a good conversation. Hopefully we have, uh, giving you some food for thought about your disaster plans.
And if your disaster plans didn't work out the way you wanted to, maybe this is your opportunity to have a meeting and kind of discuss that. Uh, and when you do, you should de definitely check out some of the stuff that my, uh, co-hosts are working on. Fernando, what's something that you've got coming up that people should be paying attention to?
So, I, uh, I just finished a report on, uh, cyber Physical Systems. The, the, the, the public version. I think it's out, if you're a foot return subscriber, you have access to the full report.
Uh, I'm starting to think about my, my next one. And, um, at the same time, we have other reports coming in. And as a matter of fact, I think that Tom, you and I have one coming up in the not too distance future.
Uh, that's another signal report. So that'll be great fun to do together. But, and, uh, prepping up for, uh, the RSAC conference.
And, and, uh, I absolutely love the, the, the conversations I have there, the travel. Yeah, a little, I love a little less, but, but seeing friends in, in, in San Francisco is great. Well, speaking of RA yeah, speaking of that, you reminded me, I have to get my slides turned in 'cause I have a speaking slot, uh, at R-S-A-C-I.
I think one of the things you'll see, of course, I spent a lot of time, um, researching and talking about ag development or AI assisted development. However, whatever term you wanna use, um, and I've kind of coined it as this year, is the developer's role evolves to engineering agents in the way that software is created. That's really what the development roles versus coding, working on code.
I think a lot of the same kind of changes are happening in the observability role and world. And one of the things we've done is elevated in my practice observability. You're gonna see a lot more reports coming out from that, which is a nice intersection that, that Orlando, Orlando, Fernando, and I get to work together on.
Oh, boy, there's a, you know, a winner thought, let's go to Orlando that Fernando and I get to work on together. So, um, it's, it's changing from being an operational tool to, it's actually moving left. Like we kind of talked about security moving further up the chain, particularly the way AI agents get developed.
But even more so, the interesting, the interest is change changing. And that's some of the things you're gonna see, uh, coming out from, uh, our respective practices and working together. So I'm excited about that too.
We wanna thank you for listening to this episode of the Security Boulevard podcast. If you enjoyed this conversation, please subscribe on YouTube or your favorite podcast application so you don't miss any of our episodes. We'd also love it if you'd leave us a rating and a review and a comment to help the show grow.
com in the Futureum Group. com, the Techstrong TV website, or the Techstrong TV app, which is available on pretty much every device out there. Now, make sure you're following Security Boulevard on X, Twitter and LinkedIn at Security Blvd.
Uh, there's a lot more content for you to consume out there. Thanks for tuning in. We'll see you all next week.
Hi, everybody. So my name is Rashed. I'm at Fabrics ai.
I won't, uh, bore you with a long introduction. I, I help make things, I sometimes break them. And I have a, a few minutes to try to go a little bit deeper on what the, uh, uh, what we're doing with the middleware.
Um, uh, first of all, I just want to talk a little bit about motivation. So this, my, my piece is really about trying to convince you that we're very focused on, uh, one specific problem, which is how to take the agents from, um, from the prototype to production in a large enterprise environment. Um, I was, prior to joining Fabrics ai, I, I, I was an independent person building agents, and I, like everybody else, got extremely excited about what you could build really quickly.
Um, and then, like everybody else, I started running into the wall of making this work in a real enterprise environment. So when I joined fabrics, I was actually surprised to find that there was already something in place, uh, to help deal with that. And so I'm gonna talk to you about that.
Um, so ba basically, one of the, one of the, some of the, the key things that make this easy to, to get off the ground with an agent is you're, the main thing I'm gonna say is that you're working with it all the time. You're sitting there, you're talking to it, you're interacting with it. Every time something goes wrong, you pick up on it, you're able to, uh, iterate very, very quickly, and then you're eventually, without realizing it, kind of paint yourself into a corner where everything works great, but, but you've kind of guard railed yourself.
Um, and then you put it out in the wild, and then the variables change and everything, uh, falls apart. Usually in our experience, they fall apart for, uh, a number of reasons. So, what's on the screen right now are 16 examples of how things go wrong when you try to take a prototype agent and you try to move it into the enterprise.
Um, these are not the only things that go wrong. Uh, a lot of the things that go wrong go wrong with any agent, whether they're prototypes or deployed at scale. And so there's things you have to solve for if you wanna build any agent, you do have to solve for, how do I retrieve information in rag?
How do I, you know, do good prompt engineering? How do I, all of these things that we've kind of like as a community have been learning out loud and, and, you know, there's the buzz of the month about how to solve this problem, that problem, whether it's chain of thought and whatnot. I'm talking about things that happen when everything's working great, but then you try to scale it and it just breaks, and it, it basically breaks down into these three kinds of categories.
Most of them are gonna be in this context, um, uh, management bucket. And so when we're talking about context purity, which is something that was mentioned earlier, what we're talking about is we're talking about the fact that the LLM needs to get exactly what it needs to manipulate the statistical probability that it's gonna give you the output. You, you desire it, right?
So you are actually playing with a statistical system, and you're trying to game those stats. That's what you're trying to do. And there's a variety of ways that this breaks down.
I'll just pick a couple here. Um, one of them is gonna be large tool responses. When you're prototyping something, you've got mock data, and it's usually, you know, dozens, hundreds lines long.
So anytime you call a tool, you get something, it, the LLM understands it and then perceives, um, in a production environment, sometimes the tools will, will return enormous data. Sometimes they'll have to return data from thousands of remote systems, right? These are things that you typically don't play around with when you're prototyping, but you deploy, especially my backgrounds in network automation.
I'm, you know, I'm telling you, if you try to get an agent to work on your network, and you've got hundreds or thousands of devices out there, and you need to do something on that network, it's going to break unless you have a solution to handle that. Um, the other ones is gonna be about operations, and how do you operationalize this? And a lot of these, uh, gotchas are gonna be around observability.
As soon as you set it out there, you lose sight of it. And when you lose sight of it, that's when it misbehaves. It's kinda like letting your, your dog go, you know?
And they're gonna get into trouble if you're not watching them. Um, and then the last one is gonna be, uh, you know, operational infrastructure. It's gonna be things like the security people, those pesky security people that are gonna come after you, because you're not being careful with the data.
You're not being mindful of what the LL m's doing with it. Um, you're, uh, not, uh, properly segregating different users. You don't properly handle the fact that the person using the agent may or may not have the same rights as the agent itself.
And how do you handle any kind of differences that arise from that? So these are not corner cases. They're going to be the types of things that everybody who tries to scale an agent is going to run into.
Um, so it just turns, it just, you know, coincidentally, completely coincidentally, turns out we've solved a lot of these things. And, um, we've solved it by basically taking three core principles, um, about how to properly, uh, build agents for scale. The first one is, don't resist the temptation to give the LLM the whole problem.
Move as much of the problem into the tooling layer as you possibly can. Okay? And fortunately, uh, in a lot of use cases, there's ample tooling in place.
It's actually underutilized. I was at a recent conference where people were bemoaning the fact that network automation, for example, which is near and dear to me, um, is, has been slow to be adopted, but it's not for lack of tools. The tools are there, it's just that people are very resistant to using them.
Agents can be convinced quite easily to use those tools, right? But designing the tools, software practices, how to, where to put the, uh, the right levels of, of abstraction so that the tools are operating effectively. There's an art to it, and we'll talk about that.
The next one is curate the context feed. Try to tee up the problem for the LLM so that it always scores a home run. That means get any of the garbage that isn't gonna help it out of that context window.
Try to use context efficient formats, uh, token efficient formats when presenting, uh, the model that information. Um, try to not make the LLM be the thing that carries the data from one tool to another, if at all possible. That's a waste, right?
And then the last thing is gonna be all of the operational stuff. You have to not just observe the model. You have to observe the agent.
What does that mean? Well, the agent has a job responsibility like a person does. So manage that and, and look for outcomes.
Is it successful? Not just did it, was there a tooling error? Or, or did it take too long?
Or did it actually do the job successfully? You have to monitor it at that level, at the business outcome level. So the way do we do this, you, between this, we, we, we call this the middle, uh, the middleware, right?
Um, and you've already heard the terms context engine and universal tooling and connectivity engines is the two functional pieces. This is a functional diagram. This is not an architecture diagram.
It's showing you the function that the fabrics AI platform provides that sits right between the age agent layer and the tooling, uh, uh, enterprise data and system layer. Okay? And it provides this three basic, uh, capabilities.
Um, a task relevant context, um, so that we're, we're making sure that the context is gonna help the LLM be successful coordination of tools and, and connectivity of tools. What that means is we let the tools talk to each other, uh, at least to pass data to each other, rather than bouncing through the LLM. The LLM can say, Hey, tool two, I just called tool one, the output's over here, go get it and then run, uh, run yourself, right?
Um, and as well as the connectivity. So make it super easy and, and uniform. Uh, some of the, the, the, the questions earlier, we're talking about normalization of data and about, um, making, uh, the, the data digestible, uh, at correlation time.
So we want to be able to handle that, and that's what the tooling, uh, and connectivity engine is doing. And then the controls, which I've talked about, and I'm gonna show you a couple of examples. So the, Yeah, I'm Gina Rosenthal, digital Sunshine Solutions.
So the tools, so the middleware actually sits, is the first thing it sits on top of, are the tools. It, it collaborates with the tool, collaborates it, it communicates with the tools, and then that is what goes, that, that output of that is what goes into the infrastructure and the data. Would that be a true statement?
Yes. I'm gonna double click on the, on each one of these boxes, and I'll show you exactly what goes on in inside each one. Okay?
Cool. Thank you. All right.
So the first, uh, I, I should take a step back. Uh, the reason that fabrics AI is called fabrics AI is because we're, our platform is, is, is, is a trifa platform. It's basically what we call a fabric, is gonna be anything that connects many to many.
The data fabric is about connecting data to data. The automation fabric is about lacing together all of the automation processes and the policy management and all that stuff. And the AI fabric is gonna be all about the stuff that happens at that level of abstraction, where you're talking about things like, um, like, uh, reasoning and, and prompts and, and, and all of that AI stuff, right?
And so, two, the two components that I talked to you about, these are functional components that are, that, that can, that constitute, um, the, um, the middleware, the emerge out of that fabric, right? That, that fabric is what they are made of. Um, the context engine is lives inside the AI fabric, and the tooling and connectivity engine lives inside the data fabric.
So the context engine in a nutshell is this, it's this box that takes in all of the dirty, dirty context and cleans it up and provides pure signal to the LLM. Um, there's four main ways that it does this. Uh, one of them is, uh, intelligent caching.
So what that means is that when a, when a, when a tool result comes in or when, uh, you know, any kind of data comes into the context, um, it is cached intelligently based on how much data it is, if it needs to be converted, it's converted, uh, in, in, and it, there's basically a, a virtual file system in which it goes, and then that becomes addressable. It becomes addressable data rather than data that lives inside the context, okay? Mm-hmm.
Um, the next one is con conversation, compaction. So we, we, we, we heard earlier, there's, there's a lot of, there's a lot of, uh, mind share around compacting the context. And a lot of our tools, like Claude Code or whatever, they, they've got that built in the, as you go, and as the context grows, it just compacts.
It does this though, through summarization and summarization is very lossy. Um, uh, especially for long operations, you end up losing something. And it's hard to tell at the time when you're making the summary what you need to keep and what you need to throw away.
And if you make the wrong decision, you're kind of screwed. It's not gonna work. You've lost the information that you need.
So the, the conversation compaction, um, uh, technology that we have in the context, it's, we have a few clever tricks to make sure that the compaction maintains the relevant information for the current, uh, uh, conversation, um, dynamic. Yeah. Is that easier to do?
Because you're specifically working with tooling for enterprise applications, so you know what information you need to maintain when you do the compassion. So what we do is we have a, we have a, a a clever trick that we pull off, which involves understanding. So usually agents have a multi turn interactions, right?
Especially if you're talking to an interactive agent. There's this multi turn things that's happening. And so you might say, for example, you know, what's the capital of France?
And it might give you the capital, and you say, what's the weather in London? And it gives you that. And you say, and what about Paris?
Right? And, and it needs to know, is it asking me about capitals or everybody's asking, right? So the trick that we pull is that we're always looking at what the current, uh, goal of the user is, and then we go back into the history and we create a summary that's tailored to the current actual goal rather than the past.
And it, it gets a little tricky to do that, but that's basically what we do. And so you end up with a series of little tailor made summaries, so when you chain those together, you actually preserve all of the information that's relevant. Thank you.
Right? So I'm not gonna have time to go into this too deeply, but I do wanna point out, you know, that this, the, the context cache, it sits there in the middle of the context engine. It's this virtual file system where we put stuff, it lives there.
And, and, and there's been a recent, um, paper that was, that was published about, um, recursive language models. I dunno if anybody's read that, but recursive language model this is, is this idea that instead of getting an LLM to work with an gigantic document, it can spawn off, uh, a bunch of an, uh, an arbitrary number of subagents, and then assign bits and pieces of the, of the document to those subagents, um, through what, uh, in our EPL notebook and our, what that actually gives the, the whole basically benefit is that the, uh, the, a incredibly large document becomes addressable data rather than data that has to be completely ingested. You can seek it, you can profile it, you can histogram it, you can do all sorts of things to it to understand and feel around its edges, and then locate, pinpoint the bits and pieces of information that are going to be relevant to the thing that you're doing before you load it all inside your Contract.
What actually makes intelligent caching intelligent? What, uh, that's what are you, what are you calling the intelligence behind that caching? Um, so the, in the, the way that, well, there's different little bits and pieces.
One of the pieces is gonna be that, um, we're not necessarily gonna cache something that is small. That's a very basic decision that we're making. If it's small, just throw it in.
Um, but if it's large, then we want to cache it. And then, uh, what we want is, as we cache it, we glean certain bits of information from the data that we've cached, and we provide that to the LLM. So the LLM has a sense of what lives in the cache before it tries to work with it.
It's almost like vectorization. I mean, you're, you're extracting some dimensionality out of the data and, and providing some sort of a, a number stream to the, to the LLM so they can understand that it exists. And if he wants more data, go for it.
Something like That. Some sometimes like that. Yes.
It depends on the, on the type of data. Sometimes it's gonna be very numeric. So you want to, you want to maybe give it, uh, some sort of a basic profile.
What are your mins and maxes and averages? What is the histogram of the data look like? How many rows of data Do you have?
Perspective? Yeah, Yeah, okay. Okay.
With a pointer to the data in case they Wanna go back with a pointer to the data. And so you, you know, for example, that, you know, there's a, you know, 10,000 lines of data and you have a histogram that says that the first quintile ends at, uh, at, at row. If Quin first quintile of this particular column, value ends at this row, and so you're able to request, you know, maybe the last three items and the first three items of that, uh, data frame, you know, so all of the machine learning type of techniques mm-hmm.
We, we give those to the LLM through, uh, a set of fairly abstract tools. So I have a question based on this slide. Yeah.
Um, which is probably bigger than the question I wanna ask, but more to like, what is actually involved in your solution. So this context cache Yep. Is an external source.
Uh, so does, is that bring your own or do you guys provide a source, or how is that working? So we do, uh, our platform does have, you know, a, a data persister and data storage, but it could be external just as easily. Yeah.
Are you building a data graph DB underneath? We do have a grave grade, uh, uh, graph db. Yeah.
So we, our platform comes with graph. Okay. Uh, as well as object storage and, you know, so Your sources are becoming nodes and you're building all of what's building the edges around all of That, depending on if the data, uh, is structured that way.
Yes. Okay. Gotcha.
Yeah. Alright. So, uh, Before you go on, I think, uh, Marianne had a question.
Yes. Can you talk Marianne? Can you guys hear me?
Yeah. Now we can. Yes.
Thank you. Okay. So I had a question.
I see you have on your slides, uh, peer and purity. So my question was around what, what does contaminate your model? And what are the enforcement controls?
Is it isolation or something else? So usually what contaminates the context is going to be, um, uh, a bunch of extraneous information. So for example, um, take, take, uh, uh, a tool, uh, that its output has a bunch of introductory text and a bunch of, you know, closing text or a lot of formatting characters.
Let's say my tool returns HTML, right? I don't need the HTML tags, so why don't I just strip those out and just keep the text, right? So that's because each one of those little characters in an HTML document is one token.
Um, whereas words are, you know, oftentimes you can get like a long word where that's just two tokens long. Uh, but, but if you have an angle bracket on each side of it, that's another two tokens. And so keeping the, the, the token count low, uh, not only lets you put more information that's relevant into the context without reaching that context rot threshold where the, the, uh, the effectiveness of the LLM drops off.
Uh, but it also, you know, is reduces the cognitive load on the model as well. Does that answer your question? You don't look like I answered your question.
I guess my question was more around the controls. Uh, you talked about data context. Is there some automated governance controls or security controls that prevent that?
Uh, no. So we do have, as part of our platform, like as guardrails, right? That guardrails is one of the operational features.
So anything that goes, uh, prior to going to an LLM will pass through a guardrails model that will do safety checks and compliance checking that. But that is outside of the, uh, that is part of the security infrastructure of the middleware, but it's not specifically part of what we call the context engine. And so is that automatic automatically deployed to do that in seconds, or is that, where does that fit in that process in your millware?
Yeah, it, it does introduce a little bit of latency between the prompt and the, and the LLM receiving it because it has to pass through. But those models are t typically very fast, and they're, uh, specifically trained to, to, uh, to find, um, you know, uh, non-compliant, harmful type of instructions. Okay.
Thank you. Yeah. I have a question, but I don't wanna subvert your presentation.
I think what some of us are grappling with is, okay, you have a bunch of slides saying you've been there, done that. What's the customer get as a deliverable code? Consulting support pilot projects?
Yeah. We're a product company. Yeah.
Uh, so our, and so our product, uh, you know, we, we, we deliver that product in a variety of ways. But, uh, I would say that, uh, oftentimes what we're delivering is agents that are going to be, uh, operating inside of a, you know, enterprise and they're gonna be effective and scalable. Uh, but for some customers, what we're doing is we're saying, here's a platform in which you can develop your agents.
And what we're saying is, as opposed to some other platform that's out there, um, this, this is going to provide the facilities or all of the services, the platform based services that agents need in order to be scalable and reliable in, in an enterprise. And I would think, from what I've seen of other products, not in this space necessarily, but folks might want a pilot project where you demonstrate how to use the toolkit on some sample problem, maybe one or two. Yeah.
And then maybe some consulting support as they go along or the Yeah, Yeah. We always help. We're always, you know, uh, we always get in there.
Um, but our business model isn't primarily consultation. We want to give people, enable them to mm-hmm. To do stuff.
So rights to code, but not so much services that's doesn't scale. That's, That's right. Yeah.
And we have a very extremely low code type of approach. I'll show you an example of that, uh, in, in a slide that's coming up here. So just before I run outta time, I'm, I, I may not have time to do, uh, the demos here, but I'm, I'm going to, uh, talk to you about the connectivity engine a little bit because it's the other important piece of this.
Um, this is the part that, uh, mentioned earlier that's connects to the tooling, right? So, uh, tools, uh, MCP tools, for example, they sit out there on a server. And MCP was a great innovation because it creates this level of abstraction around APIs.
Everybody loves them. And as we mentioned, however, they're in an enterprise environment, there's still plenty of tooling out there that isn't, um, wrapped by an MCP uh, server. But even if it is, right, you have this problem, I, let's say I have an MCP server that's a MCP server that stands in front of my SQL database, and I have another MCP server that stands in front of my, uh, my, uh, my inventory system.
And I somehow need to update, uh, do something that involves information from my inventory system and information from my SQL database, while great, there's these two SEP tools. So the LLM can easily call the tools that are sitting there, but if anything needs to happen here based on something there, the LLM has to suck in all of the responses. Understand that, construct the data that it then pushes down to the other MCP servers and you just wasted a bunch of tokens, essentially performing a shuttling process from one thing to another.
Mm-hmm. Right? So the fact that we have this MCP, um, the, the, the, our particular, our, our, our engine over here, the, the, the tooling engine is, it creates kind of like a, a, uh, a net, uh, an MCP abstraction layer.
It's itself an MCP server that wraps other MCP servers and non MCP based tools. And because we do that, you get this joint and common execution environment for them. My tool results come back and they're in the tooling engine, and therefore I can move them to the next tool as an intermediary without ever touching the LLM.
I just let the LLM know that the tooling, the tooling results are there. Right? That's one thing.
Another thing is that, um, we have this YAML based way of defining tools because we provide some primitives inside of the, the tooling engine that pretty much lets you develop any tool. And through this, uh, YAML based abstraction, you can easily develop new tools and you can take a, a tool that might be, for example, SQL Query tool. And a lot of people, they love this when they're building a prototype.
They say, look, I hooked it up to an MCP server, it has a SQL Query tool, and now the LLM can write SQL queries. Awesome. Except it breaks on production because building SQL queries, if anything goes wrong, uh, it doesn't work.
And the LLM, the bigger the job you give it, the more likely that the SQL query it will build for you, will be wrong somehow. It's too complicated. It can't come up with it at inference time or whatever.
When you have, uh, this, um, this, uh, middleware layer, you're able to define a tool for specific types of queries that are going to be common in your enterprise workflows. And now it's just simply calling a tool that has a name and a bunch of parameters rather than constructing an expression that's brittle. Right.
And we can actually, you know, you can talk to the system and have it generate those for you. We have a, a thing that makes the tools for you, so you're gonna have to bother coding all that up. It'll do it because we have something called dynamic data discovery, where you can say, I'm interested in doing these things.
The system will connect to your MCP server or to your data source, and it will not just consume the entire API, it'll explore that API and figure out exactly where the data lives. Mm-hmm. And it'll skip the tables that don't have data.
It, it will focus on the tables that do. And, and, and, uh, it just generally means less responsibility for the LLM at inference time doesn't have to query the wrong table and say, oh, I've got zero results. Why is that?
Oh, there's another table with the same data every single time it runs. Right. You mentioned early on that you were doing sort of data science type analysis Yep.
With the intelligent caching. Yep. Do you do similar type of data science work in this layer as well?
Yes. As you explore sources? Yes.
We, because data science is in our DNA, so we do, we, we do, uh, schema discovery and we, we code our agents when we write our prompts and all that stuff to do exploratory data analysis as just a matter that they work, always do EDA every time. And we give them the tools to do EDA because when you do your data analysis, you create shortcuts to good information that you otherwise would, uh, would miss. Great.
Thank you. Okay. Will you guys, uh, like to see the demo or actually we have a couple case studies.
Case studies would help me understand that. Yeah. Alright, so I'll, I'll just, I'll just, uh, I'll skip to a couple of case studies then.
Um, so two, I have two, I have two case studies for you. This is a customer, they came to us. They had, um, uh, a lot of documents.
Uh, this is actually not an IT or a networking use case. It's a straight up, we have an enormous cache of documents. Each document is very, very large.
We need to be able to, uh, to query those documents. We currently have a team of 12 people, and it takes them a long time to, uh, find the documents that we're looking for. Can you do this with an agent?
Somebody said, yes, sure. And they built a prototype and they did the demo and they said, this is fantastic. And so they tried to use it and it broke almost immediately, right?
So then they came to us and they said, can you do better? Um, where we said, well, we're certainly gonna try. And so, uh, we built, uh, this using our middleware with all of the features that I've been talking to you about.
And so some of these questions that they had, like, you know, um, we're very basic questions. How come for one time I tell you to find the matching documents? You gimme 12 results and sometimes there's only three.
Uh, this is the kind of stuff that drives people crazy when they use agents because it's probabilistic and it's, you know, not always the same. Well, we want to get to the bottom of that. So a lot of the techniques that I've been talking to you about, we developed while learning how to do this right.
Um, and, uh, and, uh, so we were able to, you know, deliver consistent results, um, and convince the customer that basically with the agent in a few minutes, they're able to get the results that are better than what their team of 12 people were usually able to accomplish in a two or three hours. Okay? So that's, that's one use case.
And it's involves like, enormous amounts of large documents. When you're prototyping, you would never use, you know, 50, 60,000 token documents with lists of thousands of them. Where when you try to find matches, you get 150 matches.
Like how do you, how do you handle 150, 60,000 token documents in an LLM unless you've got some sort of a context management strategy. So that's what, that's what we're doing. The second case study is, um, is a case study where the, the customer had an enormous number of systems.
This one came up a little earlier in a preview presentation. So it's a very heterogeneous environment. Each one of these tools provides an amazing dashboard that lets you look at their, their information, right?
And, um, and so they're like, don't replace, um, there was a question earlier about data lakes. We can do a data lake, but they didn't want us to do a data lake. They said, no, no, we have, we're very happy with the job that each one of these tools is doing, gathering information, syn synthesizing it.
It's just that in our business, sometimes issues come up that span these domains. And right now I have to go and sit in front of one beautiful dashboard after another and correlate between them. And a person has to do this work.
Can, can an agent do that? So we took that on. Uh, and, and we, we, we built, we built an agent.
Um, and we basically, we, it's what I was just talking about was the, the, uh, dynamic data discovery is what led us do this a little bit better than just federating all of the data and ingesting it into our platform, which is what a lot of people would say is they say, oh, the data exists in all these places. Why don't you just pull it all in and, um, and into our system, and then we can do all the analysis. Um, certainly there's times that we do that, but with our platform, you don't have to do that because the, uh, the, uh, middleware and the, the dynamic, uh, schema discovery is gonna be really good at, uh, understanding where the data lives and then going to get it when you need it.
There's another thing that I really want to show you guys. I, I don't have time to to to do the demo, but I do want to talk to you about the fact that we take a very different approach to, uh, evaluations as well. I, I mentioned this at, at the top right.
Um, a lot of times you'll see, uh, observability platforms that are simply, um, evaluating how many tokens, how, what the latency was, uh, whether there are any, uh, tool errors or anything like that. Um, but we have this really nifty system, um, that, uh, once a a agentic session goes idle, we'll go back and review the entire session and perform a qualitative analysis of it, and it'll pull out various themes, what the topics that were covered were any lessons learned, any, uh, optimizations that are possible. It creates these reports.
It scores the agents dynamically, it considers any user feedback as well as any interactions. Uh, and then, uh, it creates, um, an improvement strategy, kinda like a pip for each of the agents that the, uh, agent administrator is able to click through and apply on the go to update instructions, uh, over time and, and guide the agents to, uh, better performance Institutes. The first time I've heard putting your ais on a pip.
Yeah, well done. Well, you know, because we're giving them jobs that are human-like jobs. And this is kind of like, um, not only like the, the first point I mentioned where I say push as much of your job into the tooling layer as possible.
Mm-hmm. But then there's still like, what are you asking the agent to do? And I, I really think what you want to ask it to do is something that you might ask a person to do mm-hmm.
That they don't wanna do. Like it's a person like activity that sits low on the person, like enjoyment scale. Right.
Get the agent to do that. Right. And then measure them And then measure them like people, did you do a good job?
Do I need to train you? Do I need to give you additional information that you didn't have before? Right.
Yeah. Okay. And I think, uh, what, maybe since I have 15 seconds, we do have spend management and cost controls, uh, built into the platform, which is sometimes people forget about that.
Agents end up costing you hundreds of dollars. It's kind of like, uh, back in the nineties when the kids used to call the party line, um, we, we have controls to help, uh, hundreds. Where are you getting Your cheap Ais from?
Thank you very much for.