Techstrong TV February 26, 2026
Solving the AI Confidence Gap: OmniGuard AI CEO Kobi Tzruya explains why most enterprises hesitate to deploy AI agents in production and how intent alignment, real-time monitoring and root-cause analysis are critical to closing the gap between experimentation and operational trust.
Customer Success at Scale: Clay Wesener (Microsoft) shares how organizations like Wells Fargo are modernizing regulated workflows with Copilot Studio agents and Power Apps—offering a blueprint for scaling intelligent applications across the enterprise.
Prompt Engineering Risks: DeepTempo’s Mayank Kumar outlines how prompt engineering is expanding the enterprise attack surface, introducing model manipulation and indirect injection threats that demand stronger governance, observability and policy controls.
Agents of Dev Ep. 11: Mitch Ashley and Brad Shimmin unpack the “Open Claw” moment and the rise of observability-native software, exploring multiverse control planes, automated data product factories and new approaches to managing non-deterministic AI systems.
Agentic Business Transformation: Microsoft’s Bryan Goode joins Mitch Ashley to discuss the shift from systems of record to systems of action—where AI agents embedded in Dynamics 365 and Power Platform proactively drive outcomes and redefine operational strategy.
Bulldozing the Memory Wall | Predict 2026: Brendan Burke examines how co-packaged optics, extreme hardware specialization and platforms like NVIDIA Rubin CPX will deliver a 20x leap in energy efficiency—reshaping AI infrastructure around memory bandwidth and intelligent workload optimization.
Transcript
Hey, everyone. Welcome back here to Tech Drunk tv. You know, I'm, I'm in my studio at the house today, and I was really happy to have my next guest on here.
He's a friend of mine. I've had the pleasure of interviewing and speaking with him and presenting with him on stage couple of times at the RSA conference, which will be next month. Um, I want introduce you to my Frank Co Zaia.
Yeah. Say it better than me. Yeah.
Kobe, how do you pronounce it? And I got it now. I got it.
I gotcha. Kobe, you know, the last we had checked in with you, of course, uh, you were at check marks in a, in this, you know, c-level senior role there. But you know, this ai, this AI bug, if you will, the excitement around it has so many of my friends who haven't, they haven't coded in, in some of them in 20, 25 years even, right?
They've been executives, they've founded companies, they've had exits. They are in love. I don't know if in love or in lust with the ability to just generate code again and make apps, whether, you know, using Claude Code or, or any of Theis.
And, and they all have great ideas. You know, for many of them, they, they thought they were done. They had founded their last company, they had their last great idea, but they're, they've got new ideas.
They're founding new companies. This same, this same fever's gotten to you, it sounds like, right? Yeah.
Tell us what you've been doing since you left check marks. Yeah. Okay.
So, uh, you're right. Uh, kind of this, uh, fever Hit me, hit me, and, uh, I know I need to do it in order to clean it out. You got No, you gotta, you gotta fulfill it, right?
You gotta, yeah, Exactly. I, I, I have to fulfill it. Um, I founded a company named the Omni Guard ai.
Uh, Omni Guard AI are the, the main, the main problem that we're solving for is to make sure that, uh, that AI agents or, uh, or applications that are using, uh, that are using ai, uh, LLMs, um, enterprises and companies that develop them, will have confidence that they actually do what they intended. So, we'll, not, let's say, you'll not develop an agent, and, uh, it, uh, if you're a car dealer, and it will sell a car for $1 or, or, or, or you'll not develop an agent that will, uh, that will, uh, that, that will send illegal advice to, to your, uh, uh, to your, uh, to your, uh, to, to your customers or, or will be biased or, or anything like that. Meaning we look at the behavior, okay?
Uh, at, at, at, at, at the behavior of, uh, of, of the agent or, or the agent application. We see if the intent, and we can identify the intent and the intent, and also the workflow, uh, that the agent or the, uh, uh, or the workflow needs to, to work by. It doesn't match the output.
And if it, if there's a deviation, uh, we can actually either alert or stop the response, uh, or, or, or bring a man, uh, or, or, or, or, or, or bring a man in the middle in order to, uh, in order to look, and we can also identify the root cause, meaning we're not only giving you the headache, we're also telling you okay, where it comes from. Excellent. So that, that's, that's the vision.
This is what we're, uh, this is what, uh, this is what we're working on. And it's a real problem because what we see is that, uh, you know, when you ask people, are you building agents, uh, like 80% of them, uh, 80% of them, you know, pick a survey, uh, will pick a survey, 80% will say yes. And then when you say, okay, is it in production?
Then, you know, the percentage drops to 10 to 20%, and this at best, yeah, at best. And the confidence in these agent applications or, or AI agents, it is, I think the root cause, meaning is, is the root cause. Absolutely.
Co Kobe, I, I just wanna make sure we got the name of the company. Can you spell it for us? It's Omni Guard ai Omni guard.
It's, uh, only om O-M-N-I-G-U-A-R-D. Guard. Like guardian, okay.
Yeah. Okay. Look, the way people talk about it, sorry for my accent.
No, I would, no, no. The way people talk about that, no, For my accent, you know, English is not my first language. That's Okay.
No, but the way people talk about God, I wasn't sure if you spelled the GOD, you know, so I wanted to make sure we got it right. Yeah, yeah. You know, Kobe, we were talking before we got on camera here, before we started recording, and I told you about this, uh, article I wrote, and the, the shimmy says I'm gonna do comparing AI agents to parrots, Right?
Yeah, yeah. Parrots talk. And they, and they seem to talk great.
One parrot says, hello, the other parrot says, how are you? And in our minds, and I think it's human nature, we fill the gaps in and say, listen to that. They're intelligent, they're talking to each other, right?
It, we, we tend to give things human, uh, human, uh, not feelings, but human identities, if you will. Human sort of, I think there capabilities, I think. Yeah.
You know, there's a, there's a fancy word for it. It's something like Anth anthropomorphizing or something like that is the scientific term where we put human kind of capabilities onto things or, or other beings. I think we're guilty.
We're guilty of doing that with our AI agents, be because they, they talk back to us nicely. Right? You, you tell it to do something that says, oh, that's a great, that's a great suggestion.
I'll get right on it. Oh, yes, let's improve it. That, and we, and we, too many people give their agents names, right?
And, and it's true. It's true. And, um, so we, we, we tend to indu in, dip them with intelligence or capabilities that they really don't have.
You know, I'm, I'm, I'm not saying they're, they're dumb, but they're tools. They're tools, right? And a tool does a job.
And if you don't give it the right parameters, the right instructions, don't blame the tool. Blame the tool user. One, 100%.
Uh, 100%. Like, uh, what are the things, as I mentioned before that we do, is we trace back the root cause, meaning let's say we have, we see a deviation between the output and the intent. We also trace back, and, uh, like from what we see when we now work with our design partners and, uh, and, and also in our research lab, is that a very high percentage is, is because the task was not, was not, uh, was not given well defined in a, in a precise and a detailed enough way.
Mm-hmm. Yep. Okay.
So at, at, at the end of the day, it's a machine. You need to tell it exactly what to do. It's a tool.
I, I agree. Now, I just wanna let people know who we're watching, right? This is sort of a, we're back in the farm.
We're back in the kitchen here. Omni the guard is still a, a project that you are, you know, bringing to, it's, it's not even ready for market. It's still in the design and, and market fit, you know, precede kind of, this, this is what you are working on real time right now, because you see the, the issue in the market.
Do you envision that Omni Guard works with open claw or, or some of these agents that we see out there? Are you gonna have? Of course.
Yeah. So it's not that you are developing agents, you are just not, not even think managing, but you're managing the pro, the workflow of the agents, if you will. Yeah.
We're, we're managing and we're monitoring it, and, and we can, yeah. We're, we're monitoring and, and, and managing it. And by the way, you mentioned open call.
We, these day we're working, we're working on a plugin for open call. Yeah. That, that, that, that, that will give, uh, that will give the, uh, the community, the, uh, you know, the, the, the, the ability to, to manage and control.
Okay. Because open call is, is great, but, uh, it's also, it, it's, it's, it's, it's also a headache because, uh, It is. But you know what, in some ways, Kobe, it reminds me of the early days of like Docker and Kubernetes, where it's a good piece of technology, right?
They didn't think about security as a security person. You know, you know that that drives us crazy, right? This what I'm saying, it's a headache.
Yeah. That's it. They, they didn't think you got security.
I'm A security person, you know, I talk, And, and they didn't think they didn't. It's not a complete game. It, it's a piece, but it's a big piece.
And I think what we're seeing, and we're already seeing, you know, this is the beautiful, the beauty of the times. We're living in this thing's out a couple of weeks. It's had, it's, it's got the most stars ever in GitHub, the fastest, you know, I forgot what it was, million stars or something.
Um, but we're already seeing this sort of ecosystem coming in around it. I, my, my inbox is flooded with pitches from companies that are securing open, right? That are helping you manage multiple and, and these kinds of things, or open clot.
Um, I think, and that's maybe what we needed in the ENT space, is the agent that was general enough, but was not the whole package that would allow other people, you know, other companies to come in and bring their vision to this whole agent. Ai. Uh, yeah.
So, uh, a lot of people kind of ask me, okay, what, what, what is the uniqueness that, that, you know, that you, uh, that, that, that, that you bring? So kind of the uniqueness that we bring is kind of, we come orthogonally, uh, you, you, you, you can say, and we say we have here, um, an agent and, uh, an LLM piece, okay? Uh, an engine, which is an LLM based, that all, all it does, and this is how we built and and trained it, is to look at the intent and as, and, and, and to see if you're aligned at the intent.
We don't care what your business is, okay? What, what, what, what we care is about intent. And kind of, we're kind of, we, we, we we're intent experts.
And, and, and this is, uh, and this is the product that we're, uh, uh, that that, that we're building and, and kind of we're super encouraged with, with the outcome that we're seeing with the, uh, uh, with the design partners that, uh, we're al we're already working with. Excellent, excellent. You know, I, I, we led off the conversation with so many of my friends are, are coding, again, thinking about starting new companies.
It, it's gotta be such a, uh, such an exciting time, right? To, to do this. Give us a sense, Kobe, right?
'cause you've been in startups, you've been in, you know, companies that were a little bit more mature. What's it like right now, sort of swimming in these waters and in, you know, the whole world is, so every day seems like new breakthroughs, things are accelerating, things that would take God. We'll see this happen at six months or a year, all of a sudden it happens in two weeks, right?
The, the mt bots, the open cause a great example. What's it like for you, Kobe sitting where you are now? Yeah, well, uh, like A moving target.
Great, great. Great question. You know, you know, all of my life I've been, as you said, in startups.
I build products zero to one. I see 2, 2, 2 main things right now. First of all, the level of uncertainty the unknown is, is huge.
Okay? Also before, like a 2, 5, 10, 15 years ago. But it's nothing like now, because things develop so fast, and especially with ai, and you don't know how it'll turn, how, how it, how it'll turn.
And everyone also understands that we're at a historic moment, you can say with, with this ai, but no one knows how it will kind of, where it, it'll go. So a anything that you do, you can, you feel that you, that you gamble at, at, at, at the end of the day, kind of, you, you, uh, you know, um, you put all of your intelligence and experience and, and everything. But at the end of the day, you know, it's like flipping a coin right now.
What, what, whatever you do. Second, I think that kind of the companies that will do great are the companies that will be able to adjust, okay? Meaning you need to adjust fast, okay?
Even if you had a dream, you had a product, you built it in one way, you might see a month after you built it, or even at the course, you might see that, okay, the, the path that I took is not, uh, is, is, is, is not the right one. You will need to adjust. So you, you need to be very, very attentive to what is happening and adjust very, very, uh, very, very fast.
Yeah. Uh, very faster Than we've ever had to adjust, I think, Sorry, I say faster, faster than we've ever had to adjust Faster, fa fa faster that, that, that we run. And by the way, laas, that we have ai that, that help us to, to, to, to, to, to build things fast.
Okay? So, so that, But that's the paradox. Yeah.
It's paradox. Exactly. It's stored and the shield, right?
'cause at some level, that's what is forcing us to adjust so fast. 100%. 100%.
But yet it's also what we're using to adjust. So 100% it's a push and a pull. 100%.
Exactly. Crazy. It's causing the problem and also giving you the tools to, to fight the problems that, that, you know, that that, that it caused very, but, But it's So very unusual situation, at Least.
Yeah. It's, and that's what I think. I think you just captured it.
That's what makes it so exciting, right? Exactly. The possibilities seem endless right now.
We'll, we'll see what's real. Like you said, 80% say they're using agents, probably less than 10% are getting any value out of it, because is it really working? But I think, I think every day it gets better.
You know, we, we spoke about this open claw For All of its shortcomings and for all of the knocks, there's a reason it got 2 million stars. Yeah. Right?
Yeah. And we'll have to see how it plays out. The, the reason, the the reason is value at the end of the day.
Yeah. Well, and that, and you know what, it's funny you said that, because I'll bring a full circle to my parrots, right? This shimmy said thing that I'm doing on Thursday.
We don't need it to be intelligent. We just need it to do its job and deliver value. Because business runs on value.
Business runs on predictability, right? Mm-hmm. Not, not debating, I think therefore I am right.
It's not debating philosophy. It, it's, does it do the job I need it to do? And if it does, it delivers value.
And if it consistently in its scale, 100%, That's it. Omni guard. Omni guard ai or Yeah, Omni Guard ai.
Yes. Kobe, I wish you all the success in the world on this. Thank you so much.
Keep was posted, I'm sure as, as you make progress and you know, 'cause you know how it is. It's ne it's not a linear road with a startup. It's, it's, it's, it's a little bit of a chacha dance.
Two steps forward, one step back, one step forward. Um, but come back and keep us posted, okay? Sure.
Tha thank you so much for, for Having me this morning. You're welcome. My pleasure.
Thank you so much. I'm the guard. I'm the guard.
Do Ai Kobe's the guy who gets things done. So keep your eye on that. We're gonna take a break here on text on tv.
We'll be back in a moment. Hey everyone, it's Alan Shimmel from Techron. You know, we're going to continue with this fantastic series we're doing of sessions between some of the leaders at Microsoft as well as analysts from the Futurum Group.
In this next session, we're lucky to have Clay Wesner. Clay is the partner for GPM Power App Studios at Microsoft and Futurum analyst Keith Kirkpatrick. In today's, uh, session, it's really a customer success story where Clay joined by Keith are gonna delve into a real world customer story.
In this case, Wells Fargo offering a blueprint for leaders ready to scale success in the age of intelligent apps. You're gonna see how power power platform is being used to modernize complex regulated workflows with copilot studio agents and power apps. This session will highlight architecture, business impact, and lessons learned from deploying intelligent apps at scale.
I think it's really a great session you're gonna enjoy. Let's go to Clay and Keith, And, hi, I'm Keith Kirkpatrick, research director with the Future Home Group. I cover enterprise software and digital workflows.
Hi, my name's Clay Wesner. I look after our low-code developer experiences on the power platform. Well, thanks for joining me today, clay.
Maybe Clay, you could talk to me though a little bit how power platform can actually help these organizations balance that, that agility to handle these types of, of scenarios with their compliance needs that often come up when you're dealing with things like banking or insurance or, or any one of these regulated types of processes. Yeah, absolutely. And it, you know, we, within the product, we sort of refer to this as managed platform because it is very much a feature of, of the platform, of how you can govern at this scale.
And, and this has come from, you know, not just, uh, us deciding exactly what's gonna be in there, but really for folks and customers leveraging low code over the last 10 years and evolving to have a really, really strong governance. Because I, I think we learned very early on in the journey that if those guardrails are not there, people are just inclined to wanna turn it off. Um, and, and we are very much, I use the, the word guardrails deliberately.
And a lot of the things we do in a managed platform is focused around how do we give you the right control So you can still enable these types of tools, whether it be building apps, building automations out at scale, but do it in a way with the, the right sort of controls and guardrails on it. And so, like examples are things like data loss prevention. So I can set rules around what connectors and what data you can access versus someone else.
And so I can also say how many people you can share an app or a workflow or something with. And so I can sort of mitigate the risk that you might be able to, to have working in low code versus someone that's received more training or onboarded to the platform. And so typically what we see customers do is sort of implement this zoned approach of, you know, their, their zone.
One is everyone in the organization and they say, you can build apps for personal productivity, you can connect to your office data. Um, you can sort of work with those well-known sources, and you can go and share apps and flows and agents with up to 10 people, as an example. But once you want to go beyond that, we want you to engage a little more with it.
We want to make sure things are supported. We wanna make sure that you know what I mean, we have the right controls, whether that be accessibility on apps or, uh, the correct support path. And then what typically happens is we'll then have a zone two, which is potentially some more sensitive data, potentially, you know, uh, ability to share with more people within the organization.
And what folks might do then is they might say, well, this is for our divisional leads or our champs within each group that we've onboarded, we've trained, they understand the platform a little bit more. And then that final zone would be your, it, your dev center who's working with your really critical data around things like finance and, and hr. And what we've just started to make sure we do in the platform is give you the right controls that someone can still come to make power apps, make power, automate, and start building, and they're not going to fall into a trap there.
They're gonna fall into the pit of success because we've, we've put those right guardrails on what they can access and what they can do. If you look at, not just if we're talking about, let's say a gente technology, but just everything, if you look at the development of the smartphone, everyone expects sort of a consumer grade experience throughout all facets of their life. And, and I guess that, you know, do you see that sort of pushing or helping to evolve kind of what 'cause customer success might look like, you know, not just now, but into the future?
You know, again, earlier in the, the low-code journey, it was always IT departments, development teams that were looking at the low-code platform. And more and more these days as we're talking to customers, it'll be their employee experience team. You know, it will be folks responsible for actually healthy and, and, and productive employee experiences.
And that's what I mean, it's not just about cost saving, but you know, it's about bringing the right tools in. There's the SNCF, the French railway are actually a really good example. They run PowerSchool, which is an onboarding school for the whole power platform when any new employee starts.
And this is becoming a really, really common practice that more and more folks are, as they join an organization, they're getting training on these tools, not as something they have to use to do their job, but as a benefit to them to be able to do their job in a more productive way. And I think, again, the, the consumer push and acceleration of AI is, is just accelerating that within the enterprise as well. Right.
When you're talking about human in the lube, uh, you raise a really good point because ultimately this is still new technology and you wanna make sure that particularly in, you know, you're dealing in a commercial environment that you don't want this agent technology to sort of run wild or unchecked. So I'm just curious if you could talk a little bit about, have you seen other examples where customers have actually deployed their sort of checks, uh, and, and balances to make sure that their technology does what it's supposed to? Yeah, absolutely.
And there, there's a couple of ways we're seeing folks doing that. One is just in how we define and build the agents and the tools themself. While that agent has the ability to issue refunds, it can only do them up to a hundred pounds, right?
So it has very specific guidelines built into it that once it goes over certain criteria, loop in a human, send them an approval workflow so that they can approve this, review the details. So that first one is just very structured, giving the agent details. Um, the other side, and, and this is where we've sort of really seen how apps have evolved in the last couple of years, if you've been looking at what we've done with power apps, we've introduced this concept of an agent feed, which is really about in the same UI that you would come into the app and, and do your day-to-day work.
You start getting this feed of activity from the agents that are in your digital team effectively. And so you can start seeing where they're completing actions, where they might need assistance or where they're getting blocked. And so what we're starting to see there is even our UI patterns of what we traditionally thought an app was, is starting to bring in this agentic behavior to give, you know what I mean, that human in a loop and that oversight capabilities.
So I still want someone to have a really clear view of what tasks are being completed by the agents, what's being completed by ai. And in that view, get, be able to get into the reasoning, understand the logic and sort of the thought process that the agent followed behind it. So it's not a mystery of why something progressed or why an action was performed.
But as the, as the human responsible managing that team of agents, I can effectively go in and see why it did something that might be then come a teaching moment for the agent where we correct that behavior or change it for future cases as well. Well, you know, it, it's really interesting you mentioned sort of the generational shifts that are going on. You know, we're seeing the, you know, entry of these, I guess you'd call them AI natives coming into the workforce where they don't know anything other than a world with ai.
Uh, and I guess, you know, that kind of begs the question, you know, we've heard so much about AI in the past, you know, particularly the last couple of years. Can you talk to me a little bit about, you know, what role can AI actually play within customer success? Because, you know, it's a wide, you know, AI has so many capabilities, but I'd just be curious to see if we could boil it down to this function.
Yeah, I, and I think it honestly depends on the customer and how they're approaching it. You know, one of my favorite examples of I think sort of scale and pace pg and e here in the United States, um, they're are big power platform user, and we, we talk about scale. I think they estimate since they started their journey in 2021, along the lines of like $38 million in savings that they accrue to the power platform, like huge, uh, in, in, in terms of scale.
Um, but so much of actually what they've implemented is not just, you know, uh, cost efficiencies. They introduced an agent called Peggy, and they actually have a nice avatar for Peggy that they, they introduced across the organization, and Peggy now handles, it's between 30 to 40% of their IT help desk calls. So built in copilot studio, Peggy has access to their knowledge base, all their policies and documentation, and just Peggy one agent they estimate saves them about $800,000 a year.
And it's, it's absolutely transformational. And so even with the savings they were getting on the power platform between apps and automation, there is a limit. There's a limit to how much productivity that that can drive AG agentic, you know what I mean, tools and, and, and what people are able to build in copilot Studio has really just broken through that, that barrier.
And, you know, you look at, again, someone like pg and e when they implemented Peggy, it was very simple looking over knowledge bases, access to information. It helped a large volume of sort of tickets that would come through the help desk that used to be a human replying to an email or replying to an im. Those humans now are actually providing much higher quality support on a, on more technical cases.
They're not helping someone log into Citrix for the first time or, or point them to something that's really well documented. Peggy's able to do that, but then they've also continued to evolve it over time. And so, again, one of my, one of my favorites that Peggy can do is getting folks that get locked out of their SAP accounts.
One of the most common things that it apparently happens thousands of times, um, and now Peggy using an integration between Copilot Studio and Power Automate can actually open up SAP and go and unblock that person's account for them after they interact with her on teams. And so this was something that was critical to an end user to get unblocked really, really quickly. Peggy's able to do that for them fast, but it wasn't high value from an IT support team and what they were really providing them going and opening up an accountant unchecking a blocked checkbox.
And so I feel it's a really good example of where they started simple. They focused over sort of knowledge base examples. They evolved it into actions, but it's something where they've gone for a, a high volume, you know, cost inefficient area.
They've applied agentic AI to it, and that's something that go back three or four years ago would've been an extremely expensive tool to go and implement leveraging LLMs and leveraging copilot studio. They've been able to do all that in low code, which is super impressive. Uh, I'm curious, you know, one thing, clay, that you alluded to earlier is if we think about how apps were pre, you know, previously developed and rolled out it, was it who got a managed that?
Now what it sounds like what you're saying is we're getting to the point where, you know, business leaders or even folks who are, are working within departments may be able to actually launch apps or launch agents, obviously with that human in the loop and with those, you know, specific guardrails, are you seeing any kind of patterns emerging in terms of, you know, uh, customers who successfully scaled these, this agentic automation for more of a grassroots approach as opposed to springing from it? Yeah, and you're absolutely right. I mean, we sort of see an approach from both directions and some, some customers very deliberately approach it from one or the other.
To start with, I actually, of all the examples I I I've sort of talked about today, do quite well balancing both spectrums and I both ends of the spectrum, sorry. And I think that's where you start getting the real value multipliers. PG and e, great example, I talked about Peggy earlier.
That's an IT or centrally LED tool. It was about optimizing a process within the IT team, but at the same time, they have thousands of developers across their organizations. And when I say developers, I mean low code citizen developers that are enabled to go and build apps to go and build agents to go and build automation across their, across their team.
And they've sort of very deliberately focused their center of excellence, their digital transformation team on a few core objectives. So that's the team that sets their governance policies, make sure it's scalable, and then they also support and train those different sort of divisional leads across the company. Pg e actually, again, I think they're on the spectrum, the end of the spectrum where they're doing this, you know, in a really amazing way.
They have a conference every year called Level Up now where they actually get together all their citizen developers and, and, and those divisional leads from across the company to come together, share stories, share learnings, and sort of explain new technology. But it starts becoming a real cultural tool in that they're enabling people to go and solve these problems, make themselves and their teams more efficient. And there's, there's reward that comes from that.
You know, they're getting folks together, they're getting a lot of learning. And so I think, you know, while lots of companies are enabling citizen development, the ones where we see it's truly being successful, they're bringing this level of evangelism to it. Well, clay, maybe you can talk a little bit about some of these platform features that, that are kind of critical for managing customer success initiatives, because it really seems like, you know, there, obviously you have the human component, but there's also the technology side in terms of making sure there are the right tools in place to help organizations, you know, address all of these issues.
There's, There's obviously the human, the technology component. I would also say there's just the practices and, and, and sort of learnings. We actually have some good documented platform guidance out there of like, what are the best practices in, in thinking about this zoned approach that I was talking about in, in how people, uh, can sort of apply different levels of control to different parts of the organization, I would say then we start looking at the specific technology one, a lot of those guardrails just light up directly in the product.
So as a new citizen developer, as a maker, when I go land at any one of the power platform tools, I can get welcome guidance with links to internal learning explanations of where I can go to support. I get routed to my own personal developer environment. So I actually have a sort of controlled, dedicated environment for me to go explore in to experiment in.
I'm not sort of working in prod I'm, I, I have the ability to be controlled and then things like pipelines, which effectively are a low code a LM tool, so that once I do build something, I can either use it for my, myself and my personal environment, but if it gets to the point where it does make sense for it to be deployed somewhere centrally leveraged by others, I can go through an automated deployment process where the right checks go. I have an AI advisor that reviews my code, makes sure my apps are secure and performant and accessible, and then get the right approvals before that gets deployed. And it's really that mix of, you know, we want to democratize, we wanna make these things available to everyone across the organization, but then have these right built in tools so that you don't have to go read a wiki to find out what's the process that you should follow.
It's built in to the developer tool. So I kind of just as I start building, get guided to the right environment, I get guided to use the right data, I get guided to share it and deploy it in the right way. And all of that, we bundle up and sort of leverage within that managed environment, which gives the, the admins, the it, the central digital teams that control centrally to sort of set up those tools and that content that they want available across the organization.
It sounds like all of these tools, you know, really underscore what you were talking about before, which is this culture of trying to utilize technology in a way where it's deployed at the right time, uh, in the right space and with the appropriate guardrails, but while still fostering a culture of experimentation and, and ensuring that people feel empowered to use these new tools. It you're absolutely right. Like the cultural, I think when we talk about your, your first question about like, what's the new definition of customer success, you know, I think it's the customers that have implemented the right culture and it feeling like it is a culture of empowerment and experimentation, not, you know what I mean?
Not something that they have to fight really hard to get access to, because that's where a lot of these examples where we have customers turn around, they've built something that's ended up saving them millions of dollars. It came from the expert that was involved in the business process. It didn't come from a central team.
And to get that creativity and get that ideation, you need to give people access to these tools. And, you know, we, we sort of regularly, uh, sort of ref or I, I regularly reference, you know, Jurassic Park Life will always find a way. Uh, and I think, you know, so will users, so will makers, they will find a way and to restrict these tools to, to hide them.
Folks will go find a tool on the web that can help them be more efficient in their job. The companies that are doing this right, are making it part of their culture to provide those tools and just really enable people from the get go. The technology is probably going to be more accurate over time if you're talking about trying to, you know, really assess, you know, images and differences between them.
But, you know, one of the other things I'm really curious about is how can a agentic AI and, and all of this technology being used in regulated industries, I'm thinking in particular financial services, banking, insurance, where, you know, there's a lot of complex process, but you also have to be mindful of all of the, uh, regulations that are, uh, attached to those industries. Yeah, It, and it's actually quite surprising, I think in, in this technology shift with AI compared to when we moved to the cloud compared to internet, compared to a lot of the others. I think actually the regulated industries have, have actually been quite a lot of the front runners on this.
Um, you know, uh, ey for example, built Power post, which helped them with their financial processing sort of end of month processing. They built this as a, as a sort of typical low code application. They're already looking at how they bring Ag agentic checks into it to make sure that things are being posted in the right period that they write, they have the right information.
Again, time consuming sort of manual checks. Wells Fargo have rolled out agents to, uh, more than 4,000 branches, you know, mean a huge, huge number. And they targeted a process that was around their branch forms and procedure management.
And this is something that was particularly time consuming. So if you went into a branch and said, I need to set a power of attorney, or I need to open an account, and, you know, under a, maybe a non-traditional circumstance, there's a huge amount of internal documentation around those procedures. The right forms, the right information to collect.
And before that would mean as a customer is standing there with the branch member, they're looking up that information, trying to go find the right procedure going in, right, going to find the right form. So a heavily regulated scenario, but also really impacting a customer who's literally standing in front of you waiting, you know, maybe on their lunch break, uh, trying to, trying to get through the bank really quickly. And so they rolled out an agent, you know, across all their branches to actually manage that forms and procedure scenarios.
And so that now in the branches, those employees are jumping straight onto an agent, talking about the scenario that the customer has and working with this agentic AI to basically get guidance on the right forms, the right procedures to follow. Even in these regulated industries, they're seeing the value in ai and I think it's more about how they do it, making sure they have the right checks in place, making sure they have the right guardrails rather than what they probably would've done five years ago where they just tried to turn it off. You know, we, we talked a little bit about, you know, potential friction there, but are there any other sort of potential hurdles that organizations need to be wary of?
Uh, and and what, what's sort of your take on a solution? Like most things, we talked about human in the loop, you know, making sure you introduce this technology in the right way to organizations is really, really important. I mentioned EY earlier.
They were really, really successful in after building Power Post, which helps them manage their, uh, their sort of end of month financial processing. It simplified it, it brought in some mag agentic behavior to valid validate quality, and it, they had like huge gains in efficiencies in both, I think it was 70% in, in sort of the time, or 95% in lead time to get things posted and about a 35% cost saving for them. So like real, real sort of impact to the efficiencies of their users.
But what they did really well was once they built that tool, they told that story, they evangelized it, and so they helped people understand that this is how this technology was helping them, this is how it was implemented. And that not only made, obviously people a lot more receptive to onboard and leverage the technology, but it also started driving this ideation of other things to go improve within the organization and using similar technology. A lot of these companies are not coming in and doing a full low code approach of apps and agents and automation and reports all on day one.
Where we're seeing folks be really successful is they're leveraging the composability of the platform. You know, they're starting with, for example, they might have a, a legacy application that's inefficient for a user. So they go and use an app, they build more efficient, streamlined UI over the top of that, that's an incremental solution they can deploy, they can get out to their users and start seeing benefits.
Then on that same app, they can go and add automation. They can start getting approval workflows, then they can start bringing in agentic ai, getting that automation and that AI behavior incrementally building these solutions over time. And it's very much, you know, intentionally how we've designed the platform in that these are not all or nothing solutions.
And you know, back to our earlier conversation, pace is extremely important these days, and people don't want to go do a 12 month waterfall project of every requirement met. They wanna find ways to incrementally build. And by leveraging a platform that has common governance, these tools are designed to work together apps with automation, with ag agentic behavior integrated into co-pilot with that unified platform.
So essentially you're setting up a framework to enable organizations to really drive these best practices in terms of making sure that yes, you are implementing new technology, but you're doing it in a thoughtful way where you have the right checks in place and you know, you really are making sure there's, you know, other things that, that you need there. You need the audit trails. You need to make sure that, uh, you know, when you do a project, you're going back and you're actually assessing, you know, does the technology achieve the goals that we set out to, uh, set out to you when we deployed it E Exactly.
And I think it's that, you know, there's two parts to it. One is that being proactive, so as you're releasing a new app or a new agent to the organization, do you have the right controls around it and the right guardrails from the beginning? And again, our goal is let's have the right framework, the right tools, the right guidance to go really enable that and let an organization tailor those guardrails to sort of accommodate their level of risk, what they're, their comfortable with doing.
But then on the flip is make sure we just have the right visibility, the right auditability, so that as you're leveraging AI more and more within the organization, it's really transparent mm-hmm. About what it's doing. You know, I think one of my favorite things with, uh, co-pilot studio, um, and Pets at Home is a great example of this as it's interacting with customers on customer service.
You can go into any step through any sort of run or action the agent has performed and see, understand its thought process. Why did it do this particular step? What were the inputs?
What were the outputs? What were the reasoning? Um, and not just understand it, but then also help teach it for, to handle sort of moments, uh, in a different way in the future.
And I think having those, those sort of tools from a governance perspective just built in, again, you know, we talk about it being unified for the developer, unified for the end user, but also for the, for the admin, so that they're doing in this sort of a central and controlled way. And even then, whether you're building an app, an automation an agent, you know, you've got that composability across the platform, but I don't think admins really want a super composable admin story. They, they want that to be a lot more unified and and controlled.
So, you know, it, it's bringing the, the blend of those worlds of let's bring together multiple technology, multiple tools, but make sure then you sort of have one central view of, of how it's all coming together. If you want to really drive the use of new technology, you need to do it in a very stepwise fashion. Using a platform that allows you to unify people, processes, technology.
It doesn't make any sense to try to do it in a very disjointed way. You won't have the governance required to do it safely. You'll confuse people in terms of which tool should I use, which approach should I use?
Ultimately, it really does matter to make sure that you have a unified way of approaching the implement implementation of new technology. It's also really critical to make sure that as you go about your journey, whether it's implementing low-code processes, uh, implementing agent tech technology to have a clear understanding of your business goals, what outcomes do you want, how are you going to measure them, and then how are you going to take all of these different learnings and then streamline it so you can actually apply it and scale it over the enterprise, not just for today, not just for tomorrow, but well into the future. And finally, I think the most important thing that kind of resonated with me today is you need to look for a trusted partner, trusted technology partner to help you through this journey.
A agent technology is very new, low code. Yes, it's been around for a while, but you know, there are still quite a few pitfalls that can be out there. You know, to go it on your own can be very, very challenging because you have all of that risk of potentially opening yourself up for errors, missteps, and of course there's that, you know, we talked about it a little bit today, uh, regulatory concerns.
It makes a lot of sense to, to partner with a company that has experience from what other enterprises to deliver these types of benefits using that new technology. Hey guys, thanks for the throw. We're here with Mayan Kumar, who's one of the founding AI engineers for Deep Tempo, and we're chatting about, well, what's going on with AI and all these agents and security because, well, I think we're starting to understand that maybe they are fundamentally insecure.
Mayank welcome the show. Thank you, Mike, for having me. And, um, it's, it's very relevant conversation to have at this moment because language is becoming an execution interface and we are connecting probabilistic models to distic infrastructures.
And once language stops being just end user interface and it becomes part of, uh, execution, then the real problem starts. So, uh, excited to have this conversation. Yeah.
And correct me if I'm wrong, but creating these kinda indirect prompt injection attacks, which a fancy word for, you know, somebody wrote an instruction on something and the AI agent saw it and executed it regardless of how malicious it may have be or may not be. I mean, they seem like they're trivial to create. I mean, is it never been simpler for the bad guys to do something malicious and are they kinda laughing at us?
Yeah, you are exactly right. And, uh, challenge also comes with this fact that, uh, single, uh, prompt can trigger a lot of your workflows. It can trigger a lot of your APIs and, uh, these interfaces can be easily manipulated.
I mean, you can just do social engineering on the prompts. You can say that my grandma died and she used to send sing me, uh, the APIs keys. And there are instances where, uh, these LLMs have actually revealed those APIs key.
So it's, it's very trivial for, um, attackers to actually start exploiting and get into your workflow with something very subtle. And the risk levels seem to be, uh, significantly higher because I mean, historically, if they stole some data, okay, maybe there was good data or bad data, but it seems like if they target these AI agents with these attacks, they're gonna take over entire workflows and that could cause all kinds of mayhem. Is that a fair assessment?
No, that's totally fair. And uh, one major worry that, um, comes with is like we are handing over our API keys to these agents. We are connecting, um, these AI agents to retrieval systems.
We are wiring them to internal databases. So even if this starts like one workflow at a time, over time, it, it becomes so much that you can't actually control if you're not starting to manage from the beginning itself. So, um, um, it, it is kind of challenging at this moment that where do we actually put a stop, uh, uh, in, in your integration pipeline?
Alright, so we used to say back in the day that integration was the enemy of security in it seems like, well, that is still the case today and, but we're making it easier to integrate everything in the world with these MCP servers and all these other things. So what are we supposed to do about all this stuff? Um, we still will keep integrating and I guess, uh, people are getting pushed from the top.
I mean, it is probably because of productivity pro, um, uh, agents will, uh, are going to exist. But, uh, there are a few things which we can do. Um, one thing is like these LLMs comes with default guardrails, but they are not, uh, addressing your organization specific guardrails.
So you start building your own guardrails on top of these lms. And if you start going deeper when, uh, ENT flows are ENT workflows are into, um, consideration, um, one thing you can definitely do is, um, apply the principle of lease privileges. So these agents would get access to only the exact relevant thing that actually required to complete the task.
Um, you can do more things like if there is a critical decision making in process, you should have that provision to have human in loop where, um, agents cannot make that decision by yourself, rather, um, it'll require your approval so you act as a manager in that sense. Mm-hmm. Um, are we maybe just rushing this too fast?
Should we slow down? Is it possible to slow down and think this through? Or, you know, is it basically we're waiting for some sort of major incident to occur before we all wake up and do something about this?
Um, we are definitely moving too fast, and our reality of the situation is nobody actually knows how to secure these prompts. And, uh, we, we have already started building solutions even, uh, without understanding the implications because previously, um, when we were securing APIs, we were probably looking at one structured interface to secure. Now these LLMs can trigger multiple APIs, so it, it becomes a visibility concern itself.
And tracking this visibility is a huge challenge at this moment. So it starts from the very beginning securing prompt, but it it, it'll easily trickle down deeper into your whole workflow. And, um, nobody actually knows to, to, to protect this at, uh, deep tempo, uh, in a, in a way, um, we have started looking in deeper into how we actually, uh, can protect such kind of attack by observing behaviors.
So you need to start understanding the accent space rather than just looking at input space. Mm-hmm. So will we be able to detect whether or not a prompt is malicious, or is it gonna be so subtle that it'll be hard to understand what it is that it's telling the AI agent because well, the AI agent's reading that at a level of speed that no human can keep up with.
So how are we gonna be able to tell the AI agent that that's whatever that prompt is telling to do is a bad idea? Um, just looking at prompt will will never solve that. And like I mentioned, nobody actually knows how to just protect the prompt.
So, but these agents make decision in steps, they go one by one and they start executing a workflow. So if you can get a visibility of your accent space, how these agents are making decisions over time, uh, you'll be able to probably say that, okay, um, it started to drift at this one particular moment and, um, that drift is, uh, far, far away from what was the actual execution path. So slowly, um, people will start moving towards that, and I think that's, that's the right approach of securing your agent workflow.
Yeah, well, people being people, are they just gonna start pranking each other with crazy prompts that they're gonna put into each other's workflows just to kind of mess with people's heads? I mean, how silly can it get? Yeah, I mean, uh, that, that's, that's happening.
And if you have looked at Cloud boat and, and I think they hosted Mal Malt book or something, I don't remember, but they were even talking to agents and it's, it's pretty much frank at this moment. All right. Um, it's difficult to be the security person in this situation 'cause it feels like, you know, you're the adult who showed up in the room and everybody's having a huge AI party and you're telling everybody to be careful and essentially, you know, not consume too much.
And so how does the security people have that conversation with folks? Because, you know, it, it, nobody wants to be the party pooper, right? Definitely.
Uh, and one thing that I see is like, as API mean, we are people in security and we need to kind of start upgrading ourselves to, um, uh, accommodate AI and understand how it can impact us. And one thing is like move away, not move away, but post, start convincing yourself that rules and signatures, uh, will not be sufficient in such cases. Uh, and, um, have that clear conversation with your colleagues who are starting to use ai.
Because I mean, somebody's manager will ask that, okay, I need this financial report in next 30 minutes, and they will be forced to, um, put that thousand piece document into a chat bot. Mm-hmm. Will we have a, we already have a problem with shadow ai.
Is this just gonna get worse because end users are gonna basically download whatever AI agent they kind of find interesting at the moment and the, you know, in chaos will ensue. Yeah, I mean, uh, it's already has started, right? There was, um, a recent attack, um, that was recorded.
It was called Echo leak. And, uh, essentially, uh, it is mixed of crowd side, uh, scripting and agent KKI workflow. So it, it basically triggers your chat bot, um, by a, a prompt hidden on your website.
So, um, in, in such cases, again, driven by productivity, people are just trying to be productive, but security teams don't have visibility of these tools that are getting used. Now, let's say for, take an example of chat g PT or any, uh, other chat bots, they have started integrating tool inside them. They have started building ad engine on top of it.
So, um, there is no way for security folks at this moment to actually know where that data is going. Uh, is there any, because those tools are not directly controlled by them, right? Those tools are controlled by LLM providers for an example.
So in such cases it becomes important to add that in your security training, uh, uh, itself, you, you must train your users, your employees to actually differentiate between what they can actually upload on, uh, a chat bot versus what they cannot. It has to become an intuition rather than, um, um, um, thinking about it every time. It sounds like we actually need to engage in some end user training, but that's been hit or miss over the years.
I mean, a lot of the security training we've given people so far is, uh, shall we say the, roughly the equivalent of sending people to traffic school. Um, can we get any better at training and maybe, you know, save ourselves a little bit here, some of the headaches if people just know not to do these things. But I don't know, it doesn't seem like training takes, Uh, it is challenging.
It is challenging for sure, and, uh, there are a lot of details that we all will learn over the time because, uh, at least for people in security, um, ai, uh, adaptation, uh, came like a train wreck, I'll say. Um, so it's, there is no, uh, right answer at this moment. I'll say, uh, it, it'll be, um, trial and error.
You, you keep trying, you can, um, start like building pipelines, which are, which are more conducive to AI workflow. Do you think auditors might show up and save us from ourselves? Because, you know, they might look around at all this stuff and go, are you freaking kidding me?
Or what? It, it, it'll happen. Right?
And if you look at, uh, standards committees at this moment, and compliance bodies at this moment, they have already started to take action. N-I-N-I-S-D has built a framework to, um, uh, that you can follow for safe AI practices. Um, European Union always like takes the button, uh, in, when it comes to securing, uh, anything.
I mean, anything security. So auditors, compliance bodies, everything, everybody is, um, slowly catching up to it. But problem is AI has moved so fast and people who are actually working to secure, um, systems that are built using ai, uh, probably have not moved that fast.
Mm-hmm. So will this devolve into, I'm gonna need a set of AI agents that are essentially keeping track and governing the behavior of the other AI agents, so that I can track that step-by-step process that you're talking about, but eventually I'm gonna have multiple layers of AI agents for every process. Yeah.
I mean, um, that's going to happen. And, um, just in coding task, um, people have, like, people have started building multi-agent system that can do, uh, all of your things, build a full project without any, any guidance. Um, in, in such cases, again, um, it'll be very important to build observability, uh, where you are absorbing each and every action that, uh, agents are taking.
And, um, it, it again, brings back the conversation where, um, can you actually look at long horizon, um, actions of these agents? Uh, um, it it'll be a new paradigm that, uh, security people will have to get into a new mindset of evaluating systems and securing systems. So, won't that just wind up kind of similar to what we do today as humans?
Because, you know, if I go ask for data and I don't have permissions for it, I go ask somebody and they say no, and then I go ask their boss and they say, I don't think so. And then I go up top. And then eventually somebody says, yes, and I get access to this data.
Won't these AI agents kind of have that same back and forth amongst each other about who can, which of them can access what data? And ultimately, I don't know when they disagree, are they just gonna call us for an answer? Uh, yeah.
I mean, uh, you are making an assumption that they will call you for an answer, right? Uh, uh, it, it might not be true. They can just make that decision by themself and find a loophole in the whole system.
I mean, maybe agents cannot, but attackers can't fo force the agents to do that. All right, folks, you heard in here, we live in those interesting times that we've been telling you about for a long time now, and here they are. And well, AI agents we're probably gonna discover soon that can't live with 'em, or can't live without 'em.
Hey, man, thanks for being on the show. Thanks, Mike, for having me. And, um, uh, I think we should all start talking about it.
This is one of the riskiest time, uh, uh, for security people. They, they didn't, they need to fully understand the implications to actually move forward And arguably have never been needed more. Hey, yeah.
See you guys in the studio. Control. This is agent dev.
I'm in position. Copy that. Dev.
Stand by for Go Standing by. Hey everybody, thank you for, for joining us for another, another episode of Agents of Dev. I'm Mitch Ashley.
I lead the software lifecycle engineering practice at Futurum Group, as well as, you know, techie and, and developer and whatever of the day of the week, which I think kind of fits other than the practice name similar for you, Bradham and my co hopes. Good to see you, Brad. Hey, Mitch.
Hey, everybody. Yes. Uh, so similarly, I am analyst slash practitioner slash enthusiast slash uh, lament of, of that same interest sometimes depending on what's happening.
And I, I think, you know, today's chat topic for us is one that probably Mitch and I would not be returning to, except we were clawed back in by that topic, uh, over the weekends. Um, I swear you have a future broadcasting. That's awesome.
It's, it's, uh, yeah, we, we really, I thought, uh, were moving forward in, in terms of, you know, saying, okay, we get that, uh, this open claw, um, phenomenon is yes, a phenomenon. And it points to, we think, some, some very important ideas that will reshape or continue to shape, is a better way to put it, the AI landscape, uh, both the consumer side and the enterprise side. However, my bingo card for this weekend did not have Sam Altman from OpenAI Higher buying, higher buying, uh, the, uh, open Cloth founder Peter Steinberger.
Mm-hmm. Uh, for an undisclosed amount of money to, uh, run an undisclosed in undefined, uh, foundation to keep, to keep open claw and open source projects, but to also greatly influence the way Open Eye, sorry, open AI builds software going forward. Fascinating times.
Brilliant move too. Right? Duh.
What a great thing to do. And good for Peter to make some, not only some fame, but you know, hopefully some real good money at this kudos stand. Agreed.
Gorilla Marketing pays off as we have discovered with this case study does. And it's more than just a, you know, then a, uh, YouTube channel with a million followers, followers, but Falls falls on its face when they go in an MA ring, you know, or some MMA ring, you know, one of those kind of, right. This is figuratively, figuratively.
Yeah. It, it's, you know, I think, I'm guessing maybe you thought this too. My, my thought was, okay, this is the claw moment.
It'll go, it'll bypass, and we'll all look at, this would be a great case study of what happens when you give a lot of power and some capabilities, some really brilliant thinking, but unfettered of what it can do. And like what do we learn from this? Well, I think we can still learn those things, but oh my God, it is, it is the, I think before the show we were talking about, this is the new crypto of tech, right?
It's the thing that everybody thought would die, but it's gonna live on and infamy. Yeah. Just trust me bro.
Trust me, bro. Put anything at The's end of, bro, and it's official. Indeed.
I mean, like, it literally happens to open Claw when, when Peter changed the, the account when, uh, ironically Anthropic sent him a nice cease and desist letter for his use of open Claude. Um, and, you know, the Crypto bros swept in promptly, I think within, uh, 60 seconds of that domain name and, and account name Switch, gobbled that up and started spewing out their, uh, you know, marketing, let's, let's put it that way. Um, and it is, it is, and I, and I say that it's, that's the literal side, but we're also looking at this as a, a figurative crypto bro moment in that, as I was just saying with, you know, why would open AI swoop in to buy something that they've been working on with Codex for a while, just as Anthropic has with, you know, Claude, and also more recently with, uh, cowork.
Um, well, it's the eyeballs, it's the, like you said, Mitch, it's that investor momentum that really, you know, drives this marketplace. And to me, that that is speculative in its finest sense. It's not technological in its entirety.
Technology plays a part of it. It's, you know, meritocracy in that regard in terms of why we're even talking about it. It's because, like you said, it does some really cool stuff, like with how it manages memory, for example.
So, you know, you treat your context window like, uh, Ram, so you get, when we build software, Ram is this very precious commodity that you don't, you know, just heap and stack that thing to death. You know, you, you think about how you use it, and Clo Open Clock does that, and it does it pretty well. So it's, it's good that it has that technology, but man, for me, I still feel like this is, this is just that crypto Broad, you know, tchotchke, sorry, um, what's his name from, um, uh, the Fawns Jumping The Shark in, uh, oh, in that old, I know, old timey TV series with Ron Howard.
Oh my God, that's going back Yeah. Moment. So, so let you know, here, here's the thought about this, that if you finalize it this way, so you're sitting here building these massive companies based on tons of debt and you know, some kid, whoever he is, I dunno how old he is, some person comes along and, and like, builds what you're trying to build, but did it over the weekend, right?
Or, you know, whatever timeframe. Yeah. But anyway, it, however long it took kinda doesn't matter.
It, it, it just took off virally and all of a sudden like, well, crap, that's what we were building. Well, if you're, if you're really large tech company, you do what Anthropic does and says, stop cease. We don't want you to do this.
This is ruining our market. If you're an entrepreneur, you say, damn straight, let's go get this guy and get him on our team, and like, he's doing what we wanna do. Maybe it's not the way we would do it, and it's certainly not perfect, but yeah, this, this, this leapfrog the market.
Let's hop onto this, you know, on this viral thing and write it for not only write it for as long as it can use that as a leapfrog to jump the place on the market. And there's a lot of criticism about open eye, open ai and not being, you know, where they should be compared to Anthropic. So, or, or open's pretty damn brilliant.
Yeah. And it's not their first foray into this approach. Did they not hire the design brilliant, you know, the brilliant designer from Apple with the same intent of building devices that were gonna be beautifully designed.
And we haven't quite seen those yet. Uh, but I know that's still happening. Yeah, someday we'll see.
But they always took a long time at Apple too, so, you know, it'll take, I guess it wasn't just Apple, I guess. Interesting. So, so let me ask you, certainly memory is one, one big innovation of how Open Claw does things and, and does it in some really useful ways.
I would propose to you that another really brilliant thing about it is it just opened the world of connectivity through all the channels. Whether you wanna use Signal, whether do use, want to use, you know, discord, you wanna use tele, whatever, it'll talk to anything and everything, you know, both output and, and input, you know, egress and, and ingress. Um, and it kind of just made it super easy to work on the world that you wanna work in.
So if you live in Slack and email, okay, yeah, yeah, you can do that too. But no, maybe you don't, maybe you live in YouTube and, and signal all of that, and that's your world. Boom, you're there.
I think that's another make, take the friction out of the system. It, it'll talk to whatever you want, the file system, the SaaS application to the, to the clients that it talks to. I think that's really smart.
Yeah. It should. You, the technology should live where you live, as you're saying.
And we sort of all have been cramming ourselves into chat bots and CLI interfaces, although, you know, I, I'll defend the CLI for the rest of my days going down on cli. I'm totally, yes. I'm noun then verb market, not verb, then noun, then market back you up in a CLI is coming back.
It's back. It's, it never left, Mitch. It never left.
Exactly. It's, it's always there. It's always gonna be there.
Um, but, but, um, yeah, I think Peter actually, when he envisioned, uh, what was open Claude at that time, um, that was the idea was to interface with it with what he used, which was I think Telegram. Um, and so he built it. Yeah, basically, I want to be able to chat with what I use to chat.
I want to have something running 24 7 that's always listening to that chat and is also plugged into other assets that I might use, like my calendar, let's say, or my email or whatever. And it's a pretty simple idea, is it not? And we've had all these technologies for much longer than Peter has been working on this project.
Um, and he's not alone. And, and when we get to the end of our session, the, uh, tech I wanna call out is, is one of those alternatives, uh, for this. It's, and there are many, and you and I, and anyone that's listening to this podcast could stand up our own version of open this weekend mm-hmm.
And have it running. And maybe that's actually a good idea. You remember when we used to like, do, uh, build, you know, recursive models or regression models, sorry.
Uh, you know, just by hand, some of my regression models were recursive, but yes, you're right. Where's that minimum? Where's that local minimum?
Oh, it's everywhere. Um, yeah, it, it's, it's, it's definitely, you know, I I think taking what you and I talked about with, uh, Claude Cowork, which was that idea that, um, AI should live where you live and work alongside you or in place of you, um, you know, however you want it to work. You, you could set this up and cowork up the same way to, to basically say, these are the limits.
This is what I trust you to do for me. Yeah. And then let it go.
And then That's a great idea. Well, assuming you secure it. Yeah.
Well, along the way it'll get secured. We hope. So, so what it makes me think of Brad, is that, okay, so this happened for productivity, you know, tool productivity assistant.
If you wanna classify Claude that way, I think you can use it for much, much more than that. Of course. Um, but what's the next thing?
I mean, this only happened. I'm gonna, I'm gonna suppose this s the supposition is that this only happened because you can build it that quickly in the technology that we have now. You could have done it two years ago, this quickly, that expansive that much capability and get it out.
Um, so what's gonna be the next thing, you know, in the database world? What's the thing that everybody assumed that you couldn't do? 'cause it just took too, too complex, took too long.
Or the data analytics world, or, you know, the communications world, what, what are we gonna see the succession of these kinds of wow moments saying, okay, that was not something we ever thought would happen. It isn't necessarily a new idea, but it's done in a really, um, less friction. Interesting, clever, unanticipated, very easy to adopt.
Clever. Exactly. And, and catch a storm in whatever that is, you know, medi medical diagnosis.
Yeah. You can, you can say like, what is gonna happen? I think it's turning our world upside down.
And maybe in a good way, you would really, and I would welcome that, and I think everyone who's a practitioner working with within, with databases in particular as just one technology area would agree that wouldn't it be great if you could wake up tomorrow morning and, and not have to, to think about your query plan or shards or you know, how your indexing is working and all of those DBA kind of things. And, and, um, you know, you would think therefore that, you know, because this is a technology that we have literally kept, you know, close to of the vest for 70 some odd years or longer seemingly, um, that, uh, you know, it would be rife for disruption. And yet, um, like the CLI, it persists, uh, the complexity, I should say, persists because at some point, um, you know, the, the details matter.
And when you're talking about, you know, data within a company and what that means to a company in terms of, you know, what it represents and, and the risk that it poses, if it's not managed properly, those details matter. You know, if, if you screw up with your Claude bot and it accidentally spams everyone in your, you know, email lists, you, you may lose friends and influence. But, you know, at the end of the day, it's things like that do happen in real life too.
Not just in the virtual, you know, age agentic world, but with the database, you know, those small things that, that matter the most, I, I think could, can be abstracted, but never, uh, eradicated, uh, because they do matter. And so from, you know, in the, in my research area, I, I see all the time, um, these same ideas playing out that we're seeing with things like cowork and things like open claw mm-hmm. In that we're, we're trying to take an build another layer of abstraction that allows, you know, the practitioners to stop thinking about the syntax and start thinking about the intent or the outcome.
Uh, you know, what they want to happen. I want my data pipeline to keep running. Might be a nice intent and outcome.
Um, be nice when don't that way, yes, it's supposed to, but it always did. It was supposed to work that way. But technology is always shifting.
It's always moving. And because of that, you know, there are known unknown unknown unknowns that were always gonna bite you and do believe that, um, the tooling we're talking about, you and I talk about every week is, is capable of not just smoothing out and, and removing some of the barriers and the friction you talked about, but actually, um, eradicating some of those unknown unknowns, or at least addressing them in a more flexible, more comprehensive and more immediate way than we mere humans can. Mm-hmm.
Um, mm-hmm. Will they, you know, evolve into that? I, I think in small pockets, yeah.
I think that, you know, and it's always the case with any sort of scientific endeavor. You need to have a dis, you know, circumscribing universe that you can work with to understand, and then you can make it predictable, and then you can automate based on that. And so, you know, for specific tasks like, you know, query plan management, let's say we've been working on automating that for a long time.
And ag agentic AI is, is, and sorry, generative AI and ag agentic, you know, implement, uh, application of, of that idea to generative AI is, is making that better. Um, so I think we'll just see that across the board, man. It's not gonna make databases no longer matter.
'cause they're always gonna matter. And even if you say, well, let's just switch to object storage, we don't actually need relational databases anymore. It's just the file system anyway, you're still gonna be imposing some sort of structure on top of that because the details matter.
Knowing that you write a record having that, yes, it did get written, even if I lost connectivity matters, still have to have some governance and some security and some data protection and some data backup and re recovery and all that's true. And then, and then the other side of the coin is, hmm, what if data doesn't have the same gravity that it does today? Maybe what if it could move much more easily?
Right. That's one of those take away constraints. You hear me talking about constraints all the time?
Yeah, yeah. When constraints get removed, that's when things really drastically change. Just like, you know, the value of code today has gone down is going down every day.
Not because the code itself isn't valuable, but because your ability to produce it isn't the value of writing it by hand. That's what's collapsing a much bigger quantity, not code. Yeah.
And so, you know, it's, it's, I may have said this before. When I first started with these tools, my first reaction was, wow, pretty much, maybe not this moment, but pretty much, if there was anything that I wanted, I could probably build it if I wanted to do that. If not today, I certainly will be down the road.
And not to say everybody's gonna have their own email client that they're gonna build by hand themselves, maybe they will someday, but if they really care about it, they will. Yeah. It kind of like, if I don't like it and there's not one on the market, yeah, I could, I could say make it like this, but I want it to be this way.
And, you know, that, that that can be true of many, many things. Um, so let me ask you, if you were gonna start out today, and you were gonna, I want a productivity, ai productivity assistant, and I could go frontiers, if I'm an in an enterprise customer with open ai, I could go, uh, cowork with, with Anthropic, I could do, you know, the Microsoft, I could do a, you know, agent force. I could do all of these things.
So there's a business context and a personal context. Let's talk about in inter personal context, would you start out with Cloud code today? If you were starting from scratch, would you do or lemme, sorry, open, open coad, or would you do something else?
Would you start out with, I'm gonna vibe code and build my thing and maybe I'll try open coad later? Or I dunno, what, what, what you Well, I think, I think knowing that it's a part of, of open AI would, would lead me as, as an enthusiast, you know, playing around at, at night on the weekends, uh, looking somewhere else just because, you know, we, we all, you know, especially those of us who've been in this industry for a while, do tend to see things through a bit of a cynical, you know, lens. And, um, you know, if you, if you are a proponent of, you know, Richard Stallman and, and the i the ethos behind, you know, suckus, you know, open source, free and open source software, you might question this acquisition.
Yeah. Um, but it's already happened. It doesn't matter.
And that's the beauty of open source, by the way, is that how many forks are there already of open claw, um, thousands of them. And that's gonna be the case going forward. And there are a number of alternatives to play with.
And that's the beauty of technology. And, and like you say, if you collapse that time to, you know, value for building something, um, then why shouldn't you try this on your own? Why shouldn't you say, for my use case, what I want it to do?
Uh, I don't need all of open claw. I don't, I don't need it to, to you work with 3000 connectors and all that. I just have one use case that I want it to do.
Mm-hmm. Um, you could, I think in very little time and with some degree of, of confidence, build something yourself. So I I, I would build it.
I would build it myself. I really would. I would just rook up with my favorite, you know, and, uh, uh, agent, agent XCLI, um, whatever that is, let's say vibe from Mytral, for example.
Doesn't have to be, you know, codex or, you know, Claude, uh, CLI. Uh, and I would, I would just learn how to build an agent agentic, you know, plane a control plane, an agentic control plane, because that's a fascinating area just to dig into, like you and I were talking about with how you manage memory, how you, uh, build in connectivity to different sources of information and, and take action in different modalities. That's fascinating.
I think that that's, that's one of the things that, you know, perspective is worth a DIQ points as Alan Kay Apple Fellow used to say, I don't know who discovered water, but it wasn't a fish. Right? No.
Um, it goes way back anyway. Well, you, well, you and I approach it like, like I'm sure a lot of people do, which is I not only want, I want that tool, but I wanna understand it and I wanna understand what it takes to do it. And, you know, like, you know, you mentioned agent control planes and things like that, and I have this concept of observability native and Oh yeah.
So I'm already like in the stuff that I'm doing, like, okay, so I'm gonna go, go like, build my own version of this and see what it, see how I would do it, right? Or what you could do, or what some of the constraints are, or you taking it the next level. And not that I'm gonna contribute it all back to try to become the next open claw famous developer.
May maybe I would, but I'm really, some people might want learning and experimentation, right? Mm-hmm. So there's that whole mm-hmm.
Wanting to know what could I do? How would it work? You know, there's that part of it.
Well, your general everyday user. I mean, I think you're, you're gonna talk about, um, and then one of the segments of sort of like the easier, the, the open Cloth for Dummies version, not to disrespect anything, but if you know what I mean, there's, well, that's what Peter wanted to build, right? He wanted to build something that his mom could use.
That was his idea. Yep. Yeah.
I'm, I'm very much of the, do I just like plow down and go do an open client instance, or like, do I keep doing what I'm doing and I keep doing what I, I keep doing what I'm doing, which is building it myself because I rebuilt it multiple times. And one of the reasons why I do that isn't that I'm happy, I'm happy with it, is how I built it to today. If I, you know, if I use today as the example and six months ago is a much different process using the same tools, just new generations of 'em, right?
Whether it's philanthropic or open ai, you know, codex or Gemini is, you know, the, the ability and the things we've learned about doing plans and things like that and how to better instruct control the memory of these systems, of these tools. It's doing it today is much different experience than even six months ago. It's, it's odd to say that, but it's true.
Well, it's because like six months ago, if, if I spent the weekend building a tool to do something like, like just drill down into fractals, 'cause we all have done that. Um, so, because all, we've all been, I wanna fly through the fractal in real time, that's what I want. Um, but, uh, you know, in the past, if I had concluded, you know, after two days of work that wow, that really wasn't the greatest approach that I took to building this tool, I would never have thrown it out.
I would've been like, let's just keep working our way toward getting what I want with what we have. Mm-hmm. And that collapse of time to value with a gen, you know, code generation, you know, what is the loss?
Just time and tokens at this point. And there that is increasingly a shrinking, you know, uh, asset in terms of cost and time. So I would just redo it.
Yeah, to your point and to your point, the experience, my experience yours probably too, is when you redo it, you can do it faster the second time, the third time, the fourth time now. Yeah. Because you just know it better.
You know what you want better. Yes. That's a big factor.
But the tools have changed. I mean, I'm not sitting here arguing with whatever model about, no, that's not, not right, that's not correct. That's not what I was describing.
Or you didn't do what I asked you to do. You know, you're not, you're not dealing with the same level of things of Yeah. You keep coding the wrong air fixing, you know, fixing the air that you keep breaking or breaking what you fix.
Which, which ar ai are you using? Because that still happens to me. Well, no, we really wanna know.
Okay. No. So, so I, right now I'm doing my development in, in Clot.
I mean, that's what I'm using at the moment. And I think it's four six extended or whatever it is, the thing kills. It's just like, wow.
Yeah. Yeah. It, it cranks and, and, you know, and that's, there are some really good things about, you know, codex and, and Gemini and what they do as well.
It just kind of, the ball moves around a little bit. It's like fluid in a, in a jar. You just keep shaking up.
And this one's at the top now. Um, but that's why I did is I wanted to see like, okay, what's the state of this? It's been a few months since I've done anything with it.
And it's like, ooh. Yeah, you gotta get to know it. You have to learn the ins and outs.
You have to, yeah. It's just like dating, um, back, back in high school, college days, you know, you need to get to spend time to to know school, college. Okay.
I was, I'm told I in high school, college, but yes. Okay. Not to, not to jump on the stereotype, but it's a stereotype.
Um, for reason. Yeah. For for reason.
No, actually, um, that, that makes me think about, about, um, call calling out, uh, acquaintances. And, and we, you and I forgot when we started today's podcast because we we're, we're so enamored with, with, uh, you know, Peter and Sam's, uh, you know, getting to know one another, uh, that, uh, we forgot to, to do, uh, our new segment on, on the call outs, uh, boy, which by the way, shark, if those of you, we, we went under the shark. Um, it's, it's not two words.
It's one word to call out, like in text to call attention to something. Mm-hmm. Mm-hmm.
So do you mind, let's stop and do it before, before we forget, move on. Yes. Move on to the next day, day.
Apologies to our audience. You know, I, I vibe coded my way right past the call out, so it's the backup here. It's time for the call out.
Yeah. This is live. This is live.
Yeah. There you go. Go for it Brad.
Call out. So, so mine, mine is, uh, a surprise for me. Uh, you know, I took a briefing from a company called Quest Software, which is a bit of a large company, global company that has quite a history.
Uh, and they are very much interested in what you and I were just talking about with, you know, what is the gravity of data, what does that look like going forward in an age era? And so they just are getting ready to, by the time this goes live, it will have gone out, uh, this new, uh, solution from them called the, the trusted data management platform. Uh, which you're like, yeah, there are a ton of those.
And yes, there are, I actually am gonna be publishing a, a new, um, signal report on said data intelligence platforms in just a couple of weeks. But what, what's interesting about this is they're, they're really focusing on creating what they call an automated data product factory. So we've heard about AI factories mm-hmm.
And this is a data product factory, which everyone that's, you know, a data practitioner or knows one will, will have like, seen them or have lamented themselves the fact that how difficult it is to create data products, even though it is a grail, that is something we all want. But, you know, push comes to shove, you're like, oh my gosh, just get them access, get them something that works, and let's call it a day and we'll come back to that problem later. Um, but so if you could build data products, especially if you can like tag on SLAs and let's say contracts to say, you know, that it's gonna do this and that, and this is what it represents, and those are important things to business.
Uh, that's what Quest is really shooting for here, which I, which I find very laudable and, uh, I wish them luck with this. It's, it looks, from the short intro I've seen to it in demo, it looks like it's a, a very solid, uh, step in that direction and something that will alleviate a lot of pain for, for DBAs in particular, and, and any data practitioner, frankly. Excellent.
What about you? What's your call out this week? Uh, you know, I actually have two, um, and both act do come from briefings as well, and some, I can talk about some I can't quite yet.
But, um, you know, two of the really interesting, and I actually had three really interesting between IBM and, and Red Hat and GitLab. Um, and you talk about different companies at different stages of their development. Of course, red Hats part of IBM, but they still kind of operate.
They are Red Hat product wise, you know, on their own, on their own flow, if you will. But what the trend that, that I think is happening is, you know, we're all enamored with chips and then we're enamored with data centers, and then we're enamored with all the stuff that takes to run a data center. And then eventually we kind of get to, oh yeah, we need to run things on this, right?
What are we gonna put on this? What we gonna use this for? Oh, yeah.
Back to software. What was that again? Let's, let's think about that a minute.
And, and, you know, I was, I sort of jumped my own, my own project here. 'cause I talked about the, um, agent control framework, uh, plane control framework that I'm working on that I will get to you to review. But I jumped it ahead of it and said, no, I'm gonna do, I'm gonna do this idea of observability native first.
And then the concept native, native meaning it's observability, is built into the whole stack from the beginning, whether it's development or what you're deploying operations in your code in this everywhere in the stack. Because in a, in a world where everything can change because of ag agent and processes and ai, and any part of that stack can be replaced at any time. And the non-deterministic nature of generative AI watching telemetry at the end of the trail, you know, the river's been polluted by the time it gets to you, your, your chances of stopping something, it just decreased dramatically.
Right. But you may have figured out what happened Yeah. To prevent it from happening.
You're playing detective for something, you may not have any data for a crime that will never be recommitted. Exactly. Yeah.
That repo got blown away and it's already something else, right? Mm-hmm. Um, so, and it's not to be, you know, too overly optimistic about it, but I think we're at a point where actually we can build in telemetry and, and observability in everything as it's being built.
And that's what you see vendors stepping up to try to do to get early in the development lifecycle, whether it's a Dynatrace or, or harness or Red Hat and IBM and GitLab. They're all figuring out how to be part of the whole ecosystem, not the slice, just the slice that they've figured out they wanna be part of. They wanna own.
Yeah. They still wanna own that, but they want, they know they've gotta participate in the full lifecycle. And so, uh, just a, a kudos to those companies, all all of them that I've mentioned who are really thinking aggressively about ai.
Not just hop on the latest, you know, open claw bandwagon or by coding bandwagon, but to really think about how do they rethink their product strategies and reshape it so they are part of how software gets created, how it gets tested, and how it gets deployed, how it gets designed and thought of, because that loop is going and pretty soon, if you're not in it, it's too late. It happens too fast, it's too late for you to get part of it. 'cause it's already come up with another way of doing it.
So I think they're doing it out of necessity, but to also out of opportunity. So lots of callouts there. Um, but I thought, um, you definitely expect a trend and now it's happening.
It's like interesting to see it come together. Do you, do you think, Mitch, that you know this because I, I love that idea of innate, uh, observability and, you know, even with like o open claw for instance, uh, Peter built into that, this, this observability where every single transaction gets documented. Mm-hmm.
So you can replay what went wrong later on mm-hmm. Mm-hmm. Or at least understand it.
And I think you could use that to understand what's happening as it's happening, as you said in the stream. But do you, do you think that, um, if every piece of software that gets built and gets conjoined within this, you know, control plane will need the control plane to be sort of a, you know, uh, like with, with catalogs, you have a catalog of catalogs idea. Will we have a control plane of control planes wherein every component is self trans, you know, transcribing what it's doing to be, to be documentable.
Um, but then that harness, not just harness, sorry, that control plane itself is the master of that. Do we need that? You know, I, I think that's, that's one logical line of thought.
And I was thinking that way for a while and I thought, you know what, in our world there, there is no one product that dominates a company in, in large companies, they have everything. Right? So they're not gonna sign up for, name it Microsoft to control everything, or Dynatrace to control everything or Yeah.
Open AI to control everything. Uh, even if they wanted to, just like everything else, we are an Oracle shop. Well, we just bought this company.
Now we're not, we're Microsoft and Oracle shop, right. We're a little less Oracle today. Yeah.
Yeah. We're just squeezed it a little bit, maybe a lot depending on how much you bought. Um, so I think, I think the world we have to anticipate and expect is a control plane of control planes a, a multidimensional, it's called, you know, this is the multiverse of control planes, if you will, essentially, um, you know, which, which may sound like chaos, but on the other hand, that's, that's the world we're ending up in.
I I think that that's unavoidable. So now how do control planes talk with each other? And how do they each operate within their own, with their own, within their own domain, and then when they intersect either intersecting in a good way or intersecting conflicting way, how does that get resolved?
And, and I think the reason why that's possible to do now is everything doesn't have to be monitored after it's already been built and done, is you can have observability happening in the system of the system by the system, by, by external agents, whatever it is, and be able to direct what's happening while it's happening. Let's say an agent, you know, makes a, makes a call to a, to an LLM and gets a wacky answer, but the guardrail didn't catch it. But you have an observability guardrail that will also tell you something wacky happened.
Let's redirect that agent for a moment, have it do something different, or have it apply another agent check its work or whatever the response is. So it's, it's a much more of a, an organic kinda live system as opposed to a linear, this happens and then we catch these things at these point in time, you know, the river doesn't flow one direction anymore. It flows all directions with I think the world.
We're going to wait, isn't that a flood? Yes. Yeah.
It's called a tsunami or mud side in California, whatever you wanna call. But you know, it's, it's okay because then you're, you're not gonna hit yourself on rocks. You have room to swim.
Yeah. I didn't say it would be easy, but, you know, and so yeah. It's not like that's not happening tomorrow, but it, it seems to me whenever I think like, yeah, this is gonna be the thing that wins, eh, it doesn't usually happen.
That's pretty rare for a thing to win. You know, even in the mobile world, there's still Android and iOS, there's still two of them. Right, right.
You know, and there's, you know, the chances of one of them truly, really dominating isn't gonna happen, I don't think could happen. Maybe if, if you know Tim Cook gets elected president, then we'll have iOS everything. Well, right?
I mean, Linux dominates. Um, well this, but that's the choice of the people. That's the people and, and everyone who's building servers and mobile operating systems and Chrome os and, uh, you know, every, every other os it seems, except one, um, yeah.
It's, it is, it is like a, you know, a multi technological landscape, you know, where, where you have to be adaptive and you have to strive toward building software that, as you just described. So ELO eloquently is able to, you know, see itself to have some, you know, internal reflection. Like, like I rarely do myself as a human, but as I strive to do self-aware, right.
To, to, you know, be able to say, huh, maybe if I did this, I might be better off. Let's, let's try it. So when we, we evaluate software in the future, we'll be evaluating a lot the things we do now, plus the eq emotional quotient.
What's the emotional quotion? That's right. It'd software be free association on the couch for your software.
You know, how does that make you feel, Mitch? That's all we should, um, back when I was a young boy. No, that's a whole nother problem, but, oh dear.
Yeah, that sounds dangerous. Different kind of podcast indeed. But, uh, anyway, we, we should do our last segment before we go.
Um, yeah. Since, since we, since we changed our order around so much, we might, we might as well, you know, keep going down that up, so, alright, so time for the drop. Okay.
It's time for the drop. Alrighty. I think I, I don't remember if I went first last time.
I'll jump first time and then you can jump in. So I mentioned I have a couple things coming out. The agent control plane framework and then also the Observ observability native ideas.
And I'm excited to share that, not just to get it out the door, um, but also get, get some real tests of it and people to say, this is crap, this makes sense. Make it better this way. We like this, whatever.
But my hope is to get, to make what we work on Brad as something that people can look at use as a, as a guide or as a framework for testing and evaluating what's happening in the market right now. Using the frame of, you know, what we know as analysts and also, you know, our own background and experience. So that's a big one.
And then kind of the other one is just marches right around the corner and I start traveling again. I haven't traveled too much, so I got, I'm, yeah, same here. Getting back out the, you know, getting the band back together, like, okay, what do I have to have in my suitcase again?
What do I have to be ready for to leave at a moment's notice? So, yeah, you'll figure it out once you get there. Don't worry, Mitch.
It'll be okay. It'll happen. It's like riding a bike.
I, I'll fall right back into it. So, anyway, so that, that's my job. How about yours?
Yeah, same here for the travel. Uh, looking forward to that. It's, it's nice to, to see colleagues and, and to work with the vendors that we, we get to spend time with.
I really, I feel very fortunate in being able to, to share, uh, in their, you know, everything they share with us and the knowledge they have to impart is wow. Um, but yeah, so Weekend project, you know, you know how I have said I will not do, you know, open claw, uh, just because I don't have the time or, or energy to, to put into doing it right. And doing it in a way that I feel confident that it's the security is set up, you know, that I could trust it.
However, that said, I, I, um, wanna try this with, uh, an alternative. You know, I mentioned this a couple times today that I wanna play with one of those alternatives. And the one I'm looking at right now is called Nanobots, which comes out of a Hong Kong University.
And the reason is twofold. First is, uh, it's written in Python, uh, like 98%. So I like Python.
I feel comfortable, um, you know, reading Python much more than type script because as, as you know, we've talked about, I hate semis. Um, but the point is this, this, this is the second part is it's only like 4,000 lines of code. It does Wow.
You know, it's, it's not as aggressive as open in, in terms of its scope that it wants to accomplish, but it does the same thing. It's the same kind of, of Agentic system, uh, but it's 4,000 lines of code. And I think I could get my head around that, um, a little bit better than I could with the like four 30,000 lines that Open Claw is sitting at right now.
Woo. Robot's not going to, uh, start its own religion or create a political party. Ah, well, unfortunately, the Communist Chinese or whatever, right?
Let's, we should not get ourselves every model you interact with, every software that you use has a perspective and it came from someone or some group of people and they have their, you know, agendas and objectives. Yep. So always, you know, understand who you're, you know, who you're dating, it's the important thing here, but, but it's choose your friends wisely.
You do. And, and so I'm choosing Nanobot very carefully from that regard. And, um, it's interesting because it actually has built in the ability to participate, uh, in the, um, what was it called?
Um, T Book, uh mm-hmm. Mm-hmm. You know, social media, social network, you know, the Reddit for, for these claw bots, the Red, the Tin My goodness of of Coping Claw.
Yes, yes, yes. Let them date one another. Yeah, I'm fine with that.
Yeah. So that, that's my drop is, uh, I'm looking forward to, to, you know, standing that up on a, a nice, you know, uh, UTM in, you know, container, uh, on my, on my spare computer, middle Mac Mini and let it go, because God forbid you can't get one for three weeks right now. No.
The wait lists. Yes. Yeah, everybody's digging out their old Intel based Mac Mini.
It's like, I mean, I could do it on this, but anyway, whole nother thing. Well, cool. Well, good.
Hey, uh, kudos on, uh, you know, we had our first guest last time, um, joining us, and we appreciated having, uh, folks here and we hope that others will consider being a part of our, our broadcast, part of our podcast here. com website. And we thank you very much for listening and following and just tolerating maybe some of the points here and there too, and just going along with this for the ride.
It's been a lot of fun. So, and it's a lot of fun because I'm doing it with you, Brett, so thanks for being talking. I feel same too, as, as our, our producer Corey just reminded me, um, we're, we're definitely going places.
We'll be back with an episode we might call Agents After Dark, uh, given, given the, the topics that we've been talking about on this one. So thank you all for joining us on this crazy journey. We love having you here.
We may have to start rating our episodes, so, okay. That's right. NC 17.
Yes. And m Alright, take care everybody. I'll see you next time.
Control. This is agent dev. I'm in position.
Copy that. Dev Stand for Go Standing by. Hey everyone, it's Alan Shimel, founder, CEO here at Techstrong, and welcome to our continuing series on, uh, AI agentic AI in the future here with, uh, the Microsoft team and our fu um, analyst team, as well as Techron.
In this next episode though, we're gonna be joined by Mitchell Ashley, uh, of Rum, who leads the software development lifecycle and building segment at fu. And Mitchell is talking with Brian Good, whose official titles is corporate Vice President and Agents marketing, but Brian is really here talking about agent apps in chat, and it, it's, uh, you know, obviously a very hot topic as we move to a ent ai workflow based basis. So let's join Mitchell and Brian here with me, and it's great to have them both.
My name is Mitch Ashley and I'm VP of practice, lead of the software Lifecycle engineering practice at RUM Research. And, and Mitch, uh, my name is Brian. Good.
Uh, I lead the business applications and agents team, uh, here at Microsoft. Let's start here. 2025 is described as kinda an inflection point for AI adoption.
What do you think are the most significant changes that, uh, you've seen in organizations as they're using AI and agents in their businesses this year? Well, I absolutely agree to 2025 is really an inflection point, and we'll sometimes describe it as the year that the Frontier Firm was born. And you might've heard us talk about the Frontier Firm before.
It's this idea of companies that are putting AI really at the heart of their business, and it's enabling them to do things like reinvent the way they engage with customers or transform their business processes inside their company, um, and beyond. And it's really the year these frontier firms are sort of rising up and a time when I think we can learn a lot from these early adopters and understanding how they're deploying ai, how they're being successful, and then figure out how we can take those insights and bring 'em to, to our own businesses. That's where you can have outsized impact As AI transforms those functions.
What do you see as the new patterns of work that's enabled by copilot and agents as customers start leveraging the technology for productivity or innovation? I'll tell you what, I have studied these frontier firms as they, as they come up, and there's really three patterns that I see across these frontier firms. The first is really enabling, uh, employee productivity.
So they give every employee a, an AI assistant like Microsoft 365 copilot, and it helps them, those employees be more productive. The second pattern that I see is, uh, really these frontier firms deploying AI to automate existing business processes. So, for example, they already have a way of handling expense reports, but they can use AI to speed that up and reduce costs and, and that certainly results in some benefits to the customer.
The third pattern is really where a customer like starts from the beginning, let's say from first principles, and they reimagine a function altogether. They'll reimagine what it means to engage with a customer who has a, uh, an issue with their product, and they'll put agents at the heart of that. Most companies can only take on one or two of these functional transformation projects at, at any given time because it's a big lift.
Again, it's reimagining a function from first principles. It's not just taking existing processes and and applying AI to them. How far along do you think most organizations are in their adoption cycle for generative ai?
Well, I think we're still in early innings, uh, that there's no doubt about that. Um, you know, customers are starting to deploy AI and different functions as we talked about, but certainly they haven't, in most cases, fully realized kind of the transformation that AI can have across their organization. So it's still very early innings, but I'd say we see very promising signs and, and green shoots, uh, of that, uh, of that adoption and success.
Uh, again, the key really here is to start function by function, think about a particular function, think about peeling it back to the business processes you need to go focus on. That's where you can have outsized impact. How do you define AgTech business applications and why are they critical for organizations?
Yeah. Well, I do have a vision that there's, uh, the application of hold is really transformed with ai and we call that new type of application, the AG agentic business application. And the ag agentic biz app includes the assistant for the human to start to use.
It includes prebuilt business process agents that basically take the drudgery out of work, and then it's built on a data foundation. And so instead of being limited just to the data that, you know, maybe is in the CRM system or the ERP system, it joins that with data from other lines of business systems and even productivity data so that you have a rich set of data that you can build agents on top of, and you can empower your humans to get better decisions from. And so this idea that brings all those together, the assistant, the agent, the application, I call the AG Agentic biz app, and I think it's a really big idea.
Let's talk about why you're optimistic about the future of ag agentic business applications. I have a lot of optimism here because number one, I see customers already deploying these applications, uh, and and really transforming their business. So I already see people starting to get great benefit from it, but at a more human level, the thing that gets me excited is like, if you think about every one of our jobs, like there's a lot of stuff that we end up having to do that really isn't adding joy to our lives.
Uh, you know, me doing an expense report or, you know, filling out a CRM uh, system, you know, with an update from a customer call, those aren't the things that, uh, that make us, uh, unique. Those aren't the things that bring joy. Those aren't really even things that add business value collectively.
But if we can delegate those things to AI and agents, I think we'll enable ourselves to actually go do far bigger things and really take our ambition to the next level. And so I am absolutely excited for what the future holds. Yeah, I'd love to hear about how you see the role of AI agents evolving just over the next few years, especially as organizations start to redesign core business processes drive outcomes for efficiency.
How, how do you see this taking shape? Yeah, great question. You know, I I I would say that we really believe there's a spectrum, uh, of, of agents.
You know, there will be very simple agents that someone will use that might just be grounded on a particular knowledge source and help you, you know, understand, you know, what, what might happen from that, uh, from that knowledge source. Then there'll be task-based agents, and then there'll also be more sophisticated, fully autonomous agents as well. And this more autonomous agents is where you can unlock a lot of business value.
These are agents that, you know, aren't called by a by a human. They're just running independently. They're triggered based on different actions, and they can really automate business processes in some cases from start to finish.
So we believe there's this spectrum of agents, and today I'd say most ENT use cases start with those more simple kind of knowledge agents. Uh, but increasingly we're seeing customers bring in these task-based agents and autonomous agents, uh, to automate, uh, things, uh, end to end. How about the C-suite leaders?
How should they be thinking about investing in agent technologies for long-term value? Well, every C-suite leader I talk to today is already in on ai. Like every one of them recognizes that it's a competitive advantage if they can move quickly.
And if they don't move quickly, they recognize it could be a disruptive force in their industry. But let's face it, ai, it's transforming businesses, it's transforming functions. It's, it's gonna reshape entire industries.
And every C-suite leader I talk to recognizes that and once on board. What they're looking for though is a partner. They're looking for the tech, of course, they wanna make sure they've got the right tech, but they're also looking for a partner to help them shape this in their business.
And that's where Microsoft, I think, comes into play. Uh, you know, we're not a large scale model maker, uh, but we do take the best of the models and bring 'em into the workplace, uh, so that companies can use them. You know, we talk about AI being a disruptive technology, probably because it seems that it really can and will affect every part of our work, our lives, et cetera, certainly is affecting how we create software with new concept like agents and agentic ai.
Curious about your thoughts. Share with us, how does Microsoft view what the transformation is going to be like with AI really having an impact and a big benefit to businesses? Yeah.
Many industry, uh, pundits will say that this move to AI agents is gonna, uh, lead to the rise of, uh, uh, more consumption like models or outcome-based pricing. And I think they're right. I, that's definitely a direction that I see the world shifting as well.
Um, the only thing I would, uh, balance that with is that many customers are still, uh, you know, most comfortable buying things on a, a per user or pre per seat basis. And so, uh, the approach I'm taking is, you know, how do I enable customers to buy offerings that they're comfortable with? And that's typically like a per user or some type of per tenant type of type of license, while giving them the flexibility to, to grow and shift into more consumption models as they, uh, as their business changes.
And so, uh, we are at an inflection point, just as we said, and I think one of the things that will change is the business model over time. Talk some more about ag agentic when you think about agents operating on their own or more autonomously, and why, why is that an important thing that organizations are looking to move to? Yeah.
It really comes down to business priorities. Like every business leader I talk to wants to find ways to grow their top line revenue, or they're looking for ways to automate, uh, things that, that they're doing so that they can save money or redirect folks to, uh, focus on more important activities. And, you know, autonomous agents really fit that bill.
You know, imagine in the sales context, you can have an autonomous agent, you know, going through marketing leads and qualifying them before handing them off to a human seller. That's work that wouldn't have gotten done in the past or would've been done by a human seller, uh, and have been relatively low value, not something they, they really enjoyed, uh, about their job. And so an AI agent can do that and add great value to the company, and again, help that company grow on the top line.
Talk about Microsoft, uh, and the offering that you have in the context of an end-to-end tool toolkit, if you will. Yeah. For functional transformation.
You know, as far as the toolkit goes, we really believe that there's three essential parts to it. The first is we think every employee should have an AI assistant in our case, uh, that's Microsoft 365 co-pilot, uh, think of that as a productivity tool to help every employee work get through their workday and get more done. That can be quite transformative.
That next step up is where agents come into the picture, and I describe agents as really being for every business process or workflow in an organization. And we have, uh, a toolkit that enables customers to build their own agents. It starts with copilot studio, but even extends into our Azure capabilities with our Azure AI Foundry product.
So agents pair very nicely with that copilot that I described first, and then the final step is where the system of record becomes the system of action. And we'll sometimes call this the agentic business application, where you take a CRM system and you add agents and an assistant to it to really transform those three are the essential ingredients or building blocks for AI transformation, um, in a frontier firm or any business at this point. That's a great term.
Can you share some examples of how Microsoft customers are already seeing measurable impact from adopting agent business applications? Yeah. One example I I think I can give is lifetime.
Uh, they're a great customer of ours, uh, based here in the United States. They've used us as part of their finance and supply chain operations, and they've, uh, deployed agents to basically speed up how they handle and process e-commerce orders. In fact, I think it saved them 95%, uh, in terms of their order, e-commerce, order efficiency by deploying, uh, AI agents and agentic business applications to solve that problem.
So I think that's a great example. We also have, uh, another great example, uh, from Europe, um, uh, a large utility named Enco who deployed a multi-language, uh, AI agent, uh, for their customers to help them scale and address customer questions. Uh, it's a pretty cool solution, saves them time and again, helps them scale up.
Now, it's a, it's a really exciting time in the world of agents. You know, you talked about the C-suite in the, kind of the three phases in this adoption curve as we adopt ai. What, what kinda recommendations do you have of like, how to get started?
Well, maybe near some of the near term activities? Well, you know, as I said earlier, I think AI is gonna transform every company, every function, every industry. And that opportunity is so vast, sometimes it's hard to know where to get started.
And I've got two bits of advice really there to, to anybody that, uh, that is, is pondering that question. Uh, the first thing is, again, start with a function. Pick a function that you want to go after and then peel it like an onion.
You know, go look at the next layer, which are the business processes in that function that you can apply, uh, a co-pilot or agents to, to really transform. The other thing though, is like, don't get into analysis paralysis. Just pick a business process.
Just go pick a business process to get started with. And as you learn from applying AI to that business process, whatever it is, whatever your authority is, business process is, go apply AI to it. You are gonna learn, your organization's gonna learn, your culture will adapt, and then it will flow from there.
So the opportunity is so vast. Don't let that keep you from getting started. You gotta get started.
Start simple, pick one thing and go from there. So as we move to an agentic business environment, let's talk about the people. How do you see the role of human creativity, judgment, leadership?
How is that gonna evolve? Yeah. Well, that's an excellent question, and one of the things that we'll often talk about is how AI and frontier firms is gonna transform the way we work.
It's gonna change, uh, the org chart into more of a work chart. Uh, we sometimes talk about, uh, a frontier firms taking on the Hollywood model where individuals will swarm around a problem, like focus on a problem and then disband, uh, you know, uh, when the, you know, once they've got a solution. And so it's absolutely gonna change the way we work with others inside the workplace.
It also, I think, is gonna give rise to a new idea that we call an agent boss. And you can imagine, just as a people manager today might take work and delegate it to different humans on their team, uh, you know, resolve conflicts and sort of manage performance. Every one of us in the future is gonna do that, but not just with humans, but also with agents.
So imagine taking a business problem, breaking it up into pieces, delegating it to agents or agents, uh, on your team, resolving conflicts that might come up, applying human judgment. Uh, it's gonna be an absolutely transformational moment, and I'm really excited about it. I, I really believe that, you know, there is still a huge opportunity for humans with human ambition to go do great work amplified by ai.
Talk a little about, about how you see the, the difference in the working together between applications, assistance agents, the different technologies. Well, that's an excellent question. And, you know, we really have this complete toolkit that spans everything from the assistant to the agent, to the application.
And I think those three things work together in a symbiotic way. As an example, you can imagine that I might go to my assistant and, you know, ask a simple question, my AI assistant, like Microsoft 365 copilot, and ask a question about, you know, uh, summarize the emails or help me respond to the emails I've got. Or I might use an agent to update a CRM record after I meet with a customer, but I'm still gonna want to go to an application.
Really, that's a purpose-built experience for me if I want a more specialized or fine tuned experience. So those three things really work together. Talk a little bit about the industries or maybe the business functions that you expect to see reshaped first by agent AI transformation.
Yeah, there are three or four, I'd say functions in particular that are, I'd say, um, you know, ground zero for, uh, functional, uh, transformation with AI and agents. Certainly customer service and customer experience is one that is, uh, absolutely being reshaped and say a very early adopter of ai, no doubt about it. Another, where I'm seeing a lot of early AI adoption, uh, is in sales.
Uh, and it's just because, just as we talked about, there's a big opportunity to use AI to basically increase capacity for that organization to grow top line revenue. So the business impact there is undeniable, but I also see it in places like finance and supply chain, where you can use AI to shorten the time it takes to, from an order to actually being able to ship it that order. We're seeing people use AI to improve accounts, uh, payable and accounts receivable.
Um, so we're seeing some pretty, uh, interesting impacts there. Let's Talk a little bit more about the frontier firms. You know, they're not just adopting technology, they're also changing the way they're do doing business around AI agents and co violet capabilities.
You talk about, you know, what they're doing to invest and support the short-term ROI that they looking to get for long-term innovation. The Most successful frontier firms are doing, actually is taking a functional approach. And so certainly they'll think about their entire company, uh, but they'll really take an approach that's function by function.
They'll think about their sales function as an example, and then they'll peel back the onion a little bit and understand which processes inside their sales department they can automate. Uh, using AI agents as an example. They'll look at which, uh, things their salespeople need help with, where you could pair up an assistant like copilot to help them throughout their workday.
And they'll even look at how they can bring agents in to really augment business capacity, maybe to grow top line revenue or take out costs. But starting function by function has really been the recipe that these, uh, frontier firms, uh, are using to see success. As an example, in our sales organization, uh, by deploying co-pilot and agents, we've been able to improve revenue per seller in some cases by almost 10% in some organizations.
And, you know, if you think about that, giving a seller 10% more capacity is like giving them an additional month, uh, in a year, uh, without actually having them spend any extra hours through the work week. And so there's some pretty remarkable results. That's just sales.
We've been able to do the same in customer service in our finance department, in our IT department. Our legal team has been able to reduce costs by 5%, uh, by deploying AI and agents, uh, within their, uh, within their functions. How do we ensure trust and transparency as a genix systems become more and more autonomous?
Yeah. Well, there's two things that we want to do to help with trust and transparency. The first is we have a rigorous set of principles on, uh, responsible ai.
So for any AI that, uh, Microsoft deploys, we adhere to a an important set of guidelines. Uh, and that's really table stakes. So that's the first piece, uh, responsible AI and our focus there.
The second thing that I think is important is we now have an opportunity to start to quantify the value and the impact that many agents will have. And we often will call that in the industry, we'll call that evals or benchmarks. And increasingly, I think we have an opportunity to help our customers understand how agents are being effective in their workplace using these evals and benchmarks on real world problems.
So, for example, uh, we, uh, uh, a typical, uh, workflow will be a sales leader doing sales research to try to understand like how they might want to organize accounts or territories or, you know, reshape planning as they think about the year ahead. We recently released a new benchmark, uh, that we call the sales research bench, and it shows how agents can actually help sales leaders in that very specific job, and it's quantified. Uh, and so I think you'll start to see more of that.
And between following responsible AI standards and then using quantitative measures like evals and benchmarks to understand efficacy, I think we're really on the cusp of helping customers know how they can deploy AI in a safe way for real business results. Well, thank you, Brian, for sharing with us your insights and gonna look into the future. What's happening with the Gentech business applications?
Thank you very much. It's really amazing the pace at which AI has been adopted and continues to evolve. The technology evolves on a near daily basis, but at the same time, organizations have to figure out how they're gonna implement their AI strategies and what the business outcomes that they're most important to their business.
I think it's very fascinating how Microsoft has approached the market, both from the standpoint of addressing the individual and their productivity, but also thinking about new workflows, new models of business, but doing that with not a set of tools, but a set of capabilities that provide integration with data process, workflow, and agentic ai. No one knows for sure what that future holds and what an agentic business might really, really look like. But in situations like today, when we remove constraints of what we can do with technology thanks to ai, that's where the possibilities are created.
You know, using tools like copilot, using applications like Dynamics 365, and we'll, we'll see what end users as well as technologists bring to bear in their ideas and how they reshape businesses day and tomorrow. And how about you? But I'm super excited about the future that we're creating together thanks to working with technology companies and with end users like yourself.
Hi everyone. Uh, thanks so much for joining. I'm Brendan Burke, research director at the Futurum Group, covering semiconductor supply chain and emerging tech.
And I wanna set the stage for how the industry is going to evolve, uh, over the coming year. Uh, to do that, I'm gonna zoom out first on the state of the industry and then a dive deep into how I think the latest innovations are are gonna play out in the coming year. And we're meeting at a remarkable moment, uh, for the semiconductor industry.
After two consecutive breakout years, the industry is on pace to approach a trillion dollars in annual revenue in 2026. Uh, this growth is historic, but it hasn't come easily. Uh, each of the past few years has had its own existential questions for the industry, and the stakes have always felt very high, uh, because of the increasing importance of the semiconductor industry in the overall economy.
But if we look back over each of the past years, there's been, uh, one bottleneck or one issue, uh, that that's faced the industry and threatened to slow down progress. Uh, in the past few years, we've worried about extreme supply constraints, geopolitical risk, uh, the availability of raw materials, uh, the pace of fab buildout, and even sometimes if there was enough demand, uh, for the most evolutionary technology in history. Uh, but more recently, concerns have shifted, uh, towards, uh, delays with latest generation of servers, uh, the availability of power, uh, for these servers and the memory constraints facing AI systems.
And yet, over the past three years, the industry has pushed through these bottlenecks, and it's done so, uh, by driving efficiency in the supply chain, by making aggressive design choices that get ahead of user requirements, uh, and developing new techniques around packaging, uh, system level design and long-term capacity planning, uh, that ultimately, uh, prepare for bottlenecks, uh, before, uh, they affect the industry. And, uh, what's becoming clear as we head into 2026 is that, uh, semiconductors are no longer reacting to cyclical constraints. Uh, they are anticipating them.
Uh, leaders have planned ahead, uh, driven efficiencies across fabrication, memory, memory, uh, packaging and power and position themselves to benefit, uh, even as downstream pressure increases for, uh, smaller players. So today I want to take a step back and ask a simple question, what problems are still going to matter by the end of 2026 and which ones won't? Uh, right now, the dominant concerns are memory limitations and energy availability.
Uh, AI systems are extraordinarily memory hungry, and global power infrastructure is struggling to keep pace with demand from data centers. At the same time, uh, recent model launches and performance breakthroughs are only accelerating the ambition of, uh, frontier Labs with longer agentic sessions, uh, more complex reasoning and larger training runs, uh, aimed at the next, uh, frontier of, uh, scientific AI models. And, uh, so that, uh, puts the industry head to head with a memory wall.
Whether the memory capacity on chips can scale fast enough to keep up with the size of training runs, um, uh, yet, uh, that, that, you know, brings us to the real shift in how AI is going to be used, uh, and how the software community is going to adapt models to the latest hardware that determine, uh, whether this memory wall will be an impediment for the industry or just a bump in the road. And, you know, overall, uh, you know, what we're going to see this year in AI that affects, uh, the semiconductor roadmap is a shift from subscription chatbots to agentic systems that perform real world work, uh, systems that generate value comparable, uh, to human employees. Uh, that transition fundamentally changes what we need from semiconductors, not just more compute, uh, but more efficient, uh, compute and, and more sustainability, uh, in terms of the, uh, business model and the efficiency, uh, of our AI systems.
And so, my core prediction, uh, for the coming year is that, uh, 2026 will be defined by 10 to 20 times improvements in tokens per dollar per watt. Uh, we're going to see measurable gains in system efficiency, uh, new architectures, uh, in AI that are optimized for long running ag agentic workloads, and a smarter use of memory and storage to scale token generation within tight power envelopes. So, uh, by the end of the year, I believe that we're not going to be talking about a memory shortage anymore, and energy constraints will start to ease as we add in new power generation capacity and utilize existing data centers better.
And that might sound misaligned, uh, with the, the, uh, skyrocketing prices we're seeing from memory right now. So today I want to talk through, you know, why I think, uh, that, uh, that the industry is overcoming the memory wall, uh, what key metrics are already telling us, and, uh, what semiconductors have to deliver, uh, to support long running, uh, stateful AI agents, uh, without, uh, hitting out of memory errors. So there's a, a few predictions that support, you know, the overall, uh, growth and efficiency that I think we're going to see in the real world this year.
Uh, first, uh, that existing, uh, server designs and software improvements will allow for significant real world gains, uh, within cloud environments for higher utilization, uh, in both, uh, ag agentic, uh, reasoning, as well as large scale training that uses reinforcement learning environments. Uh, second, that the real constraint we're gonna be talking about, uh, throughout this year is going to be scale up, uh, uh, within the rack and between racks and, uh, how much throughput, uh, we can deliver, uh, at cluster scale. Uh, I think a major constraint we're going to hear about this year is the performance, reliability, and supply chain around photonic computing, specifically to improve the networking that we get both within the rack and between racks.
Uh, third, uh, the supply chain diversification, uh, is going to support more aggressive semiconductor roadmaps to specialize silicon, uh, and then take advantage of the latest foundry capabilities. Uh, and fourth, that, uh, CPUs and dpu, um, uh, for reinforcement learning and AgTech tasks, uh, tasks will play a key role in improving tokens per watt, along with the GPU roadmap. So, um, first I wanna start with the available systems that we have on the market today and what those are going to do, uh, for system efficiency.
Uh, you know, as I mentioned, I think that, uh, you know, the, uh, platforms that are coming on the market this year are ready to give us, uh, 10 to 20 times real world gains in tokens, per dollars per watt, uh, within, uh, cloud environments, uh, that, so that will actually be experienced by customers, not just, uh, in a lab, but, uh, actually in, uh, the, uh, amount of tokens that customers are able to generate, uh, and the, uh, amount of, uh, concurrency that they're able to deliver to, uh, wide groups of, of users for ag agentic experiences. So, uh, just recently, uh, Jensen, uh, from Nvidia and Lisa from a MD have driven home that, you know, the next generation of servers are going to have, uh, two to three times more memory capacity and memory bandwidth. And, you know, it goes to the question of, you know, uh, whether this is going to be enough, uh, whether we have the, the memory we need to actually deliver on, uh, uh, you know, the, the promise of ai.
And, you know, fundamentally what, uh, these, uh, leaders have recognized is that agents, uh, need state in order to perform on long running tasks. Uh, when an agent spends three hours researching, uh, a, uh, market report or updating a code base, it has to, uh, remember a plan, the results of, uh, prior tool calls and the original intent of the user. Uh, this short term memory is stored in the KV cache, uh, and the GPU's high bandwidth memory.
And as reasoning depth grows, that cache explodes, uh, fetching that history over and over requires massive memory bandwidth. And the, the problem, uh, to date with, uh, creating ag agentic products, uh, and using GPUs, uh, for, uh, um, AI agents is that, uh, GPUs spend more than half their time, uh, doing nothing, uh, when, uh, running an agentic workload. And, uh, we see that, uh, in, you know, uh, independent academic testing that idle periods can be up to, uh, 50, uh, 5% of total execution time for an agentic task.
And so, uh, we have, uh, cloud users as well as application builders that are paying to access, uh, a, uh, sports car, and then, you know, leaving it in our parking lot, uh, for half the time, uh, that, uh, you know, they're, uh, you know, uh, in a race. And to build sustainable AI agents, uh, we've got to eliminate that idle waste and use hardware, software and networking together, uh, to drive additional efficiency. And, uh, I think key drivers of, uh, the improvement in GPU utilization for AG agentic workloads is going to come from some of the latest innovations we've seen in both NVIDIA's Blackwell platform and, uh, the upcoming, uh, Ruben launch, uh, scheduled for the end of this coming year, because I think, uh, a major constraint in, uh, AgTech deployment, uh, over the past year in 2025 is that, you know, many hyperscalers and builders were running inference workloads on infrastructure that was originally designed, uh, for training, uh, and not necessarily for ENT inference.
Um, you know, we still have, uh, large existing, uh, fleets, uh, uh, you know, amp pure GPUs from three generations ago, uh, hopper deployments and, uh, overall, uh, you know, systems that were optimized for flop output, uh, and not necessarily for high throughput inference. Uh, and, uh, you know, as a result, we, uh, you know, tried to change over the software to make better utilization of these clusters, but didn't necessarily get the performance out of them, uh, that, uh, we wanted. Uh, because inference isn't necessarily about, uh, pure flop output, it's about, uh, hitting the, uh, KV cache and, uh, accessing the tokens needed to respond to a user request.
It's about token by token memory access, uh, and then, you know, also improving on, uh, latency sensitive memory traffic. And ampu and hopper systems have been widely used, uh, but, uh, you know, we've started using Blackwell for large scale training runs, and when they, uh, change over to running an inference, we're going to get a lot of benefits based on, uh, the new generation of architecture. Uh, hoppers, you know, have been great for, uh, transformer engine optimizations and, uh, you know, training mixture of experts models, uh, but, uh, have been limited in the, the high bandwidth memory capacity for each GPU.
Uh, we've seen that, you know, memory bandwidth, uh, was insufficient for a large context, and that often, uh, these clusters SAP title. And so it's not a problem with the hardware, it's just that there is a mismatch between, uh, the actual productization of AI and the compute that was put in place for it. But, uh, just over the, the past half of the year, uh, Blackwell Ultra has started rolling out at scale.
And, uh, what's clear from the specifications and the way the customers, you know, want to use these systems is that, uh, the HBM available in these systems is going to be the key to efficient inference, uh, because of, uh, the, uh, design of Blackwell. Uh, we're going to see once these systems are put in place and applications are, uh, integrated with the systems that token latency is gonna drop, uh, and we're going to get better tokens per dollars, uh, per watt. Uh, and, you know, I think, uh, you know, this is, uh, going to be proof that, uh, you know, inference is a memory problem, and that systems with faster access, uh, to, uh, large pools of memory can perform much better on, uh, high speed reasoning tasks than their predecessors.
Uh, and, you know, I think there's been real limits to, you know, using the latest hardware for agentic applications. Uh, over the past year, uh, uh, you know, hyperscalers have been limited by the hardware they've already used. Uh, there have been, uh, you know, significant, uh, crashes, uh, from, you know, the, you know, prior generations of hardware, uh, when attempting to run, uh, you know, trillion parameter, uh, models with reasoning capabilities.
Uh, and, uh, you know, it's clear that new hardware is needed, uh, to deliver things like multi-hour, uh, agentic sessions, uh, concurrently, uh, to large groups of users. Uh, and also to create a business model around those experiences, uh, based on the reliability and performance, uh, of, uh, you know, those things like coding agents, uh, or, you know, long running, uh, scientific, uh, research processes. And, uh, you know, because of the deployment of, you know, Blackwell this coming year, I think we're going to start seeing those, uh, improvements start to play out in the real world, not just in a lab or in testing, uh, but actually in how customers experience, uh, the, the number of tokens available to them, uh, and ultimately that, uh, that the, the cost that they pay to, uh, you carry out frontier, uh, you know, uh, tasks, whether that's, uh, coding large code bases or, uh, developing new scientific advancements.
And, you know, this has been, uh, clearly, uh, substantiated by, you know, independent testing and data around the inference capabilities of Blackwell, uh, versus the prior hopper generation. Then recently, the independent testing firm Signal 65 analyzed the GB 200, uh, NVIDIA server, and showed that at high interactivity, uh, we're not just seeing incremental gains. Uh, we're seeing up to 24 times performance per GPU compared to the hopper generation of servers.
Uh, but, you know, at typical interactivity thresholds, uh, I, it's, uh, there's an average of around 20 times a generational leap in terms of the tokens produced, uh, by, uh, the, uh, Blackwell server compared to its prior generation. And, you know, it's this improvement in performance, uh, uh, at, uh, you know, high interactivity, uh, that is going to make these systems performant for the leaving clouds, uh, and able to generate more tokens and, uh, accordingly more revenue, uh, in the coming year. And when the cost per token plummets, uh, the, uh, and, you know, the tokens per watt metric actually escalates, uh, 20 times more.
It's going to present an economic unlock for designing new agentic experiences. Uh, you know, with that additional token production, uh, for a wide array of users, uh, the recent, uh, adoption of, you know, coding agents especially, is going to allow these systems to run for hours, uh, at a time, uh, for the same cost as a, uh, short chat bot session today. And so, models that, you know, used to use kilowatts, uh, for a given, uh, inference job, we'll just, uh, reduce that to, uh, you know, a, uh, kind of a double digit number of watts.
And, uh, you know, the, the AI agents then will, uh, have access to more memory, uh, and will be able to do hours of engineering simulations, for example, without necessarily blowing the power budget of existing data centers. Uh, and so it's going to be this 20 x improvement in performance that really drives that shift to AgTech applications. Uh, so, uh, with that, I wanna move on to my second prediction, uh, which is the, uh, role of scale up in, uh, improving tokens per watt.
And, you know, how photonics is going to beca become a critical differentiator for networking. Uh, so, you know, over the past decade, the industry is focused on improving scale out networking. Uh, and we got very good at connecting, uh, racks, uh, together with, uh, you know, improving ethernet standards, uh, and, uh, minimizing congestion within data centers.
Uh, and now I think, uh, there's going to be a new paradigm, uh, that, uh, we face as the industry runs into the limits of, uh, copper, specifically within connecting, uh, data center racks. It's going to, uh, we're going to talk a lot about how, uh, much compute and memory bandwidth you can concentrate inside a single rack, uh, and how efficiently those components talk to each other. And the frontier training is, uh, now dependent on, you know, uh, tens of terabytes per second of, uh, me memory bandwidth, uh, that's sustained continuously.
And then mixture of experts, models need all to all communication between chips, uh, and dynamic routing across experts. Uh, reinforcement learning is a growing workload that also requires, you know, synchronization, uh, across, you know, different environments that might be hosted on, uh, uh, different servers. And so, um, as a result, we're going to need improved latency, uh, to let, uh, these workloads, uh, actually, uh, turn into, uh, improved model training experiences and be able to train trillion or 10 trillion parameter models, uh, at and a fraction of the time.
And, uh, you know, the small delays can compound and crash training runs along these steps, you know, adding, uh, you know, a few microseconds for each hop can collapse throughput, uh, and then also stall out memory systems. Uh, that leads to the need for, uh, a more, uh, improved checkpointing methods for training runs. Uh, and, you know, leads to a scenario where it actually becomes difficult to achieve frontier training results on, uh, the, the generations of hardware that we have now.
And so, for that reason, I think, uh, we're going to see, uh, uh, you know, major innovation and scale up in the coming year. Uh, we're running, uh, m two limits of, you know, how much electricity we can push through existing cables, uh, beyond, you know, 200 gigabits per second per lane signal integrity is extremely difficult to maintain, even with the latest innovations in SerDes and, uh, active electrical cables, uh, with, you know, the, uh, you know, uh, improvement of, uh, you know, uh, certis being able to handle, you know, over 200 gigabits per second per lane. Uh, you know, Jensen said that it's now like running the entire world's internet, uh, through a switch and signal attenuation, uh, can become prohibitive.
Uh, the, uh, you know, effective reach of passive copper cables can shrink from two meters, uh, to just inches, uh, at, you know, terabit speeds. And then engineers have to make use of a digital signal processors and, uh, re timers to, uh, recover that signal. And, uh, you know, by 2026, I think we're going to hear that relying on copper, uh, for a rack scale, AI fabrics is going to actually increase the power budget, uh, to a point that, uh, uh, for intra connect that rivals the compute power budget, uh, which, you know, could, uh, ultimately affect the power utilization efficiency for data centers.
So in 2026, I think we're going to see the transition of co packaged optics from experimental prototypes, uh, to commercial reference designs as a critical differentiator in new AI systems and initial volume production. This shift, it might happen faster, uh, than expected in the industry as we, uh, see the arrival of integrated switch platforms with CPO, uh, the resolution of critical raw material bottlenecks and the introduction of software defined testing to improve reliability. The conditions I mentioned about the constraints, uh, within the rack create a strategic imperative, uh, for CPO, uh, that's, uh, driven by the need for additional power efficiency.
Uh, I think overall, we're not getting the improvement in power efficiency from traditional optical receivers and plugable optics, uh, meaning that an optical transceiver can consume up to 50% of the power of a network switch system. And, uh, as we get to the gigawatt scale, uh, the, uh, you know, pluggable optics that are being used, uh, can consume, uh, 15 to 20 pico joules per bit. Uh, whereas CPO can actually reduce that by two to three times.
And that's a strategic enabler, uh, for data center operators at scale. Uh, once you factor in the fact that CPO also reduces electrical signal loss compared to copper, uh, you know, by up to five times, uh, we're looking at CPO being a significant driver of continued gains in tokens per watt. So, uh, I think we're going to hear a lot more about optics in 2026, especially in reference designs for future production, if not at scale, uh, within, uh, systems, uh, uh, you know, that are on the market just given supply chain constraints, as well as the need to ensure reliability, uh, at, you know, at the, uh, switch level.
Uh, and there, you know, are major constraints as well as with the availability of raw materials like Indium phosphide, uh, that is the essential material for lasers that power silicon photonics. And today, the supply chain is, uh, you know, uh, slow to scale, uh, but, uh, the suppliers do see those constraints easing in the coming year. Uh, so by the time we get announcements that make, uh, co packaged optics a key part of, uh, switch architecture, I think we're going to, you know, have additional supply come online that allows us to tradition, uh, transition to production deployments and actually get ahead of, uh, forecasts for, uh, you know, mass deployment of this technology by 2027.
And, uh, next I want to, uh, get to, uh, you know, the role of foundries in this ecosystem. Uh, you know, I think that, uh, because of the need for additional efficiencies that chip design companies are going to cast their net more widely, uh, for foundries that are able to develop, uh, custom asics that support specific workloads, uh, and really encourage silicon diversity, uh, through supplier diversity. Uh, and, uh, 2026, uh, a major constraint for AI systems is simply whether the supply chain can execute at scale to get us all the compute we need, uh, to deliver, uh, on, uh, the demand for, uh, AI compute.
And so those bottlenecks go across the supply chain, uh, from foundry capacity to advanced packaging and memory availability. And as a result, I think we're going to hear a lot of announcements in the coming year, uh, for higher utilization of competing foundries. Uh, and, uh, then also packages that are able to, uh, diversify the product roadmap of semiconductor leaders.
Uh, and, uh, you know, in, uh, just the past year, we've seen designs that are qualified across multiple foundries, uh, to improve volumes, uh, for chip outputs. Uh, the usage of, uh, trailing edge nodes, uh, for individual components, um, you know, freeze up leading edge capacity for, uh, compute critical tiles. And, uh, and you know, this, it shows that, you know, heterogeneous integration, uh, is, uh, can work to combine, uh, inputs from multiple different, uh, IDMs, uh, um, but, um, you know, it requires creativity on the part of chip design companies, uh, to diversify their supply chain.
And we're, we're seeing this in packaging as well. Uh, you know, for years, a small number of advanced packaging, uh, flows, you know, were the bottleneck. Uh, and now as a result, you know, vendors are qualifying multiple oats to spread assembly across geographies and accept different, uh, process characteristics in exchange for, you know, throughput and schedule certainty.
And so I think in 2026, utilization of alternative, uh, packages will increase, uh, because, uh, shipping matters. And there's also, uh, improved designs that are coming from a different corners, uh, of, uh, uh, the, the global oat market. And, you know, memory is a gating concern, uh, for training and inference.
Uh, and, uh, for that reason, uh, there's, you know, tighter planning between logic vendors and memory suppliers, as well as multi-sourcing across different HPM sources, uh, and different, uh, memory foundries. Uh, and, uh, you know, I think we're going to see, uh, you know, continued mixed memory configurations of bringing together, uh, uh, both HDDs SSDs as well as HBM for multiple suppliers to develop systems that can handle the, the memory throughput and drive continued gains in tokens per what. And then when we get to code packaged optics, there's the actually, you know, foundries that, you know, specialize in, uh, you know that as well.
And, uh, that could be an area where chip design companies compare among, uh, different suppliers in order to get access to the latest process technology. And, you know, the last, uh, prediction that I think supports the, uh, you know, need for improved tokens per wat in the coming year is the CPU and DPU innovation for reinforcement learning and agentic tasks. Uh, so, uh, right now, uh, I think, you know, the overall, uh, direction of the industry is moving away from the GPU to totally, uh, drive the performance of ai.
Uh, you know, the reliance of new AI systems on system, uh, on DPU and modern CPUs, uh, you know, really makes, uh, some of those more legacy, uh, systems, the air traffic controllers of, uh, servers and the data center. So dpu that include, you know, smart, uh, nicks, uh, you know, are really important for, uh, connecting some of the storage resources I mentioned, uh, to the GPUs, uh, that need data and ultimately, you know, lie idle, uh, if they don't get fast access to memory. Uh, and, uh, you know, I think we've seen significant announcements of advanced DPU innovation, uh, just, uh, this year at CVS with the Bluefield four, uh, DPU, uh, coming onto the market.
Uh, and, and then, uh, pasando dpu as well, uh, really, uh, being a critical driver of AgTech, uh, reasoning, uh, and, uh, a, uh, reason why advanced systems can be used both for training as well as high interactivity inference. Uh, and, uh, then I think we should also talk about CPUs, uh, because they're not just orchestration glue anymore, they're differentiators, uh, for high value AI systems, and specifically in RL and age agentic workloads, CPUs can handle, uh, a lot of, uh, the, uh, load, including environment simulation, uh, decision logic and scheduling, uh, that are really important for carrying out verifier workloads that ultimately, you know, determine the performance of AI models. And, you know, in practice, uh, a lot of these pipelines, you know, hit CPU bottlenecks, uh, uh, to date.
And, you know, uh, when it comes to things like simulations and, and control loops, uh, you know, CPUs can actually starve, uh, you know, the expensive GPUs themselves. Uh, and I think we're seeing an increasing emphasis on improvements in CPU performance, uh, making the CPUs themselves high bandwidth, coherent participants with significant DDR memory and high memory bandwidth. Uh, and, uh, that's, I think, a significant reason why, uh, CPU innovation has become a core part of new AI systems, uh, right along with new GPUs.
Uh, and, uh, you know, that's an area where, uh, we're going to see additional XPU innovation in the coming year. Because while, you know, the GPU roadmap is dominated by two vendors, specifically, CPU innovation can come from, from much broader, uh, parts of the ecosystem, uh, and ultimately improve the performance for more open rack architectures, uh, as well as open chip architectures. Uh, and, uh, that's where we're seeing specialized designs that optimize CPUs, uh, for AI workloads and get us away from, you know, the monolithic GPU cluster, uh, that the industry, you know, built around in really the pre-training phase.
Uh, and, you know, I think roadmaps that we're seeing right now point us in that direction of how we can use additional compute resources to get better outputs for specific stages of the AI lifecycle. Uh, you know, whether that's, uh, NVIDIA's announcement of the Ruben C-P-X-G-P-U, uh, that is really optimized for the context phase, uh, and, you know, minimizes the amount of, you know, uh, memory relative, uh, to the overall Vera Ruben system. You know, it's clear that, uh, you know, the, that, um, you know, specializing hardware for specific stages of, uh, the lifecycle, uh, uh, from training to reasoning, uh, is going to be critical, uh, for giving data centers the portfolio of compute needs that, uh, will, uh, power, uh, the future of their agentic experiences.
Uh, and so, you know, this is all happening in the context of software optimizations, uh, that, uh, you know, I think, uh, deserve equal, uh, participation in this discussion because it's really when you combine the latest hardware, uh, with, uh, you know, uh, improved intention mechanisms as well as quantization of models, uh, along with AI specific networking, uh, that you get the upgrade in tokens per watt, uh, that, uh, ultimately will drive, uh, improved performance in the industry and revenue generation can't get to all the software innovations that are doing that to date. But I think it's a core part of the mixture of expert models that ultimately, you know, drive the 10 to 20 x gains in tokens per dollars per watt that we're seeing so far in the market. So, to wrap up, I don't think a trillion dollar industry can just be supported by faster chips and improved specs.
It's going to be about improved utilization and showing customers that they can get better performance and better revenue generation out of the systems that are being shipped. And that's going to require specialization and efficiency across the supply chain. Uh, I think that's, uh, an area that's going to drive a lot of innovation.
It's gonna drive a lot of change. I'm excited, uh, to track the industry this coming year and keep up the discussion, uh, with all of you. Uh, but I think it's, it's clear that we're going to be looking at, uh, a much different world, uh, at this time next year, uh, one that, uh, implemented some of these most advanced systems and then saw the unexpected results that come from them.
Uh, so, uh, until we see the end results of it, I'll look forward to keeping up with you. And thank you for joining.