The Future of Agentic AI: Opportunities, Risks, and Society | Utilizing AI Episode 22
AI agents are already shaping how we work, think, and interact with digital systems. This week on Utilizing AI, Stephen Foskett sits down with Dave Graham of MLCommons and Frederic Van Haren, founder and CTO of High Fens, to explore the rapidly evolving world of agentic AI. They discuss the current state of AI agents, including emerging tools like OpenClaw, and examines how these systems are moving from experimental tech to everyday utility. They also take a look at real-world applications, enterprise adoption, and the growing role of AI in automating tasks and extending human capability. But with opportunity comes risk. The discussion also dives into safety concerns, reliability challenges, and the broader societal implications of giving AI systems more autonomy. This and more on Utilizing AI.
Transcript
Awareness of agentic AI has reached the general public, with millions of people, even outside the technology industry, deploying OpenClaw and other agentic systems. This episode of "Utilizing AI" features Dave Graham and Frederic Van Haren discussing the real status of agentic AI. Welcome to "Utilizing AI," the podcast focused on practical applications of artificial intelligence from the Futurum Group.
Every Wednesday, we explore news and use cases of the ways in which AI is transforming enterprise IT and the industries it serves. I'm your host, Stephen Foskett, President of the Tech Field Day business unit here at the Futurum Group. Before we dive into the conversation today, let's meet who's on the panel.
My name is Dave Graham. I'm the Director of Marketing at ML Commons Association. So we're a member-led industry technology standards organization that's designed around benchmarking and characterization, both in the AI risk and reliability space, as well as the infrastructure, AI infrastructure space.
Yeah. Thanks for having me. I'm Frederic Van Haren.
I'm the founder and CTO of Hyfence, and we provide HPC and AI consulting and services. And of course, I am Stephen Foskett, your regular host and good old Frederic and Dave here. We've known each other for many years.
We've been talking about this stuff for many years. As I said in the opening, today we're going to talk a little bit more about sort of where the industry is at, where people are at with rolling your own agents. And I want to kick this off, well, with an anecdote.
So I'm driving home last night, or the other day, and listening to NPR's "Marketplace," and there's a guy on there, the guest on there is talking about agentic AI. And his pitch was essentially, this isn't quite ready for prime time. We don't know exactly where it's going to go.
But you, dear listener, not techie, just NPR listener, should be deploying agentic AI now so you don't get left behind. And he gave a metaphor for that about how basically if you're not experimenting with AI agents in your life and in your business, then you're going to be basically behind on the football field. And it got me thinking because people like us are actively experimenting with this stuff.
" And frankly, OpenClaw has reached, especially in China, but increasingly here in the US as well, has reached public consciousness to an extent that people are out there, like regular people, normal people, are out there deploying agentic AI for themselves. And this is an interesting situation that we find ourselves in because, frankly, I have been aggressively experimenting with this stuff as well, and I'm not all that satisfied with the performance and the user experience of it. Let me throw this to you first, Frederic.
What's your reaction to somebody telling normal NPR listeners on their drive home that they should deploy OpenClaw? Well, I think the first problem is people don't necessarily have an understanding enough of AI. But let that alone for a little bit.
People want personal assistance, right? You have thousands of emails. You want to go through your email and figure out what did you miss in the last week?
Do you need to go through all these emails? So people are looking for personal assistance and an entry to an easy AI solution. And I think today what we're seeing with OpenClaw and others, it's really the beginning.
It's really easy to install it. They promote no code. The no code also comes with no understanding of what it's actually doing under the covers, which a whole different conversation.
But I think what people are trying to do is to get a taste of AI, and they're trying to do that with some kind of an assistant that looks at their own data. To me, it's the beginning. Should we tell people to try it?
I would say let's try to understand it first, what it can do for you, and also what it shouldn't be doing for you. If you give OpenClaw too much information, it actually will use that information, not necessarily for you, but in some cases against you. So it's the beginning.
Should people experiment? I guess so to better understand what it is. But should they bet the farm on it?
I don't know. Yeah. I always teeter on the edge of Luddism when it comes to some of this stuff.
Which is ironic, isn't it? Because- But- ... here we are.
But I feel- Yeah ... the same way, Dave. Yeah.
So, it's a principle thing. Did I go out and buy a Mac Mini? Yes, I absolutely have a Mac Mini.
It's sitting there right behind my head running, I think I'm running some auto research benchmarks right now, right? Because it's the fascination that Frederic and I, and you yourself have, Steven, in these things is what keeps us going. But we're like, I wouldn't say we're the 1%, but we're the 1%, right?
Well, the techie 1%. We're the techie 1%, right? It's a running joke that you go on Reddit, and you see everybody complaining about X and Y silicon vendor not doing overclocking chips the right way.
But they're the ones that tend to yell the loudest, if you will. So I think we're in that type of situation at this point, right, where the pantomimes at conferences about OpenClau or NemoClau or whatever claw, and IronClaw, I mean, everything, it's going to be an evolution. You're right, Frederick.
I acknowledge that everybody wants this concept of a virtual agent or a helper or assistant. I think this is-- I use one, Little Bird is mine. I'm on a Mac, and Little Bird's my go-to agent.
It rolls me up a journal every day of the activities that I've done and watches what I do and kind of interacts, and it's great. Because it keeps a level of logic into my day. It just starts to look at patterns.
But again, this is a patterning engine underneath it all anyway. It's looking for patterns. This is something that's been designed to do statistically and kind of algorithmically.
The next step beyond that is this, Steven, to your thing about NPR. Should everybody experiment? Yes, sort of.
You're always going to run into it. I mean, I can use Claude on Chipotle's chatbot on their website, right? If I really want to do this.
We're interacting with agents that we don't actually understand are doing every day. So I think there is this again, that razor's edge of should everybody experience it? You know, maybe.
Are we experiencing it whether we know it or not? Yeah, absolutely. And we need to kind of figure out what that looks like.
There needs to be an education cycle around it, all that to be said. Yeah, and that's an interesting point. So I mean, first, a couple things to unpack there.
First off, I had another interesting conversation this week about AI tools, and somebody asked me how we're using AI tools internally at The Futurum Group, and I said, "It's funny because on the one hand, I think if I ask most people, they would say they're not really using AI as much as they would like to. " And so it's like, but they're not thinking of them as AI tools. They're thinking of them as tools.
"Oh, I need a transcript of this video. Boom. " "Oh, I need a summary.
I need some keywords. " Boom. Throw this to ChatGPT.
Use Apple's built-in summary tool. Use Parakeet to trans-- And suddenly you realize that these tools are really pervasive in our daily lives. And so from that perspective, I think the use of AI tools really is already everywhere.
But there's a different thing when we talk about using an assistant. Now, I will say, too, Dave, I also have been experimenting with Little Bird because a friend of mine suggested that it was a really helpful advanced platform. I've not gotten to the point yet where I can render judgment on it, but I'm excited by the idea that we will have professional usable tools out there.
Because I will tell you, I didn't buy a Mac Mini because you don't have to buy a Mac Mini because it'll run on fricking anything. I'm using actually a stack of old MacBook Pro as my MacBooks Pro as my lab for OpenClau, and I have a half a dozen of them running various incantations and incarnations of this thing because I wanted to see what it's like. I want to see what the default is like.
I wanted to see a better way to deploy it. I actually, over this weekend, deployed it in Docker Compose in a very, very locked-down situation. Previously, I had just deployed it with a bunch of junk data and a fake account because I am terrified of what it can do in your name if you deploy it.
And frankly, I've been very disappointed with the default suggestions because on the one hand, everybody says, "This thing is dangerous. " But yet the default normal, whatever internet, go ahead and try this thing out instructions, give it access to everything in the world. I mean, it's the generic install is exactly the wrong way.
" It's going to give you, I'm sorry to say, it's going to give you a terrible, insecure recommendation, and people are doing that. So, have you all tried it in various configurations? I mean, what are you doing?
So, I'll go, Frederick, yeah. Yeah. Just give you a, I'll bounce it over to you in just a second.
So my initial, I was early in the OpenClau. " And I think this was before OpenClau. So whatever it was, the name and nomenclature and the repository before that.
Before they kind of put the pretty ribbon on it to try to make it useful, quote, unquote. It struck me that a lot of what it was trying to do were things that were low effort. Low effort human things, right?
It's go out there and do X, Y, and Z, right? Just real basic kind of principles. env file when you upload and things like that.
json has all your tokens and keys in it, and people upload that all the time. Well, and I've kind of come to this realization that OpenClau is the next evolution of the Amazon billing problem that Corey used to talk about, where you all of a sudden wake up one day andEverything in your bill is insurmountable. " So I approach it with a certain amount of caution, just simply because, listen, my life is pretty well boxed in and I want to keep it that way along the line.
So my initial input was, yeah, this is good. It's a science experiment. It probably has some really reasonable outputs and pretty useful for some things.
But those are things that right now I want to do. I want to continue to engage in these type of things, and I'm really looking for offloading. Again, the concept of take meeting notes for me.
I don't need OpenClaw in order to do that, and Gemini doesn't- No, all you need to do is be on any meeting in the world now, and you're going to get 12 different meeting notes My CRM has a marketing bot, Little Bird can record it, I can record it. I mean, whatever. So, that was my inception.
Frederich, what about yourself? Yeah, I think when we look at AI being a tool, the interesting piece about AI-driven tools compared to, let's say, traditional tools is that the AI tools typically that are being promoted or pushed for are not really ready for prime time. It all started way back with the large language models, right?
It's almost like it escaped out of a lab, and then the drive for large language models was really from a competitive perspective, right? It's the other large language model vendors that started pushing for it. And I feel like that's what's happening today, too, is that a lot of these AI tools are experiments, and we're kind of the guinea pigs to see what works and what doesn't work.
And so that's the consumer side. And then when you look at the vendor side, the large language vendors, I'm trying to remember what Jensen, the CEO of NVIDIA, said on stage at GTC a few weeks ago. Didn't he say something like, if the engineer is not spending 200K on tokens, then he's not doing it right?
Yeah. And so it's almost like the push for let's do it and generate as much or as many tokens as possible. But the reality is that us, the guinea pigs using all those tools, it gives a lot of data to the vendors to work on the next level of AI tools.
And so I don't think it's wrong for us to try these tools, but we have to realize- Yeah ... that when OpenClaw came out, it was leaking data left and right. And it was not considered a concern.
It was considered as, well, it's an experiment, it's an AI tool, use it at your own risk. But they make it so easy to use today that I don't think people understand what those tools potentially could do. Yeah.
And I think that's kind of the risk, right? So if we really look at AI tools and we transparently want to use them, we need a better way to make those tools ready for prime time without having to suffer the consequences afterwards. Yeah, 100% agree.
And that's actually my feeling. Again, I have now installed and configured and played with versions of OpenClaw for a few months. I have done half a dozen different start from scratch implementations with different scenarios in different ways, and my verdict on it, at least as far as it stands today, is this is in no way ready for prime time and people shouldn't be using this.
But the amazing thing that it does, and this is the wow factor of OpenClaw, is that by allowing it to have persistent memory and a schedule and iterating over itself, it is as transformative as the sort of introspective GPT models were that came out at the end of last year, in that it takes things to an entirely new level simply because it's able to autonomously act and do things. Essentially, what it's doing is it's showing us where this technology is going to go next. And that to me is the exciting thing.
So when we see NVIDIA on stage saying that they're going to create their own Claw-like agentic, and we see OpenAI acquiring the Open Claw team, we see Apple and Google and everybody else saying, "You know what? Everybody is looking at that and saying, 'That's the guide post. That's where we want to go.
'" And I think that's healthy. I think that's where the industry should be. It should be looking at ways of building systems that deliver that wow, like OpenClaw, but aren't just incredibly bizarre, hacky science projects that barely work.
And Dave, I know that you guys are working with a lot of these companies. There's a lot of this stuff coming, right? Yeah.
To even backtrack a little bit to what you just said, we look at the iPhone, for example, as the inception point for basically this turn of the wheel, this digitalization, this an actual social phenomena that turned into a technical revolution, right? Yeah. It was the first time that you had really had the ability to capture a world around you in more realistic ways, right?
Embedded camera. You're interacting with a human interface to a machine, right? And so actually in academia, we look at 2007 in the introduction, 2005, forget, 2007, and the introduction of the iPhone as one of a social revolution.
Yeah And I view OpenClaw and tools like it as that kind of second step on that revolution, right? Because it's starting to embed into our social fabric the idea of agents or non-human digital actors, right, that have or don't have, and not to make this into an entirely academic conversation, agency or the ability to go out there and fulfill a role of its own devising or fulfill a role that another or deus ex machina, right, to a certain extent, that somebody else is projecting upon it in order to kind of go and do, right? So a lot of that's kind of a prelude to why MLCommons and reason why we're interested in characterization and kind of the embeddedness that we look at here is that we're interested in kind of the entire groundswell, what you build from the foundation on up, as well as top-down, right?
So the infrastructure and silicon and all the technological marvels that make yes, make it work, right, that you run on top of is one part of it. But the other part of it is that stickiness at the top, right? So Frederick, you talked about it, and Stephen, you talked about it as well.
It's that human interface to these devices and the influence that those digital personas or those agents have on the world outside of them, right? There's a risk involved, right? You give it the wrong...
env file, right? env, it's right there in the JSON. Yeah, yeah.
So I don't know how you're running your thing, but this is something I had to learn as well, right? When you go back and do a CodeRabbit audit of a GitHub repository. Anyway, but all these things out in the open, right, there's a certain risk and reliability aspect of it.
You can call it safety if you want to, but risk and reliability kind of focuses in on this, where you start to see where this shapes markets, it shapes cultures, it shapes society, and that comes with its implicit dangers. We look at the mental health crisis that we already had that's in our teen, in the pre-formative brains of our youth worldwide, right? And you now exacerbate that by access to tools that build digital cliques and build these digital communities that are highly focused in on presentation, what you look like, and all these kind of invariable, immutable attributes of yourself and all this stuff.
Anyway, all of this stuff kind of starts to become this really crazy foment of everything. So again, it's delving into sociological and anthropological phenomena here a little bit, but that's the reason why MLCommons becomes so interested in these things. We're working on both ends of the stack to try to understand.
It's a characterization. It's a benchmarking as well. But it's all a means to an end to understand better, ultimately, how these tools will impact both infrastructure, the ecology, the power, the fundamentals of things from the silicon side, as well as the social side and cultural side from the risk and reliability perspectives.
Yeah, and I think it's important to kind of realize where the focus is nowadays. At some point, almost all of the focus was on training. How do we get the greatest and the best large language model?
So everybody expected the large language model to do all of the heavy lifting. Today, the shift is completely different. Certainly, there's a bunch of vendors generating large language models, but they're roughly good enough in order to get going.
And so I feel certainly from our perspective and certainly MLCommons too, where originally there was only a focus on training, now it's all on the inference and more consumer, let's call them tools if you wish, but that's where the focus is, is to make it a lot better and improve and innovate. And I think that's what we should expect moving forward. Me personally, I still think that large language models would benefit from more specific markets, more market-driven as opposed to being generalist.
But in the end, for now, it's not bad to focus on tools that can deal with those large language models and continue to innovate. Yeah, that's a real good point, Frederick, and that's been my experience as well. I'll just say that I have been actively hopping back and forth between OpenAI, Anthropic, Gemini, as well as the open-weight models, Qwen and Kimi.
And what I've found is that frankly, these models are really good. They've gotten very, very good today. Yes, some of them are better than others, but at least as far as agentic applications go, at least as far as my experiments have gone, the defining factor, honestly, is the cost per token, not like this model is far and away better.
I agree with you, Frederick, though. I think that there's probably headroom for somebody to develop better tool-calling models, lighter tool-calling models, especially as we deploy more and more systems that have agents and sub-agents and so on. I'm using, for example, I've got a Gemini Flash instance and a sub-agent that's basically only doing security scans and tool calling.
And so that, I can use a very lightweight, quick little model for that little guy, whereas you might want to use a heavy-duty GPT for your real agent, that kind of thing. But a lot of that stuff is still fiddly dial turning, and that's why I'm more excited and more interested in where this goes. There are companies that are going to be developing these things and rolling them out, and people are going to be experiencing these things increasingly I don't think that that's going to be a situation where people are deciding, oh, do I want to use Anthropic or OpenAI?
I think it's going to be a situation of I signed up for such and such application- Yep ... and it just freaking works, and it just does the thing. And then it's more of a question for that service provider to decide how and when and which models to develop and so on.
And then that kind of leads us to the next point, and I want to jump into where Dave was going here with: what does that do to us? What does that do to society if we've got sort of a silicon persona that exists in our lives that is assisting us? I love the idea, especially if it works.
What does that do to everything? And Dave, I know you've got some thoughts here. I have lots of thoughts.
As I expect, you have lots of thoughts about things. Well, I always like to joke. Or actually, even to my executive director, we both come from the humanities side, not the computer science side, right?
And I have two degrees in psychology and counseling, and worked as a time as a social worker here in the state of Massachusetts, right? So very human-based, right? So the impact of tool of things on bureaucracy, even on society and a lot of these things was part and parcel of what kind of I grew up in terms of my knowledge with.
So I'm very, probably pragmatically more interested in looking at the effects of these tools on society, right? How does this increase a person's agency, their ability to function in a society, right? So if we look at, and I have huge issues with the way this was expressed, but if we look at this concept of universal basic compute, which a certain individual from a company that I will not name talked about, versus a universal basic income, right?
I live off money. Universal basic income is a good thing. It provides a common platform.
If I look at universal basic compute, and we look at kind of the analogous relationship between those two, if I'm providing this level of computational credits or tokens or whatever you want to call it these days, I feel like I'm going to Chuck E. Cheese. " Like, "Throw it in the ball pit and see what happens," right?
Is there a positive benefit that comes out of this for somebody that doesn't know how to use these systems? Going back to the original thesis of MPR, talking about how everybody should experience it. If I look at my mother, who, congratulations, has made 80 years old this month, right?
This is a huge milestone. But she's not interested in an agent telling her what to do. She's interested in being able to function and engage with her doctors and the services that are required of the elderly mindset.
So a lot of these things, and as I get older, I care less and less about the fancy stuff and more about just being able to live life, right? So I think this is where there's a possibility here where some of those offsets become taken over by digital actors. I think there is, let's solve the bureaucracy problem.
How do I get to my benefits if I'm on Medicare or some sort of social benefit? How do I solve going to the DMV or the RMV problem, which everybody universally hates, right? So a lot of these things.
How do I understand bills? And having these type of things become a positive, I believe, net benefit to society over time, and this is where that capability of models that are increasingly trained on public data sets that start to understand common systems and processes and problems that we all encounter in our daily lives, whether we want to believe it or not. Solving a tax problem or interacting with these things.
I think this is where there's a lot of net positive effects. There's always going to be room for the 1% or 20% or however you want to call it in the digital space, where we can kind of push the envelope and figure out, hey, now you can do stock trading with your open cloud instance, and you can give it your portfolio and let it run wild. That's an exception to the general rule, and I think this is where we're seeing these kind of step functions go into place, right?
Where it becomes useful. That's the part I'm excited about. I'm less excited about a compression from the top down, where you must use in order to function.
I think that's the wrong way that this thing, it should be a socially augmentative type approach versus a compressive or oppressive approach of you have to use, you must use in order to participate in society. So... Yeah, I think it's interesting.
I think those tools can help us do some tasks. There is a risk is that you built a digital twin that pretty much takes over, right? There's a lot of services today that are relying on interaction with assistants, and so the last thing I want is to have a large language model that makes decisions for me because it believes it has all the data for me.
So I still want to consider AI as a tool. I still want to be in control. I want it to improve my life, but I'm not a fan of creating a digital twin that can speak for me, so to speak, right?
So I think in general, it's an exciting time to live. I had another conversation with a colleague this morning. The market is growing so fast that it's very difficult to try it all out.
There's a clear shift from training to consumer/inference. It also means that the audience got a lot wider. It's not difficult to install these tools.
But I still feel like we're experimentin And to consider those tools as ready for prime time, I think we're not there yet. But that doesn't mean that people shouldn't try it out, experiment, learn from it, and also understand that AI is just more than buying a Copilot license. Yeah.
And for that matter, I will echo my friend from NPR and say anyone listening to this show really should be trying this stuff out. You know what I mean? I would say that anyone who is motivated to listen to an enterprise AI show, like "Utilizing AI," if you haven't gone out and tried installing OpenClaw, you should.
I will say that you shouldn't do it on your primary Mac with all of your data and applications on it. I strongly recommend doing it in a sandbox and doing it slowly. The installation will ask you to connect to everything.
Don't. In fact, start by connecting to nothing, and then add connections later as you feel more comfortable. Worry about security.
Make it read-only. That's a very good idea. Don't let it send emails or post social media or do stock trades on your behalf.
That's just a terrible idea. But basically, it's gotten to the point where you need to try it. That being said, I don't think that this is going to be the end-all, be-all.
I think that, again, OpenClaw is really a surprising transformative use of this technology, and I think that what we're going to find is that there are going to be companies, and I'm going to just say right now, I look forward to what Google does. I look forward to what some of these companies that have a lot of experience building these things do in order to build agentic experiences and agentic companions for people that actually deliver the goods. Like Frederic said too, I'm going to throw this one out there too, I think that there's a looming conflict and question about whether we will have a digital twin that acts as us, or a fake person that supports us and interacts with the public, or an invisible agent that gives us superpowers.
And we'll see where we go with that. I've actually been experimenting with all three scenarios to see what would happen . The first one, obviously, in read-only mode because I don't want to go out there and share some crazy stuff.
But it's been interesting to see to what extent it's able to duplicate me or duplicate a person or just assist me. So we'll see where that goes. We do have to wrap.
We're getting toward the end here. One thing I'm going to do is I'm going to give you each a chance to tell us a little bit about how we can continue this conversation, because as you can tell, I think that we all have a lot to say. We are all going to be at AI Field Day coming up next month.
I can't wait to see you in person and have these conversations in person. And those of you listening, of course, we would love for you to join us for those videos. Check out Tech Field Day website for more information on that.
You can watch them on YouTube and LinkedIn and Techstrong. But of course, if you're listening as well and you're saying, "I'm just as good as Dave and Frederic. I should be part of this thing," we can also have you join us.
So reach out. I'd love to have you join us in the future. So before we go, Dave, where can people continue this conversation?
" I think number eight. Yeah. Eight, next month in- That's right ...
Santa Clara. Otherwise, I am on LinkedIn like a gnat, mostly simply because I'm doing a whole bunch of benchmarking characterization. You can find me on GitHub, Elemental Collision.
I have an auto research repository that I'm working through right now. It's a good acquaintance that I've made through this benchmarking and characterization process, and I have a little sneak peek of a project coming up that I'm hinting at on LinkedIn as well that kind of talks to a lot of what we've talked about today. What happens when you let an agent just become?
And so that's out there, too, and more revealed in time. Not quite ready yet, but yeah. Frederic.
Yeah. So I will also be at Tech Field Day AI. I don't think I missed one, so I think I went to all of them, so hopefully I didn't jinx it.
But I'm looking forward to see everybody face-to-face, no AI assistants interfering. And you can find me on LinkedIn as Frederic V. com.
And as for me, of course, I will also be at AI Field Day 8, which again, is May 13th through 15th. And Frederic, I believe that you have been to all of them. Also this week, as you listen to this, I am at Qlik Connect in Orlando.
Qlik is a great data company, and we're going to be talking a lot about the relationship between data and AI and how that all works out. And so I look forward to the experiences there. Of course, it's going to be at a great event.
Frederic, you're at Qlik Connect as well, so folks can learn more there. " It seems like that's my experience every week. " So thank you all for listening, especially Dave, Frederic, thank you so much for joining us.
It's been great to have this conversation. I do look forward to having more of these conversations in the future. If you're listening to this and you enjoyed this, please do subscribe.
The best place to do it is YouTube, where you'll see our lovely faces. But of course, you can also find us in your favorite podcast application. And reach out.
We'd love to hear from you. This podcast is brought to you by the analysts and experts from the Futurum Group, where insights meet AI. ai, the "Utilizing AI" YouTube channel, or the Techstrong TV app.
Thanks for listening, and we will catch you next week.