“Full Frontal Chaos” of AI: Unpacking the Current Wild West of AI Tools | Utilizing AI Episode 28
The artificial intelligence market is currently experiencing a period of total chaos, leaving businesses and developers scrambling to keep up with an overwhelming flood of new models, frameworks, and technologies.
Recorded live on-site at AI Field Day 8, host Stephen Foskett sits down with industry experts Guy Currier of The Futurum Group and Frederic Van Haren of High Fens Inc. to discuss this rapidly evolving AI landscape. They dive into the rise of “neo clouds” engineered specifically for low-latency AI performance, the shift from massive large language models to specialized small language models, and the skyrocketing token costs associated with multi-agent workflows.
Ultimately, the guests share actionable advice on how organizations can conquer the fear of missing out, establish stable business baselines, and embrace constant market change as a competitive advantage.
This and more on Utilizing AI, part of The Futurum Group Podcast Network.
Transcript
The AI market is experiencing a period of total chaos with new technologies, models, and products introduced every day. On this episode of the "Utilizing AI" podcast, Guy Currier, Frederik Van Haren, and myself, Steven Foskett, discuss the current constant state of change in AI. Welcome to "Utilizing AI," the podcast focused on practical applications of artificial intelligence from the Futurum Group.
Every Wednesday, we explore news and use cases of the ways in which AI is transforming enterprise IT and the industries it serves. I'm your host, Steven Foskett, president of the Tech Field Day business unit here at the Futurum Group. Before we dive into the discussion, let's meet who's on the panel today.
I'm Guy Currier. I am an analyst at the Futurum Group and also the chief analyst for Visible Impact, which is a division of the Futurum Group, and I study AI and AI infrastructure, among other areas. I am Frederik Van Haren.
I'm the CTO and founder of Hyfens, and we provide AI consulting and services. We're recording this episode at AI Field Day 8, which was attended by Guy and Frederik, as you probably guessed, considering we're sitting in the same room together. And much of what we say is informed by what we learned at AI Field Day, but it's not necessarily directly attributable to it.
Now, the topic of this week's episode, it actually did come from a comment that Guy made during AI Field Day, and that is just the observation that it's a period of, I don't want to sound too much like the introduction to a "Star Wars" movie, but it's a period of great chaos in the Galactic Empire. That is AI, yeah. There's so much going on right now with new products, new models, new approaches, new technologies, and every one seems more transformative than the last.
It can be so hard to keep up with everything that's going on. Guy, what's your impression of the total chaos of the AI market? I just think that there are so many entry points right now to AI.
We've talked about it since the ChatGPT moment at the end of 2022 as the most pervasive, in and outside of technology, the most pervasive trend that we're going to be experiencing, and are experiencing now. So, that means it applies to every application. It applies to every layer of the stack.
It applies both within it as something like a feature or capability we're seeing more and more in SaaS and software, just to pick one layer, as well as a way to manage these stacks and applications and their interactions with each other. So I'm just starting to feel like I don't know how to advise myself, my clients, customers at this particular point in time as to what the best avenue of pursuit is for them, strategically speaking. Because we are seeing so many different ways in which AI is being added or used, and so many different ways of adding or using it too.
Big models, small models, agentic, working together, working at the edge, working at... Well, anyway, talking about it makes me feel chaotic. Yeah, it's like drinking from a hose, right?
If you see what's happening in the market, there's so many new models, new frameworks, and it's very difficult for people to start working on a project only to realize that something new came out, a new framework, a new model. And I think from that perspective, it becomes a little bit chaotic for the simple reason that people don't know exactly anymore what's right for them to use because there's so many different options. And the other challenge is you don't know if it's good for you or not unless you use it.
And so it's kind of, I think NVIDIA is calling it the flywheel. It's almost like it doesn't stop. Absolutely.
Cats and dogs living together, right? I know. What's the world coming to exactly?
But it really doesn't stop, and I think that that's the thing that's really been challenging for us because Futurum is a very AI-first company. Hyfens is obviously out there on the cutting edge of all of this technology. And I think the three of us are trying to keep up to speed with everything that's happening, and yet it's challenging for us.
Can you imagine what it must be like for people whose job isn't focused on AI? Right, exactly. And even if you kind of know what you want to do, the supply chain is not really helping you, right?
So even if you know that you want to build something and it's going to take a certain amount of hardware, the supply chain is adding to the chaos, to the sense that not only is it more expensive, you don't know when you're going to get it. So a lot of people are trying to find different approaches. For example, some of our customers who typically were on-prem are now going to a Neo Cloud, because a Neo Cloud has the capacity, and it's somewhat predictable.
But it's a little bit like during the COVID days, people also went from on-prem to the public cloud and then eventually came back. So it's an additional level of chaos if you ask me. Well, I want to jump into that Neo Cloud topic as well, because that's a term that is thrown around a lot, and I was actually having a conversation yesterday with a few of the AI insiders that were presenting at Field Day, and there is a little bit of, I don't know, confusion in the marketplace about what is a Neo Cloud.
How does that differentiate from, I don't know ... an oldo cloud. And why do we have that term?
Can you explain that? Yeah, definitely. So everybody's aware of public clouds, and public clouds, the main feature there is the flexibility.
So you can compose supposedly your own infrastructure, your own kind of storage, and compute, and so on. So the challenge with that is that AI needs very performing compute, very performing storage, and it has all- Networking ... networking, yes.
And it all has to work really well together. Additionally, models need to be involved in the picture. So if you're on a public cloud, you have to download all these models, you have to upload all your API keys.
And so it becomes very difficult, but not impossible. It's not impossible. What the NeoClouds are doing is they're delivering you a platform where you have access to the models, where you have access to an AI framework with compute, storage, and network all built in, and you can also define what kind of AI you're doing.
Is it training? Is it inference? And so it's almost like your time to market, or time to development, or time to deployment, if you wish, is very small with the NeoClouds.
And I think it's really important because it's part of the chaos in the sense that people don't have the time to build all these frameworks and make it work. Right? Installing software by itself is not necessarily that difficult, but making it work fine together, and I think that's where the NeoClouds are coming in.
So they deliver a platform that is fully focused on AI, and it allows you to choose and pick the size of the infrastructure. Can I offer an analogy, though? The CPU on a system of whatever size is a general purpose processing unit that quarterbacks and runs, manages, and orchestrates as a central point, even if it's only a small area of what a particular application is doing.
It might be more storage focused or networking focused, what have you. But it's general purpose. That to me is like the hyperscalers.
They are designed with what at this point is a giant variety of services who are doing all sorts of things, including AI, including agentic AI. Meanwhile, you have specialty processors, typically the GPU, but it's not the only one. And then there's even a whole category called application-specific integrated circuits, ASICs.
And those are particular. They can do fewer things, but they do them very well, and some of them are tuned particularly for AI as it happens. That to me is the NeoCloud corollary.
Tuned, created, managed specifically for all the different AI usages, and really exist as a corollary to the general application state that an enterprise has. Does that make sense? " They abstract the hardware.
They provide frameworks. They provide platforms. They actually go much higher.
So why do we need NeoClouds when we have Amazon? We need NeoClouds because of the volume of processing of inference or training, or what have you, across what are multiple life cycles for any given AI service application or agentic host. And those have to be done usually within a really specific power envelope for the region.
I don't mean for a particular rack, for the region. This is a build-out that really has just begun and that needs to be concentrated on. I'm trying to put this in really general business or societal terms, not specific to technology.
And so that focus is really helpful for that build-out. That's what I would say. Right.
There's no doubt that the NeoClouds are a niche play within the public clouds, but we also have to look at the under- It's a huge freaking niche. Right. It's kind of my point, but go ahead, yeah.
Yeah, and I think the underlying infrastructure is also slightly different. When we look at public clouds, the hardware, or at least the hardware from a consumer perspective, is composed. So your storage and your compute is not necessarily next to each other, and the networking they provide is not necessarily the highest performing, low latency infrastructure, while the NeoClouds specifically go after that market.
So, there's definitely overlap. But I think what's really important there is the time to market and the ease of use. Not a lot of people nowadays understand how to set up infrastructure, download models, and make them available, and the NeoClouds make it really simple.
And I think that's helpful because many people might be wondering. For example, I know that Microsoft actually works with the NeoClouds as an additional data center provider, and it can be a little confusing wondering exactly how these puzzle pieces fit together. If we move a little bit up the stack, of course, though, you talked about platforms, application platforms.
This is another area that's rapidly just diversifying and exploding in AI. All the different ways that you can compose and build and deploy applications, AI applications. It's just, again, every day it feels like there's a new AI platform or there's a new one on the rise or a new one on the fall.
We had a conversation, another "Techfield Day" hallway conversation, where people were joking about how great one platform is, but nobody uses it Instead they all use this other one. That can be very confusing as well. I don't know if you can mix and match.
That's the problem. I think that it's in some ways getting us back to the old days of application service provision. That's what I'm really trying to get at.
So the old days. Why did ASPs appear? They appeared because the architecture of enterprise applications at that time, this was like proto-internet days and stuff.
The requirements of the application started to get fused more and more with the requirements of the infrastructure. And so it was a useful model to have someone else concentrate on all of that because you had to run your business. And in the same way, I feel like there's so much chaos and variety in- Would you call it total chaos?
Total. I would call it full-frontal chaos is what I would call it. But there's so much chaos.
There's so many possibilities and variety. You land on... You're starting with some mandate from your board or something.
Add AI everywhere, or something like that, or your vendor's doing it. But now you need to get down to the specifics. What are you trying to accomplish?
That leads to so many different possibilities. Frederic, I know this is like what you do all the time. Yeah.
Right. One of the big challenges is that a lot of organizations that are not doing AI or basic AI are basically looking at their competitors, and their competitors are supposedly doing a lot more with AI, and it becomes a time to market issue. Mm-hmm.
Where they feel left out, and then they want to even move even faster, and they have no way to achieve that. And I think the additional challenge there is, is people don't know, should I build my own models? It sounds crazy if you want to build a large language model.
That's the thing, though. So in the last three months, or maybe six months, there's been more interest in small models. Right.
Because people are understanding how to use them. Right. But it also has a different effect, right?
If you have a large model or a large language model that contains way too much and you know it, deploying a million copies of that is not cost-effective, right? Mm. That means you need to buy a lot more powerful GPUs.
As a matter of fact, if you look at what NVIDIA is doing, the NVIDIA RTX series, their latest card is not a faster, better card. It's a card that is less capable, less power, but to accommodate people that are trying to downsize, if you wish- Mm-hmm ... and build large language models that are more focused on their vertical.
Mm-hmm. But it's a very good point. I think I've been hearing that for a long time.
The challenge in the past was that people were using RAG to supposedly put guardrails on the large language model. So you generate all the tokens, but then you put the guardrail on to make it look smaller. Now we see more and more trends towards a large language model or like a medium large language model, I guess, whatever you want to call it.
Well, SLM, small language model- Small, yeah ... you know what the- But it's all relative, right? Mm-hmm.
When I started with large language models, that was kilobytes. Yeah. Well, that's the funny thing, that now what's called a small language model.
So when Frederic and I started utilizing AI in 2002 or something- Mm ... we were saying, "Oh my gosh, will we ever get a billion-parameter model? " Right.
And now that would seem laughably small. Mm-hmm. I remember that that was one of the specific questions I used to ask.
But you're totally right that people, I think, are starting to evolve away from the number of parameters because essentially, as an industry, we already have models that surpass the entire record of written human communication. That's long since in the past. I think there's a rising understanding that the progress made on models is sort of going to be going asymptotic to 100% general intelligence, that we'll never quite get there, and that there's sort of a real diminishing return of having bigger models.
And in fact, people are now focusing on quantizing- Right ... and on reserving memory for context and key values and tool use and other ways in which the biggest model is absolutely not the best model. And I think that that's confusing to many people because in popular culture, all the talk is about the latest, biggest, giantest chatbot and how smart it is.
But to implementers, that's almost irrelevant now. Right. Exactly.
Yeah, I think the large language models and the billions of parameters was really driven by the people building those models. Mm-hmm. Right?
But people indeed are more interested in accuracy and on topic. Yeah. If you're a hammer manufacturer, it pays to be a good hammer salesman.
Yeah. But there's other tools out there. Right.
Unless you can turn everything into a nail. Yeah. And I think there's a little bit of that in this...
I just made the easy joke, but there's a little bit of that in the general trend and still the popular discussion around AI. Namely, the hammer is more and more and more compute jammed into smaller and smaller places more efficiently for power utilization. Right?
And that is the hammer. And so the large models, the billions and billions, billions and billions of parameters that's the nail. Right.
And we're starting to see that not everything's a nail. Right. And I also think that the conversation is changing a little bit In the past, it was all about billions of parameters.
Now we're talking about the amount of tokens being generated. Yes. I mean, in- Well, now it's billions and billions of, well, however many tokens, right?
Right. Well, and it's considered like a productivity thing, right? Yeah.
How many tokens did you generate? I mean, it's really, that's really- I was struck by one of the examples that provided at AI Field Day recently. That's where we're recording now.
I was struck by one of the examples at AI Field Day that walked through a given prompt and query and response. Mm. And what struck me about it was that this is a typical, it was essentially a recursive process that multiplied the number of tokens required.
And I was looking at it, I was thinking, wow, a little bit of process flow design here, and you actually don't have to be as recursive and as repetitive, reducing the amount of tokens. Now, that was a simplistic approach to it that might not work at all, give it that. But that's the kind of thinking that, like human thinking, and AI-assisted human thinking we need to bring to these in order to have really efficient utilization of resources.
And now we are out-- And it's token model size, it's token usage, it's flow of agentic thinking, quote unquote thinking. A lot of options there. Yeah.
Too many options. It's just chaos. Yeah.
I think we should come up with some kind of a factor that includes accuracy. How many tokens did you generate, and what's the accuracy and on- Accuracy per token, practically. Something like that.
It's very difficult, of course- Neat idea ... to figure it out, but- Yeah ... it's something like that because- Effectiveness per token.
Yeah. I mean, generating a million tokens, what does that mean? What did you just generate?
Yeah. Did you just figure out what the weather is going to be in five different places? Well, that's a good point because right now, I think that the trade-off is we're still in the phase of AI where it is better to burn more in order to get a better result- Mm ...
without really worrying about efficiency. And to me, that's another thing I think that demonstrates the sort of total chaos that we're surrounded by, is that efficiency, despite the talk about power consumption and water consumption and-- Efficiency is very much an afterthought at this point. Right now, I think that when looking at, if we want to kind of turn the page here to-- When we look at agentic applications, what I'm hearing from people that are developing and deploying these is that they are very happy to use bigger, more effective, and more token-heavy models in order to get a better result.
And not only that, but that they're starting to do, frankly, what I do when experimenting with agents, which is stacking models on top of each other. So essentially, I have an orchestrator model, a coach, and then I've got a whole team that are all doing different aspects of the task, and many of those are actually checking each other. And so, this one's doing the research, this one's filtering the input.
This one's doing the, quote, thinking or processing, and then it feeds it to this other one, which criticizes its thinking and feeds it back- Mm ... and so on. And I'm seeing a lot more like that.
And every step burns more and more and more tokens. And frankly, I see it in myself, but I also definitely see it in the industry. There is no concern, apart from the financial concern, for token effectiveness and efficiency.
Right. Exactly. It's because the context has to move from one agent to another agent and, yeah, there is no stopping to it.
And I think at least companies like Anthropic, they have now tutorials where they're kind of explaining, well, we have different kind of models. But you don't have to use always the biggest, meanest model. You can start with some other model, do some planning with that, and then move on.
Isn't the biggest, meanest model just the easy button? Well, it- In a scenario where admittedly, you don't go from easy to less easy. You go from easy to freaking difficult in one step, practically.
But it's like the easy button. There's a lot of other things we'll figure out later. For now, you know.
Mm-hmm. Right, but I think the- Bigger is better ... " Yeah.
And so- Is that even possible? Yes, there- Well, we all try to figure it out. No, there is documentation.
I can't remember the name of it. There's explanatory documentation. Is it- Yeah, that would be like 1,000 books maybe.
I don't know but- I don't remember what it's called, and frequently it's seen as inadequate. But it is out there, yeah. But the- We'll put it in the show notes.
But yeah, but the definite challenge there is that with these large language models, you don't really know exactly what's happening. So the reason why people are excited about a new large language model is to find out, will it do better than before? Will my prompt finally work?
Yeah. Maybe, yeah. Yeah.
Well, it does seem sort of a, let's just throw the kitchen sink at it- Right ... kind of approach. And if you think too about these reasoning models, the way that they're doing that is just by running again and again and again.
One of the things that also gets me is that the token usage of these more and more advanced models is just off the charts. 3 and three times as high. And of course, those models are more expensive, which means that running the same query yesterday costs maybe three times or four times less than it does today.
And there's very little clarity to the users that that's going on. Can you imagine if tomorrow you woke up and got in your car and it got one quarter the gas mileage that it did yesterday? My car.
And everyone was totally fine with that. Mm. Okay, that was a drop-the-mic moment, I guess.
So what do we do here? So as we wrap up this discussion, how are people going to deal with this total chaos? I'll start.
Normally, I want to say normally, this is a trend, or this is a stage in one of these trends, especially big ones, that's normal and natural. It happens all the time, going back through technological history of humankind. So one thing is enjoy it, live in it, revel in it, ride it out.
But usually what I've advised customers is pick something and work on it. And avoid fear of missing out, or FOMO, avoid recriminations about what you could be doing, and just pick something. But I got to be honest, I feel real hesitant about that in this particular case, simply because of the multiple dimensions of options involved.
Yeah. So I think what I would say, just to close it out, what I would say is, it's an evergreen statement, but really critical right now. What do you think AI can do to help?
Start with that and spend a little bit of time with human beings cogitating over what it is you think you can accomplish and how you think you can test that you can accomplish that. Maybe do that for four things and then take it to the next level of these two seem the most promising, or this one seems the most promising, and then try that. That's the best I can come up with right now.
Yeah, my answer's kind of similar in the sense that first of all, it's all about education. It's to educate people that want to do AI, what their goals are, and I'm talking about business goals and where they want to be. And then secondly, if they want to start applying AI, that they should start with a baseline.
AI is about learning. So AI today is not going to be the AI of tomorrow. So there will be changes.
So change is a default. There's nothing you can do about it. So you're saying turn and face the strange?
Is that what you're saying? I guess so. But yeah- That was for you ...
but they have to start with a baseline, right? And the baseline is how accurate can I achieve with what I'm doing today? And it doesn't have to be a large language model.
It can be something that you download. But try small, but do it with business considerations and don't get derailed by a new platform or a new model. Just create the baseline, know where you are, know where you stand, build your business around it, and then increase the infrastructure and different types of models.
And the other thing I think that I would point out is that this is still very cutting edge stuff, and I don't think people should feel bad about finding that they are behind the times. One of the most valuable lessons I learned when I was a consultant was the value of saying, "I don't know" and the value of asking people and admitting the limits of your knowledge. I know that's difficult sometimes for the nerds among us, but the truth is you don't know.
Nobody knows all this stuff, and we can learn so much from each other if we're only open-minded enough to listen. And so my advice to anyone that's trying to get their heads around what's going on in AI is to continually listen and to listen to a diversity of viewpoints, not just the hype that's out there, but listen to some of the naysayers as well, because they have a lot to teach you, too. And ultimately understand that none of these people have the answer.
None of these people know what is going to come next, and nobody is yet at the point where they are building a fully realized and lasting AI infrastructure. We're all still in the building phase, not the living phase. And so as this becomes more real, and as this matures, what we're going to find is that what today is a very useful and productive AI application, next month or next year is going to be woefully behind the times.
" Mm. On the plus side, what is today expensive and difficult to run will in next year be actually much more affordable and approachable. And so this is the time to experiment, this is the time to look for new things, and this is the time to reach out and try to do something really cool with this technology because it is going to change and there's going to be really valuable things that you're going to uncover as you're investigating AI.
So thank you so much for joining us at AI Field Day, AI Infrastructure Field Day, and on the "Utilizing AI" podcast. Before we go, where can people continue this conversation, Guy? Well, you can find me at LinkedIn.
LinkedIn is a place where I put a lot of my activity and thoughts and ideas. com. I'm one of the analysts there, and I publish there with some regularity.
Yep. com. And of course, you'll find me here every week on the "Utilizing AI" podcast.
So thanks for joining us for the "Utilizing AI" podcast today. If you enjoyed this discussion, please do subscribe on YouTube or in your favorite podcast application, and consider giving us a rating or a review. This podcast was brought to you by the analysts and experts from the Futurum Group, where insights meet AI.
ai, the "Utilizing AI" YouTube channel, or the Techstrom TV app. Thanks for listening, and we'll catch you next week.