Kamal Ahluwalia on the Rise of Domain-Specific AI Models for Enterprise
Kamal Ahluwalia explores the shift from large language models (LLMs) to smaller, domain-specific models (SLMs) for enterprise applications, where cost, relevance, and data privacy are key priorities. Kamal Ahluwalia emphasizes that while LLMs offer broad capabilities, SLMs provide tailored, efficient solutions that align better with the specific needs of businesses.
Transcript
Hello and welcome to the latest edition of the Techstrong AI video series. I'm your host, Mike Bazar. Today we're with Kamal Al, who is president of iki AI Labs, and we're talking about the rise of well small language models.
Kamal, welcome to show. Glad to be here, Mike. Thanks for having me.
Everybody and his brother probably at this point has heard of a large language model, but I wonder if, um, that was just kind of the first phase and we seem to be moving towards the next phase, which are these small language models that are a little more domain specific, shall we say. Where are we on this journey right now? Uh, it's coming fast and I think it's surprising, uh, that already at the beginning of the year, uh, Gartner had started to write about what LLMs were good for and what they were not good for.
And there were a lot of use cases where they were, uh, not a good fit. And I think that, uh, what is driving and accelerating this need for better, cheaper, faster way to do this is, uh, the cost is very high with LLMs, so it's not yet affordable for everyone who wants to adopt it. Second is starting with the foundational model that's trained on internet scale data also requires models that actually can handle all the noise and irrelevancy that comes with internet scale data.
So when you bring it into the enterprise, you actually need accuracy and you need relevance, and you need the context, which comes from enterprise data, not from gathering everything that's out there. So I think all these things are sort of necessitating that. Uh, we've become more focused, more relevant, and that's the promise of s SLMs.
In fact, even Microsoft, uh, I think months ago had announced their own efforts to build s SLMs. And uh, the other thing that I'm seeing is with enterprises, there is a very strong push that whatever utilities they're gonna bring from the outside, they should be deployed inside their VPC on the hyperscalers. So what that means is they don't want their confidential data to be going to different people's, uh, tenants in the cloud, which basically means you need to, whatever innovative stuff others are doing, it needs to come to that enterprises instance.
And that clearly is not feasible with foundation models that are based on anonymized aggregate learning. So I think what we are seeing is really the evolution of something very cool and breakthrough like LLMs, which was fine along with the hallucinations in a consumer setting, but there's a much higher bar when you bring all these things to enterprise. That's what's driving this As they say, uh, small is beautiful, but will small supersede the large language models or will they kind of have two different use cases?
I think to some extent both. And uh, what I've seen is when the two go together, meaning that you're first setting the context because what in between right? There was this whole thing about rag that if you want the relevancy for your enterprise, then you use the rag model so that now you are feeding your, uh, domain specific data to LLM so that now they can answer specific questions.
So I think there, the experience that LLMs provide the chat experience is great. I think clearly the UX is moving in that direction, but within the enterprise you would rather start with a SLM approach, set it up for that domain and context and for your enterprise use cases, and then intersperse that with externally available data that is relevant to the domain. Am I going to daisy chain all these s SLMs together or am I gonna put agents in front of them to create some sort of workflow?
How is that all gonna come together? That's exactly right. It is going to go morph into the agent AI architecture where there'll be dozens and dozens of these SMS that are good at specific tasks and uh, you will be orchestrating across all of that.
You're absolutely right. That is where things are racing. Mm-Hmm.
In some instances, would I not put an SLM in front of an LLM and kind of have the SLM manage that process out to the back end? So, 'cause there is some information I imagine where that proverbial large language model might add some value. Absolutely.
Mm-hmm. Both in terms of experience as well as, uh, for, you know, creating your, crafting your email, creating video multimodal stuff, all those things. I think LLMs will be fantastic.
And I think the overriding thing around this is don't underestimate the cost that everybody's incurring Mm-Hmm. Even in the POC stage, right? So what you're describing is exactly the right way to do it because I think that'll optimize the cost versus trying to use LMS for everything.
I feel like when I look at large language models, part of the cost is traced back to, well this is just a brute force methodology for doing something. So are we getting a little more sophisticated in our approach and a little more nuanced in a way that might continue to drive down the cost of ai? Yes.
Uh, for sure. And also this, uh, 'cause there's also this, there now basically four or five large companies that are doing very well with LLMs, right? Open ai, Google Andro, et cetera.
Uh, then the GPUs are in, are scarce commodity and this cost is through the roof. So right now the ability to use LLMs extensively is in the hands of a few very large tech companies. And that is not good for the industry as a whole.
So I think that will require a far better both to bring down the cost, make it affordable so that you, when, and I've worked with a lot of companies, that is the overriding concern because suddenly the gross margins are getting shot. Mm-Hmm. Right.
So this, all this combination of both the accuracy part relevance, part cost of inference, all of these things are going, that's why it's moving so fast because, uh, otherwise adoption is going to be curtailed. Mm-Hmm. We have this sense of fear of missing out FOMO that has been driving much of our AI activity for the last year and a half or so.
But, um, how long will it take, you know, organizations to actually harness s SLMs, train them and kind of build agents and do something, you know, that runs in a production environment? That's a great question. And uh, interestingly, and I have some data points for you.
Uh, there is a company called Next G ai. They are in production with a couple of very, very large clients and it took about three months, four months to solve some very complex use cases using their s SLMs. And they do have that agent AI framework, what's called swamp, so that the agents talk to each other and get the task done.
But, uh, it only took a few months. So my current company in Qai Labs, we are into forecasting and planning. Our core technology is not based on LLMs, it's based on large graphical models.
So it's a probabilistic distribution of data and we also can be live trained and live and accurate in three months or so. So time to value is an issue because, uh, people do want to see results, measure results, feel comfortable before things are pushed into production. Mm-Hmm.
So what does the future of our workflows look like then? Is it gonna be kind of this mixing and matching of small and large LLMs and a bunch of AI agents that we're orchestrating, uh, somehow or other through some sort of command and control mechanism and us humans sitting in the middle of this kind of environment trying to make sense of it. But, um, do you think we'll have the skills to figure out what to do when and what to hand off to what?
So Here's what I talk about. All our jobs think of it in terms of a third, a third, a third, a third will be automated slash eliminated because of AI performing those tasks. The middle third are where the jobs will change dramatically on what we do on a day-to-day basis, again, because of ai.
And a third will be net new. These will be the new jobs that will be created that don't exist today. And we've already seen some of those like prompt engineering.
Uh, I didn't know two years ago, what, what the hell was that and why do you need it right now? It's a thing. Mm-Hmm.
And same thing around observability and how to test the models because that too is non-trivial to make sure that they are accurate and they're giving the right answer. And, uh, there'll be more of these. And then now if we talk about the agents on how to actually have this library of agents that perform small tasks and how do you orchestrate all those, they will be layer of those and the middle layer will involve human beings also.
'cause in some regular industries, you just cannot let an agent make a decision and execute as well. Right? So you do need the human in, in the loop.
And so it's not like we won't be needed and we are just sitting on the side, but what we do need to do on a daily hourly basis is going to be dramatically different. And I think with a lot of these other things, uh, uh, they will, as confidence grows, it will be a little easier to then just focus on managing by exception rather than having to look at everything. And you've seen that right?
With the Uber, which was basically that you don't need to own a car because it can come to you and you can get to the place and it's like, uh, cost, uh, TCO is about the same or less. So now those thing and uh, then we were worried about the taxis and their business. But now you have driverless cars coming and way Waymo is there in San Francisco operating and in other cities where there is no driver.
So that thing will happen in more and more functional areas. We will see this and uh, we'll adjust to it. So will the nature of competition as we understand it today, basically.
Now come on to, you know, my agents can beat up your agents and hopefully my agents are smarter than your agents. Yes. And that is basically the equivalent of the talent war that everybody is talking about.
Right. And uh, what I'm actually surprised is how fast the programming agents are moving that you can ask, actually ask the agents to program some fairly complex, uh, use cases. And I know of one company that recently got funded by Axel, uh, and the small team is, comes from a consumer background and research background and they're actually very capable of moving SAPs customized post code to clean core.
Mm-Hmm. So Right. So these are fairly complex 'cause you don't even have the talent anyways to actually do all these things.
And that will start accelerate this thing that Yeah. And it all needs to be tied to outcomes. If you can deliver very meaningful results, explainability is a table stakes, you will actually see broad adoption.
And also keep in mind the wallet share fight will drive some of the this, because if you look at the tech stack, most of the enterprise applications are process applications. Mm-Hmm. And as you start to automate more of these tasks, you will see the dollars move either below to the hyperscalers or above to the agent layer and the experience layer on top and the middle when people, vendors in the middle layer will lose out on, uh, their subscriptions.
I feel like I'm all good with all of the things you just described, but I wonder if we're in danger of getting to a point where we don't really know how anything works because some agents been doing it for a while and you know, ultimately, um, that may not be good for us. So that While true, if you just look at, all of us have been using, uh, PCs and the latest stuff and uh, when I went to school and few years after school, like everything, I was opening my PC all the time, right. To put in a bigger, uh, memory replacing this graphics card, this or that.
And now everything I can do with my MacBook, but I've never opened one in a long time. Right. So it's not like I don't need to, and the curiosity is not there, but curiosity is going somewhere else because it does do what it needs to do and you don't need to keep fixing it yourself or having to take it in for something else.
So that's what, where you'll see a lot of these things, yes, there will be automation, it gets to the point where it works and, uh, the next generation won't know how it ever worked, how it ever started to work. Uh, but that's, uh, I think evolution. And I think the main part is to actually, uh, instead of worrying about it, I think let's get into the details.
Let's understand how things are working so that it's not as abstract and scary. But yes, there will be a level of abstraction that is inevitable once these things are solid and robust and doing more and more of these day-to-day tasks. Mm-Hmm.
I remember when I was a child and the television would break and my father would open it up and we would take the tubes out and walk to see and figure out whether it was which tubes actually worked and didn't work, and then walk the two miles back and then put that in. And that was basically the entire day. I think ultimately we found better things to do with our time.
Right, Exactly. Exactly. And now, uh, probably, uh, current generation doesn't even know what those vacuum tubes are.
Not at all. But you, you can see them in a museum somewhere if you wanna check them out. Hey Kamal, thanks for being on the show.
Thanks for having me. Enjoyed it. Thank you for all watching the latest episode here of Techstrong AI video series.
You can find this episode and others on our website. By all means, check them all out. Until then, we'll see you next time.
Take care.