Enterprising Insights: – Salesforce Announces an LLM Benchmark for CRM, Episode 33
In this episode of Enterprising Insights, The Futurum Group Research Director Keith Kirkpatrick discusses Salesforce’s LLM benchmark for CRM, and rants about the lack of transparency around the collection of data from Hyundai, Honda, and GM.
Transcript
Hello everyone. I'm Keith Kirkpatrick, research director with the Futurum Group, and I'd like to welcome you to Enterprising Insights. It's our weekly podcast that explores the latest developments in the enterprise software market and the technologies that underpin these platforms, applications, and tools.
This week I'd like to talk about Salesforce's recent announcement of what it claims is the world's first LLM benchmark for CRM. So, we're gonna get into what that benchmark is, why it's important, and assess whether or not we're gonna start seeing others like that. Uh, benchmark start to crop up in the market, particularly as we are kind of entering this new phase of generative ai where, you know, we're moving away from just sort of utilizing a large language model, you know, sort of a generic large language model and starting to see, uh, more purpose-built ones and ones that are being tuned for specific applications, specific tasks, and so on and so forth.
Then, as always, I'm going to move my rent or rave segment where I pick one item in the market and I either champion it or I criticize it. So, without further delay, let's get into this week's topic, which is Salesforce's announcement from a new benchmark for LLMs, uh, you know, tailored toward the CRM application. So, yes, so Salesforce did announce this, I believe a couple weeks ago.
Now I'm a little late to get to this in terms of my schedule, but, uh, I, I do think it's an important topic. So this new benchmark, uh, it, it's essentially according to Salesforce, it's, uh, a comprehensive evaluation framework that measures the performance of LLMs against four specific measures, LLM, accuracy, cost, speed, and then trust and safety. So, uh, what this benchmark is really trying to do is evaluate common sales and service use cases, including things like prospecting, lead nurturing, and sales opportunity and service case summers.
So it is not sort of a generic benchmark trying to assess the LLM performance against sort of, um, you know, generic factors or just overall speed or overall accuracy. That's not what it's trying to do. What it's trying to do is look at an LLM in context being used for these very specific use cases within A CRM application.
Now, this is really designed to help people who are using Salesforce really evaluate, uh, other LLMs and choose which ones are going to work best for their specific use cases. Now, why, why is there even a need for this? Well, there are existing LLM benchmarks, but they're really sort of academic in nature, or they're focused on consumer use cases.
It's not really relevant for business users. Uh, I think one of the other things that, that they mentioned is that it really existing ones are very, very much focused on sort of pure, uh, benchmarks with respect to compute and how quickly it returns a result. They don't really have sort of these expert human evaluations going in to really address some of these other issues and, and making sure that the, that there really is accuracy in context.
And, you know, the other thing of course, is trust. Making sure that the LLM doesn't go off the rails, uh, and he'll hallucinate. Uh, this is really about, you know, looking at something in practice, uh, to, to give you a, um, you know, sort of an analogy.
If we look at how automobiles, you know, if you read any kind of car enthusiast magazine, whether we're talking car and driver, MotorTrend, whatever, they'll give you sort of the, um, you know, the specs that the, um, that the manufacturer provides. Then they'll run their sort of road tests and, you know, give, let's say a zero to 60, uh, you know, assessment on how quickly a car can accelerate. That's great, but that doesn't necessarily translate to sort of the real world of, you know, driving a car.
That's where you need to actually conduct other tests. Things like, you know, what is the speed to accelerate from 55 to 75 to pass, uh, you know, how long does it take to do that? Because that's really looking at the vehicle in context.
And then of course, you would need to do that, you know, to, to really get an accurate benchmark. You'd really have to incorporate other factors such as, you know, the drivers' confidence and skill level, uh, all of that kind of stuff. And that's what this LLM benchmark is trying to do, is trying to look at LLM performance within the context of very specific sales and service-based use cases as it pertains to being used with Salesforce's CRM.
Now, I think what's really interesting, and I'm gonna try to include a screenshot of the interface, which actually has this benchmark, and basically it will line up a number of different, uh, LLMs and it will go through each sort of attribute across the different four different categories of accuracy, cost, speed, and trust, and assess each LLM based on that four specific use cases. And I think what is, basically, let's just take a look at some of the, uh, attributes here. So, accuracy, this is probably the, they're all important, but accuracy is a big one.
Uh, so this metric actually contains four different subcategories, factuality, completeness, conciseness, and the ability to follow instructions. So if you think about, um, accuracy as it pertains to a model, it really is about combining all four of those subcategories together. Because even if, you know, uh, an LLM were to provide a factual answer, it may be not necessarily, uh, following the instructions of the prompt as accurately as we might want it to be.
Same thing with, uh, completeness. Uh, if you think of, you know, a, if you were to ask the LLM to return something and it return, it returns a result that is partially correct, but it's not fully complete, that may not be yet useful in the scope of an actual business use case. So, uh, this is interesting that that accuracy measure, uh, contains a number of different subcategories, and that's why it's important to consider all of them together to, you know, to really assess the performance of a particular LLM and its value as a business, uh, tool cost, this is another one.
Now, cost is sort of, this is interesting because it's categorized as high, medium, and low based on a percentile basis. Uh, this is an estimated operational cost that will vary, uh, by each use case. And really it's more of a comparative, um, uh, scale where you're comparing different LLMs against each other based on a specific use case, uh, and is really a relative measure as opposed to a, uh, you know, a, a, you know, a hard and fast, um, you know, cost of X amount to, to operate, because that's very, very difficult to, to ascertain because, uh, you know, it, it will vary based on exactly what goes into each prompt and, uh, what type of data is being queried, all of that kind of stuff.
But you can get a relative idea on the cost of a model versus one versus another. So I think that's interesting. Now, uh, the third one is speed.
This is a metric that assesses, that assesses the LLMs responsiveness and efficiency in processing and delivering information. Obviously, faster response times will improve the user experience and reduce the wait time for customers when deployed in certain customer facing applications. And of course, it enables any kind of service or support teams to address inquiries quickly, uh, more quickly if the LM is responsive.
And, you know, again, this is, uh, interesting because certain things are going to take the l lm longer, uh, to respond to. And that's where it's important to look at which l and m works best for a specific, uh, use case. Uh, so you can determine which one, uh, might be most suitable.
And then of course, you're gonna also have to balance that against some of the other factors in there as well, because again, uh, all of this is about, it's never, it's not necessarily going to be about, Hey, let's just pick the one that is, you know, has the top numbers in a specific category. It's really about evaluating it in context. And then, of course, the fourth category here, and this is something that a, uh, that Salesforce has been sort of pounding the table on, uh, really since generative AI became a thing about a year and a half ago, uh, that is trust and safety.
And this metric is designed to measure the LLMs capability to shield sensitive customer data, adhere to privacy regulations, secure that information, and, uh, refrain from incorporating bias and toxicity for various CRM use cases. Now, this one is interesting because it's, uh, it's really looking at, uh, the model and, and framing it in the context of adhering to those specific, uh, criteria that's gonna be really important for organizations as they start to evaluate other LLMs. Uh, and, and obviously even, um, you know, as we move to more of an open model for generative ai, ai, um, you know, I expect that that's going to be increasingly important.
Uh, you know, we're gonna get into a little more, a little later on in terms of looking at, uh, these issues, particularly around, uh, the use of data and how that really will impact an organization, uh, you know, and its relationship with this customers. Now, I think there's a couple other things here that I think are worth mentoring. Uh, I, I do think that Salesforce, this is a good move for them in terms of kind of being out in front of the issue of trying to assess the performance of LLMs in context.
Because really if you don't assess an LLM using very, very specific or standardized metrics, it's very hard to compare what is the best one for a specific use case. Uh, I do think that we're gonna start to see other platforms also develop these types of benchmarks, and perhaps we're gonna see some third party, uh, uh, companies provide the solution as well. I still think though, it's going to have to be an evaluation that's done in a very, uh, specific manner.
And what I mean by that is you can't just have this one sort of, you know, okay, uh, let's compare these models because you're gonna have to look at what platform are you using, what data are you using, uh, or type of data are you using? What use case? And, and that's gonna be very, very specific to the organization and the tools that they have in place.
Uh, I do think that, um, looking at this within the context of data available is going to be, it is not an easy process. Um, but, you know, handing a tool to kind of start that process really kind of gets you that much further down the road. Uh, particularly as organizations, end user organizations start to incorporate things like, um, uh, you know, small language models, that type of, uh, you know, those types of very, very specific models, uh, that's gonna be important to them, uh, to, to really ascertain what is the best approach with generative ai.
Is it taking a large, you know, open model with, you know, however many billion uh, parameters like, you know, open ai, uh, or is it going to be looking at a model that might be a little more expensive, uh, to use because it has been purpose built? And if so, if it performs better on some other metrics, perhaps that's the best use case. So that's gonna be very interesting to see how that kind of pans out over time as organizations get more familiar with generative AI and really start to crystallize their vision for what they want to do and what business outcomes they're looking to, uh, to actually enable using generative ai, obviously in conjunction with other technology.
So, uh, I think to kind of wrap up on this, uh, it'd be interesting to see where we're headed, but I think this is the start of a trend. I do think we're gonna see some more of these tools, uh, come into the marketplace over the next several months and years. Okay, so with that, I want to move to my, uh, rant or raise segment.
That's where I pick one item in the market, and I'll either champion it or criticize it. And this week I have a rant. So this is interesting kind of piggybacking on the data security, data privacy issues that I was just speaking of.
So apparently the New York Times just reported that senators Ron Widen of Oregon and Edward j Markey of Massachusetts recently sent a letter to the FTC on July 26th. And in this letter, the two senators called that General Motors, Hyundai and Honda for collecting, driving data, uh, from customer VE vehicles. So when we're talking about driving data, that's things like, uh, how fast the driver accelerated, how hard they brake, uh, how often they went over the speed limit, you know, are they, uh, breaking hard a lot, all of that kind of stuff now.
And they're saying that this data was sold to insurance companies so they could better gauge driver risk. Now, uh, basically this is, uh, really interesting because at the heart of this, it's, um, Senator Wyden said it was deceptive on how it got people to opt into this. And just as a little bit of background, you know, obviously we all have insurance, auto insurance companies, and many of them have similar programs where, uh, one of them, Allstate, I think it's drive wise, where the idea or the way it's market is, it says that they will collect data from you, the driver, and then, you know, it gives them a better picture of your risk.
What they don't always say is that obviously if you engage risky behavior, your rates were probably gonna go up because you were a higher risk. So what this lawsuit does though, is it is saying that it collected data from vehicles with an internet connection. And, you know, it sounds like they weren't very, you know, at least the letter says that, uh, GM and Honda actually gave the drivers a choice to opt in, but it was deceptive.
Now, I think that, um, really when we're, we're talking about this, there's, there's, there's two issues here. One is, um, you know, were these auto companies clear about the fact that the data would be collected? Number two, were they clear about, you know, what this really meant in terms of, uh, the potential for that data to be sold to another party, essentially insurance companies.
And then what would the impact be or potential impact be from those insurance companies then using that data? Uh, I think, you know, regardless of whether or not you agree with, um, you know, that type of data being used to assess risk or not, I tend to think, hey, it's kind of fair game in the sense of what are insurance companies really, you know, they are basically risk managers. That's all they are.
Uh, and, and obviously their job or, or the only way they really are successful in the market is that they can accurately assess risk and then price their services based on that. Um, I think the bigger issue here is really looking at, uh, you know, again, disclosure issues and deceptive language when it comes to the collection and use and potential sale of the data. A as you know, as we become a data-driven world, it is really important that organizations, uh, you know, are transparent about this process.
You're gonna get some people who are perfectly fine with this. I mean, I actually just spoke with my mother about this, my nearly 80-year-old mother who is, you know, participates in one of these programs, and she says, I'm fine with it because I don't drive very much and I don't wanna pay more, and I know that I'm, you know, uh, gonna not be driving aggressively or crazy or anything like that. So I'm happy with that.
So she's okay with that. But I think the issue is, is again, letting people know, being transparent about it, because obviously the more data that's being collected, it's also going into these algorithms and, you know, it can have an adverse effect on people without them knowing. And the fact that the organizations are not being, uh, clear about it or transparent is, is kind of a, uh, a major impact or can have a major impact on consumer trust.
And that's why you get people who are running around, you know, you know, obviously worried about their data, how it's being used, is it being resold and so forth. So, uh, I do think that that is a negative, um, you know, mark on, uh, or signal to the market in terms of data. 'cause obviously if their intentions were completely altruistic, I don't think they would have that sort of, uh, what the senators say is deceptive language.
Now, the other thing that was very interesting in this piece that I read about this case, um, these automakers actually didn't make very much money from selling this data. Now, according to this letter, uh, uh, s paid Honda $25,920 over four years for information on about 97,000 cars. And if you look at the math, that's about 26 cents per car.
Now, Hyundai was paid just about $1 million, or 61%, 61 cents per car over six years. And, you know, to me it seems like a heck of a lot of negative potential publicity for not a heck of a lot of money. Um, but, you know, again, it's possible that they were looking at this in terms of something that, um, you know, as an incremental revenue, uh, stream, or they were looking to maybe at some point expand it.
But it just seems to me that it's, it's strange that they would engage in this and potentially use deceptive tactics for really what is just honestly, uh, you know, pocket change. So, very interesting. Uh, again, I think the big takeaway for me with all this and why I'm, I'm ranting about this is again, we live in a data driven world.
Everyone knows that the, that companies are looking to collect data at this point. It just doesn't, to me make any sense to violate any kind of trust you have with consumers by being deceptive or less than transparent about the data that is being collected and what it could be used for, or that it could be sold. Because in the end, um, you know, when when we think about, uh, consumer choice, that could have a very negative effect if people feel that they were duped.
So, uh, that is my rant for the week. Uh, so I wanna thank everyone here for joining me at enter, uh, on enterprising insights. I'll be back again next week with another episode focused in on the happenings within the enterprise application market.
So be sure to subscribe, rate, and review this podcast on your preferred platform. And I will see you next time.





