Democratizing Data is the Answer for the Future of AI Strategy | Utilizing AI Presented by Starburst
As enterprises seek to utilize artificial intelligence technology, they would be wise to bring their data and analytics professionals to the table. This episode of Utilizing AI, brought to you by Starburst, features Senior Vice President Jitender Aswani discussing the opportunities presented by AI to democratize data insights with Futurum analyst Brad Shimmin and host Stephen Foskett. Businesses have been dealing with data silos for decades, yet we are institutionlizing this approach with AI applications. It would be better to democratize access to data and enable everyone to directly query data rather than spending time on dashboard development. Data analysts and engineers are ready to move forward and deliver true business insights. The goal is to bring context to data to build trust in AI applications.
Transcript
As enterprises seek to utilize artificial intelligence technology, they would be wise to bring their data and analytics professionals to the table. This episode of "Utilizing AI," brought to you by Starburst, features Senior Vice President Jitender Aswani discussing the opportunities presented by AI to democratize data insights with Futurum analyst Brad Shimmen and myself. Welcome to "Utilizing AI," the podcast focused on practical applications of artificial intelligence from the Futurum Group.
Every Wednesday, we explore news and use cases of the ways in which AI is transforming enterprise IT and the industries it serves. I am your host, Stephen Foskett, President of the Tech Field Day business unit here at the Futurum Group. Before we dive into this discussion, which is brought to you by Starburst, let's meet who's on the panel today.
Hi, Steven. I'm Jitender Aswani. I am Senior Vice President of Engineering and Security at Starburst.
And hi, Jitender. Hi, Steven. I am Brad Shimmen.
I take care of our data intelligence, analytics, and infrastructure research practice over here at Futurum. Excellent. Well, thank you both for joining us.
Now, one of the things that I have been very focused on lately is trying to bring in learnings from other fields to the AI community. And one of the fields I think that has a lot to teach us in technology is data, because of course, in the data and analytics space, they've been trying to understand many of the core concepts that AI people are really struggling with. And they've spent years working on building up expertise and processes around this.
And of course, AI is also impacting the data analytics and business intelligence world in many different ways. So let's kick things off. Jitender, I'm going to throw it to you.
Talk to me a little bit about the ways that AI is changing and challenging and opening up new doors in the data and analytics space. That's a great question. I've had the good fortune of being in this industry for roughly 15 years, starting with SAP, building a lot of these BI products and ETL products.
And I've seen the interface to data evolve over the last two decades or so, starting with SQL as the primary interface to get access to data and then make meaning out of that data. And when that data was represented in some visualization form to be shared with business users, dashboarding became the main interface. And over a decade or so, maybe more than a decade, I've seen dashboards have remained the primary instruments business users have used to get some meaning out of their data.
Eventually, businesses that call themselves data native, they need access to data. They need to understand how their business is performing, what sentiment consumers have about their products, how products are performing in the market. A lot of that requires data preparation, data movement, and then eventually transforming all of that into visualization and sharing that with business users.
So that has been the primary interface. And since 2022, the ChatGPT moment, a lot of us have now become so used to firing a natural language and expecting a response back. And we've started to see a similar behavior emerge with business users with respect to accessing data and gleaning some insights on that data.
So that's where I've seen a pretty strong reaction from business users who expect similar interface from data applications as well. Dashboards haven't evolved for the last 20, 25 years. They've remained static.
Internally at Starburst, we have learned from our customers that dashboards have a very long lead time. It requires what we call a chain of yelling. What basically it means is that if you have a question about something about your business or product or a customer, first and foremost, you need to know as a business user, who do you go and talk to?
And that person you are talking to could be a data scientist or a data engineer. That person now needs to know if they have access to the data. And if they don't have access to the data, they will spend a few days, sometimes a few weeks, try to get access to that data, transform that data, and then ask the question you asked, and then bring you back the result.
We used to say that by the time you bring me back that response to the question I asked you, I might have already moved on. Or if I am still stuck on that question, I may have 10 more follow-up questions, which means you will take another two months to come back to me. And that's exactly how analytics have worked traditionally.
There are some dashboards that allow business users to drill down, but they generally end up creating a lot of boundaries around questions I can ask. And that static form, in my opinion, has lived a much longer life than anyone anticipated. So natural language, instantly firing a question to an agent that is powered by a very smart foundational model, and then expecting it to come back with a response, and then it serves my curiosity, then I can ask 10 more questions because I don't have to understand...
the language of data, which is SQL. I also do not have to be concerned about how to discover data, how to get access to that data. The governance is very well taken care of.
So that's exactly how I've seen is natural language is so friendly to business users, and I think it unlocks a new level of productivity and insights for business users. Yeah. It's funny, Brad and I both laughed about the chain of yelling because, I don't want to spill any beans, but internally in the Futurum Group, we have dashboards.
And we have certainly experienced exactly what you are describing, Jitender. It feels very real, doesn't it, Brad? It does.
It makes me feel like we're back in the 1970s, '80s, and just doing those green and white printouts, the perforated ones. And I feel like that's kind of where BI started and where, just as you said, Jitender, it has remained, and that is this sort of living but not breathing entity. Meaning we set it up, we spend all this time making what we often call pixel perfect reports, and we set those, don't touch them, walk close, softly by them and hope that we never have to go back.
And I don't think businesses work that way anymore. And as you say, Jitender, it's the expectations that companies and the business users have are not the same as they were in 1970 at all. And AI, and generative AI in particular, has shown them that indeed they should be conversing with their data.
The challenge, of course, is how do you equip them with responses that they know to be true, that they trust, and that the decisions they make upon those can be, I want to say, operationalized so that others in the organization can benefit from those conversations so that the insights they gain from their data become a part of the business itself. That's where I think we're going and what excites me about this. I think, Brad, you're hitting a pretty solid point here.
Trust and governance have been the two central pillars of business intelligence. And with AI displacing BI, trust and governance, if anything, their criticality, their importance only elevates. And we have seen evolution of dashboards, but not as much as where the whole cycle time is compressed.
I think the expectations, as you very well said, with business users is that if I have a question right now, I want answer to that question immediately because I pretty much know that I will have 10 more follow-up questions. And I've seen, I've been a business user myself, I have built dashboards myself, that business users as well as technical developers, they are both drowning in the number of dashboards that they end up building. Each dashboard only serving one very small, narrow use case a business user has.
And these dashboards are very well governed. They're audited, they're authorized, as you said, that these are pixel perfect, but they're also built on trusted and governed data sets, which means that I can ask a question, and I can trust that the response I'm getting is accurate, and it won't be misleading. Yeah.
It's a tough challenge, though, isn't it? Because what we've gotten used to with generative AI and agentic processes more recently is these become intent engines or engines of disambiguation for our intent, which are often very confused. But users look up at a chat interface and say, "I'd like to know how my group is going to perform this quarter.
" for instance. And they expect but don't know that what they're going to get back is actually representative of their business, for one, or even the time period they're talking about is accurate. And because of that sort of unknowing, there's a huge sort of trust gap that I think they need to look to companies like yourselves to really fix for them.
We did a survey recently of 818 data professionals, and we asked them the exact question of, what's the biggest challenge for you in adopting generative AI, in replacing dashboards with a conversation? And the biggest one was, of course, hallucinations. Yes.
That's their biggest concern. So, Jitender, how do companies really kind of tackle that? Yeah.
And it's a great question. This is the top concerns our customers and prospects also have. They have seen large language models which powers these agentic systems hallucinate.
Sometimes they come across so confident that you start believing that they're telling you the truth. Only a trained professional or a trained eye is able to understand that we didn't have the data, and you just made up that answer. So now how is that traditionally tackled?
In companies, customers have invested a large sums of money and time in building what they called is a semantic layer. And a semantic layer is basically bringing a large collection of related data from multiple data sources, whether cloud or on-prem data sources, and very carefully curating... the relationships between those data, applying some business rules, and as I just said, business rules sometimes include a very simple glossary.
What FY, which is a fiscal year, what does that mean? FY27. Yeah.
It starts different for different companies. Or what is monthly active users, or what does MRR, monthly recurring revenue, mean, right? That's a business glossary.
So, companies have traditionally built these layers of context in different systems. And then they have struggled to bring all of these together when the applications they have built, or even the more recently, the agentic applications that they have built, they have struggled to provide the right enterprise context to these agentic system. What we have been able to do is because of our open source nature, we are built on top of Trino and Iceberg, and most recently we have integrated MCP into our platform.
We have the capability to be able to discover all data an organization has, and since we have access to all data, we have access to all the metadata, and then we have given a lot of primitives to our customers to be able to build a very rich enterprise context layer. And that enterprise context layer includes things like a data product. A data product is basically, as I mentioned earlier, you bring semantically related domains.
Let me use an example. If you're an e-commerce customer, e-commerce company, you have customers, customers have reviews, customers have support tickets. So just these three domains.
If I bring these three domains together, I can build a very rich set of data products which allow me to understand questions about how my customers feel. What kind of support tickets have they opened? Based upon these type of support tickets and reviews they've given, what has been their ticket size?
What products have they bought? Our customers who have large basket size, I presume they're filing less tickets, things like that. So you bring all of these very different domains within an enterprise together, and then the primitives that we have allow you to now bring all of these together and provide all this rich context to the agent.
Because at Starburst, we learn from our customers, your AI is as good as the data it has access to. Right. So now we have given you access to all your data, regardless of where that data lives.
We don't force re-platforming. A lot of the vendors in the industry force you to re-platform. What that means is that they say you have to centralize all your data into a lakehouse before you can access all your data.
Our biggest differentiator is, no, that's a very expensive undertaking because 25 years in the industry, I've seen most detailed projects. They run very late, and there is a huge amount of cost overruns as well. So instead, I was at Meta when Trino was first founded, and this was our value proposition is that we would not ask you to move your data.
Let data be where it is, and then we will give you access. And as a result, we have access to all the data. We have given you all the metadata and business rules.
Now your agents, which are pretty clever at reasoning, but they are only good at reasoning if you give them very rich context. And that's where we believe that this is where the value proposition of our platform shines, and agents are more trusted as a result. And going back to your original premise is that hallucinations come down by a large percent because of the context these agents now have as a result of Starburst platform.
Yep. Yeah, and I think that this is such a much more customer-friendly, consumer-friendly, business user-friendly approach. Because, as you said, ETL is such an old-fashioned data concept, but yet it's not in the past.
It's as if we were all running around in covered wagons with horses or something, because that is exactly how most organizations work, and more importantly, that's exactly how most AI applications work. It's maybe not called ETL, but essentially what we're doing is we're extracting a subset of data, we're transforming it, and then we're allowing the AI to see that data. And in best case scenario, it is only moderately out of date.
There's almost never a situation where it is actual live data, and yet that's what end users want. They want to be able to access real data in real time without having to fight with it. And as you said earlier as well, without having to talk to a wizard to rearrange the chairs to make it work.
Brad, I know that you've seen this. Yeah. Oh, yeah.
Yeah, it's interesting, isn't it, that we've been living in the wagon train for so long, it's hard to really just let them burn, but we should becauseIt's an outdated idea, and as Jitender talked about there, the notion of having these data silos, which we've been dealing with for decades, is not something to champion. It is something to do away with. But we don't do away with them by centralizing.
We do away with them by surfacing the knowledge about what's in the silos, and that is what we're talking about with the semantic layer in particular here. And what that does is it enables you as a company to really do what BI's been trying to do forever, which was to democratize access to the insights sitting in business. If you can flatten your data estate and make it so that whether I'm in AR systems or I'm in accounting, I'm sorry, well, that would be accounting.
Or if I'm in HR, or I'm just looking at a chat interface in the executive boardroom, that I should be able to see what's going on across the enterprise, and I should be able to ask whatever questions I want without having to sort of babysit, or as you were talking about early on, Jitender, do a two-week POC to sort of set up a brittle ETL pipeline to get my answer. Those have gone away. That survey I was talking about, when we asked data professionals if they were still working in Syntex, or if their jobs had changed to be more business-oriented or business-facing because of agentic AI, 75% of them said, "Yep, already there.
" So it's here now. The wagons are burning, and I'm happy about that. Yeah.
I could not agree more. I love the definition, Stephen, you gave about ETL. It's a new ETL.
It's not moving data from one repository to another repository. It's actually moving data and turning it into context for these agentic applications. I think that's where the power is.
Moving data doesn't truly add value. It just copies data. It multiplies risks.
It explodes governance nightmare. Instead, if I were to now truly democratize the access, I should just allow access to all of this data through agents, through a very rich context layer, and let agents allow the business users to be able to kind of overcome all the data silos that have existed and address the business needs they have. And when we say allow these agents to have unrestricted, unlimited access to all your data, it doesn't mean that these agents are not governed at all.
As a matter of fact- It's the opposite ... yeah. The governance rules that were applied to tables in SQL world, coming out of any data source, they get applied to agents as well.
In fact, at Starburst, we have a brand new control plane just for AI. In that control plane, just like in data plane, data users have governance rules based on maybe the departments they come from, maybe accounting or finance or marketing or HR, and they have access to certain data sets. Same set of access policies get applied to agents as well, as agents are representing users.
So role-based access controls, Brad, also get applied to agents as well. Yeah, whether it's policy, asset, or any form of control or governance that we have for people is being applied to agents, and it's really curious to me to watch how quickly the industry is pivoting toward this sort of mirror approach to all of our software, where it reflects both the needs of a user and the needs of AI as an agentic system to access them. That's one of the reasons why MCP has played such a huge role.
And people think, oh well, it's just something Anthropic made up to bring data to agents, and it's really a lot more than that. It is an actual sort of layer of abstraction. As a market, we're always looking for these, are we not, to enable our software to not deal with the underlying piping of, well, what was that API call?
I was trying to get to Qlik or Tableau, and I couldn't get to them because they changed something upstream and now my pipeline's broken, and my dashboard isn't working, or I don't have the data that I thought I had. It does away with a lot of that, and I'm very pleased to see how it's being applied specifically for analytics here. Yeah.
It is a protocol to enable the sort of semantics that we need in order to move forward. And I am actually excited about MCP even beyond the world of AI because we have needed some way of specifying context, as Jitender was saying, for a long time, and now we have it. Do you agree?
Absolutely. Before MCP, the data democratization has been going for almost 15 years. We have seen what data democratization truly meant.
In 2010, it was data engineering, and then data analyst, and then data scientist, and then came the machine learning apps, and machine learning scientists then further democratized the access to data and insights. But what has been lacking is theAccess to all of these data and tools and context for business users. Ultimately, business users always relied on some expert.
They know that data exists. They know where that data comes from. They most likely have an understanding of the data dictionary.
They know what questions to ask, but they always relied on someone. And the expertise required, sometimes if you need a more deeper understanding of the data, like if you're trying to understand campaign effectiveness, you have to rely on a marketing data scientist. But now with MCP tools, a lot of the data has been opened up, and again, in a very governed way.
Governance does not get compromised when you put an MCP interface in front of all your data and tools. But now, sitting inside any agentic application like Claude and use Starburst MCP server, now all of a sudden I have access to all the data Starburst has access to in a very governed and trusted way. And I can actually start asking some very advanced and sophisticated questions.
I sit in Claude Desktop and sometimes I have no domain knowledge of marketing, but I will be curious about what kind of users have visited our websites, what have they done, where they've come from, what kind of pages they've clicked on. And then, okay, why don't you now query Salesforce data and see if anybody showed up there? It's truly unbelievable that how democratized the access have become to all the data sources through MCP.
That was not fathomable roughly 15 months ago when MCP- Yeah. Right ... was introduced.
This still required a lot of preparation from us. We were building agentic applications without MCP servers, but then MCP comes in. I was like, "Great.
" But now all of a sudden, so many other business users who have access to just Claude can now talk to all data in a very governed and trusted way. Yep, and you couple that with a nice semantic layer, and you get some assurance that the agents are going to know where to look and how to interpret and treat that data. And so you get that build-up of trust, because it is so hard for companies, especially right now, to sort of retain that institutional knowledge and domain expertise.
" And when we asked in January this year, it had doubled from last summer. So companies are really struggling with what does sales close and fiscal year mean? You're going to have those experts, but they're not sitting there everywhere ready to go.
If you can institutionalize some of that domain expertise and some of that knowledge about how a company works within that semantic layer, and then open it up with open standards like MCP to these agents, then you can start to really look at your company not so much as a collection of data warehouses, but instead just as an ambient knowledge base of value of what it means to be this company. That's a real sharp point. And I think Jitendra, this is another thing, of course, that Starburst is really keen on, is making sure that data teams and analytics professionals are able to really evolve their roles and better serve their customers, right?
Yes. So with these agentic capabilities in Starburst platform, a lot of our users are now starting to focus on other higher value projects. The amount of wasted energy has significantly gone down.
Going back to the previous example of chain of yelling, is that now all I need to do is actually point you to an MCP tool, and you could sit inside any agentic application, including Starburst agentic application, and you can start to interact with the data. Conversation analytics has truly become the main interface for many business users. Ad hoc exploration used to be for data scientists.
No more. I can work inside any agentic application, and I'm doing ad hoc exploration. And if I have any question, as you said earlier, I can actually use these skills at Starburst, and many of our customers are also now starting to build agent skills.
There are many ways to bring enterprise context. One is you curate enterprise context and you put it inside some kind of a relational data store. But you could also start to think about this enterprise context come to you in the forms of agent skills.
So Steve and Brad, I presume you have certain skills, and you have been thinking about sharing those skills with others in your own professional environment. And now with agent skills, we're all codifying our own knowledge, and then that codified knowledge is made available to agents. And as a result, I have a lot more headspace to be able to go after more complex problems as opposed to getting hung up on very simple and trivial either data access issues or some very simple dashboarding iterations.
So that's where I see that many of our users, they have been able tohelp their business users close the skills gap. And as a result, both business users and technical users have a lot more freedom to be able to go after complex business and engineering problems. I love what you said there, in particular about the skills gap and about the skills as the anthropic standard, which has been adopted widely now.
And, yeah, I always look to keep mine a secret because that's my IP. But it is something that we all should think about. And what this reminds me of is that if we can't think of our businesses the way we used to, what you just described is something that most companies would have been trying to do with an internal Wikipedia.
Remember those days? Exactly. Yes.
Yeah, and it was the same for the data estate when we were talking about building a data fabric, which was basically just trying to set up an API-based access to your data, or with data meshes, where we tried to empower the people who were those domain experts. And those objectives haven't gone away, but we don't have to think about them the same way because of the way that agentic and generative AI are changing that data estate. We don't have to try to strive for these top-down ERP scale transformations.
We really should be looking to surface the value of that data through means like we've been talking about, with a semantic layer, knowledge graphs, and other means. And once we do that, you could just think about your business in ways you had never thought about before, and that's truly exciting to me. Yeah.
Absolutely. And I think that that's the positive outcome of a lot of what we're working on here, whether it's AI applications, agentic AI, or data analytics, business intelligence dashboards, is ultimately we're all trying to move the business forward and help people make better decisions and have better insights. So thank you both so much for this incredible conversation.
It's been really enjoyable. I know the folks listening would love to continue this discussion as well. If they wanted to, Jitender, where can people find you and continue this conversation with you?
So I'm on LinkedIn, and I share a lot of pieces on agentic applications and MCP and open standards on LinkedIn. io and search for Ada. This is our new agentic platform built on open standards, and that basically is where you can continue to learn more about Starburst AI capabilities.
And for me, I'm also on LinkedIn, and like to parcel out skills and insights from our research that we do as I go. com as well. And as for me, I believe the day that this episode is being published, I will be presenting AI Field Day in San Jose.
com to learn more about that. ai website as well as, of course, the Tech Field Day channel on YouTube to see recordings of those presentations. And I think if you enjoyed this conversation, you'll enjoy that one as well.
So thank you for listening to "Utilizing AI" today. If you enjoyed this discussion, please do subscribe on YouTube or in your favorite podcast application, and consider giving us a rating and a review. This podcast is brought to you by the analysts and experts at The Futurum Group, where insights meet AI.
And this special episode was brought to you by Starburst. ai, the "Utilizing AI" YouTube channel, or the Techstrong TV app. Thanks for listening, and we'll see you next week.