AI Leadership Insights: Maintaining Data Competency with Nitesh Bansal
In this Techstrong.ai leadership interview, R Systems CEO Nitesh Bansal explains why attaining and maintaining data competency has now become critical to ensure success in the age of artificial intelligence (AI).
Transcript
Hello, and welcome to the latest edition of the Techstrong AI video series. I'm your host, Mike Bazar. Today we're talking with Natasha Al, who's CEO for our systems.
And we're talking about the greater appreciation there is nowadays, or should be for data lineage and data competency in the age of ai. Neth, welcome to the show. Thank you, Mike.
Pleasure to be here. What's your assessment of what's happening here with ai? I feel like a lot of what we're doing out there is, um, embracing AI enthusiastically without understanding exactly how these models were created, and then we're being surprised when the data doesn't quite align and the output winds up being something we didn't anticipate.
Um, is, is that kinda where we are at the moment, and how do we kind of move beyond that? You know, I would say that's a fairly apt summary of where we seem to be. However, you know, I think we need to look, uh, beyond what's happened.
Look, with any technology wave, it is very clear that there's an initial exuberance to try things out, to look at what's the next big thing, next shiny object that we can go towards. And clearly, you know, there are those, uh, tech forward companies who will obviously want to try take the lead and get those things in place. So this has been the case with iot in the past.
This has been case with cloud. In the past, this has been case with any new technology wave that has come in, in ai. What is different is people have a lot more latitude and a lot more space to try various things.
Now the models have been put out there, um, by some very large tech companies and, and, and open source models out there to try things out. There is a promise of the outcome, the, the insight, the additional revenue stream that can be created. And when people implement those, that's when the real devil in the detail, which is about your own data that you're consuming to build that model, to build that insight that comes out.
And that's why I continue to, you know, whenever talking to my clients, I continue to, uh, advocate that the power of AI also lies in the power of the data that you have, the power of the data that you feed it with, and how clean, how curated, how well managed that data itself is. So I wouldn't really call it that we are seeing something like we've tried, but it hasn't really delivered. I think we have tried to realize that to make it deliver, we actually need the data behind to get the real power outfit.
Is this creating some seminal moment here in the history of data management? 'cause I feel like, um, I've been kind of banging the drum about data management being a mess for more decades than I care to admit. And is this all coming home to roost now?
We've been there for sure. I mean, um, old times of master data management, getting our customer master and uh, and product master and all those things in line. We've been there.
But I think the context has shifted. And the context was what was largely an enterprise's own internal context of getting their processes in line, getting the, uh, everyday MIS or or ability to act and decide put those things together, has now suddenly gone outside the organization. Because where AI has the largest promise is in developing the customer intimacy, in creating new, uh, cons, customer experiences in coming out with products, with your ability to monetize them in a micro monetization manner, create new revenue models.
And what that does is the, uh, the amplitude of, of, of, of impact has suddenly gone up. Right? So is it a seminal moment?
Well, like you rightly said, the importance of data has always been there. Uh, the realization has always been there, but the value impact today is higher than ever. If we want to leverage AI to really differentiate our company, our offerings and our services to our clients, and we want to leverage it to the best of its capabilities, then the realization that we need to leverage our data well, and to, for that we need to have good data behind has become really front and center In a lot of ways.
We've been talking about things like data lakes even before the rise of ai, but now that we kinda seem to have data lakes that work, and maybe they're not data swamps, but how do I find the data that I need out of that to drive the, to the right AI model? 'cause it seems like the challenge here is I gotta get the right data at the right place at the right time. And that's a lot easier said than done.
Absolutely. Absolutely. Because, uh, you know, having data lakes in order to pull out some of the organizational data and to be able to respond to some queries is probably easier.
But when we look at, um, the models that are driving revenues, models that are driving customer experience, uh, a recent study, I think it was, uh, probably a McKinsey study, um, or something, two factors that came out. Number one, the use of AI in driving customer experiences has almost doubled over the last two years, right? So that means there is bigger imperative in driving some of those, uh, differentiation using ai.
Second, the, the, uh, data point says that in order to build a proper insightful a AI engine, you're probably leveraging more than eight data sources, right? And that's where this whole element of have I put a data lake together and is it comprehensive enough actually comes into question, is it really a data lake that can feed that model? Are we going to be, um, requiring external data sources that need to be plugged in?
How are those pipelines getting curated? How are those pipelines getting tested? Some elements that were not so important earlier, which were, uh, uh, relating to how that data has been tested for bias or accuracy, all of those have become to the forefront, right?
And that's why I think, uh, the whole imperative of data first approach on platform designing or data first approach to look at what AI models we are putting forward has become very important. We hear a lot about the general shortage of AI expertise, but when I look into it, it seems like we even have less available skills for data management. So how do we kind of get at all this in a way that enables us to build these models?
'cause I kind of feel like this issue is the, the, the thing that's holding back all these applications from being built in the first place. Skills wise, I think the data skills have, uh, probably always been there. Uh, the focus of data engineers always trying to, you know, then move towards becoming data scientists or analysts has driven them towards roles that have, that have promised them, uh, you know, that kind of exposure we have to pull back and create a new set of, uh, uh, data engineers who are more focused and more worse on data quality, on data management lineage, uh, and and more importantly, testing because testing that data for keeping it, again, like I said earlier, free of bias of understanding the, uh, the variety of sources that that data has come from.
Is it representative enough? Is it wide enough? And in addition to that, the ability to integrate external data sources has to be taken into account to build this, uh, from a talent perspective.
Clearly, uh, the focus is there and, and the, uh, the talent is being prepared to, uh, to do that within our own organization. For example, uh, we take a lot of, uh, we put a lot of focus on training people on these disciplines. And to do that, we have to also train them on general AI literacy.
How is data, this data going to be used? Because people who are preparing or curating data, people who are testing those pipelines also need to have an understanding of how is this this data going to be used to finally deliver the outcome that we are, we are planning to use it for that enables them to actually test it better, to prepare it better, and to create those pipelines better. When we wind up using AI to essentially automate the process of training the AI models using AI agents that are specifically designed for data engineering tasks, We would, I don't think we are, um, um, completely there yet because, uh, we do use AI for, uh, doing certain amount of testing.
But even for that, the AI model needs to be trained on how to differentiate between, uh, the right kind of data sets that get through versus not. And the question of, again, you know, bias and completeness com continue to spring up over there. Um, I'll give you a an an example.
Uh, working on a platform where we are using AI to develop dynamic content, and this is dynamic content marketing content being developed based on the profile and the behavior of the user that that platform is interacting with. So there is a library of contents that content needs to get stitched up in real time to create a very specific personalized marketing content as an output. And for doing that data training, the combination of using an AI model to train the data or, or to rather curate the data and having human in the loop both become equally important, right?
So the human in the loop part is something that we cannot really take our focus away from. While the model will become better and better at curating the data, because it's coming from a library that has already been built, but continuously looking at the human in the loop factors of was it right for the audience was did it really get the nuances of what the user was looking for and what other inputs need to be provided back so that next time in the model does this curation and select the right data, uh, data points or data sources to create that content. It is actually corrected for whatever was noticed by the human in the loop.
So does that mean we need something that feels like a feedback loop? 'cause I think models, whether on purpose or inadvertently we'll get exposed to additional data after they've been deployed, and then we might start to see some drift. Um, that risk is positively, uh, possible.
And, uh, we all need to embrace the data cleansing, data curation as an ongoing priority, not as a one-time job. Because you're absolutely right. As we consume data to serve our customers, we actually generate even more insight and even more data that gets in.
And if that is not subjected to the same rigor with which the original dataset was prepared, it will show adrift and hence the human in the loop. The feedback process, the, the whole revisiting the entire, uh, what I would call the, uh, the filter set on which the data was filtered or curated, that itself needs to be revisited from time to time. So this whole, uh, discipline of establishing the lineage, curating the data, cleansing it, uh, staging it for the right purpose is not a one-time activity.
One, it is a continuous activity. And second, it is also an activity that needs to be revisited and the assumptions need to be revalidated periodically. Speaking of revisiting, do we need to maybe take a look at the way it teams are organized in the age of ai?
'cause I feel like there's data scientists and data engineers, but then I gotta add in application developers and security folks. And, um, by the time I get the, um, not to mention DevOps and engineers and a few other folks, but by the time I build and deploy some sort of AI application, it takes a village. So does the village need to be reorganized, Uh, if it has not already been done?
So my, yes. I mean, we work with leading products and platform companies in helping them design and develop their platforms, uh, most of which are SaaS platforms, which are, which are basically plugged in with ongoing understanding of every interaction. That interaction feeds back into creating their customer profiles, creating the personalization.
And there is absolutely no way we can think of a team that is delivering that without having a combination of those skills and things, which are, uh, which probably could have been done after, does done as an afterthought in the past, which was like, okay, you know, I'll do a security testing towards the end of this. No, these days we are working on this continuous loop of, of doing integrations and deployments. We are continuously working on data getting added and, and new, uh, insights coming out.
That security has to be an integral part of the team, right? So whether it is, like you mentioned the whole DevOps side of things, absolutely the front end backend engineers to get the right backend data in place, the right experience in place, the developers which are developing it, but also the quality engineering, the security engineering aspects of it have to go hand in hand. One area that I also think is overlooked here is this notion of data observability.
And while the regulations may be still in the realm of the wild while west, I think customers are demanding to know where the model was trained and how it was trained, and they wanna see and they wanna have that validated. So do we need to, uh, invest more in observability around these models and where the data came from to satisfy that transparency requirement? Absolutely.
I don't think we can deny that. Now, there are two ways of, uh, obviously approaching it. One is, uh, where we look at platforms or or services being offered, uh, by companies where they're relying largely on their own historical client data or, or company's own data, which is where developing those observability, uh, capabilities inside being able to understand the lineage of that data and, and how it was, how long back does it go?
How was it, how was it assembled, what does it contain, et cetera, become important. But there's also the factor of any third party data, anything that is coming off from a public source versus a data which has been purchased or, or subscribed to. And getting that, uh, observability information in is also very important.
I perhaps foresee that, uh, we will probably see a lot more of, uh, data exchanges coming up where, uh, companies or, or third parties who will invest in building up that curated data sets with a proper lineage and observability behind it. We'll be able to, uh, for profit or as a, as a stream, be able to supply that data to various, uh, parties who can use that data in various combinations to create their own unique experiences for their customers and be able to work with that. And some of that we should start seeing happening or, or, or, or beginning to happen.
So what is that one thing you see organizations doing that just makes you shake your head a little bit and go, folks, we need to be smarter about this. So we see a, uh, you know, if you look at generative AI use cases and the initial rise of everybody trying to put a chat bot in place, uh, in order to, you know, handle some of their customer services or, or some of the other things, naturally, those are some of the most common use cases. One of the, one of the capabilities that generative AI has, has suddenly made very feasible because of its, you know, multiple tokens.
It can produce, it can, it can process the whole, uh, natural language processing and, and all of those things. But going back to the same problem of what data this chat bot is feeding off, sometimes those responses can be fairly off, right? And, uh, with a, with a, with a direct customer interface, uh, you know, sometimes it makes you wonder, well, was there a rush to do this?
Uh, why it has not been curated, but ultimately I still, you know, look at it from a positive angle because, you know, adoption obviously brings maturity, right? So as adoption increases, as companies continue to look at trying those things out and putting them in place, maturity will follow and will happen as a reverse is actually worse. Hey, folks, you're hearing it here.
Even in the advanced AI era, there's some comfort to be taken in the fact that, well, when it comes to data, garbage in is still garbage out. Hey, Esh, thanks for being on the show. Thank you, Mike.
Really appreciate it. And thank you all for watching the latest episode of the Techstrong AI video series. You can catch this episode and others on our website when you invite you to check them all out.
Until then, we'll see you next time.