Managing Dark Data Risks with Soniya Bopache
In this Techstrong.ai interview, Soniya Bopache, vice president and general manager of data compliance at Veritas Technologies, explains how unstructured, untagged and unused data, otherwise known as dark data, can, when not effectively managed, lead to sensitive data being inadvertently exposed to a large language model (LLM).
Transcript
Hello and welcome to the latest edition of the Techstrong AI Video series Army host Mike Bera. Today we're talking with Sonya Bopa, who's vice president, general manager for data compliance and governance at Veritas. And we're talking about dark data, which turns out to, there's quite a lot of it.
Sonya, welcome to show. Hi, Mike. Thanks for having me.
I think people don't understand how much dark data there is, or for that matter exactly what makes up dark data. And so from your experience, um, I know we're all trying to get our data act together for the age of ai, but what is all this dark data out there? Where does it come from and, and what kind of issues is it causing?
Sure. Um, so think of the, uh, dark data in the context of old attic at home. You store many things there, some of which, uh, you might have forgotten about.
That data is like those forgotten items in the attic. It's stored away and not active used, but it still takes up a space and could be valuable if properly managed and analyzed. So that data often refers to vast amount of unstructured un unpacked and unused data that gets accumulated over time, but fin to properly managed or processed or analyzed into the system.
This data often includes emails, documents, log files, and other forms of information that are stored but not actively used. And oftentimes it gets overlooked because it's not always immediately visible or understood, as they say out of mind is out of sight. So that applies to dark data.
Mm-Hmm. And as data grows, um, exponentially, it becomes increasingly difficult to monitor and manage every piece of information leading to significant amount of data being ignored or forgotten. So now if you think in the context of ai, AI also relies heavily on the large volume of data, and that also includes star data to train and make predictions.
How do I sort through all that dark data to find the data that's relevant for my project or whatever it is I'm thinking about, uh, showing in AI one? Yeah. Um, and that's a very good question because as we all know in the AI world, and as more and more businesses are adopting to lot of large language models and in general the generative ai uh, workflow building, it's very important to put a greater emphasis on understanding and managing the dark data.
And that can be done by various means in the sense that really being proactive by implementing the Rob robust data monitoring system to track the data as it is getting created, that also includes that how you classify the data, how do you manage its sensitivity, ensuring that the policies are in place to handle it properly. Regular audits, data governance practices also help in understanding and controlling the data data. It then potentially it reduces the risk by not having any unintended problems.
So it's all about having that robust data governance framework ingested into the overall ecosystem makes it having a better control on the data, uh, on the data data. Who's going back to kind of sort through all of this? Is it the AI team or is it more of the data governance and data compliance teams?
How do they collaborate to make all that happen? Yeah, so oftentimes we feel that, okay, if there is a one team in the organization who has an ownership and complete responsibility of maintaining the sanity and performing the hygiene on the data, but in my opinion, it send responsibility on every individual in the organization purely because the data is getting produced almost every day in an exponential manner coming from various data sources. Now in such situations, every individual who is producing that data also need to make sure that that what data is getting processed through the system, how are we analyzing the data, what are the patterns are EMA emerging are are very important to have that visibility from that data.
So I would say that the owners is definitely on the compliance team, the infrastructure team, but also an equal responsibility of an individual to make sure that the dark data is not accumulating every day. Are there things that we should be doing now to limit the amount of dark data that's being generated from here on out? I mean, should all data be in the light so to speak?
Or will dark data kind of always be with us? It's just a question of making it easier to find it when we need it. Uh, I think that dark data will be with us.
Um, it is going to be very interesting, uh, to see how the industry is going to evolve with more and more tools are coming out, but I feel that that data is going to stay and that's a fact. But what are the best practices that we can really apply in order to really, um, adopt to data minimization approach is the key here. Like for example, if you see that some of the best practices in order to managing the data data by having a very clear and concise, um, uh, data governance framework, like for example, developing, um, advanced tools to perform data, data assessments, um, then monitoring and classifying the data when it is getting processed or indexed, and regularly updating the data policies to align with the regulatory changes, tracking the data as it is getting created and applying controls, this will help us to minimize and having a control over that data.
That way we can minimize the footprint of that data and potentially help in really refining that data and then ingesting that to your processing as well as to the AI systems to train it. But I feel that it is going to stay, it's just that how we handle it carefully and minimize that is going to be the key. What is your assessment of the level of sophistication of data governance that's out there?
'cause a lot of organizations I talk to, they haven't always been especially good at managing data. They kind have a lot of data that's duplicate, some of it's erroneous. Um, can I go back in and kind of impose a framework on that to govern it all?
Or do I kind of start from at this moment, from now on, we'll get it right. Yeah. So I think sure.
Um, proactivity is crucial and the time is now for these type of assessments, um, for organizations that are new. Um, sure it would be a great step to have this mindset from a day one. But for the established organization, what would be very important is to have that mindset fostering the culture of making sure that we do have the robust data monitoring systems to track the data as it is getting created.
I would say that this is not the project which happens just one time. You assign your tiger team to it and then be done with it. It's truly an effort.
It's an ongoing effort where we need to make sure that by having this entire data management lifecycle and having really a proactive approach towards it, will help organization to really create that culture of handling and maintaining this data. As I said, um, there is, there is no, um, uh, no way you can just wait and like really wait for the time, create your team, do one-time execution. No, it's all about creating and fostering that culture and making sure that that data assessment, the data profiling, having the data management lifecycle is the key to make sure that we have a complete handle and control over, uh, the overall, uh, data footprint.
Will we ultimately need AI models to help us govern all the data that we have out there so we can build AI most? Yeah. So all the data, um, that is getting fed to the, uh, generative AI as well as to the, um, more and more s these are all very data hungry.
As more businesses allow to AI technologies, the role of that data will become increasingly significant. Uh, we need to place greater emphasis on understanding and managing the, that data to ensure that AI systems are trained on accurate, relevant, and secure data. The potential for that data to introduce biases or security risk in AI will drive the development of more sophisticated data management tools and practices.
So as it comes with a certain risk, but having the robust framework, having a control over that will help us produce a lot more efficient and better outcomes from the, um, ai uh, tools. Um, as, as we all know, um, technology is definitely going through a very, um, very, um, exciting phase of AI evolving data potentially growing, and a lot of industry regulations that are taking the birth. I think the role of AI in decision making expands the need of a comprehensive strategy to monitor, classify, and understand our data becomes increasingly critical.
Uh, this AP approach not only just safeguards AI integrity, but also ensures compliance, fairness, transparency, and security, all of which are essential for the sustainable and ethical development of AI technologies. It seems to me when I go look at anything involving AI regulations, there's a, a stress providence and understanding where the data came from and how it was used. So do you think that these AI regulations as they become more stringent and actually come, maybe the law of land will require organizations to have, you know, a much more, uh, elaborate approach to data governance and data management?
Yeah. So as, as we all know, regulations vary significantly across industries and regions, but aligning the most stringent standards is a smart strategy. It's important to stay informed about relevant regulations and adopt best practices that exceeds minimum requirements.
Um, I feel that this approach not only helps in ensuring compliance, but also makes it easier to adopt to regulatory changes, AI related changes, industry regulations, changes in safeguarding against potential and legal issues like, for example, industry specific regulations such as GDPR in the EU or uh, HIPAA in the US or even California Act of privacy set stringent requirements. These regulations often mandate the organization implement specific controls to manage and protect all types of data, including that data failure to comply can result in hef fines, legal actions and impacting customer trust, making compliance really a key aspect of a dark data management. So I think it all boils down to really making sure that how we truly stay ahead by aligning with the most stringent standard.
And I, I feel that it, it is truly a, a game changer and it's a smart strategy. It may sound daunting that okay, there is a lot of things that we need to do and develop, but a step is definitely in the right direction. Um, As we kind of continue to revolve, well, the way that our IT teams are kind of structured needs to change.
And I asked the question because we have a lot of silos and people, businesses generally, or at least the business people, uh, create the data and then they just kind of store it with the app and they don't really think much about where that data is gonna reside. And a lot of times the IT people are, um, you know, they're managing the data, but they don't distinguish between the data. It's all data to them.
So, um, do we need to kinda have some sort of, um, reassessment of the way we manage it in an age where the data itself, um, the type, how it's stored, where it's stored matters more than et Yes, the detailed chain of custody is truly the key for that. As you said, um, a lot of uh, different departments, groups in the organizations are contributing into it, but ultimately it boils down to having that really defined data governance framework where the advanced tool to perform that data assessments, monitoring and classifying the data from where it is getting generated, who is owning it, really having a good data taxonomy will definitely, uh, change the game here. Um, because I feel that whenever you write the data classification policies, it's no longer just a job of a compliance officer or for that matter infrastructure, um, uh, leaders.
IT is now truly evolving and making it far more effective by involving cross-functional teams in policy department because then you can write more robust data classification policies, create more transparency around that, classifying the data, and then ingesting that into the system so that you can have a complete view of your data set. I think that's definitely is, is where the shift is required. Significantly, We have seen simultaneously the rise of the data engineer, um, but I feel like there may be not enough of those folks out there.
So, um, is part of the goal here to figure out how to democratize data engineering or we just need more of, Uh, as I said, I think it's just the onus on each and every individual in the organization. I think having that culture, fostering that culture to make sure that I am being very, very conscious about how the data is getting generated, stored, what is the pattern, what are the data sets and the data sources as associated with it, and having a complete control of that is very important. It's no longer just the data engineer who will be like performing the data scrubbing and all of that.
I think there are a lot more tools that are out there, but it's truly an onus on every individual in the organization to make sure that they are contributing in a various ways to keep their data management lifecycle framework intact. Alright, folks, you heard it here. Hey, it all starts with the data and maybe we should all think like data engineers because if we do well, we'll have a better AI outcome.
Heys, thanks for being on the show. Thank you, Mike. All right.
And thank you all for watching the latest episode of Textron AI video series. You can find this episode and on our website and we invite you to check them all out. Until then, we'll see you.