Cybersecurity with ChatGPT and Big Data – Michael Rinehart, Securiti.ai
Michael Rinehart, vice president of artificial intelligence (AI) for Securiti.ai, explains how cybersecurity teams will be able to leverage AI platforms such as ChatGPT in combination with Big Data to even the odds when it comes to thwarting cyberattacks.
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're talking with Michael Reinhardt who is vice president of AI or security and we're gonna be talking about how AI will be applied to data security because as everybody's heard the bad guys are using AI so why not the good guys the question is how and when hey Michael welcome the show. Oh, thank you. Thank you very much for having me.
What can people realistically expect from AI for the purposes of good because there's been a lot of hype around the subject for several years now, I would argue but not everybody seems to be using it or using it equally. Well, some people are downright skeptical. So what is the reality of AI as it applies to data security these days?
That's a great question. These security is a complex field the hackers continue to adapt to defense strategies. They're also employing AI like you say as a means to exploit new defenses in security and specifically, you know, our area in security is in data protection AI has the ability to give visibility into critical data that a company may have on its data systems and the ability to discover that information and Discover it with high accuracy is important and important and important first step to then tackling or at least being able to manage that data so that if and when A hacker or the bad guy gets access to your data systems.
They're not able to access that information. They're not able to subject it to ransomware. They're not able to expose it to a the internet and cause your company reputational harm.
And so at least for the area of security we're in AI for us is a powerful tool that allows us to do an effective job discovering sensitive information and making that accessible to our customers so they can take that important first step. What is our problem with visibility into Data because you would think that we created the data and we put it somewhere that we might have some idea where it actually is but the turns out maybe that's not the case. So, where is this kind of gap between the volumes of data that we're creating and our ability to keep track of where it is.
Another good really good question. So for very small organizations or for individuals me maintaining your own data is really not too big a deal. You've let's say one data system or just your hard drive you put files on you know, where they are, you know at a later time, but we're talking about large companies and you have companies with multiple organizations many divisions.
They're under time pressure. They're trying to get things done. It's easy for data to kind of proliferate into new areas.
Sometimes they copy it to an S3 bucket. They're gonna use it for a part of a data pipeline or for some machine learning experiments. They pull down to a hard drive safe data scientists or data and data analytics team will information down to their local computer to do some some analytics.
They forget about it later that can happen. HR has access to employee information. They may store it in spreadsheets or collect PDFs that people filled out.
Upload to Google Drive, and sometimes you just kind of forget about it. And this information gets copied it gets moved and eventually continues to multiply as new data systems are brought online companies or individuals might move information to those new data systems for getting that they had them in older ones and database tables. I mean data sales are rich in information and database tables get copied frequently.
Sometimes you create a new table from an existing one because you're doing some analysis and they forget to drop the table later. So data spread is is something that occurs very frequently and it's it's easy for it to fall into the cracks and and that's where that visibility that the ability to discover. It really comes in.
How am I discovering that because years ago? I think people were talking about metadata and catalogs and that's an interesting way of classifying things. But can I just throw algorithms and all this stuff and the algorithms will tell me what's where?
I mean through a degree the algorithms can do a really good job. So if you're working with thousands of databases and each of them contained many tables algorithm scale and I think that's really the primary advantage. If you have an act an algorithm with good accuracy, it detects the type of data that's of interest to you.
You have the ability of customizing what it's capable of well now you've a tool that those are pretty good job going out to all of your databases all of those tables and saying hey, here's what I think is in the look. Look you have a customer table here at customer table contains Social Security numbers credit cards. And maybe that wasn't part of the catalog no one it was just a temporary table that had been created and that can happen thousands of times S3 buckets people download information there out we know it if you have connectors, they're able to connect to these different Data Systems algorithms can then classify that data readily and give you visibility into it and from there you can take action.
what kind of actions what kind of policies can I apply can I wipe out that data that I know shouldn't be sitting on somebody's drive or in the cloud or in some insecure bucket or am I just kind of Yeah, making a copy of it somewhere just in case that data gets encrypted by somebody who shouldn't be encrypting my data. If let's say the company has a good ticketing tool what they can do is establish policies around data, they find and it may be Data Systems that have concerns on them. And then from there issue alerts and those alerts can then be fed into their their workflows and they can take action from there using existing company processes.
So as part of that therefore I can kind of decide my level of response right? I can give people a gentle nudge that says hey, you know, you got this thing going on to I could you know, maybe execute some more. yeah, you know comprehensive and more exact policies that just wipe that date out or how far can I go?
You could probably go pretty far. I actually there's there's an interesting twist to this and I probably should have mentioned this. Let's say that you discover data on a data system.
You're not crazy about and in the course of examining it you realize well. Hey you I I sure I can delete this but how do I prevent this from occurring in the future? You may do some root analysis and discover that it's a data analytics team or data science team that had pulled production data for the purposes of Performing analysis exploratory analysis.
And so then the question is, okay. Well instead of being reactive can also be proactive. Can we use these alerts to inform us on future actions and one interesting tool for this actually that might come up in the question of how far can we go is tools like even like synthetic data.
So if a company realizes hey, I know I have this important data system through visibility. I've discovered that a certain teams need to make use of production data, but maybe they're not controlling it particularly. Well we can do is make available synthetic versions of that data actually using AI algorithms and then after they complete doing analysis analysis on that synthetic data, that's synthetic data is is safe for them to use they can then deploy code in production where they have access where both access to this production systems to do the real job.
But the team will never have to access the real data in such a way that can cause an increased data exposure. So so maybe that's one way you can really carry these things, you know, pretty far and even leverage AI in the process of doing so Who's in charge of data management and security these days because in my experience it was always kind of like business units created the data and it people stored it. But from their perspective all data was relatively equal and they didn't really have much perspective on what date it needed to be treated differently and and it just fell between the cracks.
It was like it's kind of like watching, you know, the opposition hit a single every day between the second baseman and the center field because nobody knows who's job it is to get it so we gotten any better. I mean we have Chief data officers, where are we? You know an interesting tool that I think helps to.
Answer that question. So, you know every company, you know has it has its own strategy for organization but a good data catalog can go a long way to closing the Gap that you've mentioned who's responsibility is it and so let's say you had the data system owner. They have data that might be of interest to the organization.
They know with the sensitivity of that of that data is that data owner can be it people who spun up the data system a data catalog provides a convenient way for these groups to manage access to those Data Systems and even manage visibility to those Data Systems by other groups. And so let's say data analytics team data engineering team needs access to production database that has customer information because they want to do explore touring work, you know as part of testing a pipeline they could Requested the data catalog. The owner can say well, hey, you know you've requested this data.
I can give you a masked version. I can give you a synthetic version or maybe you know, we can provide access to some other way, but a tool like that probably helps a company regardless of how it does that organization to facilitate that type of access and perhaps make it easier. Very few organizations that I know would get a Good Housekeeping seal of approval for the way that they manage data.
So what's your best advice to folks to kind of bring some order to what has been arguably as somewhat chaotic process for the better part of two or three decades? When you say two or three decades, I think that really puts in perspective. I used to work at it a number of years ago.
I would say there there's two sides to that point. There's the proactive side and there's the reactive side on the reactive side visibility is is first tools that allow you to get access to all of your data systems have visibility into their permission systems gaining insights into whether or not they're over privileged and then take actions on what they discover. I would say that's number one being able to know what's out there and to take action on it to decrease the chances of the daily exposure.
The second side of that is the proactive side be doing a making it easier for teams to manage access making it easier teams to share data. I mean data is the new oil. We we have to share this data within companies or many companies do have to share data derive out to derive value from it.
And on that proactive side you being able to use tools to mask that information before they share it or to generate synthetic data or even differentially private synthetic data. And then this way they know that hey, you know on the reactive side and cleaning things up on the proactive side. I'm reducing that data proliferation and I think that can go a long way to addressing.
The two to three decade problem that yeah you you've rightly brought up. Do you think also in that two to three decades we were just far too focused on network security and we were playing, you know castles and moats is kind of our strategy for defending when you know the roof to the castle is wide open. Then the data was going out over the top of the wall.
So do we need to kind of rethink our approach and maybe focus more efforts specifically on data security versus the net worth the perimeter and whatever else we got going on. I I think so it you know, it may be the case now that data so abundant and it's so rich so valuable the tools for analyzing have become so import a so effective that companies now have a lot of it. There's a need to spread it around and it is a it's it's a new source of risk and so like you say networking security has been you know, improving over a long time.
Yeah. I think this is this is the new area of attack because they're probably just hasn't been the program of data that we've seen in Prior decades. All right, folks, you heard it here not all data is of equal value, but you should figure out which data actually does matter and then proceed with.
Zealous security of that data because that's usually the thing that's going to get you in the most amount of trouble. Hey, Michael. Thanks being on the show.
Thank you so much for having me. All right and back to you guys in the studio.