Data Management Challenges with Stephen McNulty at OpenText World 2024
Data management challenges are on the rise as unstructured data outpaces structured data. Organizations face difficulties in securing sensitive information, especially with the added complexity of AI. Security risks from unencrypted data and increasing regulations highlight the need for better practices. Understanding data ownership and privacy rights is crucial. Effective technology can analyze video data and monitor events, but best practices require clear problem definitions and validation of AI outputs.
Transcript
This is Textron tv. Hey everybody. We're back here at OpenText World and we're here in Las Vegas with Steven.
And we're gonna be talking about, well, what's going on with customers? Because one of the challenges that I hear is everybody is struggling with data, but you talk to a lot of customers more so than just about anybody in your role. Yeah.
Um, what's the common themes that you're hearing and what are the issues and challenges that all the customers kinda have in common? 'cause they don't talk to each other all that much. I think there's two major ones.
One is data and finding data. Uh, people have data in nice structured tables. That's the easy bit.
That's about 20% of a company's data. The rest of it is all over the place. It's in videos, it's in social media, it's in pictures, it's in, you know, whatever you can think of.
And that, that's a challenge as people try and build. We're all talking about AI to, as you try and build an AI model, if you can't find your data, how valuable is the ai? Uh, but the second part is securing that data as well.
When you think about an organization, you've gotta login to a certain amount of data, but you stick AI over the top of it. What secured the gaps for you? Potentially opening up by giving pe people access to an AI model that's accessing data that they're not allowed to see.
So, a few challenges, but a lot of opportunity. You're not quite sure what those models are doing with that data. 'cause if it's sensitive data, it might end up training an AI model and it pops out somewhere down the road.
Right. So They, they call that the Swiss cheese security gap. Right?
I mean, it's like, I've got an opening here, but it leads to an opening here. And before you know it, someone's got access to data. Like I said, they shouldn't have.
It's, it's unencrypted. It's unmasked the same stuff. They shouldn't be seen.
One of the issues that I hear about is that the amount of unstructured data keeps growing a lot faster than the structured data. Yeah. And I don't really have the mechanisms in place to manage it.
So like the problem just gets exponentially worse almost month by month. Yeah. Well, if you think about what unstructured data is, the, the, the, one of the best examples is video.
What's in a video file? It could be a facial image, it could be a a car number plate. There could be personal information.
I was talking to a customer just two weeks ago in Australia, government customer, and they've got the Freedom of Information Act there probably somewhere in the us. And you can go to a, an entity and say, I want you to give me, or tell me all the data you have on me. How can you do that?
If it's a picture, a video, you're in it. They don't know you're in it. Mm-Hmm.
The other challenge, I mean, it's hard enough to manage the data, but a lot of folks are also struggling with the, in the age of ai. I've gotta get the right data at the right place at the right time. Yeah.
And this is not easy with unstructured because it's largely unmanaged to begin with. Yeah. I I think there's a huge challenge all over the world with AI around the, the complexity and the size of potential models.
And everyone talks about, you know, chat GPT, and if you look at what that it's got the entire data of the internet that it's got access to, but organizations don't need all of that data for all the data within the four walls of their organization to make decisions. So what they've got to do is define what the business problem is and then define just the data they need or the types of data they need to get into it and build a model that is scalable, that doesn't cost too much, doesn't drive so much power, and doesn't need so much storage. And, and, and that is a challenge.
So finding the data only the day you need and getting it into that model. So you can actually get an ROI from the money you're spending on the particular model. I talked to some folks and they're basically hoarding all the data they can get and they think it's all gonna be valuable, but at some point I kind of choke on the cost of storing all that data.
Right. Well, yeah, I mean, I, again, talk, talk to a lot of customers, you, you'll find that a lot of them will have thousands of copy of the same file. Right.
So I've got Poppy, I send it to you, I mail it, I mail it to you, or I transfer it to you, send it to other people. So just looking at the, the, the amount of the same data that exists in an organization is huge. I mean, not only does that create a challenge of just simple, uh, complexity, but also cost as well.
And then you get version control. If I tweak my document just a little bit, but you haven't tweaked yours and then you send it to someone else, you've got a document that's 98% correct. Across 50, 60, a thousand copies in an organization.
What is the real copy? What is the copy that matters? There has to be one golden copy.
Where is it? How do you find it? I don't know.
We do, Well, let me ask you, what is the maturity of the organizations that you see when it comes to data management? 'cause at least in my experience, very few of them would get a good housekeeping seal of approval for the way that they manage data. Yeah.
In Asia where I live, I live in Singapore. Those developed countries with very mature organizations and developing countries with maybe not so quite so mature organizations in the developed countries, banks, and, and other types of organizations have got a fairly good handle on, on, on data. I say fairly good.
No one's perfect. A classic example would be if I, if you're in a bank and I email you a, uh, send you an email and I attach my driving license, my credit card, my passport, whatever, only, you know, you've got it. But the bank has rules around the what types of data they're allowed to ask for the types of data they're allowed to keep, how long they're allowed to keep things for when they should actually, or what data should be redacted.
How do they find that? If it's attached to an email I sent as a customer I send to you, and even in mature organizations, that is a big challenge. Mm-Hmm.
We don't all love regulations, but it seems like one of the outcomes of all these in increased regulations is I have to manage data better. I have to be better at data stereo. Are people changing their mindsets about the way they think about data?
Oh, very much so. Especially when it comes to things like personally identifiable information. The, the fines can be massive if you have that type of data and it's not secure, uh, if it's unencrypted.
Uh, we've all heard of many data leaks, uh, throughout organizations. Hackers or nefarious employees are taking data that they're not supposed to have. And that, that is a, an absolute challenge for organizations.
But you can only secure it if you know what it is and where it is. Speaking of security at the keynote, your CEO said that, but in two years, you're gonna mask all the data in the platform. Yes.
Is that, is that feasible? And what, what will it take to Achieve that? I'd once if our CEO said it?
Listen, I, I think data as, as more and more data gets opened up to any form of AI model, making sure it's only the people that, that are allowed to see the data can see the data. You know, you're leaving an organization, you've put thick in a thumb drive, you take away whatever you want to your next employer, obviously very illegal. Uh, but you, you need checks and balances to make sure that that absolutely cannot happen.
Not only is it giving your competitors an advantage, but you also get fined for not securing the data properly in the first place as well. Mm-Hmm. Have we kinda somehow or other inverted our thinking in the wrong way, or maybe putting carts before horses when it comes to security because we have all this data and then we spend a fortune collecting it and then we spend an additional five, 10, 15% more trying to secure it.
Yeah. Should it just be secured by design? And is that kind of where we're headed Now?
That's exactly where we're, we're headed. So I think, you know, just looking at OpenText, there's a couple of models. There's for data that already exists.
We've got technologies that can help you find it, identify it, and apply risk scores like GDPR or other data privacy, uh, type regulations to it so that you can actually understand the data you've got. That's number one. But secondly is data.
'cause it comes into your organization putting rules and regulations around and, and educating your employees. What type of information are they allowed to accept of what, how should, how they should store information. Uh, that, that, that's obviously for the, for the new information.
But that, that's very, very important as well. Who's in charge of this these days? And I'm asking this question because while we saw the rise of chief data officers for a while there, and then we saw, uh, digital CXOs and, you know, all kinds of people like managing data is, is all this coming full circle?
'cause there used to be this person called the CIO and maybe they're running this, Uh, listen, I think everybody's responsible. When you've seen the large data breaches over the last number of years, these are board level discussions now. So everyone's responsible, every organization should be, should be educating their employees on all aspects of data.
So it's not just resting on one person. Maybe one person writes the policy, but everyone's in, in should have full knowledgeable data privacy, secure to regulations and comply with it. Most end users, I know they can't keep track of all the rules and regs.
Yeah. And so they're download stuff and it's unintentional. But I can't help but wonder, are we also on a path towards not just using data to train AI models, but using AI to help us manage the data?
No. We, we, we, we, we absolutely would do that. Uh, using AI to again, understand what the data is and put rules and regulations as data's coming into the organization.
Uh, you, you would've heard on the, the keynote with Mark this morning, these agents, uh, where you've got, you know, something that's a very intelligent, very smart rules driven column by itself, making these decisions for you as well behind the scenes. 'cause you're right, no single person can manage every single aspect, every single thing they do. And you'll also get the bad employees, the ones that don't care about the rules and regulations.
So how do you actually monitor them and take the appropriate action out, shut them down, or, you know, send someone to someone nice to see them to make sure that it's not happening again. And escort them out the building in a, in the, in the right, the right manner. These days you might have to go visit their house 'cause it's probably on a spreadsheet in their laptop at home.
It, it Could very well be. And I think work from home, that has absolutely opened up data privacy issues. There are laptops, not just at home, but in Starbucks.
You know, you, you fortunately living in Singapore, I think it's one of the few countries in the world where you could actually leave a laptop open and walk up to your Starbucks counter in order your coffee and come back into a laptop will still be there. But that's not the same in every country in the world. So it, it, it's not just home, it's, you know, work from anywhere.
And how do you make sure that that's secure as well? So technologies that look at, like if I did leave my laptop open and someone ran away with it and I I, I was stupid enough not to secure or lock it before they ran away with it. Something in the background monitoring what this thief is doing, you know, they talked about bi billions of, of events per second.
Looking at those outlier events, understanding that that's not Steven McNulty's normal behavior. Something's not right with that laptop and shutting it down on this spot. So they can't actually do anything.
And we can obviously remotely like a laptop, we can remotely lock a laptop, but we can only do that if we know that something bad is happening. Mm-Hmm. Do you think that, um, as we evolve the way we think about data and the, and the management of it needs to change because historically it people, you know, they charge of storing the data and then the business units created the data and the IT people didn't have a lot of insight into the value of that data.
And which data might be more sensitive other than the fact that there might be some social security numbers in there. But do we need a just a better, um, data understanding, data hygiene, a more mature approach? I absolutely, I, I do think that most mature organizations typically have quite robust data management policies in place.
I mean, OpenText, we go through training every single year on what, what is data? What is personal identifiable information? What is private data?
What data should be encrypted, masked, et cetera. I think most organizations do it. I think that the challenge at the moment is that the, the, the, there's so much data being added.
We talk about social media, you know, companies are capturing social media feeds themselves now to look at, uh, consumer sentiment. You know, that's unstructured dating data sitting on a server somewhere. And again, just the absolute explosion of data that the world faces.
Yeah. I mean, in fact, one of the things that's always struck me is, is a lot of times you'll talk to execs and they'll start referring to our data. Yeah.
I'm not sure it's our data, right? Yeah. I think it's data that's on loan to us and we need to kind of treat it as such.
Yeah. I mean you've got a responsibility. Your data could be things that you create with that data, like a customer profile.
So you've got customer information that is their data. They've given you, you might create something with it, but, but you're responsible for all of that data. Some of it you might own, like your intellectual property, your patents, and, and some of it, like you said, is, is borrowed and, and people have got in, in some countries, uh, you're, you got the right this right to be forgotten.
Uh, I, um, if you've got some skeleton in the closet, you've got the right to say, you know what, I don't want that to be available anymore. How that gets done, I have no idea. 'cause once it's in the internet, it's in the internet, but they've got the right to be forgotten as well.
So, you know, it is not just the borrow of the data, but people can for it back. And you've got to be able to find it and give it to them. You know, I, I was driving, so I live in Singapore, but I was driving around Australia uh, just a few weeks ago.
And they've got cameras everywhere. They've got these AI cameras that can detect if you're holding a mobile phone, it's highly illegal in Australia to even touch your phone as you're driving. So they've got these cameras that can recognize someone if they're holding their cell phone.
But again, they're also capturing facial images as well. And you know, just rules and policies around what are you allowed to keep, who is it? And if I've got the right to understand or ask you from the freedom of information, like what do you have, Mr.
Department of Main roads, what information do you have on me? Uh, and I'd like you to remove it. You've got me to find me on video as well.
'cause I exist on video. Not me personally 'cause I wasn't using my mobile phone. Lots of other people.
How do we analyze all that video? Because, um, you know, I was talking to somebody once and they were telling me about how if they wanna analyze all the video the organization had created, it would probably take them, you know, a year of 20 people watching all this Stuff. I could believe it.
So I mean, the open text plug is, we do have technology that can look inside video files about two and a half thousand file formats and find anything we want. So we can recognize a facial image. We can recognize if a bag's been left alone on a train station platform, it has been left alone for a period of time, a minute, 30 seconds that there's something not right.
So we, we, we can do it, but you're right, if it's a lot of data, it's a lot to scan. You know, obviously the quicker you get started the better. But you can scan things coming in and even if you've, it takes you well to scan everything you've historically got, you can start scan things coming in quite quickly.
Do I need the index all that? Is that kind of the secret? Or how do I kind of know what data is what?
So our technology categorizes it, auto categorizes, it can auto you as a person, a dog, a cat, a bag, a bus, a car. It can auto categorize it, uh, without you even asking it what types of categories. There should be that obviously you can tweak that so that you can define your own categories, but it would be too much for a human.
There are classifications built in is I ingest the data, I get A block. Exactly. Exactly.
And, and apply a risk score to it as well. So what do you see in the customers you talk to? What are they doing well?
I mean, the ones who are at the forefront of this thing, what do you wish that others would do that the, the best in class of doing? So obviously innovation. We, I speak to lots of organization, department of defenses, uh, banks, et cetera.
They're all playing around with different AI models. More, more to automate than anything else. Uh, and the ones that are doing it well, I've got a very clearly defined business problem they want to solve.
And they're only putting into their model sufficient data to do that. So they, again, back to the Ella point, if they don't need a lot of power, they don't need a lot of storage and they're making fast and quick decisions and they're, there's some really cool things that they're, they're, you know, they're looking at doing it. You saw the demos earlier today where, you know, as an employee we can use it to answer RFPs.
The, the example is an 80 page RFP. We get RFPs hundreds and hundreds of pages that span so many dimensions of our company, uh, all the way from not just functionally what something does, but the security aspects of it as well. And we can, uh, very quickly summarize that, get questions answered.
What we don't do is just take that verbatim. And I, I would coach that to all organizations. Whatever answer you get out of an AI model, don't just take it verbatim.
We have heard stories on, on the, uh, you know, public media people that have done that and it's, the data has been wrong. Uh, you hear stories of people asking a question of chat GBT and then going back and asking the same question three minutes later and getting a different answer. So it's not full prep, it's not going to, uh, it'll make you smarter quicker, but it's not going to alleviate and remove humans from every step of the equation.
You've still got to validate that data, make sure that it's, uh, that you're not revealing something you shouldn't be doing. Make sure it's written accurately. Make sure the question has been understood properly.
Because as you know in English, there's 20 different ways you can ask a question. Uh, so alright, Well folks, there's two things we learned here. One is garbage in is still garbage out.
The second thing is don't boil the data ocean in the a JI 'cause it's probably gonna be counterproductive. Hey buddy, thanks for being on the show. No, thank you.
All right. Awesome. We'll be back in a minute.