How a Data Diet can Help to Achieve Essential Sustainability Targets | The Six Five Summit
Cohesity has analyzed the issues faced by the exponential growth of data worldwide, resulting from cheaper storage, cloud growth, and AI. Data center energy and efficiency are not keeping pace.
While AI is data and power-hungry, alongside machine learning, it can help defuse one of the most complex problems, “unknown” data. Predefined filters immediately fish compliance-relevant data such as credit cards or other personal details out of the data pool and mark them. Once loose on the data, the AI develops a company-related language, a company dialect. And the longer it works and the more company data it examines, the more accurate its results become. Companies will be enabled to automatically identify obsolete, orphaned, and redundant data that could be deleted immediately.
Transcript
All right. We're joined right now by Mark Mulino. He's the, uh, CTO of an interesting company called Cohesity, facing one of the most interesting problems in our society right now, yet that big, these ideas about sustainability data centers and, and the the ecological pain that AI could give us.
Mark, so glad to have you on. First of all, tell us about what Cohesity is. Yeah, so, um, Cohesity is a leading company in data security, data management, and ai.
Um, we've got a mission to protect the world's data, and, um, I think we're doing a very good job, though. Uh, well, I certainly hope so. It's a big job.
Um, uh, but the first thing that comes to mind when we talk about, uh, protecting data isn't sustainability, uh, our, our environment, uh, the uses of, of water and power and, and, and climate change. Uh, explain to me why this is such a concern in general for data centers, and then let's go on and talk a little bit more about how, uh, AI changes that. Yeah, so I think the, the, the, the challenge around sustainability is that most companies today don't actually figure sustainability as a big picture item within their vision and strategy specifically for data.
So they have it in their company agenda, they have it in their goals to do things like reduce energy, turn off lights, et cetera. And some go a stage further and actually start properly looking at scope three emissions and looking at, um, downstream's providers and transport and things like that. What they don't do is look at the data.
So what we are seeing is we're seeing data grow by 50% exponentially every year, and that data content isn't categorized. People don't necessarily understand what that data is and why they're keeping it, and it grows and grows and grows. And because they don't understand much about that data, they didn't do anything with the data.
So if you think about that in data center terms, that's filling up storage, that's filling up media, that's filling up full space. It requires, uh, electricity to run the technology. It requires water to cool the technology.
All the technology produces CO2, which obviously contributes to greenhouse gases, which is what are the aims that the company as a bigger target is trying to reduce. So what we're seeing is we're seeing data growing exponentially with very little care taken about it, and that directly impacting what a company's trying to do for its sustainability goals. But isn't the promise of data that the data has value and you don't know which data has value or, or that AI will give you the ability to access and extract value from data that you didn't know was valuable.
You know, old credit card receipts from customers that you may have thought you lost. Now you've realized there's something in the patterns that may tell you something about your business. And the best thing to do is we were told is, is keep that data around you might need it later.
Absolutely. Yeah. So that's the key mpromise behind keeping the data.
Everybody keeps the data because it's this big ticket item that they can then get insights from in theory, from, yeah, in theory, but it's also in fact that data is valuable. It can provide those insights that you need it to provide across a variety of use cases. However, AI isn't, uh, a magic wand.
It just, it doesn't just work. You have to train it. And unless you know what the data is that you are putting in the, the, the language models that AI use are only as good as the data that goes into them.
You have to train them effectively, otherwise they hallucinate they don't bring back the correct results. So the, so it's largely useless. So to be able to make AI useful, you have to understand what the data is.
And if you don't understand what the data is, you've got this circle now of, of the fact of, of these things just won't work because you don't understand what that data content is. So we all come back to understanding what the data is. Now, AI itself drives exponential use of technology.
So as I said a moments ago, technology gets more and more intensely used as data goes onto it. If you look at it from an AI perspective, your average server rack is probably using eight kilowatts per hour, but AI is 30 kilowatts per hour. So it's a huge difference when you're starting to push AI through that.
3 kilowatts, but if you are doing an interaction with a large language model, alphabet's chairman said it was gonna be 10 times the value. So you're now talking three kilowatts per large language search chat, GPTs responding to 195 million of these a day. So now you start to see 560 plus megawatts of electricity are now being consumed by ai.
So AI has a dramatic effect on what you do with sustainability, but it still all comes back to the point that we were mentioning a moment ago around data. If you don't understand what that data is, how useful and valuable can that data be to you? And this is partly where we're advocating as Cohesity that you need to understand far more about your data.
You know, we're, we're, we're coming to the party from, it Seems that you're also arguing, you gotta get rid of data that if you don't know what it is, toss it. You haven't used it lately, toss it. That's exactly what it is.
Yeah, I mean, the term that we use for it is defensible deletion. So can you make a defensible deletion decision against that data? So you are probably holding Cory's shopping list from 10 years ago, maybe pictures of your dog, you know, um, videos of not an old conference call that you did that you needed to keep for a few weeks.
There's no intelligence around how much data is kept on an individual basis. And then if you expand that out to a department or to a company, it's huge amounts of data that's kept hand over the fist that isn't needed. Now every company has a relevant record strategy.
So strategy where a record for a particular purpose is kept for a particular length of time. What we don't see, especially in unstructured data, is anyone categorizing the data. So classification and indexing of data is absolutely critical.
If you classify your data, first of all, if you index it so you know where it is, if you classify that data and then say, well, I know that that's co shopping list, but I also know that this is a mortgage record, or this is dental records healthcare. This is a wing blueprint from an airplane that was made 10 years ago. These are all relevant records that have to be kept for a period of time.
And in some cases they have to be kept immutable for a, for that set period of time before they can be deleted. But you can start bucketing up in this record strategy to create this defensible deletion to be able to say, I can make a decision about that data and not keep it. And this is where you start to reduce that volume.
This is where you start going on a data diet. This is where you start reducing the volumes of data, which then contribute to what you are doing with sustainability directly. It, it's a fascinating conversation I'm thinking of, um, I couldn't, in preparing for our interview today, I could not help but think of, uh, one of my ex-wife's stories, and I don't like to tell these public about, I'll share one with you.
She saw an episode of Oprah that said, if you haven't used something in three years, throw it away. And then went through and threw away all of our old tax records and bank records, uh, because she had never used them. And so her decision facing that pile of data was, this is useless data.
My position, when I returned home and found out that the garbage man had taken away the, uh, many years of tax aid and banking data was different. We'll, to say, to keep it, uh, um, you know, keep my language outta the four letter word zone, my, my take on the data that someone else decided to toss was very different. How does an organization deal with that?
Uh, understanding they've got sustainability and climate goals and wanna, they wanna limit how much data they're paying for to store, but recognizing that somebody else might have a different view of the value of that data. That's it. Exactly.
I mean, that's, that's where the relevant record strategy comes in every company, and I'm Not asking for marital advice No, no, but it kind of is. I've been married for 30 years, so I wouldn't even dream of giving it. Um, yeah, every, every company should have a relevant record strategy, and if they don't, they should be creating one.
And this applies to every record within their business that has a material value. So every business unit will know what's valuable to them. And in some cases on the industry you're in, it's prescribed if you're in financial services, you know, you are keeping those loan records and tax records and mortgage records, et cetera for set periods of time.
That's your relevant record strategy. What I tend to see when I'm talking to customers is there's no correlation between the relevant record strategy at the top of the company that most employees do training for and attest to every year to what actually happens down at the bottom of the company. In a backup strategy, what tends to happen is the backup c uh, group are told to backup up the servers or backup up this storage array or backup up this piece of data.
There's no correlation between the backup policies and the relevant record strategy. And if there was, they would know exactly where all that data was within backup. They would know where it was within storage because it'd be clearly marked.
It would be classified and it would be indexed. Now, you can materially get value from that immediately through reporting anyway. Product, for example, reports on data content, it reports on ownership, um, last access size type of data, and also, um, elements within the data P two data, for example.
So you can already go a stage with backup and and recovery if you like. If you then put an AI layer over that because you've classified an index that data, you can now use AI to start u to train your large language models or to use something like, uh, retrieval, augment generation to actually create a bucket of that data that you can then query with natural language. So you can go in there and say, Hey, I want Corey's tax records for the year 2018, and it'll come back with everything, or it'll come back that time you, you under page your tax between, um, one month one and month six, and it'll come back with that.
So you've enabled that data mounting to start materially giving you really strong insights into what you've got on the shop floor. Then you can make a defensible decision against what you do without data. And there are also legal ramifications about, uh, what data is kept and what is, and I'm thinking about, you know, I think, I think the financial services industry is very strong about this only because they've gotten in so much trouble from, from eliminating, uh, messages.
I'm thinking particularly the messages about traders and things that have led to big lawsuits over the years so that they've have some really rigorous policies about what can be destroyed and what cannot. Absolutely. Yeah.
Yeah. I mean, yeah, tick data, all sorts of things is, uh, are kept as part of financial services. But I think after the tr uh, I mean, I was lucky enough to be part of the 2008 financial crisis.
So I was on a group where we were actually doing data retention and yeah, look at me. And we were, um, we were, we were in a retain all mode because we really didn't know what data we needed to keep at that time and how long we needed to keep it for. So we just kept everything because that was the path of loose resistance.
And in a company, uh, you know, in the company the size of a major financial, it's a risk weighted decision to say, well, okay, the risk of of not having this data and being found by the regulator is actually lower than the risk of keeping all this data and just spending money on it. Remember back in 2008, sustainability wasn't the big ticket item that it was today. It was only after COP 26 that we started and the power, the Paris Accord that we started to see companies worldwide signing up to their government strategy for sustainability goals and agenda.
Now it's a material difference. Now, you can't just keep going out buying data centers, sort of filling it with storage and filling it with data. You've gotta be intelligent about the way you deal with that and how you, how you classify and index that data and tag it for AI use because then you can get, get genuine insights from it.
And as I was saying, you know, it's not just about the insights that you can get from the data are hugely strong for your business, but as you said, it can go back to the regulator. You know, if you have a security breach, you have to report to the SEC within 72 hours. Well, I'm Thinking of, I'm thinking of 2001 when there was an investment banker in, uh, in Silicon Valley.
I'm pointing to Silicon Valley behind me here. Uh, there's an investment banker in Silicon Valley who upon being informed, uh, that the firm was under investigation for some, uh, uh, potentially bad practices of, of paying kickbacks to executives to steer banking business their way, instantly told one of his colleagues to send out an email to everyone saying, remember our policy, delete all your old emails right now. Uh, and that became a crux of the case.
Yeah, Yeah. And that's it. And that's, you know, if you had a proper strategy in place that linked your relevant records to your backup policies, the guy there in that example could have gone and had everyone delete their emails and it wouldn't have made a difference because there would've been a copy in a clap in a, in an indexed fashion that could be queried.
And in modern times now with ai, you can query it with natural language. Let me ask you finally, it seems that there's always gonna be an inherent, um, conflict between the, the idea of, of storing up lots of data because it has value or potential value down the road and not storing data because there's the cost of doing so is becoming greater. And as you point out 10 x and ai, what's, is there a simple solution to figure out what do I do when I'm at that fork in the road?
I want to keep it, I need to toss it. What do I do? It's a, it's a material decision, isn't it?
I think when you are, when you are at that point, you are gonna understand whether you need to delete the data because it's obsolete. Whether you need to keep it for a particular purpose. I think if you follow this mantra of, I mean you've gotta start somewhere.
So you might as well start now. If you can start moving forward and classifying and indexing the data going forward, you can make those material, those decisions immediately because you know, you are not gonna keep my shopping list for the next 10 years. You're gonna keep it for five minutes and then get rid of it.
And that's, that could be a personal choice, it could be an enforced choice, but companies can put this enforced data drop within their policy so the data doesn't exist after that period of time and you start to manage that mountain. So if you, if you know you are growing at 50% a year in your data, but you also know that probably only 30 to 40% of that whole data is actually of any value, why wouldn't you start making those decisions now? Why wouldn't you start driving those numbers now going forward and then kick up something in parallel to go back and look at that other data to make those decisions against it?
Because it's all gonna go towards your sustainability agenda. You know that what, as we said before, if AI's gonna drive um, 30 kilowatts per rack, you need to be starting turning some of these racks off and reducing data is the answer to that. Alright, here's the Marie Kondo of data.
Mark Ow of Cohesity. Thank you very much. We appreciate your time.
Thanks Very much.



