Generative AI and the Data Revolution | Digital CxO Summit
The enterprise is experiencing a tectonic shift and increased pressure on innovation budgets due to economic headwinds – IT leaders are looking to understand their priorities in moving toward key transformation initiatives or risk being left in the dust by competitors or leaner startups. With all eyes on generative AI, the sheer volume and velocity required to build, train and operate these models requires a robust IT architecture that extends beyond legacy platforms making today’s investments critical in creating the foundation for the next decade of growth and progress toward sustainable tech practices (e.g., power and cooling). In this session, Shawn Rosemarin, VP of research and development for customer engineering at Pure Storage, will discuss how generative AI will impact storage requirements and the opportunities and challenges of generative AI across industries to achieve their digital transformation goals.
Top 3 Takeaways:
-Why the industry needs to modernize its data infrastructure to build, operate, and train generative AI models
-The impact generative AI and other emerging technologies have on the storage market – as it can strain compute, network, and storage
-How to keep sustainability top-of-mind as IT leaders look to scale their infrastructure and the dynamics of their data center environments
Transcript
Good morning. Uh, good afternoon for some of you. Welcome to the digital C X O summit and thank you for taking the time today to sit down with me and talk about emerging AI architectures and what we at Pure are seeing in real world deployments.
First off, let's just take a step back. We're almost through 2023, but wow, it's been an incredible year for innovation and what a renewed energy and excitement across organizations, specifically on how they can use technology to create and expand in new markets. I won't beat around Bush Chat.
G P T was definitely a tipping point late last year and after decades of talk about ai, I personally can date it back to my uncle's thesis in 1977 about neural networks and expert systems. And now we're in 20, 23, almost 50 years later, chat G P T emerges and shows all of us the potential of how we could leverage our collective data and knowledge to drive increased productivity and solve problems faster. This consumer exuberance has driven boardroom excitement to AI with most large enterprises, starting to build out focused AI steering committees to look at what problems can be solved and in response to this bubbling interest, public clouds and hyperscalers are doing whatever it takes to acquire and build a capacity of GPUs and get ready to harness that demand.
It is absolutely an arms race as it relates to GPUs currently, and we are seeing that even bubble up in the public markets in lockstep with AI discussions is a conversation around sustainability. Why? Because up to this point in history, the finite nature of electricity and power has never been an issue.
But with AI it is, and it will become the biggest issue because we are simply running out of available power and we are probably decades away from alternative energy sources. And on top of that, the cost of electricity continues to rise. So as we get into today's discussion, we're gonna talk about three main things.
First of all, why your data strategies foundational to AI success? So what we're seeing in the data space, what's holding customers back, what challenges are they encountering and how sustainability goes hand in hand with innovation. So how we can look at that particular problem as it emerges.
So let's just look at this market holistically for a moment. 3% from 2023 this year to 2030. 8 billion by 2030.
It's a pretty large market, right? But let's keep in mind that early interest is emerging. We're starting to see the thoughts percolate about where this could take shape across automotive, across healthcare, across customer service.
And what's most important is that we really look at the use cases and we really look at what problems are each of these organizations looking to solve. So when we dive into that, we can fundamentally see we've got healthcare looking to solve imaging, right? Specifically, can you help me read MRIs?
Can you help me interpret pharmacology? We look at medical diagnoses. Can we look at doctor's notes across multiple patients and start to find, you know, similarities, start to steer a diagnosis quicker, steer a prognosis better.
We're looking at clinical trials for drug testing. Can we actually get drugs through quicker using simulations in financial services? We're looking at quantitative trading or quants as they call 'em.
We're looking at know your customer. Can we actually get through the process of understanding who our customers are, what products to market them, how to personalize their experience with the bank and fraud detection? Can we root out those fraud deal and transactions quicker?
When we look at the public sector, our governments we're looking at visualization lytics, being really able to emulate certain situations, natural language processing and citizen protection. And then we even go across, we go across all the way to regular enterprises looking at chatbots and customer support, marketing, personalization. And lastly, in automotive with autonomous driving, predictive modeling and factory floor optimization, allowing us to build more cars faster, but more importantly, more quality cars faster.
So having worked in many of these projects across pure, the fact is all these use cases are vastly different, but what the customers are asking for is actually quite similar. First and foremost, they're asking for performance, performance and training and inference training and inference being very different sides of the coin for ai. One, I build the model, the other, I leverage the model and ask questions.
But in addition to that performance, they want flexible scale. We are early in this journey. We don't know exactly how this is going to scale, where it's gonna sit, what data gravity's gonna look like.
And so giving us the flexibility to say this workload should run on this infrastructure and this one should run in this infrastructure, in this place, in this country, is very, very important. Sinking costs into large infrastructure from a CapEx basis is less favorable than looking at things as a service and even as an operating expense, allowing us to consume on demand. Obviously, security is big.
We'd all love to build our AI models on synthetic data, but synthetic data doesn't have the same value as real data. In fact, synthetic data can cause unintended consequences in our models. And so security, intrinsic security becomes critical.
Customers are also looking for minimal operational overhead. None of our enterprises have additional opex to run massive farms of infrastructure. They need simple, they need intuitive, they need automated so that the majority of their people assigned to their AI initiatives can be out figuring out how to build and leverage the models and not run the infrastructure at the backend.
And lastly, long-term sustainability. Nobody wants to be on the front page of the paper on the basis of the fact that their AI initiative has put them into the largest, uh, energy consumers or the largest e-waste, uh, creators out there. And so everyone is really thinking from an E S G perspective and a sustainability perspective, how do I do this in a way that is sustainable?
So when we keep going, what happens is we really get to this, how do these projects start? How do they begin? So here's what we're seeing, although there's a lot of talk publicly about sort of, you know, I'll call it three, four, and five.
We read a lot about training and models and parameters and tuning and operating, but let's not ever forget that this all has to start with a strategy. And that strategy has to be linked to a problem, a problem that is large enough that it is worth investing money to solve it. And that problem becomes the, uh, the enigma that folks build these AI initiatives around how do I get a patient to diagnosis, prognosis, and treatment faster?
How do I allow myself to catch a fraudulent transaction at a thousand dollars level before it becomes a hundred thousand dollars problem? How do I look at the quality of cars coming off the line and maybe increase battery life by looking at where the consumption is and better managing it throughout the automobile using ai? How do I build better cars coming out the other end?
All of these start with a problem, and then we have to look at the data. And folks, I'm gonna remind you that we've been through this data discussion for decades. We went from business intelligence, one of my early careers back in the, in the days of building data cubes.
And then we moved to this concept of data lakes. And as we created our data lakes, some of them became data swamps because our data quality wasn't there. And now we're looking at taking those same data sources and actually consolidating, integrating them and building models off them.
So let's not forget that ingesting is part of this, but cleaning and sanitizing is essential. Machine-driven data is typically more accurate than human entered data, but we have to think about how clean is the data that we're building the model on. Because ultimately that model is the library on which our brain is being built.
And if the books are not clean and the books are not valid, then the brain being built off, it will be the same. Once we've clean and sanitized the data and integrating it, then we can start to train it. We can start to think about the models for training.
We start to think about the infrastructure to train it, and then the complexity of the model otherwise known as parameters. How many different questions, how many different vectors are we gonna put into this model? Is it gonna be very specific to one industry in one use case or is it gonna be more general like chat G P T?
Obviously the more specific the less parameters, the more general, the more parameters at a high level, we can then tune it for accuracy and experience. We have to ask it millions of questions to see how it responds and see where the issues are and tune it apparently, and potentially go back and feed more data. And then we have to operate and measure it.
We have to not just sit back and say, wow, we built it, but does it solve our initial problem? Does it deliver the value that we intended for it to deliver? And so this is what we're seeing across the customers out there.
Um, but we've also learned where there are pitfalls. I've put at the top here a workflow and you can see the ingestion coming from cars, coming from genetic sequences, coming from smartphones, coming from factories. You see the persistence of writing that data into memory and then processing it using Spark and other platforms and then going and analyzing it, whether through Jupyter Notebooks or Tableau, and then actually going and training, building our models with PyTorch, uh, or slum or et cetera.
But what our enterprises have found is that across this workflow, and this is a high level workflow, we can see that inefficient and siloed storage is a major issue. Having little islands of disparate storage actually creates complexity because each of those islands has to find its own conduit in which to talk to this model. We're finding unpredictable performance and long turnaround times to be very painful.
We'll talk about this over and over again, but GPUs are your highest cost item. And if you don't keep them busy, you are wasting that money. Think of GPUs as a PhD or a data scientist.
You wanna keep them busy. You want to have enough sitting on the edge of their desk that they never are without work. That's the same thing with GPUs.
We need to not just allow them to talk to each other at the highest rate possible, but we have to also allow them to continue to get work from the engine, our AI engine as quickly as possible and never be waiting. We also have enterprises that are struggling with high costs and complex upgrades. They were given a budget, did they spend that money in the best way possible, not just the CapEx of equipment, but the opex of running it?
Did they look at whether they could sustain this operation for the long term and what the operating costs would be as it went live? We also see organizations struggling with manual integrations, workflows, and orchestrations. The majority of what is happening across this environment must be automated.
It must be must be orchestrated through code. The more humans you need to modify and change and feed this environment, the less humans you have working on your actual problem solving, your actual model, your parameters, your tuning, et cetera. And last but not least, containerized environments.
A lot of these are getting built out in Kubernetes because microservices allows for tremendous flexibility. But those container environments are getting huge, a thousand, 10,000, a hundred thousand containers and actually managing the data services, the individual volumes of storage sitting behind these containers, once again is consuming manpower but also creating risk. So as we look across it as pure fundamentally, I wanna talk about some of the individual projects we've worked on and what that's looked like.
So, you know, we've worked on over a hundred plus AI projects, specifically in ai. Uh, if we look at kind of edges of AI numbers far larger than that. But I wanna talk about some of those projects individually today.
So let's start with meta previously known as Facebook. Well, meta back in 2017 told the world they were gonna build out a research super cluster. And we love this reference because of how large it is.
If we look at the scale of this, right, we have to really think about scales that we haven't seen in traditional infrastructure before. But meta came to us with a problem. They wanted to increase their production training speeds by 20 x on a dataset that scaled up to an exabyte.
And so ultimately they were looking for 16 terabit, uh, terabytes of training data to serve. They were looking at a scale of hundreds of petabytes growing to exabyte scale space and power constraints. You would think Meta has a very large data center.
The challenge was using yesterday's infrastructure, they could only fill roughly a third of every rack, which left two thirds of every rack empty. How do they get density? What kind of technology could they put in that would allow them to get density to a full rack?
Also, this needs to be reliable, right? The beauty of AI is it compliments human beings and it augments our abilities, but we become reliant on them. Imagine for a moment if I'm using augmented reality powered by AI to do surgery on a patient remotely and that service goes down.
Imagine for a moment if I've been using AI for the last five years to do a particular task and now that service is not available, it's very possible the professional never got trained in a manual way to do it. And so as we build these out and we augment human capabilities and we bring everybody forward, we have to think about making sure that these are reliable services. And last but not least, security is key.
As we give up more and more of our data in return for convenience, we will look for companies to leverage that data in the most appropriate way and for them to protect our data, which puts security at the forefront. And so when we think about what this really translates to for Meta, it translated to their ability to build out a research super cluster. And I will tell you with many, many folks at Meta to evaluate technology, pure was the clear choice based on density, performance, availability, and security.
We can also look at Chung Buck Techno Park out in South Korea. And you know, ultimately when we look at Chung Buck, I'll tell you folks a little bit about them. Um, they're an innovation hub that supports growth and economic development in the province of South Korea.
They provide a development environment for deep learning. So they basically help the local South Korean community to, um, help those companies further enhance their innovation. And their integrated development platform includes training data, data cleansing, AI training, and an AI model service environment.
Their challenge fundamentally was how do they process more workloads efficiently? A lot of their local constituents, the local enterprises in South Korea wanted to build out these models. Their demand skyrocketed, but they had a limited number of GPUs and they wanted to figure out how to maximize the GPUs in a way to achieve fast AI training.
While with Pure, they were able to get a two times increase or double their storage processing for faster AI performance, they actually got their G P U utilization from 30% to 80%. 6 times more workload on the same infrastructure and they cut their operating expenses. They were able to manage and operate this in a sustainable way by using Pure one and the rest of Pure's tool set.
If we look at another reference, let's take media Z believe it or not, another Korean company. Um, but they're delivering a cutting edge AI enabled voice recognition. So you've used their technologies within your cars, uh, within public service, within retail voice to text, uh, voice to voice, voice translation.
That's what Media Zen does. And you can think a lot of that's done in AI because it's not just about translation, it's about context of where translation, which actually is very similar to what we see in generative ai or tools like Chat G P T. Well, as they moved from a traditional translation technology to AI power, they lacked the performance.
They couldn't process large unstructured data. We think about videos, we think about about audio files that are very large and very complex and maybe with accents, AI gave them tremendous power to expand. But the problem was compounded when the company expanded its G P U cluster and couldn't scale its storage.
They made this major investment in GPUs. They spent many millions of dollars getting this into their infrastructure and they weren't getting utilized. So as the data volumes grew, the complex storage environment was increasingly difficult to manage.
So not only could they not use their GPUs, but the amount of storage they put behind it to feed, it was so complex in terms of its operating, um, amount of people necessary to manage it. They fundamentally did not have a balance system working with Pure. They were able to shorten the voice recognition modeling cycle from between six and 12 months to two weeks, right?
That's linear and exponential in terms of benefit. They were able to scale into new market, uh, and they were able to enhance their r d capabilities that really enhanced their competitiveness. So they were able to bring balance to their system.
The storage and the operating expenses dropped. That storage became more performant and was able to feed the GPUs faster, which increased their utilization, which allowed them to actually bring their projects to value faster, allow them to bring new products to their customers and potentially more premium services to their customers, faster generating a better revenue stream. So all this is great and really exciting, but let's touch on the other side of the coin long-term sustainability.
5% of the global energy is consumed by storage today. We know that US electricity prices are up by about 25% over the last three years. That's about 8% compounded average growth.
So ultimately, if we continue to turn on these systems and we continue to scale AI initiatives and we continue to create more and more content that needs to be stored, we will run out of energy. We already have countries like Ireland, parts of London, even parts of Virginia and the US that have fundamentally said, no more data centers can be built. We are out of available power.
6 terawatts. That's the same amount of power to power our data centers that would otherwise power nine and a half million US homes. 8 million gas powered cars.
We believe it's important to think about how we are going to deal with this. So across pure, because we're all flash because we've been here for thir, you know, just over 13, 14 years, we are leveraging and we believe that the power of flash, specifically writing directly to flash with what we call our direct flash modules or DFMs, allows us to deliver not just performance and efficiency and reliability that's tied to flash memory, but density. We've already told the world that by 2028 we'll have 300, uh, excuse me, by 2026 we will have 300 terabyte drives or uh, direct flash modules.
And we also deliver our evergreen, which means we upgrade controllers rather than discarding entire arrays. We consolidate storage, right? We leverage, we leverage usable flash rather than drives and we allow our customers to consume as a service so that they don't need to throw away.
They don't need to look at investments of yesteryear. They can buy and consume across block file object as they need to. What this fundamentally does is it allows us to deliver 85% less energy consumption, which means the projects of the future will have the energy necessary for them to be turned on and for them to deliver value.
This is very, very important. Nuclear fission is exciting. It's likely a decade or moral way.
In the meantime, we're gonna wanna power the initiatives we talked about earlier rather than decide who gets power and who doesn't. Or even worse, watch the price curve go up to a point where only the richest of the rich can afford it, or we each get a quota of energy that we're allowed to consume per employee or per capita or whatever it happens to be. We need to start thinking about consuming less today.
And that's exactly where pure is. We also generate 97% less e-waste if you look at what the components are within spinning hard drives versus flash. If you look at the throwaway cost of millions and millions and millions of hard drives across their usable life of roughly three to four years versus flash with a usable life of about 10 years, there's a significant amount of e-waste there.
And obviously with density of 300 terabyte drives, you can imagine the amount of space we take up a 30 terabyte hard drive versus a 300 terabyte flash module. You're talking about 10 times density difference. And so that's where we're focused.
So let's kind of tie up today and then we'll go to some questions. But let's just sort of touch on what we talked about. First and foremost, performance is key.
AI's all about pumping massive amounts of data into GPUs over and over again. The faster organizations do that, the quicker and better results they get. AI resources, our GPUs are PhD students, they're expensive and they're in high demand.
So keeping them waiting can lead to a hefty bill. In addition, we need flexibility. AI is the easy, easily the most rapidly evolving space.
We don't know what the best architectures are gonna be yet. We're in the first inning if we're not in spring training. So these tools, these techniques, these platforms, these data sets, they're gonna evolve.
We move from a hotdog versus not hotdog to medical imaging and diagnosis to self-driving cars to chatbots. So we're gonna need to be flexible in terms of how we consume and how we build to allow us to take advantage of what's out there. Last but not least, enterprise reliability and controls and organization relies on are more important than ever.
We talked about how these will become an integrated part of the fabric of our lives and we need to look at the reliability of them because any downtime can lead to exorbitant costs. So we're gonna look at automation orchestration to not just keep our systems running, but allow us to really have control over what's out there and make sure that they keep running. Last but not least, let's think about operational sustainability.
Let's think about what we're gonna do to ensure that we can get that 85% reduction in electricity consumption so we can say yes to more projects and so that we can effectively be a partner to the business who's asking us to get going and harness the benefits of artificial intelligence. I thank you for your time today. I'm gonna move over and take a few questions in a few minutes that we have left.
Okay, first question. So this is from Sam in Springfield. What's the first step most take when training or building generative AI models?
And do you agree it's the right one? At what point do folks typically realize infrastructure modernization should have come first or been more robust? Look, unfortunately, if we look at the common paradigm of people process technology, we tend to get gaga over the technology, especially us, right?
I'm a geek in this stuff. I think some of the folks on the phone are probably technologists on the conference are probably technologists as well. We tend to get gaga over the technology and how cool it's like we did with chat G P T, we ask a few questions, we're like really cool, but weeks later or months later we say, ah, I'm not really sure what I would do with this.
I'm not sure what problem it solves for me. I'm not sure how to leverage it in my day-to-day life. So it sort of slips off the table, it slips off to the left.
So I think we have to think about first strategy. What problem are we solving? The second thing we have to think about is what if we're right?
What if this does solve our problem? What would it do to our infrastructure costs if this became the standard? What if this new service became the single most important service in the company?
Could we sustain it? What's our monetization model? And so this is why we're starting to think about companies looking at their digital foundation, looking at even some of the infrastructure that they have on their floor that they may have considered a commodity like storage and say, wow, maybe I've let this thing grow a little old in the tooth.
Maybe my operating model is out of check if my data volumes grew by 10 x or even a hundred x based on what I'm now gonna bring into light, what was already a major operating cost is gonna become a boat anchor, it's gonna kill this project. And so, you know, in to this particular question, folks realize typically too late that they've built their nice shiny new bathroom fixtures on top of really old plumbing. And the right idea was to think ahead and say, okay, while this may not be the sexiest thing to focus on, how do I go and make sure that I have the right digital foundation, that I have an operating model and an infrastructure model in place that will allow me to build without coming back later and having to rip out walls without having to come back later and tell the organization I can only serve a third of what I said I could or I now need 50 more people.
Uh, so absolutely infrastructure modernization is not typically looked at first, but it should be part of the feasibility of any of these projects and should be something that customers are looking at early on as they think about what if this project is successful, what would it mean in terms of the revenue coming in, the cost going out, the risk to the organization. Uh, and it is a big piece of that. Alright, second question.
For businesses looking to make the business case for sustainable infrastructure, how do you foster collaboration between IT and E S G program managers if there isn't regular alignment? That's an excellent question. You know, E S G is typically a boardroom initiative.
It's something talked about by companies in their quarterly earnings. Uh, in the case of pure storage, we actually author and publish a third party E SS G report. You can find that online.
Uh, and I can tell you that that goal is actually fed all the way through and through all levels of our organization. Not just environmental sustainability but social governance as well. And that includes things like diver diversity, equity, inclusion, uh, and everything else around how, how we operate.
This is a top down initiative. If companies are serious about sustainability, they have to take the boardroom initiative and they have to set those goals and bring those goals down to each individual operating unit. You'd be surprised how many times I hear infrastructure people say, I don't pay for my power, I don't care.
Well, the fact is somebody pays for the power and at the end of the day it's in the financial statements of the business, whether it's blown, borrowed, built. And at the end of the day, if you're gonna get to a point where you can't further support any further growth without major investment or major cost, um, then it is everybody's problem. And I do think that starting to look at what power are we consuming, how can every part of the organization start to take their part to bring that down, uh, is the only way that this can come to bear and it's gotta be part of a K P I or balanced scorecard across the organization.
Alright, I think we got time for one more. So this is from Kevin Liu. Do you think the mass adoption of AI from businesses will bring some infrastructure management issues like the current DevOps world?
What are your thoughts on this trend? So I'm gonna come back to what I said early. We're in the early innings.
AI is real. It is not a fad. It's not a trend.
It is absolutely one of these major initiatives that just like we saw with the advent of the internet, just like we saw with the advent of micro computing, this will carry us through the next 15 or 20 years. At the end of the day, um, mass adoption I believe is several years away, but there will be a lot of individual projects and testing. There will be a lot of figuring out, right?
You talked about DevOps here. If I think about DevOps back in the early days when we were even talking about, you know, dot net and WebSphere, I think back then, um, we learned a lot since then. We tried what was called bimodal it, we separated traditional apps and modern apps.
We completely separated the teams. And then we said, well that's not really working very well. We actually need one team.
So we got to this model of more like, you know, maintain what you build. Uh, and this concept of everybody having a joint responsibility to keep code clean from the time that it's built, to the time that's operated, to the time that's delivered. As we think about ai, um, ultimate ultimately what I would suggest is we think about what is working and we build better collaboration to talk to those that are a little bit ahead or kind of on the pioneers and the early adopters and really spend the time to glean to the things they wish they had known when they embarked on this journey.
And what we're finding more and more is thinking about the long-term sustainability, the model, thinking about what happens if this is the next big thing and this is successful before we go and make major investments to turn it on, uh, becomes more and more critical. So, you know, absolutely we have a lot to learn here, Kevin. We are early, but I do think there are many, many companies ahead.
If we look at research, we look at even what we learn from H P C, AI and H P C are very different, but we look a at a lot of the lessons we learn there, um, there is a lot for us to glean. Um, okay. And then last but not least, one quick end of end of time question.
Is there a blueprint for governance aligned with long-term sustainability? Absolutely. So what I would suggest is do take a look at Pure's e s g report.
Uh, it is clear, it's available online. It's just published as of two or three weeks ago. Uh, you'll see it's quite extensive, not just in terms of what is our, what are our plans in terms of sustainability, but how have we broken everything down in terms of phase one, phase two, phase three, where are the areas that we see ourselves driving sustainability in terms of greenhouse gases, terms of energy savings.
Uh, i, I do think it, it provides a really strong model, uh, and one that we can learn from. Folks, we're out of time. I really appreciate all of your energy, uh, excitement and participation today.





