Processing Data at Scale with Jeff Denworth of VAST Data
In the wake of picking up an additional $118 million in funding, Jeff Denworth, co-founder of VAST Data, discusses the impact of processing data at scale in the age of cloud computing.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Jeff Denworth, who is a co-founder for Vast Data, and we're talking about their recent edition of about 118 million in funding that brought their valuation up north of $9 billion.
But more importantly, where are we headed here as we go into 2024? It seems like the way we managing data, the amount of data, the velocity is all changing faster than we can handle. Jeff, welcome to show.
Thanks for having me, Mike. So give us some assessment of where we are. I think we've been managing data for ever and a day, sometimes, well, sometimes not so well.
But these days it seems like it's exponentially increasing. AI is changing the landscape. What are the challenges we're really facing?
You know, it, it's funny about, uh, about 15 years ago, the term big data came into the market and everybody was talking about how this is the new oil and things like that. The reality is, if you kind of look at that market event, it was really around the modernization of the data warehouse, right? And this is a largely addressing data types that can fit very cleanly into things like tables.
Um, so typically numerical data, things that can be very easily queried upon. And if you look out, if you take two steps back and you look at all the data that's being created in the market, about 95% of the data that we create is, is called unstructured data. Unstructured data is like videos and files, uh, in the form of imagery or, or maybe like free text.
If I'm working with chat GBT, that's like a little text file, uh, or maybe like instrument data. And as this stuff is being created, you don't actually, from a systems perspective, historically, you haven't known what's inside of it, right? You know, if you go back 10 years ago and you said, okay, I want to go find a photo, you'd literally have to look through every photo to, to kind of try to determine what you're looking for.
So, um, a few years ago, uh, GPUs and neural networks in the form of, of AI and deep learning came around. And for the first time it's, it's possible to make sense of unstructured data. So we think, um, what's happened is that this new AI revolution has really opened up the, the aperture for the big data market and, and essentially grown it by a factor of 20.
'cause now we can start to process on all this unstructured data that's largely just been sitting in your archives. Do you think we're gonna have to clean up our data management practices a little bit before we dive into this whole new world of ai? Because the models are somewhat sensitive to the quality of the data, and we're not so good at managing the data.
So do we need to take a pause here and figure out how to, um, mature data management processes? Yeah, for sure. Um, I, I read an interesting blog post.
I'm, I'm not sure how accurate is, but it was from a employee of open ai and they said, you know, if you kind of look at what's happening in terms of the quality of the models, more and more it's the quality of the data set that's being used to train the models that impacts the model's accuracy more than the architecture, architecture of the training model. And so the interesting thing is that data is, is is definitely the fuel of ai and quality data tends to be what determines what is and is not good ai. So, um, so AI actually comes to the rescue with respect to dirty data too, because now you have these tools that can actually go and, and like curate and label data before it actually gets exposed to different training algorithms.
So, um, it kind of becomes a self-fulfilling prophecy. As we get better at ai, AI gets better at the data, and then that kind of process goes on and on. What does that do to our storage systems?
Then? What kind attributes do we need to manage or handle all this data? Well, um, so, so once you, once you unlock the ability to process on the data that's been in your archives forever, uh, it does challenge the definition of what you could think of as archive storage or data store infrastructure to be like forever ago.
You know, if, if you kind of think about how people consider archives, you know, conceptually we always think, okay, that's the slow place that data goes to die, right? Um, but now if you kind of look in the dictionary and you say, what, what's an archive? And you not talking about it terms, it's basically like a, uh, a government or some organization's source of truth or they're like, you know, information repository.
And so we see a changing definition of the term archive where it is that source of truth. And if you can, if you can get real time access to it, then you can actually bring AI to your data and, um, very quickly your models will improve as they're just exposed to more and more of the, the real world. Um, you know, deep learning is basically just applied statistics.
So the more that you can represent the real world and what you train deep learning engines with, the better these models become because they understand more and more what's happening in the world. So from a storage perspective, that goes all the way down to kind of rethinking your relationship with data, right? Because now you need to go back and train and retrain these models as they start to drift, and as you have improvements in your algorithms.
And so the idea that, you know, the archive is the place that your data goes to die is kind of also dying. Um, where now we're seeing that customers want fastest access to the largest amounts of data in their environment, and that's one of the places that vast really specializes at. So what exactly makes vast different, we've had storage platforms forever and a day.
What exactly is differentiating you guys? What's the actual architecture? So, um, so, so what we built is something that we call the vast data platform.
Um, and storage is, is the foundation of this, but it's not the entirety of the solution. And for the last 20 years or so, if you think about distributed systems, um, they've all been kind of born of an architecture that was, was developed initially by Google in the form of what's called the Google file system. They wrote a white paper around it, it was a seminal white paper, and it created hundreds of billions of dollars of products in the market that kind of tried to copycat or replicate the success that Google had, um, that gave birth to all of the file systems, all the object storage systems, all of the NoSQL databases, all the hyper-converged platforms.
This notion where you have like a commodity node that is stuffed with a bunch of storage devices and you build clusters out of these styles of systems. Well, the problem becomes as these systems now need to kind of keep in lockstep with each other, there's a lot of internal traffic for these, these, um, these data systems. And what happens is that they become really limited for certain applications, right?
And what we saw was an opportunity to reinvent the way that people store, process and manage their data. Um, and so we built a new style of systems architecture. We would argue that it's the first new one in the last 20 years.
We call it, uh, disaggregated and shared everything or our our days architecture. Um, and what it does is it, it essentially decouples compute and storage and allows you to kind of have an infinite amount of compute in a data center, all talk to a shared global volume as if it was directly attached to every machine that's running our software. And so you can kind of think about it more like a web scale computer that we've built that's completely parallel.
And the objectives of the product were, um, first and foremost to kind of kill the hard drive. We said, okay, if AI is coming and people are gonna want to get this like real time relationship with their data, you can't have hard drives in the way, there can't be any mechanical media. So we spent a ton of time really building these algorithms that could bring more efficiency out of flash infrastructure.
And it's counterintuitive, but the only way to build a system that's cheaper than hard drive based storage is use flash if you can implement these new algorithms that we built. And then we said, okay, well let's, let's synthesize structured and unstructured data. Uh, and what I mean here is that we've integrated, uh, a file system and an object storage system, which is where people typically put their raw data that they may want to either process store a train upon.
Um, and then we've combined that with a next generation approach to building a, a essentially an exabyte scale distributed transactional and analytical database. Why would we do that? Well, as systems running, as data's running through this system and you have the ability to capture these new learnings on unstructured data, well that has to be kind of cataloged somewhere.
So we put it in this new database system that you can essentially query upon. And what it is, is it becomes a query engine for unstructured data. Um, and then the last thing that we've done, if you think about this, this concept as a web scale, um, very low cost flash powered computer that's combining structured and unstructured data is we, we essentially add, are adding triggers and functions into this system so that you can think of it actually as a runtime environment where programmatically you can now have, um, data flow through the system and it will go and do that data refinement that I talked about earlier and, and kind of derive, um, structure or understanding from unstructured data.
And so all of that comes into one package that you can deploy from edge to cloud. Um, and where customers really like it is they can just start with like classic enterprise storage applications or classic enterprise data lake applications. And then they start to kind of grow into the capabilities of the system and they realize that not only is it modernizing their data management approach for their legacy applications, but it's also readying those applications for, um, the additional enhancements that are coming with respect to either machine learning or deep learning.
So it's a big insurance policy against the future. Um, and for that reason we've been kind of selling like hotcakes for the past four years. Are we kind of coming full circle here?
It almost sounds like, you know, after years of trying to bring data to the compute, we're actually creating an architecture here that allows us to bring the compute back to the data without having to always move the data. So that's a, that's a really interesting question. And um, that, that data to the compute idea was actually started by Google, with the Google file system.
Um, and at the time, um, storage devices were much faster in the network, so it made a lot of sense to bring the compute to that part of the data and, and partition systems. And it it created a lot of, um, complexities as well. But if you fast forward 20 years later, networks are now about 800 times faster than they were when Google originally, um, invented that architecture.
And so from a performance perspective, it's okay to move data around networks 'cause you have so much more bandwidth than you ever had in ways that it, where it wasn't in 2003. And so, so yeah, we believe in in, in, you know, kinda very high performance networks within, uh, within customer data centers and across data centers. And if you can get there, then you can get a much more fluid data processing model than what, uh, I think that bring the compute to the data model supported.
Is this something the average storage administrator can handle, or does it require somebody who feels and acts more like a data engineer? What's your sense of the level of skill required? Um, so historically we've been selling the product.
You know, when we came outta the market, it, it essentially was just a, a file and object storage product. Uh, and that was sold to, uh, I would say a class of storage engineers or storage administrators typically that kind of rotated on the, um, the higher capacity scale than the average customer in the market would. I think vast is unique where we decided to come into the market from the top down and work with the biggest of the big customers just to start with the thinking that it's a lot easier to scale down a product than it is to scale up a product.
Um, and that, that model's been very successful. We've got a lot of customers that have deployed hundreds of petabytes of infrastructure with us. Um, now as time goes on, when we add things like database services and we add, um, actually a programming environment that you can use to go and refine your data, that has us also working with data engineers and chief data officers as well as developers that just wanna bring their code to the systems.
And you can think about it as some ways as like programming your storage. Um, and so it's, it's, it's been evolving. We've talked about how AI is changing the way, uh, we need to access and manage data, but will AI also be applied to the storage systems itself and will that make things easier to manage?
You know, how will these algorithms kind of evolve on that side? I think they already are. Um, we have been shipping for, were probably the last year or so, uh, a variety of different machine learning, um, engines within our system for different purposes.
So a a simple example is a few weeks ago we announced something called anomaly detection. And so, um, it's kind of a nebulous term, right? It has, um, different definitions for different people, but as, as users are interacting with the system, what we start looking for is patterns that are, um, you know, from a heuristics perspective are not common.
And so if some actor is like overriding or deleting a ton of data, or if some performance metric goes a little bit wonky in ways that you wouldn't expect to happen, we send off alerts to engineers that are managing the system saying, Hey, this is a point of curiosity that you may wanna look at. And so we do that for data reduction, we do it for anomaly detection, we do it for a variety of different, um, ways to kind of, uh, enhance the data flows through the system and make sure that the system is more and more secure. And so, yeah, I think the age of, um, AI infrastructure being infused with AI is already here.
Will we also see generative AI capabilities to, yeah, help me figure out what's going on in the storage architecture. I can just kinda ask a question and see what happens. Our system's pretty easy, so you know, that's kinda like a 100 level class.
If managing our systems a one-on-one level class, uh, yeah, prob probably in the future. I think, you know, the interesting thing that you will see it in is, um, you know, this, the synthesis of a structured and an unstructured data system, um, you know, SQL as a, as a language is pretty complicated for people that haven't been programmers or database engineers. And so there is a, a trend now that evolving where you see, um, just natural language search starting to replace or sit on top of sql that can translate just standard requests or questions that you and I have into SQL commands.
So that's also coming, but I think that's more of like a user or, um, potentially a developer interface more than a, um, a systems manager interface. All right, so we're at the tail end of the year. Looking into your crystal ball for 2024, what are you expecting to see happen?
Well, what I can say is I don't see things slowing down. Um, I tend to think of infrastructure as a forward indicator for how people kind of view the adoption of new AI tools and applications in the marketplace. So we are seeing the action before the models get trained and things like that.
Uh, and you know, spending time with some of the largest organizations in the world like Nvidia, who are also investors invest, I can tell you that, um, things are very, very active right now. And so I don't expect things to slow down for quite some time. All right, folks, you heard it here.
It's always been about the data. It's just a question of matching the architecture to the data depending on the use cases and well, the use cases are changing. Hey Jeff, thanks for being on the show.
Thank you for having me. All right, and back to you guys in the studio.