All-In-One Data Lake – Ram Venkatesh, Cloudera
Cloudera CTO Ram Venkatesh explains why the time has come for an all-in-one data lake as a service.
Transcript
This is texturing TV. Hey guys. Thanks for throw again.
We're here with ram Bank. Attach who's CTO for cloud Derek. We're talking about the launch of a new offering they have which is an all-in-one data lake house, which I guess you can access in the cloud or wherever you need it to be and we're gonna be talking about what exactly is a data Lake these days and and what goes into that and how much heavy lifting do I need to do to manage that whole Space so Ram welcome to shop.
Thank you. Thanks for having me on the show. Happy to be here.
So explain if you would what exactly do we mean by it? All in one data lake house is that you know, TurnKey platform or is everything becoming automated what's going on? Sure happy so the Legos is really this is the it's a very simple idea.
I think that it's hard is being able to analyze all of your data efficiently. That's what the lake house is all about. So it brings together, you know a data Lake think of this as infinite storage, right?
So you can land on different kinds of data streaming data real time data batch data tabular data right flat Files video music what have you so the lake gives you a place to land all of your data the lake house part of it is let's you actually buy. By enabling sql-based queries over the lake. Now, you're made all of the data in the lake both the structured the semi-structured and the unstructured data that's in the lake accessible to all the analysts in your the business owners that you have in your company.
So that's what the the framing for a lake house is essentially, you know structured analytics at scale across all of your data. And the all-in-one piece of this is with CDP 1. What we are doing is we're making the lake house now be available in a new deployment form factor.
To your question about how much work it's going to be for my customers, right? So CDP the Clutter our data platform. We have traditionally had two form factors one.
We call customer managed. This is the CDP private Cloud offering where the customer runs the whole thing themselves. Then there's CDP public Cloud, which was a platform as a service offering so CDP 1 is the software as a service version of the same thing.
So it's the same platform and now available in three different deployment choices based on what our customers particular requirements and it might be We've seen people build data warehouses before and then we saw the emergence of the data Lake and then that led to issues around what became known as the data swamp and a lot of people couldn't figure out how to implemented or maintain it. So what is changing here to make the whole thing more accessible to people and maybe help guarantee that they're going to be successful the first time It's good. Good question.
Look, I think that the conflict between the rathaus on the lake has really been one about you know, the red house was all about like an end to end vertical stack that you could use sequel to analyze all of the data as long as it was in the warehouse. And the lake was as horizontal thing, you could land any kind of data, which meant if you were not being very rigorous about it. It could quickly degenerate into a place that just had a lot of data that was not easily accessible.
So what's changed is the the piece that links these two together is what we call the Open Table format. Think of this as a way to Overlay our tabular skin using metadata on top of the data that's already in the lake. What this this new architectural piece in the stack then lets us do is you can come in through the warehouse and you think you're dealing with tables and rows and columns and you can you can do your SQL efficiently.
You can learn the data into the lake as if you were just riding it to this infinite store. So in a way for a lot of our customers this is it this the lake house lets them have their cake and eat it too. That's it.
They can land semi structure data and then they can analyze it through through a sequel based front end. And so the lake house this is a pattern. This is not new, right we discovered this pattern probably seven eight years ago because it's much easier to analyze data using SQL than almost any other way of playing with large amounts of data, especially in interactive fashion.
And so this pattern in the Standard Time, what's new about it now is this new open source protocol and standard that makes us even more easier and even more interoperable than in the past. This is Apache eyes. And so injecting Iceberg into this lake house conversation is the piece that has made the lake house truly open.
That we can launch queries against larger volume and diverse set of data. My question to you though is are we making it easier for developers to access data as they build their applications because that's always been a hurdle. It's we build the applications faster and then we wait for the data to put into those applications.
So we bridge in that Gap. Yeah, so and for increasingly for our customers, it's all about velocity, but it's like how quickly can we stand up a use case on top of the data that we already have how quickly can we get to the time to Value if you will that's that's the core metric that our customers really want to optimize on. So what we realized is that there are three contributors to where time gets spent on a project that's worth optimizing right.
The first one is actually operational, you know, standing up the platform making sure that it's secure connecting up to data sources and Landing the data inside the platform. That's the first piece of it. The second piece of it is how do you make it be you know, self-service for practitioners to massage the data to fit the analytics or the use case that they are looking for and in the third piece is how do you actually produce business consumable dashboards and reports and so on again in an interactive self-service way.
So with CDP one, we're actually attacking all three parts of this problem. It's fast so by definition it's zero Ops, right? So the amount of operational can and feeding that you have to do is significantly low.
In CDP 1 we are also introducing new what we call low-code connectors to bring data in from different sources without adding a lot of code. So for example, if your data isn't your relational backend or increasingly even in a SAS application like maybe Salesforce? So you want to bring Salesforce data into into a particular use case you can do that just by configuring the Salesforce connector and selecting what's upset of the data you want to work with but without writing a lot of code to accomplish that and in the third piece is around self-service analytical exploration, which we can do with that visualization framework that let's customers the business analysts themselves.
Not not the the development stuff that they have but the analyst themselves can assemble reports. They can they can make these things look exactly the way they want them to and then they can publish them out. So it let's them automate out the entire workflow and by addressing this holistically, this is how we plan to make it more.
See more efficient for customers to deploy their use cases much faster and to end as an example cwt where an international travel firm. They've been working with CDP 1 as a you know, when it was in private preview for a while and they had to roll out a fairly complex travel use case that touched sensitive customer data. You can you can see that they work in multiple GEOS and there's different regulatory concerns when you go do that.
They were able to deploy this ammetical model based on top of CDP in a matter of days to a small number of weeks to get the interproduction and that represents like significant ease of use over what they would have had to previously previously could take somebody three to six months to operationalize in use case like that and now they can do this in a small number of weeks. So I think that's how the platform has gotten easier for customers to work with. We hear a lot about digital transformation these days data is the new oil and yet we hear a lot of people are having trouble with this whole Paradigm in this whole shift.
So did we maybe underestimate that challenges associated with trying to manage data at this level of scale and velocity? I think that you know, there's too facets to the problem. We naturally gravitate towards the the technical part of it.
You know, what features can I enable in mind in my product or service to make it be easier for product customers to get value and that that train will that'll keep chugging along but I think there's enough knowledgement now in the industry that there's another piece to this which is cultural right which is you know, how to be a data driven Enterprise but our customers they're looking for ways to build trustworthy data sets. That's a you have a data set that you know that this is accurate. That's good Fidelity.
I can trust it. They want to build datasets that are shareable. So that if one business unit produces a data set the second one doesn't have to go through pipelines and all this other magic for them to get access to the data.
They need to be able to discover it and use it. Right and the third piece is around collaboration and sharing all of these things are are cultural. We can enable a set of things in various products and services, but I think that data data as a product mindset I think that has taken a lot more time to evolve than we thought and I still think we are at the beginning of that Journey that's going to take us another six to eight years.
I think for the industry to get really mature and thinking about You know, how do you build data as a product? How do you how do you move data around so conveniently so that you can actually deploy lots and lots of use cases, you know, when you talk to our customers, they don't want to deploy like three use cases or five uses a 10 use cases. They want to every part of their business has to be territory.
So we're talking hundreds. We're talking thousands of use cases. That's what I think the scale of this product mindset thinking is still I think in the early part of the journey there's recognition and now we need to see that actually yeah.
It almost sounds like yeah, maybe I'm dating myself a little bit here. But yeah back in the day we had XML and one of the cool things about it was it was self-describing. Are we moving down that path now where the data actually does describe what it's about and maybe where it can flicks with other data and then we can automatically resolve all that throw some algorithms at it.
Maybe we're not going to be spending all our time manually massaging data. Yeah, that's so self-describing data and comprehensive robust metadata. That's a big part of it for sure.
But I think you probably don't need to go quite that far back in the past. I think the the piece that is starting to get really exciting now is you know, people are applying I want to call like conventional software architecture microservices like thinking to data. This is the whole data mesh Revolution that you're starting to see in the industry.
Right? I think this is where you know people need ways to like think of data problems as small team problems things that you know, three or four people can go and put together and stand up in a matter weaks and they can realize value and then somebody else can go to another one of these someone can discover the things the artifacts produced by the first two teams and build out a third one, but it's that sort of viral composition of data apps and use cases which will be built on what you talked about on metadata and on being able to describe the relationships between data, but it's really more systematic. Applied to how you actually build a data application.
I think that's where some of the you know, the the 10x improvements and productivity are going to come from Do you think then we will? Bridge The Divide between the business and it and I know we've come a long way but frankly, there's a lot of business users that don't trust the data and they don't want to make decisions based on it because the date is conflicting or it just doesn't jive with their hands-on experience. So are we getting the point where we can have that Kumbaya moment because we do have a set of data that's reliable.
I think that to me the way I look at this problem is I feel like it is an enabler. You know ID can make sure that your infrastructure is reliable secure, you know, your your being pragmatic about cost. There's a whole bunch of things that it can do to make sure that things function properly and all of those are still very valid and relevant when it comes to data use cases the piece that you talked about trusting the data, I think that is actually if you start to think about the data as a product then every product has a product owner, right and that product owner need not be necessarily be part of it.
For example, if you have a marketing product owner, they know the marketing data sets and their company they know which which campaigns are being run at what time they know which data sets are current which ones are out of date which ones are gonna get replaced three months from now, right? So, I think that that's where the I see the the way teams collaborate to produce data. I think this data product owner is going to be increasingly a very important person to our history how the use case gets.
Out and this is very similar to how apps get rolled out today, you know an application a product owner or a product manager is so critical in the app space for you to get an app. That's well put together and can be released in a very experient way the same kind of thinking I think will get applied to data and then it will have a role to play in making sure that you have consistency you have you know security you have compliance your risk and cost all of these things thought about what the actual use case the business needs to be able to build out that use case with ask little friction as possible and We have seen the emergence of data Ops. Is it discipline and a lot of the principles that are borrowed there are from devops.
So when we see devops and data Ops kind of converges we go along here. I think increasingly, you know, both of these they have this underlying theme of being policy based. Right, like being declarator being able to specify exactly what you want and then let the systems whether it be sass whether it be pass whether it's customer managed, you know, you have processes in place to automate the things that you want to automate to achieve that desired and state that you're looking for.
I think in that sense that devops culture I think increasingly is so critical to be part of the data landscape. The other piece there is around making sure that you have really robust ways of testing. And releasing the software to make sure that like cicd like methodologies can get up like properly when it comes to data applications.
So I think this is this is where definitely our customers want to go with the platform site. This this is consistent with their view of lots of data apps faster cheaper better and that's I think devops is the critical piece to solve as part of that job. What is it you wish organization's new going in about building a data Lake?
To last the scale for you know multiple business units that you see them kind of not appreciating or mistake that they make over and over again. That just makes you shake your head and go if they only know to do what If that's one thing right, I think that comes to mind for me it is. Starting off with a very concrete use case.
I'm getting that use case stood up in production all the way making money for you. That's the best way to actually deploy a platform. Right.
So you're just starting from the customer. The use case has to be solid. It's got to be something that is like impactful to the business.
It's got to have a timeline that's in weeks not in months or years and it has an outcome. It's very binary. You know, you got there.
What did you learn along the way I find that when companies embark on this journey to be data driven the ones that are really successful are the ones who have that bowling in use case knocked down and then they got the next one and the next one and the next one and there's a lot of momentum on how they're delivering. It's when companies look at platform first that I think that they get in trouble where they might be like, oh first let's all get all the data and the company in one place. Then let's build a data model across the company for what does a customer mean to us?
And these are like deep and relevant questions. But you're you're path to Roi is a lot longer than the first one that I described to you. So usually when we you know work with with customers who want to be data driven, this is we encourage them really to start from the use case in rather than from the platform out if that makes sense.
All right, folks. You heard it here first repeat after me. Do not attempt to boil the data ocean thing Ram.
Thanks for being on the chef. Absolutely, my pleasure. Thanks.
All right back to you guys and student.