GenAI and On-Premises Data Lakes with Starburst’s Justin Borgman
Starburst CEO Justin Borgman explains how the rise of generative artificial intelligence (AI) will lead to more investments being made to create data lakes in on-premises IT environments.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Justin Borman, who is CEO for Starburst, and we're talking about, well, the resurgence of On-premises.
It, and well maybe never left, but we're certainly paying a lot more attention to it these days. Justin, welcome to the show. Thanks for having me.
A funny thing seems to have happened, I feel like for the last 10 years we were focused on, um, moving data into the cloud, and now folks are saying, well, um, maybe we should actually be bringing compute to the data, which is how it was when I first started out in this business. So have we kind of come full circle in what's driving this conversation? Yeah, a little bit.
I, I think the truth is there's always going to be data in, in both places, if you will, both cloud and on-prem, but I think you're absolutely right that there's, uh, a renewed focus on those on-prem assets, and I think there's one big motivation for, for what's bringing that to surface to the surface. And I think that's AI and AI workloads and, you know, if AI is this incredible new engine, then data is its fuel. And it turns out a lot of that fuel is still a on-prem to begin with.
But b, from a TCO perspective, from from the cost of processing perspective, there's a pretty compelling argument to be made for, you know, hosting a lot of that yourself in your own data centers. And so I think that's, I think the future's a hybrid world to be clear, but I do believe that AI is bringing renewed focus to, to on-prem. I think you're absolutely right.
I don't think all this data exists in isolation though. And we have the concept of a data lake, and it's been mainly up in the cloud, but is the future of the data lake gonna be more federated where I can have an understanding of what data is resigning where, and what relationship it has with each other, regardless of where it's physically located? Yeah, absolutely.
I, I think that's exactly right. I mean, again, in the context of ai, your models are only as good as the data that you train them on. And so having access to more data is going to create a better model and allow you to, you know, train that model on the most relevant proprietary data that you own as a business.
Um, and that's gonna create competitive advantage. So having access to all of the data within your enterprise is gonna be a huge strategic advantage for you, uh, and give you a lot of flexibility. Um, so that's certainly the way that we see the future playing out.
Data lakes are great for storing as much data as you possibly can, just because again, the economics make it very attractive. You can essentially use relatively low cost object storage and open formats to store your data in those lakes, be they on-prem or in the cloud. Um, but being able to connect to all these other data sources is gonna help you enrich, you know, the analysis and, and the machine learning that you're doing.
I also feel though there's always a fine line between, uh, collecting data and data hoarding. So how do we kinda figure out what the right balance here is? Because a lot of times we have massive amounts of data and from years and years and years, and nobody's really using it.
Yeah, well, there, that's certainly true. I think, again, economics play a role in this. If you can store that archival data at a low price point, then you can potentially keep it forever.
And, and so I think that's also what's driving renewed interest in data lakes versus traditional data warehouses. Um, you know, traditional data, data warehouse like, uh, let's say Teradata or, uh, maybe Snowflake in the, in the cloud, um, you know, our proprietary database systems that, uh, are quite expensive. And so you wouldn't wanna store, you know, 30 years worth of data in a traditional data warehouse that would probably be cost prohibitive for you, but storing it in a data lake, again, on relatively cheap storage and storing it in these open formats that really you own and you control, uh, allows you to store more data, more cost effectively.
And so that's the way that we see a lot of organizations going. Of course, you know, the, the folks that pioneered this movement, you know, 10 to 20 years ago were really the internet companies that, uh, achieved a level of data scale greater than the rest of the world at the time. I think the rest of the world is catching up though.
And, and, and hence, you know, the need for, uh, a more data lake centric, uh, model to, to how you manage and analyze your data. Does there need to be a conversation between the business units creating the data and the IT folks that are storing and managing the data? Because I feel like there's been this disconnect going on forever and ever, and the business units just keep creating more data.
And a lot of times from an IT perspective, you know, the data types are different, but they don't have a lot of insight into what data really matters more than other data. So, mm-hmm. Um, do we need to have a growing up conversation about data management?
Yeah, I, I think that's an excellent point. I I do think those, that relationship between the tech and the business needs to continue to, um, become closer and closer. That context is very important as, as you mentioned.
And certainly that was one of the observations I had, you know, pretty early in my career, um, you know, selling software into, you know, data organizations was, I, I was surprised by how often, you know, the data teams themselves had no idea actually what the real use cases were, you know, for, for analyzing that data. And I think that's just, um, putting a finer point on, on, you know, your comment about that disconnect. Now, having said that, I think there are, uh, movements underway today to help remedy that.
One of those is this notion of data products, which is to say, um, that the business owners, the, the product owners, if you will, uh, should be playing a, a role themselves in the curation of high quality data products to be consumed by others in the organization. And I think if you take that type of, uh, mentality of sort of data products and data product ownership, uh, you're gonna create a closer relationship between, uh, it and the business where the business can really enrich that data product with the context that they know, uh, including a whole host of business, you know, metadata, uh, about what that data product is for, how often it's refreshed, um, what types of queries you might run on that data. Uh, and so we see, uh, increasing momentum for large organizations in particular to, uh, treat their data sets as, as really first class products, um, for others to consume and, and therefore, you know, bringing those relationships closer together.
Is there some tension in that equation with the security folks though, who are often running around going, get rid of as much data as you possibly can, because that's the stuff the bad guys are after? Yeah, yeah. So I mean, access controls are critical to all of this, right?
And having very fine grained, uh, access controls, be they role-based or attribute based, where, you know, you can tag sensitive or PII data where you can do, you know, role-based column-based, uh, uh, or, or mask data as well. Um, you know, like social security numbers and, uh, you know, bank account numbers and things of that nature. Uh, I think, you know, access control is essential with, with this, especially as we try to move towards democratizing greater access within an enterprise.
Uh, if you're gonna have a single point of access to, to all of your data, you definitely need to have, uh, very good, uh, access controls. And so that's one of the biggest investment areas in our product, uh, that we bring to bear with, with customers who work with Starburst. Are we getting better at automating this data?
And I can remember back in the day, you know, 20, 30 years ago, you know, people were talking about tiered storage and automatically moving data from, um, the hot area to cold, but I felt like we never quite got there. And so how sophisticated are the management capabilities inside of Data Lakes these days? And can I move data to the right place based on its access and what it means to the organization?
Or is that still something we're working on? Well, it, we've made progress. You're right that in, in many organizations, it's still, you know, something that data engineers are scripting themselves in terms of, uh, when they're gonna batch load data and, and, and you know, what data they're gonna load.
And doing those transformations and creating those, those, those new tables and so forth. Um, but where we've seen progress, and this has been an area of focus for us, is we're seeing more and more customers adopt, uh, a particular open format called Iceberg. And Iceberg has sort of like won this format where war where it's established itself as the defacto standard.
And what we're doing, uh, with, with our product is, uh, trying to make it very easy to get data into these iceberg table formats, um, by, uh, connecting to, for example, a Kafka stream and streaming that data in and ingesting it into this iceberg format automatically. And then adding retention policies and doing data profiling and collecting statistics, and automating all of that process to try to create a very seamless turnkey, you know, iceberg lakehouse, if you will. Uh, or as we like to say, an ice house, um, you know, by, by doing that and, and, and automating that.
So to us, that's, that's the current frontier. That's where we're trying to, you know, make these lakes as easy to use or even easier than a traditional data warehouse. To your point, we are, um, streaming more data lately, and, and I can't help but wonder if that's gonna become kind of the, the dominant way we think about, um, and processing and analyzing data in flight versus when it gets stored somewhere else.
The history of it is largely around batching and processing, but are we moving more towards, uh, processing and analyzing data at the point where it maybe is created or consumed, or at least as it's moving between points? And has that changed in the way we need to think about managing data? Yeah, I, I do think that, uh, streaming data is becoming a more popular way of moving data and, uh, you know, a more popular data source.
I think like the popularity and prevalence of Kafka is a good indication of exactly that. Um, but I also don't think it's an all or nothing, you know, zero sum game either. I do think there will always be a good place for batch, um, uh, you know, for batch transformations, for example, uh, you know, data quality processes that you're doing that, that may make sense from a batch uh, perspective.
But, uh, certainly to try to get closer and closer to real time, which many organizations are, you know, for certain use cases, streaming, uh, offers a lot of benefit there. And we do see, particularly a lot of digital native companies, you know, that are maybe born in the cloud or, or, uh, you know, their business is very digital from, from the start, uh, employing these, these more streaming oriented, um, data architectures for sure. What Is the future of data engineering look like?
Because, um, back in the day, you know, we had extract, transform, and load, and we didn't call it engineering, and now it's like much more programmatic and it seems hard to find these people. How automated can data engineering ultimately get and, and where are we on that journey? So I think the goal is to make data engineers as productive as we possibly can, and that really means automating to enhance that, that productive output.
Um, automating the tasks that are tedious, that are not necessarily adding value. Um, that can be both on the transformation side, uh, as well as, um, you know, ingest like we spoke about, right? So, uh, the more that we can automate there, I think the, the easier we can make their lives.
Um, you know, we've also invested a lot in the cluster management aspects, so we really have two products at Starburst. We have one which is called Starburst Enterprise, and that's self-managed, where the, uh, data engineer is responsible for running the cluster. Uh, and that has a lot of flexibility, especially for deploying an on-prem like we talked about at the opener.
Um, but we also have a cloud offering, and that's called Galaxy, and Galaxy manages the cluster entirely for you. So, uh, that's really taking the work away of, you know, how many machines do I need to run? How are, am I configuring them for a particular workload?
How do I scale those up and scale those down? You know, there are just a lot of, uh, man hours or, you know, person hours, uh, invested in those types of cluster management activities. And, and so we're trying to abstract that away, take away that work so that those data engineers can focus their time, uh, elsewhere.
I also remember the rise of the chief data officer, and I wonder if we still see that trend, or has data just become so critical to the organization that, uh, you know, we're just managing it better collectively, per se, within the context of our existing job functions. Yeah, we still see the, the chief data officer being a, uh, important function in organizations who do decide to, to hire someone in that position. I think it basically says to the organization that data is a priority here, uh, and becomes maybe a, a a forcing function or a a means through which to create focus around the data strategy.
But we also see maybe a, a varying, uh, scope for those, those chief data officers. Some of them are more focused on the data consumer specifically, you know, the tools, the data science, uh, you know, elements. Others, uh, also own the data infrastructure and the data management.
Um, and so different organizations have, have deployed that differently. I think the key thing is just making data first class citizen and recognizing that, uh, again, data is probably the most valuable asset outside of the people that you employ. Um, so let's treat it as such.
And, and I think that, you know, movements like data products, like we were talking about before, uh, are a good, um, you know, investment in that area. Uh, and, you know, sort of, um, clearly defining your, your data leadership, uh, is another, uh, great investment in that area. So Ultimately, what do you see among organizations that are really managing their data well?
What are they doing? What are the patterns? What are the things that differentiate them from everybody else?
I think first and foremost, they're trying to create architectures that can stand the test of time. I think that's what every data leader should be focused on, is how do I build something that will adapt to the changing needs of the business as well as the changing technologies. Um, for us, this is one of the reasons why open formats, uh, open engines, open source products in general, uh, are particularly, uh, useful because you're not locked into a particular vendor.
Uh, and so switching and, and adapting your architecture, uh, is generally easier, but also building an architecture with various layers of, of abstraction. You know, in the case of Starburst, we are functioning as a, uh, a query engine that queries the data that you have. Um, but that allows you to change what data sources you're storing that data in.
You know, I mentioned Iceberg is a very popular format today, but you might choose Hootie, which is a different format, or you might choose to store some data in an Oracle database or SQL Server database or Postgres database. And having that flexibility of having a common query entry point, but being able to change your data sources gives you adaptability to, to changing environments. So to me, at the end of the day, um, people should be thinking about how to create optionality for the long term so that, you know, you don't get stuck or, or encumbered or held back by the technology choices you made maybe five or 10 years ago.
And I think that's, that's where I believe we're at in history is people have sort of learned a lot of the lessons from maybe 20, 30, 40 years of database history and are now, you know, getting to the point where they're, they're saying, Hey, you know, we can't put all of our eggs in one basket that's gonna hold us back, you know, for what the future looks like five years from now. And so we need to build these more flexible, modular open architectures. All right, folks.
Well, you heard it here. No matter what computing era we're in, it's all about the data. Hey, Justin, thanks for being on the show.
Thanks for having me. All right. And back to you guys in the studio.