Bridging the DevOps and DataOps Divide – Sean Knapp, Ascend.io
Ascend.io CEO Sean Knapp explains why more progress needs to be made bridging the divide between DevOps and DataOps in an era where applications are consuming massive amounts of data.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Sean Knapp, who's c e o for Ascend io.
And we're talking about where all this work around data engineering and data ops is gonna maybe intersect with developers and DevOps and how it all comes together, otherwise known. Hey, what's it all about? Alfie.
Hey Sean, welcome to show. Thanks for having me. So it's a big topic, right?
We seem to have developed, uh, pipelines on both sides of the equations. The data guys are building out their pipelines and managing processes and workflows through that. And they even have best DataOps processes that kind of look a lot and smell a lot like DevOps processes.
And then we have developers running around building code and applications that's supposed to invoke that data. But I feel like there's a gap and we don't necessarily have a, an easy way to kind of integrate these things just yet. That winds up making it easier to build an application that consumes data.
So where are we on this journey and what do we need to kind of close that gap? Um, well, whether we like the answer or not, I think we're pretty early on this journey. Uh, I, I think we got ways to go.
Uh, I think the overall industry is still in the sort of early era of mass explosion of technology, emergent best practices, uh, when it comes to data, data engineering, data pipelines. Um, but what we do see, and I, I think what's really exciting is we are starting to see the convergence around some best practices in the data realm that, that we can even call around data ops itself that mirrors, uh, a lot around the DevOps evolution. I think there's, we're on a, a time shift and, and so we're, we're, we're trailing behind from a, a maturation perspective.
Uh, but I do see that, uh, data ops and data engineering teams in general are behaving more and more like software engineering teams, uh, to, in a good way. And I think we're just, we're early. Seems like every time I turn around what we used to talk about as terabytes as being a massive amount of data, as now people are like playing around with petabytes and stuff, and I'm like, well, how big is too big?
Yeah, I, I think, so I do think that scale is a challenge, but I would actually contend that over the course of time it's a different kind of scale. It's becoming a challenge for us. And, and I would contend similar in the, the evolution of, of DevOps.
DataOps is now trying to tackle a different kind of scale. So to, to first touch on your point of, you know, petabytes of the new terabytes, uh, in, in many ways the, as we've seen the industry able to move up stack and really phenomenal, uh, innovation and, and scaling of modern day cloud data. Cloud vendors like Snowflake and Databricks, we're seeing that they are taking off the table the, the pains associated with sheer scale of data.
We run an annual survey, uh, every year and year over year over year for the last three to four years, we've seen a systematic decline and the number of respondents who say they're struggling to handle the scale of the, the number of records, the size of the data itself. So we're seeing that drop actually, which is great. What we see climbing, uh, on the other hand is dealing with the scale of complexity.
Just as we saw through a lot of the, the modern day software movement. Uh, we enable so many more people to write so many more software products. We end up with more people able to write more software services faster than ever before.
In the interdependence of all of those, uh, makes it increasingly complex. We get this exponential expansion complexity. And in many ways I, I love that to simplify things down, but I think of the entire DevOps movement in many ways boiled down to how do we enable more, uh, software teams to, and more individuals to write more software faster, but safely.
And that's really what a lot of it boiled down to. And that's why I think in the, the modern data era that we're seeing is that same thing. We've given this tremendous amount of empowerment to all sorts of new bi engineers, analytics engineers, data engineers, software engineers, turn data engineers.
And, and so there's so many more people. They can write data pipelines and data, create data products faster than ever before. And on top of all of that, the interconnectedness of all those systems, it's like a, if you think of a, a pipeline no different than a microservice, you get all these interconnected dependencies.
So you get this massive explosion and complexity. And that's where we see that the, the next big scale challenge and the importance of DataOps is how do we enable people to build so many more of these data products faster and safer than ever before? And I, I think that gets solved in really similar ways to, to a lot of the DevOps challenges, which are better normalization, better productization, better automation, uh, over the course of time.
Speaking of managing data at scale, we're seeing all these AI models start to run around. And the question then becomes, well, if we were having a hard time figuring out how to manage data at scale before, how are we gonna cope with all the data for the AI models? 'cause these things are, shall we say, on the massive scale.
Yeah. Uh, but don't worry. The AI will manage the ai.
We're gonna be good. Um, I'll be on a beach sipping a my tie, so I think we should be okay at that point. Um, no, I think in all seriousness, it it is a, a bit of a challenge.
Um, again, I think a lot of the, both the hyperscalers and the data clouds are continuing to invest in these really remarkable capabilities to, to bring access to modern generative AI and ML capabilities more easily. I still think the data engineering aspects of creating these models gets increasingly hard and, and cumbersome. Um, so I, I do believe that, that we're at the really, really early stages of it.
Uh, a lot of the focus at Ascend has been not just, Hey, how do we help you build these highly automated data pipelines to fuel, whether it's an ML model or your BI dashboard, both require highly durable, highly reliable, uh, data pipelines. But one of the areas where I think the data engineering realm will benefit greatly from is using Gen AI in better managing and automating the tools we use ourselves, uh, to create these products. And so I do think there is a, um, a circular benefit there, uh, to better embracing ai.
Should we at some point put the DevOps engineers and the DataOps engineers in the same room together and kind of say, maybe if you guys could figure this out, we'll go from there. But it seems like this is as much a cultural issue as it is a technical issue. It is.
I I, you know, we, we continue to see cultural challenges and, and oftentimes we, we see, especially in, you know, what I'd call the, the most bleeding edge innovative, uh, companies, oftentimes their earliest data platform and data engineering teams were the DevOps teams. They were the people building out a lot of the tooling and the frameworks to help embrace and adopt this new technology. Um, we see oftentimes even see that they're the same people, uh, if not at least the same teams who are helping with that.
Um, I think the, one of the, the fundamental challenges that currently happens is that there's a bunch of different constituencies that, that all participate as part of the data lifecycle. And so you have data engineering teams, you have analytics engineering teams, which are really becoming the citizen data engineers, if you will. Uh, and then you end up with the, the bi uh, developers and, and engineering teams as well.
But pulling these folks together, I think if we can get some, uh, if we can get teams together from the DevOps background and really spend time thinking through what does that unified layer that, I don't wanna call that a measure of fabric 'cause that's, those are are ambiguous terms now in, in the data realm, but what does that layer that get, they can bring in different teams and different individuals with different highly valuable skill sets where they can still collectively build and, um, contribute together, uh, playing to their strengths. That I think is a, an enablement strategy for any large team or organization that will pay tremendous dividends. I'm old enough to remember that there was a time when you brought compute to the data and then this cloud thing started popping up and people were bringing data from all over to the cloud, and that kind of was the opposite of where we are.
And now I wonder, I look at the edge and we're trying to figure out how to, uh, process and analyze data at the point where it's created and consume. So maybe we're going back full circle here, but do we need to be smarter about where we process data and when, Um, I do think that we're going to find federated processing of data, uh, is going to become an increasingly large thing. Uh, I think there's very real world limitations.
Uh, one, you know, the laws of countries that say you can't take this data, at least in this current form outside of the, the geographical boundaries. Uh, and then I think there's also the laws of physics, which say this data volume is too large, oftentimes iot of what data to try and actually stream every raw piece of sensor data in real time somewhere. Uh, and so I think there's, there's some fundamental limitations there.
Um, I also see, and we see this across a lot of our customer base, is increasingly really large organizations are looking to run this data mesh strategy where you do want some of your data and your processing your compute to run in one cloud versus the other cloud versus your on-prem data center. And, and you actually want geographical density there, uh, which I think makes sense for a lot of of different use cases. And that's why I think the, over the course of time we're going to see the architectures really start to emerge that are more flexible that it, it's largely we've, I think we've largely seen that it's really hard to just say, I'm gonna just consolidate all of my data into one lake or one warehouse.
I think some have been able to figure that out, but, uh, to me, this is no pun intended, but it's always like trying to swim upstream where the amount of data being produced and the number of things that are producing data and the number of people who want to consume that data is just growing so much faster than a, a centralized organization's ability to say, let's get it all in one place. Mm-hmm. I think that's what's really paved way for this whole data mesh movement.
If you'll We're living in some interesting economic times, do you think that the economic headwinds are getting people to think through the cost of all that data? Because I think we've ignored that for a little while, but it seems like, you know, that's all coming home the roost as well. Oh, absolutely.
Uh, literally just before, uh, this, I was on a call with a customer who, you know, they've been spending a lot of time rationalizing, uh, their, their data sets all the way up to what is the dollar cost for this dataset, and does it produce enough signal and their machine learning models that it produces enough revenue lift that it actually justifies the, the processing cost? And so I think we're starting to see rationalization really of twofold. One is rationalization of different data sets themselves, and does it make sense to process all of this data and does it produce lift for the business?
The other one that I think we're increasingly seen, and it's a bit of a, um, uh, just the natural pendulum swing back from the, the crazy buy every technology and download every open sourcing thing you can, uh, from 20, you know, 20, 20 to 20, late 2022. But what I, I, I think we're also seeing is this rationalization of, Hey, I played with a lot of these new technologies and working on, grabbed a ton of this stuff off the modern data stack, and now I need to rationalize those investments. And so what we're starting to see is in these economic headwinds, most technical leadership is pressing hard on, we have to get value created from these, uh, data investments, and we need to produce valuable data products that fuel the business and fundamentally move the needle forward.
And so I think that's a, a great thing that we're seeing currently happening in the market is, um, short teams are being pressed to put points up on the board. That's great. One other metric that maybe people are paying more attention to has to do with sustainability.
Are we looking at the, how much energy it costs to run a data set, per se? And is that gonna factor into people's IT strategy? Or is that just more, you know, it's political issue of the day, but it's not something we're actually gonna monitor?
Really good question. Um, I would say I, I would readily admit we don't see that concern much from our customers yet. I, I think it happens at probably a, an even more senior corporate level than, than oftentimes than the, the data teams.
Um, I expect it will come, uh, I see from the, the clouds, uh, creating a lot more visibility as to the, the environmental impact of the, the workloads that you're running. Um, I think again, kind of on that classic, we think of a, a maturation cycle of organizations. The, this probably is a terrible analogy, but we're still kind of in the like old dirty coal burning era of, uh, of data engineering.
And I think we're, we're starting to get more, a little bit more sophisticated and refined, but this is in the, the early eras of folks are making big strides forward on producing the outcome, but it's not always the most efficient. So all that said, you know, what's your best advice to folks right now? What should they be thinking about?
You know, where do you, because you know, if you look at all the data and how it's processed and the whole thing, it's pretty overwhelming. So where would you kind of focus? Uh, I'm a really big believer that that teams should continue to invest in automation as there's, there will be no slow down in the, the need and demand for data products.
Uh, I think that the business need for, for data pipelines and data engineering is gonna continue to climb and, and demand around generative AI is only gonna contribute to that. The what, what is going to end up going hand in hand with that is the expectation that teams are getting more efficient. And I think that's gonna prove to be a, a pretty big challenge for folks, uh, without leaning more aggressively on just raw automation and more advanced technology as it's, it's just simply too difficult to scale.
Uh, and for just doing bespoke pipelines and, and bespoke engineering, to me it's, it's like watching those telephone operators in World War ii or just constantly plugging things in and, and that doesn't scale when you're trying to support organizations with hundreds, if not thousands of data products that are constantly flowing through. That's a, that's a software game, not a, a manual operator game. First, somebody talks about how often they wound up in the wrong party line, but you get the idea.
Alright. Yeah. Sean, thanks for being on the show and sharing your insights.
That was great because, you know, we live in interesting times here and we talk about building apps, but we don't always think about, well, where's the data gonna come from? The feed, all those apps, right? Yeah, exactly.
All right. That.