DataOps, Observability and Data Journeys | DataOps Day
Data professionals live in a complex, chaotic world. The modern data stack is too complex. New cloud data toolchains are fragmented. Data architecture patterns are diverse and complicated. And data itself, of course, is diverse, always multiplying and forever changing.
On top of all this, the process of ingesting, storing, transforming, predicting, visualizing and governing data is highly distributed across various people and departments in your organization.
It’s no surpise, then, that data professionals’ jobs are chaotic and stressful. There’s a constant fear that somewhere along the journey data takes from source to value, suddenly, everything will break, and you will be left holding the bag.
Something is missing from our data systems; the ability to identify, measure and quantify expectations versus the reality in production data systems. What is the variance between what is happening now and what should be happening? Is it on time? Late? Is it trustworthy? What is happening now? Will my customers find a problem before we do?
That missing piece that connects data system expectations and reality is a ‘data journey.’ In this session, Christopher Bergh explains what the data journey is, where it’s broken and how organizations can use DevOps principles (DataOps) and observability to make sure data is truly driving an organization’s success.
Transcript
Uh, welcome everyone. My name is, uh, Chris Berg. I'm c e o of a company called Data Kitchen here in Boston, Massachusetts.
And I'm gonna talk about something called the Five Pillars of Data Journeys. Uh, and I want to thank the sponsors for inviting me to DataOps days. I have a long history with DataOps.
I wrote the first, uh, Wikipedia article on data ops about seven or eight years ago. I wrote the, uh, or was the key author of the DataOps Manifesto, which, um, some of you may, may have read. And so we're gonna talk a little bit about that today.
Um, and the first thing is that we're gonna talk really about the first step in doing data ops, which we, we call the data journey. Um, and I'm gonna focus on that and we're not gonna talk about a lot about the other parts of data ops as a whole. And to go through that idea of data journeys, I'm gonna talk about the five pillars.
Um, I'm gonna give a customer example. Uh, I'll give a little quick shout out to my company's products, and then we'll end up with some resources that you can use to, uh, to learn more about data journeys and data ops. And so, um, you know why, uh, you know, I'm a technical guy.
I spent sort of 15 years, uh, building software at companies like M I T and NASA and Microsoft. And then, uh, actually 17 years now, uh, ago now, I did, um, went into data. And, you know, data building data and analytics systems are hard.
Most of them fail. Most of them are just full of errors. Um, and customer data trust is almost at an all time low.
Um, and people who work in data and analytics are just incredibly stressed and, and, and want therapy. And, and we've got actually, uh, surveys, statistically relevant surveys to back all those things up. Um, and so that's not a great, uh, uh, a great situation.
And so for us, we have a very different view of the world, um, that it's not really about the tool that you use. It's not about a better database or a better e t l process, uh, e t l tool. It's really about the system that people work in.
And the things that we wanna affect in that system is, number one, helping you run things with low errors in production so your customers don't blame you when things go wrong. Um, and then help, also help you put things into production quicker, because the faster you get feedback from the customer, the more likely you're to learn exactly what they want and the less things that you're gonna waste. Um, and, and so that's really in some ways the business value of data ops.
Um, but we've come over the years have really said, let's just focus on one of these things. And, and we've come to this, uh, first step in, in the data journey, first data ops, which really means just focus on your production systems and just focus on making sure the data is right, the integrated data is right, the dashboards are right, everything's on time, um, that you're running a really good factory that does production. And so a lot of that talks about, that's what this talk is about.
So it's not about data ops in general. Uh, I, you know, we've written kind of two books on data ops. io.
So if you're interested in the overall idea of data ops, but just after working with dozens and dozens of customers, most people don't have this, uh, handle. Um, and focusing on this first is the good first step to talk to, to start. And so why is that?
Well, I guess the first thing is little boxes everywhere. Like we build these analytic systems in the cloud or with tools, and then there's just lots and lots of little boxes and little tools. And so looking at any system that has a lot of little boxes, you can pick things that go wrong.
And so here's, uh, a slide from, uh, Andreessen Horris that shows all the little tools or an example pattern of tools. And, and likewise, if you go to all the cloud vendors, well, they all have tools and, and now, uh, you could probably add Azure, uh, you could add Snowflake and Databricks. They also are trying to have every tool into the sun.
And everybody wants to be your one-stop shop for every tool that you do. And, and that's never the case, right? Because people have as-built systems, they have existing tools, it's just tools and tools everywhere.
And then there's these design patterns that people put them together in tools, right? We've got hub and spoke patterns, we've got streaming versus batch, we've got data mesh, we've got producer, consumer, um, you know, sort of hub teams and, and self-service. And there's are these design patterns of all the tools.
So if you start, start to think about how data goes from where it comes in your company, all the way to where it gets delivered in a dashboard, in a application, or in a, uh, data export for your customers, it, there's a lot of patterns going on. And plus the complexity of the systems themselves that, that we build is that, uh, a lot of people are doing a lot of this new term of analytic engineering with D B T. A lot of people are creating a lot of tables, uh, just like people created a lot of E T L programs or a lot of sort of, uh, backing stores behind dashboards.
And so there's a lot of, there's just a number of D B T customers have 5,000 tables, which is just incredible. Um, and, and so the, the problem with lots of little boxes, the problems with complexity in those boxes is that you have an across problem, right? You have to figure out when something goes wrong and your boss or your customer says, this looks weird.
Um, you gotta go across all these tools and find out where it is, you know, when it's, if it's in batch and streaming, if it's in your e t L tool, if it's in your vis tool, if it's in the model, if it's sent to some other system. And, you know, you've got multiple tools, multiple dataset, multiple paths, multiple methods, multiple architectures, multiple customers, um, and multiple people looking at it. So it's just complicated, right?
And, and so that's why, um, a number of people, and, and like myself in, in 2007, I'd have the morning dread. I'd dread if someone, um, would find problems in what I'm doing. 'cause I'd have to spend the rest of the day chasing it down.
And number one, you've got a lot of little boxes, and two, you've got a cross problems. And then number three, you've got a down problem, right? Because all those little boxes are made up of many things, right?
You can have, um, the software that's running in them, you can actually have the software itself. You can have the server it's running in. You could have the tool, you could have the thing that's orchestrating it all.
Um, some people have, uh, data tests. And so this across and down problem makes a big challenge. And so at the end, that's really why we started to put this idea of a data journey together, because data journeys are problematic and there's too many errors, and those errors are causing too much time and, and hair pulling for mostly everyone who does data as, as, as work.
And if you could, uh, fix those first, um, you'd have a lot more time and honestly, a lot more fun doing your job. And, and that's what this is about. In the next, uh, section, I'm gonna gonna walk through this five pillars of data journeys.
And so this is really, um, meant to be a best practice presentation. Um, of course we have some software that can help, but there's other software that can do this as well. And so, um, the, the problem with unmonitored data journeys is that things go wrong and reducing errors on the path that data takes.
So if data comes from, uh, a system that's outside your company, so maybe you're analyzing website traffic, so it comes from Google or some third party, you're putting it in, uh, maybe a bucket store, you're putting it in a database, in another level, you're having an e t L tool, an ingest tool. Maybe you're having a data prep tool, maybe you're visualizing it, maybe you're running a model on it. Maybe you're reversing EL ing it.
So you've got where is the problem, right? And is it the server? Is it the database?
Is it the data itself? Is it the tool? Is it the code acting on the tool?
Um, and you, a lot of these things are, are hard. And I've had just too many teams who take, you know, large teams where the best people are spending all day trying to run with their hair on fire, trying to fix this. And, and the shame and stress is just, it's just not fun.
And, and I've always hated this. I've always hated having teams that I've led and having to answer the phone call with some irate customers sort of threatening to cut me off or berating me because you know, the data's wrong. Um, or the worst case is the data's been wrong for three months and you've kept, you've, you've said it's good and it's obviously been wrong and, and you're first learning about it.
So the this, um, is meant to solve that problem. And so I'm gonna talk about these five pillars, one at one, and, and I'm gonna go through each piece. And so, you know, the, the, the first one is I'm gonna talk about it sort of across the steps.
Uh, then I'm gonna talk about down the stack down your tools. Then I'm gonna talk about data at rest. Um, and then I'm gonna talk about how that data is used, and then I'm gonna talk about setting expectations on all of it, because, uh, your journey is really a, a set of expectations on what should be.
So the first step is, is, uh, we're talking about across the steps. And the thing of it is, is that your tools run in order. So you may have data land, it goes into a level one in a database, it goes into a level two, it goes into a dashboard, maybe into a model, maybe into an export.
There's an order of operations that has to happen and a schedule. So you have an assembly line that data goes through to deliver value. So you have to monitor the order and the steps in that assembly line.
And when things go wrong, the first thing is like, which piece went wrong? Um, and you can see in the lower right here, a picture from one one of our software products saying something is a warning in, uh, Azure Data Factory. And so really it's about the entire process reliability.
Um, and so that's, I think, uh, an important part. And for us, we think that monitoring this assembly line that produces insight is also about monitoring the assembly line, but it's also about monitoring what goes on inside the assembly line. And that's why you have to kind of go and think about, um, across the steps.
And so across the steps could be from one of those tools, lots of tools acting on data, and we talked about that they could be that you just got lots of journeys in your organization, it's across the steps, but there's not one data journey that, and they're everywhere. Um, you know, companies will have hundreds of interconnected data production systems, data assembly lines going on, and those data assembly lines are sometimes broken out into who owns them, right? Um, organizations and teams, um, and hub and spoke.
And then there's this, uh, law called Conway's Law that says, when you build something, you should kind of follow the natural contours of the technology. And what Conway's Law says, that doesn't really happen. It's actually broken up by how the organization is, and it's meant to be, um, uh, a derogatory term.
Um, but that's really the case in a lot of organizations. You have hub and spoke, you have one division and another, you have producer, consumer, uh, you have data mesh. All these different ones are really organizational patterns that end up in how the data journey's constructed.
So they're composite. There's not like one uber journey that everything fits into there. These, uh, assembly lines, data journeys, uh, relate to one another.
And so lastly, the, how this works, timing, relationship order, uh, are kind of tribal knowledge in organizations. They're just not in anyone's head. And so when things go wrong, you have to start to pick up the phone.
And who knows how this, well, where is it? Who's, who knows how this works? And so I talked to, uh, recently the 15th largest company in the United States.
Uh, a an external report was empty. And the c e o of the 15th largest company in the, in the United States called up his data team to yell at them. And of course, it's a huge data team.
So 26 people spent all day trying to find out where it was, because everyone owned a piece. There was this tribal knowledge. And of course, it ended up being something completely innocuous that was fixed very easily, um, that, uh, you know, and, and therefore, 26 people of their team lost a full days of work chasing their tail trying to find out where this is.
Um, and so tribal knowledge and location, uh, where it is in the journey across is important. Now, you may find where a problem is, but you may actually not know where it is down the stack. So you could say it's in, something is wrong in your e T L process, but is it in the code that's acting upon it?
Is it in the data? Is it in the server? Um, is it in the C P U?
Is it in the ram or the disc? Like where is it in this? Um, and it's kind of a technology status because we've got these layers of technology that we build to support our little boxes everywhere, architectures and sort of drilling down into it, right?
Between the tool and the data and differentiating where these levels are down the stack, trying to find where the levels are. So we have a, and a cross and down problem, there's a problem. And then where is it in it?
And so that's why the idea of a data journey is this sort of composite thing of a cross and down. And because you want to correlate, right? You wanna say, well, there's a problem in a table that shows up in a tool that is caused by software that was scheduled by a schedule, by a scheduler like kron, that shows up in a data journey error.
And, um, finding out that correlation, looking for root cause looking for patterns over time. Maybe you have a bad disc drive or maybe you have, uh, some code in an E T L process that creates a problem under certain conditions. So that, um, we're gonna talk a little later about that process of, of this is sort of correlation across, down and across, but then there's sort of digging into root cause and that's also an important thing to do.
So the third pillar is kind of data at rest. And, and, uh, this has gotten a lot of focus over the years, right? 'cause we have data systems and, and so where the data lives and, and is it good and do you trust it?
And data quality, I think it's, is really important. But it's not just about things like schema and freshness and in volume, um, and things, uh, does it fit a profile or vary from the profile sort of drift? There's a lot of terms out there.
There's the Dema five dimensions of data quality. The way I think of it is you've gotta test, um, data based on the syntax of the data. And you've gotta base test data based on the semantics of the data.
And, and I purposely use the word test. Actually, what I mean here is a automatic in production data quality validation text that checks row counts, that checks co compares previous and other versions, checks percentage growth, um, and these kind of data checks that death check tech data, check data at rest, and the integrated data at rest, because you could have perfect data from your data suppliers, but someone messed up the data integration or a joint condition happens in another table, and suddenly you've got a problem. So validating quality automatically either based on the syntax of the data or based perhaps in partnership with your data stewards, um, or people who know the business test that represents sort of the syntax of the semantics of the data, the meaning of the data.
I sales growth should be 50% quarter, you got 20%. Is that unexpected? Is that a data error or is that really the way sales are?
You got 200%. Is that, again, again, trying to differentiate what's, what's signal and noise here in having, in partnership with people who actually might know the business better? Um, and, and why is data rest a problem?
Well, I guess a lot of data engineers who focus on data at rest, they're just stressed, right? They have a lot of problems, a lot of unfound data en engineers. Um, and you know, we did a survey that 52% of data engineers, uh, said errors are a significant, uh, source of burnout.
And 78% of data engineers in the survey of 700 people two years ago. So they wanted a therapist. Um, and then data engineers don't really know how to test data.
You know, they, they're just, they're super busy. Technology's rapidly changing and they just don't have a concept of, or don't have the time to learn the concept of the business and how it fits. Um, and so, uh, these two problems make sort of testing and validating data, raw data and integrated data at rest of a problem.
And then, um, you know, what we think of it is, is sort of these three buckets on the left here, is that like looking at kind of profiling the data, looking at syntax level checking based on that profiling. And that could be, um, the schema of the data. It could be the freshness of the data.
It could be, you know, this column had three values in it. This column had only three types of values, and I suddenly got a fourth. Um, and then these sort of business rule, more semantic tests that I talked about.
Now the fourth pillar, and it is really data at use. And so that's where most people actually get at the data, right? They see it in a dashboard or an export.
Um, and so they actually see the results from a model. So somebody is using the data. Well, there could be a problem there.
You could have perfect data, you could have perfectly integrated data. Everything is a hundred percent right? But some model that someone's doing a prediction of suddenly goes wonky.
Or someone put a configuration in Power BI that makes the, you know, quarterly sales count wrong, and you're getting a call from the sales VP saying, my numbers are off, and, and you're having to chase that one down. And so looking at, um, these tools through APIs, through Python, talking to directly to the use and testing those is very important. So test the models, test the visualization, test the delivery and the utilization of, of, uh, of your data assets.
And so that, that's really kind of looking at it completely. You can have, you have data at rest and data at use, and you could just have problems everywhere, right? It could be because the files posted late or some automation didn't trigger, or there was data consistency or, um, data process failure or the dashboard didn't get refreshed, or the dashboard itself is wrong.
And so checking data at use and data at rest is, is sort of what we believe is an important part of, you know, trust, but verifying your data systems. And then the last thing is, you know, you're trying to look for problems sort of across the process, down the steps, down the stacks, data at rest, data and use. And so what happens when reality doesn't meet your expectations?
So I guess the first thing is, what are your expectations? Um, how do you know, where do you keep those in your data architecture? Um, and so we think the data journey is a logical place to set these expectations of what the world should be.
And so, um, you know, if you could look at this diagram, think of 'em as a data journey, as a thing that says it should be delivered by this time. It should have passed all these data quality tests. Things should have happened in this order.
Um, um, this server shouldn't have gotten over this C P U, um, this cost shouldn't have gotten over this barrier. And so all these sort of collection of expectations of running, um, even utilization, this should have been utilized by someone, a dashboard, all these things are ways to judge the reality and say, okay, maybe it, the data quality test didn't happen. Your row count was zero or your row count was a hundred.
When you expect a million, you should find that out before your customers see it and you should be notified. And to do that, you've gotta set the expectation and judge the variance between expected and reality. And that variance actually is a, a set of events.
Like if something's going, uh, down, you should be able to know, and is it trustworthy? Is it on time? And because the data journey itself is kind of a shared concept, like there's a lot of people who care, right?
Your data customer who's calling you up and asking, is it ready? Um, or can I trust it? Well, they care.
The person who's your manager who runs your data team, well, they're kind of in charge of the factory. Maybe they've got, um, their teams and, uh, contributes to 20 different data journeys. Um, maybe there's an individual data developer who's working on the E T L process of the model, he or she caress.
And then there's a person who is in charge of production running all this. And maybe they've got, they're running data and analytics systems, maybe they're running sort of your production website. They all care.
And this shared concept is sort of looking at the world judging. The, the expectations that are embedded in the data journey against it is, is, is all a source of contextual information. And I think this is, um, can be very productive once everyone's sort of looking at the same thing.
And then lastly, we talked about this is, um, briefly is when things don't meet expectations, trying to look for patterns over time, trying to sort of find a needle in the haystack, trying to find out what the root cause is, um, is keeping, is really a basis for analysis and trying to keep track of the analytics of what happened and sort of longitudinal data collection. Um, and the journeys become the context, sort of looking at journey expectation versus reality over time. So let's just talk about a quick, a quick customer example of someone who's done it well.
Um, and so this is a really interesting, uh, customer of ours that they do cancer cell therapy. So what they do is they take some blood, they ship it back to the company, they take this blood and look at it, blah, blah, blah, put some cool s**t back in the blood, cool stuff, excuse my French, back in the blood, send it to you and you get cured from cancer. Like, it's amazing.
Um, and so this process is real. Everyone cares about this, right? Marketing, sales, production, c e o.
And to do this, you gotta put together, the team had to put together 70 different data sources, right? And everyone's on it. So the, the velocity, everything has to be updated quickly, like within 30 minutes.
Um, and data integrity and is incredibly important, right? If things are wrong, order of operations is very important, right? Did, uh, where the status is?
Is it in shipment back? Is it shipment forward where it is along the line? And of course all the aggregates, how many went through today?
How many are in process today? Um, and so there's a very small team who built this whole end-to-end system, sort of three and a half people over a year, um, that did this. And it's all fully automated.
So the data flows in all hands off. And there's this layer of expectations along the top that says, is everything right? Did it, is it on time?
Is it on order? Is it meeting SLAs? Is the data right?
Is the integrated data right? Even is the visualization, right? Um, and all this sort of automation actually happens at every step of the process.
So, you know, there's just going through each column here, there's sort of a file demon that happens, a file processor, there's ingestion and data prep going down one column. There's operational analysis and data prep, and there's finally Tableau. Every step, every one of these red boxes on the bottom is there's checks and then there's alerts that go off.
So the system runs updates, data every 30 minutes, um, and literally it'll tell you if something's wrong. And so, um, that kind of automation, that kind of factory that runs kind of hands off is, is we think very possible. And if you invest in it, um, it, it really makes your team productive.
'cause I've seen other organizations where it's not three people in a year, it's 30 people in a year and they get less done. Um, and this is, uh, um, the value of kind of monitoring your data journey and actually even doing more data ops automation on this is sort of a full data ops case. Um, and so just to, uh, finish up here in my discussion.
So one of the reasons we set out to build Data Kitchen was, um, you know, our experience and the application of these sort, sort of lean manufacturing techniques and DevOps techniques are general across, uh, a lot of data companies. And I, I'm pretty glad to see that a lot of these things about observability and testing and using Git and having data scientists and data engineers act a little bit more, like more software and engineers and a little bit like running, uh, their factory aligned. These ideas are getting more and more common.
And so we built a couple of software products that does that, do that. One is we've got an observability software that actually implements this idea of a data journey. Um, we've got a test software that does testing at, uh, data at rets testing, and then we've got an automation product that does data at use testing.
And so, um, really easy, really fast to set up, don't have to spend a lot of time. You can get something, uh, going in in 60 minutes, um, and then start understanding your data journey and improving it. Um, and so this is what our products look like, kind of played out on that picture that I have before.
Um, and that's about it. And I think the, the full benefit of all our products is that, you know, people are incredibly productive. And that, that example down at the bottom, really the idea of DataOps is that teams are incredibly unproductive and there's just a lot of waste and weight and, and, and re re things that being redone, um, and things that are being done that aren't necessary.
And so by focusing on at the big picture, focusing on data journeys and reducing error in production, and then secondarily, uh, focusing on how fast you can put things into production with low risk. So you can do a, do little bits of work and learn, um, that actually drives a huge amount of productivity. And Gartner's published two reports this year.
If you do all these things, you get an amazing 10 times productivity. Um, and I guess we, we've seen that. Um, but most people have a lot of systems that are built already and, and we say start with first with the data journeys.
And, uh, in conclusion, uh, we just have a, a bunch of resources for you. So, um, we're a big believer in manifesto. So we have the DataOps Manifesto and the Data Journey Manifesto.
Um, we have two books about, um, kind of the DataOps Cookbook, which we've given away, I don't know, 20, 30,000 copies. Um, we've got a book on DataOps transformation, and then we've got a whole a set of, uh, discussions on data journeys and, and, and, and where they go. And so I want to thank you all for, uh, for taking the time to listen to me.
Um, but appreciate it. And, uh, if you have time, check out the Data Journey Manifesto and give it a, uh, a sign, 18 points. Um, uh, and that's it.
Thank you much.





