Modern DataOps: Expert Insights on Evolving Technologies, Trends, and Future Challenges – The CD Pipeline EP14
In this episode, hosts Alan Shimel and Lisa Cao are joined by an All-Star team of experts, including Barry Hart (Owl Bear Hart Consulting), Karn Wong (Data Cafe) and Dadisi Sanyika (Apple). The panel discusses their experiences around what modern DataOps has become, perspectives on evolving technologies, as well as future trends and pitfalls. DataOps is the intersection of DevOps and Data Engineering, which includes MLOps and AIOps. The discussion covers concerns such as pipeline orchestration, containerization, observability, as well as how the tooling should evolve in the space.
Transcript
Hey everyone, it's Alan Shimel for Techstrong, and you are watching another episode of CD Pipeline. For those of you who may not be familiar, the CD pipeline is a monthly, uh, video series, uh, done in partnership between us here at Textron Group and our good friends at the CDF part of the Linux Foundation. And, uh, we work with the CDF, some of their members and volunteers and ambassadors to talk about, well, to talk about anything we want really, but no, we try to talk about topics that are related to the, you know, to continuous delivery CD pipeline, hence the name and what's going on in the world of DevOps and cd.
So, today's episode is actually a really good one. It's called Modern Data Ops, or Data Ops. I always call it data ops.
We could talk to Tomato, tomato. We'll, we'll ask our panel what they, what they say. Um, but we're gonna jump into that.
And before we do though, let me, let me introduce our panel to you. First of all, he is a repeat, remember, he is always on our show. He is also, is it the chair at this point point?
I'm the chair of CDF. Yeah. Right.
Our, our, uh, friend Dei Ika. And if I mispronounce it de dc I do my best, but say it right for us one time. No, that was perfect.
That DC Ika and I am, I'm getting here. Chair of the Klu Foundation. So, Awesome.
We took me a year. So as you mentioned, you are the chair of the Continuous Delivery Foundation. Anything else you want to tell about your background?
Um, I've been, uh, working in continuous delivery, uh, building tools. I'm a tool builder. I'm also on the, I I've been working for about 15 years in this space.
And, uh, I am a, also on the Spinnaker, uh, TOC. Um, and I am a tool builder and, uh, someone who very passionate about it, just here in the space trying to do, trying to contribute. Excellent.
Next up, I'm very happy to introduce you to Karn Wong. Hi, Karen. How are you?
If you wouldn't mind introducing yourself to the audience. Hi, everyone. Uh, my name is Ka and currently I work as a platform engineer.
Um, I've been working in a data engineering and machine learning and develop SRE capacity, so I understand how, uh, these teams work together. So what I like to, what I like to do is that to help people achieve their goals and overcome obstacles in first along the way, either the, the, what they've known or unknown to help people reach their potential. Fantastic.
Thank you. Thank you for joining us. Next up, I wanna introduce you to Barry Hart.
Hey, Barry. Welcome and give people a little bit of your background. Okay.
Um, currently I'm a consultant, uh, do part-time work for a number of different companies, mostly small and mid-sized with, uh, MLOps and data projects. Um, I, let's see, I've got my start with data related work. Really back in the nineties, I worked with some companies that were doing, uh, it wasn't called data science then, but operations research Optimization.
So over the years I've worked on applications evolving, uh, vehicle routing, retail forecasting, and predicting, uh, cell phone signals, if you remember the, can you hear me now? Commercials, things like that. Uh, cell phone location data.
I worked, uh, in the, around 2016 to 2020. I worked at MailChimp for several years, which is, uh, email marketing application. A lot of interesting data work there.
Uh, since then, I've switched into consulting and mostly have one, uh, customer that I work for about half the time, and then occasionally take on other customers. Um, I really like, uh, email work. I like working with data scientists because I find that they're a little more freewheeling, I think, than, uh, maybe sometimes we engineers can be.
So I think there's a lot of creative, uh, potential, I don't wanna say conflict, because it's not conflict. Um, also I wanted to do a quick shout out for SQL Fluff. Uh, that was a SQL Linting and formatting tool that I, uh, was a heavy contributor to for the first few years.
Um, less so now, but I'm still a, a core maintainer. I think it's pretty widely adopted. I just wanted to mention that.
Fantastic. And thank you. Last, but certainly not least, I wanna introduce you to a woman making her debut as our co-host of the CD Pipeline show.
She's filling some big shoes of my friend Lori LaRusso, but I, I don't wanna say she had big feet, but, but I think she can fill the shoes. She's the, uh, she's a CBF Ambassador. It's Lisa Cao.
Hey, Lisa, welcome and thank you for stepping in as co-host here with me. Give people a little bit of your background. Yeah, thanks for the introduction, and I'm, I'm so happy to be here.
We'll, miss Lori. Uh, so my name's Lisa. I lead product at an organization called Data Straight Out.
We work on an open source data catalog. And so a lot of what my background has been in is looking at metadata and how it can enable all of these different frameworks and bridge, you know, DevOps teams, MLOps teams, and data ops teams all together. And I think it's a really exciting time right now.
There's tons of new open source technologies being created and being adopted across worldwide and in all the different organizations at all these different sizes. So this is gonna be an amazing discussion, and I'm so excited to to hear everybody's takes. All right.
Welcome. And it's great to have you on board. So let me, I, I mentioned this when we, we, we came out, today's shows on modern, I call it data ops.
Other people say Data ops. Let's consensus on the panel. Is it data or data?
Data, Data, yeah. We got one for data. Did see, yeah, I, I I say data too.
So data also, Darren, what about you? Not in my language, we say da ta, but like in English. But it was like, really with data, Lisa, Definitely in data, but I use both interchangeably.
I, I have no loyalty. I, I often get, like, I, I'll say both in the same time. So I, I cover all my bases.
Um, but it, I guess it's a tomato to model, but when we talk about modern data ops, right? Well, today, modern, I think has become a code word for ai, right? Because everything modern seems to have either some sort of AI or automation or, you know, something like that built into it.
And, um, I'll, I'll let, let, I'm gonna kick it off to the panel just with that, when we talk about modern data ops, right? How, what, what, what's, what makes it modern, right? com, but data ops has been a term that we've, you know, seen rise over, let's say the last five, six years.
But what makes it modern? What, what, what's going on in modern data ops that's so important. Karin, I'm gonna ask you to kick it off because you're a real life practicing DevOps engineer and data ops.
What do you, what do you thinking, what, what's modern about Modern Data ops? Actually, what, what defines that? Well, to me, the, the word data ops itself is in a way, modern, because the thing with the word opposite is kind, is somewhat, uh, happened in the last maybe five or 10 years.
And, and by, but by keyword, modern, I would say is modern by the fact that, uh, uh, there's the development stage and the, until the production stage can be tracked and inspect along the way, and that, uh, people, uh, know which version of data is in each, uh, uh, uh, storage layer, and we can pinpoint exactly, uh, which could produce the data and who owes the data. So, in way the same as a practice with the, uh, number DevOps, where we can track, uh, which revision introduces change for certain features or something. So in a way, I will say the keyword for being model.
Instead, you can track the changes and, and trace respect to how it must be integrated or generated. Interesting. Did, I guess we should make sure our audience understands, what do we mean when we say data, data ops, right?
To the data ops is the intersection of DevOps and data engineering. Yes. Right?
Yes, absolutely. Well, I, this, if I could add to the last point. Sure.
Um, the, I think when we talk about modern data ops, we're really talking more about how companies are using that data as, you know, using this practice as a part of their everyday cycle, it doesn't seem as isolated. And so as you're making products, you're thinking about how do I use data? Um, that data incorporation is being layered in, um, all across an enterprise where it, in the past it had been more like in the analytics side and more in the building and contributing and coming to, like, how are we gonna make decisions?
But now, as we've modernized this and integrated into tools, not only is the need for the data, um, you know, in these, in these, uh, analytics, it needs to be available to the tools in a different way. It needs to be available to the products in a different way. So we're thinking about that whole process now as a part of how we're making products.
And I think that's what I think of when I say modern. It's not about like AI or anything like that. It's like, how are we trying to get this information and this process integrated into the rest of the organizations that we work with from an enterprise scale?
Um, and, uh, I, I, as a tool builder, my question is, you know, and as, and you're DevOps guy, Alan, I mean, nobody is more of the DevOps guy than you are, right? Uh, so it's like, what are these challenges? How do we incorporate this process with a lot of specialized things into the tools that we have now?
Like, how do we, how do we bring this together in a way that, um, we can actively, uh, incorporate that into the enterprise? And that's, and that's pretty much, I hope that also covered the question that you were asking. Absolutely.
So I, you know, building on what you said though, and, and Barry, I was interested in your thoughts on it, is it's not just the technology, it's the people, right? DevOps people, culture. And, and so what we see with data ops is we're bringing some new guys, new, not guys, excuse me.
That's terrible. Some new people into the, into the tribe, data scientists, data management folks, data engineers who are now working, it's part of our DevOps teams or with DevOps teams and all that that entails to that make modern data ops tick, right? Barry, you're a consultant.
I'm, I'm wondering if, what, do you see what, because you're dealing with many organizations where a lot of us deal with the organization we work at, what, what are you seeing at that? Mm-Hmm. When I think of modern Data, one thing I think of is, if you go back 5, 10, 15 years before data science became a big thing, uh, often data was used only by, uh, the, the main purpose was to power a website or an application.
So it was, it was used by the product engineers. And, um, there was often some use of it by the business, but it often wasn't very well supported. Maybe you just throw all your data onto a Postgres server and people do the best they can.
Um, and then, uh, data scientists came along also. And early on, I think there was a lot of a tip to try and maybe use whatever infrastructure or tooling you already had, or, Hey, you go figure it out. And it made projects really slow and inefficient.
I think now, when I think of modern data ops, I think that all three of those parties are, you know, have a seat at the table, and they all need to be supported. Well, and they have some overlapping needs, but they also have some different needs. And so, uh, a good modern data ops platform will support all of those.
I, I feel like we have an addressed elephant in the room, and I'll throw it out to all of you, right? One of the things about modern data, data ops and modern data is just the sheer size of the data, right? You know, it's not a, someone was telling me yesterday, you could buy a little, you know, we're talking about edge computing.
You could buy like a little brick this big now that can store petabytes, petabytes with a P of data that, that, that much data would bogle our minds five years ago, 10 years ago, we couldn't get our heads around that much data, but today we're using those kinds of data models, that much data to, to, to see answers, to see patterns, right? You can't have, you almost can't have like, sort of MLOps, right? Without having that big data set to, in order to, to empower that.
And that, that's a, that's a big change to me. That's, that's, you know, the biggest difference between modern, or maybe what happened in data beforehand is the, the sheer size of, of the data sets we're working with DC you, you've been around, I mean, what, how big a change? That's a sea change in my mind.
No, you're absolutely right. Um, I, it, it's, again, it's more of you. You have these large sets of data and the size of the data and is, is because of the data science, right?
That that's, that brought that into place in a way that, um, you had to really change the way you thought about it in the DevOps process. Um, what I see the biggest challenge as is how do we get the traditional DevOps guys to integrate with the data engineers, not just the scale of the data, but in the relevance of how you need to use it, right? Right.
And so it's, it's not about just like, hey, you know, the size. It's about how do we integrate this into, um, how do we integrate this into the, the corporation and, and the products that we're doing. As Barry mentioned, this was really more like, um, uh, uh, uh, a, this started out as more of like product analysis in trying to get to the right space.
But now there's so many extra things to think about. Uh, and I began to hear about something called like, metadata in the space, and I was like, okay, well, what do you do with that? What does that mean?
Um, so I, yeah, it's, it's huge in, in scale, but those, those changes are definitely regulated. Yeah. I'm curious too, to hear some of car car's takes, because I think in response to this large data revolution, we've seen architecture change as well, from warehouse to lake to lakehouse to mesh, and all of these things.
And I'm curious how you think that data or ops has evolved with these architectures as we've seen it growing complexity. Um, so speaking of the data velocity, uh, so they are new technologies, but, uh, I think it should been to, uh, to refer to a network from uhb where this, they said a big data is there, and I will say that me data is not dead. It's actually that it is, this is not a fabricated thing, but, uh, uh, the scale that people actually reach that point is, I mean, there's a lot of people in the world that reach the scale.
So to me, the challenges that are how we can bring the, um, um, proper practice to teams while we make, make enough room for them to prepare for the future skill scaling AppSec, because what I've seen often is that, uh, if you go for the, or out on the big data solution for the get go, uh, the teams may not reach the full potential of the platforms. So in the end, it, it results in basic computing money and resources. But if you start so small that there's no enough room to scale up, that they are kind of locked into the security operating at.
So in a way, I think partly the challenge is that, uh, how we choose to educate users and stakeholders on the data size and the technology that we should use. Because I think at this point, we can come up with ways to couple up technology together to support the skill of the data we need to operate that. But the question is, do we, uh, have enough support for the people we work with?
Because, um, so for, for example, um, we know that at a certain scale we can do the centralized material. So we need to decentralize. So it turns to a, which allows the teams to operate with velocity.
But the problem is that maybe sometimes we need to maybe consolidate the tools or say, Hey, I know that before each team you refer this and this tool, but now we are kind of need to use the same tool for every department because we need to, uh, conform to the, uh, single connector for the data source for the BI or something. And what happened is that, uh, the, this team, and they grade what they do, but when they went through, uh, month long, uh, bootcamp to upskill, uh, to, for them to use a new VI two at the end, the bootcamp they design, yeah. So in the end, uh, uh, it comes with, uh, the skills people you have, the security data you have and what you need for, for a business wise.
Because what I often see is that you have that much data, but operationally you don't use maybe 10% of the volume. So most data can be moved to core storage, and you get to use less compute and less to work with data you have. Got it.
Make that, sorry, Alan, if you wanted to, to chime in there. No, no. I think the DC board gonna say something.
Well, no, I just was gonna ask a question 'cause I'm, I'm, I'm listening to this, this, uh, as a distributed match of data and this transition to this distributed mess of data. But there is also this, this, uh, you were asking, you mentioned the tooling, um, and getting people to standardize on tools. Now, I, I am very passionate about the CD space.
That's, that's why I'm here. Um, and so is it a need to say, Hey, we need our data tools, our continuous data tools to kind of be standardized on how we use them in a way that improves this, the data ops process? Or is it, like you said, each team was using different tools you wanna bring into one tool, or are there already standards that are like written down and say, Hey, this is what this process is, this is how we inspect the data, this is how we make it ready.
I mean, like, what's, what's the transition in that space? Well, uh, anything can happen because, uh, before, uh, or actually at the start or anything, uh, when times go on, there can be silo opening in organizations. So maybe there are some pockets of your teams that they have kind of shadow team going on where they use, uh, they own two and another group in the, the DevOps team or the infra team.
So whatever is, you need to send someone in to gather the situation and then, uh, assist it and then come up with the plans to, uh, uh, to kind of convince them to use the tool. That, uh, platform thing. It's great idea because I, I think it goes back to, to the fact that, uh, uh, so this one, uh, my friend, he was in a consulting firm.
He worked with data mesh for three organizations, major one here. And what we found out is that, um, uh, if you do data mesh platform central, these centralized data, and you want to do, uh, test testings, SCE for each data set you have, it doesn't make sense to replicate the, uh, test infrastructure for each department, because that's where compute and, and, and you need to replicate the, uh, central data set anywhere. But if you share the, uh, testing infrastructure, you can reduce the cost, certainly where tooling standardize can help reduce the cost as well In this matter.
Okay. Lisa, sorry, could Lisa, Lisa, did you, I thought You had something to say, or you, uh, no, I just kind of wanted to, to see what Barry's kind of like takeaways, because I feel like sometimes in consulting, especially with a small and mid-size companies, you don't necessarily get that straight delineation of DevOps and data engineering. Almost everybody has to kind of have ownership of it themselves and their teams.
And it might not necessarily, like the DevOps teams might not have capacity to learn the tooling, uh, of data infrastructure. And how have you seen that kind of play out? Mm-Hmm.
One thing I've seen is, um, moving more and more of, uh, like infrastructure management to like the, uh, MLOps team or DataOps team itself. Um, having one central team that tries to manage all cloud infrastructure is just way, way too slow. So, uh, I've done a lot more terraform the last couple of years than I ever had before.
Um, see, can you repeat the rest of your question? Uh, sorry. Uh, just how have you noticed like the, the teams kind of converge in terms of their tooling, but also when you're working at a smaller mid-size organization, they might not have that capacity.
So how do you see that play out oftentimes? Mm-Hmm. Yeah.
One, one thing that's common with small or even mid-size companies, I think, is that, uh, there'll be teams that build data related jobs, but they don't, they either don't automate them, so they run them manually each day or something like that, or they run it unlike a machine in a very ad hoc way. Um, I, one company I work at is kind of midsize, and the teams have taken a number of different approaches to just like scheduling things. And those things probably work, but it means the team is responsible themselves for, you know, making sure that it continues to work.
And that can make the organization very fragile over time, because sometimes teams have a lot of turnover. Sometimes teams disappear completely. So if someone's built something valuable and it's not handled in any sort of standard way, then that thing may, you know, that job may fall over, uh, in six months and nobody knows where it's even running.
They just know that the data was magically showing up every day. So is this more of like a get ops process at this point? Like, is, is this transition of moving this data from place to place, or selecting the data that you use, or is this, is this closer to a true DevOps process, like a get ops process?
Um, or is this really like still, you know, handpicking these isolated data sets and growing them? When I first started to see data ops happen, you know, the, the data engineers, the data scientists were very, very selective about what they wanted in those data sets. Uh, and so automating that as a process was, was very, very difficult.
And when I think, and this is kind of pulling it back to the original statement about modern DevOps, but, uh, what I am seeing, or at least sensing is this need to kind of figure out, okay, we know we can control the flow of data. We understand now how when we are attaching to, you know, what, whatever the products are doing, how to collect that data. And so, is the process in, in your opinion, for like medium size and smaller, medium sized companies, is that moving in the direction of a more like GI ops kind of plane?
Or is this like, you know, it's people, like you said, writing jobs and then disappearing? Mm-Hmm. Yes.
I think, uh, things are definitely moving more towards GI ops and having like defined, you know, tools and infrastructure. So, uh, uh, the, the primary company I work at, it's very much GI ops. We, we have, uh, like a centralized airflow instance, like a job scheduler.
And, uh, those jobs are defined in this one central place. Um, one thing we're starting to do more and more is that, uh, individuals, we, uh, for jobs that need to do heavy processing, especially data science, they need to run custom code outside of airflow, but airflow still orchestrates those jobs. So more and more, whether a team is data focused, like data science, or it's some other team, maybe analysts or maybe a product engineering team, not a data team, we're, we're trying to onboard them to having the same standard approaches and, uh, where possible make that, uh, easy.
So, you know, they don't have to learn a complete complicated tech stack just to get things done. Um, and, you know, get an airflow can both be kind of a high bar for people to learn, uh, if they're not exposed to it before. So some of that is, um, you know, building nice tools to get them started.
Uh, some of it is, you know, providing good visibility. Uh, I spent a couple of weeks just working on making our Slack notifications nicer, so if something went wrong, it took people to exactly the place they needed to go, so they didn't have to, uh, navigate a complicated ui. But, so here you, I'm sorry, go ahead.
I was gonna say definitely, uh, get first and foremost, uh, one trend I've seen that I remember, uh, many years ago, it was very common for data tools and other kinds of tools to be like a black box. And all you could do was click to configure. And those tools demoed great, and people would buy them, but they weren't usable.
If two people were clicking on the same screen at once, you'd corrupt everything. You wouldn't have a history of what had changed. So I think more and more things have moved towards that.
Like I mentioned, Terraform, and that's one example, but also airflow. It's not that everything has to be code, but, uh, more and more things should be code and configuration, and those things should be managed in source control, so they have a history and they can be reviewed. Sure.
I, I, I think we're seeing how data ops can play into pipeline orchestration, right? GI ops and stuff like that. And when I hear GI ops, of course, what close from there to me is cloud native contain containerization, and then orchestration of containers.
And, you know, Karen, I'll come back to you, but any, anybody on the panel if you wanna speak up, how, how does this play with cloud native and containerization, and maybe even like observability, which is, you know, another big thing often associated with MLOps as well. Um, how, how do these things kind of play together nicely, or do they, or don't they? Uh, so a again, this is, this is the question.
I think that, um, you know, Karen, of course, being that you're the, the, the data ops and the SRE of the group, um, that is what I think about most is like the observability, containerization. How are the tools being made available, or how are you using the tools in the space to, um, you know, is there a standard for that? Um, so what are you doing in the space to make that happen?
Well, so speaking of data ops and in terms of observability and monitoring, in terms of data, so I think at this, uh, for this stage, we can agree on that. Uh, what we mean by observability is observability on data sets and assets and not the infrastructure itself, because that's because the observability for infrastructure. Say we have airflow, and most of the time we use managed airflow, say automat or GCP or a strong number.
So they have already have observability dashboard built in, so that's kind of standardized. But for, uh, the data metric itself, so, so this data, data dashboard, uh, quality and we see to like multi carlo, so that's, um, uh, sales, but then, uh, we can customize it in a way that, so what is C to B common pattern is that, uh, when you create a pipeline that, and then, uh, you add the per method to something and then you push it to say, postcode or something, and then you have, uh, dashboard that we should ask those metrics and then send alerts and do a, a routine when maybe hourly or, uh, 15 minutes into wash and then alert to the proper, uh, channels, say Slack, email or opportunity or what have you. Uh, so I see some people also push this into and display it on graph as well.
So I think in the end, uh, if you're talking, if you're talking about, uh, the infrastructure itself, so, and maybe airflow, the, this database, then those are kind of, that does, but if it's about data quality, the access itself, then, uh, mostly, uh, is the DY route where people customize the themselves and create their own dashboard and set up everything themselves because the needs, uh, depends solely different organizations. I don't say they standardized in any, uh, other places I've been in or what I've hear from other people. So that, that, that, you know, and again, tool builder.
So is there a gap between, like you said, you people monitor the data itself and there's observation of the data where it lands and, and the quality of it, but the infrastructure is a separate set of, of, of monitoring. So if the infrastructure itself, and, and we see this in more like in, in, if there's a problem with the delivery pipeline and it causes either a lack of data or, um, you know, a repeat of data, a data corruption, but that was in the infrastructure set, how does that translate to like, okay, it gets to the data you find a problem, but how do you know which part of the system to go back and look in then if there got separate? So I think this same problem as when we are doing pressing on production deployment for Apple Services, and then we have to, uh, to pinpoint exactly, uh, in which perimeter issue originated.
And so I think in this case, it makes sense that if we, uh, have key persons from each domain and then have them, uh, have say, discussion and, and kind of, uh, guesstimate on where the problem might alternate, and then each team investigate their own perimeter and come to a conclusion and then, uh, do a sit rep and update each shoulder. Because in a way, uh, it's like you say that a small problem or a small change can manifest itself in the other side of the, of the system, but it originates in the other side. So, um, I think in a way, uh, we can, uh, uh, add all the, uh, checkpoints of selves for the system to alert exactly on what the issue is, but we can communicate and help each shut it out because at, at, at the end of the day, when you are starting out the platform anyway, you still have to get the deployments and, uh, assess how the system behaves.
And if we don't communicate agile, then the tooling, uh, the, the tooling dream what happened. All right. Agreed.
Guys, we're almost outta time here, but I, I wanted to touch on, on one more, uh, aspect that, you know, we, the MLOps, AI ops versus generative ai, we haven't even discussed what generative AI may or may not do the data ops, right? As I said before we got on camera, I'm hearing more and more people move away from AI ops as a term, because I think when they talk about AI ops, they're referring to MLOps. But if you don't mind, and I know it's not in our abstract here, but what about generative ai?
It seems to be affecting everything. Why wouldn't it be affecting modern data ops? I'll throw it out to all of you.
What do you think? Yeah, and in my work, uh, I personally am not working on teams that are using, uh, generative AI models, like in the product. But, uh, it, in many cases, I'm using it extensively like chat, GPT or GitHub copilot for my own work writing code, uh, researching new technology, things like that.
Uh, I do know every company is exploring, uh, generative ai. And, uh, you know, I think the, in some ways, the jury is still a lot like how valuable it's gonna be. I think there's gonna be some very valuable use cases.
I think we're near the peak of a bubble where, uh, people, some people would have you assume that it's going to replace everything else, but I don't think that's the case because it's not as reliable as other things. It's very hardware intensive. Um, and now I think at, at the moment also people are trying to decide, do I use these commercial APIs or do I self-host like an an open source model?
Um, and you know, there's, there's ups and downsides to both, but if you host it yourself, you still, you know, you're still paying the price just in a different way. Um, so I think it's gonna be important, but I, I, I think it, it's gonna be some time before we really see, you know, how much of this will kind of fall away. People often want to assume that the new replaces the old, but you know, just because there's snowflake and BigQuery doesn't mean people don't use Postgres and MySQL anymore.
It's just that they have, you know, they have their content. Postgres is probably more popular than it's ever been, I think, these days. Right.
Sharon, what about you in your world? Gen ai? So I work in a consulting firm, and actually we have worked, we work on a few projects with, uh, gen AI is the main feature and the, okay, let's recap a little bit about MLOps.
MLOps is basically, uh, how you, uh, train and evaluate and improve the body. Uh, but you can cut the changes and which gen ai it is STEM you. But the challenge is that, uh, how do you, uh, track and monitor the results?
Because it's all takes. So if it limit to, uh, test duration like she GBT, then uh, the problem is that, uh, it's very hard to, uh, evaluate the dis of gene AI unless you use human labor. And that's really time consuming.
And the part that, uh, when you, uh, uh, stack the, uh, JI, uh, LM calls together, it's very hard to track exactly, uh, whereas the, the point where the results start to go off track. So there are many handful of different, uh, products that, uh, aim to solve this by say, just add this function or tracking agent, and then we can track the AIM calls for you. But they are not standardized.
And there are many tools and even more fragmented. And currently there are no, uh, based in solutions from major cloud providers or any major providers that are, that are, allows you to track gene or lot of the gene results. So in a way, it is still, um, uh, not kind of measure.
And what I noticed is that with chain ai, uh, mostly it works if you attach it to an existing feature and a sprinkler chain AI on top to make it more attractive or help use it in some way. But if you use an AI as core feature that it's very hard to to to sell it because I mean, people want it, but people don't understand that it's generative ai, which means the resource are not the same every time. And it's very hard to communicate to people that, uh, the answers to AI is not from human and it's not trustworthy, even though, even though, I mean, that's happen with Google, but with generic both There, Lisa, the DC actually, the dc why don't you go and then we'll come back to Lisa, we give her the last word.
Of course. Um, I still, I, I'm probably gonna say something really unpopular. I I think that, um, from depends on of course, how you're using it.
If you're using it like, uh, Karen said to add value to something that's already happened, um, in the process to make that more, um, human accessible. I think there's a place for it in, in that regard. But I, I still don't think from, um, monitoring, uh, a system that anything has is going to come out that's gonna be like a man Whitney, uh, uta, right?
If you are doing the data analysis, like in flight and looking at errors or trying to figure out where there's a problem in your system, that's really what you're trying to do. And that, um, doesn't really fit the catchall of, uh, uh, uh, generative AI kind of thing. It's more like machine learning, um, where you're like trying to collect this data and make some analysis.
But even that, you need a huge data set to do that over time. Um, if you're doing it for code that you're writing, that's always changing where those errors are. And so you really need a really good, you know, uh, analytic process.
And that already exists in like a man Whitney, that in my opinion. Um, you know, I, I know I've got the, I'm here with the email ops and, uh, um, the, the data ops folks, so they probably aren't gonna agree with that, but I said it, you know, I'll take a risk, man. Hey, it's out there.
It's out there. Once, you know, the internet is once it's out there. Hey, Lisa, I'm, I'm going to come to you and, and don't, not just on agenda ai, but why don't you, you know, this was your first co-host.
Why don't you wrap it up and what you think, what you saw? Yeah, no, today was a, a really interesting conversation, and I think it's rare to see like the data ops people and the DevOps people get to have this conversation. You can see where the holes in people's knowledge are too, right?
And with respect to generative AI and where we're moving with data architecture, you know, we're dealing with so many large amounts of, of unstructured data that we've never had before. We're dealing with multimodal, you know, data and, and trying to process it and, and serve it in the same way we're having to deal with dynamic context as we're trying to serve to different, like age agentic architectures and like LLM and RAG systems. And so things, as I noticed, uh, and based off this conversation as well, our data itself is moving into this dynamic system, but we have to kind of make it predictable and we have to make it manageable in a way.
And I think that's where a lot of our trickiness is. Will we see generative AI be able to be used to, to replace some parts of DevOps systems or DataOps systems in the future? I definitely think so.
I think that it's so early now where we were two years ago. It's a complete night and day. You have to give these things time to develop.
And, you know, when I think of things like, you know, predictive load balancing or, you know, dynamic configuration, I think that these things are going to be enhanced and, and, you know, really, uh, in an exciting way. So yeah, today was an amazing conversation. Alan, thank you so much for hosting us and asking Thank you.
Amazing questions. And of course, Mary, I gotta tell you that For coming. Yeah, Thank you.
No, thank you. Thank the DC He, he keeps it moving, judge, you know, great job there. All right, we're gonna call a wrap on this version of CD Pipeline.
Uh, we'll be back next month. We'll, with another great topic of, of interest to CD engineers. I, I will just quickly tell you, you know, when I first started doing DevSecOps at the RSA event about nine years ago, it was very similar to this.
We were bringing a security community into the DevOps tribe, and they didn't always see how they fit together. And now today, DevSecOps is DevOps. I think we'll see the same thing with DataOps.
We'll see data engineers and data scientists as part of just part of the DevOps umbrella. So let's see. You know, and it's all, all in CD For now though, this is Alan Shimel for Techstrong.
On behalf of the CDF and the Linux Foundation, thanks for watching. We'll see you again soon.

