Demetrios Brinkmann on Building Reliable Evaluation Systems for MLOps | swampUP 2025
Standard benchmarks often fall short and can be misleading. Leaderboards can erode trust in model claims, as they rarely address specific, real-world needs. In this talk, Demetrios Brinkmann will detail how MLOps engineers and developers can build and continuously update their own evaluation systems to create a strong competitive advantage. He’ll cover how to build a reliable “golden dataset,” optimize data collection, labeling, and utilize the right tools to ensure evaluations truly reflect their intended use case.
Transcript
Hey everyone. We're back here. Well, we're not live, unfortunately.
We were live when we recorded this, but you're watching it on recording. Let me introduce you to Dimitrius Brinkman. We are here at Swamp Up.
If you couldn't tell 2025 Swamp Up. And we are thrilled to have you tuning into our coverage of this year's Jfr Swamp Up Demetrius, first of all, welcome to Text on tv. It's great to have you on here, Demetrius.
Looking at my notes here, it says, uh, founder of the ML Ops community. Great title. Talk to our audience a little bit.
What, what exactly is it and what do you do there? Yeah, so we're a community of around a hundred thousand developers right now that's primarily focused on bringing AI and ML into production. That's the main thing because there's a lot of research, there's a lot of demos that you see out there, but then actually getting use out of it and bringing it into production, that's what we focus on.
And we do that in the various, in various ways. One being we've got a Slack workspace, we'll do in-person events like meetups or workshops or conferences. We do virtual events and like meetups and workshops and conferences.
I also have a podcast myself. We have a newsletter. There's various ways to engage in the community.
We'll do like one-on-one matches, curated matches of people in the community. So in general, we just are trying to keep the education and the understanding of this field as high as possible because it is moving so fast. It is.
Hey, just saying you have a podcast isn't enough. Look into that camera. Tell them where they can get you podcasts.
What's the name of it? Yeah, you can find it on anywhere that you find podcasts. It's called the ML Lops Community podcast.
And right now we're on the 314th episode. Really? Yeah.
So we've been How Often do you do 'em? Twice a week. Really?
That's fantastic. Good stuff, man. So you're also keynoting or on stage tomorrow doing a session.
You know, by the time people see this, you probably have already done it. Yeah. So tell 'em what they missed.
Well, by the time you see this, it could have gone horribly or it could have gone wonderfully. Let's hope for, I'm sure It went. But really what I'm excited about talking about is the idea of how there's, there's almost two big ideas that I wanna present.
One is how the chat interface isn't necessarily the best interface for us to interact with machines. It's very low bandwidth and we're used to a much higher bandwidth when we interact with humans. And then the other idea is what I am thinking about how all these companies that are putting agents into production, they all want to be an agent.
They don't want to be a tool. And the way that it could shake out is you have a master agent that goes off and is using tools, but right now it's very fragmented. And this ecosystem that we live in to today is, I go and I navigate to one chat bot, and that has agentic capabilities and it goes off and it does some stuff.
Maybe it has access to some tools, but it's not like there's this ecosystem, this homogeneous ecosystem that I know this one chat bot can do anything. I have to then go, if I want something specific done, navigate to another website and use their agent capabilities to do something. So a perfect example of this is when I wanted to file a claim for a delayed flight that I had, I was talking with my LLM of choice and saying, you know, can I get money back on this?
And am I in the right to file a claim? And it said, yeah. And instantly what you wanna do is say, okay, go file it.
Go file It. That's the user experience that I want. And, and, and you know what?
That, let's call it the dream, if you will. And and I thought we were getting at least when you talk about travel. Yeah, right.
I thought that was part of the, uh, and I'm not knocking them, don't get me wrong, but that was part of this chat GPT agent, like, Hey, chat GPT, I gotta fly to Flagstaff, go out, find the best fare and book it for me. Yeah. I haven't used it yet.
I don't know if you have No, I I don't, I don't trust it. 'cause I'd have to go look at the flights myself and make sure that it, I'm not stopping over in Chattanooga or some, some place where, Well, you bring up something fascinating. There's two pieces of that.
One is the trust aspect, and the other is this UX cliff that I've been thinking about where a lot of interactions with machines, we don't necessarily need to type everything out. That's a much slower experience than if we just do two clicks and we get what we want. Yeah.
So there's almost this valley that we need to cross before an agent is even useful. The task has to be quite complex in order for us to do that. And I think the reason that the flight bookings have captivated our attention is everybody has done that, and it's way more than two clicks and it's cumbersome.
And so when we think about that, we think, wow, it would be nice if I could just say, I want to do it this time, this day, I want to go to this place. And then it goes and does it. And we don't have to go and click through and do all these multi clicks, which is, and then look back, ah, is this the price I want?
I don't know. And that's not fun. Yeah.
But to me, it, it sounds like a pay me now or pay me later kind of situation. Right. Because how do I set all those up?
It, it, I, look, I don't pretend to be a, an AI expert, but like I've gotten to the point now with my ais of choice where it knows me, right? Yeah. It knows my style and voice for when I'm writing it knows what I want out of the tasks, the usual tasks that I ask it to perform.
It would be great if somehow I could train my ai like, hey, I like to fly our first thing in the morning. Yeah. I like to fly home first thing in the morning.
I will do a direct flight. I don't care if it's twice as much money. And no matter what I want direct, if I could help it, um, you know, all of these little kind of, this is me kind of thing.
Yeah. And I think that's where we struggle, right. Well also, if you think about that I'm not the same person today as I am tomorrow.
Yeah. And maybe that's, there's certain things that I have hard rules on, and then there's other things that I'm a little bit more flexible on. And so that as a problem is a very difficult one to crack.
Yeah. I also think that, you mentioned something fascinating earlier about the trust, which is we have to be okay if we do have this master agent world that is some kind of a hybrid chat interface. So it's not only us with words, but maybe there's other kind of UIs that we can take advantage of.
So let's explore that. What do you mean? Well, I look at different ways that we interact with programs already.
And if you take a little inspiration from video folks, you have histograms. Like these guys are used to dealing with histograms for the colors. So is there a world where we can deal with a histogram like experience for what we want as opposed to trying to really get into the minutiae in the words, because words aren't as easy to develop or as easy to tweak on that very small scale level.
And then on the other hand, when we interact with humans, we're interacting at a very high bandwidth. And I'm sure you've been in a meeting where you end up diagramming things to get your point across. When we are just restricted to text, we can't diagram anything.
Yeah. And again, that brings us down in the bandwidth that we're able to convey to that LLM. So potentially there's some kind of a whiteboard or it you can think of like your, your tablet that you're able to diagram with and it's recording your voice as you're talking to it.
That could be a world. But at the end of the day, right now, what we're funneled into is the experience of just chat. And then you're getting some inkling of when the chat bot will respond to you, it gives you these new UI elements.
Right. So sometimes you'll get a scroll, sometimes you'll get a photo or you'll get a code snippet, some data visualization. You get that, which is great.
And I think that's the first step. But for us as input, we need to up the input levels. Well, so I'm a little older than you.
Yeah. I'm gonna guess. But, uh, look, I'm a child of Star Trek, right?
Yeah. Man, my whole life I wanted to be Scotty and just say hello computer, you know, and, and, and tell it what I want. But I I, I thought we were getting there, right?
And then I realized in like doing videos like this, right? So I can't give the video to the AI and, and tell it do it. You gotta transcript it.
And you would say, okay, transcribing is easy and it's word for word. And even if you, you know, fact, uh, uh, copy, edit the transcript to make sure it is in fact word for word, it's not enough. Because the way humans communicate, we communicate with our eyes, our eyebrows, our hands, nuances, tone in, in speech.
Yeah. Right. And ouris just aren't up to that yet.
No. So, I don't know. I mean, one of the, one of the things that really they say separated humans, let's say from Neanderthal or Dan, so the Nasos or whatever that uhhuh close relative of the Neanderthal is, is our, our, the, the, the depth of our communication.
Even if we didn't have a huge big difference in vocabulary, all the nuances in human to human communication. And I think that is, that's a job that we need the AI to solve. Yeah.
We don't have that No. Anywhere near that, right? No, no.
And it's, you don't realize how important it is until you just look at a transcript. Yeah. But there also is the whole idea of, I know there's probably people out there that are gonna be thinking, oh, well, voice is trying to tackle that problem.
Voice AI is the next frontier. But I am not sure, have you played around with the voice tools? It's not that they're bad, it's that us in a work setting, what am I gonna do?
Go put myself in a cubicle when I wanna work and speak to my Yeah. Computer. You know, it's funny you brought that up.
So I met a guy I interviewed last week, and I'll give a shout out to them. This guy, Dr. Allen Becker, his PhD is in voice to text, text voice and ai.
He started a company, got sold to Snapchat. He ran Snapchat's text to voice for a while, but now he has a new company, I think it's called E Self, E Self ai. Check it out.
When we're done for me, you can sign up for a free five instance thing. They've developed avatars. Yeah.
That look at you, that watch you and talk to you hooked into LLM in the backend. And they do try to pick up nuance Yeah. From your voice and from your gestures.
Mm-hmm. It's early. I played with it.
It's, it's freaky. Right? It really is.
It freaked me out, but it, you know, it's not perfect yet. Yeah. But, um, I am, I'm bullish on, on that happening.
Yeah. But you still have this, it's like we're in meetings, right. And then we have to have a moment where we get work done.
Right. And so if the way that we have to get work done is by talking to our, It's still cumbersome. It's like we're in a meeting again.
Yeah. And that's exactly it. And really, whether you're talking to the computer or typing to the computer, there are people who type really quick.
Yeah. And a lot of people are really not good communicators verbally that Like me. Exactly.
That There are, I mean, that, that's an issue. That that's definitely, it's a big issue, You Know. But let's, let's look at it from the other side of the coin.
Demetri, you know, the windows mouse clicking kind of interface that is dominant today. Look, this was like 1960s, early seventies out of the, uh, park. Yeah.
You know, Xerox Park out here, we haven't really, I mean, it's been 50 years. Yeah. And we haven't found a better mouse trap.
It's time. It is time. You could see that with like the touch screens.
We have these gestures, you know, the pinch to zoom mm-hmm. The swipe. Mm-hmm.
And the other thing that I think is a big problem with us having to use chat and take what's in our mind and put it into a chat bot is how, right now we're very used to being fed things. It's almost like a passive experience A lot of the time when we're on the internet. And you can think about Netflix or when you're scrolling on Instagram or TikTok, these are passive experiences that we have become accustomed to.
And now chatting is very active. Right. We have to really define what we're looking for, what we want and put it into the chat bot.
And so we don't have these passive gestures anymore when you're trying to work with chat either, which I find fascinating too. So is there a way to bring in these passive gestures into the chat experience? Or I guess at a certain point, once it evolves outside of chat so much, we probably won't call it the chat experience, we'll call it just the AI and Communication experience.
Yeah. So you're not trying to tell me we gotta get passive aggressive with our ais. Do you?
Are we? No, not that, not on that level. I mean, you might see it, you might see It Better.
I don't know. I haven't tried. I'll Tell you one of my biggest things that I've had to teach myself is you don't have to be polite.
You're only making it harder on them every time you say thank you and please. And all of these things. Nice burning energy.
Exactly. I wanna turn a little bit Demetrius and, and talk a little bit about security. Right.
So look, I, I think everyone agrees that we're all gonna have agents, digital workers, whatever you want to call 'em mm-hmm. Who are gonna go off and do these tasks for us. Whether it's booking flights, writing code, or, or what have you.
And we're gonna need either, we're gonna need a crap ton of agents, right? One, like almost an ephe ephemeral, disposable agent for every task we do. Mm-hmm.
Or some sort of master agent that's able to clone small parts of itself to do specific task. Yeah. No matter how, no matter which way we go, there's security issues.
Yeah. How do you view that? Yeah.
There's a few different issues that I'm looking at. And these are like the most basic of the most basic. If we get some of these DevSecOps people in here, I'm sure they think about it on much different levels, but in a broad strokes way, if we have this world where we have a master agent that helps us go out and it's our gateway into the world, and it can use these tools and it's a big if, because like I said, everybody wants to be an agent.
They don't wanna be a tool because you're giving up your distribution, you're giving up your relationship with your customer. If now Chachi, BT, or Gemini is what chooses to use you or not as a tool, that's a big vulnerability for your company. So that's a big if right there.
But if we do get to that point where I go to my LLM of choice and then I sink in with the tools that are out there on the internet. So Amazon is a tool. So buy something from Amazon can be many different types of tools.
Uh, look for something on Amazon, whatever, search Amazon. Now, are we okay with the context just flying around the internet? This data potentially sensitive data is now gonna be going to different tools and going to different LLMs.
And I'm not talking on the l are we okay with our data going to the LLM provider, but just data flying around the internet. That's one part that I think about. All right, well we need to get the context and we need to have a way to securely do that.
It's not necessarily a new problem because we've been transporting data across the internet for a while now, but now it's a little bit different because there's agents that are interacting with each other and maybe one agent thinks that Conte isn't that personal. But then the other agent, when it summarizes it, it sees that, oh yeah, actually it will say something that is personal and you don't want that. Right?
So you have wild cards in each agent, agent to tool call or subagent, whatever you wanna call it. And then next you have the authentication issues. So I want my agent to be able to understand everything about me.
That means it needs to look at my calendar, it needs to look at my Gmail, it needs to look in all of, everything that I'm privy to. It needs to be privy to in case it needs to act on my behalf. So you need to off into all these things, but it's not just OAuth because you then get to the next piece, which is the actions.
You don't wanna give it permission to take any kind of action. No. You wanna give it permission to take the action that you said was okay, not anything else.
Because if you give it a lot of scope, it can abuse its privileges. So I've been in, I didn't tell you this, I've been in security for 25, 30 years. You're only describing what I would say are innocent security issues on the agent.
That's true. What about when the bad guys say, oh, he's got an agent. Let me exploit that.
Well, did you hear what happened recently? There was I think some output from an LLM that had a nefarious link. And when the user clicked on that link, it then was able to take control of the system.
It happens all the time. And so yeah, you have, again, you have this wild card in there that the nefarious actors can hijack this agenda. They're Not dumb.
They're as smart as we are. They're well funded, well organized, and they, and that, that's the truth. And if you do end up having everyone as a tool for your master agent, how do you verify that this tool is okay to use?
Agreed. It's it, look, here's the good news. First of all, no one's gonna waste.
Everyone's running as fast as they can anyway. And they're gonna keep running as fast as they can. But as these things, and I, I've seen cycles before, right?
We never lead with security. We just don't. Yeah.
We, The sad thing. But it's True. It's a sad, but as a security person, you either gotta come to terms with that or, or you know, you're gonna be depressed.
Um, we will catch up, we will put the guardrails in, we will come up with processes around it, but people are gonna run as fast as they can. And, and, you know, you can't put your, you can't lay down in front of the tracks and say, stop the train. Yeah.
You get run over. Yeah. And, and so I always say it's more of a yes we can mm-hmm.
Kind of thing, right? Yes, we can. You wanna run as fast as you want?
Yes, you can. We'll figure it out. Yeah.
And I, and I think that if I had to leave us with one thing, that's what I'd leave it. Yes, we can. We'll figure it out.
But man, thanks for the work you do, Demetrius with you community. It sounds great, man. We appreciate you.
Thank you for presenting and swamp up and for being here on Tech Drunk tv. Thanks everybody. Alrighty.
We're gonna take a break. We've got more swamp up coverage coming your way, so check it out.