Operationalizing GenAI in Organizations with Maria Apazoglou
Maria Apazoglou, head of artificial intelligence (AI), business intelligence (BI) and data platform for Thomson Reuters, delves into what it takes for organizations to successfully operationalize generative AI.
Transcript
This is Textron tv. Hey guys, thanks with, we're here with Maria Aplu, who is head of bi AI and data platforms for Thomson Reuters, and we're talking about how they're going about operationalizing ai. And Maria, welcome to the show.
Hi. Thank you for having me. I think a lot of people are struggling with this issue.
I mean, everybody's kind of very enthusiastic about all things related to AI these days, but, um, it's a bit of a challenge to get it into a, a process that every that's repeatable and that people can take advantage of. And I know you folks are pretty far down that path. So what did you encounter in terms of early challenges that, you know, you wish you knew now versus then?
Um, I think I would say the way that I am thinking and we are thinking about AI in, in two different ways. So you have the AI that we use internally to make our kind of like our working, uh, ways better and like, you know, improve the way that we work and make it easier for us to perform like our activities. And then there's the AI that we have into our products and really what we're trying to do with the capabilities that I'm responsible for within the AI platform is scaled for both of these two, uh, user bases and, and communities.
I think from, um, I dunno if I could say there was something that I knew ahead of, ahead of time early on, like from the beginning of the, of the journey within AI and the platform and the things that we were doing, we really took some decisions that I think have paid off. Uh, one of them would be to be multi-cloud. Uh, so I think we understood, and I knew kind of early on that um, there are capabilities that exist out there from multiple providers and being able to be multi-cloud and be able to exploit that in a way that you enable your teams to experiment and easily understand what is the best use case, the best model for each of the use cases that they want to build was, was crucial.
And the second one is that, uh, when we're looking at our, our use cases and when we're looking at the lifecycle, the things that were true and we're holding through within generative ai, within ai, the whole through within generative ai, uh, and that really is around how you get data, uh, how you create your model, how you evaluate your model, how you then put into production, and then you continuously monitor. So we, what we were really approaching and thinking is a similar way of, of working and uh, kind of addressing generative AI as what we would with ai and then tweaking some of these capabilities and the steps. So now you're not necessarily like training a model, but you are more kind of, um, training or like you're utilizing a prompt and you're creating prompts and you're creating like ways that you can use the models instead of necessarily training as much, uh, smaller models as you did before.
So we, we made some changes and we had to add some tooling to address the kind of likeif from AI to generative ai. But we held through a lot of the things that were fundamental into the journey, which is, you know, how do you get access to good quality of data? How do you evaluate and how you productionize, What kinds of models are you using and how do you decide that process?
'cause some folks are using, well I open AI or AWS foundational or, uh, other folks are taking open source models and trying to customize those. How did you go about that? I would say our approach is a mix of everything and uh, that's something that I appreciate the most.
So as a platform team, so we as a platform team, we consider ourselves to be an enabler. Uh, as an enabler, my job is to make sure that everything is available, uh, in the same at ease and securely. So that's what we take pride on and that's the thing that we kind of really, uh, push forward.
We also try to make it easy to compare models. So we have like low code capabilities that will allow you to compare models across 10 or 15 different kind of like open, uh, like, um, foundational ones or one that are, are others that are coming, that are open source, uh, that again are can coming through different services and enable our end users to compare these models in, you know, one prompt or like one user a query. And you see the output of everything.
So we have these capabilities to help you with a choice, but at the end of the day, we also rely a lot in our data scientists and our TR labs team that is like experts in the fields. And when it comes to solving a specific problem or a use case, they use our tools to experiment and find which of the model is the one that provides a better, a better experience and the better experiences around, you know, the accuracy of the response, the performance of the model in terms of how speed it is, how speed it is, and how fast it is and the user experience overall and the economics of the model as well. So like it's really around thinking everything when we are thinking about the evaluating them.
How did you minimize the hallucinations that get created sometimes? 'cause people are struggling sometimes with the probabilistic nature of lms and a lot of end users get it in their head that, you know, it's magically deterministic. And how did you kind of walk them through how to use these things?
So we do a lot of training. So internally for like the users of internal, uh, platform capabilities, we have trainings that we kind of, uh, we certify people on. I think we've certified over 2000 people or maybe even more at this moment in time in the tools that we have within the platform so that when they're using large language models, they know that, uh, to your point, the answer is not deterministic and there are actually parameters that you can't even tweak to get the model more, more or less creative.
So we do trainings to help people understand that the model is something that is a tool that you can use, but you still need to have your critical judgment at the end of the day when it comes to our internal tools. Um, we then also, when we are looking at our products and the things that we, that we have in, uh, like in as a Thomson Reuters offering, there is a, a, a much more rigorous process of basically evaluating the models, continuously testing the models, and most importantly grounding the models in, in data, right? So that we are relying on Thomson Reuters data and like the quality of that data to be able to make sure that the models are performing at their best capabilities on grounded on the data instead of grounded on the knowledge that they've had from, uh, you know, from the public domain or whatever they were trained on.
But it's really around like, you know, continuously monitoring, continuously listening to customer feedback, evaluating and incorporating all of that into the, into the life cycle of the model and improving the models and the way that they work to ensure that what we are using and the way that we're using it, uh, meets the needs. There's no substitute for the work. How did you determine which kinds of models to build?
Are there a set of return on investment metrics that people are tracking? How do you know, uh, which effort was worth the time and effort? Uh, so again, like for, for us it's really like for me as a team, it's really around, uh, like as I mentioned before, like the enablement.
So we didn't decide what to build. We decided that, you know, we need to make sure that we offer an ecosystem where teams can access models across the multiple different vendors that exist out there, as well as provide them with environments to hook up to, you know, open source frameworks and, and tools and models like hack face and things like that. So we have like a that approach to enable teams to utilize something that is off the self but across providers and then something that they can build on their own.
And the way that we try to, to kind of expose these capabilities is to obfuscate as much as possible that nuance to our end user. So, um, whether they are operating in our high code environment or in a no-code environment, they could basically from a notebook be able to access the models from all different providers and they won't necessarily have to be switching between environments. And that's fundamental to be able to compare.
Uh, and really like to the, to your point earlier about the, the economics, there are some things that we do. So whenever a new model gets released, which is, you know, more efficient and, and better, like as you'll see from the market when you'll have, you know, different versions of cloud coming up or different versions of TPT coming up of different vessels of Germany coming up, really what we do is we upgrade every time like a lot of our, um, low code solutions to be reflecting and coming off that so that, because all of these models tend to be more efficient and more capable and, um, across multiple domains. And that's one of the things that we do to encourage and make sure that we always remain current, but we also remain as efficient as possible and we still allow the maintaining and we still basically allow people to use the models, which are previous versions.
Uh, but it's like what is gets exposed by default is the, the latest, if that makes sense. And that has helped quite a bit with like economics of scale on, on how people are using the models to augment their day to day kind of life and, and, and work. And again, with our like products, it's a more kind of like involved process to figure out, you know, which version and which model is performing well.
Like is it outperforming a model that is smaller, that is trained versus something that you have where you are using at ease with like different versions of ragan and prompting. How is the relationship between the data science and IT teams evolving? A lot of organizations are trying to move more responsibility for the inference engines and deployments to the IT team and the data science folks focus more on the training.
Um, what's the relationship that Thomson Reuters between these squads? Uh, we have a, we have data scientists. Uh, we have a team of data scientists, which is called TR Labs, which I would say is a, a flagship team within Thomson Reuters that has like an amazing talent, uh, of people that, you know, they have studied n lp, they're like experts on the field.
We also have data scientists like dotted within different like business units for some of the work that might be needed, uh, to support these business units. And then we have a product engineering team that collaborate with the tier labs team when it comes to taking models and moving them into production. And I think the thing that we are seeing and the shift that we're seeing is really having these things working more closely together.
So that's like a collaboration of the same code base collaboration, like across like the lifecycle of, of the model, uh, that makes it much better on and like the path to production rather than being something which is, you know, it's just you create something over here and then it's, it's kind of moved, uh, over the fence because it's, you also want to make sure that the, the engineering teams are also responsible, are also understanding the mechanics of how to maintain the models or how to retrain the models. So it's important for them to be part of the process early on. Um, and then we also see a lot of teams which are within engineering, being able to create models or create initial, I would say versions of kind of AI solutions on their own.
Like they could, you know, a lot of the engineering teams are also starting to become prompt engineers and they're able to utilize some of the tools that we have to write prompts and choose a specific model to specific data and create a solution. And that gives a very good head start for then our data scientists if we need to do something which is more complicated than a, you know, a simple rug and a simple prompt and a simple model. So there are ways that we see like, you know, the, the start of the journey coming sometimes from sometimes even product teams, uh, sometimes product engineering teams as the idea that then our data scientist teams might perfect and move a little bit closer to like a, a, a standard that is required or more higher or higher.
And then the engineering teams then continue to kind of like be part of the journey and follow on moving that into production. Right. One of the things that I think we're seeing is the language models themselves now come in various sizes.
They're small, medium and large, and, um, and they're, are they being continuously updated? I mean, how do you manage that process as, as you, as we need to retrain models and there's drifts, um, it seems like that requires some sort of ML ops plus DevOps workflow, but how did you guys go about it? So the first one is about obtaining, uh, obtaining feedback and sometimes the, and having a very good test, uh, test suite.
So I think these are the two fundamental ones. So, uh, in one of the products that I am, uh, for example, responsible for, which is our coun on the engineering side, uh, one of the things that we have is a, a, a big test suite of, you know, use cases that we then run our models every day to see the performance of the models and, and, and kind of like plot that and monitor that over time. And that's really around, you know, the cases that we have that we create when we are initially evaluating and putting something to production that then are maintained and augmented as a mo as a, an AI use cases in production and, and performs.
And you want to see how it is. And then you then have like a human evaluation where we'll have SMEs coming in and evaluating, uh, like the models in like not just the models, but really the AI use cases, right? Because you're not just evaluating the model itself, you're evaluating the AI use case.
So we have SMEs that are evaluating, um, like the, the use cases and they can, you know, evaluate them over various different criteria. It could be around like completeness, relevance, you know, coherence of answers. Like, you know, there's a multitude of criteria that our SMEs can evaluate, um, the models on.
And then there's also like what's a bit more, I would say industry practic practice at this moment in time or one of the industry practices that we see with is having an LLM, we call it LLM assisted evaluation, where you will have one large language model being used and prompted to evaluate another, uh, another AI solution. And then again, you can have hybrids between having the LLM and then you have human reviewing the evaluation of the LLM, of the LLM. So, uh, there's multiple of things that we are basically are seeing and, and we are using.
And I think the reality is that, um, taking as much as a, I would say comprehensive and like open approach is ma is I think what, what is important and not relying on a single thing, uh, you know, you could have like traditional things like rules, metrics and things like that to evaluate difference between ground truth, but we know that's not necessarily always the most effective way. And you could have the LLM assisted evaluation and you know, that's not always the most effective way. And then you have the human evaluation.
And again, that, that could take a very long time and like require a lot of people. So in our view and like the things that we want to enable is all three of them so that you can have like the best of both. And when you're judging or when you're looking at large language models and AI solutions, you're not looking them from one angle, but you're looking them across multiple angles to ensure that you, you've covered as much as possible.
All folks, well you heard it here, AI models are like children. It takes a village to raise them and train them and make them responsible. So there you go.
Maria, thanks for being on the show. Thank you so much. All right.
And back to you guys in the studio.