The Rise of Synthetic Data in AI – The AI Times EP 5
Datasets are becoming bigger and more unwieldy, and as they do, they are susceptible to greater copyright challenges and activism as well as security issues. As AI becomes more mainstream, some organizations are turning to synthetic data to train tomorrow’s AI and machine learning models. But how mature is this practice? Is synthetic data up to the challenge of replicating real data, or is it just a sub-par substitute? In this episode, hosts Mike Vizard and Lee Baker are joined by Etienne Martin (Algola), Ram Ramamoorthy (ManageEngine/Zoho) and Ryan Berry (OneStream) to discuss the opportunities and obstacles around synthetic data, explore what’s state of-the-art in synthetic data and predict where it’s all going.
Transcript
Good morning, good afternoon, and wherever you are in the world. Uh, welcome to, uh, the latest episode of AI Times, uh, brought to you in conjunction with the AI Infrastructure Alliance and Techstrong. Uh, uh, my name is Lee Baker, uh, and alongside my co-host, Mike Baard, uh, we address some of the key challenges and issues, uh, within ML lops today.
Um, and we have a illustrious panel for you today to tackle, uh, one of the most, uh, uh, interesting aspects of kind of data management to feed MLOps that of synthetic data. Uh, almost everything in our lives today is recorded and stored digitally. Uh, every interaction with TE technology leads, leaves a digital trail, each credit card purchase, medical diagnosis, Google search, Facebook Post, or Netflix preference is another recorded data point, a breadcrumb that leads back to us.
Uh, however, as we produce more and more data, it becomes harder and harder to create truly anonymized data and the risks of companies releasing potentially re identifiable personal data grows. Uh, so how do we reduce the RI pro privacy risks and burden of legal compliance, uh, that need to be in place for sharing and processing real or personal, uh, or private data? I think we've got the right panel to address some of that challenge today.
Um, so, um, uh, let, let's get straight into this. And guys, if you can just introduce yourself very briefly as, as you answer your questions so that, um, our, our guests know who's speaking today. So let's kick this off.
Let's start with kind of starting with, you know, where is synthetic data come from? You know, the idea of kind of data theft is nothing new. Uh, but, um, uh, where, how has synthetic data emerged as a technique and approach, uh, within today's kind of modern approach to, to data processing?
ram, do you wanna, do you wanna take us off on that one? Surely. So, uh, I'm Pam Prash.
I lead the AI research for Manage Engine. Uh, we make tools for IT administrators, and, uh, I've been working on AI since 2011. And, uh, back when I started, we didn't have much of this concept called synthetic data.
In fact, the first project that we started off was on sentiment analysis, where we were trying to do sentiment analysis in enterprise help desk. But the only training data sets that we had were from Amazon customer review dataset, imdb movie review dataset. And, you know, you're looking at B two B emails and you're training a model on, uh, consumers data very, you know, very casual.
You don't write very formal email, like movie reviews, right? So, uh, the first thing we did was we had to annotate a lot of our own data sets. But again, the challenge is being a, a lot of it had to be done in-House because it had a lot of sensitive information.
Uh, we just used our own organization's data. So even within our organization, that was fine. So a lot of sensitive information, uh, was there, and slowly as AI came, uh, into development, and, uh, all of a sudden we saw that the whole AI narrative, uh, after 2011, the summer that we're in, is primarily due to three things.
One is, uh, cheaper data collection. I mean, everybody wears a smartwatch, uh, where heart rate, uh, food staffs sleeping, uh, head HR dip and whatever, right? 10 years back, had you told me every aspect of my life would be recorded in a digital, uh, place, and it would be stored, that would've been something super fancy.
But no, it's, it's bare minimum. And then cheaper data collection, you have the infinite power of cloud to actually have a lot of compute. So you have ton of crossing power and then having revenue models based on data, right?
We don't pay for our search engines, we don't pay for our social networks, and these companies are successful trillion dollar business models. So all these three things put together led to the rise in ai. And one thing that we realized was, especially in enterprises, because today I can monetize my personal information, but is it same to monetize my company's personal information, right?
It's my business secret. Can I use that? Can I have it monetized for the benefit of somebody else?
Now, this is where the whole need for synthetic data comes in, where, you know, the quality of your predictions are going to be very, very important, right? If I'm going to predict a fake outage in my monitoring system, then they're going to turn it off or probably go look at another monitoring vendor. At the same time, I don't want my customer's data.
Let's say I'm an insurance company. I don't want my data being used to the advantage of my computer. So to fill the gap to, and, and the way the AI algorithms have exploded, they all need more and more computing power, more and more data, right?
Especially after the advent of large language models. The quest for data is like crazy. So that is where I think synthetic data can come in.
Uh, it really augments real life data very well. You can extrapolate real life data, and then you can identify the patterns. You can generate a lot of synthetic data, and that kind of satisfies the hunger of modern day AI algorithms.
I only see the need for synthetic data growing further and further given these algorithms are getting more and more data hungry. So that would be my pointly. Very good.
Very good. So, okay, so we, we understand where kind of synthetics data has been born from, um, and, and kind of what's kind of stimulating kind of that use. Uh, but let's understand a little bit about what is the kind of the data sets that are most prevalent, uh, in regard to kind of synthetic data.
Uh, are, are we talking about structured, you know, tabular data, uh, or, or are we talking about more unstructured data in the sense of kind of aspects of my, of my life that may be I image-based or, uh, uh, you know, certainly don't fit within a kind of a column format? Um, Ryan is, is where are we seeing kind of the, the, the, the vanguards of synthetic data are being applied? Yeah, that's a good question.
And the, you know, from, from my point of view, you know, um, uh, you know, the organization I work at, being in the, the financial services, uh, organization, you know, we tend to, to need to make use of more of that structured data. So, you know, we have, uh, um, AI capabilities in our platform that our customers use to be able to do things like, you know, protect their earnings and, you know, for forecast, um, you know, year over month or quarter over quarter, quarter over quarter, uh, um, you know, earnings based on historical, you know, findings with their own dataset. So, and by say their own data, you know, they, the training that is necessary to be able to allow us to, to have any sort of, uh, forecastability into the future, you know, requires us to you to have, you know, uh, uh, that customer, that very specific customer's actuals in play to be able to, um, you know, generate that, that, um, you know, that that structured model that, that we, that, you know, then in turn use to, you know, show, you know, show the customer that crystal ball into the future.
Um, you know, however, I, I, I think it, and the reason I I went into that level of specificity and kinda in detail is because I, I think, you know, the, the answer is, it, it it really depends. And I use, um, you know, autonomous vehicle driving, you know, that that might be, you know, a case where, you know, more of that unstructured training data is, is more necessary to be able to have all the, the variables and permutations of, of conditions that could be encountered when, when you're driving on the road, you know, pedestrians that jump out in front of your vehicle, you know, just started snowing here in, in Michigan, um, you know, adverse, adverse road conditions, you, all of those things become very challenging to represent, um, or, or to, to, uh, to represent in, in the form of actual data. And that's kind of where, you know, some of that generative capabilities tend to, uh, um, you know, come into play to be able to, to, you know, help improve the predictability of those, of those, um, you know, models that are working against that type of data.
Um, and, and then you, you've got a hybrid of the two, I think, as well. And, you know, maybe in, in, uh, um, maybe, uh, you know, an example that comes to mind is, you know, in the medical space, you know, I, I for sure would wouldn't want to ensure to, to make, uh, um, you know, to to have any, any physician diagnosing me, uh, you know, making use of any sort of AI based capabilities to be able to, uh, you know, uh, um, be ensured that, that the data set is being used to generate that prediction is based on some real world models. So, uh, you know, so that that's, you know, some unstructured data that, that probably needs to be trained with more real data and not, you know, generative data.
So, so there's lots of variations of that. But, um, I just wanna kind of compare and contrast the, the, you know, some of the, the, the areas, you know, a alongside your question, um, with respect to industry from my point of view. So, so at the end, Ryan touched a little bit there on, on kind of, you know, use case appropriate application of synthetic data.
Um, and, you know, the, the, the, the topic, the, the space of kind of medical records is, is, is very emotive because no one wants to feel like their, uh, their, their medical records are being, um, uh, being surfaced. However, let, let's be a little bit more pragmatic. Um, it helps when we have the most useful and valid data to improve our model quality and performance.
How much do you see that synthetic data compromises our ability to tune and pro can produce quality, uh, models and, and prediction? Hi, thanks for, for your question. Um, I'm VP of product at Algolia.
So Algolia is a ai, uh, search and discovery, uh, online, uh, serving 17 plus thousand of, uh, customer and, uh, trillion of, uh, experiences. Uh, to your question, uh, can we actually, uh, use, uh, rely fully, uh, synthetic data, uh, to train model, to do tests, uh, and so on and to, uh, anonymize, um, customer data? I think it, it, we really depend on the industry, uh, and, and the type of, uh, of technique you are using first.
Uh, they are structured data and unstructured data. So when it come to structured data, we are talking about table that are very structured. We can use simple techniques such as whole base system to generate new data, and that tends to be fairly mature.
I think, uh, we, uh, generate fairly realistic realist, uh, result today, uh, no, uh, it's used a lot in finance, healthcare, but I will be worried of using this sort of data in maybe some more live critical situation. Uh, typically, um, you know, airplane and so on, uh, where it's still prone to error. We create rules, uh, that are often, uh, defined by human, and there will be errors there.
Uh, another area is, uh, when we use these data, uh, initial data that have bias, uh, where we need to be careful, and again, there are way to go around that by creating a rule-based system, but this is not perfect. So I think we have to be careful and we have to create the right, um, guide in place to test the result and making sure they are really delivering the result we want. Um, no, we also have unstructured data, so it's much more difficult to create rules around that, to create new data.
Uh, we generally use a deep generative model, uh, everyone can experience today, AI creating new images and so on. And I think it's fairly clear that the technology is great, but in term of reliability, consistency, it's still not there. So it can be used in vertical such as fashion and so on.
I will be worried using that in, in the medical space, uh, at scale today. Hey, Rob, let me ask you a question. Most of the organizations that I know, or shall we say data management challenged.
So how do we manage all this synthetic data alongside all the data we're already struggling to manage? And do we keep this data, or is it just get dumped when we're done? Or what is the process there?
Sure, Mike, one thing that we have been doing at is, uh, uh, you know, we treat data like code the same way we have code repositories. Uh, we have access controls, we have regular dataset reviews and even configuration parameter reviews for the models. And whenever a particular data set is shipped along with a model, it is tagged as a milestone model.
So we can always look back at it. And then, uh, in case something needs to change, we use distributed version controls, uh, across different teams. They use different, uh, DVC uh, frameworks that are available.
But one thing that we have seen is we never mix real data and synthetic data. Uh, see, a lot of times the synthetic data might not be newer data as such, it could be just augmented real life data, right? So for example, you have a bunch of images, and then you add some white noise to it, you zoom in, uh, you crop it a bit, and then you keep, so this is the first level of synthetic data.
This augmented original data actually stays along with the real dataset. And then there are totally synthetic datasets where we use techniques like GaN, uh, where we use, uh, so for example, uh, this case where we use an AI model to, uh, to do OCR and invoices and extract, not just OCR, also extract a lot of information, uh, out of it. Now, uh, we had very little invoice information, uh, and these information actually varied very differently across the organizations that we had, depending on the regions they are in, depending on the kind of suppliers they have.
So I know you are doing, uh, document processing on invoices, and then you only have very limited sets. Of course you can do OCR and extract all the information, but, uh, you know, identifying classes like what is the pay by date or what is the PENALITY plus or what is the mode of payment? So that needed a lot of synthetic data to be generated.
So, uh, we have a policy where one organization's data is just used for that particular organization. So these are the specific original data is stored right within the organization storage. And then there is synthetic data, which is common for all.
So basically what we do is we bootstrap our models with the synthetic data, and then the r specific data gets added to it, so the model gets better for that particular organization. So the idea is we keep it separate, um, and then that helps us with the level of privacy. Also, we don't want one organization's data being used for the other organization, but the synthetic data is common across organizations.
This is how we have been doing, Mike. Yep. Yep.
So, so, uh, if, if, if we take that, so Mike alluded to, you know, a lot of organizations are challenged by kind of their data governance and, and processing, right? It's a, that's the, that's a significant overhead even before we even get into the machinations of, you know, best practice MLOps. So they're gonna be looking at this episode and thinking, yeah, you're just adding another hurdle.
You're just adding ano more friction to my data governance process. Um, you know, do I need this, you know, it is scrubbing my data good enough. Do I really need to engage with, uh, the synthetic data providers and tools that are just gonna create a, a more headache for me?
Um, r Ryan is, is that, is is, is, is that reasonable to expect from enterprise? Yeah, that's a, that also is a good question. And, and I had to, uh, I I'm gonna use, you know, one stream's own experiences with this as, um, you know, to, to kind of, you know, to talk through this.
Um, you know, because for, from my my point of view, you know, our, you like what ROM had mentioned, um, you know, the, the, the, uh, the training data we use for our models is based on the customer's actual data. And that's, you know, very highly regulated industries. Um, and our customers tend to be pretty sensitive over, you know, data, uh, not just data, data sovereignty where their data, you know, what, what region of the world their data lives in, but also, you know, ensuring their data's not co-mingled with, um, you know, with other customer data.
Um, you know, so, so that that notion of data privacy becomes, you know, hugely important there. Um, and, um, you know, for, for us, we also have, uh, shared data, you know, like rom that, that we introduce, you know, to be able to, um, you know, bootstrap those models that, and I, I like the, that, uh, that, that phrasing. 'cause that's, that's precisely what, what, uh, what we're doing as well.
Like, I'll use a a for instance there that we have LLM uh, capabilities in our platform that we're looking to, to roll out here soon. And in, um, you know, one of the simplest capabilities is just to be able to ask questions about how to build, you know, scripts in our platform, our product, you know, how to, uh, searching through documentation. You know, that's, that's an example of, of the, uh, of actual data that we use to train that, that LLM.
But we weave that in with the customer's actual data, you know, so now when they start engaging more deeply about, you know, their aspects of their financials or, you know, what, what their, you know, what their, their, you know, top line revenue was last quarter. Like it is, it is coming from their receptacle, their, their data domain. Um, and then you have, you know, the, the, um, the, the notion of, of, uh, generative data, which, which quite frankly, you know, at OneStream where it, it is, it is a new era for, or new area for us to kind of tee tee, uh, tease into.
Um, and, you know, we're realizing, you know, particularly in things like the LLM world, you know, that, that, um, you know, that there's, uh, a, a need for being able to improve the accuracy of some of our AI engines. And, you know, that absolutely will be a necessity. And, um, you know, we're looking at, at, you know, things like, you know, generative adversarial networks and such to be able to actually generate that, that data, uh, and, and, uh, you know, warehouse that, you know, accordingly and, and separately from our actual common data and certainly from the customer's data.
Hopefully that answers your question. But, but, um, you know, it's, it's an interesting world because that aspect of it is, is sort of new to us that, that we're, we're venturing down, uh, uh, uh, you know, pathway kind of, you know, in the early phases right now. Well, I'm gonna, I'm, I'm gonna unpick it 'cause you, you said the R word first, uh, regulation.
Uh, so what does this mean for where regulation needs to happen? You know, are we talking about regulation of the application layer? This is one stream's problem, or are we talking about a regulation of the data ownership layer?
And, and that's obviously your financial, your, your fp and a customers, you know, where does that happen and, uh, and are you even a believer in that level of, of regulation? Um, absolutely. Yes.
I think for our customer's sake, IWI would be remiss if I didn't say, you know, didn't answer that affirmatively. So, um, you know, the, and, and I think it's a, a little bit of both. You know, we have, um, uh, you know, attestations we abide to, you know, or I have to uphold to when it, with regards to, uh, things like, you know, FedRAMP high or, you know, ISO 27,000, you know, those sorts of, of, of, um, you know, certifications that, that we've obtained and have to maintain.
Um, and then there's an element of, of control that's in the customer's arms when it comes to publicly traded organization that has financial data, um, that they, um, that they have to control, you know, they, they, they for, and there's a lot of, of, uh, of controls that the customer has to maintain and put into place, um, you know, for instance, to make sure that, that there's no, uh, you know, shotgunning of earnings or, you know, that there's no, no leakage, you know, before they actually finalize a number and actually disclose it to Wall Street. So they have to be very protective of that data themselves. Um, and certainly when, and, and then our part on that is, is kind of conjoining those two to make sure that the, you know, we help keep the customer out of jail.
That when, you know, say I, I'm, you know, a low level manager at a customer asking for information about, you know, top line earnings and at, at the, you know, end of Q four, which hasn't happened yet, that, you know, there's some security trimming that our product would have to, um, you know, to apply to that data to make sure that, you know, even though that might be predicted and, and maybe the CFO has access to it, I don't, so, so there's, there's an element of, um, of control that, that we have to apply as well to be able to, uh, you know, help keep our customers out of, out of jail. So, so there's lots of different moving pieces to that. And, and I'm just talking about the financial space, it gets even broader when you start looking at, you know, medical and, and safety controlled industry.
You know, like I mentioned the automobile industry. Um, you know, it gets, uh, um, um, much more, uh, challenging. And I would also say, you know, regulation becomes even more interesting.
I'll say, you know, when it comes to generative, um, works, can you copyright that? You know, for instance, if you're using a generative model to produce, you know, some, you know, work of art and you're like, you know, gee, this looks fantastic, and you know, you, you, you, you know, you, you print it and it gets put in a museum with your, your tag, with your name, but it was created by some AI model based on other works. Like, is that something that, that you can actually put your name on is coming from you even though you, you hit the go button on the, to actually generate that work?
Or, uh, and and the same holds true in, in generative or, or code that's generated by ai, you know, when what, what, what are the, the, you know, copyright and, and IP protection rules that apply to that? And I think it's really, we're in early phases to kinda see how that plays out, is I think the industry technology is advancing way more rapidly, uh, rapidly than what the regulations can keep up with. And, and I think it is gonna turn out to be some, some interesting roads that, that we, we venture down, uh, in the, the next year or so.
Ryan, can you clarify a point for me? If you create synthetic data off a customer data, is that quote unquote new data that you've created, or is that something that the customer has to give you express permission to go do? And who owns that data?
We have. So again, speaking to, to, um, you know, to the controls that we have in place at OneStream, we have extraordinarily protective rules on the customer's data. Um, you know, no, no single human being, um, even in, in our operations team that actually manages customer infrastructure, has access to customer data and that, and that is, you know, by design, uh, to be able allow us to uphold to, you know, some of the attestations that, that, um, we have to abide by.
Um, so, you know, that said, um, you know, we ourselves have put rules in place, um, and we honor to ensure that, um, uh, you know, that we can safely or safe straight face to our customers that we're highly protective of, of their financial data. Um, so that said, it makes, you know, that's why I mentioned a sort of early phases when you, you think about, you know, the, the, the generative world being fairly new to us, we have a lot of handcuffs in place to be able to make use of, of some of that actual data. We have immense amounts of, of generative data that we use ourselves, you know, for in engineering world and, you know, to be able to build our, our product.
Um, and, uh, you know, that's, that's largely been, you know, created by, uh, you know, accountants and, and audit, you know, tools and automations that, that work for us to be able to have data that's representative of real world situations that our customers encounter. So we have to, um, uh, you know, we're relying on that for the time being to make sure we have some accurate, uh, uh, you know, data that we actually QA a product with, uh, and then we're looking at, at dipping our toes in the water as well and, and that generative world and kinda seeing what we can do there, uh, to be able to maybe make use of that, that, you know, QA test data that we use and be able to elongate that and, and make for more, uh, you know, a a wider surface area that, that our product can be tested against and, and our models can be trained against. I think, uh, there's a, an interesting point on, uh, on using, uh, personal data to create synthetic data because, uh, at the end of the day, the big argument and that's fairly advanced, is that we create synthetic data to anonymize the user data.
And this give us agency to the QA to train new model because we say these people are not identifiable, but in my opinion, uh, there is still a gray area, uh, which go back to copyright a little bit because this solve one problem. Yes, we can't identify this user mm-hmm. But we are still creating value out of this data without necessarily, uh, having concern from the user to, to create value.
And, uh, and, and yes, the, the data are anonymized, but we are still using all this value to create advertising and so on. And, uh, and to me, there is a big gray area there, uh, data is because gold mine and we are saying we take your gold, but it's fine because it's anonymized. Yeah.
You know what's interesting? Yeah, I, I, I agree it is very much a gray area and, you know, for, for us, our, our, you know, customer privacy agreements are written by, um, uh, you know, written by our lawyers, but then also enriched by our customers lawyers, and we have some very big customers who have very large legal teams. So, um, you know, those have, you know, that's, that's been in effect our digital handcuffs of sorts to be able to really restrict our ability to do that sort of thing.
Um, and, and arguably, you know, a lot of those contracts were done, you know, five, you know, 10 years ago, um, and, and you know, and, and renewed kind of along the way. But, you know, I, I think that too is something that might, might change, perhaps, hopefully to be able to give a little bit more flexibility. But, um, but it, it is definitely an interesting in, in gray world, you know, particularly when you think about advertising data like you mentioned.
And, um, uh, you know, we're all using products publicly, um, that, that are free to use. And really the, you know, the currency that we're using to pay for that is our, our, our data that we provide those, those organizations. And that's where, um, you know, we're kind of giving express consent for our data to be used in that way.
And it's, it's a little bit different when you have, you know, a customer depending on you to warehouse their data and kinda keep them out of jail and apply and abide by all the regulations that they have to abide by, and they want our product to also, you know, help them do that. So it kind of gets, yeah, it is a, it's a tough world to navigate. Yeah.
We have just adding, go ahead. Just adding to that, uh, I mean, even across geographies, I think, uh, uh, there's been a lot of, uh, talk about, uh, from, especially from the technologist perspective, there's a talk about, oh, I, I tweeted this, can this, uh, be actually used to train an LLM? There's a context in which we say things, right?
And that cannot be generalized. So, but all the regulations seem to be addressing the application of AI and the risk that is being associated with it. If you look at all regulations, not many talk about how was the data sourced, or, uh, is it okay to, so I mean, we saw Reddit and Twitter cutting out on their APA access by crazy amounts when the whole LLM wave hit.
But, uh, for example, Japan has said anything that is available on the public internet can be freely trained by a, there is no, I mean, you're not copying it, so the copyright is still valid. You're not reproducing it, but you are learning from it, and then you are putting it out on a different environment. So that is Japan's way of looking it.
But, but then there are other places which, uh, again, huge gray area. Uh, but, but like we rightly pointed out, we are using a lot of free tools. Um, the whole idea being, let's say there's a terminal cancer patient.
Now all of this data, any, any inference or any insight from this data is going to make a day night difference in this quality of life. But now let's say I have my health information and then I have my food order history now that is being correlated and sent to my insurance provider to jack up my insurance premium. Let's say I drink a vanilla vanilla milkshake at 11:00 PM every day, and then my insurance provider knows that.
Now that is not intended, right? My data has to be used to my benefit and any data that is derived from my data. So, for example, uh, people living in this zip code, uh, are generally obese or have less physical activity.
So you can generalize, uh, a synthetic data from the actual data. Now who's the owner of it? Is it okay to do that is still a very big gray area.
I think it's, I just learned two quick things there. One is lawyers are gonna get really rich who specialize in licensing agreements and secondarily stay off social media. But, But I think you make an important point because when we say data anonymization, uh, yes at individual level, but sometimes there are unhealthy segmentation that can be made.
And, uh, while the data is not directly linked to a user, you don't want to, for example, uh, do segmentation on a specific postcode and draw, uh, wrong conclusion or easily are, are creating wrong solution for, for these people. So, uh, that will not benefit them. And I think that's where the danger is as well, is using this data, uh, anonym anonymized, uh, pretending that it solve every problem around the, around the, the use of the original data.
And quite often we use personal data to train models that anonymize this data to train the final model. So we are adding a layer in the middle that is doing exactly the same thing. Et is there concern, is there another concern that, that where synthetic data obfuscates my identity, it also shrouds inherent bias that, that, that is to say, you know, some of the codification, negative codification that's baked into the original dataset also gets hidden by the application of synthetic data.
Yes. Uh, I think it's, it's a very important point. I, I actually did a test recently, not necessarily on synthetic data, but they're all using, uh, generic model and, uh, uh, I went on one of the famous, uh, you know, image generating website, and uh, and I asked, can you, uh, draw me a picture of a powerful person?
And this was returning only white men. And, And, and therefore, uh, that's, I mean, everything is said here, here is, here is the danger. It is strain on data maybe that are old, uh, where at the time, uh, powerful women, uh, were not necessarily as visible and therefore is not a good representation of the reality of today, and it's showing the wrong conclusion.
So here's an example. At the end of the day, all this model are trained on a selection of data that will be biased. I can think of models that may be trained in the US and use in a different continent, uh, and other results, uh, are not providing results that are relevant for that continent.
And I think there will be a lot of, uh, example in this area. It's very important to be vigilant because they are not so obvious all the time as well, and they're slowly creeping into the different AI system. And by the way, the use of data, uh, if the rule are not properly set, could even amplify them.
Yeah. So we use synthetic data in a lot of places just to avoid these bias. For example, uh, you know, you look at a supermarket purchasing pattern and you're trying to identify pregnant people and, uh, normal people.
And then in real life, only 1% of people are pregnant. So if you look at the actual data, it does 99%, the data is skewed on one side. So now we generate a lot of pregnant shopping behavior, and so make it some 60 40 so that the model is able to, uh, detect with lesser false positives.
So I would say, uh, the need for worrying about bias with synthetic data might not be there because synthetic data is only used in tandem with real world data. And, and, and a lot of times this bias filter is very important. In fact, people ask, I mean, you're an IT management company, what do you have to do with bias?
But something like prioritizing a service delivery ticket or assigning it to an agent, we ensure there is no unintended bias because of the shape of the data. It does not creep inside and just because of the pure expertise and the past relationship and all of that. So I would say, uh, synthetic data can potentially reduce biases and then using original data to train them.
Yeah, I think, I think, uh, it's important because, uh, if you use rule-based system data, you can actually on purpose, uh, uh, compensate these biases. But as soon as you use, uh, deep generative model, it become more and more difficult because the system in itself finds relation and there may be bias, you may not be aware. So I think it's, uh, it's important to constantly test this system and not just, you know, your initial assumption may be wrong.
You may unconsciously introduce new bias and so on and so on. So it's always a balance and they need to be, uh, countercheck to make sure that this is not happening. Hey, Ryan, how elaborate are these efforts?
I mean, am I creating like a full boat digital twin of something to generate the synthetic data? Or is this something I can do, um, relatively easily? Yeah, that, that's, that's a good question.
I, and I, I, I was gonna comment just, just something real quickly too on, on the notion of bias that I, it, it, it becomes, you know, I think there's a lot of challenges with that as well when you think about it. Because if you're, if you're commingling, you know, even if it's a anonymized real data, I'll use it in an example, uh, you know, kind of clicking down on what ROM had used on around support cases, um, you know, maybe there's, you know, a thousand support cases and you know, a, you know, 250 of them just use some rough, rough numbers. Easy math, um, are for, you know, the executive team and maybe the support organization, you know, gives priority to those particular support cases and kind of how they respond or what they do to them, your, or maybe they, they engage with them differently.
So now you've got, you know, 25% of your dataset, um, that you're using to potentially train a model to be able to, uh, you know, help support agents be more effective at their job, has got some, um, you know, some slants in it that maybe the, the, the person training the model isn't even aware of. Um, you know, likewise, you know, you, you hear about, about, um, you know, situations where, uh, you know, people of color are, uh, you know, like going back to my automotive example, you know, that, that, you know, that the, the autonomous driving systems, you know, have, have issues recognizing, um, you know, the, you know, people of color because the, the data sets that are used to train a lot of their models didn't, didn't incorporate enough, uh, you know, of a sample size to be able to make the, the, the model that's actually driving the vehicle autonomously be able to recognize things. So, so, um, you know, that, you know, to, to me, uh, you know, when you, you mentioned, uh, digital twin, I, I do think that, that for safety critical systems, I, I, I think that that generating, uh, and I'm using kind of clicking down on the vehicle example, I was just talking about that I, I think in those, those types of applications in particular that are using, are making more use of, of AI-based based models, um, that they need to, um, you know, that, that there has to be a lot of attention and detail applied towards, you know, the type of data being used to, to train those models.
And, and I think up to and including, um, you know, maybe the case of a vehicle building a digital twin, you know, and I know, you know, my son works in a software space, uh, you know, for, for a company that makes automotive, um, you know, camera sensors for, for autonom driving ironically. Um, and you know, for, for them, they have loads of drivers that go out in the road and actually are recording, you know, real world telemetry, you know, dr and, and different, you know, conditions on, you know, cloverleaf pa patterns on highways and things that, that are unique to particular municipalities that maybe, you know, the, the Department of Transportation decided to, you know, try in an area that doesn't exist elsewhere. Like you, you, you can't model that, um, easily because it, maybe, there's only a couple real examples of that that exist.
And, and, um, you know, I think that you've seen, and, and you know, the, we have seen, you know, companies in the automotive space as an example, like, you know, general Motors have super crews that is only able to drive on roadways that they have trained, you know, that they have actually, um, you know, driven down themselves and actually have acquired training data, real training data, um, you know, to be able to help aid the vehicle and, and driving in a, you know, totally hands off mode. So, you know, using that as an example, like that's an extreme. They've taken it all the way to the, you know, to the, the point of, you know, even though they might've been making use of generative AI data to train that model, they only allow the vehicle to go on rows that they actually have incorporated real data as well.
Um, but I do think that, um, uh, for safety critical data, uh, situations, you know, medical situations, I, I, I think that a, a digital twin is, is not out of the realm of, of, of, uh, you know, possibilities to kind of include and, um, you know, the, the, um, uh, you know, the degenerative adversarial networks that are used to, to create some of that data. Okay. So, so we've talked a little bit about kind of techniques.
Um, uh, we've talked a little bit about use cases and applications. Uh, let, let's talk a little bit about stakeholders. Let's, let's, let's try to understand what are, um, our, our data governance roles should look like.
Eddie, and you are probably engaging with some of these people already, you know, what is the, what is the new data officer, CIO, you know, what does that look like and, and, and who should be on their team? Uh, I think, I don't think there is one size fit all answer on that. It'll really depend of the type of organization.
Uh, if you, uh, if you look at, uh, very technical base layer organization, you will have, uh, a lot of engineer, uh, data science and so on. Uh, and maybe on more customer facing organization. I've seen organization who actually have merchandiser, uh, in their team.
Um, and, and, um, when I think about, uh, synthetic data, uh, this is, uh, a lot of the backbone of their work. Uh, and, and one of the big challenge, uh, around this area is, and especially to to scale this sort of technique, is explainability making everything business friendly. And more and more, uh, people working in AI will have to speak to a broader audience than just, uh, engineer who maybe understand the fact that AI sometime can be a black box and so on.
So I think it's important, uh, to have, uh, the entire organization understanding what it means to be AI and data driven, but it means that we need to bring the tool in the organization to engage everyone with, uh, the ability to explain the result, to explain how this, uh, this, uh, synthetic data are generated to build trust in the organization and really deliver that at scale. I'm sure you all have this sort of challenges, but, uh, when we, we talk to customers, they, they believe in ai, but they also want to make sure they understand what is happening so they can make decision. Uh, one example is, uh, to remove bias.
Let me, let me dig in a little deeper on kind of, uh, on et n's response on, on this one ram. What do you think? Uh, well, we've talked a little bit, or we've, we've danced around generative AI and, and LLMs and, you know, an observation certainly from kind of i's point of view is that that has, uh, abstracted away a lot of the infrastructure challenges for organizations.
Um, but those challenges haven't disappeared. They have just kind of been wrapped up in the, uh, in the Pandora's box that is large language models. Now, what we're also seeing is that there is appetite for enterprise to put, kind of push this away to their tech providers, their, you know, their, their services organizations.
What does that mean then for synthe synthetic data as a, uh, as a function? Um, I, is this going to be, is this gonna sit under kind of monitoring et alluded to kind of explain as being critical to understanding synthetic data? Does, does, does synthetic data then become part of monitoring, or is it part of our ETL process, or does it, is it pervasive throughout our entire data pipeline?
Surely. So one thing, uh, this crazy risk, in fact, uh, sometime in 2015, 16, I used to joke, uh, the, the thing about ganis, uh, it is, it's not really useful anywhere in real life, but it is really good for somebody, uh, who is trying to understand the concept of neural networks, understand, uh, what a loss function is, understand how stochastic radi works and all that. But then we saw increasing applications of GaN because on the other side, these algorithms were becoming more and more data hungry.
And then here we are with LLMs, but even then, uh, do synthetic data. Can they standalone go change things? And it, it's all about the quality of the synthetic data, right?
How, how is it, how is it a reflection of the real world data? Let's say a concept drift happens in the real world data, does the synthetic data also shift in that direction? And especially when you're doing things like reinforcement learning.
So for example, again, reinforcement learning was widely used in robotics, but today as AI has made an inroads into a lot of everyday procedures, so for example, uh, we use reinforcement learning in tuning our databases. And, and if you see that the factors there is constant flow of input to the model, and the model keeps getting adjusted as in when, uh, the data comes in. So the model is always light.
It's not like you have trained the model, you put it there for inference, and then every week you go update the model with the latest parameters. Now everything is, uh, you know, real, it, it's happening real time. So same with robotics, right?
The robotics, when a neural network model is getting trained, a real robot is not going and picking up an object. But then the reward function is built in an imaginary space where now let's say your center of gravity is low, you'll fall down, your center of gravity is very high, you'll fall down. So you reward is to maintain it in a particular range.
So these simulations are happening real time and also there is a lot more emphasis. In fact, recent times I'm seeing a lot more papers on improvements on the algorithm to reduce the need for data consumption, right? So, so the elements are based on a 2017 paper, and that's a far six year away.
So, and, and they say we have exhausted all real data sources, right? So now most data is LLM generated or LLM summarized, especially the ones that we see on the open internet. So we have exhausted our data sources and there is some new hope, some shoots, green shoots coming out where you lay more power on the algorithm than the amount of data that is needed to train the algorithm.
So summarizing that, I would say synthetic data is here to stay, but will it be a game changer? Will it, uh, will it really shake up things? I am not sure, but it still will be a very, very important component to your AI pipeline.
Hey guys, guess what? We are, uh, coming up on our full hour here and we're gonna have to check out at the moment, I think I just learned that I'm gonna need an AI model to manage the data I'm using to create the AI model. So that's about how I got that far.
But we'll see where we are. Um, I wanna thank all our panelists for being on the show today. This was great, Ian and Ram Brian, awesome insights.
We could talk about this for hours to come. I wanna thank you all for watching this episode. You'll find it on, uh, tech Strong tv.
We'll also find it on Techstrong ai, and I'm sure our friends at Lee Baker and the Alliance folks will also be promoting it as well. So we look forward to seeing you all again next time. Take care.


