Dangers of Data Bias – Techstrong AI Podcast EP9
Amanda and Mike speak with guest, Cindi Howson, the chief data strategy officer at ThoughtSpot, about the dangers of data bias and the importance of including women in building more inclusive AI models.
Transcript
Hello and welcome to the Techstrong AI Podcast. I'm Amanda Ani, and I'm here today with Mike Vard. How are you doing?
I'm doing great. How are you? Doing well.
We have a lot of AI news and topics to cover today, as always. And so with that, I'll start with our first topic, which is the United Nations has been making some AI related resolutions. And basically, to sum it up, they're saying that if the AI cannot be controlled and stay within compliance, uh, then they're not to use that AI tool.
So, um, just trying to ban the use of systems that can prove to be violators of human rights. So, um, that seems a little difficult. What are your thoughts?
Yeah, that's the thing. It's gonna be too hard to actually execute that 'cause, um, you know, who defines what human rights are? I know the UN tries to, but, uh, human rights in China is dramatically different than it is here in the States and other countries, and nobody has a consistent approach.
So it's unclear to me how I'm gonna ban an AI engine to prevent it from doing something that, um, theoretically it, um, uh, may be allowed to do in some countries, in other countries, not so much. So I feel like this is the latest example and, uh, regulatory concerns around ai, and I'm kind of a thankful that they're interested. But b um, it's pretty clear it's early days.
The story in question here up on the site on Techstrong AI also talks about, um, this document that's being circulated by the United States State Department, where it's worried about, um, the rise of AI that we wouldn't be able to control. And what are the implications of that? Now it looks like it was written for them by an AI consulting firm.
Um, there's, you know, definitely gonna be some issues here in terms of, well, yep, things may spin outta control, but it's not like we don't have the controls in place to monitor these systems. And I know it's a, it, it sounds a bit flippant, but, you know, worst case scenario, there is a power cord somewhere that can be pulled. Um, and then, you know, second part of that question is, um, these systems also have controls around them, right?
So we monitor them and we see what's happening. We may not understand everything that's going on in those systems, but they, you know, it's not like they're just being spun up and tossed out there and saying, good luck. There are, you know, people actually monitoring these things and staying within reasonable amount of controls for them.
And I'm not quite clear, there's multiple flavors of where AI is gonna go, right? There's advanced systems and then there's, you know, super intelligence and super intelligence is theoretical as far as I can tell. There'll be more advanced AI models, no doubt.
And we're seeing some of those already where the reasoning engines are getting much more sophisticated in these large language models. So we'll be on, you know, we'll be able to automate more tasks per se. But, um, I don't think I'm at the point yet where, you know, terminators around the corner.
Yeah, and, and as you said too, it's very hard to get, um, worldwide acceptance of this because who's to say what's acceptable, what's not acceptable across the different countries? And so that's gonna be difficult. And then again, how to con how to really control it.
It, it's, it's still a, a big work in progress. Now that doesn't mean that, you know, the regulators should ignore what's going on because, um, honestly, you know, left to their own devices and their pursuit of greed. Some companies will do some incredibly stupid things.
Um, so we, we need somebody to kind of provide a little oversight on the other end of this. But I think the challenge is, is always trying to strike that balance between reasonable adult supervision and the motivation folks have to drive investments in innovation and, you know, the things that drive a capitalistic economy, which, you know, we all stand to benefit from. So, um, I'm hopeful cautiously about where these things are going.
And, um, you know, you're, they're are folks who are coming together to have these conversations about how to apply AI judiciously. There's conversations between, you know, Microsoft execs and, um, a various agencies. But yeah, is there a cadre of AI hardcore folks probably running around northern California somewhere who were like, devil be damned, and we're just gonna throw out at everything.
Absolutely. And do they host their own little parties out there? Absolutely.
And do we need to keep an eye on what they're doing? Most definitely. For sure.
Well, with that, I think we'll move on to our next topic, which is unconscious bias. And we talk a lot about bias in the AI models and unconscious bias is, uh, quite the con quandary, because if it's unconscious, then how do you keep that bias from entering into these models when people are training the ai? And so, um, the suggestion from this article on our site is, um, to have a good diverse team.
That way you're bringing in a lot of different viewpoints when training the models, having, um, uh, knowledge, uh, sessions and, and classes to kind of just focus on, um, bias as an issue. And, and, and then the other idea is to use tools, automation tools, but again, any kind of technology you're using or automation tools that you're using is made by a human. So again, how do you remove that, that bias?
So there's still some concerns how to do it, but it says 66% of organizations experience data bias in the models, and 78% are concerned it will just get worse. So what are your opinions? The devils in the data.
Uh, the models themselves are, you know, a mechanism. The where things are going wrong is that the data being used to train them is stuff that's coming out of existing applications. And a lot of those, um, analytics applications or whatever it may be, or deeply flawed.
'cause the way we created the app or the way we collected the data was fundamentally biased. For example. Um, it is not uncommon for, uh, there to be some sort of redlining of real estate areas, right?
So if you're unaware of that activity and you're just throwing a data analytics application together, you'll conclude that you know, you should never buy a house in this area because, you know, it's just valuations are horrible. And, uh, and then you start making business decisions based on that, or your personal decisions and demand for housing drops. And suddenly the valuations of those housing declines 'cause demand went down.
And it's all baked into the fact that, you know, somebody decided somewhere that, you know, they wanted that outcome. And, but that was happening long before ai. It's just being, um, highlighted and maybe augmented or accelerated in the fact that, you know, we're using these AI models.
Uh, as one person once joked to me, he said, you know, it's one thing to be wrong. It's another thing to be wrong at scale. And that's the issue with ai.
Absolutely. And, and the problem is that, uh, data can always be kind of manipulated to, to read how you want it to and to give the outcomes you desire. So that of course is a problem.
So I think having a good diverse team and having, um, you know, maybe some training in place to really focus on the issue, I mean, that's a great move in the right direction From your ellipse to God's ear. All right, moving on. So we have another interesting article on tech strong ai and it's, um, the Chief AI officer for Solace, and he gives his five AI tech advancements, uh, that he thinks, uh, are really useful u utilizing ai.
So the first one is transforming analog events into digital events and using event mesh to distribute the event information among applications which we'll streamline and, uh, make faster intelligence and decisioning. Um, so moving on, his second one is productizing AI into applications for customers. His third one is shortening software delivery cycles, which we've talked a lot about that one.
Um, AI for platform engineering, which is gonna require more platform engineers to design. This is kind of in the future. Um, AI being able to do this and then integrated enterprises because people, the problem with this though is that people are too reliant on siloed or legacy systems because only 12% have all their data connected across all their departments.
So, um, using AI to integrate enterprises. So that's his top five. What are your thoughts?
I actually love this art and it is very rare that I kind of get to that 'cause I'm getting old and crunchy, but, um, it was really well thought out series of steps that you gotta look at if you're gonna operationalize AI in a meaningful way. And far too many of us are just kind of like building or customizing an LLM and throwing some data at it and marveling at the outcome without understanding the implications of what it's gonna take to run those at scale consistently in a way that delivers inaccurate output. And so he did a nice job of kind of laying out all the things that gotta get addressed and modernized and updated in terms of our IT operational processes and the way we think about the interrelationship between everything from software artifacts all the way up to, um, you know, ultimately what is gonna be real time communications in an event-driven model.
Um, how are we gonna monitor that because we're already seeing a scenario where frankly, you know, events are happening faster than we're able to comprehend as humans, and that's even without ai. So with ai, um, things are gonna be happening in, I don't know, super sub millisecond. So it's not gonna be possible for humans to keep track of the interactions that are occurring across multiple systems in a way that would be enable them to intervene in real time to prevent an outcome.
So we're gonna need other systems, maybe we're gonna need AI to manage the ai, but we need to think all these things through. And I thought this article was a great step in that general direction. Yeah, me too.
I mean, humans only have so much capacity, so if they wanna go past that, they do have to use some sort of augmentation tools. And AI is a great one for speeding up a lot of these processes and, and catching things really quickly. So, um, it was really great.
So I encourage you all to go to Textron AI and find that article. Um, and next on our list we have so much to talk about today. So next, um, I'm gonna kick this one to you, Mike, but, um, am I saying this correct Tab, tab extend Tab tab nine, I think?
Yeah, tab nine. Okay. You wrote an article about tab nine extending gen AI platform for writing code using natural language interface.
Can you share a little bit about that? Well, what's interesting here is they developed their own large language model early, but now they're also adding support for the open AI one. And I think the other one is mistrial that they built together or some customized version of that.
Um, what it points to though is how rapidly LLMs are becoming commodities, right? We're gonna be able to, um, swap these things in and out and maybe even orchestrate prompts across multiple LLMs that are optimized for different tasks and they're just, they're all gonna be either a, an API call away or b embedded in the application somewhere. Um, it's just been remarkable.
I mean, I think this time last year we were like, LLMs are the coolest, most advanced scientific thing in the world, and now we're kinda like, you know, there are an API call away for, you know, prompting and prompt out, and it's, it's what you do with it that matters. And the LLM itself has become kind of, uh, you know, oh boy, here's another one. And it's got more parameters than the last one.
But it's not like I'm going, holy moly, Batman, there's a whole new way of thinking about this. It's more like, you know, uh, a rapid evolution. But it's just been an amazing advance to the point now where, um, everybody's gonna invoke an LLM.
You'd be silly not to have it in your application. You might even be considered archaic if you don't have it. Oh, absolutely.
I mean, there's thousands and thousands of them now. It, it's, it's really amazing. But I love this because it, it does, um, uh, it does reduce the level of skill and I know that's one issue that we've talked about is, um, the skills gap.
So by using the natural language and AI to write this code, that's great and it's, it's gonna save a little bit of manpower down the road. Yeah, well I hope so. It's not clear to me that these LLMs are gonna, I know some folks are like, well, it'll just be like the cloud and we'll have three at the end of the day and we'll, they'll all be delivered by a cloud service provider.
I think if you look at a lot of these things, there're uh, there's a lot of open source ones. You can pretty much start using them and customizing them if you wanna host your own or if you just wanna expose your data to it, you can do that either in a data center you run or through a cloud service, or even, it doesn't even have to be one of the big three, it just has to be anybody who has enough compute power. Um, so I just think that, you know, this is all gonna be standard issue equipment before, before, well before the end of this year.
Yeah. And that's really incredible considering, you know, I feel like it was only about a year and a half ago we start hearing about OpenAI then from there everything just explodes technology wise. So it's pretty amazing.
Yeah, I think what's really, what's going on there is the, uh, the innovations were occurring long before OpenAI showed up is, you know, what they did is aggregated everything into a large, um, compute service that made it accessible. But, um, you know, it wasn't like they were the only ones working on these models. And, uh, and as a result, we're now aware of many more of them and their capabilities are all slightly different.
And you know, I think ultimately there's gonna be a lot more reliance on, um, narrowly focused LLMs that to automate a task. And I'm gonna orchestrate a lot of them to together to manage a workflow. And I'm, I'm not sure that, you know, the world is gonna centralize on two or three really massive models that, you know, do a hundred things, uh, shall we say imperfectly, Right?
Yeah. And, and I don't know that we want to centralize on only a few, so, you know, use the tools that are best. Alright, so the last topic of the day is the schools and educations as it re in regard to ai.
Uh, so there's still some issues with AI in the school system and how teachers are gonna respond to students using ai and then how teachers themselves are using AI for work and different tasks. Um, so some of the, um, some of the concerns are data privacy and you know, what students are putting out there, what, what staff is putting out there when they're using this ai. Um, and then does there need to be some policies in place for if students are using ai, how and when they can use this ai?
And then how do teachers respond? What's gonna be the repercussions if they use AI in a way that they're not supposed to? Mm-Hmm.
So what are your thoughts? I'm trying to navigate this one. 'cause there's nuances here, right?
So it is a concern that if I'm just using chat GPT to help write my essay, am I really learning that topic or am I just kinda quite literally cutting and pasting something, using it in advanced technology to create something that I will use to pass the course. Um, and so, uh, you know, some folks I've talked to on the professor side are like, you know, well we're just gonna kill the essay altogether because it doesn't really reflect anything and it doesn't reflect the knowledge of the person that we're trying to test. So they might just move away from that and we'll use other mechanisms to test the level of knowledge.
That said, um, there are tons of folks out there for whatever reasons, don't have the best writing skills. Um, we'll never be great writers. They're just not designed to be that way.
I mean, you and I write for a living, it's like breathing, but for a lot of folks it's like, you know, they'd rather have their nails pulled. Um, so, but I think these platforms enable them to articulate ideas that previously they would not have been able to articulate. And I think that's an important service.
Um, a lot of folks have some great ideas, they just don't know how to express them entirely and this gives them a tool to do that. So in that regard, um, I think it can be awesome. And I also think it has huge implications for areas that can't afford the best teachers and the kids aren't always at the highest level of academic performance.
And this gives them a mechanism to kind, you know, take the, take the imbalance out of that system a little bit. So it's not perfect, but, um, I, all things considered, I'd rather have it than not have it. I agree 100% because I feel while yes, there are always gonna be people using AI that are just kind of essentially using it to cheat and they're not learning, they're only hurting themselves if they're doing this.
But I think for the most part, most students are embracing it as a tool. And it is teaching them, it's teaching them how to write better. It's helping them to put their thoughts out better.
Especially when you think about English as maybe it's their second language, it's not their most proficient language and they need, and they're really, really smart, but they, it's not as easy for them to write in English. So using AI as a tool, um, for writing, using it for research and for information just like any other tool or any other way that you would do your research or get information. Um, and, and I think there's also, there are some tools out there that teachers can use if they are concerned about overuse of ai.
There are some tools, I don't know how well they work that, you know, apparently, um, can show if they used AI for the whole thing, um, yeah, how well they work. I'm do, I'm dubious as to whether those things actually work, but I think if you're a professor, you should be able to tell if somebody, you know, just quite literally went to chat GPT and Complete, had to complete the assignment for it. 'cause it, it, it, you can feel it in the copy and the flow.
It's very mechanical. Now do, if somebody's using that and then they're extending it and customizing it and they're just using it as an outline to get started and then they create something from there, I'm okay with that. I mean, it's essentially a, a research tool at that end.
It's not much different than, you know, running a Google search and copy and pasting articles and citing them in your paper anyway. So, um, in that regard it's okay. But, um, you know, the, you should have some sort of, uh, sniff testing your, in your mind as you read something and keep it in mind.
And, um, I guess the, the other side of that too though is outside of school or anywhere else, it's just too easy to create content that is, shall we say, less than reputable. And, you know, we all gotta get better at kind of just assuming that whatever's presented to us, or at least be skeptical of it immediately as to where it came from in its providence. And I think schools will probably be the, uh, first place that might need to apply that skepticism on a daily basis.
And, um, and then hopefully, you know, maybe we'll teach the kids not to believe everything they read and see either. Absolutely. So do you think there should be a blanket policy across the school or should it be left to the teachers?
I Think we should teach these kids how to use this stuff responsibly and, and, and what makes sense now. Um, if they're gonna just copy and paste something, to your point, yes, they are hurting themselves, but they're also hurting other kids because, you know, if there's a grading system that's on a curve or whatever, you know, suddenly they're higher up the curve than they should be. So there is this, uh, issue about, you know, who really knows what, based on how we test and grade, but who knows, maybe the way we test and grade needs to be looked at because uh, maybe we're not really having the right mechanisms in place for doing that anyway.
So ai, we may force a larger issue that is probably long overdue. Very true. And with that, I think we've come to the end of our topics for today.
Do you have any final thoughts, Mike? No, I just think that, look, it is not like the genie's going back in the bottle ever. It's out there.
So the challenge and the opportunity now is to figure out how best to apply it responsibly. I guarantee you that things are gonna go wrong somewhere. And the question is, is you know, can we minimize the impact of those things while taking advantage of the good things?
'cause it's no different than any other innovation that we looked at in history and past, whether it was, you know, firearms to the invention of fire, but there's good things and bad things about all these things. It's just a question of how we control them. Well said.
Well, I wanna thank our audience today for tuning in and if you miss last week, go in and catch last week's. We'll be here every week. Have a great day.
See.