How AI-Generated Deepfakes Are Fueling Executive Impersonation Attacks
Jim Brennan, chief product and technology officer at GetReal Security, explains how cybercriminals are using AI-generated deepfakes to impersonate business and IT leaders. As synthetic media becomes more sophisticated, organizations face heightened risks from targeted social engineering attacks, requiring new detection strategies and stronger identity verification controls.
Transcript
Hey guys, thanks for the throw. We're here with Jim Brennan, who's chief product and technology officer for Get Real, and we're having a little chat about, well, deep fakes, but they're coming to the enterprise now and it looks like they're targeting specific individuals who maybe have a, a lot of power and a lot of net worth. Jim, welcome to the show.
Well, thank you Mike. Really glad to be here. I appreciate the opportunity.
I think when most people think about deep fakes, they're like, well, some politician somewhere was impersonated or that it was, uh, you know, maybe some actor or somebody who was saying something about something. But it seems like lately the bad guys are getting pickier about who they go to, the trouble that create these DeepFakes about, 'cause they're after, well, you know, senior level execs and companies that have access to money and workflows. So what's changing here?
Yeah, that's absolutely the case. You know, it's actually, there's a couple different types of deep fakes to think about in this conversation. So one is, is a case where you've got maybe an image, audio or a video file floating around, and maybe to your point, it's of a celebrity or an executive saying something or doing something they shouldn't be saying or doing.
But now what we're seeing is actually the use of this, this same type of technology, but instead taking place in real time forums, much like an interaction like this or a phone call. And so now these, these are actually attacks focused on the enterprise because things from candidate fraud. So, you know, fake job candidates showing up in an interview to somebody calling the help desk claiming they got locked out of their, their Google or Microsoft account.
These are ways in which these technologies are now being utilized. And so they do represent a real threat to the enterprise. How good are they?
Because a lot of folks would assume that there's, through a level of interaction and conversation, it would become apparent that that was a deep fake. And yet we hear about people being fooled for an extended period of time. So what's changed?
Well, it's changing dramatically by the day, Mike. That's, that's the reality. New tools are coming out, the existing tools are getting better.
There are some differences if you're talking about audio or video. So, so for quite some time, for about a year now, audio has been to a point where most people, including you and I, if we get a phone call, likely can't tell the difference. Video is a little bit more complicated.
Of course you've got more signal to work with, but even that's getting better. There's new tools that have come out recently that are just uncanny in terms of how close they can mirror somebody's resemblance In the future. Will we have to validate every interaction?
I mean, before you and I joined on this call, uh, you were introduced to me by somebody we both know and hopefully trust. And, um, as that comes together though, it seems like a little awkward, but is that where we are? Yeah, Well that's just it.
Exactly. You know, and if you think about how, how much time we spend in a video conference or, or taking phone calls and most businesses conducted in these forums, and to your point, there's really a missing authentication layer here. So not only the things you mentioned, but well, what's to stop me from sending the link to this meeting to somebody else that could then pretend to be me?
The reality is you and I right now have no, no real reason to believe that we're talking to who we think we're talking to. Right? I don't really know that I'm talking to Mike.
You don't really know you're talking to Jim. I can assure you that you are. But there's, there's really no technical reason that we should have that confidence, but we're all conditioned to trust what we're seeing and hearing.
That's this, that's human nature. But we're now in this, this realm where that you can't rely upon these signals. So it's a fascinating change to the dynamics and for the last 20 years, business has transformed to again, be utilizing these types of communication forums for very important transactions and conversations and, and every type of business.
And now we're at a point where you really can't trust who's on the other end of that interaction. So what's to be done about this? Do we need some sort of missing technology or is there some service somewhere that will validate our representations to each other in a way that we can at least reasonably trust, maybe never perfectly trust.
Yeah. Well there's a couple different questions that you have to be able to answer though when you're thinking about trust, you know, so, so one question is, is the person I'm speaking with or interacting with, are they actually a real human being or are they some type of synthetic creation? A deep fake, but that's only one, one question.
The other question that is equally important in an interaction like this is, is the person I'm talking to who they're claiming to be, those are two very different questions, right? So you take an interaction, like an interview, a job applicant shows up in a Zoom call or a teams call, right? Oftentimes that candidate fraud that's taking place is not involving a deep fake.
It's involving somebody claiming to be somebody who they're not. You know, maybe they don't have the right skills. Maybe it's a state-sponsored actor hoping to infiltrate a company.
So two questions. Is this a real person and is it the person that I think I'm speaking with? And those two questions, to answer those, you need different types of, of techniques or solutions.
On one hand you need some ability to, to detect synthetic content, audio and video. But then on the other side of that coin, you need some way of verifying consistency of things like facial biometrics, voice biometrics, behavioral, those are, those are some of the techniques that need to be brought together. Then along with bringing in threat intelligence, because these are security incidents.
So what can I know in advance of an interaction about somebody I'm going to be meeting with? Are they exhibiting, are they going to exhibit a face or a voice that is known to be associated with a threat actor? These are all the things that have to come together.
Yeah. Um, when we get to some point where there'll be tells, and I, and I asked this question because early on there used to be kind of the sense of, well, if there was something that somebody wanted you to do and it was urgent, it was kind of a tell that maybe this is the wrong thing to do. But to your point, it also seems like there's a lot more patience being exercised now on the bad guys and they're willing to pretend to be something for an extended period of time before they strike.
So what can, what can I look for? Yeah, so there are some tells today, and I'll go through a couple of those depending upon the, the situation, but increasingly those tells are, are going away. So in the case of of an interview for example, there are certain things like obviously not being on camera or, or using a virtual background and refusing to take off that virtual background.
Um, also in the context of an interview, you know, a lag between when the questions asked or an answer is given could mean that somebody is being fed responses from somebody else or looking something up. So these are things that exist now. They're all gonna go away very soon because all the technology is getting better.
And so really what's gonna be needed is a much deeper lower level detection of the tells and the artifacts. And here at Get Real, that's, that's really our focus. We're taking a very low level look at the actual processing pipelines of these platforms, zoom and teams, et cetera, and, and corporate telephony platforms and understanding what does normal look like.
And by the way, normal's changing every day. 'cause those platforms are themselves utilizing AI to do things like noise reduction. So it's not just a matter of it is AI present, it's what does a normal non nefarious use of that platform look like.
And then you have to be able to couple that with in-depth knowledge of what do the common generator tools that are out there and that are emerging, what do those impart upon that processing pipeline? So you're essentially looking for deviations or anomalies from what healthy, normal, non nefarious activity looks like. And that's, that's our focus here.
We think that's the only way to solve this problem, but it does require a very low level of expertise and knowledge because again, the tells that exist today, some of the things I mentioned, they're not gonna be here long. Mm-hmm. Does that also include, I don't know, measuring latency?
Because, uh, theoretically if I'm talking to somebody and they're supposed to be in Green Bay, but if I'm measuring latency, it's pretty clear that they're not responding and the amount of time window that you would normally expect between New York and Green Bay instead it feels like, you know, New York to Africa, maybe something exists. Yeah, that's a good example. I would, I would probably take that example and, and, um, look at it somewhat differently.
So, so the idea of location being an important indicator, absolutely spot on. That's where threat intel can come in as well, and that's where other context and signals coming from the interaction such as IP address for example, can be really handy. Also, things like understanding if somebody is using a virtual camera driver or audio driver, things that are associated with the use of these nefarious tools, very strong indicators.
And then to, to another extent, maybe a somewhat lesser extent, just again, knowing some, something about the person on the other end of that interaction. You know, somebody's email address, you know, no, what can you, what can you tell from that? Um, are they associated with, with a company?
Do they have a history at a company that you can trust? All these are signals that can definitely play a role. Mm-hmm.
So what's your best advice to folks? 'cause I think, uh, taken to its nth degree, you know, this whole communications revolution that we've been counting on for the last three decades or so might just unravel. So how do we think about this?
Yeah, well, and, and, um, I don't wanna sound alarmist, but, but internally here at Get Real Security, we, we think about the idea of a zero trust approach, but a zero trust at the human layer. So we're all familiar with that term, zero trust in terms of infrastructure and assets and so forth, but we need to apply a similar thinking to the human layer. And again, the human layer is anywhere a human being is being represented in a digital signal.
It could be, again, interaction like this, it could be a profile picture, it could be a voice recording landing in your inbox in WhatsApp, anywhere where a human being is being represented. We think of that as the human layer, and we really do think that we need to apply zero trust to that back to the conversation we had before. Really, there's no inherent reason in today's climate where we should be trusting every interaction.
And so you do have to take an approach like that and then therefore that implies that you have to have some tools that can answer those two questions that I talked about. Is this a real person and is this the person that I think it is? But then you also have to have tools that allow you to respond proactively when there is an incident.
And then ultimately you want to know something about who's on the other end of that attack. These are attacks and, and as is the case with any attack, you wanna understand the intent and the actor behind that attack because if they're knocking on your door once, it's not only once they're probably, they're probably targeting many people within your enterprise. What should law enforcement be doing about any of this?
'cause I guess fraud is still fraud, but, yep. Um, I wonder if the technology has just moved far beyond their capabilities at the moment, but what would you like to see happen? Well, I would say it hasn't, it hasn't moved beyond their capabilities yet, but, and we do a lot of work with government agencies, law enforcement, that does tend to be more in the realm of file or content analysis.
So perhaps evidence to be admitted in a court of law as one example, a similar case with, with intelligence analysis within governmental agencies. And so they're very focused on this problem. Again, we have a lot of, uh, conversations and some relationships in that area.
It's very, it's very quick moving as we talked about. Right. Um, and that one in that, in those cases, it's primarily about that deep fake detection.
Is this synthetic content? Can I, can I trust this? But if it involves a person, back to my idea about the human layer, that's where the approach I mentioned can be very relevant.
Not just understanding is it synthetic, but is this the person that they're claiming to be? Mm-hmm. What's that one thing you see people doing today that makes you shake your head a little bit and go, folks, that's no longer gonna stand.
We gotta be smarter than that. Well, certainly, and this may be as a no-brainer, but it's just answering, answering your phone to an unknown number and engaging in any type of serious conversation or taking any action as a result of that. You know, that, that, uh, maybe that that's a no brainer, but, uh, people still do it.
I don't answer my phone if it's a number. I don't know. Um, and, and to be totally candid, even if it is a number that I know, I'm, I am much more careful these days than I used to be because numbers can be spoofed.
Of course. Uh, so that's one thing definitively, but then I, I really do, I wanna emphasize this area of video conferencing because we all, we live our lives in this little window, right? We live our days and we operate our businesses on the information that we get.
Uh, we can no longer trust that. And so, um, my advice right now is to, to start a, viewing it as such, again, this idea of zero trust. Mm-hmm.
You know, to your point, I don't even remember all the numbers I'm supposed to know. So when I do see a number that calls me, I always wait to let it go through, and then I check to see if I've actually texted with somebody on that number. Just to know that's enough.
Exactly. That, that's exactly the right, the right approach. I think that's a healthy thing to do.
Yeah. All right folks. You heard it here.
Better to be safe than sorry. And all those little things you can do to protect yourself will make all kinds of difference. But we might need more tech no matter what.
'cause tech, fights, tech. Hey Jim, thanks for being on the show. Thanks so much, Mike.
Enjoyed it. All right. And be you guys.
And Steve.