Understanding Tokenization for Data Security with Leon Bian
Leon Bian, vice president and head of product for data security at Capital One Software, explains why tokenization has become a critical tool for securing data in a way that still makes it easily available.
Transcript
Hey guys. Thanks for the throw. We're here with Leon Bian, who is, uh, a vice president and head of product for Data Security Solutions at Capital One Software.
And we're talking about tokenization and the age of AI because, well, it looks like we gotta figure out how to secure that data better than we have been. Leon, thanks for being on the show. Thank you for having me, Mike.
The concept of tokenization has been around for a while, so, um, what makes it difficult? And you would think maybe we should have done this for every piece of data out there, but it seems like people are challenged with managing it. What's going on here and how does AI kind of exacerbate this issue?
Yeah, so, uh, it's a great question, and if I may step back, uh, what we are seeing today is, uh, there were three trends. The first is the explosion of data, as you mentioned, with the explosion of data, we are seeing the explosion of data breaches. Uh, the, the second f uh, force, uh, or trend that we are seeing is that, um, in know worldwide, we're seeing so many different complex, um, privacy related regulations and laws.
And the third, um, trend that we're seeing is the wider adoption of AI and generative ai. With these three forces that we're seeing, protecting our data, sensitive data is becoming more and more important. Tokenization, as you said, has been around for quite some time and different, uh, there were different use cases in tokenization.
Uh, there's one use case, which is what we are going to talk about, protecting sensitive data, uh, by tokenizing the data. The second use case for a tokenization is really in the blockchain space, uh, which is the tokenization of digital assets. And then there's a third context, which is, uh, tokenization in the, uh, in GN ai, but we are not gonna deal with the, the latter two.
What we are focusing on is the first use case, which is protecting sensitive data with tokenization. And, uh, I wanna bring up three key aspects of tokenization, why it's an effective tool, uh, for protecting sensitive data. One is tokenization has been tested, um, over the years, and it, it is secure.
Uh, the second is tokenization is reversible. Uh, unlike some of the other tech technologies like, uh, uh, masking redaction. Uh, if you need to de tokenize the data to get the original value, it's available.
But the third, um, uh, characteristic is it's format preserving, uh, in the sense that if this is a social security number, we can create a token that looks exactly like a social security number, uh, thus, uh, pre preserving the, some of the analytical value, um, in data analytics. So, for example, if you're trying to do a search on tokenize the data, uh, you don't necessarily have to de tokenize the data, unlike some of the other technologies. If you want to do, join two tables, you can join two tables on tokenize the data as that.
So in, by, uh, preserving the analytic value, we don't have to de tokenize the data every time. So we believe tokenization is a very effective data protection technology, but, um, in the grand scheme of things, it's one of those technologies. We also have to, uh, have effective, uh, uh, access control, for example.
Right? So it's one of those technologies that we believe, uh, will help us safeguard our, our data, especially in today's AI world, Is that the primary reason we don't make greater use of tokenization is that other approaches, I had to kind of basically de token it every time I wanted to use it. So everybody kind of decided that was a little too much overhead.
And how did we get past that issue? Yeah, so I, I believe there's a bit of education, uh, uh, that, that will be needed, right? Uh, there are different types of, uh, data, uh, cation technologies.
One is, uh, redaction, or masking, uh, that's one way street. Once you redact the data, it's not coming back. Uh, there's encryption, which is also very widely used.
The problem with encryption is, uh, mo in most cases, it's not format preserving. So a, a social security number could become, uh, 12 characters, 16 characters, and it's completely unrecognizable. So you have to change the schema in a database, right?
Um, and then, but tokenization is a, a technology that has, has been tested. It's, it's widely used in the payments industry, uh, to tokenize credit card numbers, for example. And at Ca Capital One, we make very wide use of tokenization.
So, uh, as we are talking to more and more of the enterprises, uh, you know, more and more enterprises are realizing, hey, tech tokenization is a very effective tool, uh, for protecting their data. And, uh, you know, in my view as, as we discussed earlier, uh, some of the attributes could help, uh, a company, uh, unlock more value in tokenized data as opposed to using some of the other approaches. You mentioned encryption, so let's just go there for a minute, but we're all kind of worried about the post quantum world.
Yeah. If I move to tokenization, is that a way to kind of deal with that issue without necessarily having to go back in and kinda replace every encryption algorithm we ever created? Yeah, so post quantum, uh, it, uh, in of post quantum world, obviously we'll have to be very careful, uh, you know, about, uh, attacks on encryption, right?
Um, but NIST has led the way to create some of the, uh, post quantum crip, uh, cryptography algorithms. Uh, so we also incorporate, when we use, uh, our tokenization algorithm, we, uh, looking into, hey, you know, how do we, uh, make quantum proof? There are two types of tokenization, uh, solutions at the high level.
One is a voltage solution. So in a voltage solution, a token is kept in the vaults. There's, you know, you have to go back to the vault to look up the token, right?
So you just, it's just about, Hey, I have to safeguard the, the vault. But the tokens in the data lake, or in other databases, there are tokens. You can't, even with a quantum, uh, computer, you, you cannot reverse engineer it.
Uh, but another, um, tokenization algorithm, which is called vol list. In that particular algorithm, which is what we are using today, um, we have also incorporated post quantum cryptography to safeguard those to make sure that we are safe from quantum, uh, computing attacks. So it's gonna be a mix of things ultimately.
And the more we have, the more secure we are. Yes, exactly. The more, more layer, right?
So even without organization solution, we also have another layer of encryption on top of it, which is a quantum proof encryption. I'm not sure a lot of folks are familiar with the Capital One software business unit. So kind of give us a little history of that.
And how did you guys come to be, because, well, most people think of Capital One as strictly a financial services organization. Yeah, absolutely. Uh, capital One software was officially launched exactly three years ago, and our first product was, uh, Slingshot.
And, you know, capital One software is the B2B enterprise business, uh, from Capital One. Uh, the main entity, and, and one of the reasons that we, um, decided to launch Capital One software was that ca at Capital One, we have developed a lot of great data management and data security technologies for our own consumption. And Capital One, if, if you say, you know, this is a bank, but half of the Capital One, we have over 40, 14,000 engineers working on different technologies.
And at that time, we realized that, hey, you know, some of the technologies that we developed in-house could actually benefit other enterprises. Uh, so that's why we launched the first product, which is Slingshot. Uh, it's a cost optimization product on top of Snowflake at the time, but now it's also working with Databricks and some of the other, uh, data processing platforms.
And then the second product that we just launched in April is called the Data Bolt. It's a tokenization solution that was designed to address the, uh, enterprise's most pressing data security challenges today. When you think about all of this, especially in the age of ai, do you think there's maybe a newfound respect for data management, maybe by extension data security, because we've had these issues forever, but I always felt like they were kinda, you know, swept under the rug a little bit.
Yes, uh, absolutely. We are looking at this area a lot recently, right? Was the explosion of ai, especially chain ai.
We were talking about, uh, chat, GPT and other, other large language models. One of the changes, uh, between now and three years ago is we're using more data, right? Uh, we're using more data.
The data has to be ready, the data has to be protected. And the question is, how do we unlock the value of data without compromising data security? Uh, and that, that is why we're, we're looking into, hey, you know, we need more data management.
We need more, uh, data security, more data governance. Uh, we need the tools, we need the policies, uh, and and so forth. Uh, in order for us to, uh, unlock that value of data, uh, to be used for ai, How are you seeing the folks who manage data and the folks who manage security kind of bringing or converging their efforts?
Because historically, they kind of, you know, in a lot of ways just were two ships sailing in the night past each other, and they didn't always, you know, collaborate. Yeah. Now they actually work very, very closely together.
And, um, previously I worked at another large software company in the financial services sector, and now I work at Capital One. Um, and, and more and more on the data data side that, you know, the CDOs and, and the data platform heads, uh, and even the analysts and so forth, uh, they have more and more awareness of the need for data security. And especially when we use gen ai and when we run machine learning models, we wanna make sure, uh, our, uh, customer data, our data sensitive data is actually protected, right?
So, so, so that, that is why, you know, this is a trend. Um, maybe five years ago we were seeing, uh, cybersecurity folks and, and data folks, they were working in different silos. But what we are seeing right now is they're working more and more closely with each other.
For example, if you are a enterprise, you're developing a data platform, you need to have the data security, uh, the governance, the workflows, uh, to make sure that we embed data governance policies and standards into the data workflows. And how do we protect our ai, uh, models, uh, the pipelines, the APIs, and the underlying training environments, right? So all of these are new attack surfaces and that we need to protect.
And, and that, that is why the, on the data side, folks are more and more realizing they have to work very, very closely with the cybersecurity side. So what is your best advice to folks about how to get started with tokenization? And I ask the question.
'cause if you're not familiar with it, this whole area of data management and data security is a little intimidating. So where do I get going? I, I think I, I wanna make a recommendation of three steps, right?
The first step is, uh, inventory your data. The second step is embed, uh, security controls in your data governance workflow. And the third step is to continuously, uh, monitor and observe, uh, data access patterns.
So let, let me just come back to the, the, the first step, uh, which is inventory. We have to know, uh, where our, uh, what kind of sensitive data we have collected. Uh, who does the data belong to, and where the data resides, right?
So we have to inventory all of that and also, uh, classify the data of what the sensitivity of each piece of data is. The second step, as I said, is to, um, embed, uh, uh, security controls into, uh, our, uh, data governance workflow to make sure that the data is protected, and we only give access to the people who needs access and to ensure that we follow the principle of least, uh, privilege. And then finally, we need to leverage, uh, you know, AI automation to continuously mo to monitor our data security posture in our environment, detect anomalies, and making sure that we are always on top of data security in our data environment.
To your point, do it, people need to have a better understanding of what the data actually is, because historically, you, they process it and stored it, but I don't think they actually thought too much about what the data was and how sensitive it was and where it needs to be. So is this just a broadening of their horizons? Yeah, and that's a, that's a very, uh, big challenge, obviously.
Uh, you know, my understanding is, uh, you know, five years ago, a lot of the data, uh, was still manually, uh, uh, processed. And, uh, we, we wrote, uh, you know, we classified the data manually and, and, and so on and so forth. But, but today there's a big need to know the sensitivity of the data, especially in the AI world.
So, uh, my understanding is more and more enterprises, uh, US Capital One, uh, included that we are trying to inventory the data, classify the data, um, upstream, and then when we use the data, now we can put control, um, into the data governance workflow, to, to make sure that people have access only to the data that they need to access to do their work, right? So, uh, it people, uh, absolutely now need to, uh, understand better, understand what the data, uh, uh, that have, uh, they, they have in their databases, in their data lake, and, uh, what kind of data, who the data belongs to, and what's the classification of the data. And to your point, the cyber criminals today seem to have access to everything and anything, and they're just logging in.
So if I don't secure the data, I kind of don't really stand the chance. 'cause it's depending on all the stuff that we did at the end point, and the network edge alone is no longer enough. Yes, absolutely.
Uh, you, you are right. And that's why we need a, a multi-layer security approach, right? So even though we, you make heavy use of tokenization, we also need to have, uh, the proper, uh, uh, identity access management, right?
Uh, to make sure that the, uh, only, you know, as, as I said, only the, the, uh, uh, employees who need to access the data, have access to, to the data, uh, we might need to add, uh, disc level encryption on top of tokenization. So my point is we need to leverage multiple layers of the data protection, uh, methods and, and technologies. And then on top of that, uh, we have to educate, uh, the employees and the humans.
And I always say there's this human hack, uh, factor, human risks involved. So how do we train, uh, our employees to make sure that they don't fall in, uh, to, uh, uh, phishing attacks, for example, or social engineering attacks, right? Because if, if somehow, uh, you, you give away your credentials, then, uh, no matter how, how much access control, how much technology we are using, uh, it, it, it won't matter.
We still, you know, the door, door is open, so, uh, we have to employ a multilayered, uh, approach to, uh, data protection and data security. All right, folks, you heard it here. Data security.
It's not only the last line of defense, arguably it's the first line as well. Hey, Leon, thanks for being on the show. Thank you, Mike, for having me.
All Right. And back to you guys in the studio.