Advancing Data Security for LLMs with Bedrock Security’s Pranava Adduri
Transcript
This is Textron tv. Hey guys, thanks for the thrill. We hear with PMA Auri, who is CEO for Bedrock Security, and we're talking about how to secure all that data that's finding its way into all these fabulous gen AI platforms, but perhaps not being as circumspect with that data as we should be.
Prana, welcome to show Michael. Thanks for having me, and, uh, excited to, excited to jump in. Alright, well, you guys just, uh, raised some funding to go solve this problem, but I guess my first question to you is, how pervasive is this issue become and what makes it more unique of a problem than per se any way we otherwise think about securing our data?
Yeah, I would say that, um, I would say that, uh, if you think about data, um, there's really been two upticks in terms of, um, the growth in data and the acceleration of data being created and how it's being changed. Um, if you think, if you think back to the nineties or you think back to the early two thousands, most of the data that was being analyzed, um, that was being processed, it was mostly being done on premise, right? People would spend a giant clusters on-prem, giant analytical warehouses on-prem, and these were fundamentally constrained in terms of how big they could get, how much data they could process, right?
Uh, the transition to cloud marked the first transition in terms of the amount of data that was being generated and analyzed. And, um, the platforms like Snowflake, et cetera, people could generate a lot more data. They could move it around, they could analyze it.
So if you think about the growth in data, the first accelerator was the migration to cloud. The second, uh, is now Gen ai. And the reason I say that, um, is that when you, when, when an organization decides to implement, arrive, for example, or they decide to train, uh, a, a new foundational model, there's a lot of data going into these models that they're being trained off of.
And then when they're prompted or people ask questions of them, they're generating data based on the inputs as well. So what we're really headed towards, and why I call this the second arc, is we're really headed to an era of a data echo chamber, right? Because there's the data that's already in the enterprise.
Those are gonna get sucked up into these models. When people ask questions, those models are gonna output data that's gonna go back into new documents, right? And there's kinda a cycle that's going on over there now, right?
So we're really headed into an era now of much more data creation. Uh, it's readily available, it's readily queryable, uh, and there's gonna be a lot of downstream derived data that's also created as well. So, coming back to your question, then, um, there's a lot of data that's currently there, and pretty soon we're gonna see a lot more data being created with these models, right?
And so come tying it back to, uh, what, what do enterprises really need? They need a way, they need a way to make sense of it, because while that data is growing exponentially, the resources, uh, the, the, the teams that are responsible for making sure that that data is being used safely, that that data is being used for governance policies, for regulatory, uh, guidelines. And that ip, for example, is being handled correctly.
Those are linear, right? Those teams might be growing at a very steady pace or, or for a lot of companies that are actually not growing at all right now. So there's that gap in between, right?
That's, that's fundamentally the, the, the risk and opportunity that we, uh, that, that, um, we see to help companies out. So, am I trying to encrypt that data or am I trying to prevent sensitive data from being shared without permission? I mean, what are the tactics and procedures that organizations need to think about to secure all that?
Yeah. Uh, that fits all of the above, right? Um, when I think about data, there's really three broad categories of, of controls for, for an organization.
One is, um, you know, the, the, the, if you can, the best is to decommission the data and get rid of it, right? Less data equals less problems. Um, so I would say the first knob really is identify the data that you can get rid of and, and put a tombstone on it.
Um, uh, help people get rid of that data. The second knob is, to your point, exactly as you mentioned, um, the axis of that data. If there are certain data that needs to be kept for retention purposes, because there, it's, it's, uh, bound by certain regulatory needs, um, you can keep it, but you can also effectively air gap it, and you can make it so that no person has access to it.
Uh, you can make it so that there's no network reachability to that data, right? So you can effectively air gap and quarantine data. So that's, that's definitely another control that people have at their disposal to manage the data risk.
And if you can do, uh, neither of the first two, the third is, uh, being able to harden that data. Encryption is definitely one mechanism. Um, different teams have different levels of, uh, different, um, layers at which they apply that, right?
Some people will do it at rest, other folks will also bring it to the application level. So that's definitely one control. Being able to, uh, tokenize that data, uh, being able to create a synthetic variant of it.
All of these can help harden the data that third category as well. So, uh, all of the above in terms of what an organization can do to, to reduce some of that risk as they're consuming that data. Um, it seems like there's a lot of people talking about this.
So from a vendor side, what will differentiate one approach from another? What should people be thinking about? 'cause you know, in the last two months alone, I can't count the number of people or vendors who are showing up and said, you know, we have a way to secure this.
So, um, what should people be looking at here to discern which way to go? Yeah, it's a great question. Um, there are, there are a number of philosophies out there and, and vendors as well, to your point, that are, that are trying to understand data, uh, or trying to work with organizations to help them understand their data and protect it.
A key. The, the, when we look at the problem, there's two, um, there's two attributes to the problem that make it particularly gnarly, that make traditional approaches really not, uh, that make it so that traditional approaches can't really handle the problem. In today's today's era, you know, gene, AI and cloud.
The first one is, uh, the growth in, in the, and the amount of data, right? So there's just the sheer volume of data at play, uh, and being able to keep up with that is, um, is the first, is the first thing. You can't protect what you can't see, right?
So the first bit is can you keep up with the many petabytes of data that enterprises are generating in, in, in today's day and age? The second is the change and new data that's being created, right? And I'll go into an example for this, but you have rapid cr uh, rapid vol increase in volume, and so you need to be able to keep up with that.
And the second is, as new data is being created, you also need to be able to understand how does this data relate to the business? Uh, how does this data, um, matter to the business At the end of the day? Is it important?
Right? And without that, you can't really make a judgment on what the risk is to the business, right? And so both of those, combining both of those and getting those right, that's where we see enterprise is still struggling with a lot of solutions that are out there today.
Uh, and so I'll, I'll get into this. Um, for example, when you're dealing with a multi petabyte environment, a lot of the traditional approaches have been appliance-based architectures, right? Um, they don't really leverage the, um, the, um, distributor architectures that cloud has to offer now.
And so what we find is that when we go into very large environments, very large enterprises that have multi petabytes of data, the amount of time it takes to go over that large volume can be in months, if not, if not bordering on years, right? So the time to value is very poor right now in, in the industry, uh, based on what we're seeing, people are spending a lot of time even getting basic coverage. On the second part, which is the nature of data changing, that's another area that people struggle with because out of the box, what a lot of tools do is they'll try to look for things like PII, they'll try to look for things like, uh, standard regulated data types, and that can help, but it, but in terms of what, for example, um, what a lot of, uh, solutions and data security, um, get a bad rep for is they have, um, they have a lot of false positives, right?
Or they can miss things. And so what a lot of teams are doing is they're spending a lot of human capital calibrating these tools. It's like, um, it's like playing whack-a-mole, if you'll, right, you have a out of the box, the tools might ship with some stencils, circle shapes, square shapes, triangles, right?
Um, and if the business has star shaped data, they're gonna miss it. They're either gonna, they're gonna miss it, right? And someone has to teach the tool that there's a star shaped thing that they need to go after now, right?
And so, both in terms of being able to keep up with the volume of data and legacy architectures not being able to do that. And the second bit is they're rule-based architectures, so they can't keep up with the, um, with the different types of data that the business is generating. So just all around a very large time to value is what we're seeing out there in the market right now.
And at Bedrock, we're changing that, right? We realize that into that today, um, teams that are in GRC and security, their headcounts are not growing. They don't have infinite time to go calibrate these tools.
What they really need to focus on is quick time to value and maximizing return on effort of the time that they do put it. And so that's what we designed bedrock for from the ground up. Uh, how do we rapidly get visibility into the large scales of data?
And as that data's changing over time, how do you, how do you redu, how do you dramatically reduce the amount of effort required for people to keep up with that data? So those are the two core design tenants that we built Bedrock on. Seems to me this issue is what's holding up empty operationalization of ai, because a lot of organizations have gone so far as to, you know, say their employees, thou shall not use these things.
And I'm sure a lot of people aren't paying a ton of attention to those rules, but, um, it does seem like we need to figure out a way to safely use them with the data that is highly sensitive. And so, are people kind of struggling with this particular issue as their first step to figuring out how to take advantage of all of this stuff in their workflow? Hundred percent.
Because in the absence of knowing what you're dealing with, it's hard for people to fully, uh, fully, uh, trust that employees are using that data correctly. And so, in the absence of, of knowing and being able to be confident that the right controls are in place, what organizations typically revert to is, um, just being very guarded about what they do enable and open up, right? And what that does is it creates friction, right?
Coming back to, uh, why we talk about why we envision a world where frictionless data security, uh, if you don't know what data you have, um, opening up that access, people tend to be more guarded, right? And therefore, people have to go through more friction, more process, uh, more approvals to get access to things. And ultimately what that does is it slows down, um, the rate at which people can experiment and create new value with that data, right?
So, um, what we find, uh, and what we're seeing with, uh, with our customers is step one to doing that is, one, keeping pace with the data. And two, as that data is changing, helping the business, uh, have a way of ensuring that that data is protected. And so, I'll give you an example.
If you are in the business of managing HR data, you might have a governance, might put in a policy that says all HR data must stay within the United States, right? Uh, just as, as, as part of our policy. Well, in order to enforce that policy, you need to first figure out what is HR data, right?
Um, well, is it w twos? Is it offer letters? Uh, is it ten nine forms?
All of these can be HR data, right? So someone has to go actually enumerate all of that, and then some other person has to create rules for what matches W twos, what matches 10 90 nines, right? And so where the friction comes in as let's say the product team, uh, um, starts, uh, um, starts building, um, a new product line that captures a new type of employee data.
Maybe it's, um, maybe it's a new type of, uh, uh, form for taxes, uh, or OP or for, uh, equity grants, right? When that new data shows up, if that data starts leaving, or if it's going into an ML environment, right? Or if it's going into something that's not, uh, in the US or not a production environment catching that becomes really hard because the existing rules don't match that new type of data, right?
And so there's work involved, but the product team having to tell the security team, Hey, we're creating this new data, create a new reg export, or create a new rule for it, right? And that creates friction, right? Um, and that might not always be done as well.
And if it's not, that's how data can leak into these models, for example. And that's why you see organizations being very careful about like what they, what they open up, or they might just completely wall off entire data sets because they don't have confidence that the policies are being followed. Um, and so when I talked about how bedrock is changing this, a big part of our air platform is instead of using rules, we use reasoning reasons.
We reason the latest in AI to actually learn what type of data is showing up, how it relates to data that you already know about. So even if there's no explicit rules, you can still make sure that that data's being protected properly. So in the case of, uh, that governance policy that employee data cannot leave, production cannot go into models or cannot leave the US in a rule based engine, it may have missed that new form that, let's say the new form is 83 bs.
I may have missed it with air, we would look at the contents of that, of that new document, and we'd say, we've never seen this before, but based on the contents, it's very similar to other employee data that we know about. Therefore, it should not go into a ML model, right? So that's the key difference in terms of, uh, opening up this data.
If you can have policies that keep up with the data, that reduces the friction in terms of opening it up and enabling people to experiment with it. Do you think auditors are gonna come looking for this real soon? It seems like, um, given the popularity of all these platforms and must be on their radar screen by now, I think we're definitely starting to see, um, this whole landscape has evolved so rapidly, right?
Just three months ago, um, the, just three to four months ago, the audit, uh, people really talking about that, right? People are just focused on how can we rapidly experiment with these models and get something productive? Now the tune has changed, right?
People are starting to get concerned about, well, what's going into my models? And you can you speak about a model as effectively an open database, right? Once something goes in, it's really hard, hard to govern the output of it.
And so people are starting to think about what data's leaking in. And you're starting to see, um, the very first, uh, you're starting to see the first bits of legislation go into effect as well with the eu, for example, and then the iso, uh, 40,000 series that just came out as well. So pretty soon we will start seeing, uh, some of the hard questions being asked and enforcements, um, being, being asked of companies trying to operationalize these things.
Mm-hmm. We've seen some theoretical white papers talking about how it is possible to remove data from these models, but it seemed like it requires a major amount of effort. So, um, is this one of those cases where an ounce of prevention is gonna be worth a pound of cure?
Because by, even if I had to do it two or three times, it would take forever, and sometimes I'm gonna have to replace the entire AI model anywhere, right? I couldn't have said that more, uh, couldn't have said that better myself. Um, in terms of, uh, an ounce for a pound, um, you can try, um, you can try methods to try and anonymize the data before putting it into, uh, some of these models.
But if there's one thing we've learned about, uh, these LLMs, they're incredibly good at piecing the piecing together information back together, piecing information back together. So, um, even with safeguards in place, whether it's inadvertently or there's actually an intentional adversary that's trying to, um, that's trying to prompt the model in a certain way, there is always gonna be a risk of that data being reconstituted, uh, with different sources of information being joined, right? Like an example is a rack.
In a typical enterprise rag, there's different data sources that are, that are available to that rack to synthesize an answer over. And so you might have taken care of anonymizing one data source, but if there's fragments of data available in a different data source, both of which that are available to the brag, the right word at prompt, could potentially re-identify someone or could potentially leak data that shouldn't be leaked, right? And so really, um, the way you phrased it is, is on point, which is the best defense, uh, the best offense here in terms of governing the output in making sure nothing gets leaked is a good defense.
Make sure that what you don't want coming out is never going in in the first place, right? And that ties back to do you know what data you have? Do you know what's most material to the business?
And do you have a way of keeping up with this with the pace with if that data is being generated and how it's changing? And only once you have that and you really put in place those assurances that, um, nothing material did make it into the model. And I can put a guarantee on this model, but it's not gonna leak so to data, Right?
So what's your best advice to the folks who are trying to apply some adult supervision in this when everybody's clamoring to use these tools? Well, um, I think, um, there's a, the defense in depth strategy is really the, is really, um, a good analogy here, right? Because there's different controls at play.
Um, what I see some security leaders doing is rather than, um, uh, what I see some security leaders doing is redirecting their employees to, to using internal deployments of models, right? So that's one, one level, level of control, uh, making sure that that, um, potential sensitive data is not getting leak to public models. The second, um, right, in terms of, uh, in terms of training that new models, it really does come to, uh, net new models or operationalizing your own rights.
It really does come to knowing your data, right? So start with, start with, um, if you, if you have a practice already in terms of data classification, data labeling, um, definitely invest in and try to understand for the various, uh, classifications of data, especially restrictive data, which of that data is making its way into model training playgrounds, right? So for example, um, uh, one of the customers that we worked with, we found that, um, we found that there were, uh, certain unmasked, uh, forms of employee data that were ending up in a SageMaker environment, right?
So really try to understand where the restricted data that you have is ending up and where that might inadvertently be getting picked up downstream by people training about models, right? Or if that data's ending up in Snowflake, for example, where it might be inadvertently getting amplified into some of these, uh, some of these models. Um, again, it starts with knowing your data, right?
So really double down on, uh, any strategy that you have and, uh, and if not, we'd be happy to help you work with that as well. Alright, folks, you heard it here. It always comes down to the data at the end of the day.
So it's all about how you secure it. Hey, pna, thanks for being on the show, Michael. Thanks for having me.
It's a pleasure. Been, And back to you guys in the studio.