GenAI Secure Data Access with Redactive’s Andrew Pankevicius
Andrew Pankevicius, Redactive co-founder, discusses the challenges of enabling access to generative AI LLMs and live permissioned data sources. Andrew shares how Redactive accomplishes this while utilizing existing access control models and how data pipelines won’t store customers’ end documents. Visit redactive.ai for more information.
Transcript
This is Textron tv. Well, we have the great pleasure of being joined by Andrew Victus, who is a co-founder with Re Red Reactive Welcome. Couldn't be talking with you Andrew.
Yeah, thanks Mitch. Appreciate it. You bet.
You bet. Um, you know, we've talked before, but that was at another part of your career, a prior step, I think, when you were Atlassian, but great to be having you on here with, uh, red Reactive. Tell us a little bit about the company, why you started it, the problem space you're going after.
I'd love to hear a little bit more about that. Yeah, absolutely. So, so at Redacted, we're kind the only developer folk, uh, platform focused on accessing, um, live commissioned data for generative AI applications.
We fundamentally feel like it's that missing part of the, um, like that kinda LOM application stack, which is a lot of people focused on that kind of governance aspect around that infrastructure layer, looking at lms. But one of the biggest blockers to kind of AI or gen AI application adoption or fundamentally building them and getting into production is how do you access permission data that's live from sources like Confluence, slack Notion, OneDrive, um, another kind of custom databases that might be governed by, um, identity access management systems. We see so much kind of in market right now, this risk around bleeding information internally when you bring in generative ai and we're kind that one stop fix there in order to make sure that integrating RD RX K into your product ensures that you can really seamlessly connect, um, to those data sources and then leverage them in real time if you're building AI agents, uh, chat bots, workflow applications, et cetera, any of this kind of next gen style of, uh, gen AI application.
And the reason why we, we jumped into it is my co-founder, Alex and I, uh, we previously used to work at Atlassian and we saw that kind of trend there around challenges around enterprise adoption of, of kind of capabilities. And that, um, it took a long time in order to move customers from kind of server environments into data center environments and data center environments into cloud. And there were so many permissioning data access, data sovereignty challenges there that need to be solved.
And, and not just in Atlassian's case, fundamentally across the industry around cloud adoption that we saw a very similar trend probably starting to emerge in the general AI space as well, is that, um, there's so many cool open source projects that they're attempting to stand up, small proof of concepts, um, for interesting use cases, be it agents, chat bots, et cetera. But we really felt like, hey, this really complex and very clear part information security review requirements inside of enterprise environments such that these proof of concepts that leverage gen AI are probably not gonna get into enterprise environments anytime soon. So we wanted to become that kind of that key arm block in order for large groups of internal employees or even their customers to be able gen to, to leverage generative AI capabilities.
One things I'm curious about with generative AI is, you know, it's relatively new, it's been around since 2017 or so, but most of us till the last year or two kind of burst on the scene. Are, are those, are the models the LLMs, the, the smaller, uh, LLMs are, is there a security model very mature? Is that one of the things we're looking for generative AI to kind of come along with to what an enterprise would need for a security model?
So you can't fully depend on what's in an LLM. Yeah, it's a great point here and it's kind of like if you look at the, the last few even kind of quarters of kind of generated AI kind of, um, changes and trends in market, it's the kind of last year obviously Gen, um, generated AI get very interested from like a productivity use case point of view with chat GPT and then organizations looked at, well, how do we bring this internally inside the data organization? So then they looked at open source models and fine tuning, which obviously comes at great cost and removes a lot of that permission layer of data access.
That's actually the information that you wanna pull on. And there's two kind of core components there, which are kind of security risks. Um, when you look at kind of just fine tuning or training and model is that the information that you want AI to leverage, um, needs to come from an end tool.
So you think about Confluence like notion one drive, you're gonna fine tune it. You're actually gonna strip that information away from its source of truth and now start kind of mixing in the concrete, um, into the model that you want to use, which means that that information at any point in time is fundamentally becoming stale. Um, which means that, right, your strategy now is how do you continue to fine tune over time and bear that cost on the organization.
The second component is from a security point of view, which is not only are you leveraging stale data now in order to leverage an LLM, but you've fundamentally stripped away the permissions or that fine grain access control from those documents. So there's kind of two layers there in the fact that not only have you removed that core access of an end user to a tool like Slack or Notion, but then also think about the pages that you can see in some of your coworkers can't. You might have some locked documents.
There might be some things that might be one-on-ones with your manager or HR information or finance information that that that department's dealing with. All of those fine-grain controls inside of those tools have now been stripped away when you start putting them into a, into a train in a trainable model. So we've really lost that, that ability to control, um, data access and create that more kind of risk vectors inside of an organization.
We leverage the RAG framework as a core component of how adaptive enables applications to leverage generative AI securely. And what we do is that we fundamentally pass through those security, um, kind of access control layers through to the end application. So when you leverage redacted, you're passing through no longer prompt, but also the access token or the token to that end user such we're only pulling on information that that right user has access to.
And you don't have tokens embedded in applications, you have your applications all using the same security model for accessing that data, not 25 different ways. I can imagine there's a lot of benefits to standardizing that process plus the maturity of it. Are there, are there certain things about LLMs generative AI that present unique challenges to enterprises that are adopting them when it comes to access control and security?
Yeah, absolutely. And I think number one is that sense that in so many CTOs and sizers that we speak to is that there's a real kind of interest in starting internally with their core use cases. So they're looking at, at from a, from a use case perspective, number one is if they're trying to, uh, provide permissions aware q and a across multiple different tools.
So if you connect up your HR application in into your, um, confluence environment, you can actually pull permission data and more general information across the business for the right user as long as you respect their downstream commissions at any point in time. And then the second component is like, you know, um, like kind of that enterprise knowledge base q and a is really where people wanna start on their generally AI journey in enterprise environments. But very quickly, once they get comfort there, they're looking at how can they augment their customer service agents, um, and really deal with high touch customer service, the human in the loop style customer service environments.
And that's pulling permission data that's maybe very sensitive about the customer, especially in, uh, financial, in, uh, financial services industries. So thinking about banks or insurers we might be dealing with customers claim or finances information, you wanna ensure that that customer service agent can speak appropriately about the business and their product information or, or kind of their policies, but then also speak very specifically about what is that individual's end challenge or their situation. Um, and really right now, customer service applications can't do that in the generally AI space because they are making these, um, other kind of, uh, these gaps around permission access.
It's kind of all or nothing. Here's your token, now you can, you know, we have access, right? Absolutely.
Yeah. For better or for worse, you, and it seems like there's a number of use cases, use cases that come to mind. You've talked about the chatbot and, and customer service.
There's also automating workflows, you know, more behind the scenes backend kind of functions, um, you know, automated forms, things like that. I can imagine a number of places where you want to hook in, you know, gender A AI may not be the main application, but it's part of the ecosystem of data, data sources that you're using that you want to be able to hook into those processes and applications. Absolutely, and I think this is really where, it's really interesting to look at the application developer market that's emerging around generating ai, trying to build these native AI applications.
We are working with some right now who are building AI agents, and one of their big challenges is that, you know, they're looking at customer service roles, researcher roles, they're looking at, uh, ITSM desk help desk environments or even kind of risk like automated risk reporting as being this kind of core use cases, um, inside of businesses. But the big challenge that even these kind of AI agent builders have is how do they access these live commission data sources as well? You can access, you know, kind of traditional data lakes and, and, and kind of more stale data in order to start to personalize how, how it all, how a agent might act, but really where you do your work every day is probably where an agent should do their work as well.
Um, and those are in those collaboration tools environments where data is changing consistently, where permissions are being, um, on an almost hourly basis being applied and removed and is where the most recent source of information that you would want an agent to respond to, especially in a risk reporting case. Um, so unlocking that component with RSDK, something that we've been helping some of those AI agents actually do in market right now as well, such that they've got more data coverage to power their m these cases, Are there things you have to do differently for various LLMs, you know, there's something from an open AI I AI versus something you get off of hugging face. Um, are they all different and you have to kind of, you have to adapt 'em to what you're doing?
Or is it you providing more of a one way to access all of them? Yeah, no, it's a great question. So, so we, uh, redacted our SDK now underlying platforms, fundamentally a large language model agnostic.
Uh, we can allow, we allow you to plug into any model, be it some of those big ones in the space or if you're rolling your own, um, there's many different use cases there. So we really encourage our customers, especially enterprise environment, that they have been fine tuning a model, uh, and they wanna start there. They can benchmark its performance whilst additionally providing live context through bio model.
Um, are ways for them to kind of leverage what they might have already built or models that they feel comfortable about from, um, like AWS or GCP looking at Gemini and bedrock models, um, in order to feel really comfortable about the security constraints of other parts of their stack, whilst also using redacted to solve that that fine-grained access control challenge. It could also be helpful too then for, not saying it's a small deal to do, but if you want change out LLMs that you're using, right? You may be using one today and now you've, you know, tuned or trained a different one or you're using a different source.
I know some, some people are actually kind of acting as a front end to multiple LLMs and then depending upon the prompt will feed parts of it or, or all of it to a different LLM to get the best response from that particular model. So like you have a lot more flexibility and less direct ties to un unhooking, you know, your connections into a specific LLM to talk to another one. A a Absolutely.
And I think one of the, the things that, you know, we go, I miss not to mention here as well, is that when you start having a lot of flexibility around different models, especially if you're an enterprise environment, you want a lot of transparency around that data pipeline. Um, such that we, with redacted when you're kind of moving information through in this live retrieval environment, um, and pulling relevant business context from, so like Slack Notion OneDrive, for example, for generated ai, we allow you to put in any kind of DLP provided your choice. So double loss prevention are obviously really large standards things that go through large procurement cycles for enterprises.
So they get really attached to their underlying DLP provider and we allow you to basically pass through any of the context alongside the prompt through A DLP provider before then hitting that large language model of your choice as well. So that additional layer of kind of egress security, um, if protected within your, your VP C environ, What's, what's the onboarding look like? What, how, what does it take, you know, developers always looking for, you know, make my job easier, don't make it more difficult than it already is enough challenges.
Um, what does that onboarding onboarding model look like and how, how much time or effort does that take? Yeah, look, we, we've tried to at redacted really simplify that down, um, to the fact that you can do it within about a day. Um, we have a SDK download off our website, contact us to get access.
Um, you select the data sources that you'd like access to for your end users, and then, um, you fundamentally just put, um, insert a a a connect button into your application usually during your onboarding flow. Um, and if you've already got system administrator access and a trusted app inside the enterprise environment that Connect is, is as simple as like an SSO screen that you would see on Google. Let's get that one encounter with redacted in your entire journey, and then you're fundamentally connected securely and managing the, and being able to pull from downstream app applications with its various levels of permission aligned to, to your level of access.
Um, what's really great as well is like sort of really simple on onboard side of enterprise environments in that regard. With application developers, it's exactly the same as well, which is, you know, there's so many challenges around building those chunking, embedding Vector Store and live fetch pipelines that are taking away time from your engineers solving problems for your customers and your end use cases such as the exact same experience, jump to a platform SK, embed it into your application, and then passing through prompts plus token allows you to pull live business context from those end data sources. I think the last thing I here as well to this is that for our enterprise customers as well, is that we're actually providing a lot of, um, kind of template use cases as well, um, to, to enterprises to really just starting their generative ai, um, adoption journey.
So whilst the redacted developer platform is really powerful to support any use case, we're building out those kind of permissions aware q and a and high touch customer service kind of template applications such that they can know in that front end. It's something we intend to open source in time and allows them to get really jumped into their general AI journey with their, with their internal and customers within the, or sorry, internal employees within about two weeks, if not less. Right.
I, I would think your experience too, working with enterprises, you, you understand the process you have to go through to get provisioned through the security teams, and the easier you can make that, the quicker it's gonna happen For sure. A hundred percent. And, and one of the, the, one of the great points where, where Redacted comes into its own is, um, we've often found, uh, business buyers internally or someone who's in charge of generative AI and taken on that helm inside of the business and might have built a really small rag application.
They've got a vector store, which is stripped away the permissions that they're pointing to, um, they've got maybe 10 internal users leveraging it, and then they hit information security review. They get really excited about the idea of scaling it out to the customers and their, and that information security review, that security team fundamentally rejects that solution architecture. Mm-Hmm.
Um, because they say, why does your, your generative AI application be allowed to have a different permission system, different identity structure compared to, um, what we've invested in in a long time, how all the rest of our applications work, you're introducing new risks, we're blocking your application here. So they really start to feel that pain internally around that like that, that permission kind of management journey and that, that challenge there about how to solve that part of the stack. Um, but once we kind of come in and say, Hey, this is what Redacted does, they go, great, how can we integrate you tomorrow?
Because we need to unblock, um, our pathway through information security review. So we very quickly become kind of the trusted friend or kind of champion of, of the security, um, team as well, um, which has made it really simple to kind of then make that progress into production environments with our customers Where the hurdles the better o one of the hurdles that can happen in those reviews is, uh, retention of data. Is there any data that's passing through some of the third party service?
Is it getting left behind? So now we go down the road of the II and other, you know, sensitive kind of information? Is that at all an issue for you?
Yeah, well it is actually something that we're, we're kind of actively solving around, again, kind of going to the origin story of redacted, looking at like what are the enterprise requirements of production grade application kind of, um, adoption and then reversing it back. We knew that how redacted manages or potentially stores customer data is gonna be incredibly critical to our adoptions and developer platform ourselves. So in those environments, what we've done is that we've really rolled this back to the idea that the big challenge is likely around how indexes of information are managed, which is a lot of applications or in the traditional enterprise search space, um, would be trolling different, different systems, pulling all that information into a, into a new vector database or just a fundamental new database before, um, large language models were really popular.
Um, and they would store kind of chunks of information about potential customers, the company, et cetera. Um, redacted doesn't store chunks, and this is part of our unique solution architecture in order to do live retrieval, which is that we only ever storing beddings of information, which is obfuscated versions of the actual end documents such that you can never reverse them back out into being a, a record of, of any kind of customer, et cetera. Um, we leverage that plus a pointer system, so we identify where relevant context exists across the business, and then we always go to the source to do the app, um, app query time permissions, check that end user to make sure that live permissions are applied, and then also pull the, the live version of the document in.
And what's really great about this is that because we act as that pass through layer there and we're not storing any kind of version of the end document ourselves or information on customers, it reduces that third or fourth party risk, which security and information security team to really look at as being that kind of additional kind of vendor risk that that third party risk aspect. Excellent. Excellent.
Sounds exciting. It sounds like you're attacking a, a very valuable space to go after. It's hard, hard for apps to do much without really good data Yeah.
And secure access to it, whether it's, you know, generative or, or traditional applications. Um, what's the best way for folks to get ahold of you and start to look at, uh, redacted? Yeah, look, we we're, we're active w with, with many customers across Australia and the US at this point in time, across application developers trying to build, you know, Silicon Valley style startups and, and accessing commission data and wanna kind of really remove that pain around, um, those kind of, that new data engineering skillset.
Um, and how to kind of, um, get customer trust, uh, by leveraging personalized information in their environments. Um, and on the other side as well is that enterprise is looking at how do they manage their journey of AI journey? How do they build a platform that can they feel comfortable in building multiple applications on top of, in the machine AI space?
Um, and for both of them on say, contact us, our website, um, has a Calendly link on it, you can get in touch with us directly, um, and we've got a full team to go through assessing your use cases, how you can leverage the red reactive platform and then get you up and running within the day or leveraging our template applications to be ready and going with internal use cases within two x. Very good. ai, correct?
That's correct. Yeah. Awesome.
Makes sense. It wouldn't be that. Yes, it's been great talking with you and uh, again, I think this is a really fascinating area and certainly something every enterprise, every business has to address around secure access to data now with, uh, generative ai.
It's interesting. I was just doing a panel here earlier today and we were talking about testing and access to, and all kinds of, the myriad of issues that come up with that when you start introducing generative ai. So wish you the best and, uh, keep us, keep us in touch as things progress and uh, love to hear more.
Yeah, definitely be speaking to you soon, Mitch. Really appreciate it. Thanks for the time today.
Alrighty Andrew. ai.