Securing Enterprise LLMs with Knostic’s Sounil Yu
Knostic, the world’s first provider of need-to-know access controls for Generative AI, is celebrating an $11 million investment to secure enterprise large language models (LLMs). The funding will be used to bolster Knostic’s offering, supporting enterprises in their AI transformation and adding a customizable safety layer to tools such as Microsoft 365 Copilot and Glean. This additional $11 million investment brings the company’s total funding to date to $14 million.
Transcript
This is Textron tv. Hi everyone. Welcome back here to Techstrong tv.
You know, this next company I'm gonna introduce you to is a company I've been following since one of their founders first announced it. I'm not sure if it was on his Facebook page or his LinkedIn page. Probably both.
Um, but they, they've had, you know, in a relatively short time, maybe two, two and a half years, something like that, they've had a tremendous amount of success. They're here today gonna talk about, uh, raising some, some money, but more than the money, let's talk about what they are and who they are. I want to introduce you to Sunil Yu.
Sunil is a founder, CTO, and the company's name is Gnostic. Hi Sunil. How are you?
Hi, Alan. Doing well. Thanks for having me.
Uh, it's a pleasure to have you on here, and we're thrilled to finally have Gnostic on. I should mention that IL's co-founder at Gnostic is Gati Evran, and Gati is someone I know from the security space, probably longer than him and I want to admit, but 25 years, something like that. Mm-hmm.
Um, Sunil, we've not had the pleasure of meeting as we were talking off camera, but you also have a, a distinguished career in security as cyber as we call it now, as well. Why don't, why don't you share with people a little bit of your journey? Sure.
Yeah. Well, uh, I actually started at Help desk, doing help desk for, uh, at, at my university. And yeah, it's, it's a great starting point.
Did most people in security, they were network people, help desk people. Yeah. No one went to school for security, but that's right.
Now they do. Now they do. Now They do.
Yes. But anyway, go ahead. Um, yeah, so anyway, that's, that was my starting point and, and in it, I've actually been using computers for a long time, but Help Desk was my first job, so to speak, and, uh, spent, uh, many years in help desk.
And then I did consulting at Booz Allen for a while. Then I went out to, uh, my background is my background. So I went to Bank of America as their chief scientist, uh, spent some time at Venture Capital, was CISO at, uh, Jupiter one, and now, um, the founder of, uh, this company.
Um, somewhere along the way, many people know me for a couple other things, including having created something called the Cyber Defense Matrix, which is a mm-hmm. Nice framework that a lot of people use to help organize a lot of things in cybersecurity and, uh, some other things that have, uh, gave me some recognition in, in the community. So, yeah.
Fantastic. You know what, well, we've got your name on the lower third, so they could Google, LinkedIn, you or whatever, to get a better idea of, of your background. You know what, I've interviewed a lot of founders over the years, Sunil, and one thing that I, I've learned is that if a founder isn't passionate about the mission of the company they founded, chances of success are low.
Right? Because you gotta be a little crazy 'cause you leave a little something of yourself in every star. I've done four or five venture backed startups.
There's you, whether it's successful or however you define success, you always leave a little bit of yourself in, in these things. 'cause when you're a founder, you, it's your baby, right? You, yeah.
You are doing it. Talk to me about your passion that led to you being a, a founder AG Gnostic. Sure.
Yeah. I mean, I think you're always giving up a lot of yourself in a startup and, um, you're committing your life towards a, a goal that, uh, does suck away time from everything else. The, the main problem that I was, uh, that I was really interested in, in fact, Gotti would tell me, um, that he took me outta retirement.
Uh, I describe it more like a gap year. I took a gap year, okay. And, um, during that gap year, I was really trying to understand like, how can I tackle a really, really hard problem in ai and specifically around AI safety.
I felt that it was a, a really important problem that we needed to tackle, but I wanted to come up. Uh, but it's such a, a lofty sort of goal that I had to come up with something, some intermediate way to tackle that. And so, uh, gnostic is actually, I, one of the stepping tones towards solving that particular problem, or at least coming up a avoid a baseline.
What we mean by AI safety, uh, I'll talk a little bit more about that later, but the, but the big challenge that we have with AI safety is what do you mean by safe? Whose definition are we using and how do we know what definition, uh, is being adopted when it comes to how someone interprets or manifests AI safety? And the system does?
Systems that they use are the systems that they build. You know, I, I did a, I, every Thursday I do this little LinkedIn thing called Shimmy, says, couple of Thursdays ago, I think two or two Thursdays ago, I, I did a, just on this, you know, it was maybe a year, a year and a half ago, you know, 100 top names in technologies were clutching their pearls saying, we've gotta go slow with ai because the safety issues, the threat to humanity itself. And, and we've gotta put some guardrails on here and we've gotta do some things.
Here we are, a year and a half later, the same people who signed that letter are trying to buy open ai, and we're investing at $2 trillion in AI and AI data centers. And it's full speed, you know, damn the torpedoes full speed ahead. Mm-hmm.
Have, have we just said the heck with safety or, or are we just blissfully ignorant or some, some of something else? Well, unfortunately, money and profit tends to drive behavior, right? An alliance incentives in ways that are not necessarily best for, uh, for society as a whole and, and just somewhat divergent.
But the open AI was originally started as a nonprofit focused on AI safety as a core mission. And part of the reason why there's this rift between Musk and, and, uh, Altman is because a lot of people feel that they've moved away from the mission. Um, but at the end of the day, you, you mentioned something along the way, which is putting in guardrails.
Uh, the question here is, the question that I've been working through is which guardrails are appropriate for which audiences? Okay. And, uh, with the release of something like Deep Seek, there's a lot of guardrails that are, that the Chinese government or the, you know, at least that society believes should be walled off, so to speak.
Um, and we have a different view of that in, in the US Today. Yeah. Yes.
Today. Um, and you can apply these guardrails maybe at a society wide or a system-wide sort of view, but inside of an enterprise, you can't actually have a one size fits all guardrail because there are some things that are entirely okay for some people to know and others who shouldn't. And so, for example, uh, if I went and asked about salaries, well, if you're in hr, it's okay for you to know that.
But if you're asking, uh, if you're an engineer asking about salaries or Lao, it's not, okay. So how do you design these LLMs? Or how do you give LLMs discretion?
Because right now LLMs aren't able to keep a secret. It doesn't know what I should tell one person versus another person. And that's the fundamental challenge that we're seeing with large language models today.
And that's the problem that Gnostic is trying to solve. Excellent. I just wanna mention from the outset, Gnostic is spelled K-N-O-S-T-I-C.
That's K, yes. Silent K. What?
Yeah. When, when, when dealing with Israelis, they, they tend to make it a hard K. Um, Yes.
I, I know, well, you know, I'm not Israeli, but Yeah, I know he a little bit. Um, Lemme give you a quick, uh, history on, on the word gnostic. Sure.
So, um, the word gnostic with a G is the Greek word Yes. For knowledge. Yes, of course.
People know the word agnostic, which means not knowable. And so know, uh, gnostic means knowable. And with the K it kind of hits the, the same sentiment that we're trying to make, that which is know, uh, that which you should know knowable.
Love it. That's a great, that's a great genesis of that name. So Sunil, we, you know, we kind of bid around the pie here, but let's sink our teeth into the heart of it, right?
You guys recently, uh, announced an $11 million round or an investment, I don't know if it was a full well blown round, but, you know, to further the mission of securing enterprise large language modules, let's start with what do you mean by enterprise large language modules? Yeah. So when we talk about enterprise LLMs, we're talking about those that are being used for, uh, the enterprise itself to share institutional knowledge.
So we have a bunch, a lot of enterprises have knowledge workers, right? How do you make a knowledge worker more productive? Well, you give them access to knowledge tools.
And these LLMs like copilot and glean and Gemini for workspace, they, they're all tools that help us, uh, tap into institutional knowledge at a speed and a rate that we, up, up until LMS are, are released. We would have struggle. Uh, we would struggle finding the knowledge that we need to be able to do our job effectively.
But these large language models also have a tendency to overshare. They, they tend to, um, give you content that you actually may have access to but didn't realize. And so if someone has, uh, accidentally shared, let's say, uh, I mentioned earlier, you know, uh, a salary spreadsheet or layoff information or m and a or something, if it's misfiled in the wrong place, well guess what?
These tools will help accelerate the discovery of that institutional knowledge as well. And it doesn't really understand discretion. It doesn't know, oh, I shouldn't share, even though you have access to this, I shouldn't share it with you because you don't have a need to know.
And so that's, again, the focus of our company. How do we, um, secure AI inside the enterprise by creating these guardrails, uh, that are specific to the enterprise so that you can share institutional knowledge broadly, but do it in a way that allows you to have proper discretion of what, based on people, what people should or should not know. I get it.
I get it. You know, look, our audience is, everybody's in AI person today, and our audience is pretty tech savvy, but just in case you didn't catch the subtlety here, right? Well, a lot of people when they're talking about LLMs today are talking about these large LLMs that companies like OpenAI or, or, uh, anthropic, you know, or or Lama are, are creating, you know, and, and it's basically, you know, some subset of the internet at large, right?
And then they train their AI on that LLM. But where we think it's going and who, who knows, you know, 'cause Gentech AI and everything else coming down the pike, but where we think it's going is that a lot of companies that create their own LLMs, and maybe they're not all large LLM, maybe they're FLM small language modules. Mm-hmm.
Um, but this is their tribal knowledge, their institutional knowledge. And you know, I, I once had someone explain it to me, Sunil, that think of that as short term memory that the, an AI query uses first. And then if that doesn't work, it goes back to a, you know, one of these real big LLMs, internet wide LLMs to kind of fill in the gaps and think of that as maybe long-term memory or something like that in, you know, in the way a brain works.
But, um, nevertheless, that information in these LLMs are make talking about the smaller LLMs. Now, the enterprise LLMs, a lot of people don't realize that that information's being sucked up Mm-hmm. Into these institutional LLMs, if we could call them that.
And a lot of that information is not, you know, should not be made public or is not, you know, it's, it's private information, proprietary information. And so is that part of the mission to make sure that that information does its sign, doesn't find its way up into the institutional level? Or is there another aspect of AI safety here?
Yeah, so what you're talking about is if I use, um, a public model like open AI and I transmit information to it, will that mm-hmm. Information gets sucked up into, let's say, the next model. And I think the, um, there's a misconception about the, the dangers and risks there.
Um, the simplest way to describe it for me is, uh, imagine you're a teacher teaching geography and you prepare, uh, the raw facts of the world to teach your class. But you come into your first day of class and realize everyone's a flat earthers, your students are all flat earthers. So you give out a homework assignment, Don't make, make fun of me in Florida here, but go ahead.
Anyway. So you give a homework assignment, they, and the, uh, students transmit to you all this most sensitive secret, uh, copi conspiracy theories about why the earth is flat. You as a professor wouldn't retrain your, your content based on the homework assignments and then what they transmit to you, you wouldn't teach next semester's class based on what the students send to you.
So this concept that's belief that when we submit content to open ai, will they use that to train next, next semester's model? Um, it's, that's not how it works. You would, however, retrain your syllabus or you would fine tune your syllabus to make sure you're addressing some of the key misconceptions.
And that doesn't really divulge the homework. That, that, again, just reshapes your content that you already have, um, at, at your disposal. You're not using someone else's content and regurgitating that to somebody else.
So that's type of Problem. What what about proprietary information though? Well, I mean, you still have the risks of how, uh, that that's no different than the ai that, sorry, that ai, that's not an AI problem.
That's just a third party risk management problem, right? Data. But the problem that we're trying to solve is within the enterprise itself, uh, because inside the enterprise different, different people have a different, uh, types of need to know.
Not everyone has a need to know for everything that's out there. And these large language models, were able to have this amazing ability to connect the dots across your enterprise. You're trying to share institutional knowledge so that your knowledge workers can be more productive.
And whether it's, uh, um, it's oversharing because that knowledge worker has access to a file that they shouldn't, or because it can be inferred, we have a problem either way. And to give you the, uh, example of the inference problem, let's say for example, someone had access to building information and future equipment purchases. Could you infer layoffs based on that?
And these large language models are also called inference models, right? So that's, they're inference engines that allow us to infer those things, and they have a really powerful mechanism for doing that. So those are things that we are trying to control.
And what we don't have today, anywhere, I, I mentioned this notion of need to know, need to know is a foundation for access control, but where do you have that codified today? Where is that written down? Where is that referenceable?
It's all in our heads. And if we have these large language models, and especially if, if we have agent based ai, where are they gonna reference what the individual or where's the reference point for need to know within the enterprise? And that doesn't exist today.
We are building that, and that's the basis of what the company's all about. So in my mind, need to know sounds a lot like zero trust. Yes, it's a form of zero trust, but for knowledge, right?
Mm-hmm. But it's zero trust. I have some, I have a little bit of problem with using the word zero trust, because the goal here is not just to say, let me only share that what you have should have access to, but also there are a lot of things that you don't have access to, but is within your scope of need to know.
Uh, but you don't have access to it. And you don't even know you have, you don't even know it exists because you, it's blocked off. You're siloed.
And this was one of the biggest problems with, let's say the nine, the nine 11 commissions that, hey, you and the intelligence community, you have all this content there, but you've just failed to connect the dots. That's because we were all siloed, but the people who were doing the analysis had a need to know for these things and was unable to get that. The same problem we have in an enterprise, we have all these silos.
So the zero trust notion is oftentimes saying, look, you can only get things that you explicitly are supposed to have access to. And that's usually done at the file or server or folder level. But what if you can also gain access to things that are, are inside your need to know boundary, but you don't actually technically have that level of access today, we would want an LLM to actually get that for you and deliver that to you as long as it, it's within your need to know.
Okay. And, and that's a great use of AI actually, right? Because I'm just thinking, had we had an AI on the nine during, you know, leading up to nine 11 mm-hmm.
It would've connected the dots perhaps, Right? It would've, we would've had a better chance of doing that, right? And, um, and that's really actually why organizations really want to use these tools.
That's of course, the business driver for that is immense, right? Um, there are security concerns of course, and the oversharing, but the way that we balance that out is by again, understanding need to know. And the, um, going back to my, the earlier conversation around AI safety need, need to know in AI safety are a lot more closely tied than we realize.
Uh, because going back to the question of what do you mean by safe, it's the same basic question of what do you mean by need to know? How do you define a need to know boundaries of an organization? And it turns out the connection is a lot more, um, there's a closer connection than, than I think your listeners would be able to realize.
But that's it. I've, I've found a connection that help us connect the, uh, ti these two concepts together. Excellent.
This is enlightening. I appreciate it, Sunil. So how do you determine, need to note is, is data in different buckets?
Are people in different buckets? Are they both in different buckets and it's kind of a double helix? How, how do you, how does that play out?
Yeah, so our, our, so you mentioned zero trust. So let me give you an analogy real quickly. Um, when we first rolled out networks, they were flat.
Mm-hmm. And because they were flat, a lot of things were reachable, which made it powerful but also dangerous. We realized, hey, we need to segment our network.
So the way we segmented it was first by looking at net flow. 'cause we don't know what your segmentation, you don't, no one knows their segmentation policy upfront. They have to look at net flow.
So when you roll out a large language model inside of an enterprise, your knowledge base becomes flat. That makes it very powerful, but also potentially dangerous. So we need to segment our knowledge, that segmentation policy is the need to know policy.
But if I went to you and said, Hey, what's your need to know policy? You would say, I don't know, where do I start? So we start with knowledge flows and, uh, knowledge pings.
So what, what, what does that mean? Well, what we're doing is, uh, one of the first steps that we do is we help people find, very quickly, find where content is overshared by basically asking, uh, uh, through the perspective of a user, a user with a particular profile or a job function, and asking for a whole bunch of sensitive topics to see where, uh, uh, content is exposed. And then we overlay our need to know policy to determine when it's overexposed, when, when a uh, organization says, yes, that's correct, or no, it's, that's not correct.
That's effectively defining the need to know policy as well. So that's the process to re update, need to know policy for every organization. That's, um, it, it's very similar to how we do network segmentation.
We show them, here's the flow, here's the, here's what the, what's communicating. Is this correct? And if it's not correct you, that's no longer part of pol policy.
And if it is correct, it gets codified as a part of policy. Got it. Sunil, I talked to you another half hour on this, but we're outta time already.
ai Dot That makes sense. Dot ai. Yeah.
ai can check it out. Um, you'll be at RSA in a couple of months. Yes.
Um, I I, I'll be in San Francisco, um, that week and I hear there's a conference going on. Yes, there is. We'll, we hope to see you there.
Tell, tell Gotti I hope he feels better. I know he, he was under the weather a little bit today, but thank you so much for coming on here and shedding up one or two layers off this onion. I'm gonna need you to come back and we'll continue peeling the layers.
Absolutely. Okay. Happy to come back.
My pleasure. Sunil, you founder, CTO Agnostic. Here.
I'll text her on tv. We're gonna take a break. We'll be right back.