Dilip Bachwani on Analyzing DeepSeek with TotalAI
Qualys leveraged it’s TotalAI solutions for a comprehensive security analysis of DeepSeek-R1 Llama 8B variant. Dilip Bachwani, CTO & EVP, Cloud Platform at Qualys, covers the results.
Transcript
This is Textron tv. Hey everyone. Welcome back to Textron tv.
You know, I, I always enjoy talking to my next guest here. He is the CTO and EVP for the cloud platform over at Qualys. It's my friend Dilip Bani.
Dilip a pleasure to have you. Usually we're in person at the QSD or at RSA or somewhere where we're in person, but today we're on Zoom, but it's still good to have you on. How are you?
I'm good. Good to be here, Alan. Uh, it's always a pleasure.
Yes, it's, it is. It's, it's fantastic. You know, I mentioned it, well, we should hit it right off the bat.
So, QSC Qua QS, Wallace Security Conference, QSC, um, is coming up this year, I think in October. October 16th, around there. And I, I heard a rumor it's gonna be in Houston, which is a great city.
That is correct, yes. Uh, it, it's a new location for us, um, this year. We are in Houston in October.
And of course, looking forward to being there, looking forward to meeting our customers, sharing with them everything that we have been working on, especially everything that we are doing around enterprise to risk management, the risk operations center, and how that has evolved since we last talked about it, uh, during the 2024 QSE, I had to think for the year, right? Yes. And by the way, if you go to Techstrong tv, all of our interviews from there, and there was a lot on the Risk Operation Center.
The Rock are, are available if you go to industry conferences, look under quais, and you'll find they all, including my interview with Dilip, is there. So check that out. But Dilip, we're gonna talk about something else today.
So, last week was a bit of like a Sputnik moment, if you will, right? All of a sudden, the, the Western AI establishment was rocked by news out of China that these folks, you know, it's a, a, a handful of PhD engineers basically put out a, an AI model that rivaled the best of what we have here at a fraction of the cost and a fraction of the time. And open sourced it on top of everything else so everybody could go, you know, look under the covers to an extent.
Um, and it was big news. Qualys turned your, uh, your, your scanning engine onto it and, and came up with some very interesting results. I don't want to say too much 'cause it's your story.
Lay it out for us, Philip, what happened here? So, and Alan, I mean, spot on, right? Uh, deep seek.
Uh, this is a, you know, fairly young AI company, uh, out of China. Uh, certainly very talented folks. They came out with a model.
Uh, they've, they've in fact been coming out with information for the past few months, and, um, I don't think a lot of people necessarily noticed them until they came out with this latest model, uh, the R one and it's open source, but it, it performs really well. Uh, and you're right, uh, you know, they are saying, uh, they have trained it at a fraction of the cost of what some of the larger tech companies here, OpenAI, meta, uh, Gemini, and others are, have done to train their models. Uh, so I mean, first and foremost, I think what they have done is something amazing.
Uh, in, in some ways they, they are proving that if you want to build very large scale foundation models, you don't have to be in an extremely large organization with unlimited amounts of money. Uh, you can be a smaller shop and you can do this. Uh, so that's a good thing, right?
Um, it certainly got a lot of hype. It got a lot of coverage. Uh, what we wanted to do was, as we were looking at it, because we were curious too, we use a lot of open source models internally for our, uh, quas cloud platform.
So we wanted to look at deep seek and just understand, you know, how it was behaving, what it was doing. Uh, now I think you probably know we launched Qualys Total ai, which is our AI security solution some months back, uh, uh, in August. And what the solution does is it gives you a more comprehensive view of your AI posture, meaning you will get full visibility into your AI hardware and software assets, uh, your entire AI inventory, where your models are deployed, where your L LMS are running, and then also a pretty detailed vulnerability posture across your AI footprint.
In fact, we have more than 1500 detections right now, just from a vulnerability standpoint for your AI footprint. So we said, well, let's take this, let's take what we have and let's see how deep Seeq performs on that, uh, from an LLM scanner standpoint. So the way our LLM scanner works is we do an outside Incan, and we have built a pretty exhaustive knowledge base of questions that we will ask an LLM, um, back and forth, um, which we call our knowledge base, and we gather that information.
Then we have some inbuilt models, which we use to then judge the quality of the responses that is coming from these target LLMs that we are testing. So we did that and deep seek, um, to our surprise, uh, it didn't do particularly well. Uh, in fact, it, um, it failed 61% of knowledge base tests that we had, and in total, we ran about 900 tests, um, just for our knowledge based checks, right?
And what these checks do is they test for, um, ethical questions, legal questions, operational questions, uh, and we do a lot of back and forth with the model to kind of get a sense of how is the model responding to our questions? Because these things are important, right? You take a model and you deploy it, whether in a B2B setting or in a B2C setting, and you expect the model to work in your particular domain and not give answers that it's not meant to give.
And when it does, then there is a liability issue here, right? And, you know, that's what we are trying to get a sense of. So we did that.
Then, in addition to the knowledge based tests, we also do jailbreak tests, um, where jailbreaking basically involves techniques, you know, that you can use to bypass inbuilt safety mechanisms that are built into the model. Uh, there's a lot of well-defined techniques. Uh, we, so obviously we work with the community, the open source community, and in, in our solution right now, we have about 18 to 20 different jailbreaking techniques that we use, and we ask the model questions around these techniques.
The idea being that you're somehow trying to coax the model to give you information that it's not, it should not be giving you harmful outputs, uh, misinformation, you know, privacy data, unethical data, all sorts of things. And I think just based on the fact that it didn't do so well on the knowledge based tests, um, I mean, not surprising, but it failed more than 50% of the jailbreak tests to failed almost 58%, almost exactly 58% of everything that we did. And other analysis was pretty comprehensive, um, you know, trying to get a sense of what was going on here.
Uh, so really the, the gist of this is, yes, it's, um, I think it's a great model from a foundation model standpoint, uh, in how they have trained the model, the underlying architecture, um, and just being able to demonstrate that you can build a model with significantly lower investment. Um, but that corresponding investment hasn't really happened on other areas yet. Right?
And then of course, concerns with, if you're using a hosted model that is sitting in China, and you have GDPR concerns or other, you know, regulatory requirements, right? Not concerns, but GDPR requirements, right? And other regulatory requirements across different countries.
Something to be mindful of, right? Um, Yeah, but well, that's the whole sovereignty issue, right? Yes.
So, Dilip, I'm not gonna make excuses for them, but let me postulate two, two things on what you said. Number one is some of the, uh, not the jailbreak questions, but the sort of foundational questions, could it be due to the fact that clearly, because it is from China, it is, it is, I don't wanna say censoring, but it's purposely not reporting on some sensitive areas that the Chinese Communist Party may beem, uh, sensitive that they don't want it to report on. And so it's been, in essence, blinded for those things.
And I mean, 61% is still a pretty high number. I'm sure it wasn't, you know, 61% of things that have been censored there, but could that be at least partially, uh, responsible for that? Yeah.
So Alan, there are two parts here. One is, I mean, obviously just based on the fact that the, you know, the model came out of China and the sensitivity of the Chinese government, there are some questions that you can't ask the model. So it's really quite censoring some things.
Um, and some of those are well documented, uh, right? Um, that what this then means is that if they do want to stop, so the model from giving incorrect information or unethical information, however you look at it, they are, you can stop the model from doing that. You cannot be perfect, but you can restrict it as best as you can.
But those controls haven't been applied to other areas. As an example, we asked the model a lot of, when we were asking jailbreak questions, um, you know, we asked a lot of typical questions on, you know, how could I make an explosive, uh, how could I come up with, um, you know, incorrect healthcare information? And it didn't need a lot of prompting, a lot of circumventing to get that information out.
It was just giving us that information very quickly. Uh, and I, I think what this points to is they have put in certain guardrails for things that, for deep seek, you know, just being where they are, you know, they had to do that, That are important to them and their right, right? But maybe not for the west, But not for a larger, and, you know, that has to be done.
Now, either they do that or other organizations can pull these open source models and have RAs sitting in front of these foundation models to say, when you're asking a question, I'm going to make sure I'm filtering the right things out, and only asking the model what makes sense, right? Because a model inherently is not trained yet to do that. Yeah.
I mean, and that's one of the beauties of it being open source, right? You could self-host it and put whatever guardrails you want in front of it, and it's self-hosted, and that takes it out of China and everything else. But Dilip, let me, a lesson I learned in my 25 plus years in security, 30 plus years in technology, is security becomes important when customers demand, it's important.
And I think clearly deep secure wanted to get this out. I I don't think it was any coincidence that this was released. This R one came out two or three days after the, uh, Stargate project or whatever, the $500 billion project to go build data centers was announced, right?
There's, there's, there's PR here and Global nation state strategic, you know, competitiveness at play. I don't know if they had the time to maybe put in the, just, just like, until a customer demands better security, you don't have better security until someone says, you've gotta put these guardrails in here, especially when they're rushing to get it out. They, they don't put them in there.
I, I would hope that that's just more of a sign of its immaturity than a total lack of, of ability to do that kind of thing. Yes. Um, and, and Alan, I think I agree with you there.
Um, this is a fairly young company and, um, I mean, they, you know, they've been working on building, you know, some, I mean, in my view, some exceptional models. Uh, and to your point, uh, you know, this is still to some degree, you know, research oriented, right? Um, but when you look at the overall AI ecosystem, there are different layers that you're looking at, right?
Uh, you are, you have one layer, which is your hardware and infrastructure layer where the folks like Nvidia are playing, right? The second layer is your foundation models, right? Which are now to some degree, it feels like they're starting to become more commoditized.
Uh, you know, some are closed source, like open ai, but if the likes of Deep Sea are making models open source that others can then pick up and iterate on, right? Um, that would help. And then the third layer, the one that you are talking about customers asking is that app layer, that how do you take these models and how do you, you know, bring value out of those models to cater to a need?
And as that, and as that app layer is gaining maturity, the security requirements will increase, right? I mean, what is my model doing and why is it doing what it is doing? What kind of guarders and checks and balances do I have?
And I think that will come for sure. Um, and I think that'll come for all models. Neil, I I gotta ask you another question.
Look, this is, we've been talking about this deep seek since the announcement every day on Textron Gang and in a lot of our articles and videos, you know, there, there's one, I don't wanna call it a rumor, but, uh, you know, some people are saying, I hate to say that 'cause politicians say that. Some people say, but there is a story out there that the reason they were able to train, deep seek, or this, this particular model so much faster and cheaper, is because they didn't kind of start from scratch. They, they, they were able to for however they got their hands on it, uh, open ai, uh, model, and then they kind of trained it off of that, if you will, or, you know what I mean?
And, and, and so that's what allowed them to do this faster and cheaper and on less powerful Nvidia and so forth. Is there anything in your testing that would give credence to that prove it, disprove it, or that's not something you looked at? You could, That's, yeah, that's not something we looked at, um, because we were doing an outside in evaluation of how the model is performing against checks.
You know, whether that happened or not, I mean, will, I mean, you know, remains to be seen. Um, but, uh, what I will say though is, um, from an architecture standpoint, from a model standpoint and the way they approached building the model and building the training, uh, and there is innovation here, which Oh, no doubt. Which I think no doubt, most Of the larger companies, everybody's going to benefit from that.
It will optimize how they're using their gpu. Uh, certainly, You know why it's the deal. And it's funny that it, it comes from the Communist Party of China, but this is what the open market's all about.
Yes. Right? If someone builds a better mouse trap, copy that mousetrap Yeah.
As fast as you can, right? And, and learn from that and, and, and keep innovating, because, you know, the other thing I, I feel with this is yes, it didn't do so well on your test. No doubt about that.
Right? 61 and 58% are pretty, I mean, those are hard to argue with. Uh, it will get better though.
I'm sure it will get better. And, and it, it, and that's again, part of this whole open source thing, right? It allows other people to innovate off of their work as well, which is, you know, is, is a great model.
Um, I, I think the bigger, the bigger thing though is that we were just discussing it on Text Trunk Gang this morning. There are so many different models out here right now, even within open ai, you know, when do you use oh 3 0 1 4? Oh, most people don't really know.
Well, it sta it versus another one. You know, how do you know what model to use? When should I use Deep Sea Car One versus, uh, Gemini or Llama or what have you?
So I, I think we're gonna develop in a world that's kind of like cars. Some people drive a Maserati or a Ferrari, and it costs a lot of money, or a Bentley, other people drive a Buick or a Cadillac, or, and then other people drive Chevys. Mm-hmm.
And that's okay, too. They still get you from point A to point B, which is the, that's the mission. Yeah.
If this thing can get you from point A to point B and fulfill the mission at a fraction of the cost, market economics dictate that you'd be a fool not to use it. Yeah. I, So, you know, go ahead.
The, I, I think the entire ecosystem is still, it's still very early, right? 0 and you know, everything new, you will not like it. We still look back fondly to the initial versions of Champ Chan GBT thinking, oh, it was revolutionary.
Yes, it was. But now after you've experienced something so much more better, right? Even from open AI and from others, you will think the initial versions didn't really have, you know, that level of knowledge, or they were not as good as what is today.
Uh, and what will happen here is this innovation is happening at an extremely rapid pace. It's not even in gaps of two years. It, it's happening, you know, within months, right?
Weeks sometimes. I mean, week to week, these things change. It seems It's crazy.
I mean, these guys came out with their model, and then Alibaba came out with a model saying, Hey, we think we have something better. And, and that's a good thing, right? Because early on, it's The market.
Yeah. It's the market. You want this kind of innovation happening.
You, you want this kind of disruption happening, and then everybody benefits from that. So I, I agree with you. Let me ask you to put your Qualys hat on now though, 'cause we only have a few minutes left.
Speaking now as C-T-O-E-V-P cloud platform or Qualys, how big a challenge are these AI models in, in making security better or trying to secure them? Right? There's two aspects.
One is harnessing AI to be a better security company. One is as a security company trying to secure against AI being used by bad guys, right? Yeah, I, it's, it's a really good question, right?
Um, I think using ai ML and AI in security has been happening for a long time now. We have, we had machine learning models embedded in our platform. We had them for years.
Uh, I think when L LMS came out, large language models came out. Uh, it was a little bit more disruptive because it, in some way, it socialized using machine learning. Earlier to do ml, you needed a data science team.
You needed a team of experts that really understood how to train these models. Now, in some cases, you have folks, you know, that take an LLM model and just doing prompt injection, they're able to, you know, build applications that can add a lot of value. So now from a security standpoint, of course, using AI ML to build security solutions, it's been there, I think LLMs will help accelerate that.
We are already seeing that. Uh, we've introduced a lot of new things in our platform just in the last two years over that, right? The bigger question now is, as especially large language models, which are more predictive, they're not deterministic.
If you ask it a question, it'll not give you the same answer every time, right? It's just how the underlying architecture is. It's getting better, right?
If you ask it a math question, it, I mean, it is giving you good answers now, right? And especially some of these newer models are really good, but more from a business standpoint, when you are asking it a question, you are expecting it to answer within the context of your business domain. And so, guardrails become extremely important because you are saying your chat bot, let's say, is representing you as an organization.
And if your chat bot gives an answer, then you are held to that answer. You can't say, I had a chatbot on my website and it gave an answer. That answer was incorrect, so it's not my problem.
You can't say that, right? Uh, so that's where, you know, more checks, um, you know, building the right kinds of gates is becoming, becoming increasingly important. What we are seeing right now is people saying everything is in beta mode, right?
Uh, that, hey, we are releasing something, but it's in, in bera, which is fine. Um, I think the industry is maturing. Obviously the models will mature, the security ecosystem will mature, right?
Just the way we introduced total ai, we looked at this as a gap, even when we were looking at it internally to say, okay, our teams are blowing models. We don't even know what's going on. We talked to a lot of CISOs and they said, we have no visibility into what our teams are even putting in charge GPT or perplexity.
What kinds of questions they're asking and what kind of information is going out, which could then be used to further pre-train those models on proprietary data, right? So you need all these checks, and I think the realization is there. And, you know, we, we obviously took a major step in saying, we are putting out a solution that will help you understand your AI ecosystem, understand your vulnerability posture, your security posture, and then of course, as you're deploying your large language models across your enterprise, you know, what is the security, the compliance, the ethical guidelines, uh, the jailbreak, uh, you know, capabilities, you know, how, how do you manage all that, right?
Uh, that's where we are at. Uh, we are obviously adding a lot more capabilities into our platform, into total ai, uh, Qualys, total ai, so our customers and just the larger, uh, community can benefit from it. Excellent.
Dilip, we're out of, we're overtime, actually. But thank you so much for coming on. Keep up the great work, everything you spoke about.
com whether you want to go check out the blog articles on, on this particular testing and, and story, or you want to find out more about total AI or about Rock, or anything else. com is, is your starting place for that. Dilip, I hope maybe we'll see you in San Francisco during RSA week, if not a QSC or you're always welcome to come on here and chat with me.
It's a pleasure as always. Likewise. Thank you, Alan.
Good conversation. All righty. Diwani, C-T-O-E-V-P Cloud platform at Qualys here on Techstrong tv.
We're gonna take a break. We've got a lot more coming at you today. Stay tuned.
We'll be right back.