AI-Driven Insights: The Future of Security and Log Analytics with Modern Data Lakes | Predict 2024
Tune in as SentinelOne’s deputy CISO and VP delve into how AI and cloud-native data lakes are disrupting traditional log analytics, yielding deeper and more meaningful security insights.
Transcript
Hello, everyone. My name is Steve s Seibert. I'm regional Vice President here at Sentinel One.
Uh, the topic for today is ai. Uh, it's certainly something that is, has a lot of hype, uh, I think still sort of developing in the world, and, and everyone's trying to figure out how to take advantage of this, uh, next gen technology, uh, especially in the security world. Now, some is hype and in other cases, uh, we're starting to see, see real world cases where, uh, it's being applied and applied effectively for security posture.
So, as you can see by the title here, it's the future security and log analytics. So how do we take enterprises to a modern security posture? And so with that today we have, uh, TNA one's deputy ciso, Josh Black welder, and he is gonna help us out with that, uh, or at least explain his environment a little bit better, uh, from his perspective.
And so with that, uh, I'll start with my first question. Um, I'm gonna tee you up here, Josh. The difference between, uh, Sentinel one, singularity Data Lake and a traditional sim is what?
Yeah, good question. So let me tee this up a little bit. Um, if, if you think about the richest data source within your environment, it's typically your EDR data.
And your EDR data contains a wealth of information. Uh, it keeps track of every process, every network connection, every registry change, every, everything that happens within your system. And it catalogs it in such a way that you can, uh, navigate through this data to back up any process through rollback or see a a, a process tree.
Uh, we call that storyline where you can see how, where a process originates on your system and everything that it does, um, later on, uh, through activity to your system. Um, and so the amount of data generated by your EDR is incredibly high. Uh, you're talking about two terabyte a day per thousand endpoints.
So you can just do the math, that's like a general number, but it's a tremendous amount of data. The traditional model is if you, if you wanted the full richness of your EDR data, you would ship this off to a sim and it's very expensive to ingest that amount of data into your sim and consequently, you also lose some of the navigability of that data when you flatten it into a SIM model, right? And so, uh, the newer perspective on, on, uh, security data lakes is to have your EDR data right there in your security data lake and to bring your already flattened data next to it.
For example, your Octa logs, they're already flattened. When you go into the Okta console and you look at an Okta log, there's no richness or nuance there. It's the same as if you ingest it into your sim.
Uh, the same goes for your AWS cloud trail logs. Uh, they're already flattened, and you bring them right next to your other logs within your security data lake. So now you have all of your security data right there in a single security data lake.
And these data sources can be correlated together, they can be enriched, and you can do all sorts of advanced things with the, with these logs. Well, it makes sense, and it sounds like it has, uh, some good promise. Um, when you ingest these other logs into the singularity data lake, how do you normalize the, the logs so that they can be processed like in a, say, an intelligent manner?
Yeah, great question. So we can pick up logs in many different ways. Um, the preferred way is to use one of our partner integrations in our XDR marketplace.
So for example, you can go to our XDR marketplace within your EDR console, and you can select something like Okta or Zscaler. And we've already, we've already, uh, created the partner integration with these vendors. And so you just input a single API key.
And when you do that, it configures the ingest of the logs and it formats them appropriately to the OCSF framework. That's the open cybersecurity framework. There's no engineering, there's no manual parsing required.
We've already done that for you. And in a, in, in addition to parsing the logs to the OCS format, we also give you default XDR actions that are appropriate for each vendor. So for your identity provider, we might have a button that causes, that causes the user to have to get a re authentication prompt, or it terminates them as a user, et cetera.
Right? And none, none of that requires any engineering or integration on your side. Um, additionally, we can pick up logs from S3 buckets, or you can install an agent in your environment, um, that can forward logs into your environment.
The ingest pipeline to our security data lake is very simple. Uh, incoming data streams come in, we alert on those incoming streams. So in some cases, you speed up your alert pipeline by a couple minutes.
There's no indexing or processing required before an alert can be generated. Then we parse to the OCSF format, we rotate the data 90 degrees to calmer format, and we save those logs within S3, and that's the entire ingest pipeline. Because this is a modern cloud native architecture, there's no indexing required of your data, data, that is a legacy concept in engineering.
Uh, it's not required in modern security lakes. Well, it seems simple, uh, seems like a good way to ingest logs. Um, but tell me more on how the logs are actually queried.
Yeah, yeah. So once your data's, um, you know, parsed and ingested and, and stored in, in that column or format in S3, S3 is an object storage, um, that resides within AWS. Um, we have this massive global compute layer that rides in front of your objects, and we utilize tens of thousands of containers that break up your query into small pieces, and each container will grab like a hundred meg of an object, right?
And then it will perform a search complex or simple or whatever. And then all of those results are quickly returned back together and returned to the user. And this happens in milliseconds.
So our, our average query time on our security data lake is less than one second. And obviously that can change depending on how far back you are going or how much data you're querying, but the average response time is less than one second. I guess the question is then why have so many folks given up on traditional sims?
So, yeah, good question. Um, if you look at the, the SIM landscape is that a lot of 'em are based on the same flawed data architecture. Um, the, the underlying data structure has not changed a lot in the, in the last couple decades.
And a lot of these vendors have added a lot of bells and whistles on top, like user behavioral analysis or threat feeds or, you know, special correlations or rules, et cetera. But none of the vendors really attack the underlying problem with performance and scale and the architecture of the underlying data. And many of the cloud solutions that exist were lifted and shifted from data center technology to the cloud without taking advantage of some of the concepts of, uh, autoscaling and cloud native concepts.
And so the security data lake that we use at Sentinel One was built for the cloud. It's a cloud native modern solution. So tell me more about why, uh, having a fast security data lake is important.
It makes sense on the surface, but can you dig a little deeper into that? Sure. Yeah.
Um, so we need to talk a little bit about the current landscape and security. Uh, we're what in we call the post EDR world. And this concept is, you know, in the traditional infrastructure you had a perimeter and you had posts and servers, et cetera, within this perimeter.
And EDR did a pretty good job here. Uh, but now with the rapid adoption of cloud services and SaaS services, sometimes the perimeter is not quite clear where your data has gone from your organization and where it resides. And so you need to have the ability to respond and have visibility of your data, whether it's on your endpoints, whether it's in a SaaS solution, or if it's in AWS or Azure Cloud, right?
And so the concept here at Sentinel One is that we need to be able to have visibility of all of these environments, right? And you need to be able to see lateral movement from endpoints to SaaS services, to possible cloud environments, cloud IDPs, et cetera. And, uh, traditional EDR cannot perform this particular task.
So one, you need the logs and the XDR actions for all of these environments in a centralized place, right? And at Sentinel One, we're using AI to intelligently stitch together these events from endpoint to SaaS to cloud to identity, et cetera. And so within the S one console, when you have an EDR event, uh, if there's an associated Okta event, it will actually pop it up in the same window and you'll see the related XDR event to that EDR event.
And that, that auto stitching of events is being done with our purple ai. And that is leading into a product called SDL graph, uh, singularity Data Lake graph, where in a single window, you can see basically how an event moves from endpoint to active directory to a file server, to a SaaS solution to a cloud environment. And you can fire XDR actions associated with each one of those events from that threat graph.
And that's an important concept. And we need this in order to react quickly to the current industry attacks that we're seeing. Uh, some of the current attacks are coming in with very heavy resources on the people side, and a lot of automation, uh, a lot of auto automation.
And so we need the ability to respond more quickly to events, whether they're on the endpoint or they're within SaaS or wherever they may be, may be. Um, secondly, uh, when we talk about your, your security data lake and ai, AI in order to, to answer certain questions, needs to have a conversation with your data lake, right? So say, um, you know, Steve, you asked the question of tell me the quickest exploit path into my environment, and you have all your log sources set up within your s singularity data lake.
The AI first needs to go ask a bunch of questions to figure out where it needs to go. So the first question that AI might ask is, tell me all the resources within your environment. So it might say like, um, tell me all the compute, all, um, you know, the servers, the, the endpoints, the laptops, et cetera.
So it gets a list of assets back. And then it might ask, tell me the patch level on each one of those instances, how are they configured? It may say, tell me the code that's running on the applications published on these compute instances.
And then it might ask, well, let's determine the location of these assets within your environment. So it may look at your cloud configuration to see what public network rules are published that map to these particular applications, and then it may ask additional questions of, uh, proximity to the edge, et cetera. So you can tell, like the AI needs to ask lots of different questions to get to the final answer or outcome.
And so if you think about a traditional sim, if each one of those conversations is taking 20 to 30 seconds, and the AI has to ask maybe a hundred questions to get, uh, to the, to the right outcome that it needs to present to you, uh, that's gonna take about 50 minutes, you know, to return that answer to the SOC analyst. And you know, the SOC analysts, you know, they get distracted on d different things. Um, whereas, you know, if it can return each of those questions, those answers within a minute, uh, are within a second.
That's a hundred seconds. Uh, it's only like a minute and a half, uh, response time to the SOC analysts. So that, uh, having a performant data security lake is really important, uh, to be able to make these AI analysis quickly across your data and respond in a rapid manner to your, to your analysts.
And the other thing is that it can just explore a lot more paths as well. So a performant, uh, data lake is a, is a necessary part of having effective AI technology for your security data. Um, lastly, another, uh, piece of AI that we have in our Data Security lake is natural language, uh, querying ability of the security data.
So, um, there's actually like four levels of AI associated with, uh, the ability of a, a user to ask a manual question of the, of the security data lake. So for example, you could ask, uh, tell me all the logins, uh, to my environment from outside the United States. And the first level of AI will take that natural language and convert it to a technical query.
And we call this guided ai, where the a the, the AI doesn't just go off and do things, it presents results back to the user for verification so that the user can make sure that they're appropriate and accurate and see each step. So there's no black box to our ai, and this prevents like halluc hallucinations from causing, you know, um, inaccurate interpretation of data, et cetera, from the AI model. So, uh, that first level of AI will return that technical query, um, and then it will execute it against the customer's data.
Customer's data is never goes back into the model to train it. Um, the second piece of AI will attempt to summarize the data for the user. So in this case, it could say something like, you know, the majority of your users log in from outside the United States on a regular basis, but this one user has never logged in from outside the United States.
You might want to investigate this one further. So the AI will attempt to intelligently summarize the results of the data. Um, then the raw data is presented to the user, all the outputs of, of that particular query result.
The third level of AI will provide the user with suggested follow-up queries that the user could execute and they could just click and execute those. So in this case, it could present, uh, tell me all the commands that the user, uh, performed while they were logged in from outside the United States. And you could execute that follow up query.
The fourth level of AI will present suggested XDR actions. So one possible outcome could be like, this behavior looks at, um, malicious, we suggest that you put this user in an EDR network isolation zone. So you could click that and then the EDR agent would isolate that host on the environment.
And so these, um, this AI technology that fronts the, the data lake is very useful for every level of person within your organization. For example, the entry level SOC analyst, it helps them understand, one, how to write a query, how to interpret the results and possible things they could do to mitigate actions. But it also helps the most technical users.
The more technical question you ask the AI model, the more technical response that you'll receive. So even your most advanced analysts, if they ask a very pointed technical question, it'll respond at that level. And then third, your, your management who may not, you know, keep up their, um, you know, sim query language can go in and self-serve and, and get questions answered from the security data lake without having to ask an engineer to write a query or wait for the platform to, you know, generate it, et cetera.
So it enables self-service of, you know, CISOs, et cetera, uh, on your security team. So huge, uh, benefit across your organization using this ai. Uh, makes a lot of sense.
And you talked about, um, the growing threat landscape being exponential, the heavy resource, I think you called it. Um, it, it seems like AI is gonna be a powerful weapon in this, you know, this ongoing, uh, challenge. Um, how do you see AI helping solve, uh, other security problems?
Yeah, Um, you know, some of the other things that we've seen out, um, in the industry is some of the new SEC rules, right? On reporting security incidents and, um, their possible effect, material effect to your, uh, financial, um, standing. And, um, you know, you think about like how many events, um, a security operations center processes, um, you know, are we certain that the team processes processed every event in, in a proper way?
Well, you know, this is another possible use of ai. AI is really good at giving us probabilities and of outcomes, right? And so you could actually feed all of your soc work back through your AI models and say, uh, you know, of all the cases we closed, which ones are the highest probability of being improperly closed?
And we could flag those for quality, uh, assurance analysis by, you know, the team and have another analyst review the work and make sure that they made the appropriate action for each one of those. So, um, you know, AI can definitely speed our efforts in terms of, uh, you know, performing the initial work, but also helping us catch, you know, maybe where analysts made the wrong decision or call on a particular, uh, case. And, um, you know, bring that to the attention of the team and have those reanalyzed and, um, you know, provide better outcomes to the team.
All right. Well that's great. Um, and, and much appreciated on your insight.
Uh, certainly you're a busy guy and as a security company, we're a heavily targeted entity. And so no doubt we have to constantly be looking for, uh, ways to, to, uh, safely enable the company, our products and our clients. And so, uh, I think we really appreciate everything you've shared with us today.
Um, look forward to, uh, seeing this continue to develop. 'cause no doubt we're in the early innings of all this AI and the ML and everything else, right? So, uh, but uh, certainly it's going to be integral to everything we do as we move forward.
So again, Josh, appreciate the time and once again, thank you for having us be part of the day. We appreciate it. You're welcome.
Take care.





