Understanding the Value of OCSF for Security Practitioners with Amy Pham and Anthony Johnson at Techstrong Con 2024
When every vendor in your security ecosystem uses their own data schema, this creates significant barriers to gain value from data, complicating efforts in AI and analytics. The Open Cybersecurity Schema Framework (OCSF) is designed to solve the absence of a common, agreed-upon format and data model for logs and alerts across vendors.
The experts at SentinelOne will guide you through understanding OCSF, how to leverage it in your organization, and its benefits for your company’s security posture.
In this session, the experts at SentinelOne will guide you through the following topics:
- Understanding OCSF as a Security Practitioner
- How to leverage OCSF in your organization
- Benefits of OCSF for your company’s security posture
Transcript
Hi, everyone. Thank you for attending today's Techstrong Con Virtual Event. We're SentinelOne and we focus on protecting your organization through AI powered security.
We're super excited to be here today, today to discuss the value of the Open Cybersecurity schema framework, also known as OCSF, which you'll hear me mention, uh, quite a few times throughout this webinar, and how security practitioners can leverage it to protect their business. To quickly introduce those on the call today, I'm Amy, a product marketing manager here at SentinelOne, and I'll let Anthony introduce himself over here. I'm Anthony Johnson, uh, field CTO over our, uh, data, uh, platform.
Awesome. So before we jump into today's agenda, um, a few housekeeping items. If you have any questions throughout this event, uh, please feel free to drop in the chat and we'll have our team answer as soon as possible.
Today we'll be discussing, um, a little bit about the challenges of data, what the open SEC cybersecurity schema framework is, and why OCSF is important for security. So we'll take a look at some of the challenges that are currently, um, apparent in our data space. To start off, start off with is scalability.
Scalability is a huge concern for our customers as digital footprints expand with it. The data we generate traditional SIM solutions that are built on top of a legacy database, um, or document oriented search engines such as Elasticsearch. They typically don't scale well for analytical workloads.
Secondly, um, something that I'm sure we're all very familiar with is cost concerns. We hear from our customers that cost is the number one pain point along with performance when it comes to their data, because as we all know, um, data retention really isn't just about storage, it's also about what you get out of your data. And so, you know, with budget and with data, we wanna make sure that, you know, consumers aren't torn between storing their data, but also trying to make their budgets work.
Uh, thirdly, and the next two are the two topics I really wanna focus on today. The first one being flexibility and integration with so many different security and IT tools, flexibility and integration is key, but many traditional, uh, solutions currently make it very challenging to seamlessly integrate. Raw logs are difficult to search and understand, and this makes threat hunting a challenge for our security analysts because data ends up locked in tools that are difficult and costly to adapt, and many vendors have their own data schemas.
And lastly, siloed data more than anything else, the value of data is its accessibility and use, not just its storage. As I previously mentioned, when we don't have a centralized data strategy, this same data gets duplicated across many different tools. This leads to inefficient decision making, but also results in unnecessarily high costs because the same data is continuously being stored in many locations.
So next, you know, while some data serves uniquely one use case, frequently the same data powers multiple use cases, and ideally, your data lake or whatever you use to store your data, allows for data to be stored and ingested once, but be used multiple times for various different use cases. Um, we'll jump into a few examples that I'll talk about here, and one of these, um, being cloud and Kubernetes logs and metrics, there are three types of, well, three use cases that I see this data being used for. One, IT teams, they use this data for cloud utilization and efficiency, but secondly, engineering teams also utilize this data to look at their performance and troubleshooting any data problems.
And third, um, which we're gonna focus on today is security teams and security teams also use the same data to strengthen their security posture and to, you know, find any vulnerabilities to help and to help fix that. So although it's the same data from the, just this example, we see three different teams utilizing this data and having it stored in one location would be ideal for costs. Another example could also be with, um, Okta, if your company uses Okta.
Um, it tells IT teams about possible performance bottlenecks, um, if users are experiencing problems that peak loads or connection issues from different regions or geos. But at the same time, security teams also look at this data to find any anomalous behavior such as users logging in from different locations or particular times that would not be, um, considered standard. So I talked a lot about the problems, um, in the data space, such as flexibility, not having flexibility, uh, ease of integration and silos, but also the importance of multipurpose data.
And so how do we put these two ideas together and how do we solve for our data challenges while also getting the most out of our data? Well, some of the key players in the cybersecurity industry came together, contributed and created the open cybersecurity schema framework, also known as OCSF. And OCSF is designed to tackle that same fundamental challenge that a lot of teams are facing in the security analytics space, which is the absence of a common agreed upon format in data model for logs and alerts.
Historically, the absence of a common model has created significant barriers to detection, engineering, threat hunting, analytics, and this also causes problems for data and AI projects logs and alerts that leverage OCSF use a common set of fields and formats regardless of whether it's coming from different sources or vendors. And this overall simplifies data ingestion and promotes scalability. And as I mentioned historically, you know, every vendor in the cybersecurity ecosystem uses their own schema.
And this makes it difficult for security teams to really understand the data, but also see it in all one centralized location as an open source effort. Um, OCSF aims to address that by establishing common schema, leveraging the critical mass and willingness of some of those major players that I've talked about. And in this slide, you can see, you know, here's a quick overview of all of, not all the players, there's about 150 different players, but these are just some of the key players, um, that I've highlighted here today.
And as a contributor to the open source effort here at Sentinel One, we also leverage OCSF to standardize third party cybersecurity data and our own data to enhance efficiency and help, you know, teams prioritize their security operations over data acquisition challenges. We're the only true unified security offering in the market with extensive security and data analytics, detection, investigation, and response, all powered by ai. And with our open architecture, this means that, you know, you can bring all of your data into a unified data lake, um, ingesting critical business data, um, into the OCSF framework and support your existing security investments and tools.
So now I'll hand it back over to Anthony, um, for him to talk about how organizations, um, can utilize OCSF and a few of those benefits for their security efforts. Awesome. Uh, amazing, amazing introduction to OCSF.
Thank you, Amy. Uh, so before we get started, um, I did wanna kind of touch on, you know, we're gonna dive a little bit deep into OCSF and schemas, and I feel like, um, rather than making assumptions that everybody knows what a, what a schema is and, um, what, what, you know, what these, what this is all about, figured I wanted to level set and give you guys a, an insight. So, schema, if, uh, if any of you have worked with databases or worked with data in the past, really all we're saying is how data is organized.
So on the right, you can see I'm kind of showing a, a sort of a sample, um, you know, diagram of a, of a schema for a firewall. This is obviously not complete, just an example. But, you know, in this one, we'd have a field name source, you know, and that would map to a logical type called IP, adder, source port, logical, logical integer.
1. You know, for those as an example, um, this is, um, you know, why is the schema important? Uh, is schema establishes, um, essentially data contracts.
And so when we have 150 or so, um, companies contributing to OCSF, um, what they're really contributing, what they're really contributing to is creating a data contract that they agree that they're gonna adhere to. And, uh, you know, and data contracts, you know, shouldn't be changed without proper notifications. What's great is OCSF creates something that's unambiguous as far as how, um, you know, companies go to it.
There's one way to do it, and that's documented. Um, touching on this, and I learned, uh, not too long ago that not everybody knows what schema on Read or schema on Rite is, and I feel like is, uh, it's my, I need to introduce this concept to potentially security practitioners, security analyst. And this is really just the idea of where we adapt the schema itself.
So, um, you can imagine that we have data coming in, um, and that data needs to be transformed before we write it to disc, and that is what we call schema on write. So basically on disc, it's on, you know, it's been burned into that, that disc as the schema. Uh, the other option we have is schema on Read.
Uh, whereas, you know, the data can be in many different formats on disk when you read it, uh, you actually transform the data into the schema at that point. Um, you know, our, on our Sentinel One platform, we support both ways of doing it. Obviously, we want to, you know, prefer to write the schema, uh, on disk so that we're not using all of that, that effort to transform it on the fly.
Um, but these, these concepts are important when you start talking about data systems and where does that schema get adapted, how it does writing it to disk is, uh, more efficient for queries. Um, you know, reading it is more flexible with being able to make modifications. Um, I will talk about though, you know, why we need the OCSF.
So the reality is, is that naming things is actually pretty hard. And so my example above, you know, we have a firewall. It's got a, it's got a field name source and it's an IP adder.
Um, the Jason for that might look something like we have below. And, uh, this is a firewall rule, so we're just gonna have a deny in there. And then the thing is, is, well, what if we bring in another firewalls, um, you know, firewalls event, you know, and all of a sudden you find out that their fields, you know, know, they decided they liked capital letters to begin a, a field, and then they, um, you know, they ended up abbreviating source is just s for them.
Export works. It's efficient, and that's how it goes. And then, you know, a third example might be where, uh, you know, a company or the person who created and parsed that data decided they didn't even need a port field, they're gonna use a pretty well, pretty standardized representation of an IP address with a port, is to use a colon in between the two.
And this, this vendor, this piece of data, just decided to merge them all together and represent them as one. Uh, so, you know, so if you look at this one, it's like, well, which source am I using? Which data am I gonna get?
And if I pick one, I'm probably gonna have some challenges querying this data and getting a consistent results, even though everything is a firewall. So the OCSF, you know, naming things is hard. Everybody has an opinion and OCSF community, those 150 companies basically defines the name of everything and the types of everything.
And, uh, this creates basically a data contract where you're either wor you're either adhering to it or you're not. And so this OCSF example uses a field called source endpoint, and within that is an IP and a port. Um, it's logically there.
And part of the reason they, you know, when smart engineers come together, they tend to design things that are more future proof. And part of that is that the future may not be IP addresses, and the future may be some other form of communication. And so a lot of what they're doing when they're building these comm, these committees, and they're building these standards, is they're asking those questions.
You know, we're not all data scientists, data engineers, but they're asking those questions of how is this data gonna change in the future? And that's where you see things like source endpoint, and it has a field called ip. It's not source endpoint IP as an example.
And then lastly on this one, uh, OCSF uses the concept of enum. If you're not a programmer, you know, you might not know what an enum represents. An ENUM is basically just a list of valid, uh, ident valid keys for that field, or valid, um, valid values for that field.
So the enum in this one is activity id and, um, activity ID is actually five means refuse. And in this example, activity is actually optional. That's the human way to read it.
The one that's actually required is activity underscore id, which is, you know, allows you to populate activity if it's not present in your value. So there's already intrinsic value in just having standardization, um, by being able to have, you know, an activity ID that maps to a concrete, uh, activity. So what does companies adopting in the OCSF schema really mean?
And for me, um, I like to, I'm a big, I'm a big data contract, and really we get into consistent field naming. So no, no longer are we gonna have discussions over, over source. Uh, how do we, how do we specify source?
Instead, it's gonna be source endpoint that takes that ambiguity and that takes that headache away from people. Um, consistent value types, um, this is, um, one of the biggest values there is, um, having consistent field names is, uh, is a bit pointless if everybody is just writing different values. So, um, before we had port, you know, port number that was there, and that was specified as an integer, that could also be present as a string inside of the, um, inside of the data as well.
And if you've ever, if you've ever written a script or done any programming or done any of that, you'll know that a string value of 80 is not equal to a, a number value of 80. And so we solve that problem, uh, in a way where we can actually, you know, validate the data, um, coming in. And then for me, um, as a, as a, as a data engineer, I guess, um, the enforceable data contracts is one of my, uh, biggest demands.
And so mainly because you have to think about what happens if something changes in your data. So let's say you set up a, um, a, a rule for, um, validating invalid logins or data infiltration of, of important data. Well, if the vendor, you know, supplying you those events decides that file name or, or some other field should change to something else, then, uh, any alerts that are kicking off of that, or any queries or any dashboards that populate from that data, if they just change that without regards to the end user, well, how it's being used downstream, then, uh, basically that alert will never fire.
And if that alert never fires, then you never actually know what happened. So data contracts for me, are something that is a must. And what's great with the OCSF is you now have 150 companies all watching for data challenges and, and having the same concern.
You know, not all of them will have the same concerns. Some of them will be like, we just wanna change whatever we wanna change. But the reality is, is by being a committee, by having standards, we slow things down and we have discussions and dialogues on the right thing to do, and that slowing down actually is beneficial, uh, to everybody.
So really does a name matter that much? And so, um, I have, I have managed large systems before. I have, uh, you know, had teams writing whatever they want into this data store.
This is, this was over a decade ago, and over time, um, data becomes harder and harder. And so you want to have a, if you wanna have your own single schema or you wanna adopt somebody else's schema, translating data, uh, idea of, of parsing data, either at query time or, or right time, uh, is money. In fact, I heard a statistic today that 60% of data lakes, uh, run into issues even failing because they struggled with translating data.
And, um, you know, and then searching for the right data. So translating data is a problem. We're always playing catch up.
Searching for data can be, can be an issue because if you, if for your data to be usable, if you have a translating layer, then, um, then anytime you, when a query something, either it's either the effort hasn't been done to translate that data into something you can query, or worse, the data was never even imported because, because, uh, nobody knew it was missing. And so being able to effectively search for the right data, see what's there, and have confidence is, is pretty important. And this is where systems that actually adapt, um, to turnkey, you know, bringing data in, not just in any schema, but in a schema that you, you can depend on to query data such as OCSF is really important.
Um, lastly, um, and one that's dear to my heart is searching for the same value across multiple data sets. And so we were talking about firewalls earlier. Um, you know, I don't want to have to go query every single firewall.
My company may use 15 different firewall, uh, vendors and versions, uh, you know, that we've just collected over the years. And if every one of them has some quirk to how the data is outputted, it becomes where you're not getting great answers when you ask questions of those different types of data. And essentially this, this all becomes inefficient and wasteful.
You know, not all of us wanna spend our time translating data, you know, asking questions that we can't get the right answers for, and having to dig down into individual components in our infrastructure to understand if it's secure or not secure. And so really what it, what it means is for a security posture is that time is money and time is risk. Anytime you are, you are not able to get a ask a question, get an answer anytime the data is not there.
Anytime we're spent really kind of investigating, you know, that's risk to the business. That's why hackers are spending, you know, weeks, you know, days, uh, months inside of these companies because the data is just too hard to find when their bad actors are doing something. So basically, any delays in importing is a challenge.
Finding the answer is a challenge. Delays in resolving or preventing an attack, definitely a challenge to your business. And so really what we're dealing with is data isolation across the board.
So companies to date, most of them, not these 150 50 companies that are now engaged in trying to get this, this standard out there, but they've realized that, you know, they have their own schemas and everybody having their own schemas creates an isolation. You know, e everybody, whether it's the elastic common schema, whether it's the computer information model from Splunk, or whether it's just a proprietary schema that your company has for its application data, all of it is its own schema. You know, data, you know, data itself has its own dashboards and alerts.
What do I mean by that? I mean by that fire, you know, firewall A, firewall B, web server A, web server B, since their data is isolated in its own types, to be able to create an effective dashboard, you need to have one dashboard for each one of them. That is actually the standard that happens.
Now, trying to get a firewall dashboard across all vendors has been pretty difficult, uh, to this point, because we're having to adapt 20 different schemas, um, vendor specific apps to monitor. Um, you know, because we're, you know, because we're, uh, doing things at a vendor isolated level, we tend to have these very specific apps. And then of course, the way that they respond, the way that they did adapt, the features that they support, the metrics and statistics, they show everything is proprietary, create, creating isolation.
And then really, companies up until recently really didn't want to share. It was sort of like your data became sort of a vendor lock in, if you would. And, uh, you know, and what ended up happening was, is they were like, oh, well, I've got friction.
That friction keeps you at my company. And if you do, if you, if it's hard for you to move, then it's hard for you to go to somebody who might be securing you better. Basically, we all lose in this scenario, you know, ultimately the companies that don't share collaborate, they're dealing with customers that end up getting hacked.
They end up not being secure. Bad things happen if we're, if we're basically playing business when it comes to securing, uh, our customers. So what does OCSF bring to this dire picture?
I just, I just painted well, companies a hundred, at least 150 of them and growing are collaborating. We're coming together. We're saying it's important that our customers are able to take their data with them, are able to interoperate with best of class tools, and are able to, um, you know, have something that they can query that, um, that allow, that they can take with them to the next product and move forward with, uh, data itself becomes more consistent, meaning that we are, because we're all pulling data with similar fields, now we can start to see that, um, you know, we can start to see that, uh, you know, see across all firewalls, see across all web servers, see across all endpoint data.
This, this is the promise of OCSF as we move forward. Um, and, um, you know, speaking, well, I guess we, we, I, I already hit on the, the, the spanning multiple components. Um, really the data itself is consistent.
Um, that is the key takeaway there. And then what's happening now is that companies as, as we start to take away that data friction, that is my data, and, uh, nobody else can, can really read it without spending a ton of resources to try to adapt it. And then guess what?
They change something, it breaks. Well, now we're sitting here saying, well, no, let's have standard. And this basically makes companies have to compete on value and encouraged.
And, you know, and, and the companies themselves are encouraged to innovation. Uh, innovation is really what's going to drive us towards a better security posture. And, you know, we want, we want companies to be encouraged to, uh, to innovate.
We want companies to not just stop because they've, they're like, we've done enough and we can just sit back and use sort of a, a, a job security or, or product security to keep themselves positioned in these companies. It's like no com, you know, customers deserve better. And I feel like these, these companies are really getting it, that we owe it to protect the future of, of our world.
Um, so ultimately what happens is, is we all win. Um, is because we have this new standard because we can interoperate, because we can talk to each other. Um, but not only do we all win, um, everybody wins except bad actors lose, uh, companies, you know, companies are always pushing forward to make sure that, that we're getting the best solution.
Um, we're not sitting, sitting back and just letting, uh, you know, we're not, we're giving cu we're giving customers choice. And when choice comes an ability to, to, to push for innovation, push for new ways of doing things, and ultimately that is actually what a lot of OCSF is empowering. And that's very, very exciting for me.
Um, I like to, um, to, you know, talking about, one thing is talking about OCSF is one thing. You know, we're, our platform itself is already, um, is already standardized in OCSF. This is OCSF on, right?
And it's also on read. Uh, we also have a lot of old data, and as a result, we actually have customers who are already getting advantage of using OCSF within our platform. And I think it's a good quote.
Um, I will read it for you just before Sentinel one, much of our time was spent analyzing data across multiple, multiple systems due to the technical cost of getting an integrated view, rather than focusing on security operations. Having a platform that provides OCSF ready data connectors is a game changer and allows my team to focus on detection, triage, investigation, and response rather than data architecture. Thank you, Owen.
Connolly, uh, it drives home the fact of we want our security analyst and we want our companies to be focused on finding the bad actors, not on figuring out how to get data in that we can analyze to find the, act, the, the bad actors. So this is an awesome quote. Um, so that kind of, that wraps us up.
I appreciate everybody for attending today. And, uh, please, uh, ask any questions and, uh, reach out. Uh, we have our website at the bottom.
Reach out if you, um, would like to know more. Thank you very much.

