The Security Copilot Journey with Microsoft Security
Nick Goodman, Partner Product Manager Security Copilot at Microsoft, shares Copilot’s evolution from a chat-based assistant to an integrated AI tool embedded in security workflows. Initially designed to assist analysts through queries, Security Copilot quickly adapted to automation and workflow integration within products like Defender and Intune. Key use cases include phishing investigations, SOC situational awareness, process verification, and e-discovery. Microsoft has introduced multi-tenancy, flexible capacity, and multi-workspaces to enhance usability. Additionally, new security agents will automate threat response, vulnerability management, and conditional access policy enforcement, reducing manual workload and improving real-time security operations.
Being secure is the first step towards AI innovation. Join Microsoft Secure and learn how to harden your defenses by exploring new AI-first tools, demos, and best practices. Register now: aka.ms/GestaltITSecure #MSSecure
Presenter:
Nick Goodman https://linkedin.com/in/nigoodman/
Delegates:
Karen Lopez DataChick.bsky.social
Girard Kavelines https://x.com/GKavelines
Jennifer Minella https://x.com/jjx
Justin Warren https://x.com/JPWarren
Transcript
Hello everybody. Thank you for having me here. I'm Nick Goodman, product architect for Security Copilot.
We're gonna talk a little bit today about how to use AI to accelerate progress in cyber and data security. As many of you know, security has unique challenges, which are both, uh, hard make AI hard in this space. Also, they represent the opportunity, so high volumes of data, data that changes in real time data that specific to the organization and that you have to look at in a, in an organization specific lens in order to, uh, to get effective answers.
And security requires deep reasoning. It doesn't look the same as other spaces. Security co-pilot, uh, had its first birthday, uh, this week, and we've, we've come a long way.
Initially, security co-pilot was designed as a chat interface where you as a security analyst or IT administrator would ask it questions and it attempted to assist you in your day job. What we learned early on, though, was that most people doing security jobs prefer to have AI embedded within their existing workflows. That often takes the place of a product or takes place in a product like Defender or intro or Intune or Purview.
And so we pretty quickly pivoted to having AI built into buttons, built into screens, uh, built into analysis flows within those products. And what we've seen over that time also was more people who are using the chat interface to do so for automation. So what they were doing is what we call prompt engineering, and they would iterate through different prompts to come to a security outcome.
And then rather than doing that every single time, they would save that into an automation and would run it in their environment so that they would get the benefits of ai, um, plugged into their automation systems so that they're getting it at high scale. And the top use cases we've seen in security operations have been, uh, phishing investigations. So this is like enriching phishing incidents not to speed up, uh, triage, uh, situational awareness of the sock itself.
Things as simple as who's, uh, on call, who's on vacation? Do I have the right mix of people in play for my team? We've also seen a use case that we call process verification.
That's if you have a, so you have a, a playbook for how incidents should be investigated. A lot of our SOC users are actually having a security copilot check to see whether or not resolved incidents followed that playbook or steps missed, so that if, uh, something important was missed, the team can go back and, and get that done. We've also seen security copilot customers moving into, uh, areas beyond the soc.
So, uh, open source, uh, analysis. So this would be like, say you have public content about threats or vulnerabilities using security copilot to parse the natural language from those webpage to help the analyst understand what the, what the nature of the threat is. We've also seen customers doing threat hunting from, uh, both Microsoft and open source intelligence, whether they'll look for TTPs or indicators across their cloud attack surface.
And finally, we've seen customers using security copilot to simplify the startup process for eDiscovery. So like if there's a pending legal action, the customer needs to go in and make sure the documents are held and that they're available to the legal team. Uh, that's a complex process and security copilots helping people, uh, simplify that.
We've recently introduced three major features that unlock customers, uh, to move faster and meet their enterprise requirements. So we've introduced, uh, broad multi-tenancy into the system. Many of our multinational corporations, uh, uh, that are customers have this sort of setup where you've got multiple tenants representing different areas, or perhaps they have a dev tenant and a prod tenant.
Uh, so the multi-tenancy is just allowing customers to use security copilot in ways that fit their business more effectively. We've released a new business model, um, or in the process of releasing a new business model that will allow customers to use security copilot with on demand capacity. So in the past you had to provision security copilot by introducing on demand.
We now allow customers to blend the best of both worlds. They can provision to get cost savings and they can go on demand to meet, uh, the spiky nature of security workloads. We are also introducing this concept we call multi workspace.
Workspace is just a, a configuration holder for security co configuration. It can be capacity, it can be user access, it can be which plugins are installed. And when you combine multiple workspaces with the more flexible capacity units, uh, it's really empowering customers to, to use security copilot in a way that meets their business.
And then finally, also related to this multi workspace concept, um, simplifies, uh, meeting regulatory requirements when you've got data that can't leave certain regional boundaries. So you can set up a workspace, say for your operations in Hong Kong and London and Central Europe, and make sure that you're meeting all of the, the local data privacy laws. Now I'm sure a lot of you are hearing the word agents out there.
And, uh, here at security copilot, we've actually been fairly quiet about agents and this is intentional. We've been building, uh, product and focusing on creating product truth and making agents that produce real value and not simply branding everything that's AI agents. So talk about why we are so excited about this space and then I'll show you an examples of what we're actually releasing.
So let me paint for you first my mental model. If we think a little bit about, let's go back to the concept of automation. Where does automation, uh, provide value?
Well, it creates a lot of value for security organizations 'cause it means you don't have to have a person doing every task. And agents aren't exactly automation. I think it's pretty important to understand the difference between them.
So I think of age as automation, like manufacturing line robots. They're very powerful. They do a lot.
And at the same time, they're also quite limited. So current generation automation is really a characterized by tasks that are predefined. They're repetitive, they're rules based.
They're not particularly adaptable. Anytime you, you come across the manufacturing line, you need to change manufacturing line. Well, let's bring in the engineers, let's figure out redesign it, let's build it again.
Same thing in security automation and current generation. Automation generally works for low compe complexity tasks does not work very well for high complexity tasks. I think of agents as like a self-driving car.
And you can see they're not apples to apples. Automation and self-driving cars are gonna continue to exist. That manufacturing robot and self-driving car don't solve the same use cases.
Agent fundamentally simulates intelligent behavior, which allows it to do things like create and execute dynamic plans. When I mean dynamic, I mean it will think about the problem or the task at hand, take some steps to pull threads within an investigation or a task that it needs to do reason about what it gets back, and then edit the plan as it goes. This dynamic nature of planning fundamentally differentiates agents from what's come before them.
That nature of of being dynamic in the planning also means that agents are incredibly adaptable. They don't have to be coded to, to every possible outcome. They effectively are given guardrails and, and goals.
And within those guardrails and towards those goals, they're able to learn and analyze patterns and adapt. I'm just curious, you know, kind of coming into this, the difference between you're explaining between just standard, you know, automation using AI and agents, and I mean, it, it first blush, it seems like an agent is really just a combination of several pieces of automation, you know, stitched together into something that's more holistic. Is the difference really in this adaptability and learning, or is there something else under the hood that's substantially different?
That's the, that's the fundamentally that's the big difference under the hood. If we talk about the out, the difference in the outcomes I think will be even bigger. So if I think of where we automate today, we automate the things that we know and the things that we couldn't, could afford to put a security engineer, uh, to go actually automate.
What we've seen in our evaluation of the agents we've built and the technologies is, is a really a different outcome, which is that you don't have to write every line of code, you don't have to point the automation at a specific thing. And it can only do that thing. Many of the agents that, uh, we're working on are capable of going way beyond the, um, the ability of a human to structure a plan to execute it.
And so that opens up a lot of outcomes that today are either really cognitively difficult to get to or, um, have just kind of defied automation in the past. So actually perfect segue into where we see the value of agents. So first, security organizations have threats and risk that are 24 7 and most security operations teams are not 24 7.
Maybe they have an on-call person, but they often are not staffed to deal with risk and threat 24 7. Um, and if you look at other pieces of the organization, uh, it administration more advanced, uh, incident response, threat hunting, vulnerability analysis, most of those teams also are not staffed 24 7. And so agents give us the ability to apply, uh, reasoning and judgment all of the time at the point of the risk or at the point of a threat instead of waiting for business hours.
Um, second, I said the word judgment. Security often requires judgment. This is where traditional automation does not stand the test very well.
Um, there's just too much we cannot automate because it today requires a human to look at it. And perhaps most exciting is that the cognitive abilities of agents are reaching the point that allow us to be proactive on tasks that were just so difficult that, uh, security teams generally didn't get to them very frequently. And I've got a concrete example of that here in a minute.
So Microsoft has announced six security agents that are launching at the end of April. I'm gonna talk about 'em in two buckets. So the first bucket is, uh, automatic response.
These are incident or alert triage in the defender and purview products. Uh, let me talk about defender first. So what it does is, uh, the problem space is that many organizations have done, uh, phishing awareness training and they teach people report phishing, click the button, just report it.
If something looks suspicious, report it. Well people do. And the bigger your organization, the more false positives you get from your users.
I've heard customers tell me upwards of 95% of user submitted phishing is a false positive. It's a marketing campaign, it's a spam email, it's not a security threat though. And yet each one of those still takes 30 minutes for human analysts to review.
Similarly, within the data security space, you get a lot of alerts that effectively say, go read this document. Lemme give you a concrete example. A friend of mine recently changed, uh, companies and when he gave his notice, um, the, the privacy system that they were using started tracking, uh, his computer.
He plugged in a USB stick when he plugged in the USB stick, every single document he opened after that for the two, his last two weeks ended up triggering an alert. His manager every single day got a spreadsheet of a hundred items he had to go triage. This is a great use case for ai.
AI is great at reading documents and telling them what they are. So it's very effective at reading the alert and going, well, was this document a a personal doc? Maybe it's a barbecue chi chicken recipe that you accidentally tagged as highly sensitive.
Um, it can distinguish between that and a document that's actually truly business context, uh, business, uh, information and distinguish which ones are more sensitive in which ones are not. Uh, purview's got another one related to, uh, insider risk, very similar under the hood and what they do. Some of that, it, it sounds like you're using AI to solve a problem that shouldn't exist, does it?
Are you, are you using AI to identify the root cause of why like false positive is in alerting, is is not a new problem. Um, and that in fact, it's one of the reasons why people tend to turn off security systems because they over alert. So does the AI help you identify why you are getting the false positives and tune the alerting system so that you are getting fewer false positives and more true positives?
It's a good question. Let's go back to the defender one and I'll show you, this is the one that I will be showing you all in more detail here in a minute, so you'll be able to see it. It's such a fascinating use case because while I say it's a false positive, it's actually a human generated false positive, not a system generated false positive.
So it does not mean that you're phishing detection missed a phish. What it means is the people at the enterprise believe that something is phishing when it is not, and they are actually the ones submitting it. And because it's human initiated, we're actually having human security teams go investigate each and every one of them today.
So this agent is targeted at reducing the workload so that the human analyst focuses on the, the 5% or fewer that are actually phishing attacks against the organization. It's not currently, uh, designed to go, um, back and tell people, oh, this was actually not a phishing email. Uh, that's potentially something we could look at in the future.
Uh, but at the moment the, the goal has been to reduce the toil within security teams, uh, that come from the, the over false positives that come out of, uh, in this case human humans reporting. And on the, the data loss side, these are, I I hesitate to call them false positives because they, again, they match the business rules that the organization set up. In the case of my friend who changed jobs, the organization did want review of every document that he looked at once he'd stuck a USB stick into his monitor.
In this case, we're just having the AI do that review rather than have a human having to do the review. So the idea is to meet the customer where their policies are, where their current state of their workforce is on all of these, and to reduce the, the human workload wherever we can within that, within that set. Let me also share some details of the, the latter three here I see as more proactive.
So, um, we've got the intune and the threat intel, uh, agents kind of like two sides of the same coin. Both of them start with intelligence, so either vulnerability intelligence or threat intelligence. And what they look at is where do you have risk and attempt to remediate.
So the Intune alert is designed to analyze vulnerability intelligence, figure out which software manufacturers have patches, figure out which devices in your organization both have the vulnerability and could have the patch applied to them. And then it builds a patching group. The threat intelligence, uh, agent is the cloud side of this.
So looking at threat, uh, threat intelligence, vulnerability intelligence and looking where in your external attack surface, you have infrastructure that is either, um, or at the highest risk of being targeted by threat actors. And so it makes it simpler to simpler, uh, have it teams and product teams go make the modifications to that software and configuration. And Nick, what's that based on?
Uh, great question. So the Intune one is primarily based upon open source vulnerability intelligence. The one that we're calling the threat intel briefing agent is primarily based upon Microsoft's threat intelligence.
This is one of our, we have a large threat intelligence team and we augment their work with, uh, open source intelligence analysis and those feed together into the incoming feed that we use to do the, um, the analysis of your attack surface and the, the TTPs. When you say infrastructure, what all is encompassed in that? The Microsoft infrastructure, Azure intra, what's in There?
Great question. So this uses a defender, external attack surface management, which is one of our products designed primarily for cloud, but it also covers on-prem and it's a multi-cloud system. So anything that the hacker can see from the internet, it is able to see, and it's part of our broader posture management approach.
So it's not specific to Microsoft, it's specific to things that are visible on the internet and attachable to your organization. I think there was another question. Yeah, I had one Nick.
So I just want, I was just curious and I think this is pretty phenomenal. So when you're doing that with both the Intune threat intel and the ID piece, now obviously you can go ahead, it'll alert you based on the automation portion of it, but outside of like scheduling will it, will the automation component, I guess, how do I word this, will over time, will it build out its own automation schedules? Like it kind of, if it's sees repetition, like, okay, we're seeing this many alerts where, you know, you set up a patch schedule every, you know, week or every two weeks at Tuesday at five 30, we'll eventually it start, can it build out its own scheduling for remediating those patches?
Or is that still gonna require some manual intervention? It's a really interesting idea today. It does not do that today.
Each one thing you'll notice about each of these agents, they're focused on a single task that is intentional and it means that, um, or we did that because it allows us to control the quality and to ensure that the agent is actually meeting the business need in question rather than targeting it very widely and having less control over its ability to do the job function. Um, effectively we may absolutely look into what you're describing in the future. It's a very interesting idea to have it learned from the environment like that.
Yeah, I wanna share the intro agent with you all because in many ways it's my favorite. Um, I I intentionally buried the lead here, the, in the conditional access agent here, what it does is it looks at conditional access policies and changes in your user environment. So people who've joined the organization, people who've changed roles, people who who've had permissions, um, go up or down.
Uh, and it makes sure that those are working together properly. This is a task that most, uh, security teams do. Uh, maybe monthly.
I've had some customers tell me it's done quarterly. And the problem with doing that task once a month or once a quarter, is that you have ongoing risk for that window of time, pretty much that you have users with improper access. So this is a case where the agent, by being able to run continuously, it shrinks the risk window from 30, 60 to 90 days all the way down to almost zero.
And it comes back to the, what I was talking about earlier, the agent's ability to do cognitive tasks is why this works. Human teams that do this say it takes days, it often takes days to do the policy analysis and the user analysis necessary to do this. So this is a case where the fact that the AI is able to reason over the information and the data and is targeted at this particular task means that it's able to shrink this down to a short enough time window and do it without direct human intervention.
Through the upfront work of gathering data and analyzing data that the organization's actually able to, to, like I said, almost effectively collapse the risk window down to near zero, uh, which I just love because it's, it's getting out of having to be in response mode, uh, and actually closing the, the problem that leads to a data breach in this case before the breach happens, or before it could even possibly happen. So I was intrigued by your mention of on-prem. So what's the mechanism?
These are cloud things, so what's the mechanism for that to work on-prem? Yeah, that's great. So security copilot doesn't directly have access to data.
It accesses data through other security products. So it uses, uh, products like Sentinel and Defender to get access to, uh, device telemetry. It uses Intune to access, uh, device configuration.
It uses the defender portfolio. So Defender, EASM has access to, um, cloud and on-prem, uh, server information. So that's how it's doing it, it's doing it through other security products.
It does not need to be independently granted access to, uh, your data center in order to get this sort of telemetry. And I'll show you actually how in, um, just a minute, I'll show you all how one configures the identity and the roles that make this possible. So I mentioned a minute ago the user reported phishing agent kind of at the high level, and we talked about how people just report lots of phishing when they're reporting marketing emails or emails that they've just mistaken high volume, like I said, meantime to triage across, across our customers is close to 30 minutes per submission.
Now let's actually take a look at how the agent that is designed to address this happens. So when I log in and I, in this case I'm an administrator, I see that I can install the SOC agent. So I, I see the setup button like an app on my phone.
The agent tells me what it does, what it's designed to do, what permissions it's going to need to be effective, and which plugins in security copilot terms it needs effectively, which software licenses, uh, you have to have as a customer for this agent to be effective. And again, as the administrator of the organization, I select the identity. Now this is a service principle, and so I am able to, um, I'm able to audit it, I'm able to govern it, I'm able to set rules and permissions on it, and the agent will always run as that identity.
Um, and the agent also is gonna be running within the, um, like as an administrator, I'm able to set the, the robust access controls within defender for that specific agent. So the agents that Microsoft is releasing are specific to the products that they integrate into and actually, uh, enforce the permit and, and those, so the permissions are enforced to those products. So once my agent is installed and running, I will start to see incidents be triaged autonomously by the agency.
You'll see here there's an agent tag that indicates that an agent operated on this ticket. And that way as a human, I'm still in control. If the agent resolved the ticket, I can unsolved the ticket.
If the agent, uh, made a decision, I'm gonna be able to see that decision. And that starts with what I see here in the queue. On the right hand side here, we see that the agent gives me a natural language explanation of what it did At a very high level, this is a meant to be a really simple explanation, even though the steps it took underneath the hood are far from simple.
If I have context that the agent didn't, I can change it, change what happened, and also teach the agent. So let me step back and give you all a concrete example. So, uh, suppose I had users reporting an email that was sent to the entire company.
The email contained a link that goes outside the company, it's flagged as urgent, and the text is fairly generic and, and pushes the user to click the button. The agent pretty reasonably would say, yes, this is something that requires a human to review, looks like phishing from a lot of its characteristics, but you know, the link itself looks benign, but I'm still gonna have a human review. What I as a human have context of the agent doesn't is that that's my HR training vendor, the, the vendor that my company to do its annual compliance trainings.
Well, in that case, I know that it's reasonable for this vendor to be sending this email. So I can actually enter in a comment here that says, this organization is who I use for all of my HR training. Stop.
Don't flag it as phishing going forward. And then once I save that, that's learning that the agent uses from that point onward to adjust what it does. So it recognizes going forward that emails from that domain are not phishing attacks and simply goes ahead and resolves them autonomously.
And so that's active learning that's happening within the system. It's specific to you the customer. It's not part of a general model, it's part of your specific agent and is applied to only your interactions going forward.
So Nick, the the way that, like, like I've seen that kind of feedback system before in just linear regression where you, you change the model weights in a fairly simple like linear algebra problem. Yeah. Is, is the training that kind of training, or are you, is it using the text that you've, you've got into a large language model to update the weights of the Model?
Great question. It is fundamentally different from what you've seen in the past. It's not like a linear regression, um, because that has deep limitations.
This is actually using it as natural language inputs for the AI as it goes forward. So it's very much in the same way that if you had a junior member of the team, you would teach them, yay, that's our training vendor. Don't flag that.
And it's stored in memory. Same thing here. We have a memory for the agent, and when you provide feedback, the feedback goes into the memory.
As an organization, I control the memory. So it's not the person's memory, it's not, it's the agent's memory controlled by the organization. So the organization admin decides who is allowed to train the agent.
Uh, the organization's admin can edit the feedback, it can delete feedback so that there's a way to tune it and control what's actually happening. But yes, it's, it is fundamentally an a, uh, LLM based approach that we're using. And I would imagine like, like pretty much like most solutions, there's a way in addition to, you know, machine learning, there's specific, uh, like controls or features and things you can implement, like certain tags to watch for, like how to I guess, modify what it is that it's actually scanning.
I always go back to like the old that dragon dictation software, like way back, you know, like you would speak to texts and I just remember that from like when I was working at retail. But the way you could go ahead and you'd speak, but it would self learn over time, like that first couple sentences, it would con ble it and mess it all up. But then over the course of, you know, a few days, weeks, months, obviously it gets better and it, it, it self learns.
So I mean, this is obviously at a much greater scale, but there's ways to modify those tags. Course in addition to the machine learning, correct? That's right.
It's meant to be something that the organization controls. They can modify it. Um, and, and that it does start to learn immediately and improve its outcomes for you.
Uh, are you able to go into the model and have it actually explain the reasoning so that you can see, uh, for example, once you've done some of this training, um, that you can see, well actually we're still, even though we're trying to teach the agent to do it a slightly different way, it's just not quite responding. Can we di essentially debug how it's working to understand why is this actually not going wrong? How can, how can we tune it or how can we, what, what are the signals that it's receiving that we need to adjust?
Um, how can we adjust the model so that it's actually working the way we expect? Yeah, wonderful question. Perfect segue into my next slide.
So, um, how do you as a, as a user understand what this agent did? Like I said, it's not simple. Under the hood, there are many, many steps going on in the reasoning.
And so what we do is we expose the inner workings of the agent through the, the, the product portal. So again, this is still defend our, in the activity tabs, a new, new piece of new section that we're adding. And what it shows is all of the steps the agent took with both human readable, uh, simple explanations of what that step was, as well as the details.
And, uh, this, I'm, I'm able to even drill into these and, and expand out the specific pieces. So you can use it for debugging. You can also use it if you just are curious or want to understand what's going on.
Um, I don't have a a picture of it here, but one of the, the cool things that we see it do is it will, it will have little tags here where there's a nice red suspicious tag saying, this was one of the pieces of information that made me think that, uh, this was suspicious. And so you can see how the agent came to its determination. You could see which data it looked at, you can see the specific reasoning that it did.
Nick, do you see this shifting left eventually into some of, like, Microsoft has a defender email security platform or something that does email security? Do you see this intelligence kind of shifting so that these just don't get delivered to the user? Uh, it's a great question.
So when you use Microsoft's solution for, for, um, email phishing here, the signal from the button is actually sent to a Microsoft research team that works on actively improving the detections upstream. Uh, so we already do have a shift left, uh, in place right now with that. Um, with this, your challenge and shift left is that many of the emails are coming from people or many of the submissions.
Most of the submissions are coming from people. And so this is actually a human problem, not a technology problem. It's that people are actually pretty bad at detecting phishing.
Um, and yet, um, organizations are asking people to report phishing. And maybe what we will see evolve is that organizations do less of that training over time. But as long as organizations are asking humans to report phish, this is likely to be a problem and this agent is likely to be very applicable for security teams.
And, and can I Ask you one last kind of under underlying question with these agents and that's the, are are they, 'cause you mentioned like several agents for, for different sort of components and features. Are they trained the same way and or with the same data, or could that be cross, cross-referenced? So if you, you know, if it flags something here and this agent, you reverse that, can that be looked up from another agent within your tenant or environment?
That's a great question. So today, no, today we have targeted each of these agents at one specific task so that we can get their quality and make sure ensure the very high quality on that task. What you're describing though is how we see the future that it is very likely that the, at some point in the next step or two agents will start working together to solve problems.
Or you might have one agent, uh, inspecting the results of another agent. You can imagine a quality checker agent coming back through and checking things. You can imagine a security agent checking things.
So as we move towards that world, yes, we will start to see, uh, techniques where we allow agents to share memory, or you might have also things like maybe your organization has common playbook or common set of business practices that need to apply to a wide variety of agents. So we're gonna have that sort of capability in the future. And, uh, the last piece here, uh, for me at least is, um, how do I manage these?
So enterprises want to understand the returns on their investment. And one of the big benefits of these agents being task focused is that the tasks that they do are all measurable and they're, uh, they have well understood KPIs by the teams that do these tasks today. And so the agents that we've built are tied to, um, a few KPIs, which we expose in dashboards.
And so we're allowing the customer to see what is the KPI, how is it actually doing on this task? Am I getting the, uh, return on my investment to enable these agents.