Using Streaming Analytics to Improve Security Detection – Ganesh Pai, CEO of Uptycs
Ganesh Pai, Founder and CEO of Uptycs and Mike discuss innovative ways to capture telemetry from endpoints all the way to cloud platforms and the role of streaming analytics in improving security detection.
Transcript
This is Textron TV. Hello everybody. This is Mike Rothman general manager of tech strong research Chief's strategy officer of tech strong group and welcome to another interview for Textron TV.
Today. We have Ganesh Pai from upticks and he is going to again give us a sense of kind of some of the evolution of security and detection around endpoints and security and their Innovative use of os query and really how you scale up detection, especially from an endpoint perspective as we know it is challenging to get on top of all of our devices and really just to ensure that they're protected from from all the threats that are out there. So Ganesh, how are you doing this morning you're doing okay.
I'm doing fantastic. I appreciate the opportunity to be here Mike and thank you for inviting me. Yeah, you best what are you introduce yourself a little bit and a little bit about upticks and and give us a sense of what the company's about and and what problems you're solving.
Absolutely, so a little bit about myself. Engineer by training technologists by vocation and entrepreneur by choice. I've been very fortunate to be in the company of right people at the right time this Venture uptex me and my co-founders we got going in the summer of 2016.
We were inspired by the notion of using observability techniques in the realm of cybersecurity based on our prior background the venture today as we speak is six and a half years old and what we do for our customer base today is a unique approach to getting visibility into their security challenges and security controls. They need to put in place, but in terms of measurable outcomes for our customers, what we do today is we provide risk reduction by looking into two important areas threats and vulnerabilities and we've fined unique model where we use Telemetry as a basis collect the data and then draw conclusions, which have arguably High Fidelity and efficacy as demonstrated by the growth of the Venture. I'll pause here but hopefully that sets the stage in terms of where we are today and what we've accomplished it does and a couple of the things you mentioned there Ganesh really kind of, you know resonated with me.
Right and and that's kind of the idea of both risk reduction and leveraging Telemetry in a more, you know, kind of applicable way right actionable way from from that perspective. I want to dig into that a little bit more. So what are the main data sources you guys are are Consulting and how you Gathering that data aggregating it.
What kind of analytics are you doing? Since you know again a lot of folks find that as one step remove from Voodoo, right? You know, you're kind of sitting there trying to figure out how do I make sense of all this data that I have in order to get to a conclusion right to get to a prioritize idea of what they can and should be remediating.
not great question before I get into what we do perhaps you might be interested because You know in other domains where there's lot more budgets and dollars at stake as in ad decision analytics and other realms. Streaming based analytics is the basis to be quick decisions. So that someone can decide who gets to see what ads we repurpose such high fidelity techniques where you can look at data and motion from various asset categories and draw conclusions in the realm of cybersecurities today if you look at The various attack surfaces one might encounter in a modern Cloud native organization you realize that the productivity and point is the laptop.
That's where arguably in many tech companies software the crown jewels of the organization are produced and again, they get deployed into some kind of a public Cloud as an AWS gcp but a developer undergoes the act of you know connecting with some kind of identity provider connecting with GitHub to pull this code down compiler and then push it into cloud with the crown, you know, the crown jewels the software running in the cloud generate the revenue or whatever for progression of that organization. Now, if you look at all the potential attack vectors, which are represented in this chain of the Arc of developer from his laptop to the cloud what we adaptex have pioneered is a unique approach to get the Telemetry at scale from these attack surfaces and stream it into a back end six. Years ago our fortunes aligned when we got started.
There was a open source instrumentation, which came out of Facebook called as always query what it allowed you to do is it allowed you to scrape the operating system endpoints to get information about static data over the passage of time. It could also capture behavioral data. And what you did with the data was analytics and conclusions and it allowed you to extract that at scale.
So we saw an opportunity to align with the vision that with so many attacks surfaces. How can you use streaming techniques and back in analytics which are very prevalent in the domain of AD decision CRM and other places business process optimization as an sap and Salesforce and others and we chose the structure Telemetry from these asset categories being sent to the cloud and do decision analytics there. Right?
So I'll pause you for a second. I know I said a whole lot if you'd love to pass that and I have specific questions, I can drill and further I do because you know, it strikes me and I'm very familiar with you know, kind of streaming analysis and and kind of the ability to you know, make decisions in I'll call it sudo real time because at the end of the day you always have some latency in in the process from that same but one of the things that we talk about especially when Monitoring Cloud environments that you seem to be mapping to, you know endpoints and in a variety of devices is this idea of fast path alerting so and that that and by that what I mean is you have alerts set up in your cloud provider for specific activities, right, you know use of root account use of you know, kind of an overprivileged storage account, but something that again you don't want it to take enough, you know the time to go through something like cloud trail or Azure log or you know, as you're logging have to go into a Sim do some analytics on that right that latency that you have in there that 15 to 20 minute delay May kill you right so streaming and doing that on the platform gives you the ability to get that notification or that Alert in front of somebody almost immediately. So it strikes me as you're doing a similar type of thing by leveraging osquery from a Telemetry Gathering and aggregation standpoint and and streaming approach to ensure.
Had again, you're kind of looking for certain conditions in almost real time that would warrant almost in you know, we're basically instant kind of either notification or possibly even automation on the back end of that. Yeah, first off you summarize it in a fantastic way. I should probably consider asking our head of marketing to see if you could just surmise so much from a few sound bites that I provided.
That's fantastic to your point. Absolutely true. What's really happening is that in the traditional realm you had a lot of intelligence either baked into the endpoint or at the source of telemetry and then you generated alerts which went into a traditional Sim and the correlation across the asset categories happened over there.
Now where we flipped it around is that the actual endpoint is doing virtually nothing. I mean it is doing something it is capturing the Telemetry in real time, but it's transmitting and it made available in the backend to your point in sudo real time. And what that allows us to do is to do things that cloud scale at the back end because each source of telemetry whether it's AWS service provider using the connector or using osquery as a sensor on an endpoint or a node of a kubernetes cluster.
What it does is it transmits the Telemetry to the back end and for us we have the ability to look at the Spectrum of telemetry in two different ways once while it's the arriving on The Wire we're able to extract signals and then store the data in a structured format so that we can use columnar compression and other big data techniques to do historical aggregation, but the signals which we extract from this Telemetry in while it is in streaming then allows us to use various graph correlation techniques to say that across asset categories or within an asset category are they graph based approaches which are leading indicators of malicious activity. This is work. Really well for us because if you were to look at it from the cons from the side of how you know attackers think tend to think they don't think about the assets and silos, right?
They're going to take whether it's the laptop is the weak link or your ci/cd pipeline is the weak link or your cloud service providers the weakling they're going to be continuously probing. Right? So while the attackers are potentially doing that they're not thinking in silos traditionally due to wide variety of reasons.
Defenders have not had the Good Fortune of not thinking in silos and going across that complete art from laptop to cloud and that's the kind of innovation that we've bought that that complete art from laptop to Cloud. You've got data available both in sudo real time as you said to draw conclusions as well as in a historic basis in correlate. Yeah now that that's that's interesting because that does bring up a pretty interesting organizational challenge, right, which is that every one of these teams has their own mechanisms to figure out you know, what's happening how they track It ultimately how they fix it their own kind of Operational and Remediation types of motions.
So how do you start breaking down, you know some of these barriers right? You mentioned observability up front in a corollary or trying to use some of the observability constructs within the context of cybersecurity and that's interesting to me, right because this siled thing really does impact, you know kind of the ability of folks to do that and it's getting harder right because clouds happening and a lot of business folks are starting to develop their own it oriented or technology driven stacks and applications and on you know, Cloud platforms. You have the ability to not necessarily deal with it directly.
So so you have a lot of folks that are going around those I call it business it right? Some folks call a shadow it I hate that term because you know, these folks are just trying to get their job done and and it getting in the way, you know, they'll just go around you if you're if you're getting in the way and not you know, providing productivity, but how does this approach really start? To you know in effect democratize the ability for a lot of different organizations to you know, really get their hands on some of the threat data and some of the threat perspectives so that they can start to fix some of these things themselves.
Uh, great question the last part of what you said is key and data is of the Forefront when you convert the security problem in to a data problem, which is to say that data is the biggest dimension on which you can tee off or draw conclusions. It creates opportunities for teams to come together. You can think of it in terms of couple of Dimensions one dimension within the security organization as much as their under one umbrella.
There is the prescriptive nature of security as in vulnerabilities misconfigurations auditing compliance things which are very hygiene Centric. They typically work of static information to say that you have password which is not been rotated or disk encryption is missing. It's it's almost prescriptive and cause I factual in nature you fix that then you're hygiene improves and then there is the part which is more behavioral in nature.
Someone due to like either Loss of credentials or other reasons has gotten into your machine as doing something nefarious, but it's not harder to identify because you look for Behavioral trade, right? So within the security organization, there is two pieces of data that you can get which is Information about static configuration and behavioral changes. Now, this has extremely interesting implication to the other dimension, which is what you bought up whether it's the devops and the security operations.
So the it and the security team. The reason I bring that up is if you were to take the same lens of data and say that this is configuration related data and this is behavioral data the behavioral data sometimes tends to be of value for the devops team to figure out. Is there a performance issue something acting malicious and doing crypto mining versus a piece of software doing something bad, you know, it's a legit piece of software.
It's misbehaving the constructs for observability a very similar. You want to look at State changes and raw conclusions same as the case with it as well as the compliance thing. If you look at it, right the compliance people are trying to figure out things to say that is this good practice.
That's where we start things. Into a common area, especially in the realm of what's called as zero trust networking. You want to make sure that the hygiene is good which makes the it guys happy and the security guys happy, right?
So as you start looking at it in the two dimensions of Prescriptive security and behavioral security and cross-organization. We're trying to find this this approach of looking at it by moving one step up and our team is come up with the whole notion of shift-up security. That's what we want to think about it like a way to like start looking at it from multiple Dimensions, but potentially at one layer above where traditionally Security in it and within security compliance and threat detection have been somewhat siled.
That that's interesting because you know, one of the pushbacks we get on shift left now is and that's obviously, you know, kind of moving security closer to the developer on that front is that you know, what we're kind of giving overworked developers, you know more to do and make it a more responsible for all sorts of stuff. So now who do we shift up to right? You know who ends up having to you know, kind of hold the hold that that bag, you know once we're able to do that.
Yeah, so that's a great question what the the thing that you don't want to do is get in the path of the developer the nice thing about sasp-based interactivity. And if you look at how a developer might do his work today, right a routine actors to assert your identity to provide us such as OCTA and then connect to some you know, source code provider such as GitHub and then pull a code down to your work. And once you're done you probably are using cicd to push your workload for validation and then you might push it into the cloud.
The really nice thing about this approach is that each of the service providers gives you three pieces of information and it does so asynchronously without having to do anything in the path of the developer while the developer is going through this act you can buy talking to the apis of the service providers. You can get information about the configuration topology. You can get information about audit Trail as in through the lens of the SAS provider or the cloud provider.
What changes did the developer do right and it also tell you where exactly did this come in from one is in flow logs. So in asynchronous more through API, you can actually observe so much and not be in the path of the developer and yet help them out to say that hey here's something which might be leading to like, you know bad security hygiene, perhaps we can collaborate and fix. Yeah now that that's interesting as we kind of wrap up Ganesh, you know.
Of the things again. It just kind of strikes me that having a fairly open and and widely deployed aggregation technology like osquery right combining that with with all of the data that you can get from a variety of different places, you know in that stack being able to aggregate that and doing you know modern data analytics on that and then combining that with you know, really streaming in real time to look for, you know certain deviations from what is either normal or you know, kind of would violate a configuration policy does start to get towards this idea of as you said right leveraging data in order to further our security outcomes, right or better or improve our security outcomes. And did I kind of summarize that right and I miss anything, you know be no, I think you summarize it really well and the end goal is when we help organization reduce risk because The big reason why they are investing in a product and Technology like ours but where we enable them to make that happen is it allows them to do, you know better decisions around their risk, it allows them to go across their asset categories and you know, go across heterogeneous infrastructure if you will and then the last part is, you know, this is somewhat measurable.
There's this whole notion of meantime to resolution which is been around for a long time, but nowadays there's a desire to have what's called as a mean time to know because if you want to look at your misconfigurations of S3 bucket or something, you might want to quickly have a graphical view to understand. What could be out here as configured so that I can infer from that right? The ability to know allows you to make informed decision.
It's a little bit more intangible as in a resolution. It will have a resolution but knowing allows you to like take actions and remediate things before things go really wrong. Now, that's great.
That's great. So Dinesh tell us how to get in touch with your company and other resources that they may want to check out after after hearing this interview. Yes.
com. We've got lots of resources out there. Feel free to reach out to us over LinkedIn or to reach out to us over Twitter.
You know, we've got a team who's very interested in hearing other people's perspective and that I would say is the fastest way to get hold of us. Of course. My email address is Jeep.
com should people want to pursue that back. But otherwise our website is the best starting point. That's fantastic.
So I want to thank Ganesh appreciate your time today appearing with us on on Tech strong TV giving us some some perspective on again how we can start to get our arms around, you know, more structured and and useful collection. Of data, you know really from endpoint all the way through to kind of the the cloud platforms where a lot of these applications are ultimately being hosted and with that. This is Mike Rothman and we will hand off to another great session on Textron TV.
Mike thank you. I appreciate the opportunity to be here.