Kevin Kamel on Selector’s Spring 2024 Release
Kevin Kamel, VP of product management at Selector, talks about innovations in the Spring 2024 release, which includes a GenAI-powered network-oriented LLM, native full-stack monitoring capabilities, and event correlation with root cause analysis.
Transcript
This is Techstrong tv. Hey everyone, welcome back here to Techstrong tv. I, you know, I get excited to introduce new companies to our audience.
It's it, and it, especially given the macro conditions in tech today, it's good to hear about, you know, new growth, new companies coming on and, and solving old problems, new problems, and all kinds of issues. In that vein, let me introduce you to Kevin Camel. Kevin is VP of Product Management at a company called Selector.
Uh, Kevin, welcome to Tech Drunk tv. It's great to have you on here. Oh, thanks for having me.
Pleasure. So, Kevin, before we jump into selector, and we're gonna spend, you know, the bulk of our time on that, I, I always like to give people a sense of who they're talking to, right? So you're VP of product management, um, so obviously, you know, handling product and aspects, but you know, you weren't always VP of product management at Selector.
Tell us a little bit about your journey. Um, sure. Little bit of an unconventional one.
Uh, I actually got my start as a statistician at nasa. Okay. Yeah.
There was a, a, a field that's called the Orbital Debris Analysis, and I, I sort of got started there. Um, and then from that, that was sort of, uh, in Washington dc I transitioned into an incubator in around 2001. Um, DC itself doesn't have a big startup scene, so that was sort of like the best way to get exposure was through an incubator itself.
Um, that was hosted by University of Maryland. And then, um, yeah, then I ended up working at an email marketing company, which was sort of, uh, my journey to get to selector. Um, email marketing is a very statistics based business.
My background, um, sort of lent itself to that. And what we did is that we would come in and we would sort of send newsletters for people and, you know, figure out who opened, who read them. But you're doing it at a very large scale and people are very interested in those numbers.
Like, who clicked? What are they clicking on? Why are they clicking on these things?
Um, and we started to get into monitoring, uh, capabilities at that point in time, uh, with that company. And the idea was, is like, how long does it take us to send out these newsletters to, you know, a million people? Um, people wanted to be delivered at exactly 9:00 AM So we had to start monitoring both the system infrastructure, the application, and then also the network in order to understand, you know, do we have enough capacity here in order to shove all of this email out over the wire, over the internet infrastructure.
That was, you know, very modest in like 2003, 2004 timeframe. Um, so yeah, that was, that was really my start. Um, you know, I had the good fortune of being at, uh, a company called Mailer Mailer, which exited in 2017.
And then I went to go work for a, uh, a different monitoring and observability focused company. I ran product there for, uh, about five years. And then I got introduced to the folks at Selector and, and here I am.
Very cool. They weren't calling it observability in 2017 though, right? That had to be later in the five years.
Um, They, they were, um, I'm not sure when that term exactly got coined. I'd say maybe 2013 and 2014 around there. But yeah, now interesting.
Now monitoring sort of legacy observability is, is the new hat. Yep, Yep. For sure.
You know what we're heading off, uh, actually around the time people see this, I'll be there already. I'm heading off to Paris this weekend for CubeCon, and you know, CubeCon is becoming a big observability, uh, kind of concentration there with that all cloud native thing. Anyway.
Excellent. And, and quite a journey, right? nasa, the best and the brightest still come from nasa.
I have some good, actually a good friend out of Colorado, Todd Vernon, who is a ex NASAs guy, and then became a kind of serial entrepreneur on a bunch of, uh, decent companies. Um, but let's talk about selector. So you were kind of recruited into selector, what, what's, what, give us, give us the selector story.
Oh, absolutely. Um, so, you know, what is, what is selector just at a very high level? So, uh, what selector does we offer observability and AIOps solutions, sort of hybrid solutions, um, at, at a base level, selector is a platform to help you transform your operational data to actionable insights.
Um, we come into your environment, we collect the telemetry that's there, and then we apply AI and ML in order to, uh, basically, uh, capture the leading signals within your network infrastructure and applications and point you towards problems that maybe are emerging or potentially have already happened within the environment. Okay. If you don't mind, we're gonna peel the onion back a little bit here.
Oh, absolutely. So what, what are you using to collect this data? Is it, is it like based on open telemetry, Prometheus, something like that, or Grafana?
Uh, that's a great question. So I, we have this phrase at the company, which is any data from anywhere. Um, and that's true.
Uh, as part of the platform, we have this sort of like no code, low code type approach to collecting data, um, you know, for the, uh, you know, more computer science folks out there. Uh, it's a declarative ETL. You can basically define without writing any code what the inputs and outputs are for a given, uh, telemetry interface, and then pull the data from that.
And the reason that that's important is, is for what you just described, there's a large number of protocols and technologies that are out there that people wanna collect data from. Um, and it's not uncommon to run into proprietary endpoints where there are no known interfaces or, or libraries through which to co and collect that data. So it's super important for us to be able to sort of very quickly stamp out integrations for, for pretty much any data source.
And very important for a large enterprise where, uh, especially on the, uh, sort of network and infrastructure side, they've gone ahead and built proprietary tools inside of these companies and they're still there after a decade. And it's important for us to be able to come in and collect data from those tools and then deliver insights on top of that. Excellent, excellent.
Um, you know, the, the, obviously the observability space has really blown up over the last couple years, right? And people using AI ml, it's almost mandatory. 'cause how else can you get your head around that larger data set?
Right. It's, it's what it is, but what we've seen most recently is with people, I, I don't even know how to say this, right, but too much data, right? They, they can't, it's not that they can't analyze it.
'cause unlike, let's say 10 years ago, 20 years ago, AI and ML kind of tools do let us analyze these hyperscale data sets. The problem is the cost of, of maintaining those data sets and, and the computational power. So between storage and compute, it, it's almost prohibitive.
Like, you know, you used to see, I want to, so I grew up in an era, right? Where we, we can measure everything and we, we can measure everything and we do measure everything. We're almost in an era now where yeah, we can measure everything, but do we really want to?
Yeah. Um, it's an interesting point. I mean, a, as a former practitioner, I mean, I remember when, uh, you had your entire company was built on just a few systems, and it was very possible for a small group of people to be able to come in and sort of understand what was happening with those systems at any given point in time.
And so long as you didn't make any late in the day changes, you were guaranteed to make it through the night. Like, you know, nothing. Mm-Hmm.
Nothing was sideways on there. Um, and those days like that, that world that I just described, that that's long gone at this point. Um, people don't even, younger folks especially have no idea No.
What that needs Yeah. That, that's just gone. Yeah.
Um, I would make the case. Um, so the storage costs absolutely are, are a major problem. And I think that that's one of the reasons why you see a lot of enterprise organizations now sort of reassessing whether or not these types of observability platforms, legacy ones, um, are intended to be historian products where you're keeping data for literally years on end or whether or not they're operational in nature.
I would make the case that a selector is more of an operational focused platform. The idea behind selector is that we're gonna collect this telemetry, we're gonna analyze the telemetry, and we're gonna send, uh, what we call a smart alert to your team. And the smart alert will basically say, this is the root cause of this particular problem that you're dealing with right now, and here's all of the events that happened as a result of that particular root cause.
Um, and I think from that perspective, that's very different than how a lot of other monitoring and observability platforms sort of operate. Um, I, I won't name names here, but there's, there's plenty of offerings that are out in the market where what they're really doing is collecting data and warehousing it. Um, maybe giving you a dashboard to go look at the historical data.
They're not providing like, pinpoint type guidance to a team as to like, I think you should go look over here. Um, that's not the nature of, of their product. Um, yeah.
So while we have some of this functionality that I just described in terms of like collecting all of this data, we're, we're far more focused on sort of providing guidance to the team with actionable alerts as opposed to saying, can you go back in time five years and show me the number of people that happened to be using this API at this minute? Uh, I have to question the value of that. Yep.
So, I mean, I, I would charact characterize that as it's, it's not really a forensics to it forensics tool. It's, it's what's going on right now, right? And giving you sort of just in time or, you know, live sort of, uh, it's the audience advice Audience.
So the audience for us is, is primarily operations teams. Mm-Hmm. Those are the people who will be using the platform.
Um, it's not to say you can't build other use cases on top of it, you can, but primarily we're operations focused. Where I think for sort of legacy monitoring and observability tooling, I, I'm not sure who the audience is. Like I do hear about we can keep the data for three or five years or seven years kind of a thing.
Where do those requirements come from? Um, uh, I'm, I'm not sure. Well, if they're security related, they could be coming from things, uh, you know, the, the, uh, C guidelines or I'm sure the EU probably has some guidelines.
I've always also been told it's the three to five years, but that's where forensics really security. It's a, I'm not Different space. Yeah, Yeah.
I'm not like, you know, downgrading it or anything, but it, that's a, that's it's, it's, it's a corner case in many ways. Not a small corner, a big corner, but a corner of the case, nevertheless. Um, let, let's talk a little, you know, more about, so selectors going after ops folks today.
Ops folks come in a lot of shapes and sizes and colors and names, right? We have platform engineers and SREs, and we still have CIS admins and, you know, DevOps folks who, who specifically within ops? Kevin is the target here.
Um, so the foundation of the company is largely network oriented. Um, mm-Hmm. And so I would say that that's where we have started, uh, building a customer base has really been on network driven engagements.
Not to say that we can't do other things, but those are sort of the people that we happen to know within the service provider and large enterprise space. So that's where we've, uh, focused our efforts initially. Um, but you know, it, it's kind of interesting.
Um, you know, we were talking about some legacy monitoring and observability, uh, platforms a moment ago. There's this sort of idea out there of full stack observability, and I'm, I'm sure you've heard that term, that's a very popular term that's out there. But, you know, when we use this term, uh, full stack, you think to yourself, well, that's gonna be everything, right?
It's, it's full, it's all of the parts, but there's one part of full stack observability that's not really accounted for, and that's the network, the primary way that companies go about delivering their service via the internet is ignored as part of full stack observability. And for us, we see that as an opportunity. Um, what we want to do is basically say we're going to do, you know, I'll use this term like true full stack, um, you know, we wanna start with the network and make sure that everything that's, uh, your services are riding on top of is working correctly.
Then go up into the infrastructure, into cloud, into applications, and give you this sort of like high level comprehensive view as to here's what your infrastructure looks like, here's the architecture of your particular stack, and here's where the problems are. And within our interface, it basically allows you to sort of like drill down to that particular domain in order to see where challenges or problems may be. Um, and then again, with that actionable alerting that we described, like point operators towards that being a specific issue.
So to that end, uh, I would say it's, it's broadly operations teams. Um, but I would say we wanna start with network operations teams and then sort of branch out from there, um, up towards application operations teams. Teams.
Got it. Very cool. Um, and how is Selector offered?
Is it a SaaS tool? Is it kind of traditional software you run on your own infrastructure? Um, however, which way people want it.
Um, so it's a fully Kubernetes based stack, um, mm-Hmm. And that gives us a lot of flexibility and deployment. So, uh, today, uh, we have quite a few folks that wanna deploy that on premises.
That's just a bias that's within large enterprise organizations. Um, but, uh, we are actually moving towards being, uh, made available within the Google Cloud marketplace in Q2 later this year. Um, and at that point, people will be able to come in and sort of click a button within the marketplace, spin up an instance of selector, and, uh, and they'll be able to start, you know, playing with it, putting their data in, um, and working towards use cases with it.
That's great. That's great. Um, I don't know how much we, you can dive into this, but how, how is it priced?
Um, yeah, I'd leave that, I'll leave that to the sales guys. Okay. Um, and, and who is the target?
Is this, you know, 'cause the, the problem with a lot of the observability space, as I mentioned before, you almost have to have that large data set to make, to make it go right? So is it named at large enterprises mid, you know, who, who's your target persona there? Um, I would say large enterprise to mid.
Um, that's, that's really the sweet spot for us right now. Um, it's not to say that we can't handle like small business type use cases, but it's just not the focus of our business at this space. I know you guys recently did some upgrading on features and stuff.
We did. We only Have a few moments here, but can you tell us about it? Um, absolutely.
So there's, you know, uh, three or four different features, um, that have been updated as part of the latest release. Um, the first one here is really around, um, integrating generative AI into our stack, which is a pretty novel and new feature. Um, the idea here is that we have an LLM, it's deployed as part of the application and it allows customers to come in and use plain language in order to ask questions of the system and their telemetry.
Um, and the reason that that's important is, is that most platforms that are out there today require you to basically use a domain specific language in order to sort of query the system and figure it out. And there's friction there, right? Like, you know, if I adopt a new platform, now I have to go figure out this new language in order to get the answers outta there.
So we wanna solve for that. Um, but, uh, we now have, uh, even greater capabilities than this. We have integrated a functionality that's called a retrieval augmented generation, um, rag.
Um, and the idea behind this is that we can leverage all sorts of different pieces of information in order to construct a response back to users. So this is far beyond, you know, what is happening with my CPU at this moment, and it gives you back, you know, a 58. This is basically, um, I have a router that's in front of me and it has failed.
Can you give me guidance as to what I should do? And what it's able to do is basically look at all of the existing incidents that happened with that particular router, look at incidents that happened with other routers. It can connect the dots between, hey, this is a Juniper router and this is the series and this is the particular firmware version that's on there.
And I've just run a query for you on the Juniper knowledge base. And here's two relevant knowledge base articles that you should read. Uh, maybe you should update, upgrade your firmware as part of this.
So it becomes really interesting in terms of the capabilities. Go ahead. No, no, I was gonna say it sounds great.
Yeah. Um, it can be a real game changer for people. Um, you can also sort of integrate with runbook systems like Ansible and things like this.
You know, maybe, uh, maybe the system has actually seen how somebody's gone ahead and fixed an incident in the past and it can make a recommendation. This looks like something that we dealt with last week. Here's how we went about fixing it.
So, um, this sort of gen AI capability. Um, it's still early at this point, but there's all sorts of different use cases that our customers are exploring with this right now. And we'd love to have conversations with people about what we might be able to do for them using this.
Um, so that's, that's one aspect of it. The native ability for the platform in order to do monitoring and observability, uh, really is a very nice capability for us. It means that we no longer need to live on top of existing tools that are within somebody's infrastructure that we can basically come in and operate, uh, independently.
We can collect telemetry from your devices, we can collect telemetry from your compute, the cloud applications, all of it. Um, but if you want us to live on top of your existing tooling, we're happy to continue to do that as well. What we want to do is just make sure that customers have options so that they can, uh, address their requirements as, as best as they can.
Um, root cause is another new capability. We have a type of causal ML that's integrated with the stack. It uses, uh, what, what the data scientists call associative models.
And the associative models basically allow, uh, the system to order the events that led up to an incident and indicate that this is the precipitating event, the very first, uh, issue that occurred, the root cause of the incident. Um, and then all of the fallout from that is sort of ordered out afterwards. And that allows us to do that smart alerting that I described to you earlier.
And then the final bit is that, um, all of these capabilities sort of add up to, uh, a larger capability, which is that, um, we have the ability to model the entirety of somebody's infrastructure. So for example, from a network perspective, we can create a digital twin of the network. We're working with some of the world's largest, uh, service providers in the world, and some of them have enormous networks with potentially hundreds of millions, billions of paths that these, uh, these networks are, are sort of, uh, involving.
And the question becomes is when something goes wrong, like why did it go wrong? And there's just quite a few variables that a person would need to take place in order to account for that. At selector, what we've done is that we're able to take the entirety of that network and represent that internally in memory, sort of an in-memory model of a massive network that's out there.
And this allows us to do a couple things. One, you can go back in time and you can say, well, what was happening 25 minutes ago across this particular point of the network? And we can visualize that and we can tell you about it.
And then also we can use that for capacity planning purposes. We can say, well, it looks like you've got excess capacity over here. We don't see you actually needing this in the near term, but you have too little capacity on this side of your network and potentially advise that you can re-provision equipment from A to B.
That's great. Very good. Kevin, you know what I realized we didn't even tell people the website.
ai, That's S-E-L-E-C-T-O-R. Do ai, that's selector do ai. All right.
Hey Kevin, thanks for coming, coming on and telling us about selector. I hope this won't be your last time on here with selector. Um, are you guys at CubeCon coming up, or no?
Uh, This, That's Paris a little bit of a, Yeah. Yeah. We're largely Maybe Salt Lake City in the fall.
Yep. We may, we may be at that one. We may see that.
Okay. ai, ai ops and observability. Check it out.
We're gonna take a break here on Text Drunk tv. We'll be back in a minute.