Transforming IT Operations with AI: Insights with Grokstream’s Casey Kindiger
Casey Kindiger, founder and CEO of Grokstream, LLC, discusses his journey from consulting to creating a company that addresses complex IT challenges. Grokstream, LLC utilizes AI and machine learning to enhance operations and lessen engineers’ workloads. The launch of Grokstream, LLC predictive IT operations marks a shift to proactive incident management. Casey highlights the importance of learning algorithms and the growing preference for interactive learning experiences.
Transcript
Hey everyone. Welcome back here to Tech Trunk tv. Our next guest on Tech Trunk TV today is Casey Kinder.
Casey is the founder and CEO of a company called GR Stream. Let's welcome him. Casey, you're on Tech Trunk tv.
Thanks for being here. Thanks For having me, Alan. So, as I mentioned, you are the founder and CEO of GR stream.
I always like to explore what drove you to this crazy level of craziness to go out and found a company and, and, you know, try to make a go of it here. And I don't have to tell you every day's a battle, right? You breathe life into this every day as a founder and especially as CEO Give people a sense of your background and, and how you wound up here.
Yeah, sounds good, Alan. So I think really there are three things that led me to, uh, where we are today. Um, number one is, uh, I'm a, I'm a math guy.
I was a, a math major in, uh, undergrad at U Chicago. Um, I like solving hard problems. It's sort of just what it's, it excites me.
It's what kind of keeps me going every day. Um, and the, the, the second thing I guess is I started my career at Deloitte Consulting, just kind of, uh, you know, as many, uh, quant focused people do in the consulting industry. And, but I got really fortunate and I, I landed in Silicon Valley on a project or on multiple projects when Silicon Valley was really just building out the, the, you know, sort of the first generation, the first wave of large scale infrastructure build outs.
And, and it, it's, it's sparked something. It made the complexity of it, the excitement of it, the energy of it. Um, you know, I was there for just a couple of years before I started my first company.
Um, you know, what it gave me was a, a grounding and a problem that I thought was really, really interesting and wasn't going to be easy to solve. And that is how do we manage the complexity that is inherent in the, this evolving it and network infrastructure that really powers our world today. Um, that problem, you know, excited me 20 years ago.
And, um, I, I'm even more excited today because I think we're we're closer than ever to really getting our arms around the solution. Absolutely. So this is one of these overnight sensations that's 20 years in the making, huh?
Well, I mean, yeah. I mean, you know, I, I, this is my third company, you know, I had two hypotheses prior to this on what it would take to solve the, you know, that foundation problem. And each iteration, sort of the problem domain has shifted and, and the complexity level has grown.
And with the, the advent of ai, I started, you know, down this path about eight years ago with a, a small research team with the real simple problem. Okay. Now, we've, we've attacked, we've attacked the problem in a couple different ways.
Um, I won't go into the details of the prior hypotheses, but the, you know, um, you know, our, our fundamental goal starting at Rock Stream was we're gonna, we're gonna harness machine learning and AI to solve all of the problems that were intractable for us in our prior two, two runs at the problem. And, um, and that really was the goal. So I, you know, I had a small team of researchers started to identify what the, what the kind of underlying problems are and attacking it with, uh, with AI and machine learning algorithms.
Excellent. Casey, you know, you know the old saying you can't make wine before it's time. Right?
And, and, and I agree with you, right? So many problems. And I like you, I've al I've been in tech 30 plus years, infrastructure, uh, security mostly.
And yeah, I look back at problems that seemed insurmountable then that are kind of common, you know, table stakes today, but yet, you know, things keep getting more complex and new problems. And much like the book, the Goal, you ever read the goal, right? As soon as you solve one bottleneck, the next bottleneck presents itself, right?
And that's kind of where we are today. AI has, you know, a crazy potential to help us. It also has crazy potential to make yet more headaches for us.
Right. Um, talk a little about Gro you said you, you gave us the setup into GR stream, but as it existed today, give give our audience a sense of where it's at, what it's doing mission wise and stuff like that. Yeah.
And, and the mission is really simple for us. It's so without AI and machine learning, large, complex IT, and network infrastructures with cloud overlaid, um, in order to keep that up and running and optimal and delivering the services that continue to evolve on top of that, um, requires a large number of, of engineers doing operational tasks to keep the lights on, to keep the systems up and running and optimal, right? Um, our, our goal really is fundamentally to reduce the amount of effort required to do upkeep and respon, you know, respond to incidents and fix problems and allow more of that talent to do innovative tasks.
So what we've built is a, a cognitive AI architecture that consumes sensory data, basically telemetry data, everything that's happening in, in the environment, in the, you know, um, com complex web of networked data centers that drive an enterprise, um, set of services. And we take all that telemetry data, we process it through a series of algorithms, each of which a distinct learning algorithm from their own sort of research domains. Um, and at the end of that, we predict when an incident is going to occur.
We allow for, um, operations engineers to, to prevent any occurrence, any similar occurrence in the future. So they don't fix the same problem twice. They don't have to do the same maintenance task over and over again.
And, um, once the system has evolved with the team and they're sort of functioning well, um, it takes on more and more of that operational repetitive work for the support teams. So, DevOps teams do less ops, more dev, um, network teams, do more engineering, less operational support and so on. That's fundamentally what we do.
We process it through a series of algorithms. We enable automated action at the end of that that's intelligent. That is, you know, isn't a set of, of rules that have been built, uh, you know, a runbook that's been coded, but rather a, an intelligent response system based on learning algorithms.
Love it. Love it. You know, we didn't even mention a URL people who want to go explore this themselves.
Where, what's the best way? dot com or? com.
G-R-O-K-S-T-R-E-A-M. Just like it's spelled underneath your name here on the video on the lower third. Alright, Casey, let's shift gears.
You guys recently re, uh, announced a new release, grocks, excuse me, grok predictive IT operations, uh, provides IT operations service management teams with an AIOps platform that continuously learns and adapts self-healing. That's always been a kind of holy grail. Um, talk to us about gr predictive IT ops.
Yeah, so I'll, I'll just say that the first, the first rung of the ladder that we had to, we had to climb up was, um, making our system smart enough to respond quickly when a problem occurs, when something happens in the environment that requires action to be taken. So the kind of the first step is to understood, kind of sift through the noise that telemetry data, and figure out what is actionable, what's actually just the shift in the system redundancy exists in all of these, um, you know, large networks and, you know, bubble up what act what requires action and respond appropriately to that signal to, to fix something, to restart a service, to create a ticket and dispatch it to an engineer, whatever that is. So, still reactive, but just fast and dynamic.
That was kind of the first run of the ladder. And what we've been working on for the last several years on the research team is, okay, so we've, we've gotten, you know, the, the problem has always been, you know, in our space we wanna go from reactive to proactive. So step one was to solve the problem with ai, the, the reactive problem with AI to relieve some of that pressure from operations teams.
Um, and the second part of that problem is forget about reaction. We actually want to stop the problems before they occur. And, um, and, and we do that in several ways.
I would say, let's say three, three ways. One is we have a predictive algorithm that monitors the, the, uh, the occurrence of incident patterns in the environment and the operational data that kind of precedes those incident patterns. And it predicts when an incident's going to occur within, we can kind of fine tune the timeframe, but say 8, 12, 24, 36 hours.
Instead, there's a high likelihood that this particular incident is going to occur, this problem in the environment's going to occur. So let's go out and take action to prevent that. So, you know, rather than waiting until something breaks and dispatching an engineer to fix it.
Um, so that's number one. Number two is we can observe the, the pattern of occurrences. So the problems that have happened over a long period of time, we can surface the, this pattern in a, uh, an actionable list.
Think of it as an automation pipeline. Things that are sucking energy from our high value engineers, requiring them to do operational tasks over a period of time, list it out and, and identify what items on that list ought to be, uh, automated. What can we do to automate each of those, those items to prevent them from occurring in the future?
So that's the second poll on the tent. And the third poll is enabling our, um, language models to process the information, apply reasoning, and take autonomous action to prevent the issues from happening in the future. And, you know, those three polls in our, our preventative and predictive tent allow us to really kind of take it once and for all from being reactive to being fundamentally proactive and predictive.
Agree. com website and it's right there on front. It is.
And we have some great videos too. I mean, I, I think it's, it's tough for folks to, you know, read a whole lot of text content these days. So I would encourage you, and we're we're, That's why we're on here, my friend Deny that I didn't do this is a written q and a.
Yeah. And we're working, that's an understatement. Yeah.
We're working really hard to produce videos that not, don't just talk about what we do, but actually show you a lot interactive so you can click on things and see how our user experience matches with the kind of the value that we deliver. And that's the, be that, that's sort of the really important when it, when it comes to, uh, AI powered systems because a lot of, you know, a lot of folks will deliver products, and the AI is so deep under the covers that you don't really know what's, what's a rule, what's just a rule that's been coded versus what's a predictive model? What's a, and the reason that's important is because, you know, if it's a rule, it doesn't evolve.
And all, you know, every environment changes, every company changes rapidly. And if you don't have learning algorithms kind of producing the output, then you know, it's only gonna work for a short amount, amount of time. That really is so key to the whole look, even before ai, you know, it, the time crunch is continuously accelerating, it seems.
Right. And what worked two years ago certainly doesn't work today. It's ancient history.
Um, but you know, you mentioned what people like, we, we do the same thing. You know, we do do maybe 400 webinars a year, and we survey our audience, what do you like, what don't you like consistently? Their favorite type of, we call a webinar a learning experience.
Their favorite type of learning experience are hands-on demos. Not, not a PowerPoint with, with screenshots of a product, but actually logging in and following along how the product actually does what it's doing, what, what, how do you do it? And they love this.
And, and it's almost counterintuitive because it's very vendor specific. It's very product specific. But they'd rather do that than listen to a talking head with a bunch of slides.
Yeah. You know, narrating the slides. So I, I agree with you.
That is, that's what people like today, of course, you know, get it all done in seven and a half minutes, because that's kind of the attention span. But um, right. You gotta break it up in pieces.
Yeah. I mean, we're human, humans, humans, humans learn With our hands. We're tactile creatures we have to touch in order to really learn That.
That's, I think that's it right there, case. All righty. Hey man, thanks for coming up here on Text Drug tv.
This was great. I appreciate it. com.
Check 'em out. We're here on text from tv. We're gonna take a break.
We'll be right back.