Mark Cusack on Kubernetes Transforming Scalable Analytics in Hybrid Cloud
Mark Cusack, CTO at Yellowbrick Data, explores how Kubernetes revolutionizes scalable analytics, offering enterprises agility, scalability and resilience in hybrid cloud strategies.
Transcript
This is Textron tv. Hello everyone. Welcome back here to techron tv.
I got another first time guest for you here on Techron tv. His name is Mark Cusack. Mark is the CTO of a company called Yellow Brick Data.
Um, and we're gonna find out about yellow brick data and, and talk a little bit about Kubernetes. But first, let's talk a little bit about Mark and welcome him. Hey Mark, welcome to Tech Drunk tv.
It's great to have you on here. Hello, Alan. Very nice to be here.
Thanks for the invite. Pleasure. Mark.
Um, I mentioned your CTO at Yellow Brick Datar, and we're gonna jump into kind of about yellow brick data in a moment. But before we did that, I wanted to kind of just get an idea of your, of your, uh, background and of your path to becoming, you know, sitting here today as the CTO. Yeah, and I guess Alan, my background is somewhat unusual as actually I started out in academia working as a physicist going way back into the dim distant nineties.
And so I ended up, um, doing an undergraduate, um, physics degree, um, postgrad PhD in, in the theoretical physics with a strong kind of bias towards, uh, distributed computing actually, which is kind of the link to the present day. But, so my career took me from academia into government research, into distributed simulation systems and then into the startup world where we spun some of the technology that we developed in, in government agencies out into, into my first startup, which is back in, uh, 2004 called Rain Store, which is all about, um, archiving massive amounts of data from relational data warehouse systems. And then that company got acquired by Teradata in 2014.
So I joined Teradata for a few years and ran a data, a warehousing product line there. And then I flipped out about four, four years ago or so to Yellow Brick to become CTO there. And that's where I find myself, What a great story, theoretical physicist.
You ever look back and think to yourself, man, if I had taken the fork of the road on the left versus the right, what would I be working on? Now I do think that, but I always keep my kind of eye as to what's happening in, in the world of physics. And, you know, there's a lot of crossovers when you look at this sort of advances in quantum computing and over, over the last, actually last six months or whatever.
And I look back at my work, so 20, 30 years ago, and there's a lot to kind of get excited about there. So, but what is interesting is when you start to apply ideas in a totally different area into what you're working on today and that kind of cross fertilization of, of ideas across different disciplines, I think is, uh, pretty interesting. Absolutely.
Very cool stuff. Mark. So were you at Yellow Brick data from like day one kind of thing?
Or, or you joined, they were already out and give us the background on Yellow Brick. Yeah, no, yellow Brick was actually well established when I joined. The company was founded in 2014, um, with the idea of how that they could apply the new emerging N-V-M-E-S-S-D technologies into high performance analytics and data warehousing.
And so the founders came from Flash Storage companies, um, as well as, um, established database companies. And so I, I joined kind of quite late on in the day why, why they've already been in market with a product for three years when I joined Give Us Expense then, I mean we, we get an idea of why, where it comes from and, and what they were doing. Mm-hmm.
Let's fast forward, you know, 'cause This's another one of those 10 year overnight sensations, right? You guys are added there over 10 years. Fast forward to today.
Tell us about Yellow Brick today. Yeah, Well I should just let your viewers know. I mean, yellow Brick is an SQL data platform.
We're essentially a relational database, but one that's really tuned for high performance descriptive analytics. And so you look at our customer base, it's financial institutions, insurers, telcos, government agencies, and they're typically want to modernize their existing old data warehouse infrastructure with yellow brick. And typically they've also got a lot of private data.
Some of the crown jewels of their enterprise data is stored in a data warehouse of course. And they want this data to, in some cases re be retained on premises and in others they want to deploy that data, run analytics on it in the public cloud and in some cases some hybrid combination of, of that setup or even multi-cloud as well. But ultimately they want to keep control of their data and that's what Yellow Brick enables them to do.
So we have customers from Redshift, from Teradata, from Oracle, from IBM, SQL Server migrating to us. And the end of the day, what do they get? They get, um, better outcomes from their data, happier users because the thing is a lot, lot faster significant cost savings and big returns on investment.
I just gotta ask 'cause I'm curious, has, has this whole AI thing had an effect on the yellow brick business? I don't think we would be a credible technical company if we didn't have an AI roadmap to our, to our investments. And, and so we do, and we've done a lot of work, particularly in a couple of areas where one of the most, well, about a year ago we added capabilities for Yellow Brick to operate as a vector store in similarity searches for retrieval, augmented generation applications.
So, you know, augmenting knowledge and injecting knowledge into your chat conversations. So we, we have that capability, but more recently we'd be looking at, um, converting natural language text questions into SQL and have that executed directly on yellow brick. Oh, that's nice.
So that's something we're actually be working on. There's gonna be a quite an exciting release of the product, uh, in a, in a couple of months. So there will actually be an SQL co-pilot, if you like, that allows you to Yeah.
Give the schemers Database that's envision you have, you can talk natural and it and it translates to sql. Correct. If you could do that without, you know, if you could do it, I mean it's always the same story with ai, right?
If you could do it without hallucinations or without mistakes, man, that would be so cool. But, you know, well, so I have a lot of friends who grew up being S-Q-L-D-B admins. Right?
And I don't know how how they would view that. Is that kinda taking their job or is that just gonna set the world on fire for 'em? Well, you know, I, there are, I think you have to split the use cases into two.
Here. You've got this kind of uh, uh, kind of holy grail of having any business analyst who's, who knows nothing about SQL being to ask any business question they like of their data and getting an answer back that they can make serious, you know, life altering business decisions on the back. Yeah.
That, that's one pile. The other is, hey, maybe this is a useful productivity tool to allow me to roughly generate a starting SQL point that I will then double check and verify and then add that to my, my uh, production. Not that, you know, that's not very different than you hear from testers.
From coders, right? Either. Right.
On one hand I could say, oh my goodness, the sky's falling, it's gonna replace me because when you know, we're gonna go from 27 million developers to 500 million developers, 'cause everyone could develop code, you just tell the AI to, or how am I gonna leverage this, the 10 x my my worth, the 10 x my productivity. Right. And I always want to be on the 10 x guy.
Yeah. And I think that is the pragmatic approach. I think, uh, we're, we're deluding ourselves if we think you can take the natural ambiguity of the English language or any other language and convert that unambiguously into a sequel statement that will run and give you the right result, we are far away from doing that.
But getting that 10 x performance improvement, that's what we're at yellow brick kind of thinking more about. Absolutely. Hey, before we jump into anything else, yellow Brick's, uh, website.
What, what's the website? com. Very simple.
Yep. Great. Alright, let's talk, if you don't mind, let's pivot a little bit to our topic of discussion today, which is how Kubernetes delivers scalable analytics in hybrid cloud situations.
And as we were talking off camera, I think in order to have this discussion, I think we have to first define hybrid cloud, right? When, when, when I, I, you know, cloud came on the scene, 2005, 2006 for me is what hit my radar. And um, you know, initially it was public versus private cloud, right?
That was the big thing. Where are you gonna keep your stuff? And then the answer became, well both.
And that was very easy to call the hybrid cloud. I've got some in the public cloud and some back in my data center and you know, calling that private cloud went outta style there for a while, right? It was just back in the data center.
But, um, it recently, it's, it's come back. But the other thing that really took me by surprise 'cause I didn't see it coming, was what we call multi-cloud where, you know, I got some stuff in a WSI got some stuff in Google, some stuff at Microsoft, maybe, maybe a little bit in the Oracle cloud. Yeah.
Maybe I got some stuff back in the data center too. And you know, I've got, and I I I put stuff into different clouds based upon what's the best tool for the job. Right?
Right. I don't know if we saw how big that was going to be in relation to the ne more narrow definition of hybrid cloud, which is just some form of public private. What's your take on that?
And, and, and if we could define that then let's talk about how we use Kubernetes to deliver the scalable analytics for that. Yeah. As, as far as yellow brick is concerned, we think of hybrid, uh, cloud and multi-cloud in facting.
Sometimes we com combine the whole concept into hybrid multi-cloud of the idea is I will place my data and my data warehousing workloads on the basis of data gravity, data sovereignty, um, cost security and other considerations. And so we are all about, it doesn't matter really where you deploy yellow brick and your data, we want to give you the same experience everywhere on any public cloud, uh, on a hybrid combination of those multi-cloud combination rather, but also running in your own data center as well. That's really what we're aimed at.
We want to give freedom of choice and flexibility about where you deploy those workloads. So for us, hybrid cloud is just picking the right tool for the job, as you say, and placing data at the right place at the right time. I agree with you.
I think that's a great way of looking at it. And what you call it, you called it the, the hybrid multi-cloud data sort of model. Yes.
It's a mouthful. Yeah. We need, we need, we need a, some initials there.
Well, I'll work on it anyway. Let, let's now turn it over to Kubernetes, mark and talk about how do we leverage that here for scalable analytics. Yeah.
Now, you know, we, when we were embarking on this hybrid multi-cloud journey, um, there were a number of considerations that we wanted to make sure that we tick the boxes on. One of which yellow brick software had to run anywhere in the data center and in the public cloud as well, and a hybrid combination of the two together. It needed to be elastic.
You know, we needed to be able to scale compute, uh, oh and storage independently of one another in all of these environments as well. All modern data warehousing solutions today provide that scalability that you, you scale your compute to fit the tasks at hand, for example. And then last but not least, it needed to be resilient as well.
We need this thing to, to be capable of supporting business critical operations with 24 7 availability. And so when you look at Kubernetes as an orchestration framework, it really does tick all of those boxes. It becomes this kind of cloud operating system for us where we can deploy the same containerized micro, uh, services architecture that yellow brick has in any one of these deployment options and have the same experience here as well.
And you know, what has been very interesting is there was the, a big lift to kind of get yellow brick in our first cloud deployment running in the elastic Kubernetes service EKS in AWS, but then the barrier to migrating it to a KS and Azure and then GKE got lower and lower and lower. Yeah. And we've just done our most recent call OnPrem to, um, red Hat OpenShift as well.
Yeah. And so now we have Kubernetes coverage wherever most enterprises would, uh, would care for it. Absolutely.
You know, I I, I made, uh, a reference to it earlier in, we are seeing more and more what I used to call private cloud back in the data standards, right. Whether it's Red Hat OpenShift, which is a, a dominant one. And uh, what what's the other big private cloud open source, uh, was like a whole consortium of people.
Rackspace was behind it, right? I mean this rancher and Tanzi and other kind of, well, Rancher is now part of, uh, of uh, of, uh, not of Tu Seuss. That's right.
Rancher is Seuss Ger of course is is Broad cob. Yeah. Right.
Yeah. No, no. But the, the private cloud stack not open.
Was it open cloud? I think it might have been open. I think open, open stack was The open stack.
That's it. Yeah. Yep.
You know, that kind of has been rejuvenated lately too. 'cause that had gone and a dormant for a while. So we we're definitely seeing a lot of that.
Um, but Mark talking about scalable analytics here, you know, it, it begs the, the real point, which is, look, our data warehousing is, you know, it, it's, it's consuming space, whether it's public, private, a little of this, a little of that, a lot of this and a lot of that. It, it, it, it just seems like there's never enough and, and we're, you know, we battle is it cheaper to do it here versus there? And what makes more sense?
And I need to segregate data, you know, and, and maybe store it based upon its particular value. Um, how does Kubernetes help us maybe with some of those kinds of analytics and those kinds of data points that we need to make those important decisions? Well, I, I think actually to some extent, Alan, the, the, the decisions are somewhat orthogonal.
I mean, our customers make decisions on where they're going to place their data and workload on the basis of the business problem and use case at hand, right? So, um, Kubernetes, I think doesn't impede what we do. Um, now, I, I have to say though, uh, when, when we deploy Yellow Brick and Kubernetes at the moment, we pretty much totally, um, take over that Kubernetes deployment, that Kubernetes cluster, the only workloads that are running in the Kubernetes cluster that we are running in a yellow brick workloads.
And we do that for performance reasons as well. Um, we, we do a lot of work at the very lowest levels of the Linux operating system to bypass main memory, to have direct access to the NVME drives that we access the store and retrieve data from. Um, we, we introduce our own threading models within Linux, and we do a lot of low level work, which means that we want to, within a particular Kubernetes compute node, take over all the resources on that box.
And so we don't sort of allow, we put anti affinity rule rules in place in Kubernetes to stop other pods kind of coming into our sphere of influence. So I guess that's a long way of saying that we, we kind of isolate ourselves from other considerations and, and give it an environment that's completely dedicated to Yellow Brick to run on. Love it.
com, you mentioned the website. Any particular path they should follow on the website or for specifics on this? Yeah, there's, uh, there's a ton of reference information blogs, um, follow, follow us on LinkedIn as well to, to get more information from you.
There's academic papers out there, uh, that you can get out. org site, you'll find, uh, a paper that we published at their conference about a year ago that gives us the, the full kind of Kubernetes architecture breakdown as well. But tons of information on the website.
Excellent. Mark, thank you so much for coming up here on Text Trunk TV with us today. It's been a delight.
Please do. Come back, keep us posted. You know, this is, you know, it's all about the data.
Stupid, right? That, that's the lesson that we've learned over the years in, in Act here. And it seems like Yellow Brick's right in the middle of it.
So do keep us posted. Will do. Thanks Slan.
It's been a pleasure. All righty. Mark Sack, CTO Yellow Brick Data here on Tech Trunk tv.
We're gonna take a break. We'll be back. We've got more for you.