Jonathan Symonds, MinIO | KubeCon + CloudNativeCon Europe 2023
MinIO’s CMO Jonathan Symonds discusses why object storage has been so successful in the Kubernetes ecosystem and why MinIO is the future of Kubernetes storage.
Transcript
This is texturung TV. Hello, and we're back at kubecon plus Cloud nativecon Europe and we're here with Jonathan Simons. And we're talking about storage in kubernetes.
Jonathan's would mid IO and they are kind of the experts in the space. So Jonathan welcome to the show. Thank you very much delighted to be here.
We were there with you guys last year in Valencia, and we're excited to be back. The debate around stateful versus stateless applications in kubernetes continues at this show. I'm still hearing people talk about it.
Yeah. So what is your sense of how mature is it for running stateful applications on kubernetes these days because there's no shortage of opinions on both side of the aisle. Yeah.
Yeah. So listen the way we see it. Is that as we mature the ability to run stateful applications.
Is there one of the things that we introduced in the past years DirecTV, which allows is a CSI driver that allows you to connect directly to the discs and that allows you to run staple applications inside the container storage being one of those the thing with kubernetes that I think is, you know, everybody understands but it's hard to basically engineer it into is the more you put into the Container the more valuable the orchestration can become and so by having more things in that container such as storage you have more flexibility have more resiliency and you have more basically optionality. And so that's one of the things that We think is important. We see folks who have small containers and we see people have really large containers kind of reminds me in the early days of VMS when we had monster VMS.
Is there a Best practice or a thought process and how those things should be structured? Yeah, so that's a great question and I would say that you know, there is very much of a workload. Looking at what your workload is in the attributes of that workload and how Dynamic that workload is and what the data requirements are gonna Define that for you if you're you know data sets are very large.
You do an AI ml training it may be that you need to have a larger container, right? You want to have more things in there it so it's gonna vary we don't have any solid guidance to say you should always be doing this or you shouldn't be above this size. We kind of take a more of a workload Centric approach when we talk to our customers.
Now as I understand you guys wrote your own CSI driver. Yeah. What was the reasoning behind that?
So the primary reasoning was that we wanted to get at the performance capabilities, right? So when you are going through other hops or or other layers, you don't get the full performance and so given that men IO is a high performance work Object Store to start with and we focus on those high performance high value use cases. We needed a CSI driver such as DirecTV that allowed us to get directly to the underlying disc now VMware gave us this, you know, when we did the the open shift implementation and so we basically saw the power of that being able to get directly to the discs for the best performance.
And so that's why we wrote DirecTV, you know prior to that. We were a big supporter of cozy. So the container object storage interface and we you know had a number of people working on that initiative as well, but it just didn't really kind of catch on with the community.
So we decided that like we do with a lot of things and then I owe we would just give our customers what they really wanted, which is direct. to the discs hinging storage in kubernetes environments these days historically we've had storage admins, but when you look at kubernetes environments, it's much more converged or hyper-converged depending on whatever buzzword you want to use. Yeah.
So has that function become something that's the compute and the storage all being unified in terms of management or is there still somebody who is you know King storage? So what we're seeing is that there's been such a change in how storage and who takes care of storage and who provision storage it's really more developer driven than it's ever been in the past. So the traditional it storage admin is not really has not made that leap to kubernetes yet.
The complexity of kubernetes coupled with you know, just how modern the overall stack is has really changed things. So developers Architects are really driving a lot of those Source decisions these days particularly around, you know, as it is kubernetes storage. And so that is a change right that's a change in terms of the balance of power.
It's in change in terms of how you go about looking at security across these applications. And so it is a shift and we think that over time it is going to come up to speed. And and become more Adept at managing these things.
But today just given a kind of where the stack is. It really is more the developers that are driving a lot of those decisions. How much are we seeing kubernetes in an on-premises environment?
Because you know the noise is all around the cloud but you know a lot of customers I talk to are still like, you know, we got data that camping going anywhere other than our data center. So how much are we seeing kubernetes in these on premisesments? So I we have a pretty strong position here, which is that the cloud is in operating model.
It's not a place right when you say cloud that does not mean AWS gcp or Azure manio runs in all those we're agnostic to this. 5 million IPS running today in those and so we run on AKs and and so forth. What we see on-prem is that same model and we're seeing more and more companies adopt that cloud operating model on-prem particularly as they mature and understand the parameters of their workloads.
Right? So you go to the cloud to optimize for developer agility for flexibility for speed for provisioning of the hardware. Once you understand that you start to think about what's the next level optimization.
Well, obviously the next level optimization for you is going to be around cost and it may be around some other things. And so that's when you know how this works. You know, how you deployed it in the cloud bringing it back isn't hard.
It's the same basically principles that you would use on premise that you would use in the cloud. So we see kubernetes on Prem as much as we see it in the cloud and you know, we've always rolled out the stat, but I think it's really telling basically containerization dominates the men IO workload 64% of all of our workloads that See our containerized and 46% of those are orchestrated using kubernetes. So that's the default deployment model for minio.
And that doesn't matter if it's in the cloud or on-prem. Do you think we'll get to the promise of hybrid cloud computing? We've been talking about multiple clouds for a while.
When you run into hybrid Cloud you then suddenly, you know the compute somewhere else and the data somewhere else and you're trying to access that and it doesn't always work out so well because you know in the old days we both have some gray hair you want it to bring the computer to the data and I think we're going back to that. I think we spent 10 years of pushing data. Away from compute.
But are we coming back around so I'm gonna say no. I think that our position is that What we see now in terms of kind of this next generation and listen, let's be clear hdfs and we're talking about bringing data to the compute, you know. The entire industry I just was a huge set of gratitude to that but that doesn't work anymore from a capacity utilization perspective just because you have so much storage and you have so little compute that the balance is doesn't work economically and the complexity doesn't work economically either what I think you do see and when you talk about hybrid Cloud, I think there's a definitional piece that we should probably address a little bit as well.
You know hybrid Cloud would suggest that you have one private cloud and one public Cloud. Well, the reality for large Enterprises is they have multiple private clouds and multiple public clouds. And so, you know, we think about that in terms of this concept of multi-cloud and so, you know men I always built for that specific use case.
You can run my out in both of those locations seamlessly and again, you'll find us in the market places and all the public clouds plus, you know, obviously our credentials on the private Cloud so that shouldn't matter right and the principles of cloud native architectures shouldn't matter so you should be able to run all Things everywhere now to your second Point around, you know, moving data and so forth. One of the things that I think has been most interesting this year and over the last 12 months is that you're seeing more and more what we would call traditional database players create the ability for you to query data on external tables, right? So it doesn't have to be there.
So you look at snowflakes external tables. You don't have to move the data in the snowflake in order to query it. You can query it externally the SQL Server 2022 made this a made object storage your first cast citizen again your ability to query data where it sits and again the point there is that those database companies see themselves more and more as high speed query processors at Large Scale rather than they do of you know, I'm a query processor and a data management platform because the data management problem has become so acute that they need to focus on the thing that's most competitive and their clients care most Out which is high speed query processing and effectively.
They're Outsourcing a lot of that storage capability to high-speed performance optimized object stores. So what's that one thing you see organizations doing that, you know makes you shake your head a little bit or you know, what's that newbie mistake or sometimes known as The Idiot tax? Yeah, you would share with some folks to help them.
Avoid it. Well, that's a fair question. So in terms of things that we see is that we we generally see customers think about where they're gonna go and architect.
For what? They think they can see over the next 12 months, but the time frame is so much bigger than that when you're dealing with data and so architecting for the ability to decommission architecting for the ability to expand over time in a heterogeneous way architecting for you know scale in ways that you just don't Envision today is the mistake that we see really most frequently and we try to guide clients to think think much bigger than what they are today because invariably when you plan small you have a tendency to kind of grow into different problems that you wouldn't otherwise experience. Just think about this over the last really nine months since chat DPT came out and you know, the acceleration of acceleration is at a level that I don't think certainly I've never ever seen before I would probably assume that that you share that sentiment and that's not going to diminish.
I think that we're in an Ever accelerating mode right now and it's scary from a technology perspective, but the amount of data that's going to come out of these types of large language models. The requirements that are going to come out of training on private data inside of these companies. It's going to be such that they want to access everything.
They don't want to access small slivers of that. They want to access everything in the portfolio. And I think that's when you talk about planning for scale.
If that's not a big giant red light on your desk that's flashing for you. I don't know what is All right folks they say date is the new oil. But if we don't have any refineries, we're kind of missing the point Jonathan.
Thanks for being on the show that thank you. All right, and we'll be back in a minute.





