Deploying Stateful Container Apps – Garima Kapoor, MinIO
MinIO COO Garima Kapoor explains how deployment of stateful container applications has led to more than a billion Docker pulls of company’s storage platform.“
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with grima Kapoor. Who's the CEO for men IO and they do a lot of work in the storage space and have passed the one billion download Mark for Docker containers, which you know, when I was thinking about that several years ago, I never occurred to me that we might get to a billion of anything green welcome to show. Thank you.
Thank you for having me my What are the implications of having that many downloads what exactly are people up to and what do you think is driving all the interests? Yeah, I have been a billion DACA downloads. It's a huge milestone for us, especially since we started back in 2014 from zero lines of code to now it has been just a phenomenal traction that we have gotten from, you know, data Architects data Engineers Community overall and just to put some things in perspective right to in storage where these numbers are unheard of and when and what I mean by that is that you know, whether it is a GitHub Stars whether it is downloads this number of downloads, they're always reserved for database or you know, AI ML Class of applications or applications from the compute side of things not so much, you know traditionally from the students standpoint.
So that is why this number is even more meaningful because there is I don't think so. There is any other storage whether it is object file or blog that has gotten to this level of acceptance and adoption within the application and the developer. The overall so that is why this number makes it very special and it is not just the billion Docker downloads, but also the trajectory of other numbers right whether it is the GitHub Stars whether it is our slack Community all the numbers are pointing to the direction in terms of the adoption and growth of manio.
So that is what we are excited about. Not too long ago people were saying you should only build stateless applications in kubernetes environments. So this is even more compelling in the sense that most of those downloads are probably been recent.
What exactly are we seeing when people start to build more stateful applications on top of Docker and what are the implications? Yeah. I think first thing that we are seeing and this shift has happened over last couple of years right that S3 API first objects to reach has become the primary storage of all means within the application ecosystem framework if you see so even the traditional applications like, you know, traditional databases whether it is teradata vertica even sequel, they're all becoming object storage first in their approach of doing things and also the new age, you know, AI ml Frameworks like who flow tensorflow they're all built on object storage first approach and I think that is what has happened and that is what is really pushed me now you into that massive scale of adoption because there is really no other object storage system out there.
First of all, which is software defined which is open source. So application developers can just get started within like five minutes of downloading when IO and integrating it as part of their staff and it is extremely simple and easy to use so I think those Kind of things have really helped propelled menio in terms of the acceptance and adoption and I think a lot of it has happened just because you know or most of the applications that we are seeing in our ecosystem are all object storage even snowflakes announcement that you would have seen recently, you know, they have gone in terms of people build on AWS S3, and now they have disaggregated in an approach to accept external tables or data from external S3 object stores as well. So these this kind of disaggregation of storage and computers what we are seeing everywhere data remixes all object storage there, of course, every thing that they lose in public Cloud, but it's a natural Direction in terms of these application infrastructures to become just compute only and disabricated and just standardize on object storage compatible high performance object storage as their first storage.
So I think that that has made the switch happen for us. There's a lot of people who think that there's no such thing as high performance object storage. So what exactly makes that happen and are we starting to see people migrate or standardize on one type of storage format?
Because you know historically we dealt with block and file and everything else and there was a lot of headings involved. Yeah, so vanity started. It was all about you know, third year of storage is object storage and you know, we just put the data and never touch it again.
So we have seen that word as well. And I think it was midnight that actually first went ahead and did the optimizations to ensure that we can deliver on the high performance. I think that was back in 20 16 or 2017 kind of a time frame where you know, we went ahead and meet the optimizations in terms of making sure that also because we are software defined right?
We have a responsibility to run on any commodity Hardware. So we had to do certain things certain way to make sure that our design was modular enough. Our architecture was simple enough to take advantage of any underlying Hardware system.
So that those things actually led to making me know high performance in terms of our partnering with whether it is Intel to make sure we can really take advantage of the ABS fight World instruction set and so on. So we did quite a bit of work in order to make object storage high performance in that aspect and secondly what helped us was our Focus because we have just talked to being object storage and object storage only even though we keep getting requests that why don't you add, you know, NFS, you know API to our story system and whatnot. But we have just been very focused that we just want to be S3 API compatible object storage.
That's what we are and we've been very focused in that. So I think that design and architecture decision also meet our architecture very simple in a way so that we could really take advantage of the underlying Hardware or instruction set. Otherwise when you make software complex, it becomes very hard to deliver on the performance itself complexity reads to a lot of issues at scale.
So I think keeping the architecture simple making having design goals very Is I think that is what enabled us to really deliver on the high performance. And of course the work that we did with the Intel and all the engineering effort that went into writing vinayo code in go assembly as well. So all those things are affects of making sure that our design principles have always been do one thing and be the best at it.
So I think that's what has led to when I have been high performance Object Store. Do you think we're on the cusp then of seeing some real hybrid cloud computing because we have multiple clouds these days but the issue is always been the data the data and isn't easily moved it kind of difficult to manage. And then every time we had a new platform the total cost of it goes up and we wonder why so are we getting the point now where we have this kind of consistent layer of data storage.
Can we really start to think about real hybrid cloud computing? Absolutely. I think you get the nail on the head.
I that is the future. That's what we see increasingly customers doing whether they have been born in one cloud and you know wanting to go to multi-cloud or whether they are reaching from on-prem to a public Cloud infrastructure and wanting to keep the underlying storage infrastructure scene because like you said data has gravity it's very hard to move data around. So first of all, I think as a design principle what we see the customers are standardizing on S3 object storage as the foundation layer because then what happens is that because most of the applications have become object storage first in our case, it's me now first most of the time.
So in that case you can easily move these applications across different clouds across different infrastructures without even a single line of code change. For example, you can run the same spark jobs on menio on Friends same spark jobs. I mean I know in gcp or in Azure or in AWS, right?
So the Flu doesn't change at all. Secondly what we are seeing is that the need of data to be available across different clouds and what I mean by that is because you know, there are certain customers specifically trading firms which have standardized on one cloud and they have experienced single Cloud going down and their business getting impacted and so as a consequence of that what they have been wanting to do is reach out and have a multi-cloud infrastructure so that they can even have a cloud disaster recovery, but the foundational piece is data here, right? So the data needs to be available across this Cloud, so what they do and what they're doing right now partnering with men IO so when I get deployed within AWS within Azure within whatever cloud is a choice of cloud for the customer gcp and me now, you take care of replication across all these clouds we have active active replication so that the data can be synchronously available across these clouds.
So even if SoundCloud goes down. The other cloud is still there and they are able to have business continuity. So that is secondly that we are seeing and thirdly there are another Trend that we are seeing is some of the customers.
like Banks like They have just wanting to standardize on one Cloud just be in one of cloud but they don't want to go through, you know, the security issues that someone like Capital One would have gone through multiple times with AWS not part of fault of AWS, but just because systems are complex that are always loopholes and if there is a human in Middle, there are always things to you know, expose from security standpoint. So now those organizations what they have done is that they deployment within AWS itself me now. You becomes the hot tier for getting all the data and all the security policies and access controls get tied to me now itself so that the customer if themselves can manage whether it is, you know, Key Management store or whether it is whatever I am identity management systems that they have in place.
They have complete control on that so they can have complete control on the data complete control on security and access infrastructure. And then when are you takes Responsibility of peering the data from whatever lifecycle management the customer would have set it up. So all these three friends that we are seeing increasingly and I think all these three Trends are resulting in one thing in common is that customers are valuing more and more, you know, they Place more value in terms of having a complete control on the software stack and not getting stuck with whether it is one Cloud vendor lock in or whether it is getting stuck from a data stand point.
So I think that is an overall trend that we see across the customer base no matter what the deployment apology looks like whether it is hybrid cloud multi-cloud or just private Cloud as we go along and in the cloud initially you didn't see storage administrators, but as we get to this multi Cloud world are we starting to see some new classes of storage administrators? Maybe we call them storage Engineers or storage or data operations people. I don't know exactly what we're gonna call.
Facebook but it looks like we actually need people to manage data again, and yes a full-time job. It is I think. I think more than full-time job where we see.
The requirement arise is from the operation side. I think because as the scale grows as the infrastructure extends Beyond clouds different clouds different infrastructure, there is an operational overhead and I think that is where also you need to take into consideration all the security because I think with you know more and other things going on rasm where attacks have all so increase. Oh operational overhead is something that we see across as the data is scaling and you do need a team that is very skill in terms of managing the operation side of things because if you have a very strong devops team, I think a lot of your issues can get resolved because usually the devops team is is expert in kubernetes.
They can spin off systems that scale they can manage the systems of scale. And also if you have a strong security team then everything comes together very cleanly Not so much in terms of the storage systems overall, but I think we do see the that the need for a good devops team. It is increasing within all the large scale organizations, especially as the scale of data increases because if you're thinking about like spanning like say 100 petabytes of cluster, you do need a strong devops team to make sure that it is deployed as per the topology that has been architected.
So I think that is what we see more than just the storage engineering. Hmm. Do you think data Ops and devops are going to meld in some way?
I mean, how will that come together? Because it seems like they are kind of hand in glove but to what they're how and what It is it is quite interesting question. I think data are I would say is more closer towards, you know, the ETL or elt side of things and devops is more towards infrastructure scaling.
That's how I usually differentiate and I usually see in a in the world, you know within our customer base as well. But yeah, I think they're going to a little bit separate because the responsibilities and the realm of the things that they take care a little bit different from that standpoint. All right.
So, once that one thing you see your organizations doing over and over again, that just makes you shake your head and go I can't believe we're still having these issues. But this still having dishes I think one thing that and this this is a what I have been seeing increasingly is post-covered, you know that pre-covered. It was a talk about going to Cloud but post covid it's all about you know, CI is taking it very seriously in terms of executing on that plan that we would have seen.
I think certain things that you know are come as after thought need to come more in Forefront as someone is designing the cloud infrastructure. For example, what we saw earlier was that multicloud used to be an afterthought, but now it is become like, you know front and center of the design architecture itself. So I think that's a really that that's a big change that we are seeing across different organizations.
And that's one shift that has happened tremendously. Secondly. I think there is more awareness in terms and the options that we are seeing how Customers are viewing just public Cloud itself.
They are viewing public Cloud more as a chief CPU drives infrastructure, but they do want to have you know, complete control of their software stack across. So I think that is second Trend that we are seeing within our customer base and within you know our conversations with the with the broader Community as well. So these two friends I think are going to stick and I I think are going to grow grow Beyond because data is growing at an insane rate.
I think see it published one study that it's going to be what one 25. Is that our bikes of storage by 2025. So I mean if that data is going to grow you need to have complete control of the data stack.
I don't think so. There is any organization data, I think all organizations are going to become data companies and they are becoming data companies. So that that is where the value lies.
You need to have complete control on it. You cannot just yeah. Give it to orally or give the controls to someone else.
I think that is something that we are going to see increasing within within the industry. All right last question. So when are you gonna get the two billion?
If you're counting you started the countdown. option has been phenomenal what I always say is that it's not just one number we need to do well across our different methods that we crack whether it is from open source and from commercial adoption open source, we know we have a bigger ecosystem than you know, public Cloud so and applications and developers really love me now and proof is in the pudding there and And what I always say is that AWS is last year's as three run rate for 10 billion dollars. So we are just getting started, you know, in terms of getting the Mind share and getting the revenue share.
So that is I think a place that where we are going to accelerate and that that is the focus of the company to get as much as market share from revenue standpoint with the broader Market. Yeah agreement. Thanks for being on the show.
So thank you my pleasure to be on on the show. All right back to you guys in the studio.