Building Enterprise AI Platform using Hypercomputer | The Six Five Summit
Google Cloud believes in offering flexibility and choice of products so customers can build the AI platform that suits their unique needs. We’ll share how Google Cloud empowers customers to build their AI platform to attain better performance/$ and shorter time to market. Tune in to hear how Google Cloud’s Hypercomputer enables you to build innovative AI applications without having to worry about underlying infrastructure, portability, compatibility, reliability and scalability issues.
Transcript
Welcome back to six five Summit. The conversation about AI continues. I'm Dave Nicholson and I've got a very special guest from Google.
Uh, Mr. Mullin Patel is Director of Product Management for GKE Google Kubernetes. And what does the e stand for?
Mullen Engine. Engine, GKE. We're gonna be talking a lot about GKE.
Welcome Mullen. Thank you Dave, and glad to be here. So a little more introduction on my side.
I've been with Google for six years leading the GKA team, owning very exciting areas, including AI platform, our stateful workload business, and also our gaming business. Before coming to Google, I was a general manager at ge, and before that I used to be principal data scientist. I'm super excited until to talk about all about AI today, so take it away, Dave.
Fantastic, fantastic. Yeah, it's always, it's always good to know, uh, I, I, I had a, uh, a mentor tell me once, always let people know a little bit about yourself beforehand, so they know why they should listen to you. Malin is someone who we should listen to.
He has a lot of practical experience, which always makes conversations interesting. Um, there are a variety of ways that people can build and deploy AI solutions in the Google Cloud platform, GCP for short. Um, you are on the sort of, uh, GKE side of things, but what are, what are some of the things people need to consider when they go in looking at whether or not they're going to deploy via the sort of self-managed GKE route versus a more managed enval environment, a more managed platform like Vertex?
Tell me about that. Fantastic question. So in order to build, let's say, an AI application, a customer has to deal with a lot of complexities.
For example, you start by gathering the data, cleaning the data, extracting the features, then you have to build a model. Then you have to provision resources to train the model. And after the model is fully trained, you have to put it in production with the live traffic and monitor its performance.
This requires a very sophisticated platform that allows you to do this at scale and makes the entire process simple, easy, repeatable, reproducible, and auditable. So there are many ways in which you can do this, but from the platform point of view, I will broadly classify two main approaches. So let's say if you are a customer and you are looking for one stop shop, comprehensive end-to-end mops platform, which is feature rich, which gives you the greatest and latest and which cuts down your learning curve and makes you productive almost instantly.
That will be the platform, we call it Vertex ai. The beauty of Vertex AI is it's actually fully managed platform, which means as a customer, you don't have to worry about provisioning and managing and maintaining underlying resources. All you have to focus on is like, how do I get my data?
How do I build my model? And in some cases, you can use open source models, go live very, very fast. So that will be the route for customers who are looking for convenience.
However, we also have lots of customers who want to train and serve models across multiple platforms. They may be doing it multiple clouds or even on-prem. So what they are looking for is portability.
I want to write my code once and run it everywhere. Often many customers are also looking for flexibility and full control over their entire stack. They have finity for open source, and some organizations are also very worried about vendor lock in.
So for all these reasons, lot of our customers actually choose to build their AI platform on Google. Kubernetes engine. Customers like Anthropic character, AI runway, Shopify, Spotify, snap, you name it, lots of those customers actually fit in the CL second category where they choose to build their own platform on Google Kubernetes engine.
And this gives them full freedom. Another, uh, dimension to this is often customers have built their web serving platform on Google Kubernetes engine, and now they're venturing into ai. So the natural, it is natural for them to leverage their existing investments and build a platform that gives them consistent operations across DevOps and mops so they can leverage their workflows and entire CICD pipeline, same way they have been doing for DevOps for their ML ops.
So that's another reason why customers build their self-manage full control platform on Google communes engine. So it's interesting, when you described Vertex ai, um, it sounded like a pretty good value proposition, and I thought we could sort of end the conversation there, but the, the reality is there's still a lot of people who, um, who, who want to leverage, uh, infrastructure as a service in terms of both hardware and software because they wanna have that control. Um, you mentioned portability and, and and flexibility.
Um, is this, I mean, how much of this do you think for GKE is, um, customer's desire to be able to control things and create value in their own solution, um, versus just the psychological fear of giving up control? Is it a combination of being rational, building it yourself? Uh, you know, uh, i I, I know Google is agnostic on this, um, but you clearly for a reason have both approaches managed and, and and, and, uh, customer managed.
Any, any thoughts on that? What's driving this demand for continued demand for GKE? Yes, so great question.
So there are legitimate reasons why customer would choose to build a platform on their own. We all know it's hard and you require quite a sophisticated, talented, um, machine learning platform teams to build their own platform. But in some cases, that investment is well worth, and that's what we hear from our customers.
So the reasons I outline are sometimes portability, flexibility, control, open source affinity, and also avoiding vendor locking. And these are good concerns for some of the large customers who may have footprint in multiple places, including Google Cloud and other cloud vendors, as well as open source, sorry, on-prem vendors. So for those reasons, it does make sense for those customers to take that extra effort to build their own platform.
Sometimes it also helps to be in charge of your own destiny. So if you are very doing cutting edge innovation and you want to be in full control of your entire stack, you want the freedom to basically implement the latest and greatest models or tools or open source frameworks as soon as they come out, then you want to be in control of your platform and the entire infrastructure. Sometimes it's also driven by the desire to optimize the performance.
So a lot of customers who spend a lot of money on GPUs and GPUs, they like to basically optimize their network performance all the way down into operating systems, security and lot of other things. So that's another reason they continue to invest in infrastructure as a service. So you me, yeah.
And again, you mentioned portability and flexibility and that's great in Kubernetes. Fantastic. Uh, you know, uh, easy to move things around, easy to move between clouds, between on-prem and off-prem.
Um, but why, why GKE? You mentioned optimizing performance. What about cost containment?
What are you doing? Uh, look, I can, I can, I can find a GPUs and Kubernetes, um, at every corner seven 11 these days. But seriously, what are you doing that's special with GKE that people should know about?
Fantastic question. So let me start with what are the main pain points and customer requirements when you're building the AI platform? So if you look at today, I would highlight a few things that are must have and the top priority for most of our customers, whether you are doing GKE or even it applies to Vertex.
So the pace of innovation in AI is astronomical, right? Every day there is some new technologies coming up. So if you are in the business of ai, you, for you time to market is very, very important.
So you want to be the first, as much as possible, you want to launch new features and new technologies as fast as possible. You all, the cost of GPUs and TPUs is astronomical. So again, you wanna minimize the cost fourth for training and serving.
And third is obtainability of GPUs and TPUs is still very challenging. So we are in the scars resource environment. We are finding the GPUs and TPUs still is a challenge if you want to do things at a scale, I understand like you can get it that at the corner if you're looking, uh, for fewer, but if you're looking in large numbers, then you need to optimize your platform so you maximize the availability and obtainability of your GPUs and TPUs.
So those are the primary concerns for customers. So in the GKE team, and you meant, by the way, uh, just to interject, you mentioned GPUs and tpu, um, you're Google is gonna leverage whatever the, the best fit for function is from a hardware perspective, right? Correct, yes.
And uh, the customer has full choice when they use Google Kubernetes engine, right? So you could use GPUs, you could use CPUs, you could use CPUs. We give you full freedom to pick what's best for you.
And within each category there are multiple variants. There are, Nvidia provides many different variants of GPUs, whether it's L four or H 100, T four for TPUs, we have various generations of tpu and we offer you full control in choice what works best for your customers. Now, going back to the innovations that we're bringing to market.
So our primary focus is to help customer get best performance for dollar. And that means couple of things. So if you are doing, let's say, inference on GPUs and tus, then you want to make sure that your expensive GPUs and TPUs are fully utilized.
And what we have found in practice in looking at our own fleet and our own internal users in Google, is that utilization of GPUs and TPUs is very, very low. And what I mean by utilization is duty cycle. So why is that happening?
The reason utilization is low often is that in order to run an online inference service, you have to account for dynamical nature of your workload. The workload can spike when there are many users of your service and it may go down over the night or when users are busy doing something else. So how do you account for it?
So Kubernetes is naturally designed to autoscale, which means when the load spikes dynamically autoscale the cluster and the launch load shrinks, it will scale down the cluster so you save money. However, in the world of GPUs and TPUs, how fast you can spin up a new node and how fast you can start a new workload has been a very, very slow. So in our own experience, when we were running like a couple of hundred billion parameter models, downloading them on a new node used to take hours.
And if you are auto-scaling is so slow, then you compensated by having significant or provision capacity to be able to handle load spikes and that increases your cost and the utilization goes down because the majority of time those GPUs and TPUs are idle. There are other reasons why they are all, sometimes they're waiting for data, sometimes they're waiting, um, they're not being used because they're under maintenance or there could be failure recovery scenarios as well. So in Google IES engine team, we have brought up significant innovative technologies and I would like to mention a couple of them.
We have something that we launch in the last, next in April where you, we allow customers to preload container images and model weights on a secondary boot desk. And by doing so, the boot time of the workload cuts down from hours to minutes for very large model and for smaller models it can cut down from minutes to seconds. So our Vertex platform actually runs on GKE and they actually implemented this secondary boot day space, fast workload startup technology, and for a 16 GB model, and they found that they could speed up workload startup time by 29 x, which is astonishing.
So this is how you get a better utilization. So that's the one example of innovation. We are doing similar things for training.
So in the training, the challenge here is that you have a very expensive GPUs and TPUs and they're idle because you can't pump the data fast enough. So we launched something called Google Cloud storage with FUSE and local caching. What it does is it downloads the data locally on local SSD and exposes this to as a file semantic.
So you can take any existing model, maybe you built it on your own or you got it from internet. Typically those models read files so you can get those files read directly from local SSD instead of going over the network to Google cloud storage, which means the IO speed is really, really fast and you're keeping your expensive GPOs and TPOs busy. So when you say in this context, when you say local, you mean local to the, um, to the pods and clusters that are in the g in the GKE context, not local as in on premises, correct?
Well, I mean, local is like a physical node, so localized. These are attached to the nodes, right? So when the pod boots up, the data is locally available, which means IO speed and reading and writing is really, really fast, which means you're keeping your GPUs and TPUs for training really, really busy instead of waiting for data to be loaded from a network drive.
Excellent. Very interesting. Very interesting.
Um, in our closing moments together, uh, can you give us, you know, if I, if if we were to meet, uh, if we were to meet at a restaurant and I had just a couple of minutes with you to say, Hey, give me some words of wisdom on, on common things that maybe are misunderstandings when thinking about deploying AI or, or some hot tips that you might have, uh, what might those be? Yep. So great question.
As we discussed, typically our customers are really motivated to cut down the time to market, to save on a cost and get the best performance. These are all the right metrics. However, in a zeal to go really, really fast, often, some other factors which are super important, are overlooked.
So for example, security reliability and day two operations this often take a backstage and that creates problem in the long run. So what we would recommend is that when you design your platform, make sure that this very important elements like security, reliability and day two operations are not overlooked because you will be spending a lot of money building foundation models or any other models and that has your intellectual property. You don't want them to be stolen or leaked.
Uh, when you go live with your service, you wanna make sure that your service runs and it has the desired uptime so your customers, end users get the best experience when you run the service at a scale. You also wanna to minimize the toil on your operations team. So day two operations become a really, really important element.
And if you think about the explosion in like a GPU driver version, Kubernetes version, the nickel version, and all poor mutations and combinations, so operating systems, the managing and maintaining and making them secure could be really, really tosome. So you want to design a platform that takes care of all these challenges from the get go instead of those being an afterthought. Sage advice from Molin Patel.
Uh, where should people go to learn more about Google Cloud engine and the, uh, AI platform offerings from Google? Yeah, so we have a very rich, um, like documentation and websites full of, uh, self-help guides as well as best practice guidance and videos. So cloud website, look for any particular topic that you are interested in, whether it's a AI platform, vertex, or the things I mentioned, GCS views with local caching, or even like some of the new technologies we have launched with secondary boot disk space, fast workload startup.
We also have some new cool work done on the Kubernetes, which we call queue, which allows you to basically dynamically share your GPUs and TPUs between training and serving. So also among different teams. So every team gets their fair slice of GPUs and TPUs and your high priority workload gets priority and lower priority workloads can be preempted dynamically so you get the most utilization out of your expensive resources.
All of this and many more features are well documented. They're available on our website, so please go check it out. And we are always here to help.
So you can directly reach out to any of the Google Kubernetes engine team members or Google Cloud specialists and your account teams, and we are eager and willing to help. Fantastic Molin. Molly Patel, thank you so much for joining us and for the rest of you out there on the interwebs, stay tuned for more exciting coverage here.
Six five Summit.


