Jake Burns on Cloud Strategy for AI Workloads
In this Techstrong.ai Leadership Insights interview, Jake Burns, an executive in residence at Amazon Web Services (AWS), dives into the cloud computing issues that IT teams will need to consider as they deploy artificial intelligence (AI) workloads.
Transcript
Hello, and welcome to the latest edition of the Techstrong AI Leadership Inside Series. I'm your host, Mike Biard. Today we're with Jake Burns, who's an executive in residence for AWS, and we're having a little chat about a subject that's near and dear to everybody's heart, where to run these AI workloads, the cloud on premise, somewhere in between.
Jake, welcome to the show. Thank you, Mike. Glad to be here.
Yeah, I think one of the challenges that you see from folks out there is they're trying to figure out, uh, well, if I have data on premise, can I move it in the cloud? Or does, uh, the old school rule apply? I should just bring the compute to where the data already is.
Uh, I think that you should, um, you should bring the compute to the data where the data is, but your data being in the cloud gives you so many advantages to use these cloud-based, uh, AI systems. And the reality is that moving data to cloud has become a lot easier, um, recently. So there's, uh, there's faster networking, there's other tools available that can help get your data to the cloud, um, and you don't necessarily need to bring all of it, um, in order to take advantage of these tools.
So, uh, just like everything else gets started now, uh, the best time to do it was yesterday. The best next best time is today. And the sooner you get started, the sooner all your data will be there.
But, uh, I definitely believe that having your data in a secure environment, which the cloud, uh, provides, um, is, gives you the advantage of being able to use inference in the cloud. And I think that you have the best, latest and greatest, uh, inference with those best guard rails and the latest models in the cloud as well. So it's a natural fit.
Mm-hmm. And to your point about that, I feel like if I'm not involved today, it is gonna be hard to catch up because a lot of people will have, uh, a lot more experience than I do, and they'll be much more competitive. So there is a cost of falling behind, right?
Oh, absolutely. Uh, I would say the opportunity cost here of falling behind is, uh, greater than anything I've seen before. Hmm.
So is there some way to think about where these workloads should go by some sort of pattern they exist or some sort of, uh, metrics that we're able to track or observe? Or is this all gonna be trial and error because, well, it's all brand new? Uh, I think both.
I think trial and error is good. And I think, again, those who are experimenting today are the ones who are figuring it out the fastest and they'll be kind of writing the future. Um, you know, I've run, I've run, uh, AI models, um, locally and I've run them in the cloud.
And, um, it's, it's very difficult, um, to, to have enough, um, to, to be able to invest enough in the hardware necessary to do, to do those local, um, implementations. I do see perhaps some value in it, but I also think that, um, you know, we, we went through this with cloud where there was this kind of debate between is it more secure to run workloads on premises versus in the cloud? I think to some degree that dec debate's still going on, although I think in my mind it's settled and I think in most people's minds, um, the cloud, uh, when you don't know much about it can sound scary, but the more you learn about it, and this has been my experience working with customers, uh, you know, CISO will go from dead set against cloud to being our biggest advocate once they learn kind of, um, the way the cloud works and the way you have to think about it differently in order to implement security.
But then the far greater capabilities you have with security and privacy, and I think that's true with AI as well, and it's a, a process. We're still going through an education process, we're still going through. I also think it depends on how you do the math.
If you do the cost of the cloud, but you're not including the cost of running your own data centers as savings, well, that's one issue. And if you're thinking that your existing data centers are kind of a sunk cost that you're gonna leverage, well, you might be in for a rude awakening when you gotta go buy a bunch of new servers that have GPUs in it, right? Absolutely.
Yeah. And then also these GPUs, um, you know, there's the refresh cycle on-prem, you know, and maybe you got a three or five year cycle. The refresh cycle for these, a cutting edge GPUs is just, uh, far faster than that and speeding up.
We don't know how fast that's going to get. So trying to keep up with that I think is gonna be, uh, a major challenge for those who try to do it not in the cloud. If you leave it up to the cloud providers, then well, it's cloud providers problem, and that's, uh, a good thing to outsource, uh, just always having access to the most powerful hardware and the most powerful models.
And every data scientist I talk to seems to wanna always have access to the latest and greatest. It turns out we talk about this GPU shortage, but the reality is a little nuanced. It seems like most of the data science teams just turn their nose up at the older GPUs, so there may be plenty of 'em, it's just they don't want to use 'em.
That's true. Everyone wants to use the greatest, uh, out there because everyone wants to have that competitive advantage. And, uh, those, the competitive advantage versus having something that's just one generation old, which might be three months old at this point.
Um, it's, it could be, uh, a huge difference in the type of results you get. Mm-hmm. We hear people talking a lot about small language models, and some folks are saying, well, they're gonna use those in conjunction with inference, and that'll make maybe the whole cost equation more reasonable and you might actually get better results because, well, a small language model is trained on a narrow base of data.
Is that a reasonable set of assumptions and is that something people should be thinking about? I am, yes. I'm actually a proponent of small language models and I think they don't get enough attention.
Um, but then I'd also say that, um, you, the cloud could be very advantageous, uh, for those because it allows you, uh, essentially when you're using a small language model, a lot of the power comes from being able to specialize it for your particular workload and your particular preferences. And so you still need considerable resources to train those models. You're just basically, um, taking the, um, the, uh, the expense and the time and the resources required from the inference to the training, um, in my, in my opinion.
So, um, you know, bedrock, Amazon Bedrock, for example, you know, does model distillation, which can allow you to create, uh, like smaller models based off of these most powerful models. And you may wanna do that, um, that cycle very rapidly. Um, in fact, you should.
So I still think there's a role for cloud for that. Um, but of course they are, those models can run on-prem on lower hardware, so then it gives you that freedom as well. Right.
Do I also need to factor in, you know, the laws of physics here? Because I think sometimes the, uh, inference engine wants to be closer to the foundational model rather than, say, a wide network away or four or five network hops away and that will matter to my performance in my application. Is that something to think through?
Oh, yeah, absolutely. Um, latency is a big issue. Um, you know, especially when you're talking about use cases that require real time responses.
Um, so like, you know, assistance or customer service agents, things like that. Um, even sentiment analysis where you want like on calls, um, you know, we have contact center as a service, uh, Amazon Connect, which doesn't, again, it doesn't get enough, uh, a love in my opinion, but the customers that use it see, uh, a lot of value from some of the ai, um, um, tools that are in there, like being able to tell, um, in real, real time, uh, what your customers are, what their pain points are, um, that's a very difficult thing to, uh, develop yourself. So to have an out of a box solution like that, um, is super, uh, super useful.
Um, to, to your original question though, um, I think that, uh, that, uh, insight is one of the primary reasons why cloud migration is so uh, critical today. Um, because, uh, assuming you want to take advantage of all of these capabilities in the cloud, then you wanna have your data, uh, live very close to where you run the inference. So if you're gonna take advantage of something like Amazon Bedrock and, and have all of those things outta the box and have kind of the latest and greatest available to you at all times, um, you want it where your data lives, and then you also want your data to be secure, because it's my opinion that, um, in the, in the next, in the near future, one of the most important things is going to be securing your data because your data is going to give you the competitive advantage over others.
You know, because the models are becoming commoditized to, to a degree. Um, being able to run those models against your proprietary data is what's gonna give you the, the real value that you're looking for. Is this, to a certain degree, maybe a rehash of the debate over data gravity.
And I'm asking that question because some folks will say, well, I'm putting all my data in the cloud and I'm gonna have a big data lake. And other folks are saying, well, all my important data is running in my on-premise environment because that's where it's driving my mission critical application. So maybe before we have a conversation about workloads and infrastructure, do we need to have a conversation about, well, where is the data gonna be in the first place?
I think so, and I think we should have the conversation at the same time. Um, uh, another observation of mine, and this is kind of my, my passion at the moment and my mission at the moment is that migrations, um, can happen very rapidly. And I think that most migrations that I observe happen far too slowly, and there's a lot of reasons for that.
And, um, it's something that I talk a lot about, um, and, and, and work, uh, most of my work is around that nowadays, but, uh, migrations could happen very rapidly. Uh, when I was a customer way back in the day, uh, almost 10, 10 years ago now, uh, I did a rapid, uh, accelerated migration, uh, to AWS, uh, as a customer and was able to go all in in 17 months. And if I was to do it over today, uh, knowing what I know now and with the tools available today, I think we could do it in maybe three months.
So migrations don't have to take a lot, a lot of time. Uh, it could be very rapid. All right.
There is an old school way of thinking about data migrations. It starts with a phrase that says, nothing good happens when you move data. Have we gotten better at this?
I mean, 'cause a lot of folks, you know, seen people's careers get trashed over trying to move databases, nevermind entire warehouses. Yeah. So it's interesting because there's so much nuance to this.
Two things could be true at the same time. So, uh, it's, it, what you say is absolutely true, and I'm seeing this, um, uh, far more than, than I'd like to, uh, where people are struggling to migrate their data, but it's actually, um, uh, quite easy. I think it's a skills issue, um, uh, education issue and a best practice kind of issue.
So like, for example, uh, when I was migrating, there weren't the, uh, you know, the A AWS snowball, uh, family of services, uh, which essentially is just like a huge dis array that you could put in your data center, copy all your data, and have it shipped to Amazon and, uh, AWS data center, and then have it, uh, ingested into your, into your environment. That is a super fast way to get your data into the cloud that I think is very much underutilized. Hmm.
So as you kinda look at all of this, what's your best advice to folks? Or the converse of that is what just makes you shake your head a little bit and say, you know, folks, we could be a little bit smarter than that. Uh, so much.
Uh, so I would say, um, focus on the pro like your pain points and the opportunities that you can't solve today. Uh, I see, uh, kind of an anti-pattern with, with AI right now, uh, with agen ai, uh, specifically now is it's kind of like a, a problem looking for, you know, uh, a solution looking for a problem. But most organizations, all organizations have problems today.
We can't go, we can't do these things that we want to do, or we're having these specific problems. Um, the, the best way to learn how to use these things is to solve a real problem that you have. So start small of course, and then work your way up.
Um, do it in a safe way. Do it with enterprise grade tools. Uh, you really don't want to be using consumer grade tools to do these things as an enterprise because of privacy and security and guardrails and all of those things.
So I think if you do that and the organizations that do that are seeing a lot of success, uh, there's a real divide right now, which is really interesting. Um, there are companies who are very much behind the curve on this, and then there are other organizations that are doing things that, um, are incredible that just blow my mind. So it's, uh, it's, it's kinda a tale of two cities and, uh, I think you really wanna be in that forward looking one, um, but do it safely.
Of course. I think there's also folks out there that think that maybe they wanna build the perfect architecture before they get started and uh, maybe good enough always beats perfect, but do we need to just kind of go forward? 'cause the truth of the matter is we're all kind of learning together at the same time.
Absolutely. Spot on. So one of the biggest problems that I've seen and, and I've been, uh, uh, working on five migrations for about a decade at this point.
And so yeah, I'm seeing the same thing with ai. One of the biggest problems is, uh, try waiting until you have a perfect plan. And then what ends up happening is you never have a perfect plan.
So it just becomes, preparation becomes procrastination, and then you fall even further behind. If you do happen to put together what you think is a perfect plan, by the time you implement it, that plan is gonna have to change anyway. So I'm not saying there's no value in planning, of course there is, but you want to get, um, you want to do iter, do it iteratively, and you learn generally by doing, uh, not by planning.
Um, so doing is the best way to get ready. All right, folks. I believe it's one of the laws of nature that says it's easier to change the direction of something that's in motion, and it is something that is sitting still, and that's true with ai.
Hey Jay, thanks for being on the show. My pleasure. Thanks for having me.
All right. And thank you all for watching the latest edition of the text on that Ai, AI Leadership Inside series. You can find this episode and others on our website.
We invite you to check those all out. Until then, we'll see you next time.