Vasi Philomin on AWS Nova LLMs
In this Techstrong.ai video interview, Vasi Philomin, vice president of generative artificial intelligence (AI) for Amazon Web Services (AWS), explains why the cloud service provider is providing access to a less expensive family of Nova large language models (LLMs).
Transcript
Hello, and welcome to the latest edition of the Techstrong AI video series. I'm Mike Ard. And today we're joined by Vassi Philman, who's vice president of gender AI for Amazon Web Services.
And we're talking about the new Amazon Nova large language models that recently got launched. Basi, welcome to the show. Um, hi Mike.
Thanks for having me here. Pleasure to be here Every time you turn around. Lately, it seems like somebody's rolling out a large language model and you guys already make a significant number of them available to folks.
So what was the compelling reason to go build your own LLMs and what makes 'em different? Yeah. Um, one of the things that we, um, that we're fortunate to have inside Amazon is that we have a lot of different businesses all over the company that are actually trying to use generative AI and achieve, uh, new business outcomes.
And so we get to see very closely, you know, where, where their challenges are and where the gaps are, and where there are potentially spaces where there's no product that actually meets those needs. Um, so we've been, um, we've been building our models for a long time, and one of the things we noticed recently is that, um, that there was a need for really fast models, um, that are also, you know, affordable in terms of the costs. Um, and, and they still have the same kind of performance that they need to have.
Um, so that's what, so we, we had all these internal teams, uh, you know, using, using our models that we had before and then telling us, you know, that this is an area where, uh, that they would like to have something. And, and so we found that a lot of our internal teams found the Nova models to be very useful, uh, and filling in a gap that wasn't really covered by any of the other model providers out there. And, and so when we thought that it was gonna be useful for all of these teams, we thought it would also be useful for all of our customers of AWS, which is why we launched our new generation of foundation models, uh, the Novo models that have state-of-the-art intelligence across a wide range of tasks, but more importantly, it has industry leading price performance.
So on that price performance question, are we starting to see folks become a lot more sensitive to the cost of AI when they get into production environments? Because, well, if they're for every input and output, there's an API call and things start to add up after a hour. Yeah, absolutely.
Right. You've hit the nail on the head right there. Um, it's very easy to build a prototype or a proof of concept and, you know, do a cool demonstration for your CEO, but you know, the real work begins when you actually wanna scale it and put it into production.
Um, and so one of the things that Amazon does really well is we're good at taking technologies like generative AI and applying it at scale to real world business problems. And I think the two phrases that I'll point you to is scale and real world business problems. Um, and so for us, what we've done here is make sure we identify these areas that are issues, and then make sure we kind of fill those gaps.
So once we get into what no is here, but how many models are there, what types are there, and how big a portfolio are we looking at? Yeah, when what we announced at Reinvent was a whole category of Nova models. Um, so there's gonna be text, uh, models in there, uh, or people refer to them as large language models these days.
Um, but the text models, output text, uh, and we've got four of those models just in that text category. So there's Nova, um, Nova Micro, Nova Light, Nova Pro and Nova Premier. Um, and except for the micro model, the other models not only take text as input, but they also take images as input and also videos as input.
Uh, but the output is always teched. So there are these four, uh, Nova models, uh, on the tech side, and those are the ones that are, you know, at least 75% less expensive than other models on Bedrock in their respective intelligence classes. They're also the fastest models in their, in their category.
Um, and then moving on to the next category, we have, um, you know, image generation models. So we have Nova Canvas where your input is text, but your output is an image. Uh, it's a high quality image.
And these, these kinds of models are very useful in, in places like, uh, you know, ads, um, brand campaigns, marketing campaigns. So it's very creative workflows where you need to generate images based on a prompt. Um, and so that's kind of where Nova Canvas is going to be useful.
So that's one of the models that we have. And then Nova Reel is where your input is text, but your output is going to be video. Um, at this moment in time, we have, we generate six seconds of video, um, but in the first quarter of this year, we're gonna extend that to two minutes of video that's gonna be created.
Um, and, and the, the, the, the video's a very high quality again, and one of the things, uh, the model allows you to do is also to make sure that you've got control over the camera. So you can actually, um, when you a prompt, you can specify what the camera should be doing, like whether it should be zooming in or you know, whether it should dolly, uh, and so, so on. You can give it instructions as to how the camera should move in the video.
Um, and then we've also pre announced a few more models. Um, uh, in the first quarter of this year, we're gonna launch a speech to speech model where the model is going to be able to understand streaming speech input in natural language. It's going to be able to interpret verbal and nonverbal cues.
And by nonverbal cues, I mean things like tone and cadence. And then it's also gonna be able to deliver, you know, natural human-like back and forth interactions with very low latency. So this is gonna be very useful to develop conversational applications.
Um, and then the last model that we also pre announced is a, any, do any model, uh, which is a multimodal to multimodals, which means that, you know, the inputs, uh, could be text, video, audio or images. And the output could also be text, video, audio, images. And so this is gonna be a unique thing.
There's nothing out there, uh, that, that is, um, um, any to any model or a multimodality to multimodal to multimodal model. Um, so these are all of the models that we announced with Nova, and then we're super excited to see what customers do with them. So as we kind of pull that together though, will people also be building smaller language models off of these foundational models?
And how do you envision all that getting operationalized? Yeah, I'm not sure. I think the, I don't think the smaller models are gonna now become more popular.
Um, the, the larger the, I'm sure the models are gonna get more sophisticated as we go along because they've not yet reached a point where they stop, uh, learning and becoming more, uh, smarter and better and more sophisticated. Um, so you're gonna see larger and larger models still being developed. But I do think, um, but it comes to at scale production applications, there is a spot for smaller models to play a role.
Um, and I think where things get very interesting is when you have techniques like distillation. I'm not sure if you're familiar with the, the process of distillation, but what it does is essentially you can have the sophistication of some of these, uh, larger models, but you can have it at the cost and latency of the smaller models. You can distill the knowledge for a specific use case from the larger model to the smaller models.
So you're gonna start to see these kinds of things emerging. And by the way, we do support distillation on Amazon Bedrock or Ready, and the Nova models also support distillation from the larger ones to the smaller ones in their, in the family of Nova models. Uh, so you're gonna start to see those kinds of things.
I think there's room for both. You don't always necessarily need, um, the most sophisticated model for every single task, uh, that you may want to perform. And so, um, uh, uh, we've always said this, like, there's, the choice is super important, even within a single family or category, like text models, text out models, that output text, it's super important to have that choice because customers always have to trade off between, you know, latency, accuracy and, and cost.
So how smart will these LLMs get as we go along? 'cause it seems like, um, they're getting more reasoning capabilities and it seems to have to do with the, the size of the context window and how much memory they have access to, uh, how is, from what can we expect going forward? I mean, the simple answer to that question is it's going to be a continuum.
Um, models are just gonna getting, they're gonna start getting more smarter, more sophisticated. They'll be able to tackle more complex problems. They'll be able to do multi-step reasoning better.
And while all of that is happening, the costs of these models, it's gonna get lower. And so you're gonna see, um, uh, we, you pe customers don't have to think a lot about where do I actually apply these models, which use cases can I apply them to, and which ones I don't? 'cause the cost of inference will go down dramatically in the coming months.
And how will that be achieved? Because everybody right now has it in their head that generative AI equals GPUs. But, um, is there another way of thinking about all that?
Yeah, I, I, I think this is, again, a thing that I think Amazon does really well, is the reason why we are able to take technologies like AI and apply it at scale to real world business problems is because we think end to end. Um, and so if you, if you look at all of the things that we've talked about when we, when we launched Amazon Bedrock, and we've developed it for the last couple of years, now we look end to end. And the first, I I wanna give you some examples here.
Um, you look at the hardware pieces. So we have custom chips, cranium, and in French here, uh, which were specifically created for generative ai. And these chips have 50% better price performance than, you know, other EC2 instances that are comparable.
Um, so again, right off the bat, you get some advantages with, uh, with the cost here. Um, and so that's one piece of the puzzle that you need to have when you're thinking end to end. Then we've got the choice of all these different models.
Um, and so there are different categories of models, as I said earlier, uh, models that are text, uh, um, that output text models that output, um, images and models that output videos. But even within each family, you've got all kinds of combinations of these trade-offs, cost accuracy and latency trade-off. And then the final piece of the puzzle is workflows to help you, um, reduce or inference costs.
I mentioned distillation as one of the workflows that allows you to do that, um, to, to get the sophistication of the larger model, but at the cost and latency of the, of the smaller models. But there are other techniques, uh, for example, at runtime, when you're getting a query, uh, for a particular application, not all of the, not all the, all of the queries require you to route it to the same, most sophisticated and, and, uh, you know, costly model. Uh, you could, if the query is simple enough, you could be routing it to a, um, a smaller model in that same family.
Um, and this way you can save costs at runtime as well. So there are techniques like model, uh, routing, prompt routing to the appropriate model that's gonna help here. Um, and, and so all of these put together is what's ultimately gonna drive the costs down.
Um, and of course, we are gonna keep innovating on the chips, um, and we're gonna make sure we implement optimization techniques with these models. Um, and so there's lots of, uh, inference optimization techniques you could be applying to these models, like going from floating point to integer, um, uh, competitions and so, so on. It would seem to me last year was kind of the year of experimentation and more organizations are trying to figure out how to operationalize gen AI these days.
And some of that has to do with probabilistic versus deterministic processes, et cetera, et cetera. But what do you see in organizations that are doing this well doing, or they just picking a handful of projects to go forward? Or is there some secret sauce in the process here?
Yeah, I, I think the first step is really look at the business problem, then the technology. I would always start with the business problem first and understand what the value of the business problem is. Um, and then the next step for me would be to, you know, determine whether generative AI is the right solution for that particular business problem.
In many cases, it will be, uh, and I think going forward, um, it's go, it has generative AI has the potential to, um, pretty much, you know, transform every function in every business in every industry. We strongly believe that. Um, and then the third step is look at the value, the business value and compare it to the cost of, you know, calling that model through an API every single time for inference.
And you're gonna be doing it millions and millions of time if, if your application is at scale in production. Um, and so at that point, you've gotta evaluate is the value that you're creating, uh, greater than the cost? And if it's not, uh, you understand at least where your cost needs to be for this to be, uh, a useful, uh, you know, useful thing to do.
Um, and then you can go and figure out, you know, what are the different, different things you have under your, uh, your fingers, you order to optimize and, and get to the cost that you need to. And I've described, uh, many of the techniques that are available and many of the options that are available. So it's, it's then a question of, um, you know, you picking the right options to get to the cost levels that you need to get to.
Right. And now of course, everybody is talking about the rise of Gentech ai, which is directly related to these LLMs and Gen ai. But, um, do I go and build, uh, all these ais myself or all these agents, or do I orchestrate them?
Or how do you see that all kind of coming together in some sort of cohesive fashion? Yeah. Uh, I think a good place to start would be, let's sort of define what agents are.
I know there are lots of different definitions out there, but I think there's a, I I look at it a certain way and, and here's how I would define an agent. Uh, an agent is a digital worker. When you give it a bunch of tools and resources, and then you ask it to solve a problem, and it's a problem that you've not trained it on, and it's not a, it's a problem that you've not broken, it broken up the problem into pieces for it to solve, it's just a new problem every day.
And the only assumption is that the problem can be solved with the tools and resources that the agent has access to. And then the agent goes about using the tools and resources it has in order to go and solve the problem. And today, you may give it one problem to me tomorrow, you may give it another problem.
And what the agent is gonna do with the digital worker, it's gonna automate the piece of work using the tools and resources. And when I say tools and resources, here's what I mean. Um, resources could be things like, you know, policies, manuals, things like that, um, documentation and so on.
And, um, tools could be things like, um, a computer or, uh, a a piece of software. Um, it could be, um, uh, Excel, it could be a browser with internet access. And tools can also be a whole bunch of backend APIs that you may have, and each of the APIs does something.
So the idea is to, um, give the digital worker or the agent all of these tools and resources and ask it to solve a problem. Uh, and the agent's gonna now figure out, okay, which tool do I use first? And which, and what, what do, what does this resource tell me about this problem I'm being given?
And then it's gonna figure out a path forward to try and solve the problem. And I'm sure the agent at some point could get stuck, and then it knows how to backtrack and, you know, try a different path using the tools in a different order. And ultimately, it then goes and solves the problem and automates the work.
That's what is a true agent for me. Uh, there are lots of solutions out there that just take a, a large language model and then store a prompt with it and, and then call that an agent. That's not an agent for me, an agent is what I described.
And the way that customers can go about building agents is, is using, like Amazon Bedrock, for example, I mentioned we have, uh, the best models, the broader selection of models on Bedrock, and then I also mentioned that we have a whole bunch of workflows, um, on Bedrock. One of the workflows is agentic workflow workflows support. And what customers can do is create an agent on Bedrock, you powered by, you know, any of the models that are there on Bedrock.
Um, and then they can, again, configure. Here are the, the, they can provide the, the agent a bunch of documents, um, they can describe APIs they may have internally. 5 Sonnet V two has a certain capability that none of the other models out there have, which is computer use.
So given a virtual computer and all of the software and the computer, it's, it's gonna look at, you know, the images of that computer and figure out how to call the tools, how to browse the web, you know, how to go to certain websites, how to do research and come back and things like that. So we have support for Agentic workflows on Bedrock. Um, and I'm gonna talk through some of the, some of the capabilities we added to the agents' workflows on Bedrock.
The number one thing I think agents are gonna need is memory. Um, so this was something we added about a year and a half ago, uh, where, uh, if you give agents memory, so assume there's a, uh, there's an agent that you know, uh, interacts regularly with a certain customer, and then let's say the customer comes back at a later point in time, if, if the agent had access to the memory of all of the previous conversations, it's gonna be able to support that customer really well in the current interaction, right? And so we gave agents on bedrock, long-term memory, a way for developers to configure the memory and make sure that they could build hyper-personalized applications as a result of that.
Um, we also created tools, um, like we created tools like code, um, tools where the agent can write the code and then run it in a sandbox, and then actually come back with the answer. And, and I'll give you a concrete example of where something like that is useful. Um, let's assume that you have a bunch of money and you wanna invest in real estate, and all you have are property sales, like house a sales of property across the United States.
What you're trying to do is figuring out, uh, you're trying to figure out what's the, what are the best zip codes in the United States for me to invest in real estate? And your input, as I said, is just sales listings like you would find on Zillow or some of those websites. Right now, an agent can't simply give you an answer to that question as to where the best return on investment is.
Uh, um, uh, by zip code across the United States, it has to do a whole bunch of computation. It has to figure out what's the, for every sale, uh, what was the purchase price and what the return was, then it needs to kind of sort, um, it needs to, um, you know, um, uh, a group by zip code, and then it needs to sort it. Um, and then finally, it's gonna be able to tell you, these are the zip codes where, um, the, the biggest return was over the last three years or five years, or whatever the input query was, right?
And to do that, the agent needs a place to write the code and execute the code in a sandbox and be able to come back with the result. And, and so that's a tool, the code interpretation tool that we gave agents on Bedrock. So any agent you create on Bedrock has access to that tool.
Of course, you can also bring custom tools that you write and then use that together with it. And then the biggest thing that we did was at the re at reinvent, um, um, just, just last month where we, we launched multi-agent collaboration. Um, and this is very interesting.
Uh, I think where this is going, the idea is can you, if you had a bunch of different agents that you've created, can you get them to work together to come up with a result, um, for a certain problem that's much better than, you know, any of those agents would've done independently? Can these agents work as a team in order to create, um, a, a much better result? Um, a again, I'm gonna explain this with a, with a concrete example.
Um, and this is actually a real world example. Um, Moody's is one of our customers that have, that have used this capability. Um, the idea is the following, um, assume you're an investment firm like Moody's, and you wanna now determine where do I, do I invest in a certain public company?
And let's just say for the example, for the, for the purposes of this discussion, the company's a coffee company. So the question is, do we, uh, what is the, what are the risks associated? What are the risks in investing in this company?
That's the question to answer. Now, typically, what would happen, uh, in a, in an investment firm like that is that they would call in an analyst who's maybe very good at reading financial statements, and then they'll tell that analyst, okay, go look up for this coffee company. Go look up the financial, the last quarterly, uh, earnings release, and then come back with a feel for the health of the company in terms of the debt, in terms of the cash flow, the free cash flow, and so on.
Um, so that would be one task that, you know, an analyst would get, but there would be another analyst to the company that's maybe very good at, um, you know, doing, uh, analysis on macroeconomic trends on various topics. And that agent may have access to, you know, um, uh, cutting edge, uh, trend research. Um, maybe they subscribe to, uh, certain consultants that publish, uh, trends across the globe on many, many topics.
So that's a completely separate, um, uh, person that would go and then do that research. Now, imagine each of those, uh, each of each of those jobs are given to now digital workers or agents on Bedrock. Uh, one is a financial statement analyzer agent, and the other one is a macroeconomic trend, uh, analyzer agent.
Um, and the naive thing you could do is simply get them to go do their work using the tools and resources they have, and then come back and then just simply concatenate the results. But I think what you're looking for here is can they produce a better result together? And that's kinda where this multi-agent collaboration thing comes into play.
Um, so you have now a supervisor agent that manages these two sub-agents, and what it's gonna do, the supervisor is gonna do, is based on what one of the agents comes back with, it's gonna create new tasks for the other agent. So in this case, the, the, the agent that analyzes the financial statements, it's gonna come back and say, okay, this coffee company's growing really great in the US and in China. Um, and what now the supervisor can do is instruct this, the macroeconomic trend analyzer agent, to go and look at coffee trends specifically in, in China, for example.
And, and that agent may come back now with much more information about that particular geography and say, okay, maybe the consumer, uh, in that geography isn't necessarily, you know, they're, they're starting to drink more coffee at home and not actually wanting to spend, spend it, uh, out there. And so now you can see that the ulti, the final report that, that these, this team can put together is gonna be richer. Um, and, and it, it's gonna be much better than what any of those agents could have independently done.
So we did launch this, uh, you know, multi-agent collaboration framework, um, on Amazon Bedrock, um, and you are gonna start to see in terms of what's going to, uh, happen in the future, uh, and I'm, I'm very familiar with all of the research that's going on in that space. You're gonna look at teams of agents and how they're organized, uh, whether they're distributed teams, uh, with no leader or they've got different, uh, groups with different leaders. And so it, it's gonna be fascinating to see how the space emerges.
Um, but I do believe that agents is a great abstraction for generative AI in general because different companies can build agents completely independently. Um, they can expose their business logic through these agents, uh, while protecting their ip, uh, behind the scenes. And as part of that, the agents can check each other's work, right?
So I could essentially have one taking care of governance and security while the other one is doing the primary task, right? Absolutely. And, and I, I can give you also another example.
If you look at agents in terms of coding, right? Um, if, if you look at Amazon Queue, which is our queue developer is our coding assistant, uh, that's actually implemented as a bunch of agents. And so we have a, for example, a code review agent.
You know, it helps you perform code reviews when one developer writes the code or maybe one agent writes the code, the other agent can actually check the quality of that code that's written and actually do a review on the quality of it. And then the feedback goes back to the first one, and the, the using the feedback, uh, the code quality can actually be improved. Um, so you're gonna see all kinds of those, uh, those kinds of things happening, for sure.
Yeah. All right, folks. You heard it here.
Nova's just part of a larger overall shift to more agents and more advanced workflows and challenge. Now it's gregger out how to manage it all in a way that we can afford. Hey, bassy, thanks for being on the show.
Thanks For having me, Mike. All right. And thank you all for watching the latest episode of Techstrong AI video.
You can find this episode, others on our website. Until then, we'll see you next time.