How to Control Cloud Costs for GenAI Tools and Reinvest for Growth – DevOps Dialogues EP15
Can AI be used to control cloud expenses? Host Mitch Ashley is joined by DoiT‘s Eduardo Mota and Weaviate‘s Jobi George on this episode of DevOps Dialogues, for a conversation on how companies can manage their cloud expenditures in the context of GenAI tool utilization and strategize reinvestment for growth.
Transcript
Hey everybody. Mitch Ashley here with DevOps Dialogues. I'm VP and practice lead of DevOps and application development with the future.
And Group DevOps dialogue is all about conversations about creating software in the era of ai, cloud security, all kinds of aspects of the we have to deal with of creating the best kinds of applications that deliver the results that our customers, our business, our partners, are gonna be most fulfilled from and receive the greatest outcomes. So we're gonna be talking about Gen AI and doing AI projects and, uh, some tips and information that will help you all in doing and, uh, taking on those. And maybe you've already started a, a project and looking for some good advice or information, some insights.
We're gonna do that by the two folks that are joining me on the podcast today. First is Ed Eduardo from, uh, do It. He's Senior cloud architect and Joby George, who's global head of partnerships at Wvi.
Eight. Welcome gentlemen. Eduardo, would you introduce yourself, tell us a bit more about you and also, uh, tell us about Do It?
Absolutely. Thank you. Well, um, I'm in Canada, so the weather over here is very cold.
Sorry to get cold. I'm a senior cloud data architect, and I've been working with AI mls, uh, for quite a few years now, uh, before Gene ai and at Do It, what we do is we'll help our customers really unlock the true value of the cloud. And so we help them strategize, create architectures, and really become efficient when using the cloud.
Fantastic. Joby, I'm Joby George. I run the global partnership at VVI eight and manage all our cloud technology and system integrator partners.
VVI eight is an open source vector database, uh, for building, uh, AI native applications. Without developer first approach, developers are able to go and use LMS and embedding providers and frameworks to build their AI native applications. And we, and they are using it to build applications like hybrid search, semantic search rag applications, or some new things like recommender systems or generative feedback loops.
And now the agent rags, we have around a thousand plus customers like Cisco, Morningstar Bunk, who are building all kind of gen AI and search applications. I'm glad to be here. Thank you for inviting me.
Excellent. Thank you both for being here. Thanks to your companies for being here as well.
You know, let, let's start out about talking first about, uh, kind of taking on gen AI projects. Uh, it's an exciting time to be in the industry. There's, you know, technology's changing every day.
There's announcements every, seems like every week, every day, whether it's Microsoft or AWS or you name it, you know, a new model coming out. There's so many, uh, kind of technologies to learn. I know from my own experience leading software teams, you don't always want to venture into this alone, um, for fear of it becoming too much of a science project rather than really getting value out of a project.
I'd love to hear your thoughts. Um, maybe Eduardo, if you wanna start out first, since you work with so many customers building these kind of applications, um, what, what are some of the best ideas about how to get started and how to decide what kinds of applications are best to take on with generative ai? Yeah, I mean, we, we are, there is so much going on.
There is just so many things as you were mentioning, um, and with an example of AWS there is so many services available to our customers. So a lot of the times we hear like, Hey, I wanna use GN ai. I really wanna leverage what is possible with the technology, but I don't know where to start.
And really it's starting from the place where, where is a good use case for personalizing the customer's journey? How can you really hyper-personalized that experience to get the most value? And everybody's collecting all this data, it's just very powerful data that allows us to get into the nitty gritty of that personalization.
Now with Gene ai, we are able to unlock that. Before we were creating personas, we were creating buckets that will put a customer in a certain category, but with Gene ai, you can really go dive into that custom personal experience. And if we take a, like I was mentioning AWS as an example, they have great service called Bedrock, where you don't have to create an infrastructure, you can start utilizing it with, uh, on demand and send texts.
You only pay for the amount of tokens or the amount of words that you send out and of you go. And it's a really great way to be able to start testing what is possible. Now, I do have to, uh, whenever I talk with my customers, I do have to mention that HN AI is not a silver bullet.
Like nothing in nothing in technology, right? So it is great at certain things. And when you are doing this test with Bedrock or with any other cloud provider, make sure that you test that limit.
Where does it start breaking? Because that's the iteration where you are gonna be able to start making progress and start adjusting things. So it fits your use case.
But, uh, to start with is think enough about that hyper personalized customer experience that you wanna deliver and then start testing with the, the managed services available to you in the cloud. So that's kind of like where I'll start. And That's, that's a great, great point and um, I'll, I'll just add to that, uh, that itself, like, uh, so for example, VV eight is one of the companies which integrate with Bedrock.
I think, uh, just to roll back a bit, right, uh, generative ai, this is a totally new way. It almost feels like, uh, almost like the internet browser coming into the market type of a landscape. And everybody wants to build all of these application, everybody wants to build an application, which looks like a chat g PT or equivalent of those applications.
But, uh, they, they might not have the skill gap. They might not have the real understanding of all the use cases which they can go and tackle. So kind of taking a more selective approach to how and what they can build and what is the platform kind of fits in for them as they're building these applications become super, super critical.
But the possibilities are endless. Uh, as Eduardo said, there are all kinds of applications, paper up building, starting from simple search type of applications, rag applications. We always had a search or a database application architectures to build these applications.
But machine learning and the generated AI has totally changed what the information retrieval and the quality of the information which is being retrieved from dramatically. So now we are able to get to relevant information. Eduardo talked about personalization, so kind of being able to do that, personalized apps, uh, so we have an approach where we call it as a generative feedback loop, which allows you to build these personalized application at scale.
And, and that's where the motion is. So some of the things I would kind of like highlight is that be very careful on what the use cases you're trying to, uh, approach. Do you have the skills to kind of do those and kind of quantify and, uh, approach it in a way that you can de uh, deploy and get ROI on these applications quickly?
Go ahead. Sorry. No, go, go Ahead.
I was gonna like, you know, you mentioned a lot of great things and you start getting a sense of it, uh, when you are starting to test because you get a sense of where the gap is on the skills and where the gap is on the data. 'cause the data is a great like model and the oil that will get all these going, right? And for, for all, for every company out there, you have access to the same models.
You have access to better, you have access to all these other models for other cloud providers, but the differentiator becomes the data. And so when you start testing, you'll be able to start identifying, oh, I can I start using this other data this way or, or this other way. So yeah, I think those are great points.
Yeah, I was just gonna bring up the data, right? It's really the center of the universe is also oftentimes is deciding where do I put my turn of AI application, where, where in the cloud, et cetera. Where, where do we need that data?
Uh, you know, one of the w with so many things happening in, in the a AI fields, particularly generative ai, you know, we heard a lot about rag people, augmented generation. Let's talk a little bit about what that is and why, why we use that with the importance of it. And then also, of course, we hear a lot about agents, right?
Being able to, uh, create your own agents and a lot of different capabilities that are being launched in the market. I suspect we'll see a lot more over time. Um, anybody wanna take on rag?
I think maybe, um, Joby, you were kind of mentioning that earlier. Yes. Um, so RAG is a very logical, uh, evolution from a search based application to kind of building a retrieval based approach to building application.
So think of a native rag pipeline consists of two parts. One is a retrieval component, typically basically composed of an embedding model and a vector database, which is VVI eight is one of the providers in that. And then there is a generative component, which is the LLM or the models which you can use.
You can use foundation models, you can use your own custom models, whatever you want. So at the inference time, or basically what you're trying to kind kind of, when you do a query, what you run try to do is do a similarity search, try to find an index of documents, which is a subset of the documents which are more closer to the kind of curry you are you are trying to put in there, which allows you to retrieve the most similar document and provide it as a context to the LMS to build, uh, to build their response. A generative response back to you.
This is fundamentally different than just doing a searching or a sickle query type of an approach. Now you're able to build some interesting applications on top of it. You are able to take the plain old chat bots, think of it as like four or five years ago.
And now they are really looking like they're really answering your questions in a more relevant way. You're also hitting information which is very relevant to the context because you have actually done that similarity search and that has provided that context, which allows you to now provide a response back, which is very relevant to the information you're trying to retrieve at. So that's a rag approach, and I see that the RAG is now kind of transforming to what we call as an agent rag, where you are taking the, both the LLM and the database and now adding some kind of a memory and a workflow tools, uh, as an elements to it.
So it's a very logical, uh, evolution of the generative AI applications from rack to rags and to eventual, uh, more complex workflow based applications. And to add to, and to add to this is that there is a race to be able to identify the best content for the question or the task at hand. And we need to, like, organizations are looking for ways to optimize that search.
And so databases like Aviate, uh, help allow the search to be more efficient. The longer the text that you send to the, to the model, the longer it will take, the higher the cost will be, and the more noise you will introduce into, into the inference. And so there is a lot of repercussions about that.
Um, and so we are trying to identify what is the right amount of data, the minimum amount of data required to answer this task. And there is a huge field that is going into that. Now, on the agent side, there is also great, uh, work being done around that.
And in a way I'm like, I don't like the word agent too much. I like the word tools better than agent just because it is more like a tool that LLM can use to retrieve data or to do other actions. So there is this combination of like vectors and embeddings that you have in a vector database, but also a lot of organizations have data, data in relational database like by SQL or Postgres.
And so they wanna access this data as well, right? So it's another way of being able to use tools to access data, get the specific points that they want, and enhance the, the prompt that is being sent to the model and make it as short as possible. So that's kinda like the sweet spot that everybody's looking at, how can I get there?
And these databases, vector databases are making a big difference when when doing that. You know, a couple other points too, uh, rag is a way of, uh, leveraging data as part of kind of the front end to the model without exposing the model to your data or vice versa of exposing your, your data to the model and how you can augment or add to that additional content. And to your point, SQL databases, it could be documents.
There's all kinds of different, um, uh, points of which we can pull data from. And you mentioned vector databases, uh, Joby, one of the ways we can kind of do high performance retrieval of information out of, out of, uh, generative AI models, correct? Yeah, no, that, that's absolutely right.
And um, like one of the point I want to kind of come back to is like around the data silos and the data we are talking about. Like I come from the data space. I, I lived through the big data wave and kind of see this kind of a wave come in.
So, um, and I think we are in a very early ginning of the gen ai, uh, genai landscape from that point of view. There also was a promise of going and taking the unstructured data and making it available relevant for you to kind of go and build applications, insights and analytics on top of it. I think we are kind of doing the same iterative loop in that that way.
But now the game has changed a lot, uh, from the point of view of like earlier it was a text-based, it was structured data, which you can actually go out now by taking each individual data element. So be it data sets tables have the metadata and the schema and some context which you can capture in a structured data. Taking PDFs, taking images, being able to take all of those multimodal data sets and building, bringing it all into the same vector space and then kind of doing a relevant data search makes it super, um, super intuitive and super useful for creating some very natural looking applications, uh, which are very different than a typical database type of applications.
I think that is the big jump which we are seeing. And again, it's a journey. Like if you look at the progress around like the new things which have come around Cobert Co Poly and some of the stuff implementations we have done around acon, this is a journey and I think we are just in the early leaning of it that where the lot of innovations will come in, which will make it super important, uh, for a Vector database to be a core component to building some of those applications.
And that's how we tell our story, which is like we talk about AI native stack, and when we talk about AI native stack, basically LLM is the core foundation and a vector database so that you can actually build the applications on top of it. And then the underlying layer, the architecture has to be right, like it has to be containerized, it has to be scalable in the cloud. Like AWS being able to kind of get that scaling is critical part as your data sets grow and you bring in all the data into a vector space point of view.
Very good. You know, one of the things with new, working with the new technology is kind of keeping costs in mind, right? And these are new services that we might be using from a cloud hyperscaler like AWS, um, but we don't have experience.
I mean, we had that happen early in the cloud, right? What, what does our cost profile look like as we start to consume these resources? Um, one, one of the things to cons, what are some of the things to consider, uh, in terms of costs that either anticipate learning from from experts like yourself or that you'll experience once you start to develop an application and see what its performance and consumption profile looks like?
Um, Eduardo I also like test a lot the rate fast and fail fast. I mean, it seems like, uh, there is a lot of entrepreneurs out there that are doing this to be able to get business up and running. Uh, when it comes to gene AI and these workloads is the same thing.
There is no one solution fits all. And so you need to iterate fast and understand what are the limits of the technology, right? Sizing these workloads is very, very important to reduce costs.
And that's true across entire cloud, not only on gen ai, but every workload, uh, is not simply by going to the cloud means that your bill is gonna be cheaper. Uh, rather you need to understand what the architecture is, make important choices, right? So for example, on the data side that we've been talking about, we can easily put a lot of data in S3, but if you don't have the right format and strategy there, you can be paying 90% more or a hundred percent more to access that data that if you have the right strategy, that you have the right format and know how to access these data, right?
And that has an implication not only on the cost, but on the latency of the model, which ultimately is your paying for how long that model is running. So all of this is, has a effect if you do not understand the building blocks to optimize every layer of the workload. Um, uh, and that is goes with tools as well.
Like we were talking about rack and agents, you can create a lot of tools for the model to access many different things. Great, fantastic. But it doesn't have to, does it needed in order to do that?
And just a quick tip for those that are listening and you are working with agents, when you start adding four or more agents to the model, there is gonna be a diminishing return on the quality of which tool is being utilized. And there is a tendency to use the first tool presented to the model. So if you have four more, the fifth tool is not being utilized and you are having this code there for nothing or paying for it for nothing.
So all these architectures we need to revise. That's something I love to do. I love to drive it with my customers and say, okay, what is it that you're trying to accomplish?
How do we architect this to fit your current estate, reduce your bill, and be able to get you in a path where you are able to build for the future? So, I mean, it's a long answer because there is no really super bullet to all of this, but I mean, Joey, you, you, you, you experience that, I mean, you're in the database space, right? Yeah, I I can certainly add add to that, right?
And I kind of like break it down into two parts. Uh, one is on the like development and developers and others on the deployment side, right? So if you look at from the developer and the AE and the development side of the things, everybody wants to build charge gpi, everybody wants to build the most complex generative AI applications.
And hey, choosing the tool, being able to look at the use cases and quantifying the use cases to the segment so that, what is the ROI going to look like? Uh, because nothing comes for free, as Eduardo said. Uh, you can do as fancy applications as you want.
Uh, if you're doing a e-commerce application and you're trying to put out a personalized shopping cart and you have certain SLA on QPS for you to do that serving, you cannot do all the aspects of the generated AI to kind of bring that applications to life. So being able to kind of do that trade offs, like very standard practices across any workloads, those are all the standard practices which apply costs do add up very fast. And in the sense that this technology is evolving, the, it's in that stage of the eco, uh, stage of maturation where a lot of innovation happening.
Optimization is still kind of coming in as part of the logical maturation curve. We have started focusing a lot around cost versus performance. So a lot of the features we have introduced things like multi-tenancy where if you are not you, you can break your workload into multiple tenants, then you can offload some of those tenants down, you can offload them all the way to the disk and save cost.
So be being aware of the cost consciousness about these deployment models is super critical. Coming back to the developer side, we kind of look at it like we built it on Kubernetes. We built it on AWS uh, basically to provide that choice where customers can start on a multi-tenant SaaS type of an offering with like $25 a month type of an offering and build an experiment.
But as they go bigger, they can actually go and get to more of a managed hosting. So we basically, we are able to take the same, same environment, move them into a single tenant offering from our end, and or they could start with a bring your own cloud kind of an offering and experiment with that. Most funny thing is that as people are deploying these, uh, use cases, what you're seeing is that their cost and ability to manage these deployments is, is not there because again, evolving ecosystem that they're coming back to, uh, us in some form to go and host it for them and run it because that is a lot more cost efficient than them trying to figure out and going and running it.
So there are a lot of these pieces coming in because of the state of the market it is. That's really good point, because oftentimes the architectural decisions you make the service not just the services you choose to make. You can create very high consumption, you know, types of applications, particularly in AI and not realize you're racking up a lot of costs.
And that's of course the fastest way to get your second projects, uh, canceled is have the first one go way, way outta out of control. So you wanna think about those costs upfront, right? Even if you don't know what they are yet is, is understanding what's driving costs and maybe you can make some adjustments in your design, your architecture along the way to better utilize the money that you can be spending to operate these kind of applications.
And I think that's where the cloud really also gives another benefit. All these managed services where you don't have to spend human hours setting up an infrastructure, but rather start getting a sense of how much things are gonna cost at a smaller scale. And then you can start seeing, okay, if I send this product, if I get this output, this is how much it's gonna cost me.
And starting to extrapolate on putting thresholds to be able to get there. Let's talk a little bit about, um, I mentioned earlier some of my experiences when working with new technologies is it's fun to learn it yourself, but that isn't always the most effective way to get it done. I mean, you make, make a lot of the same mistakes that other people have already made.
Some of the learnings you can gain from working with others. Let's talk about the value of working with, uh, organizations like, uh, yourselves, people that have been down this path multiple times, maybe worked in different, uh, different scenarios, different kinds of applications, and how that can help accelerate, uh, maybe a new customer that you're working with to get not only their, uh, their project delivered on time or successfully, but accelerate their own learning. Absolutely.
I mean, for, for me, I've been in the trenches, right? I was a DevOps engineer before trying to figure out AWS and trying to be the most optimized way to it. Spend hours going through documentation, trying to figure out, and at the end of the day I'm like, am I doing this right?
And that's what I love about working with DOT and myself here because we are a group of 300 engineers across the glove that have expertise in everything around AWS And so we collaborate with one another to be able to get the right answer to the customer, understand the requirement and what the customer is trying to solve for trying to build for, and then being able to cut all the noise of what is not necessary and be able to hyper focus on like, this is the best architecture that you can build today that will grow with you, with your business and will keep the cost in balance. And then that will just help the, the customer reinve reinvest all those savings back into the cloud and growing the business. So for us, do it, uh, is an extension of every team, every engineer team, and every organization.
That's True. And I can add to that, like, um, I think, uh, being able to, like we, we are creating a distributed new database and, and we are an early stage where a lot of experimentation going on and now a lot of them are now going into production. So a lot of like moving parts at this point and being able to have the help across from AWS and others being able to kind of build the up the database in the right way and e out wherever we can, the cost and optimize the performance and cost for the customer is super critical because as, as we were talking before, the cost starts adding up pretty fast and being able to go there and go under the hoods and find out like how you can go and do this particular optimization at the Kubernetes level four S3 would help you to get some of the cost advantages is critical.
We also do things like, uh, based on your QPS, like what is your expected query speeds and things like that, we can then go and find the cheapest way for you to go and run a certain machine configuration. So for example, we run a lot of our production workload for our customers on graviton and, and, uh, a lot of, when, when people talk about generated ai, they talk about GPUs, but I think we are seeing a lot of our actions on the low cost side of things because cost is a big factor. And as more and more deployments starts happening, it'll become a bigger and bigger factor.
What is the performance and cost goes are going to look like and how you deploy things in production, uh, becomes very, very important. Very good. I would just add to that my, my general advice in this particular area, things are changing so fast, so many things are being introduced seems like on a daily basis it's, it's tough to say what's the best technology you use?
'cause it's changing that rapidly. I think you're more, more, uh, on, on the good track by picking the right partners to work with because the technology will change, our understanding will evolve, our learning will certainly increase, um, as the technology changes too. So you wanna kind of pick the right horses to, to be a part of, to ride if you will, not just necessarily the best technology.
'cause they're gonna be a lot of really good things. Not just available today, but coming out real soon. Just, uh, just open up your browser the next day.
I think you'll probably see a new announcement. Well, let's, let's, uh, wrap up. Uh, I'd love to hear a little bit more about, so folks that are interested, wanting to find out more.
Like how can you, how can you help me on the journey that I'm, maybe I'm already down the, down the path on an app or I'm looking to get started in wanting to, uh, kind of accelerate my own learning. Um, what, what are some, what, what's a good way folks can engage you to find out more? Sure.
com. We have a wide range of services available. And just to point out, GA accelerators my favorite program that we offer, so you can check it out to get you started in the in AI journey.
io where we, there are a lot of gateway to a lot of ways. You can get a lot of information on documentation, slack forums. We are open source.
You can go to GitHub, play with the code. If you're of that type or you want to go and basically just get started, go to our VV cloud, uh, you can get started and get 15 days, uh, for a free trial. You can, you can run your sandboxes there and provision your clusters and get started in a few minutes.
Uh, so test it out. That's where You learn get started, right? Jump in.
It's a con. AI is a contact sport. Yeah, get engaged.
Well, thanks to you both, uh, Eduardo and uh, Joby, it's been a pleasure talking with you today. And please folks, check out other respective websites, uh, thanks to both your companies to, uh, do it. And, uh, we, we appreciate having you both on DevOps dialogues.
And thanks to our listeners, we wish you the best on your journey into generative AI and applications and beyond. Take care, everybody.



