Real-Time Apps and Multimodal Databases with Aerospike’s Subbu Iyer
After picking up an additional $109 million in funding, Aerospike CEO Subbu Iyer dives into why a larger shift to real-time applications is driving the need for multimodal databases.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Sabu Aire, who is CEO for Aerospike, and they're fresh off of a investment of $109 million that's on top of what they've previously have managed to recruit from folks.
But we're gonna talk about what's driving that, because they're a pretty big player now in this whole space of graph databases and vector databases, and they're used for slightly different things. And how all this is coming together. Sue Bo, welcome to the show.
Thank you, Mike. A pleasure. What is your assessment of what's driving all this?
I mean, on the one hand, it seems like vector databases are closely tied now to the boom and AI graph databases may be tied more towards, uh, knowledge graphs and visualizations, but it feels like both are being lifted. Are the two related, or are they different use cases, or how does this all come together in your mind? So, Mike, before I answer that question, just a little bit of a background.
So, you know, what we see is, you know, realtime access to data and realtime decisioning is actually now prevalent within every industry. You know, Aerospike was founded on the premise of really making, you know, real time data accessible at high performance at, with any scale of data, right? We started our journey with Ad Tech, but now we have customers in every industry, financial services, telco, e-commerce, retail, gaming, and entertainment, healthcare.
So really what we have seen is the, you know, the adoption and the desire to harness data at any scale in real time. And, you know, we've been deployed in, in a lot of customers grow globally already for several years within their AI and ML pipelines. So, you know, as you mentioned, obviously Vector has become a topic of conversation ever since, you know, chat.
GPT became available last year. Uh, but we've been in deployed in customers across their AI and ML use cases, whether they're fraud analysis, recommendation engines, uh, profile stores, so on and so forth, uh, for the last several years. Years.
So what we see out there is really, you know, a desire to actually act on a lot of data. So if you think about ai, I think AI begets more data because what we hear repeatedly from our customers is that more data actually drives better context and better accuracy to their AI systems and their AI results. We were built from day one to handle data at scale while delivering real-time performance on that data.
So we are, we believe AI is a perfectly suited application for us, if I can call it that. And, you know, we were, we were ready for the AI age. It's just that, you know, over the last year, as you talked about, there's a lot, lot more, you know, desire to actually implement AI within systems.
Mm-hmm. As far as graph is concerned. Uh, the second part of your question, we entered the graph market last year, and, you know, we are, as you correctly pointed out, a knowledge graph that supports both use cases like identity management and fraud.
And, you know, we saw a, you know, opportunity in the market, again, based on customer feedback, who wanted to actually deploy graph at scale. So the existing solutions that were available for graph, uh, really didn't scale much. Uh, so we came out with a solution which can, you know, support billions of vertices and trillions of edges.
Um, so this is the, you know, class of kind of graphs that we are supporting for our customers. And we've got a, got a lot of traction. It's not just the scale, but that size of a graph we can traverse on a multi hop query in, you know, single milliseconds.
That's amazing. You know, when we talk to our customers and we, when we see the market out there with ai, you know, and the desire to actually use more vectors to generate those embeddings, and whether it's within generative AI or predictive AI use cases, actually leverage vectors. We are seeing that happen, uh, as far as the adoption of vector and vector search is concerned.
But we also believe there's a congruence, if you will, where graph and vector can exist independently to power some of these use cases, and they can also come along. So if you talk about a simple example, if you're trying to look for a specific document, let's just say using a vector search, uh, a vector kind of search with vector embeddings is probably the right way to go. But if you're looking for documents which are similar in nature, for example, which is, you know, for lack of a a word, let's just say it's a corpus of similar documents, then you want to use a vector search, which is then augmented with a graph so that, you know, the graph can then actually build relationships and build out similar documents that are available.
So that's how we see this entire field of vector and graph coming together and evolving over the next several years. There's much to unpack there. So let's jump in.
Let's start with real time. So we are clearly seeing that there is more data being processed and analyzed at the point where it is created and consumed, but then it needs to be fed back to something else in near real time. Is this changing the way we think about building our applications?
Because historically everything was very much batch oriented. So, you know, do we have the developer skills to do that? Where are we on this journey?
Absolutely. I mean, you know, some of our, uh, leading edge customers are already doing that. As I mentioned, uh, before, you know, we've been deployed in AI ML pipelines for several years across, you know, uh, our global set of customers.
Um, what we hear repeatedly from prospects and customers we talk to Mike, is the desire to actually use more data in their AI or other kind of applications. And the limiting factor for them is one of two things. One is really scale, which is, you know, they look at systems and solutions, which actually hit a scale wall.
And the second is really the cost. So it becomes very prohibitively expensive for them to actually keep all that data online, if I will. So we actually, once they talk to us and they discover, you know, the art of the possible with Aerospike in that Aerospike can support this infinite scale, whether that's hundreds of gigabytes or even petabytes of data without compromising performance.
And at A TCO that is the lowest in the industry, right? We run on hardware footprint, which is 80% less than anything else that is out there now, that actually opens up a lot more possibilities for them. And as we kind of get into more and more AI applications being built for more context, better accuracy, as I mentioned, they're gonna want, you know, all that data to be available as quickly as possible.
You know, we, we are talking to customers right now who are telling us, Hey, I'm only keeping 20% of my data online because, you know, these same things because it, you know, my platform doesn't scale or it's pretty expensive, but if I have a solution, I would love to keep, you know, if not a hundred percent, at least 80% of my data online. And then it's not static dataset, right? It's also coming in where you are ingesting data.
Let's just say in the case of vectors, there's vectors which are coming in at a rapid rate. So can the platform actually support, you know, high throughput ingestion and then make it available, build the index of the vector as this new data comes in and have it available to feed the AI application or the decisioning systems? These are the kind of, you know, key kind of considerations that our customers have.
So to your question, absolutely, if they could find systems like Aerospike, which can actually help them process data online in real time, they would do a lot more online. And, you know, several of our customers have done exactly that. They've gone from nightly kind of decisioning to first going kind of, you know, several hours a day.
And then now they're, you know, some of our customers are doing it every few minutes. That's just changed the trajectory on their business. Do we have the skills to manage those types of databases?
'cause historically we've had relational databases and then document databases. Um, a lot of folks I talk to, you know, vector databases have been around for a while, but until AI came along, they had no idea what a vector database was. Do we have the folks that are required to manage that?
Is that the, a new type of DBA? Is it a data engineering team? Where are the skills?
Um, so, you know, skills, uh, we see, you know, two, two kind of, um, kind of areas of, of investment here. One is, you know, larger organizations do seem to have the talent and the skills, and then other organizations rely on partners or look to us for the expertise so that they can actually onboard and train their workforce, uh, around these skill sets. Along with that, you know, it's our job as vendors and we are investing a lot into this, which is making it easier to manage these databases and also help the developers build easily.
So let me give you two examples. We are investing a lot in observability and management. So as you know, more data comes online and is managed under Aerospike.
We wanna make it easy to really manage that database, you know, whether you're scaling up, scaling out, you know, whether you're trying to understand if something actually pops up as an issue, we should be able to be proactive, um, in, in terms of alerting the, uh, user, uh, so that they can take the appropriate action. So that's one area that we can make life simpler for the operators, if I can kind of call it that. The second part is really the builders and the developers.
So we are investing a lot in terms of, you know, making it easier for developers to find resources. com, there's tutorials, there's sandbox sample applications. We have a community forum, which is actively actually responded to by our team.
And then we also are making it easier from an API perspective, so more simplified APIs, you know, using frameworks like Spring Data if you're a Java developer, uh, using Link Pad if you're ANet developer. So that, you know, you know, our graph solution supports gremlin as a query language. So we are trying to get it closer to the language of the developer and have simpler APIs.
So for both, for the builder and for the operator, our mission is to make life easier. And, you know, as simple as possible, Might we not use generative AI to make it simpler to manage the databases that we're using to drive ai? Absolutely.
So, you know, we are investing in efforts, uh, exactly to your point so that, um, you know, we don't, you know, we want to give the operator the control, whether they want, you know, decisions to be made automatically on their behalf or based on certain policies, or we just want human intervention so that we alert the human. But you're absolutely right. I think, you know, part of kind of what we are gonna invest and build out is exactly what you talked about so that there is, you know, uh, self remediation in certain cases and or alerting, which is propped by ai, uh, which is, you know, built into the product.
How do you see the way IT teams are organized kind of evolving? And I'm asking the question because in this day and age, we'll see data scientists hanging out with developers and DevOps, and then I've gotta talk to A DBA and there's a data engineering folk, and there's probably a couple of cybersecurity people throwing into the mix as well. Um, is there a different way of thinking about managing all those things more cohesively?
'cause right now we just had a lot of fiefdoms over the years. Um, we see two models evolving there. You know, some of our, our customers actually have what they call center of excellence and a shared services team that stands up and manages kind of the database and infrastructure services and have makes them available, uh, internally within the lines of business and through the developers.
Um, other organizations, maybe, you know, medium sized to smaller sized organizations have, you know, all of that actually built together into a singular app team, if you'll, so the person who owns the app owns not just the build face of it, but also the deploy and the operate and manage phases of it. Maybe different people and different roles, but it's all part of the same team. So we see two models kind of, um, emerging, um, where, you know, when, when you have multiple lines of businesses, mo you know, larger organizations, they tend to go for a center of excellence or a shared services model.
On the infrastructure side, You guys have garnered this latest investment round. Are there any priorities for that kind of allocation here? I mean, what's top of mind for you?
Yeah, so I mean, I think we've been, um, you know, incredibly lucky Mike to actually attract, um, you know, attention and, you know, requests from a bunch of, uh, investors wanting to actually join us in the next part of our journey here. Uh, we decided to go with sumero. Uh, we love the team.
They got exactly what we are trying to do and our vision for the future. Um, so, you know, really thrilled to work with the Sumero team. George Kfa, uh, who's, who's been, you know, in the industry for a long time, has joined our board from the Sumer team.
He is a co-founder and managing director. So, you know, personally I'm thrilled, uh, to be working with, uh, George. Um, the core areas of investment are r and d, primarily around ai.
So we talked about Vector, we talked about graph. We wanna make it where the Vector database and the vector search from Aerospike is natively integrated to the workflow automation and the, you know, models that are available on the different cloud vendors, whether it's A-W-S-G-C-P or uh, uh, Azure, uh, so that, you know, our customers don't have to stitch these things together on their own. And, you know, it's seamlessly integrated.
So there is a lot more investment around ai. Uh, the second area is cloud. We released, uh, our database as a service product last year in Q4, and we continue to build out our entire cloud portfolio.
So, you know, we are investing in that. Um, but, and at the outside of r and d, um, you know, it's go to market. So, you know, we are trying to expand our go-to market efforts, both in term of direct sales, um, and also channel, uh, so that we have more reach and coverage across the globe.
So, you know, r and d and go to market are the two primary, um, areas of investment. Seems like there's this ongoing debate about whether or not I need a dedicated database for a particular use case, or if I can extend, say, my existing, uh, relational database to other data types. And we've seen that going on over the years and, and it's now playing out again, especially in Vector.
What is your sense of, you know, when do I need standalone? When can I use something that I'm extending or, or are there legitimate use cases for both? Or is it one or the other?
So I truly believe that, you know, the world is gonna go towards a multi-model database. That's what, you know, we have built out and we already always strive to actually build out. We started with key value, we added support for documents, then we added support for Graph, and now we added support for Vectors.
Um, the reason is, you know, one of the things is what you mentioned, which is skillsets, how many different database technologies, you know, can a particular developer or your team learn at some point in time? You know, you wanna make sure that you're able to leverage the skills across a broad set of use cases. Whether that use case demands that you use a vector or a graph or a document data type under the covers, I think that's gonna be an important consideration.
Uh, you know, the entire database and data field has been pretty dynamic and, you know, obviously, uh, folks have chosen, uh, the best, uh, or what they consider the best database for a particular use case, and that's why you have such a dynamic ecosystem of vendors, uh, that have evolved. Um, but I think if you can deliver the core value propositions that I talked about, which we believe we do, uh, with real time performance, you know, low latency availability, and you know, lowest TCO that's actually applicable to a bunch of different use cases. And, you know, data model happens to be just something that supports a particular use case.
Uh, so, you know, we think the world is gonna get to at, at a point in time where they will have a fewer set of databases and they'll build a bunch of different use cases on a multi-model database. Whether, you know, SQL and No SQL come together, you know, I don't know, they seem to be two different world right now in terms of what the sql, uh, databases are doing and, and kind of the no NoSQL databases is where the energy and the growth is for the most part. So what's your best advice to folks?
What's that one thing you see them doing these days that just makes you shake your head a little bit and go, folks, we need to be a little bit smarter than it. I think, you know, um, what I see out there, Mike, is, you know, we tend to resort to a path of lease resistance. So I would encourage people to do a little more research because before you choose a platform or a particular solution, you know, look at the alternatives, look at what's available and what's gonna stand with you for the next 18 months, for the next 24 months of your journey.
'cause replatforming is gonna be very, very expensive for most folks. So if you start your journey with a solution, which may look okay for you right now, as you grow and you start kind of, you know, leveraging more and more aspects of that particular solution, you'll see that you'll start hitting a wall either on a data wall, a scale wall, rather I would say, or a performance kind of, uh, issue or a cost issue, which is it starts getting prohibitively expensive. So do your research, uh, before you actually start that journey so that you look at a particular solution or a platform or an infrastructure can, that can stand its time for the next several years.
Uh, you don't wanna be replatforming in 18 months or, you know, 24 months. So that's kind of the, you know, big advice that I would actually put in front of, uh, folks. All right, folks, you heard it here.
There's an old saying that says, you know, you date your hardware vendor, you marry your software vendor, and it's still true. Hey, bu, thanks for being on the show. Great.
Uh, okay, once again, great catching up again, Mike. Thank you so much. All right.
And back to you.