Integrating Open Source Vector Databases into the Data Ecosystem with Chris Churilo at AIE 2024
Vector databases represent a pivotal addition to the data ecosystem. In this talk, Chris Churilo, VP of Community & Developer Relations at Zilliz, explores, explores how vector databases enhance existing data management systems and unlock new insights for enterprises struggling to tap into their massive volumes of unstructured data. Unlike traditional data platforms, vector databases are purpose-built to efficiently handle high-dimensional unstructured data and enable advanced similarity searches at scale, powering use cases like product recommenders, anomaly detection, drug discovery, chatbots and more. Open-source vector database solutions like Milvus benefit from community-driven development and innovation, flexibility and increased transparency not possible with proprietary systems.
Transcript
Hello, everybody. My name is Chris Trillo, and I'm the VP of marketing at a company called Zillow. And, uh, yeah, it's a really strange looking, uh, word because it is a made up word standing for zillions and zillions of, uh, of vector embeddings.
So today my talk is titled in Integrating Open Source Spectra Databases into the data ecosystem. And I wanted to start by, um, letting everybody know that, um, as a part of Zillows, we are the core maintainers of an open source vector database called bevis. And, um, with this, I just wanted to preface my talk by saying that I'm definitely gonna have a point of view when it comes to open source, because I've been on open source projects at this company and previous company, so I definitely, uh, have a bias towards open source.
Uh, the other thing I wanna mention about VIS is that, first of all, it's a, uh, type of bird. And, um, because, you know, open source projects always have to be associated with some kind of an animal. And we chose AM Elvis because it's, it's kind of like a falcon.
So it can fly really high up and really far, really fast, but it has amazing eyesight. And we chose this because, um, this is what we're trying to achieve with a Vector database. We're trying to help you get insight into the vast amount of unstructured data that you have.
And so we chose VIS as the name. And, um, about a year into the project, we then donated the project to the Linux Foundation. So even though we are the, um, the core maintainers, uh, this is definitely an open source project with our friends at the Linux Foundation.
And you can see we've got the usual set of metrics to share, uh, the popularity of Melva. So, quite a number of stars, a lot of developers that are using this project. And we're so grateful to all of them because they have definitely made the project, uh, the success that it is today.
Alright, so enough about the commercial and let's dig right into the presentation. And of course, there isn't a, uh, presentation about AI that doesn't come with the, uh, required generated artwork. And I've got this crazy looking flower, uh, that I generated on the right hand side from one of those, um, gen AI tools that are really popular.
And I would say that a few years ago, if a, uh, digital artist, you know, gave me this, uh, image, I would think, why in the world are the petals made out of, uh, water? What is this crazy thing with the lights? And I would look at the artist and think, Hmm, there must be something really interesting happening in that, uh, that head of yours.
But today, knowing that this image was actually generated by ai, I actually have a different question in my mind about this image. And the question that I have is, what is the data that was actually used with the model to actually generate this image? And I think it's something that we should be thinking about all the time, because that data is the, the data that's gonna bring the biases, or it's gonna bring the unique characteristics to the, uh, to the, uh, the models, and then ultimately to the things that we're gonna be generating with these models.
So data is just so fundamental to ai. Data is required to create these models, so then we can do incredible things with these models, with all the other data that we're collecting. So it's really an interesting concept.
And, you know, I've done presentations in the past where I always say, you know, data is really critical. It's really important. Uh, it is so foundational in ai, uh, without data, uh, we wouldn't be able to create these models.
Um, and the reason that we have these models is that we wanna make sure that we can, you know, get some insights, try to do some kind of predictions, try to understand, um, what, um, insights that maybe as human beings, we can't glean very quickly. And so, um, so in this particular case, it's not just buzzwordy, uh, data is really just at the heart, uh, of ai, and we need to keep that in mind always. So what I wanna talk about right now is I wanna talk a little bit about, uh, a search and, uh, data and a little bit of the shift that's been happening to help kind of ground, um, our, um, thoughts on vector databases.
And so we are all used to being able to use some kind of a search engine, whether it's, you know, for the public internet or internally or within an application. And so this, you know, picture should be super familiar, right? We type something in, um, maybe I'm making bread, especially during pandemic.
Uh, so I type in proofing bread, and then you can see I get two results. I get a nice little picture of, um, looks like dough, you know, uh, proofing. And then there's another picture of, uh, bread, uh, already baked.
And then, then there's probably some kind of associated text with these images, which was great, right? We've been used to this, and it's been really helpful for us for many, many years. But we, what we might've forgotten is that we actually had to train ourselves to make sure that we use keywords that would actually give us results.
And so if I had typed in, uh, something else and I didn't get these results, then I would sit back and think, oh, what are some of the syns that I could use instead? And I would put that in. And we just have become so accustomed to having to do this.
And the reason that we have to do that is semantics really matter. And sometimes, you know, when we just type in certain words, it just completely misses the context or the user intent of my query. So you can see here on the left, if I type in apple, maybe I expected the fruit, or maybe I wanted to understand the company, or maybe I wanted to understand, uh, the stock, right?
There's different, uh, semantics associated with the word apple. Back to my bread, uh, example, when I type in rising dough, maybe I actually meant proofing bread. And if I just type in, uh, rising dough, I would miss getting those articles that are about proofing bread because that is a synonym, but it might not have been associated, you know, in that search, uh, engine.
Or another example is if I put in the word change car tire, um, I might get this great image that shows me how to change a tire, but maybe what I was, my intent, my question was, when should I actually change the tire? And so I wanted information about that, that's showing in the bottom right hand graph. So semantics and user intent are really important because we wanna make sure that we can get people to that content, um, that, um, as quickly as possible that's gonna be really helpful to them.
And the reason this is important is, you know, you've seen this statistic, this has been popping up for the last couple year. About 80% of the data that, uh, is getting generated or will get generated is unstructured data. And unstructured data is essentially, uh, images, audio files, videos, um, user generated content, like in the form of reviews.
If you think about the body of your email, there's no structure to it, right? It's j just a bunch of, uh, stuff that we kind of wrote down. And, um, and I don't have to, you know, even just give you the statistic to prove to you that this is the majority of the content or the, um, the data that's being, uh, constructed.
Because just look at your own iPhone or, or, um, you know, smartphone or your laptop, you can see, uh, that this is by and far the number one content or data types that you actually have. And, um, it's really tricky to be able to get insights out of all this content. Um, as humans, we can look through it, we can watch a video, we can understand it, we can listen to an audio file, we can understand what this is, but the volume is just too big for us as human to go through all that stuff.
And so we need a way that we can be able to get through all that, uh, data so then we can use it, um, to then, you know, make our jobs or our lives even easier. But it's quite a big challenge. And actually, you know, the other thing you can think about is, um, a lot of us have like a Dropbox or a box, um, uh, folder that we've had for many, many years.
And I guarantee you'll see that you have a lot of this unstructured data in there. And, uh, if you're like me, and you'll notice that over the years, I even use different terms to mean the same thing. So if I have a bunch of notes that I've taken over the years, uh, I can't even do a really simple search.
'cause maybe I, I called something one thing and now I call it something else. I can't even remember, uh, what, what what I was, uh, referring to. Um, so it's just natural that, you know, these things are gonna evolve over time.
And so we need a way to be able to get sift through all that data. And the great news is, I think, um, we all know, um, by now that AI has really been the thing that's been helping us to be able to do this semantic search of this unstructured data. And, you know, it's primarily driven by the, um, fact that, you know, some of these new, uh, natural language processors or image classifications have really made a dramatic shift.
You know, instead of doing things in a, um, kind of like a more of a, um, statistical, more of a, a kind of a accounting way. So if we think back to that keyword search, um, issue that we described earlier, the way that we were, you know, building out these keywords is initially we actually had a keyword list, and then we thought, oh, let's try to generate these keyword lists from the, the, the content itself. And so we would just count to find out, you know, which words are used most and kind of assume that, oh, that must be an imp important term.
Um, but then when you found out over time that, um, you might have a small, um, body of text, and then you might have a large body of text, so the large body of text would have those terms in there more frequently. And so then that would unfortunately give, um, the longer content, uh, higher priority, which might not be correct. So then we made some modifications to then, you know, take into consideration the length of, uh, the content.
Um, but still those were still plagued with, you know, some problems in that they, they might not have all the semantics, definitely not have all the semantics at all the cinema synonyms. And so by going down, um, you know, by going down, um, uh, and taking advantage of some of these newer AI models, um, we can actually, um, you know, gather that semantic meaning without having to do all this extra work. So, uh, I think it was in about 2010 when Google published the paper, um, that really changed, uh, the, uh, approach to being able to, um, you know, create these models that could then be used in, in the ways that we're using it today.
And one of the things that we can do with these different models is we can actually take the data and we can actually convert it into something called a vector embedding. And the beautiful thing about vector embeddings is that, um, they actually contain, um, the semantics, semantic meanings of these, uh, items. And, uh, we can do something called an approximate nearest neighbor search.
And what's fascinating is things that are close and distance, um, are actually gonna be semantically similar. So it is a really exciting time, um, for us. And just to give you an example, we can see my crazy image on the left.
If I take a machine, machine learning model, I can turn that into a vector embedding, and this is what a vector embedding will look like. It's just an array of numbers. And the array could be just 30, or it could be, you know, 20,000, uh, um, items.
And I bring that up because what we're trying to do with this array of number is, well, first of all, a computer can understand numbers. We have to remember that, that that's how it's always been. It doesn't understand, uh, these images, so we need to convert it into something that it can understand.
And once it has it in a format that it can handle, now what we can do is look for another array that's gonna be close and distance. And, you know, if it was maybe two dimensions, we could probably do that, you know, math really easily. But when we're talking about the number of dimensions that we need for these vector embeddings, all of a sudden trying to find another, um, another embedding that's close and distant, it's gonna require a lot of math and a lot of computation.
And so, um, so, you know, it's, this is why we wanna rely on computers to be able to do that. Um, so a lot of you probably already familiar with, um, that, that kind of really simple, um, semantic search or that approximate nearest neighbor search example that I presented. Um, you know, a lot of people will share, um, examples using text, or they'll show, um, examples using the images like I just did.
But there's actually, that's just the start of this whole journey that we're on. There's actually some really interesting stuff that's, uh, um, happening. And you can take a look at this paper, uh, that was published almost a year ago.
And, um, what you can actually do with embeddings is if we look at that first example on the top left, we can actually use an audio file. And it actually, in this particular case, it has a sound of a crackling fire. And, um, what it can actually then present to us in the search results is an image or a video of a crack of a fire, um, because that's semantically similar.
And so this, like, cross modal retrieval is really exciting, right? Because as humans, like, if I heard that sound, I'd be like, oh, yeah, that sounds like, you know, a bonfire or a fireplace with a fire, um, in it. And so we are, uh, able to do this, uh, with this semantic search.
Uh, another example you can see here on the bottom left with the, uh, embedded space arithmetic, is you can actually take an image, in this case, an image of a crane and, uh, add it, add to it a, uh, audio file of waves, and, um, and we can embed those things. And, uh, the results would be, as you can see on the right of it, uh, a crane, um, in water, uh, in the, in fact the last image. You can even see that those are crashing waves.
So it's really, it's a really exciting time, and I think we are gonna be very surprised in a couple of years when we look back at all the in incredible things that we can do with these, uh, vectors, embeddings, and really being able to get the insights that we need out of this unstructured data that we have. And here are just some of the, um, the, uh, use cases that, um, you will see today with doing semantic search on unstructured data. So, uh, rag retrieve, augmented generation is all the rage right now.
Everybody's talking about it. Everyone's trying to build a chat bot. Um, recommender systems have been in place for a couple years, but these are becoming a lot more mainstream.
Um, also, um, doing search, um, for, in order to, um, you know, find things like, um, similarities with the molecular structures or trying to identify, um, new proteins are also like pretty common, uh, use cases that we see with Symantec search. Uh, and also anomaly detection are, um, some things that we've seen, um, for, uh, in the security space or even, you know, for fraud detection. So lots of really cool use cases that are, um, possible with Symantec search.
Um, so, you know, I'm here to tell you that as exciting as Symantec searches and vector embeddings, um, this does require a new type of database. And so in the past couple of years, a Vector database has emerged, and this has been important so that it can support all those various use cases that I, um, just described. But it's not just, you know, not just supporting those use cases at the surface.
Each of those use cases are also gonna have very specific requirements in order to make sure that, um, the, uh, the right kind of search results are going to appear as well as, um, in the, uh, in a timely fashion. So let me just give, um, you a little bit of context. So in that fraud detection case, uh, or anomaly detection case, um, one of the things that you're going to have to, um, uh, prioritize as in, in your requirements, you're gonna wanna make sure that you get very accurate results.
You're gonna wanna have very high recall. And it's important because remember, what we're doing here is an approximate nearest neighbor search. So we're trying to find things that are semantically similar, and so we wanna find things that are close and distance.
But when it comes to anomaly detection, we wanna make sure that we really get something that's very accurate, because we don't wanna, uh, accidentally, you know, charge somebody with fraud if that's not really the case. And so, in that instance, we need to make sure that we prioritize recall. Um, whereas in the product recommender use case, um, you know, it's probably more important for you to prioritize, uh, latency and query per second because you wanna make sure that your user isn't stuck with that spinning wheel of death and waiting for the recommendations to show up.
You wanna make sure that the recommendations show up really quickly, but is it really important that the latest, uh, pair of shoes that you're trying to recommend is in that data? Is, is recall really that important? And I would argue, uh, no, uh, you know, latency and performance of, uh, serving up the information is much more important.
So you need to make sure that you have a vector database that can help you to tune your requirements so that it's gonna fit and it's gonna work with your audience. Uh, in addition, uh, scale is going to be important, as it always is with any database, once you put it into production. Um, and when I talk about scale of vector and bennings, um, what we see in production is, uh, you know, 10 billion, a hundred billion vectors.
And so you wanna make sure that you can handle those really large volumes. Uh, in addition, you know, when everything are in production, things change constantly, right? So you need to make sure that you can, um, be highly performant, uh, when, um, when, uh, data changes, when you know, all of a sudden the number of queries is really high.
So you, you wanna make sure that you can handle all these, uh, different things that are gonna impact, uh, your database because you need to give your users an optimal, uh, user experience, right? This isn't just something that's sitting on your laptop, uh, where you might have the patience to, to wait for the search results. So, okay, now that I said that, uh, there's this new database, do you really need, uh, a vector database to do the semantic similarity search?
And I'm gonna contradict myself a little bit and say, no, you actually don't. Um, you can actually use one of these approximate nearest neighbor libraries like Face Annoy, H and SW, there's so many that are out there. Um, and just use that library with your, uh, with a small data set and, and, uh, build something on your laptop.
And this is sufficient for prototyping. And, uh, it's just really simple to do. Uh, in fact, if you go to the face documentation, it's, it's really comprehensive.
And so you could do all kinds of really great, um, build a lot of really cool semantic search, um, prototypes, uh, doing it that way. And it, and it supports a million vectors very, very easily. And, um, so you could just go down that path.
Uh, the second thing that you can do is there's a number of databases. In fact, I, I can't think of a database that doesn't support vector search today. Um, you could just use one of the vector, one of the databases that you have that have the approximate nearest neighbor plugin.
Uh, so like Postgres, you can use PG Vector. Elasticsearch has a plugin, BigQuery, Mongo Datas stack, you name it. Everybody has this, uh, uh, capability.
And, uh, if your requirements, um, are not so stringent and you don't need, um, something that's super performant, and maybe the number of vectors is, you know, uh, in a manageable space and maybe, you know, up to a hundred million, you could use an existing solution, um, and, and be just, just fine. But once you, uh, determine that semantic search is core to your business, and, um, then, then you do need to start considering a vector database, uh, for this capability. And that's because vector databases are purpose-built to handle the lifecycle of vector embeddings.
Um, you know, vectors are gonna get generated constantly because we're constantly generating more unstructured data. And so we need to constantly update our, um, our data, uh, and, um, not only do we have to update, you know, the, the vectors in the database, but remember, databases have an index. So we need to update our indexes.
Uh, in addition, when you are talking about scale, then we have to make sure that we can handle, you know, any kinda updates to our, our sharding so that we can make sure that, um, you know, at the end of the day, we can have a fully distributed system that's gonna be performing really well to make your users have a really great experience. Um, and so other capabilities that are, uh, inherent in a vector databases, the ability to do real-time search, the ability to do a number of different kinds of searches for semantic similarities. So the libraries and your typical databases are gonna support something called a Top K, uh, search.
So helping you find, you know, the top 10, uh, nearest neighbors as an example. But there are other searches that are really important, uh, to be able to add onto that. So a range search would be just imagine that, um, maybe I wanna find things within, uh, an area.
Just think of like a circle. So within the circle, I wanna be able to find these, 'cause maybe my top K goes outside of that, or, um, and so you wanna be able to have that as a capability. Uh, you also wanna make sure that you, uh, can have hybrid search.
So, um, you may want to actually have two different vector embeddings to do your search for the same body of text. One will be a dense, uh, vector, and one will be a sparse, um, vector that you wanna do a search for. And so once you get those search results, then what you'll do is re-rank them, and then make sure that you can, uh, present the most, um, uh, likely, uh, candidate.
Uh, there's also multimodal, so it's not just limited text. You wanna, as you know, we saw in some of those examples, uh, it's gonna be pretty clear very soon that, um, you know, all the other forms of unstructured data are gonna be things that we're gonna wanna search on. Um, and the list goes on.
So vector databases have a bunch of things that have already been put into place. Uh, and, and that's because vector databases are purpose-built for vector, uh, data. And it's all about, you know, these three core things, right?
Indexing that data, storing that data in a way that it's gonna be really efficient so that, you know, the query at the end of the day is gonna be fast. It's gonna give the, the kind of results you expect for the requirements of your specific use case. And like any, uh, database solution, uh, it takes time to build out a database solution.
And this is just a, um, a picture of a timeline of like when we started. So we did our first commit in, uh, 2019, and you can see I listed some other, uh, vector databases. And, um, you know, a lot of these, uh, uh, traditional databases added vector search, uh, just last year, which is great, but you can, you'll, you can probably, um, guess that it's gonna take time for them to be able to add all the other capabilities, uh, that are already part of, um, vector databases.
So, uh, so let's talk a little bit about the popular use case, which is retrieval augmented generation. I'm not gonna spend too much time, but I'll briefly go over what this is because I'm sure you've heard this from a number of, um, different conferences, sessions and, uh, events, uh, in the last 12 months. So the idea is that I wanna use a large language model because they're pretty cool.
They, they're really good at generating text, but, uh, they might not always have the right information. And so you wanna augment your language model with the knowledge base that you have, uh, within your organization. And the reason that you wanna do that is if you are building, for example, a chatbot, you wanna make sure that you can answer the user's questions in this chatbot with data that you actually have.
And the data can be in the form of PDFs or audio files or videos, uh, all kinds of different things that, that you have. And you may also wanna make sure that, um, you keep it separate. And so there are ways that you can implement this so that you know, you're not exposing your, um, precious, uh, internal knowledge to the outside world.
In any event, what happens is, on the bottom line, you convert the data, the unstructured data from your knowledge base, and then store that into a vector database. And so when the user asks a question, we convert that question into a vector embedding. And because remember, we're doing an approximate nearest neighbor search, so we wanna make sure that we can find the, uh, results that's closest to that query.
Uh, once we do that query, then we'll actually send the results, um, with some additional things in the prompt. Uh, we might, uh, we might want to, you know, add, um, I dunno, some kind of like rules around it, like, make sure you use this kind of a tone of voice, make sure the answer is in a bulleted list, make sure that you have a little citation, you know, to refer the user to maybe like a FAQ document somewhere. So there's a bunch of things that you're probably gonna add into the prompt, and then we take that prompt and we send it to your favorite large language model, which is really good at generating a really nice human understandable, uh, answer.
And then they get the, uh, answer. Uh, and, you know, this is just something that, um, you know, everybody's really excited about doing. So the thing I wanna point out here is with the Vector database, uh, hopefully you can see it's only storing vector embeddings.
Yes, you can store, um, text if you choose to in a vector database. Um, but you can definitely also, uh, store a number of metadata that you'll use to filter on, but it's not replacing the general functionality of your other databases. You're still gonna use your other databases to, uh, track, um, you know, user information, transaction information.
Uh, it's just a compliment. And, uh, and it, and they work, you know, beautifully together. So, um, you know, although you may think, oh, I could just use my NoSQL a SQL database to do this whole thing.
Remember, if your requirements, um, uh, are that, uh, the semantic search is part of your core offering, and you need to add a vector database, it doesn't mean that a Vector database will replace all your other databases. They actually work side by side. Um, so the other thing that I think is really important, and when we think about semantic search is that at the end of the day, this is a search engine and a database, but there's gonna be many, many ways that your dev teams are gonna be implementing, uh, semantic search.
So in the first option, what I've articulated here is that, um, the kind of the first portion of, um, you know, creating the, uh, rag flow, uh, your dev teams might create a lot of these things themselves. So they may be already familiar with a bunch of different, uh, machine lea learning models. They've already maybe created their own models.
Some people have, well, some organizations have some pretty sophisticated, um, ai, uh, teams that are building these models themselves. They may also have a mechanism to be able to split any, uh, text, um, because they really understand that data. So, you know, the first option is, you know, making sure that these, uh, solutions can work where you have a, a, a dev team that can, you know, basically they've already built all the capabilities.
So in that particular instance, it's important for solutions to be able to work with the environments and tools that your dev teams are, um, are using. In the second option, there's more and more, um, tooling that's becoming available. Uh, things like, uh, link Chain or LAMA Index, where those tools are really good at basically preparing your content, right?
They're going to help you to, uh, split the content into chunks, uh, depending on, you know, what the requirements are gonna help to generate the, uh, vector embeddings and then store them into a database and then follow the, the normal flow. And then the third option is, and I'm just using Zillow's Cloud our own, uh, solution as an option, but there's actually many other solutions that are doing this. In fact, um, there are solutions that do it, um, with a Vector database like I'm describing, uh, down below.
And some of 'em, uh, um, a allow you to pick your Vector database. Um, but this is for groups where, you know, they don't have time to build out their own models or they really don't wanna, um, you know, work with, uh, a number of tools and try to pull these solutions together. So there are a number of end-to-end solutions, uh, that start all the way from building out that data pipeline and then processing the data, turning into embeddings, and then, you know, putting into the RAG framework.
So regardless of which option your teams feel the most comfortable with, it's important that, um, you know, we don't force our teams to work a certain way that we should make sure that these solutions can work in the way that's gonna be easiest for them. And so, um, you know, initially I thought, oh, I should just make a giant eye chart with every single integration that, you know, or, um, uh, and every single, you know, solution that should be a part of the, uh, ecosystem. But I think that would just be, you know, a bunch of little logos and we wouldn't be able to see anything because there's a lot of different ways that you can, uh, do this.
And so instead, that's why I, I split this out, um, so that, you know, think about how your teams are approaching, you know, building the semantics search capability. Are they building it, you know, from scratch? Um, and they just need, uh, you know, a vector store to be able to do that and, and an LLM, or do they want something that that can, uh, help them do this, uh, end to end without having to build it?
Um, so, you know, of course, open source is, uh, really important to me as I indicated the start of this presentation. But I think it's important that we, um, remember that open source is really about, um, being able to foster innovation. Uh, I think we've seen this time and time again that, um, by making it open, inspires, uh, developers to look underneath the hood and, um, you know, consider other capabilities or bring other capabilities to these different projects, um, and then work together to really, you know, make these enhancements.
Sometimes you might bring in, uh, a capability that you're interested in and someone else is like, oh, I think I can build that. So it doesn't just put the onus on your own shoulders. So really encourages, uh, collaboration.
And the one thing that is really, really important, especially in ai, is transparency. Understanding what is underneath the hood, not just for a vector database, definitely you wanna be able to look underneath to see how the vector database works, how it's actually treating your data, but it's also important to be able to, uh, have the avail availability of open source machine learning models, open source lms, um, also open source toolings, um, like, you know, LAMA and digs and Lang chain, so you can dig underneath to see what is going on, um, with my data, because remember, we're using data to ultimately try to get insights from data. And so if we get really strange results, we're gonna wanna know what is causing that strange result.
Is it my core data or is it the tools, or is it the models? We really need to make sure that we have that transparency so we can troubleshoot appropriately. So, you know, uh, hopefully, you know, I was able to convince you that vector databases, um, need to be, um, or already are a core part of your AI stack.
And, um, you know, it's really important that we understand that efficient data retrieval and storage retrieval, indexing, et cetera, is gonna be really important. But, but also it's really important that you understand your own, uh, use cases and requirements to make that determination. Don't just blindly pick, uh, technology, uh, just because it's the, the hot thing of the day.
We should make sure that, uh, whatever you're striving to do, we should make sure that you pick the technology that's gonna help you in, uh, becoming really successful. And I would recommend very strongly, you know, consider open source. Um, I love a lot of the tools.
I think, for example, open AI did a spectacular job of creating, you know, an API really simple way for us to be able to access, uh, their LLM and their models. I think they did a brilliant job with that. They really, um, opened up the eyes of a, of a lot of developers to using AI and ml, which is fantastic.
Um, but when it comes to your own data, I do think you should really consider looking at open source, um, either as an alternative or maybe something for you and your teams to really, you know, dig into and learn to understand, you know, what is what's actually happening underneath so that, you know, you can make sure that whatever you're building is gonna be successful. And, you know, maybe we'll have even more and more of these kind of crazy, uh, looking, you know, images that, uh, I've got here on my, uh, slide. But, you know, I think the thing is, I wanna be able to have the confidence that I understand what's the data that's behind it that's helping to, you know, generate these things so then I know what's gonna, you know, ultimately come out of, uh, the solution that I'm building.
So that's, um, that's pretty much it for my, uh, presentation, but I do wanna end with just a little bit, little plug. So, um, myself and my team, we run a set of meetups called Unstructured Data, um, and we intentionally call it unstructured data because, um, we don't wanna just limit our meetups to, uh, vis, I think it's important that we understand the entire life cycle of unstructured data all the way from generation to creating the pipeline, the embeddings, what you do with it, et cetera. And, uh, so we have a set of meetups that, um, that we have every month.
You can see this is the typical, if you look in the little picture here, that's usually the, uh, crowd that we get at each of these meetups at, uh, the various locations. Uh, so if you are in these areas, and we're starting New York City in June, I, I, um, encourage you to come join us. We have a lot of really great talks, um, from a lot of really interesting people.
And we also do live Twitch streams, so you are welcome to join that way. And then also, if you go to our YouTube channel, you can see we've got all the recordings and we put all the presentations, uh, uh, out there. So it just helps, uh, everybody to learn more about what's happening in this space, uh, because I think every week I find, uh, a new cool tool or a new model or some new approach to Vector embeddings, uh, and it can be a little bit hard to, uh, keep up.
So hopefully you can, uh, take advantage of this resource. And then if you are really keen to, you know, try a vector database, um, I encourage you check out, um, vis, as I mentioned, it's open source and also a Linux Foundation, uh, project. And, uh, just put the links in here to make it easy for you.
So Cool. Well, thank you so much for joining me. I look forward to everybody's questions.