Building the Vector Database Behind Generative AI
Jeff Zhu, VP, Product Pinecone outlines Pinecone’s mission to enhance AI knowledge and the role of vector databases. Key topics include advancements in generative AI, the introduction of dedicated read nodes for efficiency, and the need for improved data input methods. The future of Pinecone focuses on user-friendly tools for developers and an invitation to explore new features.
Transcript
Hey everyone. Welcome back here to Tech Drunk tv. My next guest is Jeffrey Zu.
Jeffrey is the VP of Product of Pine Cone. It's his first time here on Tech Drunk tv, so let's welcome him and, and get to know him. Hey, Jeff, how are you?
I'm doing great. How about yourself? Good.
Um, I mentioned you were VP of Prime, uh, of Pine Cone, Jeff and Jeff. Is it, do you prefer Jeff or Jeffrey? Let's, Jeff is usually good, so Jeff is good.
Yeah, Absolutely. Truth be told, my middle name's Jeffrey too, so I, but not many people know that. But anyway, um, so Jeff, a lot of people know Pine Cone, especially, you know, in this era of ai and when generative AI and, you know, o uh, open AI first kind of exploded, you know, pine Cone really kind of got to people's consciousness in front of their faces.
Um, we're gonna talk about that, but before we do, let's talk about you a second. What, what, what's, what's your arc? What's your story?
How, how did you wind up here? Yeah, I'm happy to do that. So, you know, I've spent most of my time working really in large scale distributed systems.
Like, uh, I, before Pine Fund, I spent eight years at Microsoft working in Bing. So I was running a lot of their machine learning AI infrastructure platform there as a platform pm So I was doing things like, you know, training large scale language models before chat GBTI was doing things like accelerating GPU inference, uh, using Cuda and asics and things like that as well, really, but most relevant. I also built their vector search platform from scratch.
So I was actually running, you know, hundreds of billions of like vectors, you know, for their web scale search engine itself. And that was, you know, that was obviously maybe like, you know, started seven, eight years ago. But that was one of the most really exciting things, you know, before Vector Search became so popular across the industry.
You know, it was already a very, very integral part of search systems and, you know, it was locked away to a really, really small set of, you know, uh, companies. Right. But honestly, that's kind of what ended up motivating me to join Pine Con in the first place, because I saw what Pine Con's mission was, which was really, you know, to make AI knowledgeable through vector search and enabling that through, you know, enabling any developer to take it on.
Right? And so that was really kind of what brought me to Pine Con in the first place. 'cause I saw how hard of a problem it was, how challenging it was, and then, you know, and how integral of a technology it was going to be.
And then, uh, you know, you know, the stars aligned and I found a startup that was in the same mission that I wanted to do. Excellent. Man.
You know, a lot of people, so I, I've founded, co-founded four or five companies over the years, you know, and I always say people don't realize, you know, their, their 10 and 12 year overnight successes, right? Open AI and Jet GPT exploded on the scene, and everyone thought, oh my God, what a, this is an overnight success. All of a sudden it's here.
Yesterday it wasn't. Right. Well, maybe to you it wasn't, but there've been people working on, on the technology behind this, in, in, in the case of, of LLMs and, and, and, you know, uh, uh, artificial intelligence, I mean, this has been going on for 30 years, some of this stuff, right?
They really Exactly. I mean, if not more. Um, and then the whole vector search thing, a as you said, right?
Microsoft Bing was doing it. There was a small group who had that use case of a huge, huge, you know, basically cataloging the internet, right? Yeah.
And this was the logical way to do it. So, you know, when you see these things that explode on the scene, don't assume it was like some eureka moment that happened six months ago. There are people who toiled their, almost their entire careers Absolutely.
To make this stuff available to us. Um, now of course, look, tech Strong AI is one of our sites. We, we talk about AI a lot, but not everyone.
Jeff, I think, is familiar with the role that Vector databases, vector search, you know, the whole way of injecting data, whether it's Rag or for s SLMs or, or what have you, you know, the role Vector plays for those who may be my cybersecurity friends who are smart as heck, but don't know, don't play with Vector a lot. What, how would you explain it to them? Yeah, I think the simplest way is to think around, you know, what vectors really do are they allow you to represent, you know, instead of in the classic search world of like text, right?
Like, oh, I know a keyword, right? This specific keyword mm-hmm. Like, you know, is exactly, exactly what I'm looking for.
But what vectors really allow you to do is kind of represent more concepts, right? Rather than necessarily that exact word, that exact text at any given time, right? And so that's really powerful because all of a sudden, you know, you can have, you know, I think the classic, you know, example they always give in like the classic ML courses is like, you know, man and woman and queen and king, right?
As concepts, right? Where if you represent them in vectors, you know, man and king would be closer together than let's say, you know, uh, you know, woman and king, right? Right.
And so idea here is that you're able to, because these models essentially are trained on huge Es of data these days, honestly, I don't even understand how much of the train on these days that, that it's actually mind boggling, you know, how mm-hmm. They're being trained, but like, they're able to essentially learn these relationships between the words, and then they represent those relationships. And essentially in vectors, which is, you know, for most people it's just, you know, it looks like a, like a series of numbers, right?
But within those numbers, that distance between those vectors is actually what we call the semantic similarity or the distance, right? Which ultimately allows you to say, is this similar to this particular other concept? Right?
And of course, when you go into the world of LLMs and natural language and how we interact with AI today, that's super important, right? Because we don't speak as human beings. We don't interact with like perfect text, you know, recall of exactly the same language.
We operate in concepts as well. So this is what it really enables you to do, is represent concepts and speak in that language, and ultimately be a little bit more closer to how we as humans actually even talk and, and kind of interact with each other as well. Very cool.
That's a great kind, almost layman's way of, of explaining it to people. I think they get it. And, you know, and a lot of that goes towards, you know, you hear people say, well, generative AI is just picking the next word based on the previous word, right?
It, it, it's not really, you know, look, there's an automagical element to this. You show it like, to my wife's family, they think it's a person in there. You know, I and, and others say, well, no, that's not a person.
They, it just picks the next word based on the last word. And, and it, and it just goes on like that. And I don't think that's a hundred percent true either, obviously.
But it, it does give you this idea of how, how vectors work, right? And how, what makes sense and, and, and how these things go together. If you can, Jeff, obviously as we, we've talked about with the, with the advent of, of, of AI and generative ai, pine cones had a massive of explosion.
Tell us a little bit what that ride's been like and what's been going on. Yeah, for sure. I think, uh, you know, when I joined Pine Cone, I actually joined roughly nine months or maybe eight months before chat GPT.
So honestly, like before, before, right? So I will say, you know, I don't wanna call myself a hipster or anything like that as well, but, you know, I was, I was there before it was cool. Mm-hmm.
And, uh, and no, I mean, part of it was that, you know, uh, you know, at the time of, you know, when I was first joined Pine Conan and the types of use cases that we were thinking about, obviously were very in the classic world of information retrieval, exactly what I was doing at Bing, right? Mm-hmm. It was all about how I do semantic search as in doing, improving my search results and search relevance through this.
How do I build recommender systems, right? And so obviously when Chat g BT came out, that completely exploded because all of a sudden, right now, what was really interesting about these LLMs right, is that they were really good quote unquote reasoning engines, right? But they really, what they needed to do to be accurate and, you know, you know, and actionable and, and, you know, useful to people is that they needed the right context, right?
And that's ultimately where kind of rag came in as a concept, you know, retrieval augmented generation, where the idea here is obviously that you're able to take, you know, the only relevant, you know, snippets or pieces of information and feed that as part into the LLM, right? And, you know, I think, you know, that was obviously the very early days of how, you know, RAG was treated. And so, I'll be honest, I think everyone at the time of really didn't understand rag.
And so the first couple years of where Pine Cone was was really all around honestly, educating the market and helping people understand like, what are the basics around kind of this, you know, vector search capability for, and how does it relate to AI and LLMs, right? But I think over time, right, as people got more comfortable with Vector Search and capabilities, you know, I think once some of the evolution was that, you know, how do I not only, you know, there's obviously very basic ways to do RAG in terms of just like a very simple embedding, you know, retrieval just shove it into a prompt, right? But obviously as the M'S capabilities grew obvi, the, the other kind of the, the other part of the market kind of grew in, in, what do they call context engineering, right?
As in like, how do I then not only just do basic retrieval, but more complex and more interesting ways to ensure that I'm getting the most relevant information back, right? And so as part of that Pine Cone has also evolved, right? Is that of course, by the way, you know, our Core Vector database is massively scalable, massively, you know, cost performing things like that as well.
It's really, really important for that to be, I would say, a base foundation layer. But, you know, as we've looked to kind of where we see the future of kind of the evolution of Pine Cone going, you know, our mission at Pine Cone is to make AI knowledgeable, right? Is not actually to be the best, best vector database in the world.
We see it that the Vector database is a component, it is a very, very integral piece to the puzzle to actually make AI knowledgeable and to produce knowledge for them, but is, you know, only one component of that, right? And so as we kind of look towards the future, you know, we still care a lot about the Vector database and serving our customers there, but you know, now that people are building agents, right? And not only just, when I say not only just developers, I mean, just almost anyone these days can build agents themselves.
And that's kind of where we really wanna play a role in the future where, you know, it's all about making AI agents and LMS knowledgeable, which means more than just pure vector surge. But, you know, we have a, we have a product called pinecone Assistant, which allows you to essentially take in files, take in, uh, you know, PDFs directly, and actually get the relevant context out of it. So instead of using Vectors as the interface, we're actually using something like, you know, PDFs and chat and words, right?
Which is much more natural to anyone building Sure. Agents today. So I think that's kind of been the interesting ride where, you know, of, of course, the Vector database is a super integral piece of the puzzle, but I've, as kind of the world has evolved and, and the capabilities that LLMs and agents has evolved, we also want to continue evolving beyond the Vector database, but more into the knowledge space itself.
Right? That makes total sense to me. com is one of our sites, right?
And then, yeah, I started it 12 years ago, 13 years ago. Um, and we had a bunch of, like the OGs from DevOps, you know, Patrick dubois and John Willis and Damon Edwards, and of a lot of people who started the DevSecOps Movement, um, just really smart, good people, and they were operationalizing ai. Hmm.
And back then, two years ago, the only way to get anything in was you, you had to basically vectorize it, if, if we could call it that, right? Put it into the Vector db. And, and so they were here, that's when I first really became aware of Pine Cone and, and what role it was playing in there.
And I was like, wow, they're kind of the only, they're the ferry boat that gets you from this side of the island, you know, back to the mainland, into the, into the ai. But I always felt like we had to democratize that. We had to make it easier for people to get the data in it, you know, converting it into a, a Vector DB while, you know, my, my developer friends was like, oh, this is easy, and I just write a little script.
It sucks it in, boom, boom, boom, that, you know, it was going to get easier. And it sounds like pine cones kind of realize that and is making it easier. And I, and I do think that's the future too, right?
Because as you said, these LLMs are so massive, Jeff, we scraped, well, we, from a publicly available information perspective, we scraped everything there is to scrape. There's a lot of stuff that's not publicly available that we need to scrape, but I don't think the way we built the LLMs is the way we'll scrape that, right? Mm-hmm.
I think Exactly We need better input. Um, people who wanna get more information pine about Pine Cone, where, what, what's their best kind of on ramp? Honestly, uh, the best on-ramp here is, uh, to just go to our website, pine ca io and sign up.
So, so we have a very generous, essentially a starter tier, which is completely free. Uh, there, you don't need to put in a payment card whatsoever. You can get started with both the Vector database if you're interested, you know, getting deep in the Vector database, but also, as I mentioned with Pine Con assistant, with a little bit more of a simple, if you ever just wanna chat with your docs, right?
Drop in a PDF start chatting with it, or even, you know, like retrieve the relevant snippets from it, all of that is available in our starter tier. And so just come to the website and sign up. The way that I always like to think about it is that there's no better way of learning about it and actually doing, right?
I mean, I can show you the docs, I can show you the blogs and hon honestly, but I always come back to just get in there and try. And I promise you, like, one of the things that we absolutely emphasize here as a, as a core design principle, is a developer experience and user experience, right? We wanna make it dead simple, easy, and as fast as possible to see value.
So just come on through our website and try it. I promise you'll be super fast. Very cool.
Um, alright, let's, let's pivot. Pine Code recently announced something called dedicated read notes, DRN full disclosure, it's in public preview right now. Maybe you can tell us more on that timeline, but first let's just, you know, define what do we mean by dedicated read notes?
Yeah, so, you know, before I dive too deep into the weeds of it, I think one of the things around just like vector search, so this is in the vector database world of the vector search world, just to be clear, uh, the, one of the things around vector search is that it's all about, you know, in many ways the primary, you know, dimensions you consider it is around accuracy, performance, and cost, right? Those are, I would say, when it comes down to it, you can talk about some other bells and whistle, but when it comes down to what people really care about for a database, it is those three things, right? But, you know, one of the things is that for, but what we found over time was that, you know, there's a lot of different types of workloads, you know, that are especially in the higher scale area that are starting to stress exactly what vector search workloads look like, right?
So on in one dimension, we have what we call, I would say like recommender systems, right? You can, this is the classic e-commerce. Every single time my page loads, I want to send, you know, find the most relevant.
You know, you may be also interested in items, but this is talking about, you know, thousands, tens of thousands of queries per second, right? But with really, really short latencies, right? Because you're on a page load, we can't be waiting, we can't be slowing that down, right?
So like under 50 millisecond latencies, right? Um, but then, and then on the, you know, in on the other end, what we have is like customers who wanna do very, very large semantic search workloads, I have a corpus of let's say decades of news articles, right? And I wanna be able to enable search over that, right?
So you have a huge, let's say billions of vectors that you wanna search over, but you wanna be really, really accurate for those, right? And then last but not least, kind of what's been around, it's like, uh, what we call multi-tenant rag or agents is that you have maybe millions of really tiny independent, you know, completely isolated like agent memory context that you essentially wanna be constantly writing to, but you know, you don't need this to be super fast because, you know, it has an LOM in it. I don't, hundreds of milliseconds is probably fine, it's not a big deal, right?
But I think what we realized when we were working with all of our customers is that to do the optimal cost performance trade off for each of these, you actually need to serve it in a different way, right? So I would say maybe around a year ago we kind of embarked on our transition to what we call the slab architecture, which is essentially an object storage based architecture where essentially the ground truth, instead of using everything on SSCs and memory, which a lot of the, you know, existing industry does, we're gonna be object storage based, right? Instead where our ground truth lives in object storage, and then that allows us to serve use cases like multi-tenant rag and agents very, very cost effectively, right?
Because if you don't need it, we just leave it in object storage and now we have a completely serverless kind of operation, right? That allows you to be very sporadic, you know, just pay for what you use. It's a really, really good scale, cost performance offering for you, for, for our customers, right?
But what that doesn't work well for though is, is something like a recommender system. Because if you think about a recommender system, I need to send a single index, a single like, you know, node itself, tens of thousands of queries per second, right? It needs to be essentially dedicated hardware so that we can actually serve it and fully saturate it and actually give you the best latency and performance instead of like trying to overcomplicate it somewhere else, right?
And so that's kind of where, you know, this new offering is what we call dedicated re nodes, is that instead of being a multi-tenant architecture, what we're doing is that we're actually allowing customers to select and say, Hey, for this given index, I'm gonna dedicate five nodes to this one and I just want to, you know, fully saturate it. You're the only customer on it. And it really enabled very high performance and, you know, high throughput for really, really good cost performance, right?
But really kind of the intuition here is that we en and we enable people to really fully saturate the hardware instead of really like, you know, dispersing it and kind of working in a multi-tenant fashion. So that's kind of the key enablement around what dedicated Read Notes does. And you know, we actually have like pretty massive scale, you know, use cases on top of it.
4 billion vectors, uh, you know, 6,000 queries per second, right? With like a P 50 of 25 milliseconds. So like, this is really, really cost perform and high scale hardware.
But of course, as I mentioned, the thing that's really interesting that I do want to call out is that Pinecone supports both models, right? You can be really, really sporadic, very cheap pay for operation, or we also support the ability to do essentially these dedicated hardware profiles, which honestly to my knowledge, is fairly unique in the industry. No other kind of vector database offering kind of provides both at the sa at the same time.
So really it's all about flexibility for our customers and really enabling them to find the right exact hardware profile for their actual use case And just for people out there. So Pine Cones offering this kind of as a service. Mm-hmm.
Absolutely. Right? And that, that's the important thing I want people to know.
Jeff, we mentioned it's in public preview. When do you think it might be JG eight? Yeah, we're looking right now at a Q2, so sometime in April, I believe targeting roughly April's timeframe for a, you know, a, uh, GA release.
Uh, in the meantime though, I think it is something where, you know, we do have, even though it's in public preview, we do have a couple of customers running on production in it, as, as we always do. So I would really recommend if you guys want to give it a try and, and, you know, and let me know what the feedback is. We're actively continuing to improve the performance, the efficiency, the capabilities of this.
And you know, we're really eager to see this kind of powering massive scale use cases in the future. Very cool, man. Hey Jeff, we're about outta time.
I want to thank you for coming on here. It was a great discussion. I think our audience definitely learned something and that's the best thing about it.
Um, come back, you know, we don't get enough information like this. We are lucky to have you, and we'd love to hear more as Pine Cone has more news and more, you know, breakthroughs that they're making here. Thank you very much.
Yeah, thank you. All right. Jeff z VP product pin cone here at Textron tv.
We're gonna take a break. We'll be right back.