Open Search Foundation Unveils Version 3.1 with Enhanced Features with Mukul Karnik | Open Source Summit NA 2025
Open Search Foundation launches version 3.1, improving search and log analytics. The platform now supports various use cases, including observability and security analytics. New features like GPU-based indexing and advanced analytics enhance performance. Integration of vector databases boosts semantic search. The open source community benefits from collaborative innovations, while skills in data governance remain essential for AI applications.
Transcript
Hey everybody. We're back at the Open Source summit in Denver, and we're here with Mukul Karnik, who's director of the Open Search Foundation, and we have all kinds of new and interesting things to talk about. Welcome to the show.
Hey, thanks Mike. Thanks. Great to be here.
Yes, a lot of really exciting things to talk about. 1, and, uh, it's a continuing evolution of, um, you know, the powerful search and log analytics capabilities that open search has. And we have some exciting announcements, um, on, in both search and log analytics, but also in, uh, the agent space.
And so, uh, really exciting, uh, like summit and a day. Well, so let's go through that a little bit because when people hear open search, they're like, okay, search engine, I get it, full stop. But this is really a family of things that you guys have put together and they have different use cases.
So kind of walk us through a little bit of what you're seeing as some of the use cases, including the AI ones, but there's also observability and all kinds of stuff going on here, so it's almost, we're seeing it almost everywhere. Yeah. Yeah.
Great. Great point. You know, yes, open search, great search engine, as you said, and, you know, used in wide variety of use cases, but also really powerful observability use case, because if you think about it, when you are trying to solve an observability problem at the core of it, you need search.
And if you have a powerful search and analytics capability, you can get to your root cause very quickly, which is what you need to do for observability. And so really good observability platform, also an emerging security analytics kind of use case, but largely search and observability platform. And then on the AI side, well, we have vector databases, but we also need search to kind of make those applications work as well.
What's the relationship? Great, great question. You know, like open search, really powerful at search, but also, uh, we have a lot of really powerful vector database capabilities.
And you think those two are independent things, but in some ways, um, like the search, think of it as like a keyword search, uh, right where you're typing and you like get back research results that are based on the keywords. What vector, uh, uh, search or vector database lets you do is do, uh, use the more semantic understanding of the words to be able to find, um, the results that are more semantically related. So the combination of like the keyword search and the semantic search gives you more powerful search.
Like it helps you answer the questions that you are asking more precisely. And, and part more, more quickly. So the paradox in my mind is that we're using open search to create AI agents and apps, which in turn will we, the things that we use to drive observability, invoking the rest of the capabilities of the platform.
So is there some sort of like symmetrical loop occurring here? That's a great observation. I, a lot of people miss that, but I, you're right.
Like open search is a, you know, great search and a knowledge base if you think about it as right, and for getting your observability and understanding the root cause correctly, what you need is a good knowledge base. I mean, we as DevOps engineers, you know, uh, have that knowledge base within us, but it's also there in the runbooks it's there in, you know, different, uh, wikis and different places, right? And so if all of that knowledge base exists in something like an open search, then you can really make your observability debug in even faster because you can now under like, uh, open search and the agents can understand the root cause much quickly because the data is there, right.
So among the new features that are being rolled out, you know, are there any of your favorites? I know that's asking you to pick your favorite child, but are there things that you know are a little more important than others? Well, I think there's several very, uh, very, like, yeah, it's hard to say which one is favorite, but I'll walk through couple.
Um, so we, uh, open search, as I was saying, you know, leading vector database. One of the challenges we heard from users is, uh, in like indexing can take time. And so we build this GPU based acceleration on, on the indexing side where you can, you know, index four, five, sometimes even more like, so five times faster, even sometimes more.
Uh, and so you can get all of that data into open search much more quickly to be able to start serving, you know, requests. And so that's one really nice innovation, which has really good application for a lot of, uh, users in the community. Uh, we are also building, um, ability to have, uh, like load graphs partially, and that gives you, again, ability to, uh, respond to results quickly.
So it's, it's making gene I more real time, if you think about it, like at a high level, and that that is really powerful. The other, uh, set of capabilities is more on the nalytics side and the observability side. Um, we've introduced, you know, more richer analytical capabilities and analytical functions.
You know, your documents that you index into open circuit, like these are logs, right? Uh, could have any kind of structure and JSON format. And so we've introduced nested, you know, JSON support and really a lot of rich analytical capabilities, um, into open source.
Um, as we kind of move along here, it seems like the Linux Foundation itself has multiple AI initiatives. There's the PyTorch folks we just talked about, the A two A folks. Will there be cross pollination between all these groups?
And how does that all come together in your mind? I mean, in general, uh, in the, I, I think in the open source community, there's a lot of cross pollination that happens, uh, because I think innovations happening in each of the projects, um, benefit, like I think, uh, other projects like you can kind, uh, tag team in some ways, right? I mean, um, being able to, uh, leverage, like in, for example, I'll pick, uh, open search right in, within open search.
We are, uh, leveraging, uh, Apache Cal site as a, uh, uh, planning layer for open search, and then we are leveraging other projects to be able to then, uh, power some of the capabilities in open source. So it just, that synergy exists between different open source, uh, open source projects within the Linux Foundation. We are also working on several, um, like, you know, there are other, uh, uh, like, you know, gen AI and ai, uh, capabilities and, uh, and as, as a community, we are looking at, you know, all the innovation that's already happened and leveraging that to benefit the users and community of open search.
Mm-hmm. One of the things I think of observed is that we used to kind of have developers over here, and then the data managed over here, and it was kinda separate in the age of ai, that all seems to be converging, and I see more and more developers, you know, discovering data management fundamentals. So, um, do we need to kind of focus a little bit on skills and training there and, and, and what are you guys doing about all that?
I think skills and training, definitely. I think, uh, I think there's that, like, as I think, uh, a critical part of any of the gene applications is the data, as, as your observation is. And I think, uh, understanding, you know, quality of the data and, you know, the capabilities of the tools that access the data is really important.
I think the other important thing is also governance of the data and like, you know, being able to make sure that you can, uh, track where the data is coming from and like permissioning and all of that plays a big role in this world because as these tools, you know, um, become more powerful and can access different things, you wanna make sure that the go, like there's a good, like all of these projects, such as written search, uh, are focusing on the data governance part of it. Like, All right, so you just got the new release out the door, and I'm sure everybody's asking you the same question. What's next and when's it coming?
Lot of really exciting things. 1 release. 0 in April, and just within a short span of like two or three months, we have some amazing new things.
Our pace of innovation is growing. I mean, if you like, um, just we, we transitioned to Lenox Foundation, uh, in September, and since then we've seen about 46% jump in our active contributors. And so that's driving all of this innovation.
And so we'll see a lot more features come out. Uh, and continuing to innovate in this, you know, in the space of, uh, search vector databases. Um, agent tech now with all the MCP support that we have, and then also in the observability space where, um, a lot of the rich analytical capabilities, uh, making it more accessible to all these agent tools.
All right, folks, you heard it here. If you wanna stay close to what's happening in ai, keep your eye on the Open Search Foundation. 'cause I think they're closer to it than just about anybody else.
Hey, thanks for coming by. Thanks Mike. Great to talk to you.
All right. And we'll be back in a minute.