Shapeshift Milestone 3 – Vadim Ogievetsky and David Wang, Imply
Imply’s VP of product, David Wang, and chief experience officer, Vadim Ogievetsky, discuss the launch of Imply’s Project Shapeshift Milestone 3. Project Shapeshift is a major initiative designed to transform the developer experience when building real-time analytics applications. In the third milestone, Imply will add new industry-differentiated capabilities that enhance the Apache Druid experience and further strengthen its position in the marketplace for analytics database for streaming data. Imply was founded by the original creators of Apache Druid.
Transcript
This is techstrong tv. Hey, welcome everybody. I'm at the great pleasure today of being joined by two folks, uh, two people from a company called Imply.
You may or may not have heard of them, but they've got some inciting news we're gonna share with you. We get a little bit of what they do and some of the background. Uh, both David and Em have joined us here.
Uh, David, VP of Product Vadi is the Chief Experience Officer. Vadim, why don't you start, introduce yourself a little bit and then David, if you would do the same and also kind of give us an overview. Tell us about Imply.
Well, thank thanks so much for having me. Uh, my name is Vamo GSKi. I am, uh, one of the co-founders of Imply and, um, also, uh, imply as a company, as based on the JID Open source project.
And I am one of the, uh, jid, uh, project management committee people. Uh, so I help guide that project as well. And, uh, I come to you here in both of those capability capacities.
Fantastic. David. Hey, and, uh, my name is David Wang.
I am the Vice President of Product and Technical Marketing at Imply. And, you know, my team's job is really about sharing, evangelizing the great technology that Vine and the other engineering teams with an imply have created, as well as across the Apache, uh, jury community. So, um, you know, call me the evangelist and storyteller for all the great innovation that they've built.
Fantastic. Well, tell us a little bit about and tell, dive some more into imply. Maybe, maybe if you want to start with, um, not everyone might not be familiar with, uh, Druid, also Apache Druid.
So kind cuz I know that, uh, given your back background, uh, Veem is one of the starting places I imagine, for the inspiration behind, uh, imply. Do you wanna, you wanna start that off, David? Yeah, absolutely.
So, uh, imply is a company that's founded by Veem and the other co-creators of Apache drd. Um, the company was founded in 2015, but the open source project of, uh, Apache DRD was started in the early 2010s. Uh, Drood itself is a really fantastic database.
It's actually much unlike any other database out there. Um, it's sweet spot as every database should have. A sweet spot is that it is a high performance realtime analytics database.
So it's really designed for developers and architects and engineering teams that are trying to build analytic applications. So if you think about analytic use cases that need high performance subsecond query response times at like very, very large data sets, uh, or supporting high queries per second, or, you know, trying to do analytics on streaming data, uh, for realtime insights, you know, any of those use cases are where people are gravitating towards a different type of database and that's why they come across Apache Druid. So, uh, use cases might be things like if you're doing analytics on, uh, multiple data streams, video streams, or high volume transactions in a financial setting or, or or things like that.
Is that where you do? Yeah, I think the easiest way to look at it is event data. You know, event data is data that's generated from clicks.
Uh, event data is generated from application logs, event data is generated from telemetry and iot, OT and sensors. Whenever you have event driven data, then you're gonna really talk about, you know, very fast velocity. You're gonna talk about high volume, uh, ingestion, and you're gonna be talking about operational visibility.
So, great customer examples or actually user examples of Apache Drew be like confluent. You know, confluence a company that has a really fantastic cloud service for Apache Kafka. Um, by standing up this multi-tenant cloud service, they wanna ensure that it's, you know, delivering great, uh, performance.
It's great, you know, cloud service for their customers. And so they're generating a ton of application logs on their microservices based architecture, and they're generating north of 5 million events per second. They wanna have operational visibility to make sure that their cloud service is doing what it's supposed to be doing.
They want to check all the performance metrics, they wanna check for bugs. They want to, you know, ensure that they're delivering the best experience possible for, for their customers. And so they need to analyze all of those logs and metrics very, very quickly.
And this is about rapid iteration of telemetry where they could get insights very quickly and then be able to take action against those insights. And that's just one, you know, great example. So it's really like, it's a different world than kind of like your business intelligence, which is kind of your classical data warehouse world of, Hey, let's do a report for the chief marketing officer once a week.
You know, it's infrequent, it's large data, but it's relatively infrequent and performance isn't that important. Um, however, in our kind of space, the customers and users of Dr. Performance matters and it's performance at scale performance with high concurrency.
I mean, that's the name of the game that we're kind of participating in. Fantastic. Madine, why don't, why don't you describe for us a little bit about, um, you know, it's been a few years back, but why did you kinda spin out and say, you know, love to do it private project and all the great things you were doing with it?
Why start a company? What was your goal to, to do with, uh, imply? Uh, well, I think, uh, drew started like a lot of, uh, great open source projects from actually solving a very specific need in a very specific place.
Uh, we were, like, me and my co-founders, we were working at a company, we were working at the then kind of nascent ad tech, um, digital advertising that was like becoming like very transaction based, uh, selling individual like banners and individual clicks. And the volume of data that we were seeing in that industry was basically unmatched by anything else at the time. And now things, other things are catching up, but there was definitely like kind of a leader of big data and we built jewelry because we were trying to build an application that could do interactive exploration.
So like you click on stuff immediately, things come back to you, uh, you know, for campaign analysis for, uh, this kind of stuff. And we wanted to do it also consuming real-time data. So we wanted to, you know, when you have a campaign that's running, you don't care about, like, I mean, you do care about campaigns that happened like a week ago, but you care a lot more about like the campaign you just launched a minute ago and making sure it's like going as expected.
So being able to like, use real-time data and blend it together with historical data and then serve, uh, interactive queries on top of that. And we were very focused on that need, but you know, we released that into open source because that wasn't kind of like we released the database into the open source because we wanted community contribution. We don't, we didn't want to be.
Um, this is just some weird thing that just this one startup, uh, is doing. And when we, one of the amazing things about open source is that, you know, anybody can try it out and your users can find usage usages for your software you've never even think thought of before. Mm-hmm.
And it, you know, it turns out that this, these problems we were solving, they generalized over, um, all kinds of event data and all kinds of industries monitoring. Um, ba basically at the end of the day, uh, like anytime you have like a black, like something that I learned from, from the users, anytime you have like a complicated system, be it a website, a fraud detection algorithm, uh, an ai, uh, whatever it is, it's complicated the black box and the, the only way you have visibility into what that black back box does is by like collecting metrics from it and understand and then analyzing them very quickly and being able to like really dive into them. And we just saw, uh, a lot of people picking up J for all sorts of use cases.
Uh, we were doing, as I said, we were, we were coming from adtech, but, uh, the biggest, uh, you know, when we launched j the first company to pick it up was Netflix. Uh, that was really interested in using it for user experience because they were having the amount of clicks and subscriber, uh, interaction that again, was kind of unparalleled with, um, anything else. And they wanted to deliver the best experience possible.
Um, and just seeing how many different things this technology could apply to was really inspiring to, uh, kind of go beyond like where we started in adtech and, and just do our own thing, really just focused on getting people successful with this database. These are great examples too. I mean, I had spent some time in the entertainment industry, the, um, streaming business, and every one of those clicks, no matter what you're doing, that's all part of that experience that's being measured and analyzed and some of those great use cases.
So in 2022, you launched, um, imply Polaris, right? That that was your, or is your cloud, uh, hosted database service for Apache Drew? And then I know later you came out with an expansion on Drew's architecture called Multistage Query Engine, and I guess this is sort of the next wave of that innovation.
Tell us about, uh, what you're announcing today. Whoever wants to, to take that question. Yeah.
Uh, I'll, I'll take that. I think that it is very exciting. We're announcing, uh, that, uh, uh, a very important feature called, uh, schema Auto Discovery.
And, uh, it's just a fancy way of saying, uh, that JT can now, uh, uh, so JT provides performance by having is schema like that is key to performance under the hood. Uh, you need to, you know, uh, store the numbers with the numbers and store the strings with the, you know, you, you need to know exactly what type everything is to be able to compress things in the best way possible. And you wanna store things in columns because Drew is fundamentally a column or database that is mm-hmm.
Absolutely key to its performance. But, um, when people, especially if you're using in a streaming use case, if you're consuming your data from something like Apache Kafka or, uh, Essis, uh, you, you might, you the database owner, you might not be in control over what data gets in there, it's usually a different team that's producing the data from the team. That's like consuming the data.
You don't wanna necessarily, you don't have any alignment, you don't wanna have any alignment between, um, between those teams. You want to have, like the downstream team just add new fields as they feel they need to, and then you just, uh, add them in. And with this, uh, announcement right now, we're able to do that.
You can ask Juror to basically say, um, start finding new fields and then start storing them as what they are, like identify their types with that, pick the best kind of compression and, uh, and representation for those values and then store them. And that gives you, uh, the flexibility of ingestion that is only seen in document stores, kind of like something like MongoDB, uh, you know, where you just throw your data in and it just stores it as documents, like it, it just works. But the performance of a column store, uh, which is really our bread and butter, so mm-hmm.
Uh, now we've, uh, kind of optimized the, um, ease of ingestion to basically like, as, as easy as it could be with, uh, just being able to, you know, just point your point your juror at at kaka stream or whatever, and it will start writing everything out and, and we'll use columns and it'll be performant, just as performant as if you officially, explicitly declare that schema. And this is very exciting because it fits into, uh, this, um, high level narrative. Fantastic.
I just did a, a short, uh, research page paper recently on column data, high speed storage, and, uh, you know, you, there's, and that, I mean, having a specific structure that you're working with, albeit it's not rigid, but column data is one of the things, as you said, allows you to, you know, hire much higher performance than over maybe a general database SQL database or Mongo db, but there's also things like cardinality around how many, you know, values are expressed for different, uh, data that's part of that. Um, I'm curious, what does the discovery look like? What's the, what's it looking for to understand the data that it's processing the ischemic discovery?
Excuse me. Um, uh, so, so, uh, it's basically looking for, um, the, like, it, it collects columns for, um, specified chunks of time. Um, and then it, um, looks at the data, like looks at the data that came in and basically figures out the ideal type to store that data in and the ideal encoding compression, again, everything associated with that.
Uh, and it does that by looking at the, at, at the set of data that it received for each individual, uh, kind of key, uh, jewelry supports nested data that is arbitrarily complex. So a key could be like a recursive, you know, like, uh, you know, uh, message dot, uh, timestamp dot, uh, first time dot x or whatever. Uh, and each one of those gets looked at individually decomposed into columns.
Uh, one of the benefits, uh, as you will know of a column or data store is that having a ton of columns is not, um, it, it's not a problem, uh, because you only read the columns you need anyway. So, uh, we can create a lot of columns and Jarid is great at handling extremely high cardinality columns, but it's also great at handling extremely sparse columns. Um, and it, so it looks for, is this gonna be something that's high cardinality?
Is this gonna be something that's sparse? Uh, you know, does this key only exist in one event out of like the million of events that I've seen then, uh, stored in this way? And being able to pick the right way of, of storage, uh, according to this heuristics, is really key to guaranteeing the performance that our users expect.
Fantastic. Um, so David, as you've rereleased these new capabilities, you're putting this back into the open source and then providing it through your cloud, uh, offering. Is that how your, your delivery model for this?
Absolutely. So this will be released in Druid 26 0, and it'll be contributed back to the opensource community. And that's very core to how we think about innovation at APA for Apache jewelry within imply.
So we're an open source company that's core, and you know, our motivation in life is to ensure that more and more developers have really easy, uh, consumption and access to kinda this, this fantastic, uh, database, but also to provide a cloud experience on top that takes away the infrastructure and database management that gets associated running a distributed database like Druid. And so that's like the commercial side of, of what we do. And so kind of a two-prong go-to-market around open source and commercial, um, that really makes up what we're trying to do and imply, Yeah, solves the resource, the infrastructure problem, the, the ongoing maintenance and upkeep also, um, speeds up that time to value.
I can start using a cloud service very quickly. Absolutely. Am I still enjoy setting up something in my own environment, but mm-hmm.
You know, sort of like you touch it, you maintain it, right. So there's always that as well. Um, but Emmi, you're gonna jump in and say something too.
You, it looks like you I Was just gonna Yeah, I was just gonna say that, you know, um, I imply is a company that's been, that was founded by the people that created j and we deeply love the project, the open source project. Um, and one of our kind of like charters so to speak, is we, we don't, um, really have any interesting performance technology that is ever withheld from the community. We always contribute stuff to the community first and then actually consume the, like that upstream.
Uh, and we differentiate really, uh, on, on kind of like the very honest, like, you know, j is a big system. You usually run into the big scale and, uh, we provide visibility for j we provide a way to optimize, uh, not in any way. You couldn't do yourself if you knew exactly what you were doing and, and like, not by giving you some technology that you can't like, that you don't have in the project, but, uh, by just having the expertise, uh, and the monitoring tools that, that are proprietary, um, and run as a cloud service and or by providing you with a, a complete, um, cloud-based solution that, I mean, that completely takes away any need to even think about the word servers, for example.
Uh, so, uh, we try to differentiate on, uh, providing you with a better TCO and a quicker time to value And open source company that maintains the integrity and the op the true spirit of open source, which is fantastic. So where can folks go check this out, maybe get a cloud account or can do some hands on? How do they get ahold of Imply?
Yeah. Uh, the, there are two blessed places, uh, for, you know, your kind of single source destination. io, we have a ton of content there about Apache Drew as well as imply it's tutorials.
Uh, great content there. org and just download, you know, Jud right now and, you know, take it first then. Fantastic.
com. Is that right? Ao?
Ao of course. I should have said ao. Thank you to both of you.
Look forward to having you back. Keep us up to date as things continue to progress, and thanks for, uh, continuing the spirit, true spirit of the open source project and community you've been a part of for so long. Thanks gentlemen.
Thank so much. Thank you for having us.