Simplifying Data Engineering at Scale: Christian Romming on Etleap and the Iceberg Pipeline
Christian Romming, founder and CEO of Etleap, explains how his engineering background shaped the company’s mission to reduce complexity in data engineering. He outlines how Etleap’s software integrates with cloud data warehouses, supports Apache Iceberg, and enables end-to-end data ingestion and modeling within customer-controlled virtual private clouds.
Transcript
Hey everyone. Welcome back here to Tech Drunk tv. I'm, uh, happy to have our next guest on, he's the founder and CEO of a company named Etle.
It's his first time here on Text tv. Let's welcome Christian Ramming. Ramming.
Is that correct, Christian? Yes. That that, that's right.
Yeah. Thanks. Thanks a thanks so much for having me.
Alright. I got it Right on the first try. It's a good day.
Christian, welcome. Thanks for being here. As I mentioned, you're the founder and CEO of that leap.
We're going to get into that leap in a second, but let's, let's start with you, you know, where, where was, what, what drove you to found Edle and kinda where have you been and you know, that brought you here? Yeah, yeah. I, uh, I'm an engineer by training.
Um, I've, I've been know, been, uh, living the silicon battle life for the past couple of decades now. And, uh, I, uh, I started ATLE about, uh, a decade ago. Um, uh, I was CTO at an adtech company before that.
And, uh, that's really, you know, there, there's, there's essentially an, an itch from there, but I'm scratching with atle. It's, um, um, you know, we, uh, I was, uh, you know, heading up engineering and analytics and, uh, I hired, had a whole little engineering team just dedicated to building data pipelines, you know, we needed to make our analysts and scientists productive. And, uh, that just ended up being a tremendous amount of engineering work.
So, um, so that's a, you know, the, the short version of this story is that, uh, I wanted to take all that engineering work out of, um, you know, something that companies needed to do. And, and so that's really what we're doing at atle. We're building software to, uh, take all the data engineering work out of, uh, data pipelines.
Interesting. And you say you started this a decade ago, so that was before, you know, chat, GPT, generative AI and all of this stuff that's going on today. Did you ever think a decade ago that what you, what you were setting out to do would one day perhaps be enhanced or, or augmented by something like Gen ai?
The, the world is so the short answer or no, the world is a crazy place. I mean, I used to, when I started, I used to have to explain to people what a data warehouse was, you know, so it's not, yeah, I, I, there's Also True, I, I know longer I have to do that. So, yeah.
That it's come a long way since then. Absolutely. So, you know, for those of us whom don't know at Leap give us the last 10 years in a nutshell, Christian, We have, uh, a software product, um, or we've been offering a software product that works great for, uh, data warehouses, uh, in the cloud, right?
Redshift, snowflake, Databricks, and so on. And so what it does is it ingests data from, uh, a bunch of different types of sources, right? Databases, applications, streams and files, uh, brings it all into the warehouse and then makes it easy for data teams to sort of model, uh, data to, to turn it into data products, um, within, within the data warehouse.
And so, uh, so yeah, so we've been on a mission to, to, uh, reduce that data engineering effort and also reduce the time that it takes to, to get all this set up right. Um, that's, that's another big part of it. You want to, you know, get your, your analysts and scientists productive, right?
Uh, as, as quickly as you can and, um, you know, have your engineers spend time on more, more interesting things than, uh, than data plumbing. Um, and I, I think something that's, that sort of set us apart in the, in the marketplace is we, we work with, um, companies in regulated industries, right? Financial services companies like Morningstar and in healthcare, right?
Companies like Moderna, um, they need to, uh, keep their data within their own firewalls, essentially, right? Like within their virtual product clouds. And, and so that's, that's something that our, our, uh, our software does.
Excellent, excellent. Um, you know what, before we jump into Iceberg Pipeline platform and all this, let me close the loop on that leap for people who may be, say, sounds interesting, I'd like to learn more. What's the best on-Ramp forum?
Yeah, I mean, uh, uh, it's, it's pretty straightforward. com, uh, si sign up for a demo. We have a, we have a great team that's, uh, uh, you know, ready to to, to, you know, show you exactly what you need to, what you need to see.
Excellent. Cool stuff. Alright, let's move over.
Now, you guys recently, uh, introduced a, a new, uh, platform, the Lib Iceberg Pipeline platform. Yeah. Uh, that, that's right.
So, um, I guess maybe a little bit of, a little bit of background. First, we, you know, we, we, something we've, we've heard a lot over the past few years is, you know, these data warehouses are great, but, um, you know, for, for different reasons, maybe for data sovereignty reasons or because people are adopting these new AI use cases, um, they, they want to, um, take on Apache Iceberg, basically use data, uh, use Apache Iceberg as the data foundation. And, uh, for, for, for data pipelines, there's a bunch of consequences.
When you, you, you want to run pipelines with Iceberg, there are things that really should work differently. So what we're, what we just launched is, as you said, the, the Iceberg Pipeline platform. Um, it's basically a, a pipeline there that's built from the ground up to support, uh, to support Apache Iceberg and everything that's different about it, right?
So it's, you know, still doing ingestion, still doing modeling, but uh, but also tying that all together, right? The, the dependencies between the different parts of the system are, are, have become really important. Um, use cases like streaming, which I'm sure we'll talk more about, are, um, uh, you know, are now possible.
And, uh, and also there's a lot more to do, right? You need to do table maintenance and, uh, um, you know, snapshot removal and, and, and that sort of thing. And so, uh, the Iceberg Pipeline platform is essentially a and sort of an all in one system that ties all of that together, uh, and can, and can run inside of your own virtual private cloud.
Excellent. Now, of course, this is based on Apache Iceberg. That's right.
Let's say someone out there says, iceberg Titanic. I never heard of Apache Iceberg. What, what exactly are we dealing with?
Tell us, you know, if you wouldn't mind, give us a little Apache iceberg background. So Apache Iceberg has been around for, for several years now. Um, you know, it's an open table format that's particularly well suited for, um, uh, you know, analytics use cases.
And, um, you know, it's not the only OpenTable format out there. Um, it has, uh, that there are, there are other ones too. But I think that the thing that's special about Iceberg is that, uh, a few years ago, the, you know, the, the, the massive vendors in this space, right?
Snowflake, Databricks, AWS Google all essentially decided to throw their weight behind this one, right? And so they said, look, we, we, we think this, this technology shows a great promise. We're gonna make our products work really well with this specific format.
And so that, that really helped, uh, future proof it, right? Because now it became, uh, much easier for enterprise to say, yes, okay, well, the, the, the big guys are behind it, we can adopt it too, right? And so that's, um, uh, you know, so that, that's, that's a, uh, a big part of the reason why it's now seen as a default choice, right?
So, um, so you know, the more concretely right, you can, uh, store all of your data within, uh, the iceberg, uh, uh, table format, um, which is, you know, built on top of Parquet, right? Parquet files that live in S3 with some metadata layers, and then you can, uh, you can use tools that you already use, right? Snowflake, Databricks, Athena, uh, BigQuery, right?
Um, these engines work really well with the data that's stored in Iceberg, right? So you can, uh, uh, you know, you can get, you can still get your fast queries, you can still support your, your AI use cases, right? If you're using Bedrock, for example, that plugs in very, very nicely.
So you, you know, you own your own data and you continue, can continue to use the tools that you already use in your stack. I love it. Excellent.
And you know, the nice thing is we, we see it with the foundations like Apache, like Linux Foundation, eclipse Foundation is it, it fosters this, um, cooperation where companies that normally wouldn't necessarily cooperate Snowflake and Databricks, right? You don't see them around the campfire singing Kumbaya too often, but you know, they all, they're all behind and they're all supporters of Iceberg, right? As, as something that transcends individual companies like that.
Um, is, is the Iceberg Pipeline platform available now? Pipeline platform is available. We have, uh, customers using it successfully in production, and, um, uh, as I mentioned, you can, you can run it inside of your, your VPC so that, you know, no, no data's flowing out, uh, you know, while, while your pipeline's running.
Very cool. Um, you know, while Apache Iceberg is free and open, I'm sure, uh, the, the, the, uh, excuse me, the, uh, pipeline platform, is it open source, not open sourced SaaS kind of thing? How do you, how do you do this here?
So there are, um, it's, it's, uh, proprietary software. Um, we, uh, integrate very tightly with, uh, you know, other, on top other vendor, vendor space. Yes, exactly right.
So it's, it's very tightly integrated with AWS, right? So if you have, you know, if you have a glue catalog, right? If you have, um, you know, you use Athena, um, if you have, you know, CloudWatch for monitoring, right?
It, it sort of plugs right into that. Excellent. And then did we mention, I don't know if we, did I forget, did we mention the, uh, ATLE website?
com? That's right. com.
Yeah, lots of, lots of information about the, uh, uh, the iceberg integration there as well. Um, so you can kind of see, see how it's different. You know, I think that, uh, one of the things that sort of sets this apart from, um, from sort of traditional ETL, right?
You could argue, oh, isn't this just, isn't it just just, you know, ETL in a new, uh, you know, for the new, with a new sort of destination? And, and I think this is, uh, you know, it's really, this really is something else. You know, ETL tools move data, but, um, but there's sort of a fundamental difference with operating pipelines end to end, right?
There needs to be, uh, you know, tying together ingestion and modeling. There are, uh, you know, catalogs involved and also these downstream tools that need to keep their own metadata, uh, fresh about the table. So in order to kind of get a, a, a system that sort of hums end to end, you need, you need, you know, whether it's ATLE or something else, you need something more integrated than a, than a sort of a traditional ETL tool that will sort of dump your data into iceberg.
But, but that's, then that's sort of it. Fair enough. Christian, we're about outta time.
I want to thank you for popping in here, man, the time goes very quickly, as you could tell. Um, before, before we end anything else, If you're gonna, if you're gonna remember one thing about Etle, I guess, um, you know, we, we, uh, handle ingestion modeling. Uh, we operate your, your pipelines end-to-end and we run inside your VPC.
Thanks. Uh, excellent. Thank you.
A really appreciate it. I appreciate it. Christian Rahing, founder, CEO of at Leap here on Tech Shark tv.
We're gonna take a break. We'll be back with more. Stay tuned.