Microservices and Edge Computing with pgEdge’s Phillip Merrick
After raising an additional $10 million in funding, pgEdge CEO Phillip Merrick explains why microservices and edge computing require a more modern implementation of the open source Postgres database.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Philip Merrick, who's CEO for PG Edge, and they just raised $10 million for a Postgres platform that they're gonna be putting out there, and it's a crowded space.
We're gonna jump into, uh, what's going on in the whole market for Postgres, but Philip, welcome to the show. Thank you very much. Great to be here.
All Right. I'm sure not a lot of people know who you guys are, so maybe start with some background on who the company is, and then what are you gonna do with $10 million? Sure.
Uh, well, uh, along with my co-founder and, uh, the rest of our team, we've been involved with Postgres one way or another for, uh, the better part, better part of 20 odd years. Uh, uh, we, uh, co-founded a company called Enterprise db, one of the largest, uh, companies in the Postgres space. Uh, and here with PG Edge, we're doing something different.
We're taking this incredibly popular, incredibly capable, open source enterprise class database called Postgres. And we are making it fully distributed. And what we mean by fully distributed is typically, uh, having it geographically distributed.
Uh, and it's not something that Postgres supports outta the box. And the reasons you would want to do this would be number one, uh, you have to have, uh, extremely high levels of availability. Your application basically can't go down.
Uh, you certainly don't wanna go down when, uh, your cloud provider goes down in, in a particular region that you might be using, um, and, or, uh, you have geographically remote users. And the speed of light means that, uh, it's actually taking quite a lot of time for the data to actually find its way, uh, into their browsers or their mobile devices. And so this, uh, gives us a low latency way of, uh, of getting the data closer to where the, the users are.
How do I synchronize all the different instances? I'm assuming if it's distributed, there might be one halfway around the world, one a quarter way around the world. And all these things have to have some level of consistency, I would assume.
So how does all that work? So, uh, we use a technique, uh, called asynchronous logical replication, uh, building on the core logical replication capability that's in Postgres. Now, since we're wanting to do this at a distance, it needs to be asynchronous and not synchronous.
And so that means that, uh, the system is going to have, uh, eventual consistency, but for the vast bulk of applications, uh, that is, that is actually okay. Uh, if you've got an application that needs to be working, uh, uh, across a, a global network, uh, then, uh, you know, you, your application's gonna be tolerant of that. You know, say one second difference to, to get the full system, uh, into consistency On the nature applications.
Historically, it was batch processing, and we had a local data center, and then we all moved significant portions of that into the cloud. But I feel like to your point, we are starting to process and analyze data closer to the point where it's being created and consumed, and a lot of that is in near real time. So, um, is the way we think about databases and the way applications are constructed need to change?
Uh, well, certainly they have been changing. Uh, if you, uh, think back to, you know, our, uh, move to service-oriented architectures and microservices, uh, this got us away from having big, uh, monolithic applications, batch applications, uh, uh, going, going even further back, uh, uh, and that actually lends itself to having more distributed applications. The database, uh, hasn't necessarily kept up with that, though the traditional relational database is, uh, monolithic.
Uh, and so what we're doing is making it possible to align that database strategy with the distributed microservices architecture that most applications are, are built on. Uh, then if you add AI into the mix, uh, yeah, we really need to get a lot of the data and the inference in particular, the real time predictions and decisions. Uh, you need to get that a lot closer to where the users are, because, you know, these might be time critical things.
Uh, if you are say, uh, uh, running a, a factory line and you are using, um, uh, AI with, uh, uh, a, a vision capture system to, uh, look for defects in products, you know, that's, that's a real time decision that you want happening, uh, uh, pretty close to, to where your, uh, factory line is. Mm-Hmm. Who manages all this stuff these days?
We used to have DBAs, and are they still in the mix, or is it being shifted a little bit more as we build these microservices and it becomes part of a app dev team? What does that look like? Well, certainly app dev teams, uh, take more responsibility for, uh, uh, for the applications, uh, in production.
Uh, and, uh, you know, the whole DevOps movement is, is, uh, uh, a part of that. Uh, but, uh, uh, we still have, uh, we still have DBAs, but in the cloud world, we tend to want things delivered as, as managed services and database, uh, especially, so most of the growth in the database market in the last five years or so, uh, has been almost exclusively in delivering, uh, database as a managed cloud service database as a service, if you will. Uh, and so there a lot of the, uh, uh, you know, admin tasks and, you know, applying patches and, and, you know, all those sort of maintenance things along with, uh, you know, taking care of database backups and and so forth, some of which the DBA would of course, done before, that's hap that's being done by your cloud provider or managed service provider.
And, uh, uh, one of the ways we offer our product is actually as a distributed database as a service. Uh, so, uh, your, your distributed Postgres cluster, uh, we can manage that for you in the cloud to relieve you of, uh, of some of those burdens. But, uh, that frees up the DBA to do some of the higher level things that are sort of more closely aligned with what the application developers are doing.
Mm-Hmm. It seems like we've seen some extensions recently to the capabilities of Postgres. It does more than your basic relational database, and yet a lot of times we see people say, we need databases that are optimized for specific data formats, or even say, a vector search or something of that nature.
Is there a balance to be struck here? Or what makes me decide to go one way or the other? Do I consolidate everything on one kind of database platform?
Or am I gonna be managing Postgres alongside a bunch of other databases? Well, you know, we could be biased here, but we think that, uh, Postgres is the most general purpose database, but it also, through the extension mechanism that you referenced, has the ability to function, uh, as a specialized database as well. And, uh, uh, you know, the, the vector, uh, capability, uh, is, uh, is a great example of that.
So, uh, you know, we had these vector databases popping up that, um, you know, were, uh, built purely for that purpose. Uh, but, uh, in recent years, extensions for Postgres have been developed to allow Postgres, native Postgres to function as a, as a vector database. Uh, alongside all of the other data, uh, the, the most popular of those is PG Vector.
Uh, and you know, when, uh, when PG Vector came out, uh, the performance was not anywhere near comparable to the specialized vector databases, uh, such as Pine Cone. Uh, but you know, as people have optimized that, uh, it's gotten to the point where, you know, the performance, uh, for a lot of applications is, is pretty comparable. And so whenever you've got the ability to, you know, manage your data in one consistent way with, with one database management system, you know, you kinda want to do that, uh, for efficiency, uh, for, you know, for, uh, uh, best optimization of your, of your engineering resources.
You don't have to have everybody trained on all the different databases you're using. Uh, the, the success of Postgres, I think, and a lot of people think is because it has had this ability to be extended in various ways, uh, the, the extensibility architecture really, really lends itself well to, uh, supporting different data formats, different modes of usage, uh, and, and things that, you know, currently we can't even imagine. Mm-Hmm.
What ultimately distinguishes one distribution of Postgres from another, there's a lot of platforms out there, a lot of options. So what should people be looking for? Well, uh, there are multiple, uh, distributions of Postgres.
Uh, but what, uh, we believe people want and are looking for, uh, uh, is, is standard Postgres. Uh, so, uh, you know, the standard Postgres, you know, you, you can download it from, uh, from the standard Postgres sites. And, uh, it has not been, uh, uh, you know, forked, it's, it's the plain vanilla Postgres that, that the vast bulk of, uh, Postgres, uh, uh, developers are familiar with.
Uh, and so, uh, that's, that's what we believe. Uh, and what we do is we use the, we use standard Postgres and the standard Postgres extension mechanism to deliver our distributed functionality. Uh, others in the market, uh, have, have, uh, forked Postgres, or they've tried to bolt on compatibility to Postgres.
Uh, and we don't think that's the ideal approach. What impact is AI having on databases and database management? I feel like for a long time, we kind of took the whole area for granted, but maybe there's a renaissance going on because AI is forcing us to kind of reexamine the way we manage data.
Exactly. So, uh, um, yeah, it, AI has, has certainly, uh, turbocharged, uh, the, the data infrastructure, uh, part of the industry for sure. Uh, you know, the demands were pretty high before, and, and, and now they're much elevated, of course.
Uh, and if you think about, um, you know, where the data comes from to, to feed into AI models, uh, a lot of that is coming from databases. But where it's particularly important is where, uh, you are doing things like, uh, uh, what's referred to as retrieval, augmented, augmented generation rag, uh, where you are actually using the enterprise data that you have in your database to, uh, uh, actually get better output from your large language models. Uh, and, uh, you know, get better quality responses, better quality predictions, better quality, uh, output if you are, uh, generating text for some application.
Uh, and, uh, in fact, it's, uh, it's through techniques like retrieval, augmented generation with, uh, with, uh, you know, your own data that, uh, uh, solves some of the inherent problems with, uh, large language models, such as the, the famous hallucination problem. Mm-Hmm. Will we, to your example on LLMs, but it seems to me like we're probably gonna focus more on these things now, come in what I call T-shirt sizes, small, medium, and large.
And maybe we're gonna focus on a smaller set of LLMs that are trained using a, uh, narrow set of data that's particular to a domain, and all that's gonna run up through these databases. But, um, what does the future of an IT team look like? Because today I got a bunch of data scientists, developers, DBAs, and, um, it feels like it takes a village to do anything.
So can we streamline any of this going forward? What is it gonna look like? Uh, well, I think that you are going to want, um, uh, folks with, with certain specialized skills on your team.
And, you know, and it really depends of course, on, uh, what your application, what your domain is. Uh, you know, are, are you gonna need to have, uh, uh, you know, 10 data scientists or, uh, or are you, are you gonna be able to, uh, get by with one? It's gonna depend a lot on, uh, on your application and the scale of what you're doing.
But to your point about, uh, uh, models and, and, uh, consolidating things, uh, I think, uh, already a lot of, uh, a lot of organizations have, have found in their practical application of AI that they can get the results that they need with, with smaller models that are trained on their data for the specific, uh, problem that they're trying to solve. Uh, and, uh, the other advantage of that is, you know, you, you don't need some giant compute cluster to, uh, to train it. You don't need a ton of resources for, uh, for inference in comparison to the, you know, the very large, you know, what we call frontier large language models.
Uh, and so there's that advantage, but if you can consolidate, uh, uh, around a particular technology base, uh, then we think there are some tremendous advantages. Uh, and again, we're, we're, uh, biased and wanting to see lots of things happen in Postgres, but, uh, our friends at, uh, at a company called Postgres ML with an extension called Postgres ml, uh, are actually able to give you the ability to run models, even train models and, and do inference against models, uh, inside Postgres itself. And so you're actually in the same process space, uh, as, as Postgres Postgres in your, uh, your other data, um, uh, that your application's using, uh, it all happens, uh, in that one place.
And you are not having to make network API calls to ask the model for a vector embedding. You know, you, you can all do, you can do all of that locally, uh, which is way more efficient. And so, uh, so yeah, we're working with our customers to, um, to build these, uh, uh, these applications that are, that are centered on, on Postgres, uh, not just for the traditional data, but also for, uh, what they're doing with AI and the models.
Last question. What does it take to migrate to Postgres these days? I mean, a lot of folks have existing relational databases and, um, some of them are open source and some of them are not.
But has that gotten any easier? Uh, it certainly has gotten a lot easier. Uh, you know, first of all, uh, we're helped by the fact that, uh, uh, with relational databases, uh, uh, we're, uh, we're all speaking the same, uh, language for data access, that being sql, uh, but there are, uh, of course, uh, different, uh, uh, variants, different dialects of that, if you will, and, uh, proprietary extensions, uh, which might be in the form of stool procedures and how those are done, uh, over the years, uh, Postgres, uh, has, uh, has evolved, uh, to, to make migration easier.
Uh, you can actually even run something very similar to the Oracle stored procedure language in Postgres own stored procedure engine, uh, is actually something that, uh, uh, one of our engineers, he here at PG Edge, uh, developed some, some years ago, ya Veek, uh, and then there were, uh, migration tools available, uh, migration tools available from companies, uh, uh, such as your cloud provider, AWS uh, uh, uh, Microsoft, uh, with Azure and, and their Postgres offerings. Um, and, uh, and there was some open source, uh, migration tools available as well. So, so this is not nearly the heavy lift, uh, that it was say, uh, uh, 20 years ago, uh, back when my co-founder and I, Denis Lucia, uh, uh, founded Enterprise DB to actually take on Oracle compatibility for Postgres.
All right, folks, you heard it here. Postgres is everywhere. The question is, is how to harness it in the age of AI and whatever other interesting near real time applications we're about to build and deploy.
Phil, thanks for being on the show. It's been a pleasure. Thank you so much, Mike.
Have a good day. Alright, And back to you guys and.