NoSQL Performance and Security – Tzach Livyatan, ScyllaDB
Tzach introduces a major release with performance, resilience, and elasticity enhancements that help power instantaneous experiences with massive datasets. Tzach and Alan chat about NoSQL performance and security, the shift to database-as-a-service (DBaaS), the performance impact of Kubernetes, new AWS instances, observability, and more.
Transcript
This is texturing TV. Hey everyone. Welcome to another tech strung TV segment here.
I have it's a fair. He's a first-time guest on Deck strong TV. We've covered the company before it's a Scylla is the company, but our guest today is sock.
Olivia Dawn sock. Welcome to Tech strong. Thanks, and thanks for having me and thanks for a pronouncing.
My nana correctly appreciate. You know, I I work at it. I work at it.
I'm not perfect but I work at it anyway, so, you know a lot of folks out here. Maybe they've heard of Sila. Maybe they haven't maybe they're not quite sure what you guys do.
Why don't we start with that? Give them a little company background and maybe a little bit of your own personal background? Okay.
Sure. So let me let me start with myself. And so I have a background in computer science and I've been working in the industry for many many years.
I try not to count anymore started the developer then move to product management. I spent many years in the Telecom domain. At the startup and later it Oracle communication unit.
And after that switch to the nosql domain and database domain a test here DB. Actually still had to be the company or let me focus on the project or the product mainly the nosql database. and Similar in a sense the category there are many many nosql database out there each of them like have a special unique sweet spots seal of fall roughly into the white column databases.
You might hear of databases like dynamodb or pachika Sandra kind of similar to be in that sense. Hey, Sarah DB is focus on availability High availability meaning it's Can stretch from a three note cluster to a hundred node cluster, although you you probably don't need as many notes. and it's based on other databases can allow a very highly available database solution and let me give an example one of our customer have a three region cluster in three different region.
And one of these region actually burned down in a great fire. It was published in the customer name was was publisher don't want to repeat his name here. But one of the cloud provider which is a well no cloud provider gut is data center completely burned down but instead of deployment just continue to work as is because they have a deployment of 10 nodes in each region.
So 20 notes continue to work one region is burned down and they will able to replace it. So this is the highly available property of Celine action. And so Yeah, no children.
Yeah, no good. I don't. Yeah, so what I'm saying, no, no sequel have now a long history is 15 years or more than that of no sequel databases and the first generation if you will of no SQL database was very much focused on highly High availability and algorithm.
But performance was an afterthought some of them. Here it is. Well for a long time I used to think no sequel mad.
No security. You're right generation. The first Direction was like I don't want to trash out the databases because all database made a progress since then it's they didn't stay in place.
But a lot of them will Design in the beginning is a project and not as a product if you will and performance and security and other in administration and general was not highly prioritized back then. Generation of no SQL database which already took all of this. Performance Security Administration into the design itself.
So Sila it was a reinvent reinvention. If you will offer Patrick assandra in a sense that we took the high of the high availability properties, but reinvent and design everything from scratch in simplest class for performance. Yeah, and by performance by the way, people performance is a sometime is a hard to understand concept.
So I think of performance in two main metrics one is throughput and the other one is latency. And each of them can be optimized almost independently. So why through put why throughput is important if you can for example, give 10 times the throughput of another database which we do you can reduce the number of server that you're using by 10.
And you save Administration much easier to administrates. Of course. It's a much less expensive if you're working on AWS or other Cloud just Shrinking the number of server by 10 is a great Improvement.
The second parameter is latency. This is a little bit more tricky because it's harder to improve. While you can always spend more money and add more nodes or server to to handle you throughput problem.
You cannot do the same for latency. So because latency is very much coupled with our literature. And Sila was designed from scratching and approach it's called sharper core.
Which is a very I wouldn't go too deep into the technical part, but it's designed from scratchable for a good performance a low latency. So we are able to provide a less than one millisecond on average and less than five. Millisecond P99.
Regardless of the throughput so you can throw up to a million requests per second, which is a lot on one still a node. And still get a sub-millisecond latency. And this is very significant.
And if you want I can explain why significant but this is a great improvement over the first generation of nosql database that I mentioned. Yeah, I mean look I don't so our audience may not be familiar with silver but they're very technical and I don't think we need to explain to anyone how important both throughput and latency are quite frankly, right? Well, you know because today today more than ever it's about scale.
Right. It's a company's can't scale up and scale up economically right given the current market conditions, right? If you've got it.
It's scale up economically or die right that that's the you know, there's no room anywhere else or in it. so stock when we look at silly though. Right Apache Cassandra an amazing product an amazing project.
Not a product an amazing project. Of course, they've been several companies that have tried to monetize it and productize it and there's a couple of models, right? There's the hosted model.
There's the the support and training that you know, the usual open source kind of stuff and then this sort of an open core model right where you have freemium functionality. Let's talk about you know, well sill is special sources, obviously around latency and and you know, reducing latency increasing throughput but what other kind of in the business model there how seller go to market with? Yeah.
So as you mentioned we actually offering all of this real options, so we have still open source, you can go to to get up get the code competit yourself and we have many many users. We don't exactly even know how many because you can just download it. And that's definitely an offer that we are supporting and even encouraging we encourage people in the community to use Zila and contribute to Sila even sometime.
We have an Enterprise version which is a closest version closely following the open source, like 95% of it is the same code but it do have some unique feature for Enterprise. A lot of them are around security. For example will give you a quick example encryption at rest encryption address is the feature that we offering only in Enterprise and not in open source.
And there are other feature like that for performance Security in other and like many other we have still as a service or City to be Cloud which basically fully managed Sila Enterprise version. So it's the same Enterprise database but we manage it for you either an AWS or in gcp and you just consume the database and that's basically it and we we take care of monitoring upgrades security patches, whatever it means to manage a database. We do it all everything for you and In the last year, it's become more and more popular.
It's now the most popular still option out of out of this reaction Sila Cloud seems to be more and more popular and I'm sure it's not surprised for you. It just following the industry Trend everyone want to move to the cloud and I work we live in a sassy world. Right and so we can offer it is that I think people like to consume it like that because they don't.
You know, they don't they they pay it they go there's not that big upfront cost and you know all of that and and quite frankly they don't have to deal with much about upgrades and and you know patching and all that. It's almost it. It's all good stuff.
All right, I think we've done a good job laying all this out now sock. But tell us what's new what's news Well Taylor that's good questions. And because I have an answer for it.
So that's okay. Those are always the best question. Yeah.
0 which is our yearly major release. So we're very excited about it. And if I can squid the history of Sailor, which is almost eight years now in into three phases.
The first phase was about performance and architecture for performance, which I mentioned earlier the second was about complete functionality and we are now fully compatible with both Patrick and AWS dynamodb. And the third phase which are starting now is about innovating and go outside of the compatibility and introduce new features. 0.
There are many small one. But let me focus on the big one. The first one is we introducing a new consensus algorithm with raft.
And Sila basically is eventual consistency database. But the database and this rule for any database of two parts of consensus or consistency one is for the data itself and one is for the metadata. What is metadata of the user the node the node in the cluster the schema everything which is not a data itself.
And so far this data was propagated in the cluster through gossip, which is eventual consistency that that was fine. I guess a few years back. But today it's not fast enough and not consistent enough and we still have five or two we're introducing a new consensus algorithm for that that's mean from for the user perspective that you can safely update both the schema and the topology of the cluster in a safe way whenever we want to and what do I mentioned apology, for example, Ed notes.
Let's say that you want to scale your cluster from three notes to 10 notes. You can now do it in a safe way and very fast so they serve not only still open source later. but of course our own cloud offering which could now scale much faster and and even shrink if you want to all of these features including the performance gain also applying by the way to kubernetes.
We have a strong offering and kubernetes. And one of our offer user Palo Alto security company, I'm sure you heard of very use actually use Sila on kubernetes and we have we put a lot of effort because Varian kubernetes everyone is running on Cooperative this day it's not new but we made a lot of effort to make sure that still actually perform on kubernetes in a very close performance to how it's run on barometer on VM and this is not trivial but they happy to say that we achieve that and we are very happy with that. would say I would say that that is one aspect of the Improvement.
The second aspect is performance. Will keep although we are already that the no secret with the best performance out there were keeping pushing the limit on performance and I'm happy to say that we achieve a lot of improvement in this release and in particular we work with the AWS to optimize seller for the new iPhone instances. I For Eyes the The success of the I3 and I2 family which are instances at AWS offer for databases.
So we work with them in particular on the ifori we got Surprising result on this instances, which is more than two and by two performance compared to the I3. So for the same number, of course, you get more than two two and a half times the throughput which even lower latency. And this was a combination of a great Harvard Hardware that AWS offer both the storage and network and the CPU and Sila was designed from scratch to work with this high-end instances.
0. We improved safety and elasticity with raft and continue to push and performance with the eye for eyes so that I would say the two main thing on on this release and by the way, all of this feature are the way that we are rolling between our version first we go. We put the feature open source.
let it stabilize a little bit then we Back 40 to Enterprise and then Betport to the cloud. So we are now in the first phase of it and in a few weeks, it will roll back to Enterprise and to the cloud. Actually, excellent.
So I worked the kubernetes that you talk about two and a half times performance sleep and that attracts your attention because it's such a eye-catching. You know number but the kubernetes functionality is important because look, you know, I got back from cool card. I guess maybe two months ago now.
But this whole cloud-native stack with covid center of it service match containers, obviously. You know microservices. It's really become the new compute stack right in many play.
Especially if Greenfield if you got a new application, you know that you're not just updating something from before. 75 80% of them are probably using this new this new stack of kubernetes and and Cloud native. So It's an important and important thing.
Let me ask you another question speaking of cloud native. So observability is all the rage. Everybody wants to get their hands around, you know, better insight into how my apps performing and so forth.
The data requirements are massive and growing all all of the things that it seems sill is perfect for right late through put all that. Companies using seller in observability situation. So you are actually we are actually playing the to side for for the for the story because we are a monitor and our own application.
So we generate our silic Cloud deployment generates a lot of metrics and we also act as a backend to collect metrics for example from iot applications. And so Sila is a very good match to keep iot information. And so but we also our customer ourselves because we do generate a lot of metrics and we are using a stack of grafana like everyone else I guess sure it present all of this metrics and we connect Sila directly to garofana in some cases to quickly export all of these metrics and and visualize them.
And by the way, we're presented at kubecon as well. I don't know if you saw presentation, but we was there as well. No, they had me stuck in the broadcast with doing interviews.
I was doing 20 something interviews a day there. I never got off from the my little area that I had to set in on there and you know, it was all over the I mean you you were there you can't with huge and I I never I never got to see more work the floor a lot, but hopefully in Detroit when I think it's Detroit in October November will be We'll be back at kubecon maybe right over there as well. Yeah.
Did we? lose it and I think we covered the main point and I I do want to mention that we are working with a lot of companies and everything. I say it's public in our side.
So I'm not exposing anything new but companies like Comcast Disney hotstar and other who want to reduce latency and save costs with high throughput our increasingly using Syria because of course some bias, but that's that's the only no sequel database out that really really work with both High availability in high performance. There are many different solution out there. But this combination I think is unique to Sila.
Agreed so I want to thank you for coming on our show. We ran a little over but that's not unusual here. I hope to see you soon.
If not a group con. But as I mentioned I'll be yalla devops next week. Maybe if you're in the area stop by.
I wish you well and and you know continued success with seller and to all our friends. It's Allah. Thank you.
Thank you very much for having me and have a safe flight. Thank you. Alright filler ciller Cilla DB here on Tech strung TV.
Go check that. com check it out. We're gonna take a break on take strong TV.
We'll be right back.