Major European Bank’s Enterprise Cloud Built on RING & S3 APIs with Scality
Aurelien Gelbart’s presentation at Cloud Field Day 23 highlighted a major European bank’s successful deployment of a private cloud built on Scality RING and S3 APIs. The bank sought to consolidate disparate storage solutions, aiming for a lower cost per terabyte and easier adoption for its users. Leveraging Scality’s offerings, the bank achieved these goals, fostering widespread adoption of the new platform and significantly reducing costs.
The bank’s private cloud architecture comprises six independent RING clusters across three geographical regions, each with a production and disaster recovery platform. Utilizing Scality’s native S3 API support, the bank implemented multi-site replication and object lifecycle policies to meet stringent financial compliance requirements. The implementation allows the bank to run hundreds of production-grade applications, including database backups, financial market data storage, and document management.
The success of this deployment is evident in substantial growth, with the platform scaling from one petabyte to 50 petabytes of usable storage within seven years. This growth brought new challenges related to performance and resource management across the many different applications. Scality addressed these challenges by improving software performance through reconfiguration and architectural improvements. The results are impressive: the bank now operates 100 petabytes of usable storage, manages 200,000 S3 buckets, and processes 300 billion client objects, achieving a 75% reduction in the cost per terabyte per year compared to its previous solutions.
Presented by Aurelien Gelbart — Sales Engineer, Scality. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/scality-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Hello everyone. My name is uh, Ian. I am a sales engineer in the French uh, sales team.
And, uh, I'm here to tell you about a great customer story that we have with, uh, uh, the state states, a leading European bank, but there was allowed to tell you that it's a major French bank that has a black and red logo type that on Google. You will probably find out who I am talking about. What was the customer need at the beginning?
Uh, this is a bank. So what do banks do? Uh, banks love to store money and to keep money.
So of course, here the goal was to save money by aggregating the different storage solutions that they were using at the, at the time. So they wanted to have one single solution that will become the go-to solution for most, if not all of their users. Uh, the two main ideas, uh, to reach that goal were to propose a lower price per terabyte.
Uh, this particular bank, uh, does not provide the internal storage solution for free. Uh, they bill their own internal users. They bill their colleagues based on usage just like AWS does for the, the rest of us.
Uh, so here, the goal of the s scaling platform was to propose a lower cost per terabyte per year. The second method to drive adoption of that new solution was to facilitate adoption. Uh, those of you who are customers at AWS may have noticed how easy it's to use their solution.
You want to start an instance, click a button, you want to store some data, click a button, you want to own resources, click a button. It's always easy. It's always available, and this is how they make money.
So the idea for our customer here was to facilitate adoption for their own users so that they did not have, they would not have any choice but to use the simpler, cheaper solution. This, uh, this bank built their private cloud on a s product. They use the Ring, uh, product, which is an infrastructure product.
Uh, and the Ring is a great solution when you need to store several applications or many applications with many different workloads on the same platform. Uh, the Ring is a product that scales very well, uh, almost limitless. And so this customer was able to move more than 2000 applications.
So this is to date, uh, today in 2025, they were able to move more than 2000 applications to one Ring platform. So it includes database backups. This is by far the biggest use case in terms of volume.
It includes financial market data. Uh, as you know, banks and other financial institutions are held accountable by governments and other regulation agencies, uh, to keep the history of the transaction that happened where they, where, where money, where money moved, uh, by who for what reason. And the history is stored on the s platform, on the s software based platform.
Another use case is Data Lake. So this is to aggregate data, to train AI on document management system. Uh, so I'm talking about bank statements.
I'm talking about official documents, things that need to be kept for a long time and are that are kept with integrity. Thanks to the core five, uh, uh, features that, uh, Paul mentioned at the beginning. Uh, fifth use case is developer artifact artifacts.
So like any, any, uh, IT company, uh, they have developers since they need to, to store data. And last but not least, uh, media, uh, media. This is a use case that we love at Scald.
Why? Because it's a high volume and low constraint use case. Uh, a video is going to be buffered to a content delivery network most of the time before being played.
So we don't have a high, uh, uh, we don't have aggressive constraint for that use case. Plus it's high volume. And since licensing is made on usable terabyte, this is great for us.
The architecture that they choose to, uh, for, for their internal, uh, cloud has actually six different platforms, six independent platform, two for each of the geographical regions that they have. So one, two in the us, two in Europe, and two in Asia. Two for each region because for each region, they have one production platform, which is the one that most, uh, applications are going to use and one disaster recovery platform.
So just before, uh, a question was asked about, uh, the difference between synchronous and asynchronous replication here for each region between the production and the disaster recovery platform, asynchronous replication is used. Those six platforms are integrated together. Uh, they are not ac accessed directly by end users.
Instead, uh, the customer build, uh, a software system around the s skill platforms that will direct the workload of a new application a automatically to towards the correct, uh, platform. So basically, it detects where the user is and is going to send the data to the correct platform, the nearest platform, uh, based on the location. On top of that, uh, I mentioned asynchronous publication between production and disaster recovery.
Uh, those of you who are familiar with S3 APIs know that to enable replication for particular bucket, for example, you need to make some specific API calls to enable the replication, to allow reading on one platform and writing on the other one. So it's not complicated for a tech savvy user for an IT person, but it can be a little bit, uh, hard for someone who is not used to working with APIs. And this being a bank, of course, a lot of people are not going to be a power user.
So what the team responsible for the s platforms did was to build as part of, of that integration, some check boxes. Do you want your bucket to, to be replicated? Check the box.
Do you want your bucket to be immutable? Check the box. Do you want your bucket to be version so that in case of human error, the file is not actually deleted, check another box.
And based on the box that are checked upon bucket creation, uh, some API calls are going to be made in the backend and are going to enable the features that the, that the end user needs. Uh, so now let's talk about a little bit, uh, about the challenges that, uh, this customer met. Uh, the challenges were pretty much what you would expect when you make a new solution available to a great number of users at a competitive price.
Uh, the platform grew, uh, grew from one petabyte usable, so taking into account data protection to 50 petabyte usable in seven years. Wow. So that's from 2017 to 2024, he said.
Wow. So, uh, you know, I, I was mentioning it earlier, the, the, the almost limitless capability, uh, scaling capability of the ring. Uh, this is a good example.
Um, uh, the, the customer that my colleague Nick talked just before, uh, the nest, they started already quite high and they weigh even higher. But here, for that particular customer, they started very, very low. Well, one petabyte is not that low, but comparatively low to where they are today, and they were able to do that online with no certain interaction.
Uh, I have a question, uh, regarding the seven year, uh, window. Uh, yes. So I assume that they start with one, uh, type of hardware, and then over those years, this hardware was changed, right?
Yes. And, um, how does s ring, uh, uh, handle this movement of data between old and new hardware? Thank you very much for that.
Great question. Uh, so, uh, s our first job is to protect your data and keep it available. Uh, so doing everything online is part of our development philosophy.
Let's, let's call it that way. So all operations are done online, and that includes software upgrades, that includes operating system upgrades, that includes expansion, and that includes hardware replacement. What we do when we need to replace some hardware is that we are going to list the components that are running on the servers that the customer wants to get rid of, and components by component.
We are going to move all the data, all the configurations, everything from the old server to the new one. This is done not even server by server. This is done in parallel for all servers at the same time.
So actually the not, not only is the operation done online, it's also quite fast. So we do it component after component with all servers in parallel. And when no more components are running on the all servers, uh, basically they pull the plug.
Thank you. And, uh, you, uh, in a rink you provide as well. Uh, this, um, let's say, uh, global load balancing, uh, networking mechanism or needs to be a external hardware like GSLB, I dunno, F five hardware, or you j or you have some implementation inside your hardware already At the moment.
Uh, most our, of our customers choose to, uh, to use the either HA proxy or F five load balancers. Uh, okay. Okay.
Thank you. Manage at the same time, high availability and, uh, load balancing. Alright, Thanks.
You're welcome. So, uh, 50 fold increase in seven years. And, uh, the cost of success with the ring is that you get unpredictable use cases.
Some applications are going to be, uh, very well suited for object storage. Uh, some applications that this customer, uh, uh, saw appear on their, on their private cloud were actually really part of our list of validated applications. So of course it went away, but some other applications we did not know.
Some other applications had performance constraint. Some other applications had very aggressive workload that have actually disturbed normal operation for other users. So it, it was really the, the cost of success, uh, to so many different application and it, it was almost a mess at some point.
So how did we solve the different issues that we had? Two things. The first, the first part was to improve our software.
Ring is a powerful software. It's already powerful out of the box, but it's also powerful because it's a very modular architecture. It's easy to scale.
Ben was mentioning it when he was talking about the leading US bank. It's easy to scale in different, in different, uh, uh, dimensions. You need more storage.
You just add storage. You need more compute, you just add compute. You need more throughput, you just add throughput.
So here we were able to handle more complex workloads with higher service quality by simply scale one, one part of our software, one sub component. We increased from eight threads internally to 32 threads internally, four times more, same hardware, same CPUs, and we were able to provide better service. So actually the, the cost of that software improvement for our engineering team was quite low.
It was only a reconfiguration. The second, uh, improvement access was architecture improvement. The ring is a product that starts at only three servers.
When you have only three servers in one cluster, you need those three servers to take part in all components. Monitoring is done by the three servers. Storage is done by the three servers.
Access is done by the three. Everything is done by everyone. When you have a 50 petabyte platform, of course it does not work anymore because some components are going to scale badly above five or 10 servers.
Some components are going to require more forward than others, so they need to be actually on the 50 or 100 or 150 servers. So, uh, we had to do some work actually. So sometimes we have to, uh, to figure out which components have had to run on the entire cluster and which components had to run on a subset of servers.
And by doing that, we solved the maybe, uh, three quarters of our issues. And for the last components that really, really need horsepower, we had the customer buy extra hardware, uh, for those components just to, to dedicate hardware to those components. Uh, but as you will see in the slides after that, uh, the customer was more than happy to buy extra power for a system that was working so well.
So the results, the number, uh, 100 petabyte usable, deployed across the six regions, 200,000 S3 buckets, 300 billion client objects. This is a, this is a really big number. Uh, it's about, uh, it's about, uh, something like zero point, uh, 5% of AWS worldwide, this number.
So it means that, uh, the single private cloud of one customer is already almost 1% of AWS worldwide, 1 billion client operations per day, which is not that, that much actually, because this is spread across so many servers, across so many regions that when you compare it to operations per second, it's not, uh, when you convert it to operations per second, it's not that high. But over the course of one day, it's an interesting number. And the most important number for probably is that they reduced the cost per terabyte per year by 75% compared their previous solutions.
And I'm going to be honest with you, the main driver for cost reduction is that before they were not using one system, they were using a mess of systems. Each team had budget and bought their own system separately. So it meant multiple maintenance contracts, multiple hardware failures, multiple, uh, uh, space saved because you can, you can never fill a system 100%.
So when you combine all of that, plus the fact that ET software runs on standard servers that intrinsically local servers, the result is a, is a factor of for, uh, in terms of customer. So it's, uh, one of the customer stories we are proud of. Are we on, um, you mentioned that there were, uh, two rings, I guess per region, one for dr one for production.
Uh, are, yeah, they, are they stretched rings or are they, um, asynchronous replicated rings? Uh, Production is a stretch. So it's, it's still available in terms, uh, in, in case of, uh, data center issue, like for outage or, or whatever.
Uh, so production is fresh and disaster recovery is as synchronous with production and is single site because, uh, they don't really care if they use the disaster discovery, the disaster car, disaster recovery platform, sorry. Uh, since production. So it's a mix.
It's a mix of production, stretch and, uh, and dr. Uh, asset. And, and did you mention that there are multiple tiers as well in the solution, or it's all single tier?
No, it, as of today, it's the whole single tier. But, uh, who, who knows what the future is made of? We'll see, so far it's one, uh, one level only.