Seamless Business Continuity and Disaster Avoidance: Multi-Cloud Demonstration Workflow with Qumulo
Qumulo presented a demonstration at Cloud Field Day 23 that showcased seamless business continuity and disaster avoidance in a multi-cloud environment. The core of the presentation centered on simulating a hurricane threat to an on-premises environment, highlighting Qumulo’s ability to provide enterprise resilience and cloud-native scalability. Brandon Whitelaw demonstrated how Qumulo’s Cloud Data Fabric enables disaster avoidance through live application suspension and resumption, data portal redirection, cloud workload scaling, and high-performance edge caching with Qumulo EdgeConnect. This allows the safe migration of data and applications to the cloud, ensuring continued access and continuity in the event of a disaster.
The demo’s primary focus was on illustrating the ease of transitioning data and operations to the cloud during a simulated disaster scenario. The process involved disconnecting the on-prem cluster and, using a device like an Asus Nook, accessing data seamlessly from the cloud. This seamless switch allowed government employees to continue their work at an off-site location. This was achieved through data portals, which enable the efficient transfer of data, with 90% bandwidth utilization. It demonstrated the ability to maintain user experience by removing the need to change user behaviors or adopt new protocols.
Finally, Qumulo’s approach offers high bandwidth utilization, and integration into a multitude of customer use cases, all while ensuring minimal downtime and data integrity during the process. They showed how edits made on the cloud could be instantly consistent with the on-prem solution. They were able to quickly and effectively restore data access to users after the storm, Qumulo emphasized that the architecture allowed businesses to be proactive, moving data to the cloud days before a disaster, reducing the reliance on last-minute backups and promoting a more flexible, scalable approach to business continuity, and with the upcoming support for ARM, and the focus on multi-cloud, Qumulo allows for a great deal of flexibility in how a business manages its data.
Presented by Brandon Whitelaw, Field CTO and Cloud GM, Qumulo. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/qumulo-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Uh, in this demo, uh, just conscious of time here, what we're gonna show is, okay, in contrast to that, when do I fail over and what's going on? We're gonna show, uh, connected to the system on-prem, in the data center like the client. We kinda shown that in that graph.
And then we're gonna go, okay, hey, you know what? We're worried about the hurricane coming in. No problem, no stress.
I'm gonna go ahead and stop using cluster on-prem. And I'm, I'm gonna send my, you know, government employees for the municipality. We're doing the crime center for looting, or looking at video surveillance.
I'm gonna send them to a, a safer location. I'm gonna just hand 'em a little ass nook like this. They're gonna go home.
I'm gonna scale out the performance of the system on-prem or in the cloud, because now it's the only thing handling all that io. I don't have that, that nice hot cache system on-prem. And then we're gonna deploy.
We'll show how the connection worked for the employee, then to the Nook to get that same data and then return it all back to the office afterwards. Uh, assuming everything was there and nothing got destroyed. But even if it wasn't, they could still maintain continuity and along, again, along the whole way if they wanted to be able to do, you know, PC of IP to virtual desktops in the cloud as well.
They could. This is just an alternative way to do it. There's pros and cons to each approach.
The the nice thing is when you say that you're very flexible and elastic and you can do a lot of different things, you end up with a lot of different variations of what people wanna do. And we had to kind of narrow down what we'd show. So anyways, my driver today is Mr.
Mike's meal. Uh, newborn father. Just wanna give congrats to him there.
Uh, he didn't, to be honest, he didn't really do the hard work there, but anyway, and he is gonna help drive while I do my best to narrate on this journey through the disaster impending disaster. So here I am, uh, I'm in the local office connected to my normal SMB share, nothing, no big deal there. I have local data that could be from a crime watch center or, you know, evidence or GIS images, whatever flavor of employee you want.
Um, and you know what? Now I want to connect this on-prem cluster into the cloud cluster so that I get to see all that data. So imagine this is a new, maybe a new, uh, build out of another, you know, government building, or they changed office, whatever the case may be.
So I'm gonna go ahead and establish, uh, data portal. Uh, today, we do this very easily through the CLI, soon those will be completely through the UI and Nexus to be able to just easily connect systems together by directory. I go establish and request the portal that I go on the cloud side and have to authorize it.
Now, the nice thing about this is this allows us to not just do within a single customer of this type of relationship, but actually extend to third parties. So you can imagine a situation where I have my data and my hub, but I want to be able to allow access, maybe read only access, or even read write access to a third party entity with the ability to revoke that access at any time and have no ability to get to that data with full audit along the way, and be able to extend those portals to other entities. So this could be collaborators for a special effects shop.
This could be the FBI coming in, or it could be FEMA or, you know, the Red Cross. They want to be able to get access to the, the camera data or the GIS data, the, you know, blueprints of buildings and be able to help, uh, find people if anything's happened. I can easily give them access in a very secure way that delivers a very high performance local like system.
So faster than I could talk, uh, which is pretty fast. Uh, Mike has established a, a data portal and boom, now all of a sudden a folder has showed up as if it was local the whole time. It's actually sitting in AWS and, uh, I browse into that.
I could see old project data, new project data. So just for the sake of, of kinda showing you in real time, we're gonna go ahead and drag a an image in this and write that. Now, this is just simply showing this is a full writeable read, writeable spoke.
The on-prem system is just a spoke to that, that hub in the cloud. And, you know, to show that we also didn't, uh, you know, prep everything ahead of time and cash things. This is, uh, what you guys today with, with Doug.
Can you tell me to suck My gut in actually Looking great. That was your good side when you turn this way. Okay.
I think you were throwing things in at that point and typical, yeah. Typical employee, uh, you know, uh, boss behavior here. Um, so we've made edits on this, uh, and it's great, great.
The sarcasm of engineering comes out strong. Okay, we're gonna go ahead and close this. Uh, and that we'll go ahead and save to the, uh, through this data portal.
Now it's, it's sitting in a w in AWS and now what's gonna happen is we go, okay, hey, let's go ahead and go back in and the hurricane's coming. Let's go ahead and shut down that data portal just for the sake of make sure we flush every single bit of data that's there locally. We're gonna ahead and disconnect that data portal.
Now, you don't really have to, uh, in fact, in real time, it's continuously streaming data block by block as it writes to the system, and it flushes very fast. Like I said, the key here is not only doing it at a block level streaming data versus file copying data, but also it's getting extremely high bandwidth utilization by using kind of a wan optimized congestion control in the network. We're getting upper nineties in bandwidth utilization versus single digit on a normal s and b share to the cloud.
I used to have customers AWS be like, oh, I wanna go move this workload to the cloud. Great. What is it?
Uh, it's just a s and b share for all my users, like home directories and file shares. I'm like, Hmm, okay. So all you're just doing is moving the data away from the users and then connecting it over SMB.
That was never meant to do that. That's gonna be a very bad experience. I wish I had this when I was there at AWS, but now I'm here.
I can take advantage of it and now help all my friends back there at AWS give it to all the rest of their customers. So now, oh, um, as you can see, I've disconnected the data portal. That directory is now gone.
So we saw it appear. Now we've seen it disappear as we established the data portal and remove the data portal question. Yeah.
So if the system goes down while you're trying to deactivate a data portal, the data is lost. Any data that was not yet written to the cloud is lost, is still on the on-prem system distributed high, you know, a highly available system as you would with any scale out nas. Now, if that, you know, connection to the internet was severed before you were ever able to finally flush that data and the data center was taken out, could you potentially lose data?
Yes. But here's the key. The key is because it's so easy to do this failover, if you will, which is not really a failover.
'cause the data was in the cloud the whole time. I don't have to wait until the last minute, right? I can do this days ahead of the hurricane and not worry about if and when and how last minute I go with it.
Can I get the, you know, if I'm doing the normal approach, right? It's can I get my, my snap diff iterative backup to the other Dr, you know, other data center in time before the hurricane hits? So it could be the same thing in that situation as well.
You'd have data loss for that backup as well. Okay, so now, so what, What's the typical, um, what right back time? Mm-hmm.
I mean, As fast as the bandwidth will allocate, It's basic basically what the bandwidth can support. Mm-hmm. Yeah.
Because of that, of how we've optimized that protocol under the, that that our own protocol of data movement at the block 4K block level with WAN optimization, uh, congestion control, et cetera. Instead of getting your typical two to 8% bandwidth utilization on a, like an SMB connection over the wan, we're seeing high nineties 95 to 98%. So that really what it comes down to is how much time do I have?
What's my change rate? And then you just buy your direct connects available or your bandwidth available to accommodate that. But again, in most circumstances, it actually doesn't matter much because you're talking about a situation where people are, you know, creating data during the day, then they go home.
It's still, you know, if, if the pipe was really small, it's a really remote office, it would still be pushing that data at night when they're not there in the background, then coming next day, no problem. Keep going on. We have actually banking customers that want this simply for video surveillance data or data creating at the branch where, you know, if something happens to that branch, you know, can it still hold that data in consistency?
And then, you know, because they were missing their backup windows between branch offices. Like, and so can I have something be moving data 24 by seven without having to take the system down for a backup window? And it's incremental, right?
Mm-hmm. Yeah. It's only changed block differentials at the 4K block level.
Yeah. Yeah. So very efficient, Simple math.
Uh, a 100 gig circuit is about a petabyte a day. Mm-hmm. So if your change rate, not your total data rate, your change rate is more than a petabyte a day.
Oh, he, he said over here I was. Okay. Uh, then your Yeah, I can imagine for the most of the use cases, not nine to 9, 9, 9 9 person is fine.
Yeah. Yeah. Okay.
Corner, corner cases like yeah. What, what, what would be the corner cases? Like, let's be honest, like, you know, like, you know, what doesn't work?
Like, you know, what are those corner cases where, so, Uh, what one, there's a lot that does, but to the point of just, you know, what doesn't, so far what I've seen so far is really, really, really high, uh, like physical security videos, fans for an entire airport or stadium where they just didn't have the bandwidth to be able to get 20,000 4K cameras to push live feed into the cloud. So it is like, you know, like thousands of files which are being constantly updated. Okay.
Yeah. By the way, that's also a fairly poor use case because, you know, it's a retention rate, a retention of like 30, 60 or 90 days, so you're just constantly deleting and writing. Yeah.
And you don't wanna pay APIs to delete and write. So, yeah. Um, so that's not a good use case.
Um, that's, I had one customer I was talking to ingesting six petabytes a day. Uh, that's six petabytes of satellite imagery. Yeah.
And the problem isn't that it's a problem, it simply just needs more bandwidth thrown at the problem, right? Yeah. Right.
So they, they can ingest that into the cloud, right? That would be 600 gigabits per second sustained, right? Mm.
And since we can do eight terabits per second sustained, you know, there's more than enough capacity. They just simply have to provision to it. Mm.
Yeah. It's really interesting, this, this, this reminds me a lot of the early days and discussions around Hadoop and kind of like, Hey, because ethernet isn't fast enough, we need data locality to the compute and the distributed hoop hadoop cluster. And then as networking got better, you could have disaggregated compute clusters for data lakes, which became the stand for everything in the cloud.
Every single cloud Hadoop data lake is distributed, compute from storage with a high performance network in the middle. None of it's direct attached. Hmm.
Um, and we're seeing the same thing, especially post COVID. The world is forever, you know, uh, distributed in workloads and compute availability. I mean, you talk about AI workloads, like you wanna be able to get your data to any a any compute available in any region of any cloud as it becomes available instead of just being confined to one.
Uh, so, you know, really we're, we're already, I would say at the point for the vast majority of workloads where the networks and bandwidth are fast enough with the optimizations we've done and how we move the data, that's one big thing is, you know, typical NAS was made to work in a local area network. It was never meant to do this. So we had to figure out how do I maintain application compatibility on standard protocols, but build a, a data network that is really meant for the modern era, uh, era and, and blend those two things together so you don't have to re-platform applications to be able to still take advantage of this.
And, uh, so yeah, networks have really caught up quite a lot, which is good. And it helps us really do cool things like this. So what's Mike show, what Mike is showing now is, is actually running entirely off of this little nook.
So that's the full scale out multi-protocol, you know, uh, uh, extremely robust file system running on just this little itty bitty device. It's pretty crazy. Um, I did ask to make it more impressive if I could run it off my phone.
Um, and the memory and compute worked, but we don't yet support arm, so I couldn't yet do that. So we had to run it on X 86, but one day we'll come back and show you. But anyway, so what's happening now?
Is it, the nice thing here is, is around, we talk about continuity when it comes to IT things and business operations and technology. But you know what a really important one is end user experience. Try going, doing change management and across a government municipality of, you know, 3000 workers, that is not a fun experience of, you know, where'd my data go?
How do I connect to that share? Or no, it's in the cloud. Now I have to log into A VDI.
That sometimes is more of a barrier than the technology is just getting people to change behavior. So the nice thing is no behavior change. They go home, they plug this in, and they connect to it just like they do in the office.
This is effectively the, the cluster in the data center shrunk down into a little nook. And now maybe they've flown to another part of the state. Maybe they're in a more safe location, they're not affected by the hurricane, and they can still get entirely to their data and still run those mission critical things they were doing relative to, you know, helping, uh, figure out what the recovery plan should be.
Or orchestrating what other agencies you could actually case might be Mount that cloud folder here if you wish. And we are just like, perfectly, is that what we're talking about? That's what we're doing.
So you know what, you should just come up right now. I don't wanna do that. If you could do the honors, what would it look like to then establish a data portal from this little device that showed only the local data?
Nothing in the cloud, we do the exact same thing over again. We go create a, a data portal. Now.
We, uh, request the data portal. We approve that data portal in the cloud. That's key, by the way.
'cause I, I can later revoke it if I want for security purposes to other agencies, et cetera. And that'll go connect to it. And again, the, the fun part is you could have, we have customers with clusters with over 80 billion files in a single namespace.
And yet this still works. It blows my mind where you can show that as if it's all local and it's 80 billion files sitting in the cloud. And now here on the, the little nook, I have the same, uh, data in the cloud that I had before.
And so, uh, I have the local home data, now I have the cloud data. We saw the old project, new project, just like before I can browse into the new project and there's Mr. Kle once again, you need to start using better pictures.
Yes. And so again, now I can open this. Uh, and I love how paint likes to jumbo size there, but it's the same photo again now.
But now point out the bald plot. Now running right off of this, we were running on a cluster in Seattle before connected to US west two. Now this is on the nook connected to US West two.
The data is in that lifeboat the whole time, right? I don't have to worry about failing over doing the last iterative replication. When should I do it?
None of that. It doesn't matter. I've engineered that continuity into the system at the very beginning.
So I have a question. I assume you were joking about running it on your phone, but is, uh, arm actually on the roadmap? I cannot confirm.
Deny, but that sounds fantastic. I did say the word yet, so I kind of gave away the the punchline a little bit. But uh, yeah, yeah.
Yes. The answer is yes. We look, what I love about some of our customers is representing these little grm quats.
If I, I I think you already threw one air. Doug will once again throw things at me. So the GRM is because our customers are never satisfied with any innovation we ever give them as their mascot.
Okay? They're never happy because they always want more and they get this behavior. 'cause we keep delivering on that.
Okay. Like in the customer experience or other things. And so yes, we had customers that say, Hey, you're 80% less expensive than any other file system in the cloud.
That's awesome. But I see these graviton systems over here, they're even cheaper. Can I use those?
And what about those same versions in Azure and OCI and GCP? Okay. And then we had some customers in the military set of things and otherwise saying, you know, Hey, uh, I have this this black box box system in a tank or an airplane and it's ARM-based 'cause I don't have the power to run normal X 86.
Can I run it there and stream that in real time back to a cloud for analysis on the engine and what's happening? Yeah. Okay.
That sounds pretty cool. We should probably do that. So yeah, some very interesting things coming in that regard as well.
So now we've edited that photo once again back there on this, this nook. We're gonna go ahead and, uh, show now. Okay, storm's over.
We're moving everyone back to the office. They can come back and, you know, in this circumstance at least let's hope that the, the whole cluster is still there. But again, the nice thing is that even in the worst case scenario, I can still provide full access to the data the whole time.
And, you know, if I have to stand up a new cluster, that's fine. I don't have to replicate any data back to it. I don't have to worry that I lost any data.
Uh, it's just a cache. It's just this, an extension of the data that's in that lifeboat in the cloud. And so therefore don't have to worry about that either.
So all we done, all we did now is just show how we revoke access to that data portal on the, the nook. Again, not something that's necessarily you have to do, but it's nice for the sake of just having clean, uh, visuals of what's happening between the two. And uh, then we will show that, uh, the flip back.
Here we are in the office. We've established the new data portal, uh, and once again back and we can see that open the new project, this is cluster in Seattle, connected to US West two and AWS and once again, it over zooms. But there's the edits once again.
So, so if you, if you kept the on-prem one up mm-hmm. And you went in right now and typed something and then flipped over to the on-prem one and opened the file again, it would instantly be consistent. Yep.
I did keep them up. Yeah. They're both actually running.
We can show right now if you want. Yeah. Do it.
Since we have seven, almost eight minutes, we can show yet more onto the demo. So made on-prem. There you go.
As he's showing this, go through, uh, any other questions of how that gets put back together, You know, you asked for a customer story. Um, yeah, Come on to video. We had a, uh, a Formula one team that said our car comes into the pets and it offloads the data quickly to a four node durable right cache in the pits, which a ton of data gets offloaded to these cars every time they go past the cloud provider pulls from that data, repro determines how to reprogram the car, and then transmits that to the car over 5G reprogramming the car as it's making lapse.
Wow. It's kind of interesting. Or imagine a, Well, to be fair, you only, we only helped that customer 'cause you wanted tickets to Monaco gp, but which just recently came back from, but and Montreal all next weekend Yeah.
As well. No, it's nice when you have really fun customers, uh, doing really cool workloads and you get some perks like that Every now and then. Yeah.
Um, but yeah, the, the, or imagine one like we said, where you have 1500 drones writing data into two cloud providers that are different cloud providers all using NFS and SMB. And imagine that's in an area that might be a target. And so you put it in two different cloud providers.
So you can tell for certain, if it's a denial of service attack against the data that's the spoke, where would you put the hub? Maybe Tokyo, where would you put the other spokes? Intelligence community sites, 17 different agencies processing data in parallel each one determining different factors, different face recognition, different equipment recognition, some doing blue force trackers, some doing different weather overlays.
And maybe only at the very top you can see all 17 different metadata streams and pull 'em together. But each agency having their own view into the data, or imagine you want to deploy something forward and you can boot VDI instances off of a read-only portal in the cloud and then write the data back into a write only one, stream it back and on and on, or follow the sun editorial. Take that.
Imagine that's not just a five megabyte photo taken on an iPhone, but it's a five petabyte movie and now you have edit editors around the world all working on their different parts of the film. Mm-hmm. Saving it back, everybody having access to all the work everyone else has done, but instantly revocable.
If you have a data security issue when the whole film's ready, you don't have to build a render far remove the parts needed up to three different cloud providers and render in parallel. You keep all that pre pre rendering data in case you need to rerun any frame. Mm-hmm.
And save it all back to the same central spot frame by frame, by frame, reassemble the whole movie and ship it out. You know, it reminds me of, uh, many, many years back I was with a vendor who obviously didn't have the most cloud forward strategy before my time at AWS, which was like upside down world comparatively. And the, uh, term data gravity, uh, was, was created by the marketing group.
And the idea was to basically say, no, no, no, no, no, you can't, you can't, you have to keep the data here in the storage systems. We sell you on-prem. And then also because it's here, you have to keep the compute here.
And okay, maybe you can do those like secondary tertiary things in the cloud and maybe separate workloads, maybe the shadow IT stuff can go there, but you, you can't, oh, it's just too much data. You can't move this. Right.
And what we're finding is, is that like, again, the, if you change how the data flows, if you change how you architect it to be able to account for this, and if you build a file system from the very beginning, cumulus, cumulus cloud, right? It, that's software differentiated to be able to have the same fantastic experience no matter what, where it is, what cloud, what region, what hardware, then you can start doing some really interesting things that change how customers perceive the art of the possible and how they orchestrate their data flows. And again, networking is exponentially growing.
And if you're extremely more efficient in using it in the first place, plus it's getting faster, you can really start, uh, thinking now more. It's, it's not about how do I bring things to my data? It's how do I make my data seamlessly available to any compute, any remote team, any talent you need it no matter where it is.
We have, especially post rider strike, post COVID, et cetera, we have media customers. It's actually hard to find one now who isn't using follow the sun production or distributed teams that don't have some teams in different regions or different talent pools because the demand for content is just so high. You support any public cloud, any of the big public clouds in this situation.
Are you, can you, can you stitch together a global namespace across multiple public cloud? Absolutely. One of our biggest broadcaster that has this, uh, on a working with us today started with us on prem, then deployed on AWS for audio editorial, et cetera, to New York.
Uh, and then had a whole nother group within the company who on their own independent talking in that first group actually deployed on our Azure Native Qumulo system. Uh, which is a integrated service with Microsoft that was co-developed as a first party service as, and then figured out, oh, by the way, you all are using us in all like on-prem and in two clouds. I can just data portal this together so you can collaborate, even though you have your preferences, you prefer those suite of tools in Azure, you prefer those in AWS, you still have infrastructure that's good to use on-prem.
No problem. And actually the nice thing also is we've used this, uh, not only for across multiple clouds at the same time, but actually as a migration tool. Yeah.
Because at the hard part about migrating is like, how do I, I have to go copy everything first. How do I maintain continuity along the way? When's the cutover?
When's the last fail? No, no, no. I create new system in Azure.
You're in AWS or vice versa, or OCII can make it read right on both sides simultaneously and you can move one app at a time and everything's still running on both sides as I move them across the bridge until you get there. And it helps with m and a, it helps with collaboration. Yeah.
So it really helps stitch all that together. Um, any any good questions though? Anybody have anything?
I mean, you were talking about infrastructure as code, which I think is to, to me a fascinating topic and the ability to use Ansible playbooks with Terraform to not only deploy but then instantaneously provision all the networking functions across the stack. The fact that in the on-premises environment we can run on the customer's, you know, IA certified version of REL 10 that they want with all the plugins they want. Oh.
Oh see? Yeah. Two Seconds.
There's a question. It better be a good question. I was like the, um, but running on the customer's IA assured version with all the plugins they want, they want the cilium driver so they can participate with, you know, have I surveillance integration, no problem.
Or do they want to use, you know, a different tool for security vulnerability scanning and have that integrated, you know, or do they want to have another particular microservices protection tool or distributed host firewall? No problem. Yes.
I can ask something. Um, do like any of your customers currently do any kind of like CICD automated pipelines? Like I'm thinking 'cause I'm always forever in GitLab Yeah.
Running runners on there. Yeah. So like what are you able to share?
Kind of like those scenarios. Like I can give you a simple one to start off with. Yeah.
A in a multi-tenant environment as you want to add a new customer as a managed service provider mm-hmm. That would could automatically through our APIs create a new folder. That folder could be mapped to specific rights.
That folder could be mapped to a specific active directory instance that was used as part of that MSP or if they then wanted to extend that to the cloud, that would make a call to our Terraform provider. Okay. Which would then execute and instantiate that cloud.
And then, you know, with the work Mike was showing in the CLI that is all in the API now and while we're building a UI on top of those APIs, being an API first architecture means I bet if I wanted to spend a weekend of, uh, rainy weekend in Seattle, I could probably hack up the Ansible playbook to auto create the portals right. On both sides of that so that then that customer could have a on demand DR environment and again, build that lifeboat before the ship took off. Exactly.
That would just be, yeah. Do I wanna one click add the service Yeah. To add DR in the cloud?
Yes. And as a managed service provider, the Ansible playbook runs Terraform executes, it's declaratively provisioning the cloud, it's running until it gets it right. And then add the portal that provisions to the API call you get a positive assurance that it's created and then we deploy some data to it.
Then if I was extending that playbook a little more, I'd probably write a series of pie tests, coverage guided tests based on the files to evaluate the files were reachable by the protocol set that I had instantiated. So I'd validate S-M-B-N-F-S S3, maybe even FTP, make sure it's reachable by the protocols I've enabled for that client. And then I'd write that back out as a set of documentation artifacts, publish that to the GitLab or GitHub instance for that client so they'd have a self-documenting Yep.
Of the results of that runbook I executed. And that would go right into, uh, you know, what my me as an MSP was running for my client. And to make it even more fun, you could probably do service instantiation through ServiceNow and ticketing through that and tie it into the complete system where the Salesforce orders executed.
The ServiceNow playbook runs maybe all integrated through potential, which would then make the Terraform and Ansible runbook calls by the time I was done building the system out.