SAP HANA High Availability – Cassius Rhue, SIOS Technology Corp.
Cassius Rhue, VP customer experience at SIOS Technology Corp., discusses how automation significantly aides in incidence response and DR with mission critical SAP HANA HA (high availability) databases.
Transcript
This is texturing TV. And welcome at the great pleasure being joined today by Cassius root cassius's VP of customer experience at scios technology Corp. Welcome caches.
Thank you. Thank you for having me on I appreciate the invitation and look forward to having a quick conversation with you. Absolutely.
Well before we get into kind of the the Deets as they say tell us a little bit about yourself and tell a little bit about science. Sure sure thing. So my name is Cassius Rue obviously VP of customer experience as science.
I've been with the company for over 23 years in the various capacities and currently focused on optimizing our customers experiences with our company and with our product in the field of high availability disaster recovery and data protection and replication. So I had a lot of challenges and experiences with the space. A lot of different technologies that we involved with socias focuses on taking some of the guesswork out of high availability and Disaster Recovery doing a lot of round application availability reducing the risk of downtime and then my team and my role in scios is making sure that both product services and customer experiences match that Automation and reduction of pain and the risk.
Or when when disaster strikes that doesn't give you a time window to say are we ready? Are we prepared? Let's kind of check things out.
It's like let's go. That's it. Let's hit the button hit the plan, you know make it happen.
Right when the the waters are lapping against the building or the fires are a block away or whatever it is. You've got to execute. I'm curious.
I mean you're in this industry, you know doing this so much. I'd love to hear what's your assessment of kind of state of the yard versus where maybe a lot of organizations are at today and trying to get to Yeah, absolutely, and maybe I could just piggyback on what you said there. Down time disasters don't wait until the convenient hour and say are you fully rested?
Are you on your second cup of coffee? Hey, you know the kids all go in the school is everything kosher and are you ready? Because I'm getting ready to have a critical application crash.
It's going to require you to get in there and figure things out. It doesn't actually work like that. It's more like 2 am 3am.
It's when your team, you know, one of your your key people is on vacation and he said hey, I've got a blackout period where just not going to be reachable by sell and that's when a disaster happens. That's when an application crashes a server fails and and in those moments. the last thing you want to be doing is scrambling to find a document right or a relying on some sort of a best practice that's couched in someone's brain somewhere and then going through a bunch of manual steps to troubleshoot figure out and then resolve your your issue with the application or figure out.
Okay, how do I actually fail these applications and get these applications access to their data? That's what we see a lot of in our in our experience when we are approached by customers for a trial or proof of concept or they're coming from a different vendor. They're actually looking to get away from having this manual set of steps that they're doing.
They're looking for intelligent automation. So I kind of break it down into a couple of categories if you if we have some time to kind of do please do. Yeah.
Yeah, so I'd say Basically, there are two two or three different categories of what I see a lot from our customers and Prospects, right? There's just manual set of steps, right? There's a run book or a process document that says when this thing happens this alert goes off and here's where you start checking and then there's a lot of scramble and what we've actually had from people is kind of them telling us their pain of in the early morning.
They can't find that document or their document was wasn't up to date and so it's talking about a version that's no longer there and deployed and and things get kind of squirrely when you're still tired. You're trying to figure it all out and you're in a panic and they're 13 Executives on the call trying to figure out hey, why are we down and why are we losing money and they want to know if it's going to be up. So that's manual right?
It's all kind of paperwork. It's a lot of back and forth about you know, when you create the steps who's responsible for validating. Are they there?
Are they followable? Can you pass that off to someone and then you have to remember to do testing of those steps, right? Hmm because you're a senior a CTO level architect.
And of course the steps one through 10 makes sense to you, but you're not going to be the guy that's waking up at 3 AM to handle a call. So that's the manual side of it. It's very cumbersome problematic and then there's an automated side right and then there's automation where you you've taken some of the guesswork out of the steps.
Okay, you've automated detection of a problem and maybe you even automated some amount of failover or recovery of services on a local system, but it's not a complete automation of both the application detection and the automation of failover. And so that's that's level one Automation. And then you have what scios provides which is intelligent automation.
So we we not only have taken the guess work out of detecting when there's a failure in a high avail. Scenario or cluster but we've we've taken the guesswork out of how to respond to that and automating intelligently whether this failure was application local and we can recover that application locally or if this failure was something that's so critical or component that just cannot be restarted on the local system and it's best to move to your standby system automatically or server detection. Hey, this server has gone out the lunch and we all know in this space you've had experiences where systems just stop responding.
They stop operating and the last thing you want to do is sit there and puts around and fidget with trying to either manually get things working on a system when you have a good active hot stand by available scios is intelligent automation takes that guess work out for you and it makes it so that we can detect all categories for applications. We have a lot of intelligence around what an application. Should one of the parameters it should be behaving with what do we consider to be out of bounds?
And then what's the best practice? Recovery for that application and that server. So if the application fails that's one class then we have an additional class of Automation and where we really see this take place Mitch is in our recent solution update for our sybase.
Sorry sap Hana deployment with scios life keep And then that solution we have taken it up a notch with intelligent Automation and expand it that to cover High availability in a cluster with the addition of disaster recovery options for third and potentially fourth nodes to really highlight the difference between our intelligent automation General to node automation that some of our competitors or other vendors have and and definitely it is light years ahead of any sort of manual operation that you may have. Oh I can imagine. Yeah, we could talk about the woes of manual I think every organization I've taken over here the Dr.
Plan updated. But that's you know, that's a PTSD subject. We've got other people around support team around us, right?
What's what's really interesting about this cash. Is that you talk about you talk about in terms of h-a-dr, right? It's not just ER it's not just ha which which sells which says to me in part you're thinking about automation of not only detection and Gathering data wanting to know what's going on right and back to somebody which I guess this one for my automation, but I'm imagining kind of reading into what you're saying is there are also fell over ha actions that can be taking automatically in the software.
So you aren't having to do you know, you may get a notice at 2AM, but it's already filled over for you or taken step two or three in the process. Am I going down the right path there? Absolutely.
You're you're spot on so in the intelligent automation space you're doing Lot of that detection and you are sending those alerts say hey, we saw this we saw that you're logging that so that there's an audit trace of explanation, you know in your experience, you know, when something happens at 3am, even if it's resolved you want to know why did it happen? What was was we are all in that space where even if an event was successful you want to know what was the root cause the triggered it? So it is intelligently logging that information, but then it's doing the next step which is saying hey, I understand this application well enough so I understand Hana well enough that I can not only detect these local failures or issues, but I can do something about it.
I can take the first step of recovery and whether that's you wanted to recover locally or if you've got because you're using HSR you've got an act of standby that's already got the data preloaded in memory that gives you that performance Improvement of not having to start from cold to reload tables and get things up and going so it's got that intelligence to say hey The data is in sync between these these two nodes or between these three nodes the data is in sync. Let me automate recovering from this critical failure on Node 1. Let me automate moving everything to node to and then I will restart my synchronization to node 3, so you still have that viable option should you have a second failure in that same data center or second failure of that node a second failure in a region if you're in a cloud and you've got that third note available, which gives you this higher level of confidence that you know when you see that alert at 2 am and you realize it's been handled.
You don't have to immediately jump out of the bed and say well now I'm down to a single system and I just even though it's handled and running. I better get on my peas and cues because I'm down to just one system left. You've got the comfort of knowing you've got a second system still available for you with our expansion of Hana to three nodes in addition to that Mitch.
We've added the intelligence to go over and figure out after That failure. What can we do to automatically bring that previous system that has failed back into the cluster, right? And that could be with our Hana application recovery kit.
It's just a module we call that plugs into our core intelligence that adds intelligence about the application to the core which has a lot of built-in intelligence about servers infrastructure networking device management and things so we've got intelligence to say we've automated failover in the middle of the night. It's 2AM. We sent you an alert.
We automated the failover you're still good. We've automated recovery of replication to another notes and you've got that additional protection and then we go about doing the work of cleaning up at other server and getting it prepared so they can rejoin the cluster and that gives you again that's just a level of service and quality and the product it goes beyond any of the manual steps, right and then goes beyond what you see. With a solution providers that are strictly two nodes in their primarily either focused on ADD detected one incident and I failed over but you'd better wake up and get a cup of coffee put on the put on your slippers and your robe and and puddle up to the desk or get down to the data center.
We've kind of taken a lot of that complexity away from from our customers and then giving them a good night's rest. There's the recovering from the recovery phase right? There's subsequent steps.
I'm curious say a little bit more about sap Hana and why that's such an area a technology that you're focused on actually in memory, you know high speed transactions. So it's being used for some pretty Mission critical analytics kind of applications, correct. Yeah, absolutely.
Absolutely. So in memory database lots of transaction lots of data, there's lots of workplace processes and there's a lot of data for business analytics that are being placed in those workloads that depend on the sap database and sap Hana and because of its performant characteristics, right? And because we see that it's so prevalent in a lot of these really large environments that are doing manufacturing Healthcare Transportation where where the amounts of data are massive and we're downtime right is extremely problematic not just from the business calls perspective, but from loss of data downtime causes disruptions and your manufacturing process shipping receiving billing on there are so many parts of the financial and work product system that depend on having these databases available.
Having them perform it enough to do the transactions that are coming in and then sap Hana. For example with the multi-target replication gives you this ability that if there's a failure you've got this this hot active system available to take over and while it's not in that role where it has to be the primary they've added functionality that allows you to run reporting off of that Target system offload some some extract load, you know processes on that backup system that gives you some added performance benefits on top of the fact that it's a in memory database with, you know, High transaction speeds and can handle massive sets of data. So you get all of this scalability benefit from using sap Hana, but Mitch and I think this is why we are concentrating using our experience, you know, the 20 plus years of corporate experience and knowledge and high availability and Disaster Recovery.
We bring that to sap Hana because it has to be available. It doesn't They no good if you have a database that's capable of all of this and it crashes and can't be restored. You know, no one realizes it needs to be restarted or the server fails and and then you've got all these manual steps to actually recover it and it takes it offline.
So we're concentrating there because we see the need for applications like sap Hana that operated this massive scale. They all need the protection of high availability. They need the intelligent automation to reduce risk.
If you let me just briefly. I'll give me a really great example of the benefit of our product and real life scenario right here for names. So in terms of having a massive database like sap Hana, we had a customer come to us or Prospect come to us because the incident they encountered was they have this massive database lots of data a lot of high speed Trend, you know, a lot of transactions going and mostly they were situated around manual operation, right database crashes.
They get an alert server fails. They get an alert they bring things up on, you know, several one fails. They got a manual process where they get an alert the DBA and the IT guy say, oh, yeah.
It's no one crash. Let's start on no two will make some DNS changes so clients and now accessing no to as their primary a lot of manual steps involved there. Well, they actually came to us because their situation involved having node one crash they ran on no two for some period of time and then no to crash.
Well your manual steps. What does it tell you to do? Well, no to crash, you know to crash go to node one bring it up.
DNS settings and and saw clients are actually redirected to node one and okay. Well, there's just a little problem that they actually uncovered the hard way is node one's data was about eight hours old older than no two because nowhere in their process. Did they have that step of Getting the replication from node to back to Node 1 when a fire happens, you know, there's no guarantee when disaster strikes that it says.
Hey, I'll give you time to do all of the other checks and balances and you know, I will be gracious to you if you forget step, right so they lost eight hours a day and what they came to silos for is that Automation and protection against those types of double failures and making sure that when they had failure one. Hey this highly critical system. It's restored properly and there's you know, risk of data loss.
Right? So those are the things that are intelligent automation guards against and that's what makes it so valuable and that's why we see a lot of benefit for our solution for any customer that's using these in memory systems like sap Hana and that need to get the benefit of going beyond what our competitors can offer and two nodes. Having a third node for scalability additional performance Offset, you know, you can offload some of those reporting some of those transactions that where you're just taking the data and running some analytics against it and then loading that into another decision table or decision database you now have Beyond just those two nodes to run those types of transactions.
And then even in the event of a failure of a single node, right? We've got the automation. We've got the protection.
We've got the intelligence to keep that going and if you have a true disaster, we've done the hard work for you so that you can quickly recover from that within your sla's and your recovery time objectives. So that's where we're in your example, you know eight hours a day to last but that really means most large corporations businesses Mission critical apps, you're really talking about millions of dollars of either the last revenue or cost business costs to the customer right pretty big deal that I mean, that sounds like Yeah, I mean I lost eight hours of data. It means a lot.
Yeah, not just today. This is the downstream impact of that loss. So I'm really curious talk a little bit about sort of how folks can get started working with your technology.
I don't imagine you just download a plug-in and poof you're ready to go. There's a process you go through to to really kind of intelligently automating this setting it up. What's that process look like?
Yeah, so great question and we've done a lot of work to make our installations of the product simple. com and you'll see some links there about trials setting up a proof of concept or being reached out or contacting our sales rep and what they will do is set you up with one of our experts on our customer experience team who can assess what your needs are help you make a plan for deploying it get you set up in the best way possible to achieve all of the outcomes right that you want to have for high availability Disaster Recovery using a business critical system like sap Hana. Fantastic Who cashes thank you very much for joining us today super important topic.
You know the time you need it is too late to go get it you need it. So it's there when you're when that when that happens so disaster strike. So be sure to check out sales Technologies and thank you again for popping on with this being here today Cassius and we wish everybody luck.
Hopefully no disasters, but they do be ready. Thank you. Appreciate the opportunity and thanks for having me on.