How SREs Can Rock with GenAI – SKILup Days 2024
This session provides a concise overview of generative AI (GenAI) and its foundational elements. What does GenAI means to SRE? Why SREs should use GenAI ? My session will focus on real-life scenarios of how GenAI can play an important and differentiating role in SRE’s life. I will be discussing practical use cases with real-life examples – for example, Incident Management, Capacity Planning and Automated root cause analysis, demonstrating how GenAI can enhance the effectiveness and productivity of SREs456.
Transcript
Hello everyone. This is wna Pande. I'm from NDL Solutions Private Limited.
I'm an associate director, site reliability engineer for ndl. I'm proud to say that I'm the first certified ndl, SRE, in, uh, women, uh, certified, uh, uh, SRE in ndl. And, uh, I have been, uh, coaching and mentoring several, uh, folks to get certified on SRE as well as I help clients to, uh, adopt the SRE tenets and methodology and, uh, thereby modernizing their, uh, environment and business.
So, with that, let's start on the topic. So, today's topic, which I'm going to cover, is how SRE can rock with gene ai. Gene AI is a buzzword, as you all know, right?
Yes. Even, even for srs. It's, it's really a cool, uh, tool, which I would say, uh, and that's what I'm going to cover in this session where, uh, so I'll start off with, uh, a generic overview what genai, uh, uh, AI is and what gene AI is.
And I'll cover in short, about the architecture. Uh, I'm not going to cover in detail on that, but my major focus for this session would be how SREs can leverage Gene AI tool and how they can utilize in the day-to-day activities. As you all know, we SRE are, uh, let's say, uh, the software engineering.
Uh, if, uh, if you ask someone, a software engineer to do the operational task, what happens, right? And that's where, um, we sre tend to shift left, meaning we are supporting the infrastructures and, uh, providing, uh, looking at the client environment, maintaining their environment and stuff. Now, we are trying to move left where sis are trying to, uh, uh, look at from an application perspective.
So now we, uh, majorly look at how the applications, uh, work and, and our major focus now being in SREs, we look at the end user experience. So, uh, even though we say that our, uh, environment is up and running, meaning the server are up and there is no issues with the server, but ideally what happens is if the end user is ex while trying to access his or her application, if he, uh, if they are running into issues, it's a big issue for us. And, uh, VSR is look at, uh, the reliability of, I mean, we focus on the reliability of the application.
With that being said, what is reliability, right? It, it's beyond availability. In today's environment, we always look at the service level agreements, right?
Which majorly focuses on availability of your infrastructure. But now slowly we are trying to adopt, uh, the SRE, uh, meaning, uh, from the SRE perspective, we are trying to look at, uh, the reliability aspect of the application. How trustworthy is your application?
That is what we focus on with that. Let me quickly, uh, so we all know that, um, traditional ai, it has been in the system for quite some time, right? So when, uh, we are, so what is AI in general?
So we, it's, it's a basic thing which everyone offers know, but still, I wanted to make a quick comment here. So it's like, just like we humans do, interact, use our brain and do those, um, um, processing stuff. Now, if we ask this, uh, if we ask machine to do the same, that is what is, uh, we call it as an ai.
So here, it, it, uh, the traditional ai, now we have the gen ai, generative ai, which we call it as, and the traditional ai, traditional AI majorly focuses on processing and analyzing the existing data. It is not going to create new data for you, but it is going to look at the existing data and provide you with the information. What about gene ai then?
So gene ai, as the name suggests, generates new content. And, um, and, uh, it's, it's by looking at the data, it, it does not use the existing data and just copies as is, but it generates the new content. Uh, what is, what are the new contents here?
It could be audio, video, images, text, and even code. So even if required via, uh, gene AI writes the code for us, it can, even if we have a detailed, uh, paragraph or a detailed document, and if we ask generative AI to concise the document or summarize the document, it's going to do it for us. And that's way it always generates new content by looking at the data.
So, um, so here I have covered, uh, in brief about the architecture. So we all, uh, in Gene, I uses this, uh, large language model, um, as the base. Now, what is that?
We have several key terminologies here, prompt engineering, that is the key here. So we put in query, right? What is the weather today?
Now, it might not know, um, what, from which location I'm asking or from where I'm asking. So let's say if I put in, what is the weather today, Jenny, I will try to look at the current local, uh, area. And again, if you look at your GPS or it'll look at process the data and give us what is the current weather.
Now, let's say I have been in, uh, I'm in New York, and, uh, I, I would like to visit the places. What is the weather today? Just if we put that now, here we are.
Now you see the way I have put the query. So here in prompt engineering, it's always about, uh, context setting. We have to set a context.
Now, what I mean by context is, just like I said, I'm in New York and I would like to visit some places in New York. So what is the weather? So now we are telling the system that I'm in New York City.
Now, if I ask the weather, it automatically understands that I'm asking the weather for the New York City. So what if we do, it'll generate concise and related to my context, which I've said, which is about the New York here, okay? And the places to visit in New York.
So it'll start generating the data based on the context, which I've said there. Similarly, if you ask a generic question, prompt engineering plays a crucial part because the way we put the questions into, uh, our prompt is the key. The more concise you put the questions, the more detailed and more, uh, to the point or bulleted way you would get the, uh, what to say, uh, you would get the output.
So prompt engineering plays a crucial part. Now, what does la uh, large language models, so large language models, LLM in short, we call that as, so it majorly involves only text. It does not involve, uh, the di uh, the pictures, the other things, okay?
So if we have something, uh, if our data is purely on the text, uh, we use this LLM. And again, we have a lot of models. We have a lot of algorithms.
I'm not going to cover those in this, uh, how we train our system, how we train the corpus. We have enormous corpus of data and how the computational power. So in earlier days, we used to have the same things, but our, uh, resources, which I would say like the computationary power was quite low.
And now we have very enhanced GPUs, which is available, which we can utilize. And in that way, the query gets processed sooner and gives us the appropriate response. Now, fine tuning is another terminology, uh, which I would say, for example, now I'm interested in only healthcare related information.
Now, what I do, I'm from the healthcare sector, I am, so, uh, I'm looking for, uh, the data, uh, from the healthcare, what are the services which the healthcare provides. Now, what I have done, I have provided a context. So in this way, so in fine tuning, what happens is here we are not generating new data, but here we are trying to look at the data.
And again, we are, uh, making it domain specific, okay? So healthcare is one domain. And, uh, so what we are doing is, uh, having it as a domain specific would mean that it is going to only for the, we have sent the context, uh, as the particular domain.
And in our case, example, which I took is healthcare. Now, what, um, uh, the prompt will do is, uh, the gene, a I will start looking at only the healthcare related information, and based on the query, which you put in the prompt, it'll give us the appropriate output. Embeddings is another key terminologies, but again, it's, um, it's again, uh, how, uh, it's again, how we represent.
So it's like computers will not, uh, will not understand the text we put in, right? It always understands numerical data. So now what embeddings do is whatever the text we are put in, how it is represented, or how it converts into the numerical data so that the computer can understand, or the JDI could understand that is nothing but about embedding and depending on the output.
So even for, uh, this session, what I did was I was trying to do some research, and I got a detailed, um, uh, content, okay? Now, I wanted, uh, it in a concise way. So what I did, I prompted and I said that, could you put me, uh, this about content in a table of format?
I got it in just one second. You know, and I was so happy. If not, I had to copy the content, put it in the excel or in the table, and then I had to tweak.
But I got this work done by just a simple query, which I did, and I got it in a table, uh, table of format, which was needed. So in that way, in a table, I could see what is the, uh, input, what is the description, and what is the example? And then that really helped me in easy and even learning, you know, why, uh, we have this, uh, in a concise way and in a bulleted form, we can get to the point sooner.
So, and, and several, uh, I have just covered one of the, uh, use cases. So, LLM use cases could be purely, I mean, mainly on the text. So examples could be your q and a if, and your initial part of, uh, the, uh, the virtual assistance, which we have everywhere.
In any site you go, you will have a work, a virtual chat bot coming, popping up, right? Where it'll do the initial level of troubleshooting or asking the queries relevant answers. And based on our answers, it'll redirect us to the customer care, or it'll redirect us to the appropriate, uh, agent so that, uh, we get our, uh, answers, uh, done, right?
So anything involving text will come. And, and again, this example is whatever example I go it, it's not just those examples, but we have any number of, uh, examples there, but I just cited few of them. So with that, uh, I'll move to the next slide.
So what is Gene AI for sre? So gene ai, we could use it as an powerful tool. It's AI power tool.
Now, what I mean by AI is it may look at the existing historical data and it'll provide us the concise output In that way, it'll, uh, we can, um, look at the data, the concise report, and take actions or make decisions so that, uh, it saves our time. If, if it would've been a manual task, for example, I had to look at the last one month, uh, absorbability data from the Grafana tool, for example, I would go ahead and look at the data, try analyzing it in my own way, so it would've taken some cycles. For example, let's say I would've spent three to four hours in looking at the data, getting the data, analyzing it, and then making the decisions.
Why is that? There is a spike in, uh, the CPU at only particular time, particular part of, uh, particular day. Of the particular time of the day, yeah.
So now, GEI does all those for us, the analysis, it provides us a concise report. Now, we SREs can look at that and we can take actions automatically just by looking at the report. And, and, and in that way, we are saving our man hours, right?
Meaning we are saving our hours. So it's not that we are going to be free in those hours, but we as are going to focus on bringing in high value fo uh, activities. Like anything to bring, uh, anything we work on, bringing in improvement in the account is nothing, but, uh, it's good for the account, right?
Those strategic related, uh, activities, those d uh, those, uh, kind of, we can focus our, um, our tie and effort on those activities, and it reduces the burden, as I told you the example, right? So I spent four hours to get to analyze those data and then come to a conclusion. But gen AI would help us to do that, and we could get that in, let's say five minutes or less than one minute as well, depending on the speed and, uh, the computational power we have right now, why should we use gene ai?
We have, uh, I mean, you will be listening or hearing these words every now and then in my session, because this is something we do. So, uh, improved service reliability, as you know, we SR look at the reliability thing, right? So now, if there is a major incident or a could set happening, now, what will happen now, if we have gene ai, gene AI will look at this, uh, depending on the error, it'll already look into the database or it'll look into the web and it'll come back to us and tell us that we already have so and so issues.
And here, here are the solution steps. Now, we as ris, we are the technical leaders here, right? So we are the technical decision makers.
Just by looking at the steps, we would know whether these steps are right, right? And, and in that way we go ahead and, uh, use our, uh, analysis, and then we follow those steps and resolve the issue. Now, do you see, so I mean, we are, uh, the, I mean the expert in the troubleshooting, but Gen EI doing as the task of doing the research and coming us with a, a precise steps of what will resolve the issue is a, what is say it's, uh, an important, uh, uh, I mean, it'll help us a lot in troubleshooting, where if we would've taken 30 minutes to resolve the issue with gene AI helps, we could help it, we could resolve it in 10 to 15 minutes.
I'm just making up numbers here, team. But depending on the issue, it varies. Okay?
Now, reduce downtime. So since we have easy to access tools, it'll help to quickly diagnose and provide us, uh, the proper stuff, which I told you, increase efficiencies. So by looking at the historical data, gene EI can provide us insights.
And now what we do, we look at the insights, the actionable insights, and what we do, we take action. And so it's, for example, um, by looking at, uh, the historical data in last one month, these three servers, we had 10 incidents popping up. Now, what do you think of this, right?
This, if it was manual, we wouldn't even have bothered or cared to look at it. But now Gene AI is helping us with that report where it, it tells us in last one month, we, I see that there are three servers from where we got 10 incidents popping up from each of those servers. Now, what will be our actionable insights to look at those servers?
Why is that on those servers we are having this issue, correct? And, and why is that we are getting 10 incident tickets from those servers. So what we do, our major focus, our actionable insights, we have got the data now, now we go to that server, we do the technical health check, and not talking about the security health checks here, but the technical health check, meaning is the CPO good memory utilization, how is it or are there any memory leaks, high CPO, any issues, any events popping up?
So in that way, we try to look at, do the, um, uh, end-to-end stuff, and then, uh, come to a conclusion that, okay, uh, on this particular server, we had this, this, this issue. Now we have resolved it. And now those 10 incidents, which were popping up should not happen again.
So what we have done, so Jenny, I helped us with the information. We looked at the information, and we took action. Now, after the action, what we have done, we have improved the environment by reducing those 30 tickets for those three servers, okay?
So each server had, uh, 10 tickets, right? So we reduced 30 tickets in the next month. So it's a great task, a great, uh, job well done, right?
So similarly, so with that, now what I'll do is, um, in SRE we regularly use, um, uh, several of the keywords, okay? So now I'm going to dwell into the use cases of Gen EI for sre. So now we have several keywords you would've heard preventive.
So we tend to look at the issues and we work on ensuring that, uh, this event or this issue does not happen. So we are trying to prevent that. So patterns, we look at the patterns, repetitive patterns happening.
So at the example, which I gave you where, um, where Jenny, I analyzed that those three servers had that, uh, 10 incident tickets. What was that? It was nothing but a pattern.
So we identified the pattern that on each server there were 10 incidents, which were repetitive happening. And if it was happening for every month for or for every week, we would know by looking at the data. And that's where we are the key decision.
Uh, so we purely depend on the data. We SRE should know what is happening in the system, and how do we do that? It's because, uh, it's with the help of observability.
Now, you might ask, what is observability? We, it's, it's day in, day out, so it's nothing bad. Uh, we looking at, uh, the data using observability tools like Grafana, Prometheus, Dynatrace, and just naming few Splunk Datadog.
Um, but these, these helps us to, uh, monitor the environment. Even we have the, uh, opportunity to even monitor the URLs and, and identify a pattern or look at, at what time the URL went down. Since we focus on the, uh, reliability I spoke about, right?
First, we want to ensure that it is available. And then the, the next step would be on the reliability aspect. Similarly, prevent, okay, preventive prevent is the same proactive.
So earlier we used to be reactive based being an SME in the account. What we would've done, we would've, uh, the incident would've happened, and then you would go ahead and take action, okay? Meaning, uh, the cur it happened, and then you go, go and do the root cause analysis and identify what was happening in the system, and then you identify the root cause.
And we ensured that, uh, the issue does not happen. But what was this? We were reactive to that particular incident, meaning the incident had already occurred, but we took action only after that.
But we, SREs cannot be like that. We have to be, we are the pro, uh, we have to be proactive enough to predict what is going to happen in the system. You'll see a lot of peace coming up here.
So I wanted to give a name, but, uh, we'll give it some other time. So predictive, uh, we have to be predictive enough. So as, uh, so as we are proactive, we, uh, we, we, we even go ahead and predict what is going to happen or what may happen in the system.
Now, here, if I mean system, it could be the customer's environment or customer's configuration item or any server. So it's, I'm just trying to use, uh, in, in easier way for you all to understand as systems. So now insights, I, I spoke about the actionable insights, right?
That actionable insights really help us out in taking actions, right? Decision making, I told you right, by looking at the data, we take decision, we, we should have, I mean, SREs are given the authority, and that's one of the crucial skill, RE skills, which we should possess because we are the key decision makers and, and, uh, uh, and, uh, we make the decisions from the technical standpoint, and we come up with the how and whys. Why did I take that?
Uh, why did I make the decision and how did I make the decision? The how piece is always on, uh, the, the way or, uh, identifying or finding out why on how you took that decision. So you'll have, uh, data or you'll have evidence supporting why you made the decision and how you made that decision.
So we always go with the data, and that's crucial thing for us, right? Analyze, we, once we have the data, what we do, we analyze, and, uh, I, I gave about the root cause analysis. So these are something, uh, I would say an adjectives and nouns, which are relevant for us is to work on.
Few would be, uh, we, we can call it as a verbs here, alerts or something, which in order for us to be proactive enough, we have to set, uh, alerts at the appropriate threshold. So if that threshold is breached or if that threshold is crossed, what will happen? We, the system or the tool, we alert us and being a RS, as soon as we look at the alert, we jump in, we log into the servers, and we start looking at what exactly is happening and monitoring, I told you, right?
Observability, we have to monitor the system. And along with that, we look at the logs traces, and that's where the observability plays the crucial part for us. And that is another key skills if we are, uh, aspiring SRE or if you're already an SRE, you absorb it.
You would be one of your key trait being in SRE, because using absorbability, that is when we know what is happening in the system, and we go ahead and take actions security, that is one of the important trait which we position. We try to look at things from the security aspect as well and take action. So in that way, our environment is secure from all the vulnerabilities or any, uh, threats which comes in, uh, from the internet, right?
Reliability. That is our key goal, which we usually focus on high availability. So once we make sure that our system is available or our, uh, or our application is available, or the next step would be to make it reliable.
So now how do we make our system reliable? So that's where, in SRE terms, we have something called, we look at the SLI service level indicators and service level objective. That is another, uh, a detailed topic, I would say, I would say an advanced topic for SREs.
Uh, because once you start setting up the SLIs in the SLOs for your application, that's when you can start, uh, concentrating, meaning that's when you can say that you are focusing on the reliability. Because as soon as you set those thresholds, you will start getting alerted. And I would say that once you have the SLIs and SLO set, it's not that right in the first attempt, you'll be successful, but let's say you have to monitor the system and modify your SLIs and SLOs depending on your data, whether that is the actual, and once we confirm that, yes, this is the SLI, and this is the SLO, we can confirm or we can assure that we, meaning the SREs can assure that there would not be any MIS or credits happening.
Now, why is that? Because even if some, some of the T in your application goes on, let's say if the midway T is not responding, or if your load balancer is not responding, we would've set the SLI because those are the important control points, right? We would've set the sli and we get alerted as soon as something is, um, I mean we'll have that word, uh, out of, uh, meaning the deviation, right?
Which we call it as anomaly, something is wrong with the system, we would get alerted. And as soon as we get alerted, we will not discard those alert being a service. We are, we have to be proactive, right?
So we log into the system and look at why did we get that alert? And we, we, by troubleshooting, we get to know that, okay, this part of the component was not responding or the service was hung, or it was something, okay, we take action. So what, what we did, we got the alert and we took action, and now that system is up and running fine.
So we that alert what we did, we were proactive enough to look into the system and resolve the issue. And that's made what we did. So we avoided, let's say the similar, the same example if we did, if we did not have that alert set, okay, and the threshold was breached today, the middleware comp, so it was, let's say the slowness or the middleware component was going into a hang or a performance issue, could be a memory leak or high CPO or, or, um, hang, okay?
Now, if, if, if we did not get alerted, so that system, the performance issue does take some time to, uh, prolong, I mean, uh, to get into an actual issue, right? We did not get alert. And what we did, we took, we did not take action because we didn't know anything was happening in the system.
Slowly after three days, the same middleware, I mean the same application went down, now went down. What I mean by that is the end user tried accessing the application. It was very slow.
Instead of getting as soon as you could submit, instead of getting the response in less than one millisecond, it took, uh, it took 10 seconds to get the response. Now, what will happen? We are more concerned from an end user perspective.
Now what we'll do, what will happen, end user will raise a ticket. Right now in that ticket, he'll say, my application is responding slow. Or, so once our application is not responding, uh, it's hung in the hung situation.
Now, do you see the difference? And then what we do, as soon as, and again, since it's a critical application for the customer, he would've raised it as a MI or recruit set. Okay?
Now, what we do, I, all the leaders, all SREs will be on the call, hands on the deck, and they will start analyzing the root cause. Now, do you see, we, we said the SI mean we identified the SLI is, hello? And we said that, and uh, the first day says we got alert and we took action.
Now, if that alert was not date, the third day customer would've experienced the issue and he would've got alerted or, uh, uh, I mean, and then, uh, we would've worked on. So what was that in the first part? First example, we were proactive enough because we said the alert and we took action.
The later part of the example where customer was experiencing the issue and he or she raised the ticket, and then we looking at the issue and resolving that, what is that? We were reactive to the incident. Do you see the difference between proactive and reactive?
None. And GDI helps us in resolve, I mean, in reducing our effort to analyze the data, to look at the data since it provides the concise reports and stuff, right? So in, in that way, we are free, right?
So we can focus on bringing in the continuous improvement into the account. Performance is something we all, uh, even, uh, not even s but even everyone look at the performance of the application as well as in the, uh, of the system and ensure that, uh, we meet, uh, the performance is good enough and it responds. Maintenance is something which we have to do.
And strategist. So, uh, I I told you right, Jenny, I is helping us to reduce our time in analyzing, doing the data driven decisions and stuff so we can, we are free, uh, by, uh, from spending the actual time meaning, uh, the four hours which I spent in analyzing the data has ahead, and now I was able to take action in just 15 minutes. For example, now I have saved three hours and 45 minutes of my time.
Now, what I do as an SRE, I can look at the, look at bringing in the service improvement plan in the account and being strategic enough, meaning thinking from the long-term perspective and, and start bringing in some improvement in the account so that we are the strategist here, right? And visionaries are something we look at. What is our vision?
What is our customer's vision so that we could focus on attaining that vision and ensuring that, uh, their environment is secure and, uh, reliable, right? So this is the key part of our, um, deck, right? So now how can Ari stroke, I did cover many of, um, uh, the things here, but I want to, as I told you, right, Jenny is assisting us as two, uh, with a lot of, I mean, it is reducing our effort and time to analyze the datas, right?
So what we could do, we could look at the strategic activities, meaning our free time. I mean, not, it's not that we are free, but still we are trying to go ahead and, um, uh, whatever the time is JAI saved for us, we are going to use that for our strategy activities like bringing in the improvement plans, and we focus on innovation. Let's say if something is a toil, toil is something which we look at, right?
We try to eliminate the toil. Sari are are very good at eliminating the toy. How do we eliminate the toy?
We try automating that, correct? And again, automation is not the answer for all the toy related activities, but still, we look at steps to see how we can reduce the toy or how we can eliminate the toy and bringing in the improvement in the account. So that's where JEI assist us.
Now, how can we rock with JEI? Here are the examples. Now, as I told you, we look at the repetitive task, right?
Meaning the toy related activities. So what are, what is toy? First of all, any repetitive task, any manual task which you do in your environment or in your, um, in your account that those are nothing but, uh, toil related activities right?
Now, what do we do? Uh, toil could be a manual task. So you do, you do 1, 2, 3, 4, 5, 5 steps to complete a task, but you are doing it manually.
Now, what we do, we SREs look at it and, and again, we don't like to do anything twice, meaning repetitive, uh, stuffs twice. So that being said, as soon as we look at, um, this is what, um, what is say, uh, this is the manual task, which we are repeating every week or every month, then it's a scope for us to automate it. And that's where we go ahead and start looking at automating the task.
Now, how gen AI helps, it looks at, um, it, uh, provides us the analysis and incident reporting, and even it gives us the concise report where it gives us the steps it has performed. And if it's a manual thing, we can go ahead and automate them, right? And I have even cited some examples, and again, this is our day-to-day activities, which we have cited.
So we quickly identify and report the issue. So janea has helped us incising those, um, uh, reports, right? Uh, giving us what exactly is the issue and how to resolve the issue.
So it's, it's easing our job, right? Enhanced incident response. So now during s or during set, once, which happens, what do we do?
We look at the error message. If we are very tech savvy there, we know what that error mean and we start taking action. But let's say we are still learners and we are growing as an SRE slowly into the system.
Now what do we do? We look at error. We put that, uh, in our, uh, search, and we go ahead and search, okay, do the search, and then we look at what are the steps recommended, and then we follow those steps and resolve the issue.
Now, do you see now that we taking the error and going ahead and putting in the search and doing the search and getting the response. Now what gene I is doing gene I is automatically doing it for us, and it is giving us even that report. So it is saving our time even to troubleshoot the issues.
And even it suggests the solution along with the steps. Steps. So now there was some issues which was happening.
Now what will happen is it is even giving us how to remediate that if there is some issues happening, it'll, remediation is nothing but how to, uh, take action so that we resolve that, uh, issue which is happening so that it does not happen again, okay? And during service outage, we look at the database issue. So example is, for example, um, here we are suggesting, uh, so where gene AI is even suggesting a fix.
And if you are tech savvy, we, we know, yes, whatever the gene AI has suggested, it is a valid one, and we can go ahead and do that, predict and prevent failures. So gene AI even looks at the histor. I mean, not even, but it does look at the historical data and it'll come up with, uh, and it'll predict what is going to happen or what will happen.
Just like, um, Jenny, I I gave you that example, right? Um, that it gave us a report that these three servers, the actual insight I spoke about, right? Where these three servers, we had 10 incidents happen happening, right?
So now by looking at the historical data, now gene AI even can predict that in next one, one, this. So 1, 2, 3 is again, going to have three or five tickets popping up. And that too, from the database perspective.
Now, why database could be that database query is ticking today, one minute to process the request, okay? I mean process the query, but slowly it is taking, uh, two minutes to process. Now, do you see the difference?
It was ideal time is one minute, but now it is taking two minutes. Now, gene, I, uh, gene AI has the capability where it'll even predict what is going to happen so that it gives us the, uh, insights and we are going to take actions accordingly. Optimize resource allocation.
So capacity planning is something which is must. So we have, uh, black Friday and, uh, cyber Monday coming soon, right? And, um, so during those times, what we ideally do, I mean, um, accounts go ahead and, uh, scale their, uh, resources.
Resources could be database, could be storage, could be even the web servers where they increase the web server. Um, I mean, increase the number of web servers in the cluster or increase the number of, uh, storage devices so that during that Black Friday when there are, uh, uh, there are a lot of, um, traffic coming in, uh, it'll, I'm sorry. So it'll go ahead and, uh, take action.
Uh, meaning it'll, uh, give us the appropriate, uh, meaning, uh, if there is, uh, heavy traffic coming up, it'll, uh, it'll go ahead and allocate the resources from the new resources, which is added to the system. But now what gene AI do is we can even, uh, use the trend and, and it can even predict that what other servers or how much resources is needed in future and how much allocation is needed. So, so that is important, right?
Predicting from the resource perspective, from the capacity planning perspective, if it, if it starts predicting, it's a blessing in disguise for us, right? So what we do, we go ahead and take, uh, I mean, look at that recommendations, the predictions, and then we start taking actions accordingly. So we are ready and we never get any, so nothing bad, meaning none of the servers or the environment goes down during the Black Friday or the cyber Monday.
And how do we do that? Some examples are analyzing the traffic patterns to predict the peak usages. And sometimes now, uh, in India, we have, um, the festivals coming up.
So now in the retail stores like online business, there will be a lot of online shopping happening, right? So now all the online vendors will ensure that their system does not go down in that way. They can start taking actions, right?
Meaning, uh, so that, uh, they can allocate resources. So now Gene AI is even helping from that perspective, instead of we adding resources. And in the end, if it does not get utilized, it's a, it's like we have, uh, invested our cost there, right?
But here what will happen is, uh, gen EI predicts and we take action in that way. We scale the resources accordingly to increase to, um, handle the increased, uh, uh, load on the servers, right? So there are other things I spoke about, anomaly, right?
Anything which deviates from the normal behavior, okay? That is nothing but anomaly. Now, gene AI has the capability to look at the historical data and come up with identifying the deviations.
It can come up and say, this server was working fine until last month. Now we see that, uh, the server is taking more time to process server, meaning let's say let's take a middleware server, uh, for example, in middleware, let's take IIS server internet information server. So now in if I, I server was taking less than one millisecond to process the request.
Now gen AI does know that the default behavior of the i i server is less than one millisecond. Now, let's say something has gone wrong. We, we don't know what has gone wrong, but now it has started ticking.
Uh, one second to process the request. Now you see the anomaly less than one millisecond, and now one second. So do you see the difference?
Now, what we do, gen AI will give us that, uh, trend and it'll tell us that i i server is there is an anomaly, uh, which is detected for, for on the i i server because now it is taking one second to process the request. Now see, earlier we didn't even know what was happening in the system, but now here we are getting prompt. Here we are getting data, we are getting responses from gene AI giving us the decisions or giving us the action items, what is, or telling us what is happening in the system.
So we res have the relevant data, which is needed for us to take, uh, predict or take actions or, uh, or, uh, uh, work on those actions. Okay? So now by looking at the deviations for sure, we know that something is wrong with that IIS server.
So what we do as an SRE, we will go ahead and start, uh, troubleshooting that i i server and resolving any of the arrest, which is happening, streamlined documentation. Now, uh, I mean I have used this a lot to be frank. Uh, we have a big document, let's say 50 pages or 30 pages document, and we are reading it.
Now, what we could do, uh, we can just upload the document in Gene I and tell that that please concise this document and give us the bulleted point. I mean, give us the bulleted points from this document. It is giving us, you know, it's all about how we put in the prompt, which I told you, right?
So if we have the concise report or concise, uh, data, we, or we had, uh, looked at the logs, the incident reports, now Jenny I is insing or summarizing those logs and reports and giving us that data, what we can do, we can just copy, paste that and put that in our, uh, RC document or the blameless postmortem document, which we are going to, uh, do after our RMI, which will happen one once, one more thing. Team, blameless postmortem is another key, uh, tenet of SRE where, uh, we look at, uh, uh, we look at, uh, I mean, I mean we are not going to do the, uh, RCS in the traditional way where we used to answer five whys, right? Why did the issue happen?
Why did, uh, the servers go down? Okay? But here in blameless postmortem, it's all about infinite house.
So blameless postmortem culture helps us to identify the root cause of the issue and even take us, uh, ensure that that same issue does not happen again. So we work in that way, and here we are not blaming any system, process or, uh, the person here, but what we are doing is we are mainly focused on, so we are trying to create this blame based culture in the environment so that everyone openly comes to us and says that I did a mistake in the system, so let us take action. So as, so see, getting the information late and getting the information before something goes wrong in the system is crucial, right?
Because if someone comes and says, uh, one thing I have done, uh, I have, uh, executed this command incorrectly on the incorrect server. Now what I'll do as an SRA, I just log into the, um, server where, where we did the mistake and we take action so that even before the customer starts reporting that my server is down or my application is down as an SRE out of taken action, now that green based culture has to be brought in, facilitate knowledge saving. So today, let's say I worked on a or a server.
So what I do as an ideal SREI will share what I learned today. And again, each day is learning for us team. So today, if I worked on an incident, I will document it.
And again, Jenny, I is helping us. It is giving us all the details. So what I have to do, I have to just copy, paste it or even ask Jen, I, Jenny, I to just give me the concise report in this format.
In that way I can just share it with my team, with my extended team, so others, other SREs get leveraged or benefited from the document or the knowledge which I'm sharing. Because today I got this issue, tomorrow one of my colleague might get the same issue. Now, what he or she will do, they will remember that Ana had worked on this particular incident.
Let's look at what she did. Now they will follow, they will look at first the problem description. And again, since we are sre, we tend to document the technical steps as detailed as possible.
Okay? So in that way, we just, uh, look at what steps. One, the not took.
So the other SRA will take action resolved. So it's like instead of taking 30 minutes to resolve the issue, he or she would've resolved it in five minutes or 15 minutes. Five minutes is too, uh, what is say it is an, uh, good thing to have, but uh, it, it does take some time, right?
Yeah. Enhanced decision making because once we have the data, we can make the decisions and accordingly, I gave you the example, right? Where one of my colleague did come back and say that he or she had mistakenly done the error.
Now what? Uh, so I took the informed decision and I logged into the server and I resolved it even before MI was popping in. So that's how we make informed decision provided we get the data.
So chain AI is helping us with all the relevant information, which is needed, and it's making our life easier. So with that team, I want to say that we are the strategic vision. And so you see all those adjectives announced, which I referred, right?
And we are tech savvy leaders. So we are purely technical leaders here. We are the proactive change agents.
So anything happening in the system, we are, we are proactive enough to go ahead and implement that and embrace those changes. So with that, thank you team. Thank you for giving me an opportunity.
So with that, if you have any questions, put that on the chat and I'll answer them accordingly. Thank you.