Shift Right Testing – Debunking the Myths of Shift-Right Testing – DevOps Unbound EP 32
Shift left has more than its fair share of admirers and advocates, as DevOps and DevSecOps initiatives aim to move responsibility for building and securing applications as early (or as far left) in the development life cycle as possible. But what about its more traditional (and much maligned) cousin, shift right? Shift right, or testing in production, was once taboo and is still discouraged by many. But testing in production is becoming a more common occurrence. The main goal of shift right testing is to spot and correct any issues that may not have been detected during the preceding development process. So, why does it have a bad reputation? Hosts Alan Shimel and Mitch Ashley are joined by Chris Riley (HubSpot), Nora Jones (Jeli) Paul Bruce (Tricentis) and Shamim Ahmed (Broadcom) as they debunk the myths of testing in production and discuss why it is important for your business. Join us to learn the truth about shift right testing best practices and trends for maintaining system integrity.
Transcript
Hi everyone, welcome to another devops Unbound for those of you who are not familiar devops and bound is a semi-weekly. Show where we delve into different areas of devops. We like to explore all the nooks and crannies.
Around devops and you know devops is a pretty big area today. So there's a lot of nooks and crannies crannies. We do do this show as I said every other week and then about once a month for those who are interested.
We do a live version with a live audience of this show where we let you the audience kind of Drive the discussion via chat and so forth today's episode though is not with the live audience, but we are blessed to have an amazing panel of folks here with us. and I'm going to introduce them here in just one second. I want to just first of all throw a big shout out to our friends, which I sent this though.
So I sent this is sponsored devops Unbound for going on two years now. and couldn't ask for a better partner and a better sponsor to work with always happy to be doing work with them and many thanks to them. Also Today Show is entitled debunking The Myth of shift right testing and we're going to jump all into that.
But first, let me introduce you to what I think is I said a really really great panel. I'm gonna start off with such Charming Charlene Ahmed I tried and then of course when I said it I got tongue twisted shameen. Welcome.
Why don't you introduce yourself to the audience? Thank Ellen shamim Ahmed city of for develop Solutions and broadcom. I'm responsible for strategic Innovations in our products also act as a trusted advisor for many of our key clients as the navigator that devops from summation Journeys, very passionate about the shifting right on test.
So we'll talk about that. Absolutely, and thank you. She mean I appreciate it.
Excellent introduce you'd actually he's not old but he's an old friend of mine and I'm thrilled to have him on the show here. He's been too long since I've had a chance to have my friend Chris Riley here Chris. Go ahead appreciate that and I have a beard now, by the way, since last time we shot a name is Chris Riley.
I am a senior manager of developer relations at HubSpot. I really lead our advocacy team been in the devops space for a long time and I'm obsessed with the idea of optimizing delivery chains and visibility. So excited to be able to talk about this.
excellent next up. I want to introduce a return repeat. Uh guests on our show.
I enjoyed the heck out of having them on these great guy Paul Bruce Paul. Welcome if you want to introduce yourself, thank you. Yeah, my name is Paul.
I'm head of incubation at trisons. I get to work in some really cool stuff work with some really cool people get to learn about our organization on a day-to-day basis and hopefully produce some really great inventions into the industry. I also am Community organizer of the devops.
Today's Boston. One of the core organizers there Boston devops meet up slack group. I run Ali fasts every year hopefully next year as well and really just care a lot about engineers have an easy time of things but also learning and improving themselves and those around them so I get to do that as a full-time job and as volunteer as well, and I'm really great to to be on here.
Looking forward to maybe seeing you a cubecon as well in Detroit. Absolutely. Well, you'll see us there we're broadcast.
Life. All right our final. Panel members kind of a guest special guest star for this panel.
It's our first time here, but we are thrilled to have her and it's one and only Norah Jones. Hi Nora. Why don't you introduce yourself?
Thanks, Alan. Yeah, super happy to be here. I'm Nora.
I'm the founder and CEO of jelly jelly is a Incident Management platform that particularly helps focus on the human aspect of incidents and understanding what led to incidents and surfacing patterns in ways that you can use to improve yourself in the future. I have a big background and testing and production. I have been in reliability my whole career and I've spent a lot of time with chaos engineering I think games evolved a lot since I first started doing some of those things and I have a lot of thoughts on testing and production.
So I'm really happy to be here and share with you all. Thank you. Thank you Nora last but not least.
He's not really a member of the panel. He's my co-host and partner and Fred Mitch Ashley Mitch. Why don't you introduce yourself?
As always great to be with the panel which actually cicchio with Textron group and principal with tech strong research. I too have been doing testing in production most the time it was due to very tight product deadlines not do intentions, but no we're talking about a new kind of testing to production. Yep.
Thanks for joining us Mitchell. All right team. You as I mentioned the title of today's episode is debunking the myths of shift right testing man.
If I had a nickel for every time I've used the word shift left over the last eight years. I wouldn't be here doing this show. I'd be out on my boat and some South Pacific island.
Here I am. I didn't get a nickel for every time I had shift left but shift left is becomes almost synonymous with devops whether we're talking about shift left and security or shift left in testing just shifting left in the whole software development life cycle has become In a common place and part of what we do when we talk about devops. However, Maybe I don't know two years ago.
We started hearing about well. When we're shifting left, we got to remember about shifting right and I heard this initially with security. Now when I started hearing it with testing and I'm like what the heck is shift right testing?
Oh, wait a second. It's testing in production. You mean the way we used to do it?
and That's kind of my feeling on it. But is that just the way we used to do it or it's shift right testing a new animal all together? You know, what do we I think we got to Define it before we start debunking the myths of it, so.
You know what? I'm gonna at Paul I'm gonna ask you to kick it off. and then panel, feel free to jump from there Paul what do we mean?
Well, I think there are some paradigms that survive transition and want to say transition. I also refer to the standards-based version of of the notion of transition the transition process. I'm not just talking about like the old monocles of like a staging environment or no staging environment or one out but like the fact is some pair that like, what are you good?
How are you gonna ask an insurance company to run their big fat load test with all their data's in a production environment or even just e-commerce, right? Like now all the time you got to go over to your sales office person and say oh these records are completely bunk and these ones are real customers. No, no.
No, that's what happens when you just pour a paradigm and assume that it's gonna work perfectly in another place what we see and actually we were down on Alan you and I I think so each other in the hallway and hotel day in Austin, right the open source Summit in Austin and what oh gosh a million years ago the April May. I don't know June whatever I think. Another organs, but it's okay with something like that.
Right? Like it was a million miles away. But like, um at that conference one of the things that they were really very interested in in terms of open Telemetry was also to be able to you know, for some companies like Trace test Ken over Trace test.
And then also who was it it was Michael Haberman and and Iran Governor that run a specto right? These are these are premises that say hold on. This is not testing in production.
It's just testing production production should emit useful information to make business decisions that by the way should also be part of your pre-prod pre-roll release process. Your system should be producing information enough where you can easily just go based on the exhaust data of our systems. Is it working enough well enough or not.
So this notion of like testing in production almost puts it like as if you take testing from the old And just poured it to production as if you can do the same things the different things is actually saying hey look system has changed highly distributed lots of information coming out of them. Why wouldn't that be the evidence-based? proof that things transitioned properly and I think that's a really cool thing on the other end.
There's plenty of people who still need to prove confidence before they release. Because they might not have these perfect sort of canary like, you know, 1% or Progressive rollout processes don't have the luxury of that have to go. I don't know.
You know one one testing is a big thing in Insurance. Same thing open enrollment just happened to try sinus. The people that support us right are health care provider had to be ready months in advance before rolling it out to even a small portion of their for requirements and for compliance.
So whatever parts of the spectrum you're on. Yeah, you do your stuff beforehand but like afterwards that shift right is the question of what feedback from your systems is absolutely necessary no matter which environment you're in. Fair Nora you are someone with a ton of testing and production You know Paul gave us a lot of kind of what what we're seeing what what's new in there, but I'm gonna ask you to you know pin the tail on the donkey is shift right testing.
Just testing in production or is it testing production as Paul's alluded to what do you think it is? Um, I mean, I I, you know, I can't stop thinking about a t-shirt that charity majors made which says I test in production and so do you it doesn't it doesn't feel that controversial to me because the thing is we're all doing it and it's about like your mindset too towards it like are you accepting and embracing that you are already doing testing and production or are you? Scared of it and you know making things in place trying to make your staging environment perfect or trying to really emulate the load of some things.
I mean like what Paul was saying, some folks just don't have the luxury of doing some of that and so I think you need to work on I think orgs just in general need to work on making production feel comfortable to test in and a lot of different ways and there's a lot of ways to do that with tooling with organizational mindsets with a bunch of different things. So, That that's generally like where my my head is at for for this talk. Actually really like how you put comfortable.
I do think of a lot of it is mindset and I think everybody's kind of starting at a different spot. I actually think it's kind of unfortunate we say left and right because it sounds very waterfall to me and I don't know seven years ago like the term continuous testing almost stuck for bit and then it's kind of come back but for my perspective is you should think about quality and security no matter where the code is. It doesn't matter at the point of feature definition.
So it like Productions is just one of those places and and creating these artificial Gates. Mentally, I think in the teams is is a big part of the problem because sometimes it seems overwhelming. I know when I talk to companies and say well, you know, I can't just like Port it over.
They haven't even figured out feature Flags. Like they haven't even figured out good functional testing. They haven't even figured out good unit testing like there's some Basics that have to happen first.
So we shouldn't just assume like everybody has those Basics and they can just do it across all environments, but we should assume that everybody is thinking about security and quality throughout the entire pipeline. Sure. It's pronounce it for me one more time Sherman.
look should we have a stigma around? You know so called shift right testing or testing in production. Is there a stigma that's grown up because I'll be really honest with you.
It wasn't unusual Mitchell and I built products 25 years ago, and we never finished our testing in time. Right. We never did or on budget either.
I always blame him but You know, we used to continue testing in production even after we released. Did we did we forget that or have we buried it up in the attic with Grandma's clothes or something? Right?
When did when did that become such an evil thing? I don't think there should be a stigma like like Nora says we've been doing this we all do it and I think maybe some of the pushback came from more riskiverse organizations like the banking the financial services people feel that they just can't afford to experiment with their life systems because it had implications right Beyond systems that they could control because he's a very interconnected systems the regulatory challenges as well. But even then I know that many of these Financial Services customers or Enterprises do ship tribe testing, right?
I mean, we actually run a survey with, you know cab Gemini a couple of years ago and that really showed about 45% Enterprises are actually doing shift right testing. But but also more importantly I think as Paul mentioned the key thing for us is being able to gather insights from production data and and feeding it back to the previous parts of the life cycle that essentially is what continuous testing is because the moment you say you're continuous test. Get beans that you're testing across the entire life cycle, we reproduction and post-production.
Right? So I think that's been a big Focus for us as well as our customers. For example, we've been doing you know, for example crowd testing forever.
That's a key part of for example, you know CX focused testic customer experience focused testing, right? I mean, I believe the unity can't do a whole bunch of CX you can shift left but I think most of your best CX testing is is done in production and I'm going to remember see text is very very different than the traditional testing we do right? I mean if you think about CX the feedback is in town in Shades of Gray, right?
You don't have a password of failness. Certainly. Oh, yeah.
I kind of like that. I don't necessarily like this, you know, so it's very very different great. So the whole approach to some of the concepts around customer experience Focus testing is different than the way we look at traditional testing right and also completely agree with with Chris for example You know, when you when you have things in production your users are giving you read real feedback in terms of you know, what the like, but don't like what they see with their competitors.
For example, that's a great way for you to do requirements. Very validation, right? You know, what should be part of my requirements Suite is it only the the product owner who gives me requirements or is there insights that I can gather from production in terms of what are my users really want?
What are they adopting? What are they not adopting? Right?
What does it like? What do they see with my competitors that they would like to see in my product all of those insights. I really part of the very shift left, right?
I mean requirements definition. So so I see this as a continuous process, I think one of the key things that we're beginning to see Alan more and more is around the the generation of insights from production data, right? Not really we need that for example for for disciples like this data management, right?
You need to have realistic data from production, right? But if you think about more and more of our AI based applications, they need to be continuous. Refreshed with data so that the model can be able to recalibrate it right the machine learning models and that requires you to have those insights from the data from raw data from production as well.
Also, for example, we're building our solution around how do we you know predict reliability, right? So if you look at the latest dollar report that came out just last week, I believe right you will see that the change failure rate has actually gone up from 2020 21, you know before that, you know, even low organizations low material organizations had change failure rates of around you know, maybe I think the range was 15 to 30% now, it has gone up from 30 to 60 30 to 60% and 80% of all of the Enterprises that they surveyed have changed figure rates greater than 16% which is very high, which means that while our Enterprises are going for, you know, High Velocity deploying more often, but at the same time they have a significant number of deployments going to production that fail Right, so would it not be better if the actually gathered insights from production and and fed it insights back into pre-production environment so we can start to get a very few for what kinds of deployments fail right Can We Gather insights? Can we embed the the data from production along with the pre-production data and and you know and and get to a better decision support system for our guys so that we can actually start to actively predict and manage your change for your rate.
For example. Absolutely. Do I only thinking about the Miss?
We sort of the title of our talk there's some fundamental things that have evolved and changed over, you know, since you and I were doing product products together in the early 2000s, you know one is we have so much more automation to have a lot of data available to us that has created because of that automation. We are able to keep it now both prior to it going into production or in test or you know in the developments cycle. And the other is just faster delivery of software.
Now. Most of us haven't aren't at the stage like a Netflix to be able to deliver multiple deploys a day, but our ability to experiment in Market or if we do have a problem in software that we need to resolve we can fairly quickly. We're not waiting for the next quarterly or six month release right?
We're living not everybody is there I understand that but we are in a bit of a different world than those are some of the kind of fundamentals that have changed to allow us to look at testing in a whole life cycle of software. Not just before it hits production. Fair anyone else on panel thoughts Let me let me.
okay, I just quickly wanted to say I am curious, you know, like they they don't seem at odds with each other to me. Like they it seems like a loop in a way like and so when you are it's not an either or sure, how do you take that? How are you taking that and baking that back in over here and consistently doing it without having to be a problem.
I mean I My very first internship ever I was a QA engineer and the way that the process worked is at a hard work company every time the hardware didn't work as expected a new test case would be added to this spreadsheet in the spreadsheet like couldn't even open on your machine at one point. It was just so big and that's that's not the way to do it. Right?
It's like, how did how did this thing even happen? Why was it possible for this to happen? What trade-offs were people think about?
How are we actually vocalizing all those trade-offs people are thinking about over time and putting that in our shift left rather than just adding an additional run book or adding an additional test case and I know the example I brought up feels quite Antiquated but it's also still happening all the time. Like when I see a post-mortem where the action items are update the Run book. I want to pull my hair out like because you can say that about any incident you can say that about any bad thing that happened it's more How are we exposing the conversation around how this thing happened and baking that into just our everyday assumptions our everyday collaborations with each other?
Yeah. The Ellen the Last time was on I think was with Tammy from that was working at Gremlin and you know, we were talking in and around this problem of like, you know, what happens when there's a problem. Well, you know whether he's Gremlin, we usually in many of our large Enterprise customers, they'll use both combination of like knee load and And Gremlin in a pipeline to kind of parallel like inject fault to prove that like the resilience.
They know has to be there is there but you're never going to prove all the things are perfect. And that's not the point. Yeah on the other hand.
You do have to start in a place where you are not you are not literally like shooting yourself in the foot. you know, so what you know, though, I hope and I love the idea that people should be treating their production systems as not just an organism which it is but also the living and body of their organizations, right like the like the that that fundamentally I think Nora you brought up the notion of the lessons we learned and one of the things that I think we we talked about Tammy and I did was this notion of a like almost like a right once and read never culture. Where it's like, oh, yeah, we got to do a retro and then either it's so bias for Action that everything has to turn into a suggested action that turns into a task or people ignore it and it becomes a problem over and over again.
So, you know, one of the things I I wonder about is how do we turn what normally is sort of like Chris you mentioned the the shift left and right still feeling waterfall. Don't worry. I have trademarked the term may see if make sure I say this, right A bull shift right?
It's about it's about she feedback being right fit at the right times at the right places. And also making sure that instead of like a flat paper. It's a cylinder it always comes back around right your posterior knowledge a poor Syrian or a posterior knowledge.
Can become a pre-orienting knowledge but only for people who are paying attention to it and prioritizing what knowledge has to be in place to make a good enough informed decision. And that's a tricky place to be so like, you know, we're doing some things and tricentus drive figure out, you know change impact analysis. We've got that live compared tool that does it for sap.
What's quite frankly is a very baked concrete sort of, you know transaction codes are over here like they are over there and it's a very well known system very well known domain model even with all their changes over the past couple years, but what does that mean for things like custom maps? How do you how do you map out the domain of things that like the change stream the domain model and the usage data across things that like I can bring fart out and know JS plus like angular front end app and have it hit millions of users in a matter of months. How do I Define those things for any given application that an Enterprise might commit, you know three to six months on?
So the the question is how do we actually prioritize that knowledge that we have at the right times at the right places and I think Nora the jelly IO what little I know about jelly IO but like what I see in what I've heard and watch some videos on it seems like you're taking a particular pain point and turning it into what do we learn? This is an opportunity for Learning and I I love that right? Like that's at the heart of I love that so my question back to the group and unless Alan you want to take an at a different direction is like what do we do in these moments to figure out how to take like a moment of problem.
Let's say we test in production and there is oops. We caused ourself an incident. Don't worry, you you also already probably have some plenty of incidents to deal with but regardless of where the incident comes from what what are some dynamics that go into making sure that we've we've prioritized the outcomes that we've actually thought about should we take action on this and how do we do that?
I don't know where else to start with this group. Can I quickly just jump in real quick? Because something you said Paul was just absolutely gold.
It's you know, we're talking a lot about learning but I think people in organizations learn in very different ways. You said something about right once read never right? Maybe your organization.
Just doesn't like documentation and you kind of have to lean into that. Right? Like how does our org learn what do people receive well in our work, how does change get made because that's how you have to tailor your learning like that's how you have to tailor disseminating your Lessons Learned into disseminating those philosophies like you could write all day about the change log of a thing and how it you know gets looked at but if no one is clicking on that and internalizing and making time to internalize that does it matter.
I when I joined the chaos team at Netflix I was so excited to join and we were working on this really great tool that they had already written some white papers about like wow, we can inject failure and production without a customer noticing like that. So cool. And yeah, I looked at the application and the only people using our tooling were the three members of my team.
And so I'm like we're learning all this stuff about our system, but it's staying with us. People that are not on call for the system and so it it opened up a big conversation around how do we get other people enrolled in this how do we do testing and Lessons Learned and experimentation in a way that gets people to learn about what's happening and also share those stories with each other. I didn't want to step on your question for the group.
I just yeah great. I think you still address it. I would say that I was thinking all as you were talking like there's other symptoms of the problem in organizations for example like documentation certainly is that Right right once read never but also people who think in terms of like change management and and change controls and and all of that stuff have a tendency not to think about things in a very fluid way.
But how do we get there like? And it feels weird because devops is actually been around for a while now. It's been around for a bit.
But I still every single conversation. I'm talking about silos and culture and like I think as as thought leaders in the industry were still not serving the folks who haven't gotten some of the basics down. Like to me one of the things I'm still super shocked about is pipeline analytics is not in every single practice out there.
And like how can you have how can you think in terms of resilience? If you don't even have the data to measure like how you're doing? Like if you don't even have that information like you're not even at the point where you can start considering doing more so I would encourage anybody listening is like if you don't have the basics and you don't have the culture make sure you get that right like you're North Star can be testing across everything but like don't skip steps.
Although my my little boy is gonna walk before he crawls which that breaks every Paradigm ever I've ever known. The other thing I'll say is that what's interesting is, you know, in in the case of blue green deploys or Canary deployments You could argue that failure is a success. So a field deploy is exactly what Needed to make a decision.
to improve and so I it really to me is like people need to get uncomfortable knowing that things will break no matter what no matter how hard they try that your users are your testers that creating parity in any sort of pre-production environment is impossible. And so you need to really just think about quality. Everywhere and not like in afterthought or not like something that you just have to do kind of like doing your homework or doing a chore.
Yep, so from a cultural perspective because I think one of the key things we're seeing is more and more Enterprises adopt site reliability engineering and this new role of an SRE. So, you know many times testers very feel very uncomfortable with production. Did I mean, in fact, I don't see many cases where testers actually have access to production it at all.
Right, you know despite the growth of AI obstet colleges where they can get snapshots and insights, but I think the SRE is one of those key rules. We look at that Bridges really the gap. From a from our organization silos and roles and personas perspective because they need to be working with with all of these different groups developers and specifically testers and really bring them out of their of the shelled.
So to speak because many times you feel testers getting us somewhere isolated, right? So because they can help share in a more digestible form, you know, smoothly insights from production what they're doing with their budget. For example, how are the calculating their budgets right at the state should definitely have an input into the whole process because you know who better than it has to give insights in terms of what are the particular risks or untested pieces of the applications, right?
So I think there needs to be some cultural bridges that need to be built with help of the SRE roles and some of the other Transformations that are going on to help developers testers and other other roles and product owners and product managers to play a more continuous, you know role in the whole life cycle as opposed to say. Hey, I only in production the operations guys deal with production and sort of break down. Yeah, I I'm glad you said SRE.
I think that was the first time we said it in the most amazing practice I've ever been a part of in terms of like a high performing engineering team leaned heavily on their sres and not like I'm not meaning like overworking them, you know, they weren't the fixers. They were the stewards. They were the stewards of quality.
The only thing that they owned fully was the tools that were available and the what we call production Readiness checklist, which included Testing and so it was more about guardrails and stewardship in communication then going in and being the ones who run a script to restart a server. And so yeah, I think that functions of that sort are very important overall. I think you both hit the nail on the heads.
Alright? Much of this is just about enrolling people that have different sharpens and different areas of expertise into the conversation where things are created and finding an opportunity to collaborate together. Like you said, I think Chris you were talking about some folks don't have some of the fundamentals down but I guarantee you those folks are also still having incidents and I'm curious how they are bringing in all the necessary people that were impacted by that like testers were impacted by that Engineers.
We're impacted by that. I think a big group that we don't talk a lot about is marketing. Right like we we talk about load testing and testing and production and we talk about events like you're open and enrollment, uh story Paul, but we don't always talk to the end of the business that is really understanding those deadlines and the impact on folks.
I mean, I don't this is a kind of silly example, but I you know, I was really excited to watch the new Hocus Pocus with my nieces and nephew. And I noticed Disney plus was pretty down for a couple days on the on the date. It was released.
Right and so I'm you know, I immediately and I worked at a streaming company that I was immediately curious about how some of those deadlines and expectations around user management were communicated internally and how and if marketing and content development was working with sres and testers around this or if sres and testers were sitting in silos and just making assumptions, you know, like hey, we have all this data we can look at we can predict these numbers and it's like yeah, but you should also Hear the stories from folks on these impacts too. Well, there's there's two there's two general phrases that come to mind one is you're only as strong as your weakest link and at least from my years of doing performance engineering and then moving out to reliability and scalability and then figuring out, you know, how to help a community of people that are regulating burned out just by dealing with the crap that their organizations throw into their production systems, which is in part a lot of devops work and a lot of SRV work and we need to watch out for that. I say for my local and Global community of essays and devops folks, but the the weakest link thing is like There are things that you can catch early on use the dinkiest tool and catch the dinkiest problems and then use something else and there are different types of testings.
And when we talk about testing, I'm almost like so done with testing working for a company that has a lot of great testing products. The fact is it's like it needs to be elevated Beyond testing and these to focus on quality engineering and what information people need at the table at the right times. The number of times that test Engineers don't have access to APM tools, even for pre-production environments, no architectural diagrams, which by the way always change and you might as well go off to that apian tool and look at the actual service dependency tree like the information that we need is beyond just right a test and run a test and past fail bull crap.
So like the strongest the weakest link, you should catch a lot of your weak links before you actually bother with throwing actual Revenue generating or Even if even if you're an NPO large npos some of the ones that we support they end up having large like people's lives are affected by when their services are down, you know critical systems and stuff. So the weakest link thing. There's a lot of weak links that you can get out of the way.
And then yes, we should focus on things that just are so cost prohibitive to set up and staging environments and stuff like that. Absolutely that stuff's gonna happen. We need to be fast and learn about it and have the second part the second phrase that I would put is if a tree falls in the woods Right, and then the follow-up is and nobody's there to hear.
It doesn't make a sense. In reality, if we don't have visibility and I will use the overloaded term now of observability to describe this, but if we don't have a comprehension of what are our most important things that we need to pay attention to in production then back Port those that's where I think shift right actually plays a role and I think she mean you you kind of point to this is taking production information and informing not only the testing approaches but the dev approaches and the product and the like who's using it comes up all the time on calls for me. So if if nobody's listening because there's not information.
That's a big problem. So get the information in place and you know, you know, the game of measurement is like it's easy to measure the wrong things but iterate on that stuff and make sure that there is a parody specially between Tech and Biz about what you're measuring and production and why wouldn't you be measuring that? When you are trying to figure out, you know, your next major release has some significant changes that could very easily combine together to cause some big issues new CPU types plus a migration to Kafka plus something else, right?
So I would say that the two things are first off if if you're if you're not doing the obvious things Come on, there is a maturity factor and Chris. You mentioned new people in the like people who like what why do we have to keep bringing some things up? Because there's always new people.
There's always people like it's something's always new to somebody and so as new crops of people who are new to you know, not just software engineering but devops and SRE right come into play. They don't know all the pains on the scars from people who've gone before. So I think the Norris point we kind of got to figure out how to create conversations that help people gain context, you know personal way as opposed to a right once read never culture.
So that was a diet. So it's a Bruce, you know widely agree with you, but I think I'm so glad you mentioned the word quality engineering, right? Because at the end of the day, it's about delivering quality and qualities is as perceived by the users of your system.
Right and many times. We find that we do testing for the sake of testing, you know, we talk about test driven development and everything, but the end of the day testing is an activity quality is an outcome, right? So quality should be driving the testing in other words what your customers or users perceive of your software right should really be driving the kind of testing you already got activity, right?
It could be developed or testing whatever right? It's not just about actually try to achieve test coverage or anything like that for example, and if you're a startup good enough quality rate, so I think everything should be modulated based on the goal and production in that sense. Really, it's quality driven testing, right?
That's a key driver in my mind for this whole notion of shift. Right? Right because that's really what our customers care about.
Right? So everything in testing should be driven by the end goal of quality. Here, you know what though?
I listening to you folks. And you know, I don't pretend you guys know a lot more about that thing than I do. One thing that kind of resonates and comes home for me is like so much so many other things in devops.
See where culture. Play such a role here. Right.
I think sometimes we lose sight of the fact that failing a test isn't necessarily a terrible thing. I'd rather find out about that failure. on my testing then when Nora's niece is crying Bloody Mary because she can't watch Hocus Pocus to right or when NFL Sunday Ticket goes out on me and I'm screaming sitting there at my football jersey, you know cursing the the NFL Gods.
But yeah, and that's the purpose of of this testing a lot of times and so culturally we have to understand, you know, don't sweep stuff under the ruglets. Yes through the low-hanging fruit as Paul mentioned, but you know, there's nothing a man with with doing as much testing as we can. you know, it's stitch in time saves nine or whatever it says and so we need to culturally make sure our culture is such I mean and it's come a long way think about where we are today with feature flags and a/b testing and all these things, right?
That's all at some level production testing, right? I mean Mitchell you and I you know, we had Paul Pinckney our UI developer. He used to do all the UI development and kind of interviews with people before we put the software out then first find out when people were kind of bitching and moaning about, you know mistakes we made in that So I think that's an important piece of it and to the point where look.
Change is constant in our business. There's always new people. new ideas people need to learn you've got to kind of build that into your culture as well.
Anyway, we are approaching the top of the hour here before Mitch long. I'm going to give you the last forever. I'm going to give you a minute to think about it.
I want to thank shameen Paul Nora Chris. What a great conversation. We're going to rerun this one again.
And you know, I'll leave it to the The Producers here to fix that for us, but I'd love to continue this conversation to all of those at home. Thank you for joining in many. Thanks, which I sent us for sponsoring Mitchell.
Why don't you take it home for us? Thanks, Alan, you mentioned culture and failure. And it reminds me of we've all heard the sailing the same failure is not an option going back to Apollo 13.
Actually, that's not true. That's the opposite is true what people dying was not an option not failure. We had to fail to figure out how to get them home and we software is always going to fail.
It's just it will you know what we live it every day, but Nora mentioned and talked a little bit about learning. Every one of those failures is learning whether it happened the moment you checked in code and got our test run to we did some engineering and testing production Canary testing, whatever it might be in production. I think the main message here though is Not only is it a cycle.
We need to think about softer holistically entire all parts of how we where we develop and how we created how we test how it lives how grows in matures and some day other software replaces that We're gonna have issues all during that process. And how do we find those things the best we can and when we don't find them how to react install them quickly because they're gonna happen. So my party thoughts absolutely and that brings us right to the top of the hour folks.
Thank you. Again. Thank you everyone for watching check us out on our next devops on down and if you get a chance to join us at a live devops on Ground Round Table, there are a lot of fun.
And so then this is Alan Shimmel for Tech strong. Be well. Take care.
Bye. Thank you, bye-bye. Thanks everybody.


