Journey From Software Engineer to SRE Through Platform Engineer at SKILup Days 2024
The terms SRE and platform engineer may be used interchangeably, but in reality, the responsibilities may be quite varied across different companies.
In this session, Yevhen Zavhorodnii, principal SRE at Ticketmaster, shares his career journey, from his time as a platform engineer at Cookpad to his time as a staff software engineer at Ivanti where he helped build a culture of DevOps.
This session will explore the formal and informal responsibilities, expectations and realities of these different roles and how each contributes to the larger software engineering landscape.
Transcript
So, yeah. Hi all. Uh, I'm Ganza.
And, uh, today I would like to, uh, walk you through my journey, uh, being staff, software engineer, then platform engineer, and now principal. So, uh, first of all, uh, just, uh, just to give you an idea, uh, on, on, on what about, uh, would be the stock. Uh, so I would like, uh, to introduce myself, obviously, uh, give you some theory, and then, uh, show how theory match the reality.
And obviously then they will make, uh, some, uh, conclusions, uh, conclusions out out of it. Uh, so, uh, I obviously, like I was a software engineer then, then I, I became, uh, S3 or platform engineer, and now I'm, I'm officially like principal, SRE. Uh, so as, as you can see, I, I have quite a lot experience in very different areas, which is kind of making me, uh, trade of all Jack, uh, master of nasa.
Uh, however, like obviously, uh, those white experience, uh, kind of given me advance, uh, in, uh, being a three. So, uh, like, uh, as, as, as you can, uh, imagine from my, uh, surname, I'm my, I, I came, uh, here from, no, I'm from, no, not from Britain. I'm from Ukraine.
And before UK I was, uh, like participated in, in, in a lot heavy, uh, software, uh, algorithmic, uh, activities, uh, where, when eventually I became a software engineer. Uh, then as a, and then I started my immigration story being contractor in big, uh, companies like, uh, Microsoft, uh, for example. Uh, so like, and like, as I said, was more like a backend or web development focused.
Uh, then, uh, like I followed, uh, the general transformation of the, of the world, uh, around, uh, um, how, how industry is moving. And, uh, during my period, uh, in Ivanti, uh, back in 2016, to be honest, like more, more like 2019, uh, because like I started a software engineer, atti, but, uh, Ivanti, uh, started this firstly DevOps transformation. Then, uh, three transformation.
Uh, and, uh, that's where my journey is through DevOps three and, uh, uh, platform engineering actually started. Uh, so, uh, I would say that's starting from, uh, 2019. I'm just like actively and continuously learning new area for me, uh, which is, which seems for me like, uh, uh, and this, uh, so, uh, now, uh, let's, let's go quickly recap, uh, through the theory.
Uh, like, uh, if you, if you, uh, like aware with terms, uh, you'll, uh, you, you, it's done. Nothing new for you. Uh, however, like, let's look up, uh, so as DevOps, uh, again, like many, many different, uh, uh, uh, explanation of what DevOps is, uh, however, like, uh, the, like, uh, uh, primary characteristic of DevOps actually, like DevOps is focused on delivery.
Uh, uh, it's, and it's different from admin, uh, admin facility responsibilities. Uh, uh, so, and it's actually the time when, uh, historically, as far as I remember, like time when, uh, more like admin people or software admin, admin, uh, managed, uh, moved, uh, uh, towards, uh, the DevOps. And it was kind of, uh, uh, time when operation and development, uh, started, uh, collaborating, uh, with each other.
Uh, then kind of another term, uh, another practice, uh, uh, appeared, uh, mostly out of DevOps, so mostly out of, uh, Google, uh, kind of, uh, adaptation of DevOps, uh, uh, practices. Uh, that's why here as I, I cited, uh, uh, that's like from, uh, the site from, uh, Google three, uh, book, uh, and effectively S3 e comparing to DevOps is more focused on reliability and scalability of the product, uh, which is like different from operation, uh, perspective, and, uh, out also more, not out of S3, like in parallel S3, but again, like more, less out of, uh, based on three, uh, adoption, uh, emerged. Uh, another, uh, uh, and, and another discipline like platform engineering.
Uh, so, uh, I started my journey as platform engineering based on AWS uh, uh, kind of, uh, vision of being platform engineering, uh, as like some, somebody who is like providing capability. Uh, however, like, uh, what I found on Splunk block, uh, kind of reflect in my understanding of, uh, platform engineering. So, uh, as I, as, as, as, as it's happening always in, in, in, in, in involved, uh, like SIR is great.
However, like, uh, uh, reality is a bit different. Uh, so, uh, when, if you, if I, uh, retrospectively take a look at my, uh, Jo job, uh, during, uh, being staff software engineer 20. Uh, so, uh, at some moment we started, uh, um, being doing DevOps transformation.
However, like our, our ops team was, uh, kind of under, under hired, uh, we, we, we had like, just few DevOps for, for, uh, a relatively big organization. Uh, and those DevOps engineer was mostly, uh, came from ops background, uh, and ops background, who willing to, to do some coding. Uh, so they, they were doing mostly like strategical strategical review, uh, and some advisory.
Uh, and, uh, all of burden of doing DevOp transformation was, uh, on the, on teams itself. And we kind of, uh, made some sort of, uh, open source approach, uh, to build some shared components, uh, shared reform modules, shared, uh, uh, uh, piece of infrastructures, uh, and, uh, like why, why is it, uh, in this speech, in this, uh, uh, in this like topic? And because it's quite related to the topic of platform engineering, uh, because like soon after I left the company, they effectively maintainers of this, uh, uh, like open source, uh, local tools, uh, uh, they moved all of this maintainers to, to a new team, and they called this team platform.
And, uh, based from what I, what I'm hearing, uh, from my ex-colleagues, uh, it's effectively platform team. So like, it's, it's a team which is providing capability for other, uh, for, uh, for their colleagues, uh, for, for their teammates, uh, to actually being able to, uh, deploy software, uh, and, uh, address most of, uh, uh, concern of, um, like maintaining and monitoring and, uh, actually, uh, deploying software. Uh, so my next destination was, uh, SLV, like formerly known as Schlumberger.
Uh, so mainly, uh, so it's, it's, it shouldn't be, uh, here for, for, for this pitch because it, it has like, nothing, nothing to do with platform engineering, cos three, uh, like we haven't at all at DevOps, OS three, uh, in, in our companies, at least, at, at least in the division where, where, where I worked. Uh, but my, my major, major, uh, major kind of, uh, driver for my, for my current career towards, uh, platform or, uh, so I said reliability engineering was, uh, switching from the net to go. And, uh, as, as, as you may aware, go, go is actually kind of native language for NTUs.
Uh, and it's kind of a good entrance, uh, for, uh, for, for, for do, for doing, uh, some, uh, deployment and patch in, uh, Kubernetes, uh, things which is more closer to some, more closer to being more closer to operations. And anyway, as, as company didn't have, as division didn't, didn't have, uh, DevOps or S3, uh, like dedicated, uh, the company, uh, like it's, it was responsibility for any every team members to actually, uh, doing some ops concerns, like implement, uh, DB migration or, uh, some, uh, more related to three, uh, concern, like, uh, like improvements, uh, for reliability or monitoring. Uh, so like, we didn't have some monitoring teams, so it was again, like kind of, uh, uh, kind of responsibility of, uh, each team members.
And obviously, like, uh, the, the fair, the fair, uh, the fair approach was to do this on, on, on rotation to actually do not have, uh, just one dedicated monitoring person. Uh, so, uh, again, it's kind of approach, uh, when, when you don't have a DevOp sources, uh, sorry. Uh, and my next top was finally, uh, kind of, uh, officially matching, uh, the, the subject of this, uh, uh, topic and, uh, conference, it actually like, uh, being, uh, senior platform engineer.
So as far, uh, at, at Corp Patch. So as far as I know historically, uh, the team which I joined, uh, so it was, uh, even named S3 E, however, and it was built, uh, based on more like, uh, DevOps people. Uh, however, however, like at the moment when I joined the team, it was already pulled, uh, platform.
And, uh, it was built, uh, out of, uh, team plus newcomers. Uh, and, uh, those newcomers, uh, was, uh, mainly from, uh, software, software develop from software development, uh, background. Uh, and because it, it was built out of a three, uh, it was, uh, it was, uh, actually on call, uh, rotation, uh, for like, uh, for, for production code, uh, uh, troubleshooting.
Uh, and in, in, in the meantime, we also, uh, we are building some tools for debugging, troubleshooting. And, uh, we heavily patched, uh, exactly 1, 1, 1 of one, one of my, my, my, my task force was about, uh, uh, patching the Kubernetes, uh, with, uh, uh, operating operator separators. Uh, and still it was some sort of mix of platform responsibility and ops, uh, responsibilities.
Uh, unfortunately as a, as a company like in uk, well, wasn't, wasn't doing well. And, uh, in, uh, 2023, during all of this, uh, crazy layoffs in the, in the industry, I was kind of victim of this, uh, layoff. And, uh, uh, I had to switch, uh, uh, my job to, to, to, to next, uh, company, uh, which was, uh, Ticketmaster.
It's actually where, where I'm working right now. And, uh, now my title is like Principal site Ability Engineer. Uh, like this title, more of like, uh, it's, it's, it's more, uh, it's more similar to what, uh, kind of, uh, Google, for example, uh, society engineers are doing, uh, like, uh, introducing SLOs kind of, uh, uh, booking teams about like, uh, actually implementing SLOs and SLOs.
Uh, however, uh, it's still, uh, uh, it's still, it's, it's, it's still, uh, required, uh, to do a lot of, uh, what, what I would call, uh, like pla platform engineers like, uh, uh, my, my, my team and my myself. Uh, we are, we are producing some tools for other, uh, for other developers, for other teams to actually being able to properly monitor, monitor their, uh, application. Uh, we, we are probably, we are, we are kind of working together with security because like, obviously, like, obviously security and reliability is always kind of, uh, going together and, uh, have, have many share concerns, or at least like, uh, like usually, usually security and is kind of good, rely with each other.
Uh, and, uh, again, like, as, as it's, it's not like platform, uh, team, uh, responsibility, but it's not usual platform, platform team, but it's really, but, uh, doing incident management is, uh, purely SE responsibility, usually in, in most of the companies. Uh, again, uh, like, uh, why I went through the, through, through all of my, uh, titles as, as, as, as, as you've seen, I, uh, I was on, on all of, uh, uh, kind of areas, uh, in, in this, in, in, in, in, in the, at least in, in, in the, uh, uh, development, uh, and doing reliability and platform engineering. Uh, so, uh, if you like, take a look, uh, at my, uh, skillset, uh, like, it's obviously not T shape.
I, I obviously, I obviously came from, uh, so from, so, so software development, uh, background, and, uh, honestly say in my kind of operations and automations, uh, which is like, uh, uh, which is very ops skillset, uh, uh, are not very good comparing to my, uh, coding core analysis, uh, skillset, uh, and historically from what, from what I saw usually, uh, S3 E, it was kind of o ops, uh, people who learned how to, to code and manage, uh, how to code. Uh, but like from what I see current industry kind of, uh, doing the shift, uh, to, to, to, uh, hire more and more, uh, people, uh, from, uh, from software engineer background, uh, to, to do some, uh, ops related job or SE related job, uh, mainly because, uh, it's just my, my honest opinion, uh, mainly because complexity of those, uh, of those, uh, those systems, uh, growing up. And so you eventually, you need to patch it with like, proper languages, like, uh, patch with Go, uh, like patching Kubernetes with Go, for example, uh, like I was recently, uh, and, and, and, and, and those, and those, uh, kind of misunderstanding and, uh, uh, not understanding by, by the industry itself, uh, uh, the, the, as a community kind of reflected even like in many, uh, conferences, which I attended quite recently, like on S3, uh, conference, uh, in, uh, San Francisco, like last March, uh, likes a huge talk or like, for, for about like 45 minutes, uh, the speaker was talking about, uh, uh, how differently, uh, kind of similar, uh, set of responsibilities, uh, called and, uh, how, how, how, uh, how, how people differently understand the, the, the same, uh, the, the same skill set and how different companies naming the same skill set and with different, uh, uh, job labels.
Uh, uh, and, uh, like, uh, even WTF is three is kind of like, you, you, you're kind of, uh, getting this, uh, from, from, from the title itself. Uh, so it seems like, uh, all of those disciplines, uh, it's not, it's not like, uh, like a, like a job title. It's more like, uh, set of practices, set of, uh, set of, uh, best practices on how, how you manage, uh, and, uh, how, how you are, uh, deploying, how you are, um, monitoring how you are, uh, maintaining your application, uh, uh, and how you are operating your application, especially in cloud.
Uh, that's, I I believe it's my, it's my, my honest opinions. It's, it's, uh, uh, why, why, why, why, uh, industries still ha hasn't come up with, uh, one strong title for one, one exact title, which is covering all, all of those aspects. And, uh, like, uh, even like if you, if you take a look for, uh, like I, I'm, I'm very big fan of, uh, uh, of, uh, roadmap, uh, project, uh, which is kind of giving you quick starter with any new, uh, subject, with any new topic.
Uh, it, it's kind of giving you quick starter of, uh, uh, what you need to learn, uh, to become, uh, some sort of, not expert, but at least be aware of the, of the technology. Uh, like, uh, again, as I said previously, I'm constantly learning, uh, new, new, new topics. Uh, like if you generate, uh, the roadmap, roadmap for site reliability engineer, uh, you will see, uh, like it's, it's, it's, it's not a full map.
It's, it's like, like 60% of the map because it's quite, quite huge and, uh, uh, touch many, many different topics. Uh, uh, and if you generate platform engineering, uh, like, uh, you'll see many, many common things, uh, like, uh, obviously networking. Networking is, uh, one of, one of, uh, the shared concern across, uh, two disciplines.
Uh, monitoring and logging is shared concern, uh, between, uh, like platform engineering and SRE, uh, like performance tuning. Uh, so if you take a look again back into side side availability, performance tuning, tuning as well, uh, shared concern, incident management is not concern of platform engineering. Uh, however, like, uh, like in the, in companies where platform team has kind of, uh, evolved as, as the independent team, uh, in, as they involved into, into incident management, because, uh, with like, without proper tool, without proper tools, uh, we won't be able to properly address, uh, uh, incidents.
So, uh, oh, uh, so, and, uh, like CI/CD also, it's, it's kind of quick comparing of those, those, those slides. Uh, ci CI/CD is also, uh, concern of both, uh, areas. And, uh, from, from what, what, from what I, uh, what what I was doing during, uh, my job as platform engineer, uh, like I spent a couple of months, uh, providing, uh, like providing capability via providing, uh, GitHub actions, uh, templates for other teams to, to, to, uh, help them to easily deploy, uh, things or, uh, I know that, uh, uh, for some company, uh, I, I, I, it's my, my, my friend is working for this company.
Uh, they, they kind of provided to Kubernetes operator, which is easily, uh, deployed technical, uh, services. Uh, so yeah, uh, from what, from what, uh, we can, uh, kind of make a final, uh, conclusion around, uh, those, uh, three areas like software engineering, site ability, uh, engineering, uh, and platform engineering. And like, obviously this is the first area which I forgot, uh, the box, uh, uh, those all is about quite similar concerns, uh, about similar concerns and, uh, many companies, uh, kind of, uh, having those, uh, uh, concerns addressed, not by one title, not by one department.
It's usually collaboration across different departments or even like, uh, one, one person can be assigned for different roles across, uh, this, those four pillars. Uh, so yeah, thank you for your attention and, uh, if you have any questions, please go ahead and ask.