Techstrong TV February 10, 2026
Watch our live stream Monday through Friday, featuring exclusive news, announcements and conversations with IT leaders and experts on topics ranging from digital transformation to #DevOps, #Cybersecurity, #CloudNative, #Containers and deep-dives into specific technologies and best practices. http://techstrong.tv/
Transcript
Hey, everyone. Welcome back here to Textron tv. My next guest, he, he's been coming on, I think since they launched this, his company, uper.
Uh, it's my friend Kumar Chila. Uh, Kumar is the co-founder and CEO at UPSer Kumar. I still remember the first time it was you and your co-founder.
I think both of you are on at the same time. Yes. That's okay.
And, um, of course, upstairs has come a long way since then. That had to be, what, four years ago? Five years ago Now It's four or five years.
Yeah. Yeah. Um, Kumar, you know, you aren't, you weren't born the co-founder and CEO of Sera, that you have a distinguished career before this.
Give people a little sense of that. Yeah. So thank you for having me, Alan, first of all, and hope you had a good weekend and, uh, uh, great to talk, great talking to you as always.
Appreciate your, uh, support and, uh, in general for the community more than us community. And, uh, so yeah, thanks for the opportunity. I'm Kumar Ula, co-founder of Sera.
And, uh, prior to Sera, I had a chance to work in, uh, companies like, uh, so Adobe and Symantec, and, uh, worked there for about, uh, 16 years between both of them. And part of the challenge that ran into it in the Adobe and Symantec, because two software publishing companies, and they used to, we used to struggle and, uh, in the, in terms of like managing the overall software, supply chain, software delivery, uh, pro, pro, approximately $30 million. We used to spend 60% on people, 40 on tools.
Yet we didn't perfect the system that gave the inspiration for us to start the company and, uh, to how can we help enterprise companies, especially, everything is a software. Every, every company is a software company, and, uh, they have to manage the software at scale. We can bring the lessons learned from the experience that we had, and that, that was a genesis of Start the Company.
Our goal was to help the companies to the platform. We are the first company to came as a DevOps as a platform. And, uh, there a lot of the point solutions at that time.
Now, many companies had tried to follow the suit, which is great. And with the ai, with everything else is happening, and we, uh, it's getting the cumbersome for a lot of the enterprises to understand the value impact. And, uh, we can talk about many of those things.
But yeah, thanks for the, that was a quick background of mine myself. Absolutely. And great, great story.
Now, as we said, is four or five years old? Um, a lot's changed. Yes.
To say the least. Oh, Yeah, of course. Right.
Remember, four or five years ago, COVID wasn't kinda full bloom. We weren't sure what, what's going on next? Go not going to conferences, everything is on Zoom.
Um, huge. Everybody's rushing to do digital transformation and get, get out on the web. Today.
We're still rushing, but now we're rushing to do AI and Ag agent ai. And, and how it's, is it going to take your job? Is it going to, are you gonna make your job easier?
Are you going to, you know, what, what does it all mean? And, and if we thought that DevOps and all of that was big during the cloud, my God, this AI stuff is the biggest thing since the internet itself made be bigger, bigger talk to it, it had to have an effect on Sera, obviously, right? It's, it's affecting us all.
How has it kind of pulled your, your, uh, development cycles? How has it, you know, pulled the product development in new directions and so forth, and how have you responded? Uh, great question.
So just to gimme the company, right? So two and a half years back, or three years back, Chad, GPT came up at the moment of, uh, we can try some, you can ask something. Somebody will respond to you.
You can generate so many things. You can generate the code, you can generate the, the majors, you can generate the videos, what have you. Right?
So, but the challenge is, it's not as easy as it sounds in the enterprise world. Enterprise world. They don't have the, the shiny need to always across the board, they have a legacy.
They have a current workloads that they have to have then future, right? So you have to deal with all three of them all at once because they cannot, they don't have choice to, uh, make, they don't have flexibility. Even the offset.
I had to cater to the, we cater to Fortune 5,000 companies and we have to balance it out while we are innovating, directing the enterprises to go to ai, A-A-D-L-C, we had to continue to support the existing customers and existing workloads. So it was a sh shift in the first place. And also the important element that I will add, lessons learned from us last, uh, two and a half years dabbling with ai.
And that architecture's changed so much. So in the technology, what we used to believe that we have architecture in two and a half years back, it used to change every three months. Rate of change was so difficult to keep up with it.
Now we see that the overall AI architecture components have been settled down because that is a good news for me, at least in the market last six months, nine months, you have a ra basically, you need to have a supervisory agent, you need to have a rag, you need to have a vector database, you need to have a agents, and you also need to have a way to manage this. Models also stabilize these models have been designed for certain aspects. So combination of that with, uh, memory and conviction, everything else, you can bring our entire stack together, which is what we have done now.
It is enabling and helping with the lessons we learned. One of the unique advantage that we have is while we are in the business for five years, the five years worth of metadata, as well as the experience working with customers, what works, what doesn't work, we are able to put into the equation. Plus 150 plus integration that we built.
It's all coming together very nicely for us with the AI agents and the ai, um, as well as reasoning agents and survey. We've done with, uh, some of the cust, like a lot of the customers right now, which we're gonna talk about it and the how, AI code assistance, helping them or not helping them in some areas. And, uh, we definitely can talk to them as well, talk, talk about it.
So we shifted two and a half years back, but we continue to shift. It's not like one time shift. Even you can say that we completed the shift.
We don't think about it. It's a constant change in the ai. You have to keep up with it, otherwise you'll be left behind.
You may be the top of the pyramid on the day, three months back. And if you don't be, if you don't care full enough, six months after that, we'll be bottom of the pyramid. Just like what happened to open air to Gemini.
Open Air was leading the PR show. Gemini slowly came back behind them and they just took over. And some of the aspects of it Don't count out Anthropic and Claude either.
Yes. Andro is, uh, They're making strides lately. They got a lot of momentum.
Yes. You know, but let's face this, this is why we get up in the morning, right? Yes.
This is why you're in tech, man. It's, it's the, it's the number I Story, it's the mission that we have, right? Yeah, absolutely.
Yeah. Yeah. Absolutely.
That's the mission that we have. There is an, is a line in the movie, the Godfather. This is the life we chose, right?
Yes. And the go far the part too. Ah, that's good one.
Yeah, absolutely. This, this is what we signed up for. Yes.
Um, speaking of which, you guys recently, Sarah, recently, uh, completed an industry benchmark survey. Yes. Tell us a little about it.
So we, we've been one, we are fortunate enough, one other thing I wanna talk about it. Like, we are fortunate enough to be partnered with GitHub early on. And, uh, we started with before copilot existed, right?
So as a result, we saw some, um, what do you call, uh, pixie dust here and there, and also sprinkle people there, sprinkling here and there. We caught the, we got the wind out of it, and we were able to, uh, connect to the GitHub copilot faster than anybody else, and got the initial insights and be able to share, share the view to the enterprise customers. Okay, you bought this 10th or th thousands of licenses, what's the value?
What's the ROI, what's the time saving? What are the time savings and what's the productivity, quality, security? So that we, we, we've been already doing it, the unified ops, unified insights, but with the co assistance and co-pilot and bio is leading the pack initially.
And then, uh, how we are able to partner with GitHub and be able to show the value to enterprises. So that can, kinda like last two and a half years, we've been helping many enterprises, uh, in the journey. So we thought we could release a, based on the metadata based on the survey and based on the, uh, uh, interaction that we have and what we see with partnerships.
And we want to release some of the study because a lot of the studies out there, but there's more data, data oriented, right? It's a sentimental, uh, not a sentimental analysis. It's a data oriented.
So we talked to about 60 plus enterprises and about about two 50,000, uh, developers across the landscape and some of the excess that came out of it. A lot of people we see that adoption at the peak level, like I know adoption is there 90% of enterprises that are using, they're using more than copilot now. But Q copilot is leading the way for sure, no question about it.
And then, um, some of the things I can share, if you don't, that is okay with you guys. No, please. Yeah.
So we, what used to take, like, there was a technology and health, uh, startups that leading the pac, in this case, the finance and healthcare and insurance, and some of those things will come later, but little, little behind, not not that far behind. They're about nine to 10, 12% less than the technology companies in that adoption. Uh, because it's, it's, it's ex expected things because, uh, you, we have, we have to go through so much process training, validation, compliance and aspects of it, right?
And as a result, fast time to PR is a pull request for those people. Understand the development landscape, if you think about a Jira story and pull request is fulfilling the Jira story, right? So you, in simple terms, and that is improved significantly, about 48 to 58% depends on the size of the l size of the team, size of the company, size of the enterprise, and which, which also leads to that.
Um, next, next, next thing we identified, GitHub copilot is still lead in the pack, but what used to have 90% or 80, 85% now kind of settle down at 65% because at least in the car landscape, what we talked about, you brought up plot code brought up. Another one is Cara. They're also like, they are taking the little bit of the share, which is, uh, people wanna do, they're not replacing co-pilot.
They wanna have co-pilot plus something else wanna compare and contrast between both of them, which is expected in the most of the enterprises. They don't wanna go go with one cloud, they want to have a more than one cloud, similar to that. Not one AA system.
Couple of them as well. And definitely like, there's improvement in the quality and uh, the lines of code has been increased, but we don't measure by that quality. Definitely, uh, increased.
And uh, but at the same token, there is a little bit of that AI code assistance are also causing the problem. The technical debt is increasing number of line lines of what PR used to be, a hundred to 200 lines now 800,000 lines. It's taking more time for people to approve it, validate that.
And security vulnerability is also increased, um, at least 15 to 18% of that because it is a, there's something to watch, watch for because the more you code you produce and security, just because AI is producing the code doesn't mean that it's highly secure and it's validated everywhere. It's generating from the same code that human brain or some, somewhere it is written. It is taking inference and pray to give the context around that.
So by no means code that is generated by AI assistant is secure. That is a wrong assumption. People shouldn't, should not have that.
And sometimes we trust the systems then, and the what we, you trust in humans. But this is the early trend that we are seeing. I'm not saying that it's the trend that is gonna continue early trend.
Another thing that we noticed it lot of the times the early adoption is high people, people, uh, shell out money and licenses. We see anywhere between 80 to 21% licenses that the people purchased. They're not being utilized effectively.
If we are really put into the uses of, of those tools, uh, the licenses, they will be able to increase the productivity even more. But in a, in a nutshell, AI code assistance, definitely helping in the loop activities, which is building the product, right? They're not helping shipping the product.
That is one of the other thing. We know that because shipping the product requires your security quality, CACD, pipelines, governance, and bunch of other things. They're not.
Auto loop is, it's kind of like left alone, which we can talk about it in a minute. And my analogy is like this, DevOps is a freeway. You've been in the DevOps for quite, quite some time, mainstream, 10 to 12 years.
And, uh, it's, I put this a tool in freeway DevOps developers, operators in platform engineers and quality security, we all fight for the space. What has happened two and a half years back, code assistance came and changed the developer persona from two x to 10 X. Yes.
But guess what? The operation element outer loop is still stuck in the same mode. Nobody wants to touch it because it's a messy world.
So many tools, so many permutations and processes not there. A lot of DAY, human touch and so on and so forth. Yeah, No, I, I think as a result of AI where developers go from two x to 10 XS actually is gonna go down.
Yes. Not even stay the same. Exactly.
That's the point. We are going to widen that which we can save the conversation in later time. Something is coming in the next few weeks and we are going to open up the whole freeway for people to drive the same velocity.
They're produce in the code. I love it. Kumar, where can they go get this survey?
So they can go to the, we publish this survey in ops ai. I can send, send you the link quickly. Gimme one second.
I can post it in the, your audience to have the, uh, video and we just announced it last Friday, last Thursday. And, uh, I can, uh, share in the link so that you can put it in there. Actually even better than that, if you don't mind, a little DevOps dot plug.
I know Mike ards done an interview, uh, article. Yeah. And I believe we have it in there as well.
We, I do have it here in my chat. We'll put it in the notes. com, there's an article on it there by Mike with Kumar.
And um, I believe the link is there as well, but we will have the link in notes. I see it here in chat. Yeah.
I also post both the links in the chat and, uh, for your benefit, I posted that. So other thing which we have not, we are giving that within survey that we have done, now we have industry benchmark. One other thing I want to mention that we know the industry benchmarks.
We know the actual within the industry, what are the seven, eight verticals, technologies, startups and manufacturing and, uh, healthcare insurance and financials. And also like some retail aspects of it. We collected the data.
Now we know, if you were to understand, few enterprises want to cause based compare, right? Hey, where, where am I standing in my, uh, with my p right? They wanna see whether I'm doing something right or wrong, right?
We will be in a position to offer that with the data that we have, which is a beneficial for them. The second thing is we have a, a reasoning agent within the Sera. What it does is, for example, you are a leader and you are A-C-Q-L-N for a minute, and you have a thousand developers working for you.
You wanna understand what is my quota assistant impact time savings value, ROI, is it really helping me to improve my throughput? If it is improving my throughput, is it helping me to improve my velocity? If it is helping my velocity?
Where is, where am I, where, what is my security quality posture? What is the time to market? Ultimately, how do you connect them in the loop auto loop, which is, uh, building the product to shipping the product along with Dora Metrics, space metrics, and uh, along with your devs.
When we do that, instead of you watching the dashboard and giving the highlights of that, you can interact with the agent just like charge GPT experience. You can ask questions, okay, if I increase my 10,000 developers, if I'm getting the 50% acceptance rate, improving my velocity by 20%, can I increase by another 50 people? What is my impact?
It'll give you the summary recommendations, graphical view, everything all at once for all at once in one place. So this is the new paradigm shift in how things are going to shape up for, for enterprises not only the leadership, also managers, the and, uh, scrum managers, product managers. They can use it for the resource allocation and then they can also, for a planning purpose, they can also have a predictability way.
This is something we're excited about it. That's one of the reason we put the data together and brought the functionality into Sera and put them in the Andes averages as well. Love it.
Gomar, thank you for coming on and explaining all this continued success with of Sarah. You know, it's, it's gonna be an interesting ride this next months and years as the whole world changes. Absolutely.
I, Sarah about it. And, uh, we are looking forward to talk about Agent DevOps in the next few weeks with you. Okay.
Excellent. Thanks for having me. Appreciate.
Thank you. Always a pleasure. Kumar Cola, co-founder CEO of Sarah here on Text Drunk tv.
Thank you. We're gonna take a break. Thank you.
Thank you. Bye. We'll be back with more.
Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. I'm your host, Mike Bazar. Today we're with Alex Yay, who's the CEO for GMI.
And they are a provider of a cloud service that hosts a lot of GPUs. And we're gonna talk about, well, how complex things are getting from an infrastructure perspective. 'cause we have multiple types of AI models.
Many of them are, shall we say, multimodal. And it all needs to be run on some type of infrastructure that well is increasingly hard to get access to. Alex, welcome to the show.
Thank you, Mike. Thank you having me. Alright, I feel like there's two things going on that are kind of constraining our ability to operationalize ai.
And so let's just jump in. But the first is complexity. It seems like there's a lot of different piece parts and things that I gotta kind of organize.
People are using multiple models, they're using text and video and all these other things at the same time. And it's, it's hard to stitch it all together with a bunch of APIs. So how do we simplify this if we can?
Yeah, absolutely. So I'll, I'll give you some, some, uh, I guess some background, right? So we, uh, we are a GPU Cloud company, but having GPU infrastructure is not enough, right?
We need to have a scalable infrastructure and easy to use tools. And that's why we come up with the API platform where api, everyone can have access to full modalities, full access to every single models out there, from open source to closed source, from all the open ai, the world and traffic of the world, the Google of the world. And here's one thing we realize is having these, uh, snapshot of these models are not enough because people are building workflows, right?
That's what agents are. It's a workflow, a string of different models. And having a strings of building these workflows is the true application, what people are actually building, not the model itself.
And so this is where we come up with a solution called GMI Studio where it's basically a workflow builder where you can bring up multiple models APIs together so that you can string them together and build an actual application so that you can be, it can be used for not just professional ML experts, but uh, not just, or even coders. You can have, you can be a creator, you can be a influencer YouTube influencer, and you'd be able to build your own agents or tools. Mm-hmm.
The second thing is, there just seems to be like the entire AI supply chain seems to be constrained. I can't find GPUs. There's, wait, we're waiting on data centers to be built.
Heck, I can't even find enough memory. Um, so, um, is this gonna be a year where we kind of just deal with some of those constraints while we wait for, you know, them to kind of play themselves out? And maybe 2027 is a, a much bigger year for, um, deployments in production environments?
Or how do you see this all playing out? This is definitely a year of constraints. I thought this would be, uh, less constrained.
Uh, you know, and when I was, uh, in 25, I thought 26 was gonna be okay, but it turns out everything was rising. You know, gold price, silver price now potentially copper price, memory price, GPU, price, everything and data center, uh, everything is constrained. And if you are a, I would say ML company or agent company or JNI in general, you have to plan ahead.
This is not where, hey, I just raised and I'm gonna find a couple thousand cards. You have to plan three to six months in advance even before you start fundraising or before you're, uh, you're planning to write a check. I think that's the, the best advice that I can give to people.
Uh, and here's something that we're doing on our, our end, which we are locking in memory pricing forward 26 and 27, as well as data center capacity all the way to 27 and 28. So we're planning three years in advance in order to secure these precious supply in either power or, uh, or, uh, memory nps. Of course.
The other thing you, we hear a lot about these days is various AI accelerators that are being positioned as alternatives to GPUs, especially for inference. What's your perception of those alternatives? When do I use them?
Are they really mature enough yet? What's your sense of what's going on here in terms of our processor options? Yeah, so training Nvidia still dominates the world.
Uh, I would say 95 even more percent in terms of market share, uh, to training workloads. But inference, they're starting to see quite a lot of different usage. But those are for much larger business.
I would not recommend any smaller startups to use anything that is, that is outside of, uh, NVIDIA's system because if you have any bugs, you can just find any CUDA code, not any, but find cuda code to support you. But if you're using other ASIC systems, then you would need actually a team of people to debug for you. So this is only for the philanthropic of the world.
If you raise less than 10, less than 1 billion, then you shouldn't, uh, think about using anything that's not, uh, Nvidia because that would just waste too much time. AI companies should just focus on pushing products out and building amazing product instead of thinking about their infrastructure scaling. Let that hard work to infrastructure providers like us.
You mentioned Cuda, it's a pretty awesome framework when you look at it. A lot of thought when into it, but a lot of people are also concerned that maybe it just locks us in a little bit too much. So what is the bound to be struck between taking advantage of something like cuda or alternative frameworks or I don't know, might be one day see Kuda running on things other than GPUs just for Yeah, so, so again, right?
I think startups should be thinking about scaling, right? Building products and killing your competitions in other companies that is building, building the same same product in the same sort of same category. I think that's the number one focus, right?
When you think about diversification, that's usually for much larger businesses, right? For example, like the, well, like philanthropic or like SpaceX or not SpaceX, uh, X ai, right? These type of larger giants, then they can start thinking about, okay, maybe I should use TPU where you should, should, uh, A MD or, or uh, Nvidia, right?
Leave that infrastructure decision on us, right? We are highly incentivized, we're aligned to build you the most cost efficient, call it token per dollar, right? For, for you or else you're just gonna choose another provider.
So we want to take care of that entire infrastructure and scaling, stability, reliability, and obviously costs. That is something that we are extremely good at. Um, and so I think that decision shouldn't be the main focus for startups.
That's my opinion. Hmm. Um, are we getting to the point now where maybe I need to narrow my number of use cases that I'm gonna really push this year because, well, it's going back to that constraint question, but um, do organizations need to maybe pick two or three things that they're gonna focus on?
'cause right now, I feel like last year everybody was in this kind of, let's let a thousand flowers bloom and everybody had a pilot. But if, if I'm gonna really make an ROI case for something, do I need to narrow my choices and kind of put more wood behind a couple of fewer arrows? If you're a tradition, if you're a CIO from a traditional enterprise that is non-digital native, then I do not suggest you to buy GPUs.
I would only suggest you to start building with APIs, start building application and do your test case roll out internally, roll out to your, call it sandbox customers and see and iterate. Once you're done with that, then you can start thinking about scaling and putting things in a dedicated endpoint with GPUs or even thinking about, uh, fine tuning. I think that will be the step 1, 2, 3, uh, to, uh, to roll out your AI application.
So in that capacity, I think you should just think about how to field your workloads will be the central message to, to DCIO. Don't think about the these GPUs constraint, you can just pull APIs from us or from other API providers, uh, or even the, the hyperscalers, even though they're very, very expensive, um, on building your, the patient first. Alright, we'll come to that.
Um, the hyperscalers, of course are having GPUs and any other thing you might imagine that you wanna access. So what is the difference between somebody like GMI who's very focused on GPUs in the cloud and the service and what I might get from the hyperscalers? Um, first of all, I like to, uh, you know, I've interviewed a ton of CI and CTOs and not one has said that their cloud bills is too cheap.
Uh, so I think that's number that, that's number one. So I think cost is a significant, uh, hurdle for, uh, for companies to use. So, um, attest on a, um, particular hyperscaler and just spinning up an instance, it costs $150 a day, just one server, one task, one person.
Imagine you have 10, 10 agents building and you're rolling out to thousands of your employees where, or, or, or customers that cost would be inhibited. So, um, I think that's number one, right? Number two is we are highly optimized for ai, and that's what I was mentioning about token cost as well as token performance.
So what we do, two things really great is we will drive that cost down. Second is we will increase that throughput. So A GPU is a GPU, right?
Everyone has the same H 100, H 200, but we're able to increase that throughput by three x to six x. Why? Because we're building on the new AI native way where the hyperscale is built on the old CPU codes that VMs everything that lo they loses control over the GPU power where it we extract and amplify the same system with much better, uh, underlying infrastructure that makes the GP performance triple to six x.
Mm-hmm. Um, as you kind of think this through for a minute, is there, um, some notion of, uh, am I, do I need to be smarter about what GPUs I'm using when, and I ask the question? Because a lot of times I go talk to these data science teams and they seem to be very obsessed with the latest and greatest GPU that comes from Nvidia.
But if I look at the workloads, particularly if I'm starting to build things that are maybe smaller models, aren't there older generations of GPUs that might just work just fine, that are less expensive and more available? That's absolutely right. So, uh, it, the gps are application specific, right?
If you're going racetrack, you should use maybe a race car, right? But if you're going, going on, you know, uh, a mountains, you should use a, you know, a four by four maybe, right? So like a four wheel drive.
So it really depends on application you're building. So if you're focusing on low latency, then you should use smaller models, right? If you're building a large reasoning, deep research model, then you probably use, you have to use the, the latest and the greatest GPUs.
So it really depends on application. And I think the great thing about using, for example, our studio is you don't have to think about any of those constraints. Literally just pull API and think about what you wanna build.
That's what we want default to think about. It's like, stop thinking about all the infrastructure. Just think about like what exactly you wanna build.
Okay? You wanna build a voice agent, or you wanna build some, a customer support agent, you wanna build internal G pt. Just think about that, right?
Leave that complexity like what GP to use we can optimize behind. Like customers won't care about, like what's, what's actually powering the GPUs just as long as it's, it's, it's, uh, you know, uh, it's in, in the right, you know, regulatory frameworks like the GP in, in America or in a, in a, uh, have the right license, right? We, we handle all of that, right?
And we optimizes we may mix and match different GPUs, which we are by the way, that lowers that cost together with a crazy science. Like we, you, we, we do a cluster inferencing. You're not running GPUs, like you're not running inferencing on a single GPU.
We use a full cluster at different GPUs. But again, we, you don't have to think of any of that. Just think about like what application wanna I build?
How do I scale my customer? How do I scale my application? That's all you have to think about.
And it's like, oh, how do I market it? How do I sell it? That's it, right?
Le leave that things going. Do you think there's also, uh, the market and the way we manage buyer's perspective and the customer's perspective is evolving? And I asked the question because early on it seemed like there were a lot of these tiger teams for ai, and they included the data scientist and somebody who was the infrastructure specialist and off they went.
But as we kind operationalize this more, I wonder if we're starting to see the IT team play a larger role in the inference selection side, because they're gonna look at things like Kubernetes and things that they can scale. And a lot of folks are saying that the infrastructure conversation for AI is moving over to the IT department while the training stays with the data science folks. That's absolutely right.
So you're talking about, okay, so the company has built your application, now they're thinking about scaling and for enterprise, uh, they have a much stringent, uh, requirements than startups do. And I yes, that typically w we have as a is your observation is correct, is typically with it t and the first I would say mistake that will they'll do is, Hey, I'm just gonna buy like four servers and put it, put it in my data center or, or like in, in, and in their office typically have like a server room. And then when people are using it and then they crash, because first of all, it's very difficult to manage the servers, the infrastructure below OS level and then above OS level on how to separate those into different containers, which is meaning that different users will come in and each user shouldn't see other users', uh, workloads, right?
So you're talking about like security solutions and softwares, uh, and, and then you have to build the actual application system sits on top and which crashes, which conflicts with I think maybe the Kubernetes instance and or even directly on, on bare metal. So long story short, the traditional, so CPU is very easy to manage and IT team is pretty, pretty good for that. But when you're talking about GP scaling, it becomes a difficult problem.
And so what I've seen that they're now moving on to is basically, okay, I built my application, I wanna scale, and they'll talk to infrastructure provider like us and they'll discuss, Hey, here's my base load, here's my base usage, maybe like a hundred GPUs. And then they'll say, Hey, at peak we will hit 200 GPUs. So I would like to flex up to that and we can design those solutions for you and even cage up these servers, obviously with the base loads and flex up, it's all difficult.
Um, so, and we can build direct line, so all the regular regulational hurdles we can, uh, and data security, there are privacy issues we can handle for you. And, uh, I would say having a localized private cloud or hybrid cloud would be the best for any enterprise star. Once they have their MVP and they are distinct about scaling and use, they start getting user attractions.
You would, you need to start talking to your, uh, GP infrastructure providers and think about, uh, how to best accommodate kind of the tail, well best tailored towards what the end the, the user actually needs. So we can give you that solutions, uh, while not losing kind of, uh, uh, I'll say, um, reliability and cost in mind. So as you look into the coming year and there's a lot of possibilities and what's your crystal ball telling you?
The supply chain constraint will continue plan ahead. That's I would say supply side. On the application side, we are seeing the explosion of inferencing starting early last year.
Now towards 26, there are multiple companies building really amazing applications in agentic workflow that optimizes your kind of your day-to-day work, completing tasks to optimizing kind of customer service, right? Uh, and so this is really the year of true application as all the infrastructures are ready. I still remember, I think last year, uh, I was on, on the show and talking about GPU scaling and people renting GPUs.
And now just on our platform, we have 146 models available and workflow. I wouldn't had that last year. And just so many, so many models available, so many toolings and so many solutions.
Um, and so I truly think that this will be the age of scaling. Uh, and there will be two, I would say gen AI applications, uh, that would amaze people, uh, in I would say a year ahead. Yeah, and let me ask you one last question about all that.
So, um, how automated can all this get? Because you were talking about APIs and not having to know what's going on with the infrastructure. So ultimately, can I just kinda express my intent for my application and the infrastructure will just automatically take care of it, including figuring out what models may be to use dynamically based on what it is I'm trying to accomplish?
I mean, how smart can smart get, You're right. So in, in the past, AI is middle to middle, okay? Human is end, end-to-end, meaning that, okay, if I give you a task and you go to research and you do ask people or think about things, read books, something and you'll complete a a task, right?
Or you ask me questions, get feedback, but you'll complete a task. Human is end-to-end, right? But AI currently is middle to middle, meaning that it's a tool that people use for productivity gain.
It only exists on the, basically on your screen. But what I'm seeing a lot of companies are building, which is quite exciting, is they're making it end to end, right? It, it will control your mouse, control your screen, do research, ask you questions, and complete, actually execute the task.
And, and for example, they'll, and then they'll open your email, write the email and ask you, it's like, is this okay? Can I send this out? Right?
So I think I'm seeing the expansion of the AI task from middle inching towards the end, right? And so human can just be now approver, basically. It's like they'll give you a bunch of options.
You said yes or no, yes or no, yes or no improve case. And all you do is making high level decisions and let all the execution be done by ai, which is, which is and, and amazing, right? And so people who can use a AI tools really well, you're can excel at your work.
If you're a company, if I'm talking to a company, you're gonna crush other enterprise if you're able to use these tools. Well, All right folks, well, I think you heard it here. We're pretty much, I'm not even sure where the end of the beginning yet, but we got a long way to go before we operationalize this AI stuff to the end degree, but it's gonna happen.
Hey Alex, thanks for being on the show. Thank you, Mike, for having me. All right.
And thank you all for watching the latest episode of the Techstrong AI Leadership Insight series. You can find this episode, others on our website. We'd invite you to check all those out.
Until then, we'll see you next step. Hey everyone, welcome back here to Techstrong tv. I wanna introduce you to my next two guests.
First of all, let me first introduce Mark Helm. Mark is the principal product manager at BMC. Mark.
Welcome to Techstrong. It's great to have you on here. Great to be here.
I've been, uh, involved in the mainframe for decades and as a developer and now in product management of VI comp and BMC. And right now is the most exciting time for the mainframe I have ever seen in my entire career. Just with things going, it's a real renaissance.
It really is so great to be here to talk about. You know what, I think it's the most exciting time to be in techno. I like you, I've been involved in tech for decades.
It's crazy time to be involved in mainframe in everything. It seems where rein the world is shifting right in front of us. Uh, an interesting stat or I, I actually saw a recent survey item from BMC that more people, more younger people are working on mainframes now than almost ever, right?
It's not just people like Mark or me or Tony who, you know, have our, you know, well, you have great Tony, be happy the great year, but at least, at least you have it. Um, Yeah, but You know, it, it's all, it's all in here. And, and so, um, it, it's a great time to be in it.
Hey, I, I'm gonna take a time out, Paul. You're gonna have to edit this. Someone named Mark popped in.
It's Eric. Eric. Oh, it's Eric o all I know Eric, Eric's state.
We don't need to worry about Eric. We don't need to know Eric. He said it was Mark.
If he had told me it was Eric, I wouldn't have stopped. I, I, I At Mark State. Alright.
All right. Uh, so Paul, count me back in three, Two. So it really is just like such an exciting time and, and, and one would think, well, mainframe, what are they doing with AI and everything else?
Well, we're gonna talk about that, but first let me introduce you to our next guest, and that's Tony Anter. Tony is a DevOps architect, an evangelist of BMC. Tony, welcome.
How are you? Give us a little of your background. Uh, doing great, Alan.
Thank you for having me here. So, unlike Mark, I've been in technology for decades, but I've only been on the mainframe for maybe five to 10 years. Before that, I was fully distributed, uh, full stack Java guy, you know, e-commerce, internet, web, and I sort of got into the DevOps space years and years and years ago.
And then from there, I, I'll be honest, I sort of fell backwards into the, into the mainframe and haven't looked back since. I've only been here. But I like to bring, uh, a distributed mindset and a DevOps mentality onto the mainframe.
I think it's the best platform. And like you guys, I, what an exciting time to be alive in, in the mainframe space and in general, we get to be here at the birth of, of AI and see what it can do. And, and to your point, yeah, you wouldn't expect to take mainframe and AI and put them in the same place, but I think mainframe is poised to be a heavy contributor into the AI renaissance we're gonna see in the IT world.
I agree. Agreed. Matt?
Uh, so, you know, it's a funny thing. Look, Tony, like you, I backed into the mainframe world. Well, second time I was in the mainframe world out of it, and then I backed back into it just when I thought I was out, they pulled me back back.
Gotta throw the, Now we've gotta obligatory godfather in there. But, um, but you know, through DevOps, much like you through DevOps, because man, DevOps, DevOps ignited mm-hmm. Mainframe there, it brought the mainframe into the modern I it stack again, even though it never really left, it allowed us to do, you know, things like running Java, I mean, it always had containers, but, you know, modern containers, modern kinds of stuff, and, and really, you know, systems of record, systems of engagement and being able to run them one in the cloud and one on the mainframe and making 'em work together, applications that panned both sides.
But it was funny, all through that, there was always an a small minority of folks who said, oh no, we gotta throw your mainframes for, right? Yeah, we've invested, you know, 10, 12 billion more dollars in this, but let's throw it away to go run on on some virtual machine or something. And, you know, and these applications and, and you know, I remember telling people, penny for penny, dollar for dollar And beat It, nothing is as secure as reliable.
Correct. When you want rock steady mission critical, critical infrastructure. Why?
I mean, why, why you don't throw out what's not broke. Yeah. And so, you know, mark, are, we're going to, even though Tony and I are, you know, contemporaries of yours, right?
You, you have more time in mainframe than us. Why don't you, why don't you lead that one off? Right?
You know, I, I think mainframe really got into a rut and you mentioned like DevOps and um, I actually wrote, oh, 10, 15 years ago, a column on it that the mainframe got sleepy and we were developing like the eighties or nineties, and we were just stuck into, that's it. And the distributed world in the late nineties, early two thousands were doing these amazing things, and the mainframe people just ignored it. And suddenly it's like the mainframe woke up with DevOps and start implementing all those practices and has done a pretty good job of catching up to it, it be, and not being so dismissive and of other platforms.
It's been really exciting to see the mainframe be revitalized by this. And I think it's the younger generation coming in and bringing in new practices and saying, you don't have to do things like this. You don't have to treat new people coming to mainframe.
Like it's a hazing ritual and you have to work like the past. And I think I see the AI as being that same type of thing as DevOps was. It's just that another injection of excitement and change and that you can do all these things in the mainframe.
Agree. Yeah, I mean, I think Mark hit it right on the head there. It's like people stopped innovating in the mid to late nineties and just froze.
And then when I, you know, when the, when I started leading transformations mm-hmm. As you said, like when we backed into the mainframe, I came into it with a perspective of why can't you do this? Explain to me, Mr.
And Mrs. Mainframe developer, Mr. And Mrs.
Mainframe CI prog, explain a technical reason why you can't automate your builds. Why you can't automate your deploys, why you can't automate testing, why you can't do any of these things that your distributed counterparts can do. And the answers I got were, well, we've never really done it this way.
It's just not how you do it. You just don't understand. I never could quite get a, a technical reason.
And from there it grew. And I think the days of DevOps being some mystery on the mainframe are gone. I think now it's just like on the distributed side, it's table stakes.
Agreed, agreed, agreed. But you know, people focus on the applications, but sometimes we re we forget what underlies that application. And that's data number one.
And number two, very important business logic, right? It's the logic Yep. Behind the application that provides a lot of that value.
When you look, you wanna talk about technical debt, when you look at the business logic that is preserved and codified, if you will, right? It's the, it's the, it's the tribal knowledge in many cases from, from companies who have been using these mainframe applications for decades, the idea of tossing them out, babies with the bath water, right? For the sake of modernity, let's say, or whatever else you wanna substitute modernity, it just makes modernity better.
I, I'm, you're making fun of my funny French accent here, but, um, but, but you understand what I'm saying, right? Mm-hmm. It, it just doesn't seem to be Right.
Reasonable. Well, you know, Exercise. No, and you're, you're exactly right.
And that's something that I've been really following for a long time is you mentioned the business logic is trapped in there and there's been a movement of people say, let's just move it to Java. And I'm like, well, first of all, you don't understand it. How can you do that?
So what we have is an approach we call the A BC method, analyze, build, and convert. So first step is having people be able to analyze and understand those applications, because you're right, the business logic's there, but new developers don't understand them, and there's a fear of change. They're afraid to work with them and to exploit them and use them.
So first step is analyze it and then refactor those applications, break them apart because they're usually monolithic cobalt programs, tens, hundreds of thousands of lines. That's not a modern architecture. Re-architect them and that makes it easier to work with.
And then you could look at maybe converting, But, but Mark, I, I don't disagree with you, but to the point, and, and we saw this during COVID, right? We saw it during COVID, all of a sudden the call went forth. We need cobol, uh, program is 'cause the, you know, the unemployment application in New Jersey didn't work anymore or something like that.
But changing that, right, right. Changing it off of COBAL then was that was like switching the engine in the car in the middle of the race. Yeah.
And we need to stay away from those kinds of, you know, gun to your head transformations, let's call them. Well, I mean, I think that's why, right? It needs to be an Yeah, that's why we came up with the a, b, C methodology, right?
Because we are not about just jumping off the mainframe just because, right? You have applications that have been around since the seventies, eighties, written in the early nineties. The people that understand them are gone.
I know you, you're trying to move something to, to use your race analogy, you're trying to switch out the engine in the middle of the race and you don't even understand engines, right? Right. So at the end of the day, you need to understand what you have.
And I think even the most, um, uh, even the biggest mainframe zealot will tell you there are workloads running on the mainframe that shouldn't. There are workloads running on the mainframe that really fit for purpose, should not be running on the mainframe. But to your point, for your data, for your business processing, for your batch, you can't beat it.
You cannot beat it. Maybe someday you will, but right now you can't. So the biggest companies in the world aren't staying on the mainframe out of love, out of love for IBM, love for cobol, love for any of it.
They're staying there because they get the biggest bang for their buck. Right? And in order to, to convert that, you're going to have to understand it.
And you're gonna have to take a, a, a slow boil approach to how you do this. You're going to have to break these, these massive monoliths down into smaller chunks, understand them, build them, and then convert what is necessary, right? To your point, don't throw out the baby with the bath water.
Understand it. And I know that's a hard rub for a lot of upper management to hear, but that's the truth. You want, you want the most risk averse approach that you can take.
And we think we've come up with that. We think we've come up with a methodology that gives you the ability to understand what you're doing before you try to do it. I, I, again agree this A, b, c, analyze, build, convert.
But I think the important thing we need to remember, 'cause people, people make this mistake all the time. It's all or nothing. It's all or nothing.
It's such A mistake. A hundred percent agree with you, Alan. Yeah.
Mark, you're shaking your head too. What do you, I mean, we see this play out, Right? Right.
All or nothing is, it scares me. It's extreme risk. It's that incremental approach, deciding which things, you know, you may say, I have this logic here that we change all the time, and I have a whole pool of people who can work in Java, comfortable with Java.
I want them to do it. Okay? Find that, first of all, in this massive application, find that business logic, pull it outta that program and maybe convert that.
But the rest of it still runs cobol. And it's fine if it hasn't been touched in 10 years, don't touch, change it, leave it, leave it there. Only convert the things that make sense to convert, refactor those things where you change them a lot.
So it just makes more sense and it's a lot less risk and you get a lot of, uh, value right away just from the understanding, rather than we're gonna take six months or a year and change it. It's like as soon as you put these tools in place, you have immediate value. Yeah.
And I think you have, I think also, Alan, to tap onto what, um, what Mark was saying, I think you have a lot of misunderstanding. I think you have a lot of people who have now inherited the platform due to retirements, due to, due to, you know, your SMEs leaving and their com it, it's what it it's, I I don't wanna call it confirmation bias, it's familiarity bias. They're familiar with the, with the cloud, they're familiar with the distributed systems, they're familiar with open source.
So that's the direction they want to go. And that's fine. That's a good direction.
I've been that direction. It it, it works out in a lot of cases. But again, you have to go back to what are you trying to do?
The value of systems is data. The value of systems is, is the data that it controls, the data that it owns and the business rules that affect that data, right? The platform, the, the logic, the um, the uh, uh, the code.
It's all I, I'm gonna use the word again, table stakes. It's all table stakes. It's what you do with that and it's how you make your money as an organization.
That's what's important. I, again, good, good points, guys. I wanna bring up something else though.
And it's at the nitty gritty of this issue. I first ran into it when I, I was doing, I was hosting a, uh, a podcast for the Open Mainframe project from the Linux Foundation and, and great organization, right? Project Zoe, all of that.
John and, and me and uh, mm-hmm. Uh, I forgot the woman's name. May I believe who, who runs marketing.
They're great people. We did a, I did it for over a year and a half, two years about, and we had several panels like this around mainframe modernization. And here's what I found out, and I'm giving it, yous an outsider.
You guys are inside on this, right? But as an outsider, for too many people, especially for a company that got bought by another company, and I don't know what they're doing with their mainframe business, but for too many people, mainframe modernization was code for rip it out, right? They didn't wanna say, we're gonna rip out the mainframe.
They said we're gonna, we're gonna do mainframe modernization. And, but yet for others, mainframe modernization was more along the ABCs. It was more along, well, let's, let's be selective and build a better architecture.
Let's transform not just modernize, but let's do transformation. But not, not just to rip it out, to rip it out sake, not some sort of risky migration. And we'll see if this can handle that kind of load.
But you know, Tony, back to what you said, yes, there are some things that maybe we could do cheaper, better, faster on, on, not on the mainframe, but there are a lot of things we can't And, and knowing the different, it's like the old saying, just because you can doesn't mean you should. Right, right. And I, I mean I, this is a problem in the community though, right?
There's, there's a lot of us out there in the mainframe community who when they talk modernization, are talking extinction. You know what I mean? And I, and that's impractical.
I think it's, it's bad business, it's bad logic. It's, it's, it's just not the right thing to do. It's skipping steps and going right to a conclusion.
Yeah. You've already, in your mind, and like Tony said, it's a comfort zone and you say, alright, we're just gonna get rid of, I don't understand it, it's just there. I'm not gonna look at it.
But what we're asking is take a good look at it, run the numbers and understand what loads should remain on the mainframe. And for those, invest in understanding them and refactoring and selectively converting. And you'll get a lot more benefit with a lot less Risk.
I think when we have these discussions too, Alan, I think the mainframe, and when I say, when I say the mainframe, I mean the collective us that is the mainframe needs to be real with ourselves that we have not done ourselves any favors. Right. That whole, you know, rib Van Winkle thing we did from like 95 to like 2015, didn't help, didn't help the industry, didn't help the mainframe.
And that's the reason why some of these modernization, why the modernization path has taken the way it is. But we all know organizations and companies that are, what, 18 or 10 years into a two year plan to get off the mainframe. Yeah.
And you know, Famous last words. Yeah. At the end of the day, the mainframe is an awesome platform.
The Z 17 is awesome. It has AI built into it. They act with the Spire chips.
They have the ability to run your LLMs and AI directly on the platform. Now, think about that. You have your, your AI inches, uh, millimeter, milliseconds, however you wanna measure that distance to your data, to the data.
Not, not some set of data, not partial data, not, uh, hallucinations on the internet, the actual data, the customer data, the purchasing data, the financial data, the data that AI needs to make itself run. Mainframe is poised to be there. And it is, it would be negligent of us, negligent of, of us as an industry not to capitalize on that and, and negligent of the organizations that are, that are looking at it.
It may be a little bit of hard work and sure. But it is still, to your point, the best platform fit for purpose for what it does, which is munging millions and millions and millions of records of data in a very small amount of time, very efficiently. So, Tony, you opened the door, you let the, the, the genie out of the bottle.
You mentioned ai. I, I I noted the time because I'm surprised you made it this far without mentioning it. Well, we do, we mentioned it in the beginning, but in passing.
So like everything else in technology, AI's gonna have a tremendous impact on our usage of mainframes on the next generation, if you will, of mainframe developers, of all developers, but of mainframe developers. You mentioned the, the, the, uh, Z 17 already has GPU like capabilities built in. Mm-hmm.
Of course, you know, when we talk about ai, the move is away from like training on A GPU inference on, on chips, optimized for inference. No reason the mainframe can't excel there either. But, you know, there's two aspects.
There's the, there's the systems themselves, and there's the people using the systems. As we sit here today, talk to me about what role or influence AI has on each of those, and then collectively together. You want to take the, you wanna start off with that one Mark, and I'll chip in at the end, or, Sure.
I, you know, with ai, the first thing that came out was the explain capability. Mm-hmm. And for like two years, people are saying, could you know, I've worked on analyzing code, could you analyze this and tell me what it does?
And I'd be like, impossible. A few years ago Yeah. Became possible.
And that's been amazing because now someone new to an application, maybe they don't know COBOL can understand it. That is exciting as it is. And game changing as it is, is just the beginning of it.
Because the next thing is really having personal assistance. Things like MCP servers and agents to go in and anticipate the needs of a developer and go off and do tasks for them to really help smooth out any roughness they may have in working with the mainframe. And that's gonna be a real big change to me.
It's all about the developer experience. Yeah, and I mean, I think too, there's a lot of fear when you talk ai, and I think you, it's good, you noted the time is how long it took us to get here, but I think there's a lot of fear when you talk to developers about ai, and just, let's just be blunt, A lot of people think it's coming for their jobs. A lot of people think that they're not gonna need engineers.
You're not gonna need people. And see, I actually disagree. I think that AI is going to be a tool that is going to revolutionize how we do our work in it.
If you look back to the days when people used to make furniture by hand, you had to cut the wood, you had to cut everything by hand, you had to measure everything by hand. You had to take those hand drills and drill things by hand. How many people would go back to those days as compared to using power tools?
No one I can think of, right? And I think the same thing is gonna be looked at with ai. I think AI is going to give us, uh, an ability to move faster, to go deeper, to understand more than we ever have had before in the IT space.
And I think it's gonna be the thing that's gonna help get rid of some of the rub of modernizing your mainframe. It's gonna, it's gonna give you the ability to truly understand things that you've never been able to do. See, what it's great at is understanding massive amounts of data, parsing it, collating it, and understanding it and coming to conclusions based off of it.
And that's something that humans aren't good at. Just being frank, uh, my mom always said I was super intelligent and I was the smartest kid she knew, but I don't even think I can do this. Right?
AI is gonna give us the ability to be able to, to do that is gonna give us the ability to, to, to push this forward. And I think it's going to bring a renaissance on the mainframe if we play it right, if we do the right thing. And again, I'm talking about the collective we, not the three of us, Alan, that's great as we are.
So, You know, it is, it, it's going, I think have a huge, a huge impact. But there's two elements. And Mark, you touched on it a little bit.
I think the first part is from an educational point of view, you know, Tony, you referenced, look, we got some applications where the people who wrote the applications along, God, we don't even understand how it was written, why it was written. It's just that it works. But with ai, we actually have the ability to put that code in there and say, explain this to me.
Right? And all of a sudden, you know, now I can see, say the blind man, right? You, you understand that.
Now that gives you the ability to edit it and work on it and improve it. That's one use case. Another use case is though just a whole new generation of developers, whether they're writing these new apps in COBOL or next version, COBOL or Java or something else, it, it's going to give us the ability, you know, to make that mainframe dance, if you will, right?
Maybe in a new rhythm that it hasn't danced in before, and maybe better able to say, Hey, using ai, look at the, look at my business logic. Look at my goals here. What should I keep on the ai?
What should I move off the ai? How do I make 'em work together? I think, again, this is gonna really help not hurt in spite of, and I agree with you again, Tony.
People are afraid that it's, it is about is it gonna take my career away? But I think those are the kinds of use cases we want to see here. Yeah.
I, I think it is. And I think that, again, you're dead on Alan. Um, I, I think that it is going to give people, and I think it's gonna give new life to people that thought their careers were over.
I mean, we talk about the great, you know, the, what are they called? The great resignation. I've heard it called the silver tsunami.
I've heard it a hundred different names for it. But I think it's going to give us the ability to weather that and continue on with the mainframe. And I think it's going to make the mainframe more open and more approachable than it ever has before, which has been the biggest problem of the mainframe.
To Mark's point, I believe Mark said that earlier, coming onto the mainframe in the old days was like a hazing ritual. You know, it was like, how much pain can, can, can we put on you and see if you still survive and get through the other side? Yeah, go ahead.
And now, now you know, it, it, it's modern tooling. It's modern practices, it's modern development practices. And with ai, now I can take those monolithic applications and understand 'em in a way I've never had the ability to before.
And, and we're only talking on the app dev side, apply AI on the operations side, on the ability to predictively understand where things are gonna go, where issues are gonna come up, where we're gonna have problems. And that's a whole other side of this coin we haven't even explored Uhhuh. All right.
It's augmenting Teams, right? Yes, it is. It's augmenting teams and you know, there's this whole, whether you buy into it or not, about virtual coworkers and all of these things.
And in that way it does augment it, guys. I mentioned earlier about the, the worst time to do a modernization of a mainframe map is when there's a gun to your head. Well, you absolutely have to because things are broke and people are screaming.
I wanna bring up the concept of evergreen modernization, right? Where we're in, let's call it, you know, Tony, you are the DevOps guy. I'm gonna ask you to go first.
Continuous modernization, let's call it, where we're not modernizing right now because something's broke. We're modernizing it because we have the cycles and it's a good thing to modernize. What do you think, you know, and that, and by the way, it's not like a one time thing.
It's a, it's a continuous modernist or continuous modernization, if we could call it that. Um, what do you think about that? Yeah, so I've, I've always, so I did a talk on this years ago about, you know, continuous integra, continuous improvement, continuous innovation, continu what?
Continuous whatever. I think that it is, my mantra has always been, you're never done it. You, every day you wake up and it's rinse and repeat.
Meaning even when you get something that's good, give it a little bit of time and it's gonna be not good. It's gonna be things, you know, it will have passed it by. So I think this idea of continuous modernization is the exact approach we should take, number one.
So we learn from our lessons. So we learn from the fact that we don't go to sleep again for another 20 years and wake up and see if, if we're still relevant or not, number one. And I think number two is you need to constantly be tuning, constantly be, uh, um, striving for a hundred percent with the understanding.
You're never gonna hit it. You're never going to get to a hundred percent modernization, a hundred percent test coverage, a hundred percent, um, uh, efficiency. You're always going to have to be tweaking and tuning and end of the day, I mean, I don't know about you, Alan, I don't know you, mark, but that's why I got into this gang.
I, I got into it not to push a button, wake up every day and like, you know, in lost every 33 minutes, I gotta push a button. I got into this because I get to create cool new things all the time. And that's, that's how I look at it.
I don't know, maybe that sounds a little too much like a gift card, but, um, that's how I look at continuous modernization. And my mantra is, you're never done. If you're done, then it's time to quit.
Fair. Mark. You agree, disagree, or wanna distinguish Tony's answer.
I, I totally agree. It's, yeah. It's, Tony said earlier, you know, do you have these big projects?
And those just fail? So if you say, we're gonna modernize and we got this big project, that's not gonna have a lot of success. But if you say to the developers, make things better every time you touch the code, you know, it's the backpacker in me that you leave that site better than when you came.
Pick up any piece of trash. You find. Same thing for developers.
It just becomes part of how they work. And every day in every little way, making their code better as they're in there, when you make a change, it should be better. And that should be part of the code review.
How did you make this better? Did you improve a comment? Did you break something apart?
If you have that as part of you how you work, and it's in your DNA, then it just happens naturally and it's a much better approach. So you're always modernizing your applications. It isn't just one big project, and we're gonna turn it on and turn it off.
Continuous Modernization. Yeah. I want you to, I want you to think about this like, like an oyster, right?
You have to think about this like, like an oyster. And an oyster gets a little piece of sand down in there, and it's an irritant, right? And that modernization is gonna be a constant irritant.
But if you work it over time and you do it, the oyster turns that into a pearl, okay? And that's what we're shooting for. We're always gonna be creating that pearl based off of the, the, the irritants that get in there.
I know that's a weird analogy, but it's one that always comes up. We talk about this, That's excellent. And energy, we'll Run with it.
Um, guys, we're, we're low on time. I, I wanted to ask just each of you, you know, there's so much, some of us who, you know, working on mainframe, like Mark for decades, some of us have left, come back, left, come back. And then some of us are new to mainframes and maybe we're even new to technology.
How do you, how, what, what's your recommendation to stay on top of the game? What resources from BMC would you recommend to people to, to stay on top of this, to, to, you know, be at their front line, if you will? I'll, I'll go ahead and go.
I would think that Well, alright, good. Tony, we'll let you go. We can give Mark the last word.
Yeah, no, that's fine. That's fine. Mark should have the last word.
Um, uh, to me, it, I've always been a, a, a believer in learn your fundamentals, right? Learn your fundamentals, understand how to write code, understand your patterns, understand your platform, and the rest of it will just come to you. If we treat mainframe like it's some, and, well, let me back up.
I think we've been treating mainframe like it's special. Like it's some kind of different thing for too long. If we go into it with the idea of the, of the, of the standard fundamentals that we have for development and programming, then we can fit it into the mainframe.
So somebody new, learn your fundamentals, learn your patterns, go out and do your education. BMC has a whole set of education out there that you can use, that you can learn. You know, go out and, and understand.
Understand the platform and what can be done with it. Don't be, don't be, uh, shackled by what was, you know, what's already there. Think about what you want to make of it and then do it.
Uh, and I think Mark can bring this land, this plane now. Yeah. Uh, definitely.
And I would add that you should follow us on, um, LinkedIn, like BMC or Tony or, or me. Um, because we update you on things so often I find that people are not aware of all the improvements we've done, all the APIs, all the things with ai. They, they could fall back into that Rip Van Winkle sleep, and you really wanna keep awake.
So follow us, go through like our documentation, our blogs, uh, the videos and keep up to date. We've done amazing things in AB and date, a product that's been around for over 40 years and it really is set up to do a whole lot of web hooks and APIs, but if you're not aware, you don't know about it. So really follow us and look to each new release to see the things we're putting out.
Good advice. Good advice. Well, mark and Tony, you know, this was the longest 15 minute interview I've ever done.
I think we're over a half hour, but it was well worth it. Lot of great information in here. Uh, really important stuff.
Thank you both so much for coming on. I appreciate it. And we hope you've enjoyed it.
You know what, any questions, as Mark and Tony said, look them up on LinkedIn. Go to the BMC website. Don't Look at the Modern Mainframe podcast.
That's something that we put put Out all the time. God, another good one. Yes, Yes.
The Modern Mainframe Podcast, right? We DevOps dozen finalist. Um, go check it out.
Check it out. You know, mainframes aren't just for me, Tony and Mark and people who look like us, they're for you. And go check that out for now though.
This is Alan Shimmel for Textron. Thanks everyone. Have a great day.
Welcome to the Security Boulevard, the cybersecurity podcast from the Future Room Group. Each episode explores a variety of topics within cybersecurity and the technologies that drive it. com, security Boulevard, YouTube channel, Textron tv, and all of your favorite podcast platforms.
Before we dive into today's episode, let's meet today's panel starting with Mitch. Mitch, it's good to see you. Hey, great to see you, Mitch Ashley, I lead the software lifecycle engineering practice analyst practice at rum, kind of dabble in security, working with Fernando and team and software, supply chain, security, all that good kind of thing.
It's good to have you and, uh, the other co-host this week. Uh, Mr. Fernando Montenegro, Fernando, I, I Montenegro, I lead cybersecurity research.
And, uh, we had, we barely started, there's a correction to be made. Mitch doesn't dabble in in security. Mitch knows security, right?
So, uh, um, it's interesting because there is a, a, like, as the industry's changing and whatnot, the, the, the overlaps between, uh, security and observability, for example, right? Uh, he's amazing at this. And, and, and I rely on, on his judgment and knowledge, uh, many, many times during the week.
Well, that's in this episode right there before Fernando changes his mind. Thank you, Fernando. And of course, I am Tom Hollingsworth event lead for security at Tech Field Day, a part of the Futurum Group.
Let's dive into this episode now as we're recording this. The weekend was a little bit messy for most people. There was a huge weather system that tracked through the United States.
Uh, there were power outages, uh, schools were canceled. They even closed a couple of waffle houses. And for those of you who are not in the us, closing a waffle house is tantamount to shutting down the military, the government, and anything you've got because it is the most reliable service that we had.
But that made me think a little bit about the way that we plan for disasters, because a giant weather system may not be the kind of disaster that you think is, is going to cause a problem. But what about a security disaster? What if, uh, you know, I don't know, some of your, uh, password files get leaked online.
Or what if somebody manages to abscond with some of your data or crypto lock it? Uh, how do you respond to that? Can you respond to that?
And worse yet, is your response to that, oh, I better go check to see if my disaster recovery plan is up to date. Because we all know that sometimes people don't think about disaster until it's upon them. So I'm gonna open this up to you, you gentlemen.
Have we ever seen one of those situations in the past where it felt like the disaster recovery plan was making it up as we go? Oh, I can, I could name a few names and a few companies, but I think I'd get in trouble if I did that. No, I think, I think, you know, it it, it's funny.
We call it a disaster recovery or, or business continuity plan or whatever it might be. It, it, it's sort of like any plan that you create, whether it's disaster recovery or a project plan, is outta date. As soon as you hit save, 'cause the environment changes, something's different, something's new.
You know, the, the CEO comes to you and says, we're gonna do this now. So it's, it's a little tough to create a fixed plan. Uh, and so I think you're almost really developing a system, and the plan is just a reference implementation of your system, of what your disaster recovery process systems, people, et cetera, look like.
And I think if you take that kind of an approach, you'll have a little more to borrow a term from, from Fernando's cyber resilience practice, you'll have more resilience in that system as opposed to, well, we did that in, uh, what is it, oh, 13. I think we're a little out of date. We probably ought to update this.
Yeah, a lot of things have changed since two th 2013, so, Yeah. Well, so, uh, the, we also had, so I up in Canada and we also had weather here, uh, I here in the, in the greater Toronto area. I haven't checked the news recently, but it, it may have been like record snowstorm.
Uh, it's no accumulation, thankfully. It's, it's, uh, relatively light and fluffy snow. So I can, I can deal with it.
Don't You call our, I'm sorry to interrupt, Fernan, don't you call our, our kind of storm that we're having, you just call that Monday up there, don't you? Isn't that Sort of No, no, no. Listen, listen, I, and, and, and, and tying back, tying back to this topic, right?
I think that one of the things about preparedness is that you always have to have a, um, uh, a threat model. Like your threat model needs to have, uh, needs to be realistic to your environment, needs to meet your requirements, and needs to, to account for the things that, that you are, uh, willing to, to, to, to address. And in many places in the United States this week, uh, they're not, like, they're not statistically ready to, or it's not statistically, um, relevant to them to have preparations for this.
So my heart goes out to the people who have been affected by this. I know that, uh, um, I know that, uh, that, uh, it, it was very disruptive in many places where an inch of snow, an inch of ice is literally the community stopping kind of stuff. So, but yes, like the, the, the amount, these kinds of amounts up here in, in the greater Toronto area.
Yeah, that's, that's Tuesday, right? Uh, funny enough, uh, if you go further north in Canada, right? People who live further north think that about up here in Toronto, right?
So, so like, it's, it's fine. Like it's, uh, It's all relative. It's all relative.
But, but, but Tom, to to, to your point on, on preparedness, right? I, uh, uh, there is that, uh, quote that's always attributed to Eisenhower, uh, sometimes variations with, uh, uh, with, uh, Mike Tyson, right? I mean, the, uh, plans are worthless, but planning is everything.
And the other is everybody has a plan until they punch in the face righter. But, uh, It, it's a really good point. And, and to illustrate that, I actually wanna bring up something since, uh, Mitch mentioned 2013, uh, a lot of people lose track of the planning process because they never actually test their plan.
I worked with a customer many years ago at a previous job, and, and we don't get a lot of snowstorms in Oklahoma, but we get the other kind of weather that tends to level buildings, uh, especially in the springtime. Uh, we are in the middle of tornado alley, and, uh, I was working with the school and they had a solid backup plan. Uh, you know, we're gonna put things on tape, and if something happens, we've got the tapes.
And then one day they had to evacuate the main building because of a, uh, a tornado threat. And someone realized that their backup plan wouldn't work if the tapes were still in the building, if it was gonna get hit. And they had to literally run into the building to grab a box of backup tapes and throw it in their car as they're evacuating.
Fast forward a couple of years to 2013, and one of the largest F five tornadoes ever recorded, hit another school district administration building. Wow. And they had to sift through the rubble to pull out the drives of the servers, to insert them into different servers, to be able to run payroll for the teachers to be able to buy supplies, to start rebuilding their houses and things like that.
Because they never thought what would happen if we got hit by a, a severe weather event so big, it literally leveled the building. You don't think about those things until you're in the middle of a disaster. And that's one of the reasons why you have to test your disaster plan, because you need to know where the failure points are.
Oh, the generators will immediately fall over if the power goes out. Really? Did you test it?
Will they automatically fall over? Or does someone have to go hit the big switch to cause a failover? If that's the case, who's gonna do it if everybody's pinned at home in the middle of a snowstorm?
And how does, you know, things like DNS records work and what's gonna happen if somebody's logged into the servers when the power cut happens? Like if you don't test it and then jot down all the failure points, you might as well throw the binder out because it's useless to everybody. And I spot on.
And I would argue there's one step before that, which is what goes into the plan in the first place, right? And I think that that is an area where, uh, bringing this back to cybersecurity, we've had, uh, multiple cases where incidents have happened where it's not like, it's not that they were completely novel, right? It was the case where, yeah, if you think this true a little bit, and oh, by the way, that can happen, right?
And this is an area where I am, I am, uh, I'm really interested in the work that's being done in adjacent areas, right? So in aviation, for example, there's an entire field of study on near misses, right? So why did this almost happen right then?
And, uh, I know that in cybersecurity, I know that, uh, Adam Tack has written about near misses in the past. Mm-hmm. And I know that, um, and I think that like, uh, Wendy Na and Bob Lord, uh, we're talking about CSA last week.
So Bob was at, at CISA not too long ago. Uh, and, uh, I think that they're doing some work on near misses as well, right? Which is how do we as a community learn from the near misses, right?
This bad thing. Why did something horrible? I mean, it almost happened what got us there, right?
So, um, yes, very much about planning and very much about thinking through, sorry, where being creative about what can go wrong, right? It's risk management. My kids hate when I, when I, when I talk about risk, I use it to tease them, of course, right?
But, uh, it is risk management, right? And I think that, that, that our jobs in, in this industry is to elevate that risk management conversation. Tom, to your point, precisely too, think through about these things.
By the way, we do have an official mascot now, Nate, the cyber cat has joined us. So this is Nate, everybody good buddy of mine. You know, it, it's, it, it's interesting just kind of my own experiences with disaster recovery plans.
I think one of the takeaways for me was, you know, even a plan can be a bad plan, right? Just because you have a plan doesn't mean it's a good plan. And because you've tested it doesn't mean it's foolproof.
That the things that worked when you tested it, tested it could still fail, right? You still could have some problem where the UPS doesn't kick in like it's supposed to. The cooling doesn't happen like it's supposed to.
But even more simple, more easy, easy kind of simple things you would think about, don't assume, like one of the, one of the plans that had to just completely revamp. 'cause it was, um, really more than a decade old and all the technology was changed. Um, but when the assumption was that such and so employees live close to the building, and if this happened, if the com computer data center overheated, they could come open the door and cool it down.
Well, what happens when you camp the ice storm, the fire, the whatever? And, and that there were things like that, that we addressed and, and, and fixed and made sure that that wasn't the, wasn't the the main way we were gonna solve it. And then as it happened, within five years, and I think it was then another three years after that, we first had floods that went right up to the edge of the building.
Nobody could get to the building. And a couple years later, they had complete fire. A whole bunch of houses got built down, businesses got, uh, burned down, um, right near our offices.
Again, can't get into it, right? Can't even get into the neighborhoods to do anything. So physical access is not an option.
So I think that's one of the things when you think about the, the, uh, risk analysis, Fernando, is think about what assumptions we make that we shouldn't assume that's true, even if we've tested it. Does that ring true? I think that one of the things that rings very true is that the world is unpredictable.
The world. Like we can't guarantee anything, right? And, um, one of the, one of the concepts that I learned along the way is that I love, and, and, and it's very, very relevant to this.
It's the notion of resulting, not sure if you ever heard of the concept of resulting. So, uh, I read about it from, so Annie Duke, she's a, she was a poker player and now teaches decision science. And Annie Duke wrote a book called Thinking that, uh, basically how to make better decisions.
And the idea of resulting is when people mistake the quality of the outcome with the quality of the decision, right? And we cannot control the outcome. We can control the decision, right?
But we cannot. And and you may have made a horrible, a poor quality decision that still turned out okay, right? Oh, I'm gonna, I'm gonna go on my tiptoes on this thing to pick up that one thing over there, as opposed to getting the ladder and I still climb, right?
Um, I say that it's particularly relevant. I don't mean to bring football into this, but, uh, as we're recording this, the, they just defined the, the play, the, the teams for Super Bowl 60, right? So Seahawks and Patriots, uh, 11 years ago, they had, that was the fame, the fame, uh, uh, matchup.
And any, duke uses the example of what happened in Super Bowl 49, uh, as an example of resulting because the, the, the Seahawks made the play at the one yard line that didn't work as people expected, and they were crazy about the result. But then when you look into the decision that was that the decision, that was a found decision that led to that play? Sure it didn't work because that's life.
That's, that's football, right? That's the, but, uh, sorry, I'm meandering as usual. But it's this, this notion of what is the quality of the, the decisions that you were making as a professional.
Are you, are you thinking through like the checklists that you need to do, Tom, to your point, to make the plans, like to think through. Like, how good is your decision process to create those recovery plans, to create those, those plans that, uh, do they match reality? And I think the point that you guys are kind of bringing up that it's important to realize is that a lot of these plans are based on a series of assumptions that we need to make sure that we can validate.
So here's a good one. Let's say you have some kind of a data protection system in place that is creating regular backups that are stored offsite. What if you have to log into active directory to restore those?
What happens if active directory is no longer available because it's been violated or it's been shut off or something? 'cause we're seeing that a lot now in cyber attacks, is that people are going after backups and people are going after active directory to halt any kind of, you know, cross system functionality. Well then how do you get your backups back?
Can, can you get them back without being logged into active directory? And those are the kinds of assumptions that you have to challenge. Kind of to Mitch's point, you know, someone will always be able to go over to the building and open the door, but what if they're not?
How do we plan for that? And I'm sure that you guys are probably sitting there thinking to yourselves in the audience, you know, oh, this is just a big tabletop exercise of how are all the crazy ways that I can, I can get around this? Well, in security, those crazy ways to get around things are exactly what your attackers are thinking of, right?
They, they want to try to hit you from an angle that you're not expecting, that you haven't planned for. I mean, I saw something today about, you know, people trying to get account, um, access by sending text messages or communicating with people, uh, through email going, Hey, this is, uh, at t we just wanted to reset your password because we noticed somebody's trying to log into it. Can you give us that code that you just got texted?
It looks legit. All the links work, except, oh wait, that thing that you just got is actually gonna allow me to take over and then I'm gonna start hammering all of your other accounts for two factor stuff and whatever. Like, those are the kinds of things that you have to be ready to adjust in your plan.
What happens if an admin gets compromised? What happens if your CEO's email starts spewing out, um, you know, spam or, or attacks or something like that? You have to think through that.
And I realize that a lot of this feels like exception handling, like, you know, programming in a auto, an autonomous car. Like what happens if a 7 47 lands on the highway? Like, you're right, the, the likelihood of it happening is low, but it's never zero.
So, like how do you About this is, go ahead, Tom, about this is, it's not a linear process either. The, you know, they're, the, the, what you see as the attack may not actually be the attack. There may be a secondary action is, which is really to get you to start the backups.
And that's when they're gonna compromise them. 'cause the system that you're, you're restoring, uh, two is, is actually the one that's gonna compromise and is gonna steal the data. It's kinda like in, in, I guess in tank warfare, not to get too militaristic about it, but, uh, shells that penetrate tanks, it isn't the first, uh, contact with the tank that the shell penetrates.
It, it's really, that's just setting the stage to get through the explosives that happen that actually would then open it up to be penetrated by the follow-on, uh, secondary explosion from the, its, uh, shell that's attacking the tank. So it's, you have the added factor of what, not only what if, what this happens, but what if we're not able to do what our action, the response is what if we're suspect about whether that environment's compromised or could be compromised if we take the restorative action. So it, it's, it's a multidimensional kind of three dimensional chess, if you'll, and tie back to Star Wars and or Star Trek and Spock.
Yeah. And this is what, this is an area where like, I, one of the, the areas I work closely with is, is that, uh, part of the reason that the, the practice that that I run is called cybersecurity and resilience, is that I work very closely with the, with many of the, the, the backup and recovery vendors that are, that, that work in this space, right? And one of the things that they do is precisely this notion of, okay, what does a clean room restore look like in a scenario where we need to restore ad before we do anything else?
As a matter of fact, a number of them are now tiptoeing their way into identity protection precisely on account of that. Uh, not to throw the, the AI into it as well. And, and they're using AI to help, uh, optimize that process.
The, so it's, it's an evolution of, of, of, um, it's an evolution of the technologies, an evolution of the capabilities that practitioners have to review their plans and say, okay, alright, what is it that, what assumptions are we making? Oh, we're assuming that we need to recover ad ourselves. Oh wait, our vendor can now do that for us.
Okay? Can we trust that the vendor can do so? It, the, the, it changes the nature of what you need to do.
You're not recovering ad you are making sure that the vendor can recover ad for you, right? So, uh, the evolving nature of the, the, of the disaster recovery plan is not only that things drift for bad, like, but sometimes they drift for good. Like, I mean, there are new capabilities into your environment that you didn't have before.
How can you make best use of that? Right? You know, so, Nan optimistic, Fernando, Fernando, question for you.
You know, my thought is just like, you might use AI to analyze your business plan that you're, that you're creating, or maybe it's helping you create one, certainly ai it could be a good, uh, sounding board to run your disaster recovery plan through and test it, you know, put pressure on it. Um, but also could be for scenario testing. Like what are the scenarios we're not thinking of?
What are the latest kinds of scenarios we should start to consider how we adapt to, so AI could be your friend actually in helping you do a better job or keeping current with what's happening. And you can rule out the, you know, far edge cases. 'cause you don't think that's practical for you to even validate or test or protect against.
But certainly you're not relying on just your thinking. It's like using an external consultant to help you. Absolutely.
The, the challenge there, the, the him every, every episode at some point, AI comes in, right? Uh, uh, the challenge, the thing I I challenge people there is that, look, yes, I agree wholeheartedly, but who is running the AI and what, what are you using the ai, how are you using the ai? How much are you depending the ai, are you a backup and recovery specialist who is interacting with the AI to ask, hey, um, like precisely as you said to do the, Hey, let's, let's think through this, this scenario.
What am I missing kind of thing. Or are you someone who is not as experienced in backup and recovery and you're coming to the AI for, oh, tell me what to do, because I dunno, right? And if it's the latter, that is complicated because you don't know how to evaluate the, the quality of that outcome, right?
As well as if you are an expert, uh, and you're just using it as a tool. It, we go back to the AI as a, as a tool for an experienced professional versus, uh, uh, a less experienced person using it. And, and not catching the hallucinations, not catching the mistakes, not catching the, I say this as somebody who's using AI a lot, uh, for my, uh, uh, personal finance tracking, right?
I, uh, it's great, but I know what I'm doing, right? I, I know what to ask. You know what I mean?
I, I think it's important to realize that a lot of the pieces that get missed are kind of institutional knowledge. And I think where AI is gonna fall apart is that no one's ever documented this. Mm.
And and that's one of the reasons why I know for a fact, that's the reason why I originally started writing on my blog, was because a lot of the things that I was learning were things that maybe didn't exist anywhere else. Like we all know the XKCD, you know, the infamous one. You know, who are you Denver coder seven?
And what do you know? Because so many questions get asked that never get answered. And when someone does answer it in their head and go, yeah, that's how this works.
They don't ever write down what that, and in the old world that was, oh, that means that I've gotta solve that problem. But in the modern world, it's AI doesn't have a knowledge base to draw off of to create a solution to that problem. And so we, we kind of find ourselves with that knowledge gap, right?
'cause this is what I have been helped with up to this point, and here's where I need to be and how do I jump that gap? Because I promise you, an algorithm is not gonna be able to jump that gap no matter how many GPUs you throw at it, because it really doesn't know how to think outside of the box. That's where humans are still valuable in the loop.
What happens if these conditions aren't met? What happens if this really random exception occurs? And like, you know, it could be something as stupid as like a race condition.
The power comes back on, but the network isn't up and it starts timing out because it'll not connect. What do I do now? We don't know what the answer is until we have tried it right now, before anybody else goes out there.
Do not walk into your office tomorrow morning and just throw the master switch and go, let's see what happens. You need to kind of, you need to game through this. This is what I want to happen.
This is what I expect to happen. Let's test this in small scales to see what occurs. And then let's analyze what we tested to make sure that if something does go wrong, it's contained and fixable before we try it on a larger scale.
Because if you do that, just yank, let's see what happens. Uh, in the industry, we call that an RGEA resume generating event. You, you do not wanna be on the receiving end of one of those, especially nowadays.
Funny, like I have so many things. First of all, right? Uh, Mitch, you're based in Denver, aren't you?
Yes. Yeah, he knows. Like, who knows, he might be that Denver holder Tom, who knows, right?
Could be. Could be. Yeah.
Could. But uh, finally, finally, you mentioned that the, the, the throwing off the switch back when, when, uh, I was working on, on network security implementations, uh, I remember working with a, a healthcare customer where we built, uh, uh, multi-site, uh, uh, network with the firewall modules and switches and stuff like that and, and, and so on. And, and, and you, you can imagine what the architecture looks like.
And, um, one of the steps on the, on the, the test plan was, okay, yank the power from the, the the, from the, from from one the, from the, the, the, the rack, poof, yank the power watch, okay? For 1, 2, 3, 4 seconds pf converges. Okay, we're good, right?
Uh, but the, the state crossed over like, okay, we're good, right? But yeah, we had those things like, but to your point, that was not the test, right? That was one step of a very, very well-designed test plan for, but it was a fun one to do.
I still get nervous anytime anybody tells me to do anything destructive on purpose. And I'm like, are you sure I'm not gonna get in trouble if I do this right? It's in the plan.
You're witnessing me do the thing that's in the plan because it, it feels wrong to purposefully break something. Oh, yeah. But that should be replaced with this, um, joyous feeling of, Hey, my backup plan worked as soon as it happened, whether it's logging into Azure active directory instead of the local copy if something blows up or Yeah.
You know, we can get that data back and we can do a fractional restore so we don't lose like three months of database tracking. That's called an escalation of a resume generation event to career ending event. You'll quickly become the story that gets told at every conference for the rest of your life.
Exactly. Hey, remember that time that Fernando did x I've, I've, I've made a few mistakes that what, thankfully, I don't think any of those rise to that level. I'm sure not sure not, But, uh, but to go back to, to, to the, to the topic, right?
I think we're all circling around this notion of people needing to be, uh, needing to, to, to have the space and the knowledge and the, and the, the, the systemic thinking around creating the plans and the models and the, the, the countermeasures and so on for what they think is, is realistic. And, um, one of the areas that this is, and, and, and we need to be pushing ourselves to keep doing that. So we had the, the whole incident with CrowdStrike a couple years ago, right?
Listen, it was, uh, uh, very, very impactful as we all know, right? But realistically, at which point, how far do you go in your testing, in your assumptions before making a decision of, you know what, I'm gonna trust this, and if this blows up, oops. Right?
I dunno. I, I, as somebody who did endpoint security, I'm, I'm, I'm, I'm always nerve with around ker mold stuff. So, Well, I think it, it, it's an important question to ask because past a certain point, there's not much that you can do.
CrowdStrike was actually a really good example of that. It's like, what happens if the colonel faults and everything goes offline? Well, it doesn't matter what my plan is to bring that thing back up if it will not come back up.
Or we run into the other problem of resource utilization, right? If I only have a limited amount of time or people to bring the business back online, what should they concentrate on? Because, yeah, I don't know.
Let's say we have another AWS outage, like, I can't fix that. I, I can do everything I can on my side to make it work as well as I can, but at a certain point, you have to throw your hands up in the air and realize, this is bigger than me. And, and no amount of resources on my side, short of an infinite money glitch will allow me to fix this.
And, and that's the other thing too, because we, we've seen this time, and again, with disaster recovery plans, it's like, well, what, what's our option? Well, we can have a warm site, uh, secondary data center where we're doing continuous replication over a private link, and they're like, yeah, that sounds excellent. And then you hand them the bill for what that's gonna cost per month, and they like all the color drains out of their face because they're like, well, how important is it?
You are like, well, if you want that to be a cold standby site, that is a forced data replication. And we're not paying the license for two active active storage units like the cost to go down, but the R-T-O-R-P-O is gonna go up because we are going to have to, you know, physically transfer media over there or something like that. And it's the back to that trade off.
Like, like, I can't plan for everything. I also can't pay for everything too. And here we're 30 minutes into the conversation talking about economics again, right?
Ah, See, we almost made it through the whole episode without saying economics. Thank you guys. But, but it is.
Right? But, but, but here's the thing. Our role is to work with our principal.
So the, the, the, the principal in, in, in economics is the principal agent problem, right? Uh, when you hire somebody, right? You want to make sure that they are doing the things that you wanted them to do, right?
And that you have enough oversight over them that they're doing it. I mentioned this in the context that, you know what, creating a, a, a, a full tolerant, active active is such a pain. I'm not gonna do that.
You know what? I'm just gonna grow and, and, and put a couple of my, put the server under my desk here and, and, and, and, and call it a day, right? Uh, if the person who hired me to do that is expecting that active, active recovery, and I'm not, and I just created a, a little server under my desk, that is a, that is a market failure.
That is a, uh, and the person who hired me has to have enough oversight, has to have enough knowledge of what I'm doing to, Hey, hey, Fernando, why are you keeping a server? What's, where's our, where's our active active stack, right? Um, it, it's a, it's not a trivial problem.
I don't, I don't mean to belabor the point too much, but it's the idea that, uh, you need to have incentives play a part. And the same person who, who's faced drains, because they don't wanna pay for the active, active, what are they, what are they promising to the people above them, right? If I'm an investor in the company and the company's supposed to be bulletproof, and then, uh, that that director doesn't pay for the active active, I'm not gonna blame the engineer who proposed the active, active and didn't get funded.
I'm gonna blame the, the, the person who didn't approve it anyway. Well, you would hope. You hope.
That's how it works. You know, there, there's kind of taking that same scenario, Fernando, of, you know, someone decided, I don't think I'm not into it today. I'm not gonna do, do an active, active, the, the, there's also the how far do you take things, right?
Because this is, these are rabbit holes. We can all go down continuously how far and, and you reach points where now this is out of our control. Um, but is there a remediation or an action that we could take if that happened, whether it's a SaaS service or a physical plant thing, or whatever it might be.
Um, you know, if enough things cascade together that that's a scenario where we would not be able to handle gracefully. Um, and I think that's, that's back to your risk analysis, um, of okay, how these are the kind, we're up good up to here. We think we're good up to here.
And we know that there would have to be several things that could happen together in order for this to escalate further. Can they? Yeah.
Well, they, they sure could, you know, it could be a Murphy syndrome, but, you know, we're gonna call the question there and say, that's the level of investment we're, we're worth. We are, um, willing to make. Right?
And I think that's the, the economic conversation that engineer owns is, let me lay out the landscape for you of here's what we need to address, what the, what the cascading rings of, how far we can take it, and what's the risk to the business and the value that we would assign to how much money we spend to protect from that some point. Yep. We're gonna take risks.
That, that's the, the meteor hitting the earth is not one we're gonna protect for, Maybe, maybe not, but yes. But, uh, I, we're running away and I go back, that's why I mentioned during the, during the, the, the chat today, this notion of the quality of your decisions, right? We, we have to, to have the good enough decisions.
We do exactly as here's what and beyond that, that's the role that I, Well, I think we're gonna have to wrap it here. It was a good conversation. Hopefully we have, uh, given you some food for thought about your disaster plans, and if your disaster plans didn't work out the way you wanted to, maybe this is your opportunity to have a meeting and kind of discuss that.
Uh, and when you do, you should de definitely check out some of the stuff that my, uh, co-hosts are working on. Fernando, what's something that you've got coming up that people should be paying attention to? So, I, uh, I just finished a report on, uh, cyber physical Systems.
The, the, the, the public version. I think it's out if you're a foot return subscriber, you have access to the full report. Uh, I'm starting to think about my, my next one.
And, um, at the same time, we have other reports coming in. And as a matter of fact, I think that Tom, you and I have one coming up in the not to distance future, uh, another signal report. So that'll be great fun to do together.
But, and, uh, prepping up for, uh, the RSAC conference and, and, uh, I absolutely love the, the, the conversations I have there, the travel. Yeah, a little love, a little less, but, but seeing friends in, in San Francisco is great. Well, speaking of r yeah, speaking of that, you reminded me, I have to get my slides turned in 'cause I have a speaking slot, uh, at R-S-A-C-I.
I think one of the things you'll see, of course, I spent a lot of time, um, researching and talking about AgTech development or AI assisted development, however, whatever term you wanna use, um, and I've kind of coined it as this year as the developer's role evolves to engineering agents in the way that software is created. That's really what the development roles versus coding, working on code. I think a lot of the same kind of changes are happening in the observability role and world.
And one of the things we've done is elevated in my practice observability. You're gonna see a lot more reports coming out from that, which is a nice intersection that, that Orlando, Orlando, Fernando, and I get to work together on. Oh boy, there's a, you know, a winner thought, let's go to Orlando that Fernando and I get to work on together.
So, um, it's, it's changing from being an operational tool to, it's actually moving left. Like we kind of talked about security moving further up the chain, particularly the way AI agents get developed, but even more so the interesting, the interesting change changing. And that's some of the things you're gonna see, uh, coming out from, uh, our respective practices and working together.
So I'm excited about that too. We wanna thank you for listening to this episode of the Security Boulevard podcast. If you enjoyed this conversation, please subscribe on YouTube or your favorite podcast applications so you don't miss any of our episodes.
We'd also love it if you'd leave us a rating and a review and a comment to help the show grow. com in the FU Room group. If you wanna check out show notes and see past episodes, check out security boulevard com, text tv, website, or the techron TV app, which is available on pretty much every device out there.
Now, make sure you're following Security Boulevard on X, Twitter and LinkedIn lvd, uh, there's a lot more content for you to consume out there. Thanks for tuning in. We'll see you all.
I am Ted Weatherford. I'm VP of Business Development. I've joined with John, who's our distinguished engineer and architect on our E-Series products.
I speed here. Yeah. And I'm gonna cover the x-ER products.
Now, uh, the x-er product is a ethernet switch. And what distinguishes it from all the other ethernet switches literally is that it's programmable and it's programming model is a load store type model. You have 3072 little Harvard architecture cores inside this ethernet switch chip.
Um, all the competition is approaching the problem with a fixed pipeline. They may call it programmable, but at the end of the day, it's a fixed VLIW pipeline with fixed elements and programmable elements, and it has a fixed latency and you don't have the flexibility to do tasks. In parallel.
What we've done is run a complete model with this 3072 little Harvard architecture cores. And it gives us two things. It gives us more flexible than any ethernet switch chip that's available or has ever been designed with an actual architecture I'm showing here that allows us to scale down or up really gracefully to call it an FPGA.
'cause it kind of looks like an FPGA floor plan would be a complete miss misnomer, but it is symmetric. And if you look at all the little squares, there's 64 of them and each of those 64 squares has 48 processors. Okay?
And they're run to complete. So when you go to program this, you're not forced into a pipeline, okay? Of sequential steps where you write the code, you turn the code 90 degrees and you drop it into the pipeline.
You literally can do recursion, you can do whatever you do within your instruction space and the time you have for the minimum packet size. We've sized all the caches, we've sized all, all the clocking so that you have this amazing low power device with tons and tons of header processing. You can, you can do 11 layers of MPLS, for instance.
So all kinds of IP, inside ip, any kind of encapsulation, whether it's already a standard or something that's developed. And I'm gonna, um, so yeah, thank you. Uh, alter ethernet, things like that, these sorts of things that are starting to come outta the woodwork UA link.
I'm guessing that's UA link here, but, um, how does that play in this space? It's exactly where this shines And for alter ethernet or eon or some of the ethernet centric extension protocols that are coming out for ai, we're, we're, we're, we're there. We can ship this and some of the physical layer mechanisms like the layer two retries, there's things in there that we didn't catch with this particular tape out.
But, uh, protocol wise, uh, congestion management wise, um, the PA in flight packets for the servo loops for all your congestion management, all that's dialed in and completely flexible, you build your own server for any scheme. Um, so it, it's, uh, it's ready for that, uh, better than all the other products 'cause they're fixed in function. Um, and that was the one exception is the retry.
And then on the, uh, the UA link reference, this is not a UA link switch. That's correct. It's not, um, I'll just offer the latency of this switch while we're at it.
Um, a UA UA link is a very low latency switch. Uh, it's backend scale up centric. This is a front end scale out or a backend scale out a a product.
Um, and we can get down to, you know, 400, uh, and 50 nanoseconds, which is screaming for a traditional scale out switch. I mean, your broadcom's rate 800 when you really measure 'em. So one of the advantages of having a run to complete architecture is you, you can dial in the latency to be as good as it gets.
So Ted And Jack Poller with Paradigm Technica, can you talk a little bit more about run to complete and what that means for You? Yeah, thank you. So, um, there's really three different architectures.
There's a Von Noman of Harvard and of data flow, uh, and run to complete, uh, speaks to, um, an architecture where you have an instruction set, uh, that you're running, right? Uh, and you'll run to, you complete the overall operation of the packet frame coming and being modified and being sent out. So the frame shows up, ethernet, it gets modified, it gets shipped on its way.
Uh, and what run to complete allows you to do is have complete flexibility. You don't have to, uh, handle the packet one piece at a time, sequentially. You can move around to anywhere in that packet header or that packet within a window.
So the run to complete is a amount of time or clock cycles. You have to do work, but that work does not have to be only sequentially. That's my level of understanding.
Yeah. It's more flexibility and it allows you to do things in parallel. And there are many packet operations that can be done in parallel.
So you, you end up saving, um, time and getting latency from it, uh, and not being caught off guard if protocols change or you have something interesting. Where it really matters is all of us use standards, right? Mm-hmm.
It matters with all of the instrumentation and telemetry people, the observation and the troubleshooting, these network is getting increasingly complex. And this kind of model just allows you to build the best instrumentation of telemetry, um, with, with really less constraints. Yeah.
So it's symmetric and it'll scale either way. So each one of those blocks is a 400 gig block. So if I want to go build a 400 gig, a tiny little switch, I just do one block.
If I wanna scale up to 400 terabit and, uh, and I shouldn't have said that number, um, then, then you increase the block size. Um, so it, it does have that sort of, um, geometric scaling up and down in this graceful, symmetric way. Um, which you won't see when you look inside the, the fixed function data center switches, you'll see pipelines and they're fixed in nature.
And there's a number of them. I'm not gonna cover this, it's an eye chart, but I want to point out that we have large amount of packet buffer. We're very low power.
8 T in derivatives. It's the same D but we down clock it. You can turn the siess down.
The siess can be all different speeds. So we're software defined even at the physical layer. You can have siess coming in at a hundred and going out at 25 with different modulation schemes.
You can mix and match. And this is actually really, really, really sexy because you've got these new fabrics up there with a hundred gigs that have just been deployed last summer, and they'll be around for a while. And then downward, you've got all the legacy.
So you can run it 10 gig, 25 gig, 50 gig, a hundred gig, 200 gig with whatever series and whatever modulations you want. So it really gives you this building block for looking backward as well as forward. 8?
It's 'cause we're going after the edge. And I'll cover that later. We're going after half rack, full rack, two rack.
We're going after satellites, we're going after base stations. Okay? Um, the performance is there, the chips are here, the boxes are out by our, our Taiwanese friends Act.
In an edge core, you can just measure everything, but you gotta show people some performance if you're claiming to be programmable and lower power than than a Broadcom or an Nvidia or a Cisco or a a barbell. So we do that, and that's what this is. It's the watts on the left and it's breaking down the certis, the core, you know, and it's showing what your max and and mins are over, over that.
8 T switch that's programmable at under 200 watts. It's disruptive. Um, this is showing you efficiency of the most important thing about a switch besides its rate switches or connectivity at the end of the day and how much bandwidth and how many different connection points or ports you can have.
But the other thing that matters, and it really matters is the shared memory or not buffer the frame buffer. The packets come in and do they come in fair? And can you utilize, in times of rustiness, can you utilize the whole buffer that you paid for without overrunning it?
So in our example here, we take 127 ports running it at a hundred gig each, and we ram it out 100 gig port, and we find out how long till the packet buffer fills up, does it overrun. And then we do it, you know, at the different, over subscription rates and different packet sizes. And we saw that the utilization never drops below, below like 86 or something.
And in real world tests. That's amazing. Okay.
So if you're a switch head like I am, then you're like, oh, wow, that's amazing. The, the, the Tomahawk products that are dominating the market, Ted. Yeah.
Regular ese silver. Okay. Still trying to get a handle on this.
You're not actually generate, you're not actually manufacturing dus, you're actually manufacturing the chips that would go into D or chips that would go into switch. We have two products, they're both chip products. I'm covering the switch first.
It's a separate tape out and it's just an ethernet switch chip with 120 800 gig UR on it. It's a, it's a switch chip. Next we're gonna cover our EERs, which is A DPU.
So we have two products. It is amazing. 200 engineers doing two products of this complexity.
Is it? It's a lot. So we have two chips.
They go together nicely. They have the same SER ip. They play really well together.
'cause the highest volume opportunity is the top rack and the front end interface or the backend interface into the server. Mm-hmm. Or the GPU server.
And so we book in that with these two products, especially for the edge, um, root chip company, two chips. Um, I want to, uh, I'm gonna pick up the pace a little bit. This is showing, um, the fixed pipeline and the map pipeline approach, which is not ours against this run to complete full SDN.
We can imitate the other architectural approaches. Um, they can't imitate us. So we can make the trade-offs between latency or the amount of work you're doing and the amount of power you're spending.
Um, so this is more of a deep dive for somebody that wants to compare these products to data flow architectures or fixed function stuff. But suffice it to say, we're trying to say that we have, you can have multiple pipelines, you can have branches, you can have, um, a physical connection and a logical connection that's flexible. Uh, Challenge with something, yeah, like this in the past has been latency.
I mean, to do a, to do something that's not pipeline or not map pipeline and achieve the line speeds has always been impossible before it's Pr we're proud of it. Yeah. Uh, I'll give you a clue.
If you're designing for your worst case and you're building a pipeline, you've got a lot of stages you don't need. If you build with us and you put our 3000 processors in, in a line, if you want to pretend it's a pipeline, uh, you can and you'll just have less instructions. Or you could have one processor handle a whole packet.
You can build a pipeline like they do. Or you can have one processor handle one flow. You have this whole range.
I guess the question is what's the clock speed? And you know, how sure, how fast are you being able, are you able to maintain line speed across, you know, however, 128 ports I guess. 8.
It's, it's mind bending. Um, and I can just say that the team, uh, is basically on their eighth or ninth generation network processor or switch when you combine the people. We have people from Motorola, Freescale, uh, you know, Nvidia, Broadcom, um, Juniper, Cisco.
It's a pretty senior strong team that's been doing network processing, DPU and switches for, you know, 30 plus years. Yeah. Now it, it is, you know, if you, before we had the products, it can be a lot more, uh, interesting debate.
Now we just have the products so you can just put 'em on the test and test. In fact, we have built-in self test that we don't advertise, but you, that's a really nice feature we have too. 8 terabit of, of line rate.
Every port has, its built-in Xia tester. Um, so you can even test the device in its own print circuit board without an expensive $3 million tester. So programming model is always the challenge for programmable products.
We have an assembler that we've wrapped in Python and we provide libraries. You've gotta configure the tables for forwarding the tables for security, the tables for quality of service, the meters, the counters. You have to set all that up and they, their structures.
And we have libraries and you have to program this thing in our assembler. However, we opened up the instruction set and our first customer, which we did a PR on, uh, called oxide, uh, they're a local, uh, club as you know, field. Yeah.
They, they wrote a P four compiler on it. So this is a simple risk instruction set. I shouldn't call it risk.
It's technically a Harvard architecture, but it's a small little instruction set that you're be familiar with if you're a programmer, especially somebody that really programs deep and they just wrote a compiler on top. So our whole ethos is it's open, do what you want. And we also have plans of putting out a P four compiler early next year also.
'cause we have a lot of customers that really want that, that higher level or what I would call fortune generation language's kind of dated terminology. But, um, so today we give you courses, we've got all the examples, um, and we haven't had a customer we had to write all the code for. They've all taken the classes and written the, written the code and uh, and uh, done quite well.
And one of 'em is, uh, a foreshadow is SpaceX. So we're really excited about that. Um, the architecture and the normal stack of what you target to on the very bottom, I'll start there.
You've got the switch device, that's what you really care about. But you could also do simulators and you can also do, um, hardware emulators. We built our own hardware emulator.
We have a room full of FPGAs of design. We did ourselves, we do all our own hardware emulation. We don't buy hardware emulators.
Uh, we have our eval boards, uh, which are just, uh, rack systems, uh, you know, pizza boxes with front panel ports like you'd see in a top of rack switch or fabric. And then this just shows the network operating system down and what we provide, um, and we put all this, you know, on open, anybody can get to it. Uh, and the network operating system of choice for all of us now is Sonic.
Uh, and we do our own sonic distribution and then we have two partners that provide hardened sonic as well. Um, so that's what your, your normal stack looks like. And this is what a box looks like.
I have it right over here. I just wanna say it's real, it's available. This particular one, um, uh, is what we call the universal switch because the MPA connectors are Q SFPs and you can put whatever you want in them.
Each little rectangle is four C days and the s can run it whatever speed you want. And there's common media for four by 25, 4 by 50, two by 50, et cetera, et cetera. So that you can build a top rack here, which we call a, a top rack upgrade or a tor upgrade so that again, you can connect to any kind of speed up and any kind of speed down.
You can migrate from older network interfaces to newer ones or maybe one storage box has a certain kind of thing and it's in the same rack with a server or a GPU server. Got a got a question. Yeah.
Specifically about the, the tour here. Um, do you think, and you can theorize here a little bit, do you think once 2 24 30 comes around, you'll be able to stick with this same form factor low power and not go liquid cooling? It depends straight up on how much bandwidth you want.
Straight on bandwidth. 6 T, you could stick with this. Yeah.
2 T would be harder. Yeah. Uh, it'd be harder.
Um, that's, I that I'd have to defer to a switch expert. Um, that's the answer. Okay.
Yeah. Maybe, maybe 51 too. But, but, And, and our architecture, um, is really on par at a geometry level with the others.
So if we do a hundred terabit switch right. It, it's gonna, it's gonna be a thousand watts. Yep.
So what we've done to just be so disruptive on power Now, full disclosure sure is we went to five nanometer when everybody else was, Was still In the old place, you know, was going forward with very large rated switches. And we did it to capture the edge, the economics of the edge, the power of the edge, and the right amount of connectivity for the edge. Um, I got an example on that coming.
Um, I'm gonna just move faster. We built this for a large, uh, hyperscale for this exact thing. Their particular format is just a different cage.
This is a 16 by 800. So that's the tour they happen to use. Uh, and this gives them a benefit of half the power, a half the Rackspace, um, uh, a, a quarter of the cost of, so that product was GA 24.
Yeah. So how long has this puppy been out there? We first sampled this April of 2024 and we called it generally available.
Both our products have been first spin, no metal spins to market. We called it generally available in November of, of 2024. And it's been in mass production since summer of 25.
Yeah, no, this is out there.