Building Incremental Observability with Broadcom’s WatchTower Platform
Broadcom’s WatchTower Platform™ is an observability platform designed to provide a unified view of mainframe and distributed systems. This platform empowers users to identify and resolve issues more swiftly. It seamlessly integrates data from diverse sources into a single interface, leveraging capabilities such as alerting, machine learning, application profiling, and real-time streaming. These features enable enhanced troubleshooting, performance optimization, and overall operational efficiency.
In their presentation at Tech Field Day Extra at SHARE Cleveland 2025, Broadcom showcased how they are building on their longstanding mainframe tools by layering the WatchTower observability platform over them to streamline diagnostics and integrate across the enterprise. The goal is to relieve operators and subject matter experts from navigating disparate systems by centralizing data for problem detection and resolution. The platform caters to both traditional mainframe operators and modern Site Reliability Engineers (SREs), enabling end-to-end visibility from applications through to mainframe systems, and integrates with popular observability tools like Datadog and Splunk.
WatchTower expands on legacy products such as SysView, NetMaster, OpsMVS, and MAT, offering new capabilities like data streaming, data visualization, role-based access, and topology mapping. These innovations allow for automated correlation of events, targeted alerting with contextual insight, and easier collaboration between SREs and performance analysts. Importantly, these enhancements are available without additional licensing for customers already using Broadcom’s core products. Broadcom emphasizes incremental value by progressively adding toolsets that improve real-time observability and reduce mean time to resolution, while still allowing experts to dive into native interfaces for deeper analysis when needed.
Presented by Michael Kiehl, Senior Manager for AIOps, Broadcom. Recorded live on August 19, 2025 at the SHARE conference in Cleveland, Ohio as part of Tech Field Day Extra. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-tech-field-day-extra-at-share-cleveland-2025/ or visit https://www.broadcom.com/solutions/mainframe/observability or https://TechFieldDay.com/sharecle2025/ for more information.
Transcript
My name is Mike Keel. I am a senior manager for the AI ops division in product management. I lead the product management team at Broadcom.
Uh, so I'm going to be talking about how we use Watchtower to build in incremental improvements into our, into our tools, and bring the observability into the mainframe space. So the, the first thing that I really wanted to cover is the personas that we were talking about. Whenever we talk about observability, traditionally, the mainframe subject matter experts and, and operators really need this deep knowledge of what's going on.
So one of the things as we started looking at our observability platform a little over a year ago is we needed to service this audience, right? We needed to be able to get them that knowledge that they need the deep insights, but we wanted to bring it all in one spot. So they've really had to look in a whole bunch of tools, disparate places, go around and look, log into the mainframe, look at logs, log into different products, go through issue diagnostics and all that.
Uh, what we wanted to do was, was centralize all that, give 'em all that information in one spot, but we also wanted to service this new an SRE role, which is really that, um, application focused environments where they are looking at these, uh, needs across the entire enterprise, not just the mainframe, when, when they look at, uh, something that's going on with the application, that they want to be able to see it from website, the whole way down to that mainframe database that the data is stored in. So they, they need that end to end visibility and the ability to kind of take that insight. So we had curate it into that application view and send it o over into these, uh, tools like Datadog, Splunk, elastic, those enterprise tools that have been out there for the traditional observability space.
We, we wanted to make the mainframe part of that. So these are kind of the missions that we have gone through, keeping this persona focused on, uh, our observability as we've continued to build this out. So the first thing that we started doing was talking to customers, and this is really how they look at problem situations.
Something goes wrong on the mainframe. They've got all these monitors sitting on the screen, they're an operator, has to figure out how to devise that. They start drawing on screens and it ends up looking like a mess, right?
They, they have absolutely no idea how things are connected, where things are going. They've gotta build all this on the fly. Uh, so this is kind of that traditional view of what problem, problem determination looks like on the mainframe and what we're trying to solve with the observability platform.
We, we, we want to get rid of all this and enable that system operator to know, Hey, I just got an alert. I just got a problem. I have all my data right here.
I'm just gonna go and get that SME that I need to fix the problem. Or better yet, the operator can fix the problem because he is got everything in the skills that he needs. So how, how are we doing this?
So the first thing that we did was we started to look at our traditional products. So we have Sphe, net, master Ops, MVS, Matt, these are all tools that have been around for years. They're very well adopted, they've got a lot of great functionality and features, but they're all di disparate, right?
We, we needed to bring a layer on top of that in order to enable this observability on the mainframe. So we, we built a bunch of tools that are kind of that first layer. So this is stuff that is available for both the Watchtower platform and for our customers to use.
So if they wanna be able to stream data off the mainframe, we have a data streaming service. If they want to have a data lake, that data streaming goes right into a database. They now have access to all that data.
It's all API enabled. They're, they're able to get to it. We, we, we built data visualization on top of that.
So if they want to be able to build dashboards natively, it's all right there built in for them. And they, they can do all this themselves. It can meet their business needs and where they think they need to go as as they move forward.
Um, and to service that SRE role, it's all open telemetry and enabled. So we're able to send the data out, they can get their traces, they can get their, uh, their, um, packet information. They, they can get that problem determination logs, all that stuff all coming through this, this platform of tools that they have.
And a couple of my colleagues are going to talk about these features in coming videos as we go through, through the day here. But where we really started to build that value is we started to build our own capabilities on top of these, these tools so that now we're enabling the customers with even more functionality. So our, our topology feature, it goes through, it reads the mainframe and figures out how things are connected together.
So if you want to know, Hey, I've got A-C-A-C-I-C-S transaction, how does that tie to, to your DB two tables? It, how's it getting into mainframe? Is it going through mq?
Is it to have a TCP IP connection? Is it going through CICS connect? All this stuff is now tied together for you automatically updated every morning so that that a, that operator can look through and say, Hey, I'm having A-C-I-C-S problem.
What else is connected to that? And they can look directly into topology and see everything that's connected that CICS region, and there's no longer any questions about what is connected together? Add to that, the, the alert insights, this is our, our alerting component that we've put together.
So all the products can now send with the alert in there, but what we added on top of that is the ability to issue commands when, when the alert comes in. So you, you get that same CICS alert, you can trigger assist you to send you a mess of metrics about that CICS region or the specific transaction that created the alert, and it's all self-contained in that alert so that you now no longer have to go and say, okay, well what was going on at the time of the issue? You've got that all contained in the alert, whether you're looking at it real time or three days later, all that data is there and you're able to figure out what you need to do, where you need to go.
All the products will do that too. Net Master will be able to give you any IP information that may be interesting about that, that alert, uh, Matt will be able to trigger any measurements that you want to be, be able to do on that application level. Um, instead just once again, just pulling all that information together so that it's all there in one spot.
Yes, All of these personas or job definitions, they are, you know, they're sort of an amalgam of thousands and thousands of organizations that you would've spoken to. So it's a, it's a, you know, most average case, but a site reliability engineer in one organization would presumably have a different kind of responsibility to somewhere else. So presumably some of these functions that they would be interested in will differ anyway depending on the organization.
So is there some sort of, you know, um, variation permitted in terms of, you know, who would need to see what information? Yeah, so e everything is kind of role-based. So you can set up who has access to what and where they want to see it.
Uh, specifically for the SRE, all we're doing is doing an open telemetry feed. So however they have their tools that set up on the other end, it'll show up the same way for them there. It's just another data feed for, for them.
Yeah. I love how you're defining, uh, the SRE by showing heat maps. That's my right in the background, right?
Um, I do think, um, you know, I think maybe Derek was hinting at this, um, that there's, there's some kind of blurry dividing line between the performance analyst and the SRE uhhuh. Um, maybe the, the, uh, the day to day is, is seems significantly different. Um, but, um, I I wouldn't be surprised, um, if at a lot of your customers, um, there's a fair amount of, of of close collaboration between those two.
Yeah. And, and that's, and that's how we're trying to, to kind of enable that, right? So the, the SRE gets the information, the performance analyst is gonna get the deep dive side on the watchtower, so the two of them can, can co collaborate on the same problem and get the information that they need at that time.
And then whenever they want to do that later, uh, collaboration that's a little more deeper in the, the SRE can log into Watchtower and see all that deep dive in information if they want to also, right. And be able to drill through those issues the same level as the performance analyst do. Do you think though, that this, the, uh, it's gonna trend further in that direction?
I mean, you're opening up a gate for that kind of collaboration, but in a year's time or two years, because as quick as that, it could wind up being a lot more fused and, and with a lot more than just a feed as the source of, I Would imagine Yes. As a, I think the bleeding edge customers are gonna probably be within a year or two. Uh, some, some of the other customers will take longer.
Um, but we are definitely starting to see that trend with the customers that are really trying to merge the, that application view in, in inside of their operations rather than dealing with the tech side. They deep dive into I'm A-C-I-C-S person. I'm a DB two person, right?
They're trying to go, I'm an application and, and moving that direction. Yeah, we're definitely starting to see that in some of the customers, even even with the mainframe side, right? They're fusing all that together.
Just kind giving an example of how, how we build this, this incremental value, right? At sis view, it's got a ton of data. It's our performance monitor on the mainframe.
It has thousands upon thousands of metrics. They're all real time. Um, adding in that data visualization piece now opens up all those metrics to be able to be on a single dashboard.
And there's no reason why you, you can't in intermix it with other products data, right? So now, now you have that ability to get your system level performance, your network performance, and what's going on with the mat performance metric all in one dashboard, all real time. Here's, here's what's going on.
So it, it's intermixing all that data together and giving you that added value. And all we're doing is just streaming the data and giving you that, uh, that database of information. So it's that incremental value we're adding on.
Uh, the Net Master product has got a lot of, uh, really nice suites for being able to trace things on the mainframe. It's Smart Trace feed functionality, being able to understand what's encrypted, what's not encrypted, what level of encryption are, are they on. So we, we saw this as, as a unique opportunity to kind of give a web interface into that level of information.
So there is a network to insights piece that we now have that builds upon all that net master data and gives you the ability to run a smart trace on the fly, right? So if somebody doesn't have to go into the mainframe, understand how to navigate through all of the 32 70 screens, they just gotta hit a button and it'll start a smart trace gives you quicker access to that. And then the open telemetry data streamings part where we can send all that information out to, to the SRE tools, that is something that we have customers doing.
Just that, that that was the piece that they, they wanted from Watchtower because they wanted that enablement across, across the application space. Okay. So what, just, just to reiterate, our primary focus here is integrating a whole bunch of data, integrating all the information that they have and making that as a better experience for these, these personas and building upon how we can use those current products that we have, our core products, the ones that our customers are relying on day in and day out for the last 30 years, and just give them better functionality.
You give them that modern use cases that, that they're looking at to kind of drive forward. Presumably you are existing the, the core products that you mentioned that Yeah. That accounts for probably a significant footprint of your entire client base anyway, so you Yes.
You can have those conversations with a, you know, right. Many, many customers. EE exactly.
And, and the way that we are delivering this to the customers, right, if they're licensed for any of those products, they have access to it. They, they, they don't have to come back and, and try and get another license with, with us, they just have it in already. So if they see something that is an interest, they can quickly move into that A POC stage, Sort of a more of a functional upgrade concept than Yes.
Yep. Do they lose any, um, functionality of the core points? 'cause I know a lot of times when teams are troubleshooting, one team has the, a lot of times we'll get more information from the, the native, um, tool.
So will by looking up, do, do they still have the core tools available to them? Yes. Yeah.
So we, we never envision that the, the SME, that data expert in performance analysts or the network person is ever going to just live in Watchtower, right? There's way too much in information, there's way too much value in those core products, right? They, whenever they get into these big problems, they're always gonna jump into sis view and do that performance analysis and understand what's going on.
But what we're trying to do is give them that starting point, right? Hey, here's, here's where you need to look now, drive down in, in, into those tools and get that A SME knowledge that you need.