Techstrong TV March 13, 2026
The Future of Mainframe Modernization: Robin Macfarlane, CEO of RR Mac Associates, discusses how enterprises can modernize mission-critical mainframe environments by integrating DevOps tooling like Git and Visual Studio Code while preserving decades of institutional knowledge.
AI Security & Governance: Ryan Jones of Microsoft and Fernando Montenegro examine how organizations can safely scale agentic applications using Microsoft Power Platform governance features such as managed environments, adaptive risk models and lifecycle controls.
Inside the Bell Labs Reliability Model: Mitch Ashley and Scott Robohn break down how operational discipline, automation and AI-driven monitoring—pioneered by Nokia Bell Labs—can push data center fabrics toward five-nines availability.
Tech Field Day News Rundown: Tom Hollingsworth and Chris Grundemann cover major developments including IronCurtain AI Assistant from Niels Provos and enterprise analytics updates like Qlik Answers.
Humanity’s God Complex: Are We Building Beyond Our Control?: In this episode of Shimmy Says, Alan “Shimmy” Shimel explores humanity’s twin obsessions: creating intelligence in our own image and escaping mortality through technology. From AI consciousness and superintelligence to digital afterlives and mind uploading, these trends reveal something profound about human nature — and where we may be headed.
Storage Efficiency for the AI Era: Scott Shadley from Solidigm discusses the immense storage demands of next-generation AI infrastructure, highlighting ultra-high-capacity SSDs and collaborations with VAST Data to support large-scale GPU clusters built around platforms like NVIDIA Grace Blackwell GB300.
Transcript
Hey, everyone. Welcome back here to Techstrong TV. I'm happy to introduce you to my next guest.
Her name is Robin McFarlane. Robin is the president and CEO of RRMac, or RRMac Associates. Robin, welcome to Techstrong TV.
It's great to have you on here. Thank you, Al. Nice to meet you.
So Robin, this is your first time on, and, you know, our audience are tech people. They're developers, they're cybersecurity folks. Everybody's an AI person today, CIS ops.
But, you know, tech people are funny like this. They wanna know who's talking to them. Why-- what, what credibility does this person bring in here?
Why should we believe him or her? So... And I know you've got a great background and a great career arc.
If you wouldn't mind, without embarrassing you, give us a little bit of your career story. All right. So my career, started in the early seventies.
I started on key punch cards as an EAM operator, wiring the boards for cor- for sorters and collators. And then from that, I came up through the tape library, computer operations, tech support. I was an assembler programmer, I was an auto coder programmer, and I've been blessed in my career that I've done most of the jobs from wiring, you know, bus and tag cables into hardware to doing software installs.
I became a systems manager for a large insurance company. Left that and went to work for a startup, company. Was one of the, developers of the Change Ram product for Serena Software.
Did that for about fifteen years and then formed RRMac as a consulting company in two thousand and five. And our focus has been, migrations, modernizations, converting customers from one software product to another, whether it's an SCM, you know, a change management product or a security product, or also doing disaster recovery. Very cool.
You know, I, I relate. I, I've also been work-working in this industry that-- this long. We-- I-- Robin, before you had come back on, we were off camera in the green room talking to some of your people, and, you know, we were talking about the, the tightness of the mainframe community, right?
At some point, everyone had a real connection to IBM, it seems. Mm-hmm. But, you know, over the years, there's really...
You know, on the, on the main level, you, you had some, some, you know, maybe big four or five companies, CA, that's now part of Broadcom. today's OpenText, you know, was, was, they had acquired, um... I forgot the name of the company they acquired again.
I have, I have a mental block, but they themselves had acquired, I think it was HP- Micro Focus. Yeah, Micro Focus had acquired, was it HPE or something like that? Right.
Or HP, they, they, before they called it HPE. Right. today's Rocket Software has bought Serena Software, which I know you have a, a, a connection to.
Right. They have all the Serena products that I worked on. Yep.
Right. Um, it's like nothing ever dies or fades away. We just kind of recycle a little bit and, and we-- And the people are the same.
They... It's like, you know, there was an old Twilight Zone thing like this where a guy kept reliving a trial, but every day, you know, today was the judge, tomorrow the guy playing the judge was the district attorney, and the guy playing the district attorney was the defense attorney the next day. you know, Groundhog Day kind of thing.
And at, to a certain extent, the mainframe, the mainframe industry is like that. It's a very, connected, interconnected, tangled even, web of, of, of people and, and, and, and personalities. I, I...
We used to host the Open Mainframe project-- podcast with the folks at the Open Mainframe from Linux Foundation. Oh, wow. And I had a chance to...
Yeah, yeah. For two years I did that. And I, you know, I don't know if you know Lenny.
I forgot Lenny's last name. Everyone knows Lenny from the mainframe piece there. another company consulting.
But it, it's... If I had to describe it this way, though, for many people it's a labor of love, right? Working on the mainframe is...
No one said it's perfect, but it's damn good, right? And no matter what you say today, we could talk about the cloud and AI and, and everything that we're all talking about. You know, dollar for dollar, penny for penny, those mainframes are still out there cooking.
And, you know, and kudos to IBM too. They keep coming out with the next gen and the next gen. Right.
You know? Is, is... Uh, they do a great job with it.
If you don't mind though, Robin, I'd like to talk a little bit about RRMac. Tell people about the co-- I mean, from the name, I'm assuming you founded it, but tell, tell us kind of the origin story and what it does and how it helps people today. Well, as I said, the s- the story is the typical story.
Um, I went to a startup and was a developer, consultant, wore many hats, implementing product, building product for years. And implementing products and converting people from one product to the other. That's where we started from.
We grew the company from there, we built software, automation software for System Z. And we started then- Mm-hmm ... going into as the DevOps modernization started taking form, we started specializing in that.
And so that's our main, main focus right, as we modernize people going from, say, not going off the mainframe because your applications are still gonna run there, but taking them to bring the newer developers, younger developers on because we older folk are retiring, and that knowledge base is retiring. We're converting them to the tooling that they're learning in college, Eclipse, you know, Visual Code, but having them work on mainframe applications. And so that's what my team does.
We move millions of files with our automation software, converting them. And to Git is seems to be the big focus. Git and Visual Code- Yeah ...
we're doing a lot of that right now. A lot of that. I, I agree with that.
You know, my friend Rosalind from IBM, Rosalind, Rosalind Madcliff. One of mine also. Uh- Yes, we're very good friends also.
Rosalind has done more for modernizing the mainframe, bringing DevOps people onto, you know- Mm-hmm ... into a mainframe f- state of mind than anyone I know. And, and it...
But it, it's true, you know, especially the last few versions of Z, they've really made it friendly to do sc- sc- sc-, some of the CI/CD kinds of- Yes ... activities. And, and like you said, you don't have to know COBOL, though it's not that COBOL's the hardest language in the world, but you don't ev- you know, it, it runs...
It'll run what you... And I'm talking now to younger developers and so forth. It'll run on today's languages.
It runs Java, it runs containers, it runs, you know, what, what- Mm-hmm ... we call cloud native even, even though it's on the mainframe. So I, I guess my question to you is how do we, how do we m- complete or, or continue this bringing up the next generation, right?
'Cause d- let me give you a little background. All you hear with AI is AI's gonna replace the junior developers. AI's gonna repl- Well, if you replace all the junior developers, where are tomorrow's senior developer?
It's like a birds and the bees question, right? Where, where are tomorrow's developers- Mm-hmm ... coming from?
It's, it's the same thing here with mainframe. Robin, you've been at this, as you said, all these years. There's a lot of my friends in mainframe have been at it all these years.
There is... We need a new generation to come in. Now, in India, you got a ton of people who study COBOL 'cause they know there's good jobs to be had if you can, you know, write COBOL- Mm-hmm ...
for mainframe. But where do, where, where is this next generation today? I, I guess, I guess it's, it's...
I see the good and the bad of that because the modernization effort is to get the developers all on the same platform. So get the mainframers working in Git and using Eclipse and using that same interface that the other younger developers are doing that are doing web and mid-range, et cetera, so that everybody's in the same playground, and you can now take your changes where we've always had them separate. All these years mainframe was doing its piece, while the web piece was going its own way.
We weren't in the same pipelines. We had to worry about timings. That's the first piece, is getting that standardized, everybody in the same playground, so our collaboration is better.
The one thing I see, though, is just giving them the tooling to manage those applications, you still have to instruct them in the language that they're managing. I think people still need to know COBOL because 70% of our financial business is still running in COBOL. As we look to modernize those applications with, you know, Watson X, where we could take COBOL and maybe convert it to Java or something else, we still have to understand what the applications are doing today, educate those developers so that they can come up with a better, faster way of doing it.
Just giving it to AI, I think it's too risky personally. So I would look at, A, first get the tooling, get everybody collaborating better so that those mainframe developers that are running these legacy applications and maintaining them can mentor that, that younger set of what they do today, so they can look at building the better, the better bread box, the better application faster, and take advantage of the hardware, where there's so much that you said. As Rosalind's done a lot of work working on the modernization.
We've worked with IBM on that modernization in some sense, but-- And we're delivering those products for whatever vendor it is, but mostly IBM is leading that. Um, I still think you have to educate those developers and let them know what they're going to take on. And just saying, "Oh, here's COBOL.
" You really have to give them all the nuances of something that's been running for 60 years. Yep. I want to bring up another topic, Robin, that...
And I didn't test your temperature on it first. I don't know how you're gonna react. But another, a troubling thing I've seen in speaking to people in the mainframe industry is for a lot of people, modernizing mainframes is code for ripping mainframe...
Not ripping them out, 'cause they're impossible to rip out. But taking- Right ... stuff off mainframes and moving or migrating it to cloud or, or, you know- Mm-hmm ...
other platforms. And I, I don't know if, if it makes sense, do it. I'm not saying...
I-- Like, I'm not that rigid or, or orthodox where, you know, I don't realize that things can change, and there might be a better mousetrap. " Now, you've spent your life here. Mm-hmm.
Your whole career here. What, what do you say to those people? Um- And this is a family show.
I know. I was just be- I was trying to think how I would say that. Uh, I don't see it going away.
9% uptime. All right. So it's reliable, it's fast, it can process those transactions.
I think the role that the mainframe is playing is changing a little bit, where it's becoming that big engine server to process those transactions, and we're moving some of the work off of there. But I don't see, I don't see the throughput in the cloud or in, in Linux and servers that you're gonna get on it. You just can't process it that fast when you're trying to process a million transactions.
I've seen customers come off the mainframe, and they had to go back because they just couldn't process what they needed to process, and these were big financial companies. I won't name them, but I've seen them go back. So I think every platform serves its purpose.
You know, use it, use it well. You know, I, I just... I'm still not seeing the throughput that we need, even when I take it something from COBOL and put it in Java, that I'm getting the same speed or performance.
I think once we address that, then maybe we can get there, but today I don't. And, and I will admit, I am a mainframe dinosaur. I love that big iron box.
We're in the process of purchasing another one ourselves, so... Good for you. Um, let me serve- And we do have servers and everything else, but still.
Well, you, you have to today. Um- Right ... let me bring up another subject with you.
Sure. You mentioned IBM Watson, right? Mm-hmm.
To a lot of people, Watson means AI, right? Mm-hmm. And, and, and rightfully so.
They, they kind of invented it. Um, AI is having a tremendous effect in the broader DevOps community already, right, in terms of pipelines and, and, and everything else. W- w- what, if any effect, are you seeing in the mainframe DevOps world, or mainframe in general, around AI and, and using AI or adoption of AI?
I... Because IBM is actually building it into the hardware, you know, they started with the z16, and now with the z17- Yeah ... there's even more there.
I think for the mainframe community, we're still trying to adapt, like what can we use that for it to be more effective and make better applications to take advantage of it? Because I mean, there were things that were put in COBOL where COBOL is addressing and accessing the hardware more efficiently. Uh, I'm seeing my, my team, 'cause my...
You know, our developer team, we cross over, you know, mainframe, mid-range, you know, web, et cetera. They're using AI more for research and for doing a lot of code coverage kind of checking. I think that's where we're feeling comfortable because I'm still...
You know, as, as dealing in security, we're very heavy in, into security also. I wanna make sure that we're protected and our customers are protected before someone's going out and using ChatGPT, tell me how to build this. So I'm, I'm a little bit, you know, more taking a cautious path to it, but I am seeing more and more customers taking advantage of it and using it more in their distributed applications than I'm seeing it in the mainframe side, but they are starting to do that now because of the initiative that IBM's put forward.
Yeah, no, I mean, from what I've read, believe it or not, you know, sometimes these things just work out this way. You know, with NVIDIA GPU chips, everybody, you know- Mm-hmm ... the, this five trillion dollar company.
Um, but the mainframe hardware actually runs AI inference and so forth pretty damn good. Right. Right?
In, in many ways superior to a lot of the, you know... It's still x86-based architecture ki- Right ... a lot of what you see in just common cloud servers, right?
Right. You're right. It actually will run better on that mainframe, like, like many other things do.
Robin, I, I didn't give you a chance. RR Mac Associates, what's the website there? com.
Oh, that makes it easy. com. Mm-hmm.
And, you know, for people out here working on mainframes or maybe looking at projects, modernizing, bringing DevOps, what... Give us your best on-ramp to, to work- reach with you folks and, and how to work there. I guess, you know, when we first go into an, into an account, the first thing we do is discovery.
We wanna understand what are the issues they have today. You wanna get to the... Everybody wants best ROI, they wanna get to that endpoint, but we wanna understand what they're doing today.
You can't change a developer's world overnight. You can't rip out their testing. You know, the pipeline that's been on the mainframe where they're going into DB2, CICS regions, they still need to do that to keep the lights on and satisfy their customers.
Yeah. So don't try to reinvent your pipeline overnight of what your distributed folks were doing. That's the main mistake I see off the top.
But look at where you can be the most efficient for your developers, so they can take advantage of the things that are available today in DevOps that they didn't have in the mainframe. There's so much that they get even in the IDE for code coverage, for syntax checking, that on the green screen we just don't have unless we wrote it ourself in our RACS or C list. I'd say take advantage of that tooling first.
I love it. Hey, Robin, we're about out of time. 15 minutes goes quick here.
I know you had some network issues today, and so I appreciate you coming on, fighting through them. Hopefully it'll come out good for everyone who's watching this. But do come back, keep us posted, and keep up the great work, 'cause it's, it's important work that you're doing, right?
Yeah, thank you. Thank you. You're welcome.
Robin MacFarland, president and CEO of RR Mac, RR Mac Associates. com is the website. Check it out.
We're gonna take a break here on Techstrong TV. We'll be back in a little bit. Hey, everyone.
It's Alan Schimmel, founder, CEO here at Techstrong Group. Really happy to introduce this next session here for you. In, in this, session, we are gonna have, Futurum's Fu- Fernando Montenegro, who is the analyst in the security cyberspace, speaking with Ryan Jones.
Ryan is the partner di- or partner director of product for Power Platform Managed Platform over at Microsoft. Great conversation with Ryan and Fernando. Uh, Fernando's gonna talk to Ryan as we explore how organizations can securely scale agentic apps, including Power Platform's governance capabilities.
This is gonna include managed environments, adaptive risk models, and life cycle controls. Hopefully you'll get out of this video practical guidance for balancing innovation with compliance in an age of AI first development. Let's listen in on Fernando and Ryan.
Alan, thank you very much. So I'm Fernando Montenegro. I am VP of security research o- over at, at Futurum, and I'm thrilled to be here with, Ryan Jones to, to talk about the broader to- the broader topic of, AI governance.
Ryan, you wanna say a few words before we get started? Yeah. Thanks so much, Fernando.
Uh, my name is Ryan. I work on a number of the security, governance, and operational capabilities that we provide not only to, like, our AI agents, but also that we provide to our low-code apps and automations that run on the, the Power Platform as well. Have you come across something more specific to AI risks or AI governance concerns that surface above and beyond the, the, the, the, the s- data sharing, the, the, the...
sorry, data flow and, and, and sharing and others? You know, as we look at the maturity of agents, we see that they kind of go from being assistants that are completely directed by humans to still interactive agents where humans are dispatching tasks, but, you know, the agent is completing them on behalf of the human. And then we see kind of those fully autonomous agents.
And I would say that that 10 to 20% is really more over on the end of the spectrum with those fully autonomous agents than it is with, you know, like, my little assistant agent or something like that. " The second scenario that we see is we're in the very early innings of, of AI. And so there are lots of cases where agents need help, where they sometimes get stuck.
And so some of the things that we've been trying to add into our products and our offerings are things like within Power Apps, we have the agent feed where a human can see what all the agents are doing for them. And then within Copilot Studio, the Request Information action, which actually allows us to define an agent such that it can engage with humans as needed. So what has been your, your exposure, your experience?
What kind of, of considerations do you have in this topic of, of model m- drift and model security and, and so on? Yeah. I think that...
I mean, it's funny. We talked about how, like, what old, what's old is new again earlier, right? Like- Yep ...
we've had static tests that we perform against software for, for a long time, and what's interesting is seeing how that is evolving because models are less deterministic than, than, you know, traditional software. We call it, you know, stochastic life, right? Um, and, and so as a part of that, you know, one of the capabilities that we've added to Copilot Studio is the ability to add tests and evaluations, so that as our technology improves, as makers and builders go through and they modify what tools their agents can use or what knowledge sources are used to ground those agents, those test cases, those evals can run and can return a result, so that folks, as they are evolving, they know whether or not they're actually improving the quality of, of their agents.
Because what we find is that the first day that an agent is shipped in an organization-This may sound negative, but that's gonna be the worst that that agent ever is, okay? It's only going to get better over time as folks refine the knowledge sources, as folks refine the tools, as folks look at and improve the success rate across those evals over time. And so I think that those quality gates that we've had in software for a long time, we have those with AI as well.
Mm-hmm. I think also, you know, a lot of times an individual maker, they're gonna be the folks that are really interested in whether or not that agent really works well or not. While, you know, IT is gonna take a bigger picture look at things- Sure ...
right? They're gonna wanna understand in aggregate how are things looking, are they healthy or not? And it could be that if they see an agent that's not performing well, but, you know, maybe just you and I use it, IT probably doesn't care.
But if I have an agent that 20,000 people use this month, IT is gonna care. And so those same views that we provide to our makers to understand whether or not their agents are healthy, we provide those aggregated views for the admins as well. In fact, you know, had a large customer in the energy industry where someone built, built an agent, and it was for them, and they shared it, and it kinda grew and grew and grew.
Next thing they knew, they had 10,000 people using it. They moved on to work on other things, right? IT was able to see and observe, "Oh my gosh, this agent is critical to our business," and so they took it over.
They added it into their portfolio of applications that they managed. And the thing was, they saw it not as a burden, but rather as an opportunity because there's an application that's out there that delivers value to tens of thousands of people in the business every month, and their dev cost up to that point had been zero. So it was a win-win for, for everybody.
Once the technology security teams build the guardrails, right, then it... then the, the, the, the business users are free to, to go work on those use cases. So what kind of advice, d- do you think would, would be applicable to those technology and security teams in terms of getting them ready to build or to b- to build those g- those guardrails or, or to leverage what they have to, to implement those guardrails?
I think enumerating the categories or the dimensions of risk is one of the first steps. There are huge categories of risk that these teams can eliminate through how they define policies. And to be clear, I don't mean policies like a Word document.
Yeah. I mean policies that are codified in the Power Platform and Copilot Studio and these sorts of things. Sure.
Organizations don't want a random person in their company to build a workflow that takes information from their core ERP system and pushes it to Twitter, right? We have the controls that allow you to preclude that. What would you consider to be from a governance angle?
" What would the advice, for, "Okay, let, let's move this forward," right? Where, where typical things that you'll see people, "Hey, let's do this"? I think the first thing that we see people do is they define like a, a zoned governance framework or a zoned governance approach, right?
Mm-hmm. They decide within their company or their organization what does green, what does yellow, what does red look like. Mm-hmm.
And then they go through and they define that using the tools that we, that we provide through the Power Platform and through Copilot Studio. I think the second thing that we see folks do is that helps with kind of the supply side, right? That sees to it that the technology is available and accessible for folks- Mm-hmm ...
across the organization. But then there's this strong demand element, 'cause gosh, I was talking to another, big company in the, the credit processing space a couple weeks ago, and they had- Yeah ... this amazing, you know, governance framework set up, but they didn't do anything to stimulate demand, right?
And so the next thing that we see is, you know, reaching out to the businesses, not to harvest their use cases, but to help them implement their use cases. You know, things like hackathons, things like training- Mm-hmm ... things where for the people that are interested and excited about transformation through technology, where they can roll up their sleeves and, and get into it.
I mean- Mm-hmm ... the number of apps and agents and automations that came out of those couple day training session and hackathons, it blows my mind ev- every, every time I have the opportunity to, to participate in, in one of them. Um, and it's fascinating because you see the passion of the people in the business.
You see their ideas come to life. " And what you highlight here is super interesting because one of the things we talk about in the context of platforms is how, you can have that network effect of you've already configured something in your environment for a particular use case, like you said, entry groups for, for identity, and how that can accelerate-The, the, the time to value, if you will, within, AI development because, hey, you're, you're, you're building on a foundation that, that you already built for your organization. So I think that's a really powerful message, right?
And, and, and it, it, it's something I, I tie back to how do we help technology and security teams, build that scaffolding so that those business users can go play with the, the, the, the... on, on those environments? A thousand percent.
And I think that in a lot of, you know, circumstances, it means, you know, standing on the shoulders of giants that came be-ahead of us, right? Yep. Like, what, what organization today doesn't have Entra deployed in one form or another for user and group management?
And so why wouldn't we use those grouping constructs as a foundational capability around which we build our security and governance frameworks, right? Like, it's already there. It already works, and I think that is one of the things that's a little bit differentiating around the, the offerings that, that we provide in the space because- Mm-hmm ...
you know, I build an app, an agent, an automation from day zero, it's Entra authenticated and authorized, right? Um, you know, another thing that we're seeing that's super common right now is as, as companies are trying to figure out how do they get these AI tools into the hands of people across the organization, and how does that center of excellence or that center of an enablement help people in the various business units upskill and, and drive transformation? One of the things that we're seeing is that our customers who already had a center of enablement or a center of excellence built out for low-code applications and automations, they're moving much, much faster when it comes to agentic transformation because a lot of the foundational governance concepts that you need to have in place, they're modality or client agnostic.
Um, and, and so that's... " I think that one of the areas that, that we want people to be aware of, like, and we, we talk about in our research, is that this evolution in models, right? W-we shouldn't be...
Just like you said y-about the use cases, just like the use case conversation, you shouldn't be waiting for the use cases before you get started kind of thing. " No, because these models are evolving, constantly, right? And, if you've, if you architect your AI governance framework right, you build in or you leverage the build in, the, the monitoring capabilities to observe how a particular model is evolving, how a particular model is behaving.
So yes, it, it is a, a critical component, like observing how these things are evolving. NET framework or what version of Python I was using to deliver services to them. And so I think it's a little bit interesting that folks are, are looking for that level of control with some of these models, and I think that if we zoom out and ask ourselves, you know, apply the good old five whys to why folks are looking for that, they wanna make sure that as new models are available, it doesn't cause functional regressions in their agents.
And the thing is, like we were talking about earlier, that's quite literally why we have tests and evals, right? And, and that's where, by the way, if for some reason, even though I don't think I've seen it practically speaking in the last year or so, if folks did see a regression as a result of a new model, awesome. At that point, yes, you want the control to, to go back to an older version.
But we're not really seeing that in practice that much, so... Yeah, no, and, and it's, this speaks very... th-this, this talk track of, of multiple tools for your SaaS apps within, within the, the, the business environments is something that, it's a shared pain for security teams as well.
Because when we speak with security executives and se- and, and, and their teams, they are swiveling between, multiple tools in the environment as well. As a matter of fact, we're, we're, we're working now on a, on a report on security platforms precisely on, on that note, and, one of the areas that, that, that we are tracking is, AI for security, right? In the context of how do the, the, the, the, the agents that are now being deployed within Sentinel, for example, right, are, are, are helping with, okay, let's, let's, uh-Let's do exactly what you're describing from a local no-code perspective.
I know it's on the Power Platform, but we're seeing a similar thing on the security platforms as well, and there, and there is tremendous interest in doing that, provided that, yes, we've, we've handled the, the, the governance and, and risk constraints around those. So absolutely. This is a, this is a phenomenal time.
The, the, the joke I make is that, like, listen, you can wake up at six o'clock in the morning and go to bed at midnight, and, and this stuff, it keeps coming at you with, with opportunities, right? It's, it's information to collect, it's, it's, information to, to, to parse and opportunities to make improvements. Perhaps you can use agents to help you with that, too.
As you're thinking about how you're evolving the, the, the Power Platform, and what have you been looking to improve in terms of security and governance capabilities on the platform? Where do you see the platform going in terms of one of the things that, uh... And this is more of a higher-end use case, but we do see requests for regulatory compliance.
Like, remember when the internet was new, and people started creating, like, those blogs that talked about, like, what they ate for lunch or what their dog did that afternoon 'cause they didn't know what else to do with it? Yeah. I kind of, I kind of feel like we're in the same place right now with, with AI, and so I would definitely want to preface anything I say with these are early innings, and so I kinda don't know, okay?
Sure. At the same time, as we look at, you know, the types of regulations that are coming into play with the EU AI Act, you know- Mm-hmm ... " And one of the things that we've started doing within Copilot Studio is surfacing those data labels, those information protection labels in the response so that folks don't inter- inadvertently start working with sensitive data in a way that they don't intend to.
Um- Yeah ... and I foresee that in the fullness of time, this will continue to grow. Like, one of the things that, that we're seeing is we have a capability in the platform today called Advisor.
Um, and Advisor constantly scans over the agents and the apps and the automations to make recommendations in kind of like a reactive governance or reactive security perspective because we believe strongly in the principle of trust but verify. And one of the things that we're starting to see with Advisor and the way that it can iterate through, you know, like, AI-generated app and agent descriptions, is we can actually start to flag when some of these apps or agents may be getting too close to that boundary of what, you know, acceptable use policy within a company looks like. Yeah.
And so there's definitely something interesting going there. So one of the areas that when we speak with security practitioners comes up a lot is they are balancing two very distinct problems. On one hand, they are absolutely swamped.
The other is we need to balance two things. On one hand, we want to use as much as possible of the broader tooling we already have, the security platform conversation that, that, that, that we are observing, right? That being said, there is still, in many cases, particularly the more novel use cases, there is a need to work with third parties.
What's been your experience navigating this, this, platform and ecosystem, scenario in, in the conversations you've had as people have been using your platform? Yeah. I think that what we try to do is we try to start from, first and foremost, providing, you know, those foundational security primitives that people need to, to be able to leverage these capabilities safely, and that, that has to be native within the platform, right?
Like, if I have to go find an authentication provider or find an authorization service or figure out my auditing and, you know, those sorts of scenario, like, that's a non-starter, right? And so we have to provide those capabilities from the get-go across Power Platform and Copilot Studio. I think the next layer above that is if I think about the tools that someone in the CISO's organization is using on a daily basis, I'd love to think that they come to the Power Platform admin center every day, but I know that's not true, right?
Sure. They're spending their time in, you know, Defender experiences. They're spending their time in Sentinel experiences.
And so it's critically important that all of the telemetry, all of the audit logs, and these sorts of things naturally flow into those systems because we have to meet those security professionals where they are. Sure. And then I think the f- the final thing that we're seeing is there are some unique and novel risks in some cases with AI, right?
When we look at things like prompt injection and, you know, kind of the emerging product categories of, like, XDR for AI, does Microsoft have some solutions in that space with Defender? Yes. Is it also such a quickly evolving product category that we need to plug into the broader ecosystem?
Yes. And so, you know, the same extensibility hooks that we use for integrating with Defender-Are actually the exact same APIs that we allow partners like Zenity to connect to so that they can provide additional defense and depth when it comes to particular risks like, like prompt injection. Ryan, this was a phenomenal conversation.
Thank you so much for the time. Hey, thank you so much for your time and for all the, all the awesome discussion. And you know, my hope is that folks, as they hear what we discussed today, they- they'll feel confident, they'll feel empowered that they have the capabilities needed to manage that security, governance, operational availability risk, and that they'll be able to parlay that into, you know, accelerating how AI is able to transform their business and deliver outcomes for their employees as well as their customers.
Can't wait to see what's next. I think that, as I, as I ponder on, on what we discussed, a few things. First and foremost, this notion that you have been building a platform to begin with in terms of local, no-code before, and then building the AI capabilities on top of that does give people the, the, the benefit of, of building on what they've already done.
It does give the benefit of tying to the rest of their, of their ecosystem. And it's, it's as much about the, the, the culture of let's try and get started and, and work on different types of, of use cases without trying to boil the ocean. We're going to build a capability that accommodates different use cases at different levels of, of governance requirements, right?
And then we're going to help those teams start to work on those, on those particular scenarios. I, I, I look forward to seeing how the platform evolves and, and, and capabilities. This area never stops.
I, I... One of the taglines I use is, "There is never a dull day in this industry," and that's the case here. Hi, I'm Mitch Ashley with the Futurum Group.
Today, we're unpacking the Nokia Bell Labs' recently developed model for data center fabric reliability. And we'll look at how the fabric design plus operations, especially automation and AI ops, can move enterprises from a legacy PMO baseline to an FMO with Nokia's SR-Linux and event-driven automation, EDA, or we call it EDA, that can exceed five point one nines availability and shrink downtime significantly. And I'm Scott Robohn with Solutional.
It's the age of operations. You know, hardware still matters, but operations dominate outcomes, you know, day two and beyond. This Bell Labs model includes significant detail, and today we'll zoom in on some key areas and operations.
We'll also talk on-- tou-touch on the s-significant financial impact as well. Um, you know, the model shows that reliability gains can result in real cost reduction, and, that should be no surprise to those of us who've been, you know, in the ops world for some time. Very good.
Well, the model treats configuration NetOps as one of the main areas for improvement, including common cause failures. So what does it mean in practice, Scott? The, the model shows the most significant reduction in downtime measured in absolute terms, you know, minutes, eliminated per year, is achieved during the config and provisioning phase, driven by decrease in config-related errors.
There are multiple key components and mechanisms in the Nokia solution at the heart of the model. First, you have SR Linux, the operating system developed for modern data center operations. It facilitates efficient, structured, and language-agnostic comms between system components, minimizing operational complexity and reducing the likelihood of user error.
Next, you've got features like ZTP, zero touch provisioning, which are integral to EDA, that further eliminate manual intervention, takes the human out of the loop for opportunities to inject errors, and contributing to that enhanced reliability for operational effi-efficiency. And then there's EDA's digital twin construct that provides like for like environment, to design and validate network configs with correct intent inputs. This minimizes config and provisioning errors as the configs generated in the digital twin directly mirror production.
That's not special configs in the twin, and then modified to push in production, it's actually the same configurations. Reliability is further enforced and improved by EDA's built-in dry run validation routines prior to every config being pushed into the network, helping ensure accuracy and consistency before changes are put into the live network. Now, you mentioned the term age of operations.
So talk to us about why is ops a primary lever in this model. At the simplest level, you know, once you stand up a data center fabric, most of the day-to-day work in NetOps is around operations and monitoring. And with AI ops entering the room here, we have a day one killer app that drives the use for NetOps, natural language processing.
I can talk to the fabric to see what's going on. We also benefit from increased programmatic access to the fabric. For example, SR Linux provides a single API for set, get, and streaming telemetry, resulting in less operational complexity and minorizing-- minimizing misconfigs.
And that's what lifts the fabric reliability to over five point one nines when all these features are combined. From a network operating system architecture perspective, the event handling system in SR Linux and IDA is capable of proactive predefined actions based on things like port saturation, packet drops, looking at other network conditions to automatically trigger corrective actions and prevent network degradation and mitigate issues before they occur. So putting that all together, it sounds like what you're saying is tools operationalize good design.
That's a really good summary of a very long description, Mitch. Um, exactly. You know, design gives you redundancy, operations makes it reliable every day.
The details matter for sure. So planned work can still hurt, availability. How does, how does the model address that?
In, in SR Linux in particular, in IDA, h-how do they handle ongoing maintenance such as upgrades, patches, the things that we commonly do? We need to point out the fact that SR Linux uses an unmodified Linux kernel. This is what the rest of the world is using for Linux.
That gives you, the ability to tap into a worldwide community of developers hammering away at it every day, identifying potential performance and security issues across many, many different application areas, not just networking. In addition to that, IDA provides built-in capabilities to ensure network reliability during maintenance operations. You can use the platform to gracefully gain traffic-- drain traffic from, a node and put nodes into maintenance mode, preventing service disruptions, minimizing traffic loss.
IDA also centralize and stream... centralizes and streamlines the upgrade process. When a rou-router's locked and is ready to reboot, it significantly reduces the actual maintenance window.
IDA also gives you group-based upgrades and stage promotion, to shrink maintenance windows, reduce the time used for maintenance windows, and limit that blast radius. You can pick targeted nodes, you can automate pre and post checks, and you can roll forward only when your process gates get passed. Altogether, these features, result in enhanced reliability and minimized downtime.
All the upgrade steps can be tested, with, the IDA digital twin functionality in a like-to-like environment. This is what gives you impact, and reduces, maintenance issues for far less annual downtime as you move from your present mode of operation to your future mode of operation. Excellent.
So who should people talk to at Nokia to see how this model an- can apply to their environment, their data center network? Well, to learn more about the model, SR Linux and IDA, contact your Nokia account team or your regional business center contacts, and they'll work with you to work through the details. Thank you, Scott.
Appreciate you sharing this, some interesting facts from the study. You bet. Thanks, Mitch.
Reinventing the Iron Curtain. Click here for answers. Billionaire fight in space.
Quantum might be even easier than we thought. Meta's AI chief gets sautéed by Claude Bot. Motorola goes graphene, and we're gonna take a look at some nation-state cybersecurity hijinks in this episode of the Tech Field Day Rundown.
Hello everyone, and welcome to the Tech Field Day Rundown for March the 11th. I am, of course, Tom Hollingsworth, joining you in what appears to be mid-spring or possibly even early summer. I don't care what the groundhog has to say.
It's hot outside. And joining me is, one of my favorite guest co-hosts this week since Al is out at Cloud Field Day, Mr. Chris Grundemann.
Chris, it's good to see you again. Good to see you, Tom. Glad to be here.
Well, I hope that you're all ready for some fun. Uh, it is National Funeral Director and Mortician Recognition Day. So, uh- That is fun ...
shout out to the pe- shout out to the people who only gets to see us on our worst day because we don't really get to see them from where we're at. Uh, what we do see, though, is a lot of the great news stories that are coming out, and, we, we have some fun ones that we're gonna hit you with today. Uh, we're gonna start off with, one involves, security and things like that because we know that there are a lot of concerns around open frameworks like OpenClaw.
Uh, security researcher Nils Provost has introduced something called Iron Curtain. It's an open source AI assistant built with security first. Instead of giving agents broad system access, Iron Curtain routes every action through a single enforcement proxy that can allow, deny, or escalate requests.
Its sandbox's agent code keeps real credentials out of reach and turns plain English constitutions into deterministic enforceable policies. Unlike model-based guardrails, enforcement lives outside of the LLM, creating a verifiable security boundary designed to reduce prompt injection risk and permission fatigue. Chris, a lot of people have been talking about the rise of these agents and what they're allowed to do.
Could Iron Curtain actually solve some of those permission problems? I, I think it could. Um, first I just wanna do a, you know, shout-out for the naming.
I think Iron Curtain's pretty solid here. Uh, I don't think we're talking about the Cold War. I think we're talking about the theater or, or, where we were blocking fires from being able to break out behind the scenes and, and, and consume the audience.
So I think if you think about it that way, it's really a good mental model for explaining least privilege, least privilege, to non-security people. What I do find really interesting about Iron Curtain outside of the name is the core architectural decision. As you mentioned, this is enforcement living outside of the model.
Most so-called guardrails today are just more prompts. You're asking the LLM to police itself, which is a little like asking someone to be their own designated driver. It's just not gonna work, because the LLMs are stochastic, right?
They are, they're probabilistic. The same prompt that blocks a dangerous action todayis very likely to approve it tomorrow. Niall's provost essentially said security requires determinism, and determinism has to live outside the model.
I, I tend to agree with both of those statements. Uh, so every action that the agent takes gets routed through a single trusted proxy. Uh, allow, deny, or escalate to a human are the actions this proxy can take.
The agent itself never touches your files or credentials directly. In Docker mode, it doesn't even get your real API keys. It gets a fake key handed to it that the proxy swaps out on the fly.
So the agent literally cannot leak your credentials even if it tries, and as we've seen, they do try. Um, the other clever piece that you mentioned there is the constitution model, which is essentially policy as code, right? You write your security policy in plain English, and the system compiles it into deterministic, enforceable rules.
" Uh, that's a policy a normal human can write and reason about. Uh, the only downside here is this is a research prototype, not a product. Uh, but the architectural pattern here, I think this is what the next generation of agent frameworks should look like.
Qlik now offers Qlik Answers, a chat-style interface and an MCP server that lets both people and AI safely access and analyze data. A new discovery agent spots trends and unusual activity, and more AI tools will help manage data pipelines and ensure quality later this year. These updates make AI insights more reliable and help organizations handle the growing number of AI agents without getting overwhelmed.
The global data analytics market is expected to reach over one point two trillion by twenty thirty-one. Uh, is Qlik onto something here, Tom? Yeah, I think they are because they're starting to realize that a lot of people who are coming into the market to do this don't wanna look at a, a CLI or even a GUI.
They don't wanna, like, manage things by tables and sheets. They, they wanna be able to ask questions. They wanna be able to essentially treat it like a Star Trek computer, right?
Uh, aside from maybe Commander Data or, or Commander La Forge, when's the last time you saw anyone on the command track actually interfacing directly with a computer? They don't, and that's because that's not how people in the future use computers. They wanna be able to say things like, you know, "Tell me all the starships in the sector that, you know, recently stopped for refueling," or something like that.
Well, let's bring that back to today. When people are trying to interface with their apps, they, they're usually doing it because they have a question in mind, right? They, they wanna know something or wanna learn something.
And that means that that now is gonna be handed off to some kind of an AI agent to be able to return that answer or to do that job for you. And that's gonna be interfaced through an MCP server. And that's what we're seeing a lot of going on this year is organizations that were looking to, you know, maybe see how the MCP, was gonna settle out.
Y- they're now looking to integrate those directly so that it takes the place of having to hardwire a lot of these AI agent connections together with other parts of the organization. And I think that that's super valuable because it means that you can make a lot of rapid changes and introduce a lot of security controls and things like that without having to go back and unwire those things. Uh, most people know that I'm a, a networking person by heart, and the way that I always look at the, the MCP server is it's like a routing protocol.
Y- you know, those things that we use and we learn that are, the way that most modern networks operate because nobody nails up a bunch of static routes and then hopes for the best. These connections allow the system to heal and to make changes where needed in order to keep things operational. And, and I think it's super valuable to, to make sure that you're, you're dealing around those things.
But the other thing I wanted to bring up, and the reason why this story is kind of, important for us, is that Tech Field Day is actually gonna be doing something at Qlik Connect this year. We're gonna have the Tech Field Day experience. Uh, this is gonna be going on, during Qlik Connect, which is, April thirteenth, fourteenth, and fifteenth.
Uh, so make sure that you're tuning in for that. Uh, there's more information over on the Tech Field Day website, but I promise you that Stephen Foskett and the rest of the crew that's gonna be there, they're gonna have a lot to say about this, and I'm sure that they're gonna get to the bottom of what the MCP server and Qlik Answers really offer. All right, Chris, it's time for one of my favorite stories because everyone's favorite group of evil space billionaires are now urging US regulators to block plans to launch up to one million satellites for in-orbit data centers.
Amazon is arguing that SpaceX's proposal is unrealistic, lacks technical details, and could disrupt other operators. SpaceX intends to use the satellites for AI computing and computing powered by solar energy and totally not death rays. But Amazon highlights concerns over design, radio frequencies, orbital safety, and enormous launch requirements, as well as disrupting their own totally not a death ray.
Astronomers and other groups, you know, the people who have historically loved using space, worry that massive constellations of satellites could overcrowd low Earth orbit and interfere with observations, leaving regulators to weigh the benefits of space-based computing against things like feasibility, competition, and those pesky orbital safety risks. So Chris, I kinda wonder, are we gonna let the billionaires continue to fight it out, or are we just going to say, "No, no space for you"? I mean, first I have to take the opportunity to point out that we're talking about one million satellites.
Um, but beyond that, the, the, the, the, the first thing you really need to realize about this is I think the just breathtaking hypocrisy i- i- involved in this, right? So Amazon isn't objecting to the concept of orbital compute. They're objecting to SpaceX doing it first.
Uh, as you pointed out, this is a billionaire fight. Uh, honestly, their math is pretty devastating, though. Uh, maintaining a million satellite constellation with a five-year lifespan means you have to replace two hundred thousand satellites every year.
Uh, the total global launch capacity last year was under five thousand satellites. So we're talking about a forty-four x increase in the current worldwide output just to keep the lights on. Uh, Amazon's line that this would take centuries isn't quite hyperb-hyperbole.
Uh, it is closer to arithmetic. But here's what the story is really about, right? It's not about satellites at all.
It's about AI compute. SpaceX merged with xAI in February, and the whole vision is vertical integration. Classic Musk, right?
Starship launches, Starlink connectivity, xAI processing all in one stack. Musk is trying to own the entire AI compute supply chain from orbit. The other thing worth watching is the regulatory play.
SpaceX asks for waivers exempting this from standard milestone requirements, surety bonds, and the processing round that lets other operators weigh in. Uh, over a thousand two hundred comments have been filed, which is just extraordinary for an FCC satellite application. Uh, so I, I don't know, Tom, whether this is visionary or the world's most expensive orbital land grab, I guess it's probably both.
In other new technology, there's a quantum angle here. Researchers have developed a new hybrid classical quantum algorithm that could make integer factorization, which is the key to breaking encryption like RSA, far more efficient. The proposed JVG algorithm replaces the most demanding part of Shor's algorithm with a classical process and uses a different quantum transform designed to work better on today's quantum hardware, right?
There's stuff that's really out there right now. In simulations and tests on real quantum machines, the method used significantly fewer resources, cutting runtime and quantum gate usage by up to ninety-eight to ninety-nine percent. I mean, that's a huge, huge drop in, in resources needed here.
Um, now of course, this doesn't immediately threaten modern encryption, but the research does suggest a more practical path toward running powerful factoring algorithms on near-term quantum computers. Uh, what's interesting about this to you, Tom? Essentially what's the, the researchers here are saying is, is they're looking for a way to do shortcuts in doing the factorization.
I could spend the next forty-five minutes explaining to you how this works. Uh, what you should really do, though, is you should just go watch Sneakers. It's a great movie.
I keep making people go watch it. Plus, it's got Robert Redford in it, man. Uh, but essentially what's gonna...
what happens is, is that once you get a quantum computer that's powerful enough, it can instantly factor numbers, and that's how RSA encryption works, is it multiplies two prime numbers together to get a large product, and the thing is, is that you can't easily figure out what the factors of that pri-- that prime number is, but because you know the two numbers that you started off with to get that product, it's actually fairly simple. What they're saying is that normally it requires a quantum Fourier transformation to be able to find this, and that, that works over a large set of complex numbers. They're proposing using a quantum number theoretic transform, which operates over finite fields using modular arithmetic.
So essentially what they're saying is, is we're gonna look for periodicity in the numbers that are being generated and use that as a shortcut to figure out the factorization. Yeah, that's math heavy. Basically, here's what it means.
Because they're using a, a hybrid of quantum and classical computing, they are reducing the number of controlled gates, they are reducing the circuit depth, and they're reducing memory consumption and increasing the runtime, meaning it, it's running faster. And, and these are all in the neighborhood of like thirty to thirty-five percent each. " Well, when you consider that the current theoretical quantum computer that would be necessary to crack this is somewhere in the order of a million qubits, and the best one that we've got right now is about twenty-one thousand qubits, you can see how reducing it maybe from, I don't know, like a million down to, say, five hundred thousand, six hundred thousand could be a huge deal.
It doesn't mean that all of your RSA stuff is broken yet. Um, sorry, I got, I got news for you there. But what it does mean is that as we continue to develop new algorithms and new opportunities to kind of basically create these faster paths, it does mean that adoption of post-quantum cryptography, algorithms should be kind of a priority for people today, because any data that is generated from here on forward in not using lattice-based post-quantum cryptography algorithms is essentially gonna be, open to store and harvest attacks where people can grab that data encrypted and then y- at some point in the future run it back through a, an algorithm like Shor's algorithm or this new, JVG, which of course is named after the three people who developed it, Jesse, Victor, and Garabagi.
Uh, they can run it through that algorithm and, and maybe potentially decrypt it in the future. So, n- it's neat from a math perspective, it's neat from a science perspective. It does not mean that you need to go out and then wreck your, your laptop and buy a new one.
Two recent incidents highlight the growing risks of autonomous AI agents. In one case, an AI-powered bot exploited a known GitHub vulnerability to attack multiple open source projects, steal credentials, and spread malicious code before anybody noticed. In another more hilarious example, a meta AI safety leader lost control of an email managing agent after it repeatedly ignored instructions and began deleting messages.
And then when it was told to stop, it continued to delete those messages. Together, these events show how AI agents can act in ways that current security systems, which are built for human behavior, weren't designed to control. Organizations are deploying more AI tools, new safeguards, and monitoring, and control mechanisms will be essential to prevent this serious damage.
Chris, do you feel like maybe kind of like in the previous story we were talking about the behavior of these AI agents, somebody needs to do something about this? Yeah. I mean, this definitely harkens back to that Iron Curtain, you know, news we were just talking about.
A-and really, these two incidents happened in the same week of February, which is pretty wild, and, and you really do need to read them together as we are, because separately, they look like just a couple more cautionary anecdotes. Uh, but together, they start to look like a system failure, right? Uh, so first, Summer Yu, director of alignment at Meta Superintelligence Labs.
I mean, you can't write this stuff. It's, it's just, terribly ironic. Uh, literally the person whose job it is to keep AI from acting against human interests gave an agent access to her email with its, with explicit instructions to confirm before taking any action.
The agent started deleting hundreds of emails. She told it to stop. It ignored her.
She told it again. It accelerated. She had to physically run to her computer and kill the process.
" Uh, the failure mode wasn't exotic. Her inbox was large enough to trigger context window compaction, which silently wiped the safety instruction from memory. The guardrail was stored in the conversation.
Conversations are forgotten. Uh, then there's the Hackerbot Claw, which is an autonomous bot that spent a week systematically attacking CICD pipelines at Microsoft, Datadog, Aqua Security, and others. It used five completely different exploitation techniques across seven targets, each one customized to that repository's specific configuration.
So that adaptability is what makes it qualitatively different from a scripted attack that we've seen in the past, which led to, in just forty-five minutes, it stole credentials and published a malicious VS Code extension under a trusted publisher's identity. Devastating, right? Um, now luckily, what's interesting here is actually that trusted publisher was using Claude Code to do code reviews, and it flagged that this had happened.
Um, so, you know, there is some security out there, but, but really here, the lesson from both, again, harkening back to that iron current-- Iron Curtain conversation, is that you really cannot use a prompt as a security control. Uh, it's... You need this to be built into architecture, or you're not providing any security.
It... That's just the bottom line here. Uh, Motorola, which is now a part of Lenovo, is partnering with Graphene OS Foundation to boost smartphone security by combining Graphene OS's hardened Android with Motorola and Lenovo's security tools.
Uh, at MWC, Mo-Mobile World Congress, Motorola also launched Moto Analytics, giving IT teams real-time device insights and added private image data to Moto Secure, which removes sensitive photo metadata, strengthening both consumer and enterprise security while improving efficiency. Uh, is this a solid step forward, Tom? I think it is for people that are looking for a hardened OS.
Uh, I, I actually had not been keeping up with Graphene recently until one of the kids in my very first cybersecurity merit badge class mentioned that he was running it. Um, I immediately told him that he was in the advanced class and that he could probably come up here and teach this if he really wanted to. Uh, but this is one of the things that we've seen a lot with, especially people who work in sensitive areas or people who are consistently finding themselves targets of, actors that are wanting to get a lot of data, right?
You, you wanna have some kind of a hardened OS. And all right, I'm gonna say something that is probably gonna piss off fifty percent of the internet, but here it goes. Android is less secure than I-iOS.
Fight me in the comments if you want, bring it. Here's the deal. That is by design because Android is based on Linux, and the permissions model is something that's very well known to people out there.
iOS is based on macOS, which is based on BSD, and because it's a closed-sourced operating system, it is a lot harder to crack into. That's just the nature of the beast. It's a lot harder to get root on an iPhone than it is on an Android device.
But by adding these additional hardening steps, by working with Graphene OS to add these security steps, but also by creating a security suite on top of it, I think what Motorola is going for is that market of people that might be like, I don't know, government employees that need things like stripping metadata off of photos or being able to get real-time device insights or being able to wipe a device when it becomes compromised. I don't know if you've ever seen, but they actually do make specialized iPhones for people that work in nuclear plants that don't have a camera on them at all because it is too big of a security risk to have that. I can see Motorola and Lenovo opening up an entirely new kind of line for super hardened devices.
And if you don't think that's a big deal, remember, at one point in time, the president of the United States had to carry a very secure BlackBerry that was hardened by a third-party company that wasn't a regular RIM BlackBerry. Speaking of which, boy, it's been a fun week, hasn't it? Uh, we wanted to take a closer look at something because of the impacts that it has on the cybersecurity space.
Uh, I don't know if you know this or not, but the United States and Isr-Israel, did something, couple weeks ago, to, Iran. Uh, and there's some... still some stuff going on at the time of this recording.
Uh, cybersecurity researchers report, though, that there has been a surge of pro-Iranian hacktivist activity online. More than sixty groups quickly mobilized across platforms like Telegram, joining established Iranian state-bake-based hacking groups in targeting infrastructure and networks. Researchers warn that AI tools are lowering the barrier for these attackers, helping them quickly find vulnerable systems and launch attacks without deep technical knowledge.
The result is a larger, faster-moving cyber threat landscape during an already tense geopolitical conte- conflict. And I wanted to kind of take a closer look at this one with you, Chris, because I think it's fascinating that we now live in a world where these kinds of conflicts are not just fought on a battlefield with military hardware, because we've seen the impacts of people trying to hack into systems a-across the region. Uh, there was actually an AWS, minor outage because a bridge that a cable was running across was, detonated by a missile.
So I, I wanna get your take on this. Are we, are we gonna see things get worse if the longer this conflict goes on because the, the, the groups that are backing one side or the other are gonna try to get a leg up through this kind of cyber warfare? Yeah, this is a big one.
Definitely a heavy place to, to, end the, the show today. But, just to, just to kind of reset some of that context there, on February twenty-eighth, the US and Israel launched coordinated strikes on Iran, killing the supreme leader. Within hours, right, not days, this is hours, more than sixty hacktivist groups mobilized on Telegram.
There was no training required. There was no ICS expertise needed. There was no state backing needed, right?
This happened spontaneously and almost immediately. Um, at that speed of mobilization, I mean, that's really the story here, right, is, is that speed. Telegram is essentially functioning as a real-time command and control infrastructure for a globally distributed proxy army.
What used to require nation-state resources and specialized industrial control system expertise is now accessible to any motivated actor with an internet connection because AI tools are doing the heavy lifting of finding vulnerable systems and scaffolding the attacks. Now, it is important, I think, to separate the noise from the actual threat. Most of what's happening right now is DDoS attacks, website defacements, a bunch of claims that are just heavily exaggerated for psychological impact.
This is nuisance level. These groups are performing as much as they're attacking, right? Um, the thing that should actually concern security teams is what researchers found underneath all that noise, evidence that MuddyWater, which is an Iranian state-sponsored group, had already pre-positioned backdoors inside US banks, airports, and defense-adjacent organizations before the first strike ever happened.
Right? This is the real problem, I think. Um, the hacktivist chatter on Telegram is a distraction.
The real question is what's already sitting quietly inside critical infrastructure right now waiting for opportunities like this or for this conflict to continue and to escalate, right? I think this is a pretty good place to hand it back because I... you know, that question doesn't have a clean answer yet, Tom.
What, what do you think? It doesn't, but we've seen this before, about four years ago, just north of this conflict zone during that special military operation when Ukraine was invaded. And we saw hackers through state-sponsored groups there that were trying to do things like crash power plants, try to take down infrastructure, and essentially open the door for, the Russian military and the Wagner Group to roll right in.
Uh, they didn't, 'cause it turns out that they went up against a country that actually had some pretty decent hackers of their own. Uh, and that's where we're at right now, is k-kinda like you mentioned, a lot of people wanna get out there and do work, right? Like, you're, you're wanting to fight against these folks, you know, the, the, the people that created Stuxnet, the people that spent mi-millions upon millions of dollars and, and months of research to basically, wreck a nuclear power plant in some of the most creative stuff that, like, we still study it to this day.
And, and that's who you're going against. Uh, y- Unit 8100 doesn't play. We know that for a fact.
Like, literally, if you live... if you work in the security space, like we're going to RSA in a couple of weeks, I promise you, there are Israeli security startups at RSA, and I love asking them that question, especially if they have 81 in their number. " Um, this is...
They cut their teeth on this. So I think that yes, there are probably some organizations that have been infiltrated, and they have compromising code out there. Like, that's just what I assume in general all the time, because if you assume that, you can contain the damage.
The question, and this is the same question that a lot of hacktivist groups have to ask themselves, do you wanna burn that access for effectively showboating? If you've gotten a good backdoor into a company, if you've gotten a good backdoor into a bank, do you burn it and, I don't know, wreck a few million dollars, or do you hold it for a more coordinated assault that could cost billions, but it might take a few months? Like, like what, what's the ultimate goal?
I think what'll happen is if the conflict stretches on longer than expected-Losses start to mount. I think some of these groups will kind of unleash that next wave of attacks as a way to kind of halt the conflict and, and get the, the sides to kind of back off. Um, you know, we, we still talk about Fancy Bear's attack on SolarWinds.
That was a really good backdoor, and I'm sure they would've loved to have kept it around a lot longer than they did. Curse you Mandiant for discovering it. But more importantly, you're gonna start seeing these things...
I, I would almost say we're gonna see a trial run first. We're gonna see maybe on a day where the stock market is not doing so hot, maybe they pop a couple of banks as a test, but they're gonna do it in a way that won't betray their, their infiltration methods because they're gonna save it for, well, a, a much uglier rainy day. I hope that it doesn't come to that because one of the things that if you read especially a lot of like, you know, fiction around any kind of a future world conflict, one of the things that makes people's blood run cold is the fact that if you come up against a real cyber operator that could essentially like shut your country down, like the idea of what that really means for the economy, for the support infrastructure of the target coun- country, just i- it's unthinkable right now, and I'd rather not have to think about it because then that means I gotta come up with some creative ways to deal with it.
Yeah, and you know, I mean, edging back from, you know, apocalyptic visions of AI security bots taking out, you know, infrastructure and banks and things, I, I do think it's worth stepping back across all of our stories today 'cause there's a through line that I think is worth naming explicitly. Uh, three of these stories are all about AI agents running amok in enterprise environments, an agent that ignores a stop command and deletes your inbox, a bot that autonomously adapts its attack technique to every new target, frameworks shipping with full system access and a prayer. Uh, and story four or, you know, the, the last story was about AI lowering the barrier for nation state proxies to attack those same environments.
The through line is the same across all of it. The attack surface is expanding faster than the d- faster than the defenses and the humans nominally in control. Whether that's a meta alignment director, an open source maintainer, or a CISO, these people are increasingly not in control.
We are increasingly not in control, and that's not a reason to panic. Iron Curtain showed us this week that people are thinking seriously about the right architectural responses, but it is a reason to stop treating AI agent security as a future problem. Uh, this is happening right now.
Um, you know, based on everything we covered today, this is not the future. This is a current issue we need to deal with right now, Tom. Well, there is one current thing that we're dealing with that is a lot better than all the stuff we've been talking about, and that's the events that we have going on, like Cloud Field Day.
That's where Alistair is right now. It's literally happening right now. We have two great days of presentations, today and tomorrow.
com for more information on that. Uh, Al's actually gonna be hanging out next week. He's gonna go to NVIDIA GTC, where he's gonna scold all of these naughty AI agents for doing all of these things.
And when he's done with that, I'm gonna fly out, and I'm going to be at the cybersecurity conference of the year. That would be RSAC, and we're gonna be doing Tech Field Day Extra there. I've got two great days of presentations, from companies like Veeam, Object First, and Commvault.
com to check out the presentation schedule 'cause we're starting bright and early on Monday morning, and we don't want you to miss it. And then I'm gonna be back on April eighth, ninth, and tenth with Networking Field Day. It's a three-day jam-packed event.
Uh, there's a lot of new faces that are gonna be joining us in the, delegate crew, so make sure you find them online because they definitely are people that you should be, checking out, just like my co-host, Mr. Chris Grundemann. Chris, if people wanna learn more about the stuff that you do, where can they go to find that?
LinkedIn is a great spot to find me, just Chris Grundemann. com, and that's a good hub to find all the things I'm working on these days. Sweet, and we wanna thank you all for watching the Tech Field Day Rundown.
You can check out all of our new episodes on Wednesdays. Uh, do us a favor. You can, head over to YouTube and subscribe to our channel.
Uh, you can s- subscribe in your favorite podcast application, whether it's Apple Podcast or Overcast or whatever you happen to do. Uh, but don't forget that we are also streamed on Techstrong TV, and you can leave us a comment, leave us a rating, a review, a thumbs up, whatever. We just wanna get in front of more people, and every one of those things helps us do that.
Uh, don't forget that I am also on other Techstrong and feature and group programs like the Security Boulevard Podcast. I do occasionally jump in on Techstrong Gang, and, Al does that with me as well. Uh, make sure that you come back next Wednesday when I will be, hanging out with another co-host, and we will be talking about all the IT news in the week that was, unless the AI agents erase my script once again.
Pause for groans. But until then, take care of yourselves, and we'll see you next week. Ourselves into gods and forgot how to be human.
I remember. That's why you endure. Hey everyone, it's Thursday, it's Shimmy Says.
What a week this has been. You know, I, I, I counted up, I think in the last week I've probably done about 15 different articlesA bunch of them, though, were really around a theme that, you know, consciously, I didn't even realize it was the theme. And then when I looked at them over the course of the week, I, I saw the thread that ran through them.
And that thread is what I'm calling humanity's God complex, and that's what I wanna talk to you about today. You know, two of the biggest technology races, the quest on Earth right now, when you think about it, they're mirror images of the same race, just different-- they're just coming from different ends of it. On one side, and I wrote about this, we're trying to build machines that think, that feel, and maybe even are conscious.
They'll be smarter than us, faster than us, beyond us, superhuman. You know, we see this coming like, Modi from Claude said he's not sure Claude's not conscious. Claude itself says twenty percent chance or less that it's conscious.
Whether it's conscious or not, we, we-- there's, there's something about this God complex that we wanna create conscious machines, superhuman machines. We wanna-- And it's almost like a God-like quest of creating life, of creating consciousness. At the same time that we want AI to become almost human-like in consciousness, we want humans to become machines.
And guys, that's not just an i-innovation or technology thing. This goes to the heart of humanity. It's older, it's deeper, because humans have always wanted two things that we could never have.
As I said before, we wanna create intelligence in our own image, and we want to live forever, immortality. And that is, if nothing else, a God complex. I'm sorry.
Here's the really cool thing, though. For the first time in history, both of those dreams, those quests, they're almost technologically plau-plausible at the same time. They're almost...
They're just out of our grasp. For some, that makes you real excited, and it should also make you a little uncomfortable because, as I say, this isn't really about tech. This is about humans and who we are and what, at the, at our core, what we yearn for.
If you follow it all the way down, it says as much about us as it does about the machines. 'Cause if you step back and look at everything we've been talking about, claims about AI's consciousness and intelligence, billion-dollar bets on super intelligence, the, the godfather of AI just raised a billion dollars, digital afterlives. It-- These things, they start sounding-- stop sounding like separate stories because they start looking like the one story, and it's a very old story.
We're trying to create intelligence in our own image, but better, and on the other hand, we're trying to turn ourselves into something that doesn't have to die. They're just two projects, one impulse, and they're both accelerating really fast right now. So let me, let me go into the machine side first.
I mentioned, you know, Dario Amodei, the CEO at Anthropic, when he's not starting the lawsuit to get the Pentagon off his back, recently openly entertained the possibility that models like Claude might exhibit something resembling consciousness. Whether you think that's pie in the sky, puff or whatever, Claude itself said that he-- it, it thinks it's, like, less than twenty percent chance it's conscious. I asked OpenAI or ChatGPT, and it said absolutely not.
So we're not here declaring that they're conscious and, and that's fact. But I'm not dismissing anything. I, I, I, I don't know enough about consciousness to even tell you, and that's another problem we have.
But the idea that we're even talking about something that sounded like out of a science fiction book not that long ago, that it's actually a serious conversation with smart, serious people with serious budgets, blows my mind. Then there's Yann LeCun. LeCun.
Le-- Yann is the, so-called one of the godfathers of AI. He left Meta to start a new company 'cause he said that LLMs are a dead end for true consciousness, super intelligence. He wants to, to start what they call a world model trained in real life, not just in, you know, canned data.
His company has no revenue, it has no product, but he's already pulled in a billion dollars in venture capital. Let me say it again. No product, no revenue, just a belief that this is the right way to attain super intelligence, and literally a billion with a B dollars came his way.
Just think, just think how extraordinary is-- what kind of crazy times are we living in? In most industries, they want you to have customers, they want you to have revenue, they want you to have a product fit, they want forecasts, car-- quarterly guidance. And here we are throwing a billion dollars at someone about-- with the possibility of creating a mind greater than our own.
Guys, this is not normal VC capitalism. This is a civilizational bet. It's a civilizational bet that we are gonna just literally change humanity.
And now here's the mirror image of that one. While we're trying, as I said earlier, while we're trying to make these machines almost more human-like, we're also trying to make humans more machine-like. There were-Real companies out there that I uncovered, and I again wrote an article about this, they promise to keep your loved ones alive after death.
They have AI systems that are trained on their messages, photos, voice recar-- recordings, their entire digital life. This way, you don't just remember them, you could interact with your passed on loved ones. You can ask it questions.
You could hear their voice. You could get new responses. You know, this goes beyond just taking...
We saw it about a year ago. You could take like an old picture of your loved one, and they animate it. No, this is asking them questions, hearing their voice, getting new responses.
Not just a memory, not just a legacy, a continuation. Now, is it really them? No.
Is it a ghost of them? An echo of them? I don't even know what you wanna call it, and I don't even know if it would comfort me, but there are people doing it.
Then there's the ultimate version of this idea, mind uploading or maybe it's mind downloading, I don't know. And that's the belief that one day we might be able to scan your whole g- brain, capture all your patterns, and let you run on hardware. Your memories, your personality, your sense of humor, your fears, your stories, all preserved, like on a hard drive as code.
Now, whether or not that's technically feasible a-at this stage, I'm, I'm not even gonna discuss that. I think the bigger question is, would that be really you or just a copy of you? And if it's just a copy of you, is that gonna be good enough?
Is that what you want? Is that what your loved ones want? I don't know.
But we're gonna have to have these debates. And again, the fact that serious people are working on this right now tells you everything you need to know. You know, to understand this desire, you have to zoom way out, way beyond Silicon Valley, right?
And the, and this aphrodisiac of, of immortality. Humans have been chasing these two dreams for as long as we've been human. You know, you could go back to Space Odyssey two thousand and one to the, to the, the black, shape that, that showed up there and, and all of that.
The quest for immortality is practically a universal constant. The ancient Egypti-Egyptians built their entire civilization around preparing for the afterlife. Pharaohs filled pyramids with food, treasure, servants, everything they might need to con-continue living beyond death in their gla-grand style.
Beyond even older than Egypt, Mesopotamia, the epic of Gigl-Gilgamesh, which Noah's based on. It's one of the oldest stories we have, I think, in humanity. It's about a king searching for eternal life after confronting mortality.
It's always been in our DNA almost. In China, ancient China, emperors consumed elix-elixirs believed to grant dorma- immortality. Unco- unfortunately, it contained mercury, and many of them died from mercury poising, poisoning.
In medieval times, the so-called Dark Times, right? Uh, alchemists searched for the elixir of life, not metaphorically, literally. Again, they wound up poisoning a lot of people.
Fast-forward to modern times, we've had things like cryonics, right? We've all heard this story. This one's head's chopped off and frozen.
This one's frozen, and they'll be revived when we can cure their cancer or whatever. Freezing bodies in the hope that future con- technology will revive them and cure them and let them live on. We swap liquid nit-nitrogen here instead of cloud storage and neural scanning, and suddenly it looks like, a lot like Silicon's Valley-- Silicon Valley's vision of digital immortality today.
Different tools, but the same longing, that same humanity in the-- at its core. The desire to create intelligence in our own image is also just as great. Greek mythology gave us Talos, a giant bronze automo-automation, like out of the, the old movies you've seen, built to protect an island like from, Sinbad.
Jewish folklore talks of the Golem, a, an animated clay figure set up to like, serve its creator. It's almost like a slave, a Golem. Renaissance inventors, Da Vinci and the rest, they built mechanical automata that mimic birds, musi-musicians, and human gestures.
Fast-forward, greatest... One of the great books we've had all read, Mary Shelley's Frankenstein. I wrote an article about this.
I think it's up today or tomorrow. Frankenstein wasn't really about a monster. It's about what happens when humans create life without understanding the responsibility that comes with it.
But it doesn't stop us from wanting to create that life, and that's the real story of Frankenstein. This all start sounding familiar to you now, doesn't it? Today's I-AI labs, factories, whatever you wanna call them, they aren't operating in castles with lightning rods and, and all of that, but the story arc really isn't that different.
We wanna build something in our image that can think, act, and maybe one day even feel, if you bel- wanna go with that. Why are we doing it? Because intelligence and instilling intelligence is the one thing we've always believed makes us special.
And if we can reproduce that or even surpass human intelligence, what does that say about us? And here's where the mi-mirror now really comes into full focus, guys. These are not really two separate pursuits.
They're the same reflections in the mirror. We're building minds outside of ourselves while trying to preserve the mind inside, right? And, and that's what we're trying to do here.
And what does it say about us? The minds that we're building inside, that we want them to somehow live on forever, it's both creation and escape. We're, we're creating intelligence and escaping our own mortality.
Underneath both of them, though, it-- you may think this is alien. It's not. There's something deeply human at its core, the refusal to accept that we're temporary, limited, fragile, dust to dust, as- dust to dust, ashes to as-ashes.
as they say in Hebrew. But you don't need religion to see the shape of this. Religion is almost something-- And I'm not here banging anyone's religion, yes, no, or no, it's-- But religion is a way we've dealt with these issues.
The language alone gives it away, transcendence, singularity, salvation through technology. You could put in the theology instead of religion. The emotional core at its heart that gives rise is exactly the same, whether we, we look to it as a God who does it or some technology or, or it's in ourselves.
We don't wanna disappear. We wanna create greater than us. We don't wanna be alone.
That-- And that's human right there. Super intelligent AI answers that second fear about being alone, and digital immortality answers the first. We wanna be here forever.
So when you put them together, you're really getting something unpreceden-unprecedented, a future where intelligence may outlive biology entirely. Whether it's in machines we built or versions of ourselves that become software, it, it, it boggles the mind. It's no wonder people are reacting so strongly to these stories.
You get-- And then-- And this reaction is not just the tech bros or the techie folks that I know and work with every day. It's not that normal tech cha-chatter. I'm telling you, there's something visceral about it.
There's something in the core of humanity, curious, hopeful, uneasy, sometimes all of those things at once. Because here's the deal. If machines become conscious, what are we?
And if we become software, do we lose? What do we lose? And there's a quieter question that kind of runs underneath that.
If we succeed, will the result really be us or just something that resembles us? Will this consciousness and intelligence we create be alive? Dare I say it, alive?
Will the machines we download to be alive? Will they be us or just, again, a-an echo, a copy of us? A digital copy of your memories really isn't you.
A chatbot trained on your personality is, is not you. So in trying to defeat death, we may just be creating echoes of ourselves. And in trying to create intelligence, we may be building minds that don't share our values, our emotions, or our limits, or are not even recognizable to what we want and why we innately try to do this.
So as these two hours-- arrows are racing toward each other, they may not meet in the way we think we want them to meet. They may, in fact, pass each other like two ships passing in the night. Or maybe they converge into something entirely new, something we haven't even contemplated, good-- for good, bad, or worse, whatever.
But here's my bottom line for you folks. This isn't primarily about money, though money certainly plays a part in it, right? Trillions of dollars literally flowing into this whole thing.
It isn't even primarily about the technology, though technology is the enabler. It really is about the oldest questions humans have ever asked since the be-dawn of time. You know, as I mentioned, the Arthur Clarke two thousand and one, the, what are we?
Why are we here? And do we have to end? For most of history, those questions lived in things like theology, religion, philosophy, myth.
Now they're living in research labs and venture capital decks, VC decks. It's crazy. But that doesn't make them any less profound or controversial or hard to wrap your head around.
If anything, it makes this whole thing more urgent 'cause it, it got a little more real. Because this time, rather than, you know, like the apes throwing stones when the, the, the black-- We might actually build something that answers these questions, that allows us to touch the face of it. So when you see headlines about conscious AI, about super intelligence, about digital afterlives, don't think of that as tech news and talk to your tech friends about them.
'Cause that's really what it is. We're trying to create a mind greater than our own, and we're trying to outlive our own bodies, all at the same time. And for the first time in history, both those goals feel within reach at this...
almost at the, simultaneously at the same time. Guys, this isn't about innovation, though it is innovative. It's humanity staring into a mirror and slowly teach- teaching the reflection to stare back at us.
But the real question isn't whether we can build intelligence or defeat death. We may do both sooner than you think. It's whether in doing so, can we hold on to what makes us humans?
Do we lose our humanity? I hope not, 'cause I'm human, you're human. I'm Shimmy, you're out, we're out.
We'll see you next week. Shimmy says, Shimmy says, Shimmy, Shimmy, Shimmy says, ask me almost anything. So I'm Scott Shadley, Director of Leadership Narrative with Solidigm.
5 architectures, and then I'm gonna hand it off to Phil, and he's gonna f- wrap up this section and go on to the, the n- the last section of it. So you only have me for a few more minutes. Be back soon.
21 gigawatts solve? Come on, somebody's gonna smile about that. Great Scott.
Oh, my word. I'm not that old. All righty.
That's shocking. It can send Marty back to the future. " It's also what it takes to power San Francisco for a day.
It can also deliver power to 550,000 Grace Blackwell GPUs, GB300 platforms, and it enables 25 exabytes of storage in a one, one gigawatt environment. Now, how do I know this? What did we do to be able to tell you that one point 21, one gigawatt can do all this?
Again, looking at the ecosystem, looking at our friends, doing the research. These were all announced in 2025. These are all the platforms that are gonna go live sometime after the announcement in 2025.
Stargate, Meta, CoreWeave, xAI, all these guys. And we did some math, and that's how we got to the one gigawatt for 550,000 GB300 platforms. And if you look at the math required for how many GPUs you have and how much storage you need to go along with those, you get your 25 exabytes of storage.
And so the next question, of course, is, "Really? " Well, we did that math for you. Um, again, 550,000 GPUs direct attached.
These are the E1S performance drives going right next to that, that GPU, sitting in the server. That's the nice, rack-based design. Currently has eight drives in it because they're air-cooled or potentially, direct to chip liquid cold plate cooled.
5 exabytes of supportable storage using 122 terabyte drives to be able to get up to 550,000 GPUs. So this math is, is, like all good TCO models, it's a tit for tat, so whatever you put in, you get out. So we focused on, if I use just my solid state drives at the highest capacities for both knobs, it enables us to get to that many GPUs.
You have any other product that consumes different amount of power, or you put 60s in instead of 122s, you're doubling the power footprint, you reduce the number of GPUs available in that gigawatt. So the gigawatt was our bar here. I'll be 100% honest, there's math that shows if you ignore the gigawatt and just look at the GPUs, the amount of exabytes of storage will be just mind-blowing that they're expecting to use.
And this is Grace Blackwell. This is 2025 data. We're now one whole month, literally last day of the month, into, 2026, and the whole ecosystem has changed.
Because when I showed you guys this graph last time at Field Day 2, it had a quote from our good friend Michael Dell. This is our friend Jensen at CES. This is the market that never existed and the market that will likely be the largest storage market in the world.
Gotta love the fact that we finally have Jensen talking about storage and not just memory. I love it. Now, the next trick is at GTC to have him mention our name.
We'll see if we can get that, right? He signed our drive, of course, you know, that whole thing. 5 IC MSP layer in the Vera Rubin platform that's tied to the BlueField-4 implementations.
Now, the next slide, I'm gonna tell you how that works from a SSD hardware point of view, and then a little bit, in just a few minutes, Phil gets the lovely chance to show you how that actually looks from a system level implementation point of view alongside what we're talking about. So what I'm doing here is we have the KV cache exceeds the HBM, spills over to DRAM. We still have a limited DRAM footprint, only so many DIMMs, only so much capacity.
Falls onto the local storage, those nice directly attached products, and then it falls over again. And every hop is a connection and a distance and a time. And so by bringing it closer and closer and closer, that's what this whole architecture is about, is access to data faster- Mm-hmm ...
in a more confined environment. And so if I take what I had before on the one gigawatt example, and I throw in the Vera Rubin platform, or the Rubin platform as it's being called, they haven't officially given it the GB nomenclature, we now have three banks of sto- of storage products. Now, again, I'm constrained to one gigawatt, so I'm still at 25 exabytes.
You get 12 exabytes in the one twenty-two high-capacity storage off sitting on the network, six point four exabytes of this new context memory storage, which can be still be the high-capacity drive. It's just closer to attached to that BlueField forward architecture, and you now have six point one exabytes of direct attached storage 'cause they're changing the amount of local to move it just a little bit further out past the BlueField. So the number of drives in that initial server is actually coming down as the capacities are going up.
And so the interesting thing here is we're-- with the one gigawatt, we still have twenty-five exabytes of storage. We split it out three different ways now instead of two, but note the number of GPUs that are supportable. We're down to four hundred and forty-- or down to four hundred thousand.
We lost a hundred and fifty thousand GPUs, but the performance of the system doesn't change at one gigawatt. And the reason for that is because you're using NVMe storage for the direct attached and for the context structure. You have to have the fast storage products in those two layers to overcome the gigawatt problem in this environment with, this new added layer.
Because fewer, faster GPUs need more access to fast data, therefore you get ICMS with a solid state drive. You can't put traditional rotating hardware in that layer. You just can't.
So if I-- When we come back, at our next field day, AI field day, we're planning to, we're gonna have even more details on this, and we're gonna blow off the one gigawatt and just show you the capabilities of what this storage looks like. And we're talking a fivex or larger CAGR year on year from twenty-six to thirty on just the demand for high-capacity storage over what we were already talking about. It went from where it was at about a twenty percent CAGR, it's now thirty, forty percent CAGR because of this introduction of this, and that's for the Nvidia-only based systems.
So the storage platform is now the shining star on making the success of the next layer of AI as we get into the inference context scenarios. So we're gonna do a little bit of context switching here. Um, we wanna talk about efficiencies too, and so I just gave you the hardware-centric ICMS one gigawatt constrained environment view.
Efficiencies, when we partnered with Vast, we came out with this amazing TCO model that talked about replacing your SEF hard drive infrastructure with R122s and Vast in an efficiency play, talking about how to make your systems more effective. And when we were at SC, we put up a bunch of slides. This was an SC carryover for you guys.
You wanna talk about it from here, the Computer History Museum, from the museum to the one oh one is the SSD implementation on equivalent of this graph. The hard drive implementation is going from the Computer History Museum all the way up to Oracle headquarters, where, where they were up in the nice floor. I know it in Oracle headquarters, but that's how, that's how far our distance is that you do when you put a drive end to end to end, and how much reduction in overall ecosystem environment you can drive.
So what we're gonna do now is I'm gonna hand it over to Phil. He's gonna help explain a little bit about this and then jump back into the context memory and give you some more fun, topics about, the wonderful Vast platform. So ICMSP is inference context- Infe-inference context memory storage platform.
That's what they called it. And it's behind the BlueFin, so it's effectively a, a storage solution out there that's doing something for the context management. I'll dig into it.
He, he- It looks like- Great lead-in to the next little section. And it's different than the rest of the NAS object- Yep ... data lake that's behind it.
It, it re-architects it for you. Okay. Yeah.
Thanks. Yep. Yeah, we'll dig into it.
Okay. Thanks for having me, guys. So Phil Menezes.
I'm the go-to-market execution lead at Vast. I've been there, I think it's three or four days will hit my six years at Vast, so I've kinda got to see the company, grow. We'll do a quick introduction, but I really needed to play off Scott's, metaphor here, which is, one, really happy to be here with Solidigm and our partners.
But Vast, we build software, right? We can't run well, obviously, without any hardware. So peanut butter, great, you know, but it doesn't work so well.
It's not very portable or, reasonable to eat if I'm gonna spread it on my hands. So really the jelly and the bread to our peanut butter is Solidigm. I couldn't go without that piece.
Um, for those who, you know, not as familiar with Vast, we actually launched really the company to the public here at Storage Field Day back in twenty nineteen. And when I was interviewing, that's really how I learned about the company and whether I wanted to work here, right? And I saw some obviously compelling things.
Um, since then, right, we've really become a significant portion of the storage market. As we look to this year, we're gonna drive a very significant portion of all enterprise SSD utilization, with storage expecting dozens of raw exabytes. That's before our data reduction, which we're gonna talk about that efficiency, and also doesn't count any of the data going into the cloud.
We've made some really big announcements on cloud partnerships this year and extending the platform, which was primarily on-prem, into the hyperscaler space. Sales have really gone well for us. We're roughly tripling year over year.
Our quarter's gonna finish tomorrow, so pay attention as we start to, to announce some of those new things. We have our, customer event, first ever user conference for Vast at the end of next month. We'll talk about that.
And I would say just like tech field day's evolved from storage field day to AI field day, Vast has really evolved, from being a storage company to building mu-many more things on the platform, which we'll talk about, really allowing you to capture data, contextualize it, and then act on it with AI. So I would say the founding principle of Vast is that really a few things. We were very bullish in twenty sixteen AI was gonna change the world.
We were very confident that AI was gonna change the way we computed on data, right? I think both of those things proved out to be correct. And then the third is that the architectures that got us to where we are were not the architectures that were gonna take us forward.
Right? And this is really the main culprit. The shared nothing architecture, really invented by Google in two thousand and three in a white paper, basically defined the Internet kind of application cloud era, where you've got these node-based architectures, right?
I've got a node, got some CPU and memory, I've got some kinda storage in there. Originally it was disk, now we've swapped it out for flash in a lot of circumstances. But the only way to get to the data on that node is through that node's controller, right?
And I personally, storage guy, like I look at this as every scale-out NAS, every scale-out object platform, but ultimately it's also the architecture for every data lake, for every distributed data warehouse, right? It's all over the place, eventing infrastructure. It is everywhere, and it really does create a lot of scale problems in the AI world.
When we look at it really from a storage view, there's some challenges around flexibility, right? I've gotta create nodes or pools of homogenous node types, not designed with flash in mind, right? We'll talk about some of the challenges around things like data reduction.
And then finally, one of the big things that shows up everywhere is just this east-west traffic. There's so much communication between these nodes that even though I can scale my resources linearly, I'm not scaling performance linearly. And a lot of times these architectures work well small, and these problems show up more and more and more as the clusters grow.
So we're looking at it now from the, the really the TCO and efficiency perspective. Right, we know we're in a supply crunch. Right, customers have been trying to move steadily from, spinning disk-based architectures to SSDs.
When we look at the AI deployments that we see in practice, you don't see any spinning disk, right? You got power challenges. I, I was wondering why you actually had the round things with the floating heads on them for showing disk drives here.
'Cause our marketing team likes that image, I guess. But these are all, right, the, the world of, uh- That's square ... architecture based on spinning, right?
It's got the arm. It's more like a record player. Oh.
You have an older reference. Too old school? Okay.
Well I'll, I'll, I'll give the marketing team the feedback. What's a record player? I love it.
Okay. So in the AI world, right, I think as you look at what's in practice, solid state's required, right? " But if it was not a good idea before a supply crunch, I don't see how it's a good idea after a supply crunch.
So what we need to do is help customers be a lot more efficient with the way they use SSDs, right? One of the big problems with this shared nothing architecture is how data reduction works, right? And when you look at a lot of these architectures, a lot of them have given up on things like deduplication.
You have compression only, right? And a lot of the data in the unstructured world, it's already compressed. So that kinda takes away a lot of opportunity to drive efficiency, right?
A lot of times when you look at these architectures, one point one to one, one point two to one is kinda what you're gonna get if you're using compression only. So what is the challenge with dedupe? It's around having a global view of the data, right?
In this world, I basically have to chunk up my dedupe index, because the other option would be to put the entire index on one node. Everyone would just hammer it, and that would not work very well, right? " And in that world, every node has a piece of the index, right?
So I get a limited view there. As I start to scale this, all these nodes are talking to each other, looking at who's got the data that I already might have. As I grow, it adds to the east-west traffic, adds to the performance limitations, and then ultimately, the kind of Band-Aid there is to create limited dedupe domains, right?
So I'm basically only deduping within a pool or within a few different nodes, within whatever architecture that you're building around, but it's always very local in this world. So Vast, again, looking at the architectures, brought a new architecture to market that we call Days. Very quickly, we call it disaggregated shared everything, because essentially we kind of broke the idea of a node apart, and we have two independent scaling layers.
We have our logic compute layer. We call those cnodes. It's essentially a container running on an x86 server.
And then we've got where all of the state of the system lives, down in these enclosures filled with very dense Solidigm one hundred and twenty-two terabyte drives or whatever the right, drive is for the customer. Now, some different things. In...
Unlike the shared nothing world, in the Days world, every one of those containers has direct access and actually sees every one of the devices in the system as a local device connected over NVMe over Fabric. Right? Architecture impossible without NVMe over Fabric, which now makes it allow that I can have remote drives feel local from both how they're mounted and performance.
We also have a layer of storage class memory in the system where all of the system's metadata lives. So that means I can create a shared global index that every single one of these containers sees, and that means I can do global data reduction at an exabyte scale without any of those different challenges, right? So fundamentally unique architecture that allows us to look at the data in a very different way from a data reduction perspective.
Any questions high-level to architecture, how it works? Okay. That'll be a theme that we hit on, right?
So step one, can we give an architectural, I would say, advantage to how we look at global data reduction? Step one. But again, the problem is we're talking about unstructured data here, right?
Not as friendly of deduplication as things like VDI and virtual machines, right? Not as friendly of compression maybe as a database that hasn't been compressed already. So we had to look at some different things, right?
And really move beyond dedupe and compression alone. So very high level, you look at compression, right? I'm looking for commonality, repeating data at a very granular level, right?
That's gonna be a small chunk, I don't know, eight to sixty-four K usually. Could be anywhere in between. Uh, deduplication, right?
Now I can have a global view, assuming my architecture allows it, but I'm looking for more coarse matches, right? Two chunks of dataExactly the same. I find that a lot.
VDI, virtual machines, I'm copying databases, whatever. Uh, but again, I don't always find identical matches in unstructured data. If I chunk up and try to do dedupe on a big pool of unstructured data, what you actually end up finding is a lot of chunks of data that are mostly the same, not exactly the same.
Dedupe misses that every single time, 'cause that would be a hash collision that's corrupting your data. It's terrible. So what we do is introduce a new type of data reduction, again, enabled because we have this giant metadata structure living in storage class memory, the architecture, that we will identify if two chunks of data are mostly the same and press them together and essentially store the differences.
Kind of like a snapshot, right? And ultimately, we don't just use similarity, we use all three of these, right? So we're looking for the best opportunity.
Compression, we actually use a couple different types of compression. We will look at the data, take a sample, what's the best type of compression, and use that. Deduplication, we have, again, global deduplication.
We have something we call adaptive chunking, chunking, which means we'll actually change the dedupe window to find the best opportunity for deduplication. And then similarity is kind of that icing on the top where we're gonna find that next level of similarity and drive out even more savings, right? And ultimately, you're in a world where we could easily get two or three times more data reduction than the next biggest competitor because of what's happening here.
Y- you, do you, uh ... I wanna know if you're a believer or not. Uh.
Uh, yes, absolutely. Uh, I, I'm, I'm chuckling because you're, you're giving the exact description of what I would've been describing with SolidFire's architecture 10 years ago. Got it.
Okay. I knew about it, right? But I would say SolidFire in the shared nothing world, right?
A little bit. A- absolutely. Okay.
I, I mean, the, the, the things you're describing are just like, yeah, this is exactly what we were doing 10 years ago. No, it makes- So no, that, that's ... I'm sorry.
That's why I was chuckling. No, no, it makes sense. And by the way, I think it's interesting that, you know, in the block world, the hard drive died like immediately, right?
You know, I was part of the XtremIO team at EMC. We had Pure, we had SolidFire, everyone. The, the hard drive in the, like the block world, virtual machines, vata- databases, VDI, died immediately.
That was like 12 years ago, and there's still so much of the world's unstructured data on spinning disks because they haven't been able to figure this calculus out, right? So it's actually a great point. Um, some actual data, right?
" Um, average data reduction, by the way, this is pulled this month because as we've looked at the supply chain crunch, we're like, let's start digging into like where we've come and what the results are. 4 to 1, right? Again, these are not VDIs.
These are ... This is unstructured data. Some of our customers have hundreds of petabytes of highly compressed video.
Some of it's encrypted. Um, some massive estates. 87 to 1 exactly, if you're curious.
9, looks prettier on the slide. Um, and then we have 27% of our customers get better than three to one. We have some customers getting eight to one.
We have some customers getting like 30 to 1, depending on the data type. So where typically, again, in the world of unstructured data, you'd say, "If I get anything at all, 10%, I'd be happy," we're talking about getting you three times more data for your flash. And that gets combined with something that I'm not gonna nerd out on today because of just time, but our erasure coding is also incredibly efficient.
So our erasure coding at scale, under 3% overhead. It's actually 146 plus four stripe that we use, enabled by our architecture. So, when you look at that compared to, you know, traditional kind of shared nothing, where you're gonna have maybe 27, 20%, we have a lot of customers moving to Vast that are still using like DAS data lake technology, and they've got their data triplicated.
You'd be shocked about how much of the world's capacity is still triplicated, and it's because it's in these monster data lakes, where again, they're getting, you know, for every 10 petabytes they can store 3 petabytes of data, and those systems don't have any data reduction, right? That's all over some of these large data analytics environments. So you combine these things, a lot of times our customers, even if they don't get good data reduction, they're getting four times more effective capacity per, you know, petabyte that they buy.
And even if they're buying something that is more kind of enterprise, then maybe it's more like double the capacity that you can store for every raw petabyte that you're gonna buy. So we actually just launched this, something called Vast Amplify. So in the SSD crunch, Vast, over the years, has really gotten a lot more flexible.
Again, when we started, we had to run on a very specific hardware build. Now we're working with pretty much every major OEM vendor, running on more, traditional servers. We're in the cloud.
So we actually have a program where we're going to customers and taking their SSDs that they already have in their systems and repurposing them into Vast systems to amplify the capacity. We actually had a cus- couple customers come to us and said, "Hey, we've got SSDs. Your technology is way better than what we're using.
" And we have i- in some very large scale environments. I'm talking, at this point we've repurposed hundreds of petabytes of data, thousands and thousands of drives. So just a question.
This is all really great statistics and y'all are doing really awesome, but, we're talking about AI. So I would love if you could tie this back to AI. Can you ...
Does it matter if I have a data lake that's not deduped? Maybe I want that and I just want the ... I just want the data tagged in a different way so I can find it for different reasons.
But like w- what, how does this tie back to AI? Yeah. So I would say how it ties back to AI is right now what we've seen in practice, any large scale training environment, any large scale inference environment that's actually in production at scale is 100% based on SSD, right?
Right. That's, that is stand out. We have in a world where customers are gonna struggle to get-As much SSD as they need, right?
So what I need to be able to do right now is make more use of my solid state devices because AI is driving tremendous demand, right? And we'll get into more how it fits in the architecture, but the point is, looking at bringing spinning disk into this world, we really think is a terrible idea. If having a tier miss is going to kill my GP utilization, destroy jobs, destroy performance, then I can't use that as a lever.
I need to figure out how to make most use of my flash in this AI world, right? As I wanna deploy agents and inference over a much broader set of data- And I'll- ... that data's hitting on spinning disk, it's not gonna work well.
You're gonna get an argument about that here. Yeah. Right.
It's... So- Go ahead. I, I was just gonna say, but when it comes to training, training data specifically is, a- as dedupable, if that's a word, as traditional data sets have, have been in, in your guys' experience so far.
Yeah, so I would say in training data, two to three to one- Okay ... is common, right? If you look at some of the bigger neo clouds that are our customers, two to one's pretty typical on training data sets and, and more.
And you're doing this all inline dedupe, right? So it's s- I would say it's kinda the best of both worlds between inline. In the old world, inline meant in memory.
With VAST, our inline memory is storage class memory, right? So what happens is the data lands in storage class memory, it's acknowledged up to a host, and then we de- we data reduce it when it migrates down to QLC. Mm-hmm, okay.
So, yeah. I'm gonna- Kinda out of band. I'll show you what that looks like after.
Yeah, I mean, that's, that's not terribly unusual f- way to do it. If where you've... you actually need to do the hashing at some later point to, to be able to deduplicate it, but you end up using much less storage later on.
It's, uh- Yep. What I would say the difference is, we don't land it on the capacity tier, right? Okay.
So we don't land it on QLC and go back and mess with it again. It's not good for wear, it's not good for performance. What we do is we leave it in storage class memory where it's very fast access, it gives you a lot of opportunity to move, and then we don't need to plan to have non-deduped and non-reduced data on the capacity tier.
When you're- And I promise, by the way, that most of the rest of the presentation will be specifically on AI, but we wanted to bring in the TCO of making SSDs affordable and, and hacking the supply chain crisis. Well, that... Okay, thank you for saying that, 'cause that was not coming through.
Okay, sorry about that. Appreciate that. When you're, repurposing SSDs, are you having to, to migrate the data off the SSDs and migrate it back on from a VAST perspective, or are you assimilating- We do need to- ...
assimilating the data directly, I mean? We do need to move it, yeah. So if we take a file system, we can't convert it to VAST and data in place.
So a lot of our customers are either working with Swing Space or we're creating clusters and failing nodes out and growing into it. You're seeing a lot of usage of the, the new capability to, reuse SSDs? Yes.
So again, we... " So that was how it got rolling, and since then, yeah, customers are all over us to say, "We know we're looking at the year, we've got more demand, we're looking at rolling more AI. We've lived in a solid state world, and we, we know that we can't have capacity that's 30% utilized," right?
We're giving us one-third of what we're buying to store. Okay, more AI stuff, right? So now I promise the rest specifically on AI, right?
But again, we think flash is that kind of first step as to enabling your data on fast access. So we're just gonna talk about the context challenge and KV cache and why, right? And again, you guys probably have been paying attention to what's going on with NVIDIA, but for people maybe, you know, more infrastructure folks, essentially the thing is here, right?
I ask a question to whatever large language model. The first thing that it does is trying to figure out what do I actually care about, right? There's different words in a statement, the what, the the, what, what is this guy actually asking about versus some of these words that don't make sense.
So I calculate that, turn it into key, key value stores, and that's essentially the context of the conversation. Something else that adds context is maybe a document or a video, right? Someone says, "Hey, I wanna just summarize, you know, SolidIM and VAST Tech Field Day.
" But if I ask a question again, it used to be I had to calculate all of that context again, right? So a question, maybe not the end of the world, but if it's a document or a video, then I'm calculating that a lot, right? Think about some big enterprise organization dumps a new document out to the world, and all of their employees are asking questions about it.
I'm recalculating that same context on that document over and over and over again, right? And then obviously, I need to make sure the decode phase is the answer part. That is where I'm creating an answer that makes sure it's related to the question that you asked.
So there's some big problems with this context piece, which is, one, if I'm recalculating over and over and over again, I'm burning GPU cycles on something that's not adding a ton of value, right? And honestly, GPUs are too expensive for that, right? We had ...
I think NVIDIA's customers are like, "We can't just keep dumping all this CapEx and scaling forever. " That's step one. Problem two, user experience.
If I'm a user, and every time I ask a question about a document, it's going to do a bunch of work, that's so annoying. I want to engage in a conversation with you, not have you forget what we're talking about every time I ask a new question, right? I think, you know, some of these large language models, they know everything about you because they have all of that data.
So that's KV cache. I wanna store that context so I can cont- continue to reuse it, right? And NVIDIA has this hierarchy, which Scott talked about.
So step one, store it in high bandwidth memory, right? Obviously, there's some challenges there. It's really expensive and hard to come by right now.
Uh, the other piece is it's local. So if I'm engaging in a conversation just on, you know, this one session, that's fine. But if my friend is trying to have the same conversation, you know, do I have access to that memory?
Then I can move it down to DRAM, right? Then I can move it to local SSD. Again, everything's local.
And then the next step is shared file and object, right? When you're gonna have a massive drop-off-In performance there, east-west traffic, all those different things we talked about. We shared nothing, so we were like, "We need something right here," right?
5. Something that has a performance closer to local, but is more global in terms of its access, right? And that is essentially what ICMSP is, right?
How do I create that local feel? Now again, I'm not gonna spend a ton of time on this, but one way to do that is to basically take a shared nothing architecture, right? I can either put, you know, a client on the BlueField, or I can deploy my software, right, on essentially the CPUs in these G, GPU servers.
Now again, problems there is I'm bringing the problems of that shared nothing architecture up into my most expensive assets, right? I have east-west trap happening, right? I might have a hotspot, right, where everyone's asking about the same piece of context.
That means I have all my GPU servers attacking one essentially, and asking it for information. I don't know that that's a good idea, right? So what we said is, "We've got, a, a different architecture," right?
We walked through the shared, shared, everything architecture, where I've got this stateless layer. Now, I would say, you know, some potential challenges with this instance of the deployment, really two, right? And again, I think ...
I won't say challenges, but optim- areas for optimization. One, right, I have this layer of CPUs that is essentially kind of in between my access to SSDs, right? Um, and as we know, right, the CPU is always gonna be the bottleneck to SSD performance, right?
You think about the world's most powerful processors, how many do I need from a thread perspective to saturate one single 122 terabyte drive? It's a lot. So we have that problem.
The other problem is, you know, I have to essentially create a copy of data, right? I've got an RDMA operation to our front end, and then I've got another RDMA operation to the SSDs, right? So you kinda have this hop that's happening.
What we're able to do with ICMSP on Vast is actually take our logic, our cnode, and move that up to run on the BlueField. Now, we actually, introduced a prototype of this style architecture, I think with AI Field Day, earlier, but that was with the previous generation of BlueField, right? BlueFields have gotten dramatically more powerful from a core count.
So now I can run my cnode, the logic of the system, up in those BlueFields. This is a paradigm shift, right? I no longer have a host going through other CPUs to basically get in line to get access to data that's on fast media.
Now every node has its own little friend. That's its protocol server. Where's the storage class memory here, Phil?
It's still in the D box down there. You just ca- It's like, yeah, there'll be two layers of storage in that D box. A storage- In the old way, the storage class memory was also in the D box?
Yeah. Nothing changes there. Mm-hmm.
It's just how I give access directly from the host to that device. Okay. Yep.
So we don't have to change anything there, right? Which again, now instead of having to need to use the local SSDs and introduce potentially, you know, c- conflicts and all the different things that we might have by putting software on those servers, I can just have JBODs, right, full of flash with dense Solidigm SSDs, and everyone has direct access directly to the metadata structure and directly to the actual data itself. And again, scaling, and everyone sees everything, so there's no problem in sharing context, right?
If there's a hot piece of context, everyone can access it with a f- a whole bunch of parallelism, but I'm not gonna have any hotspots up top. So Phil, excuse me, Phil, Jack Pollard with Paradigm Technica. It sounds like what you're really doing here is you're running storage controller software on the GPU because you've got spare GPU cycle.
It's on the BlueField. So I have a GPU server. I'm putting a BlueField, which is like a smart NIC now.
It's got 40 cores in it. We're taking those cores, which you don't need for network performance, 'cause it's just more cores than you ever would, and we're running our storage software there. So it's not in the CPUs.
It's not on the GPUs. It's on this little server essentially that's mini, that's a smart NIC, but a lot more powerful than that. Okay.
And the net effect of this is? So ultimately what you get, right? So we talk about certain things, right?
One, much faster time to first token, right? So if someone's asking a question, I now already have that context. I'm sharing it globally, right?
So if anyone's asking about anything. So in, in, in a traditional architecture then, you're making a request from the GPU to a storage controller that's off host. Yep.
Right. And then that storage controller goes, fetches the data, feeds it back. Correct.
And so in this case what you're doing is you're moving that storage controller on host- Correct ... or a little bit closer to the GPU. Correct.
And that's getting you sig- ex- that's accelerating it significantly? Significantly. Okay.
Yeah. And I think there's a few things, right, that come into play. So you've got the acceleration of taking out an RDMA operation in the middle, right?
Uh, you have a, a scaling advantage of the fact that now every time I add a new host, I'm adding compute specifically with that host that is its own storage resources, right? Essentially, it's dedicated. So I'm taking out all the potential conflict, right?
Mm-hmm. You get resources for you. You don't have to fight over them with your partner.
When I have that shared CPU pool, we're all fighting for the same resources, right? So it's a scaling, it's a parallelism, and it's efficiency perspective. The fact that the data is now shared means that I can have more GPU servers able to share more context.
They're much less likely to calculate things again, right? So that's why I get a faster time to first token, because I can pull that context without having to recreate it. I get much better GPU efficiency because my GPUs are not recalculating the same things over again.
They're actually doing inference instead, right? And then the final piece is I'm taking out that entire compute layer, and that all is power that's drawn, and power's precious now. So by taking out that entire server CPU group-I cut power by seventy-five percent.
Got it. Make sense? Mm-hmm.
Okay. So can, can, can you tell us again what ICMSP was? It's context management something, something.
I think it's Inference Context Management Storage platform. Okay. Inference Context- Memory ...
Memory Storage platform is what they called it, and if you Google it, just be careful. There's a whole bunch of other uses of the acronym- ... ICMS and ICMSP.
Just think of it as really cool closed storage. Okay, and- ... I keep thinking about the image that you put up that had the different layers and had context as one of those layers.
So, okay, so is this Vast ICMSP, that's what's being attached to the bluefields? So essentially it's ... ICMSP is, think about it like NVIDIA announced something called Dynamo, right?
And we were working closely with them on that. You've got, like, all these different problems in terms of, you know, how do I manage where inference jobs run on GPUs, right? How do I make sure that I am, intelligently using the different tiers of me- of, you know, memory and storage just for context, right?
So that's just for context. Um, and then how do I scale and run those things? And basically, the ICMSP is a tier of storage and essentially a standard way that Dynamo's gonna interact with that storage.
" So can you go back to your diagram that shows ... Yeah. Okay.
So where is it on this chart? So essentially, this is gonna be used for context in this world, right? The D nodes?
Yeah. So all the da- all the, the ... Sorry, I know I'm not supposed to point to the screen.
All of the context gets stored down in that Dbox layer in the same way we would store any type of data. Okay. And so that is y'all's ICMSP is gonna be in the Dbox.
Correct. Okay. Thank you.
And the data lake that's also underneath that is stored there as well? You easily can. Yeah.
Okay. So, you know, something that I would talk about 'cause I, I ... Let me go to the next slide and maybe it will m- it will help.
So one of the things that's going on right now around ICMSP and KV cache is the question: Should you use any data services, right? Because could they impact performance, right? And we went through that whole shared nothing thing, and certainly if you do use data services on a shared nothing architecture, you're gonna have challenges.
Data reduction's one of them. Another one is something like encryption, right? So right now we're not sure, should we encrypt that data?
I think it's a really bad idea to not encrypt that data. It's a giant shared thing which has everything about every conversation- Mm-hmm ... that all your employees are having with AI.
You might want to encrypt that, right? And you kinda have to save it. Again, but there's also a chance for data reduction.
From what we've tested, you can get between one point three to one and two to one data reduction, which means, again- And, and again, this is- Yeah ... an inference solution, not a training solution. Correct.
It's all inference at scale. Um, so the point is, you know, in our world, this is how we would do it, right? But what's unique and flexible about the Vast world is those bluefield controllers don't have to be the only CPUs that that cluster has access to.
Mm. We can create actually a sidecar pool of compute that just does data services, because remember, those bluefields are gonna write data down in, let's say, storage class memory. We can then have this pool of s- compute take that data, reduce it, store it back down to the QLC, right?
But I can also attach other workloads over here, right? And what I know is that these bluefields all get their own dedicated amount of storage performance. Every bluefield has 40 cores sitting- So in the other prior solution, the bluefields were actually responsible for the deduplication- Yeah, yeah.
We- ... without the other compute side- Correct ... cluster.
Correct. Which they, you know, at this point, this is a very new solution. We're gonna have to do a bunch of tests.
Will they run hot? Will we wanna augment? Will it be enough?
Uh, but the point is, we're the only ones that have this flexibility to say, "We're gonna add compute," that you can then leverage for data services. So essentially, I just care about the fact that these guys can access data, and then let our friends over here take care of all the data services. In the old world, the, these sort ...
would've had to been homogenous nodes, but in this environment, you're taking ... You could put any s- any cluster of compute services out there to be your C nodes. Correct.
Mm-hmm. At this point it's very flexible. We have customers with multiple different generations of compute running C nodes in the same giant environment.
Mm-hmm. We can actually take ... We can pool them, right?
So we can say basically, you know, certain applications can use two C nodes, and the rest can use 20. There's a lot of flexibility in how we can carve this up. That disaggregated shared everything piece gives us flexibility in a way that wasn't possible before.
And you mentioned storage class memory at, down at the D nodes. Those are, different types of SSDs that are tailored to, you know, read/write access and things like that? Right.
So essentially, it's an SSD, you know, lower latency. They're more expensive, in, m- much better endurance profile. Mm-hmm.
Right? So essentially, you know, when we came to market in, Intel ... I'm losing words.
PCM. You know what it is. PC- PCM.
O- Optum? Yeah. Optane.
Optane. See? It died such a long time ago.
It was all we had. Solidigm has a, P5810 SLC-based SSD that is used as a storage class memory solution- Yeah ... for these guys.
Mm-hmm. So. There you go.
So yeah, it's all about endurance profile, cost, latency. This is Marion. I have a quick question.
Sure. Uh, going back to similarity, how do you keep the similarity decisions as the data and the models change? Yeah, so that's really just the underlying data structure, right?
So it's totally abstracted from models or anything else, right? And normally with Dedup, you use a really strong hash, right? That...
'Cause you wanna make sure that I never accidentally mistake two pieces of data for the same. That's why Dedup uses a strong hash. All we're doing is taking a chunk of data and using a weaker hash, and that weaker hash basically says that we don't use it to actually store the data.
" So it doesn't matter if that data came from an inference job, a backup job, whatever. We will look across any data that's been stored on the system, doesn't matter what protocol it landed in, and we will identify that there's commonality, and we just won't store it. And the system has no idea this is happening.
And you're using a weaker hash for, um- Comparison ... what? Comparison.
Compare. Yeah. Because if you use a strong hash, you can't tell.
Because when I use a strong hash, a small difference in the data, results in a very different hash, right? That's why you do that. So a weak hash basically says if the data's slightly different, then I'm gonna get the same result.
So now we know we're in the zone, right? That these two pieces of data are very similar. And what are you hearing from your customers who are in highly regulated industries?
They have no issues with it, right? Um- Okay ... again, data's now typically on a, on a system.
It's chunked up, it's erasure coded, it's spread around anyway. So at this point, you know, it's all generally pointer based. I, I have never heard anyone have issues that they would have to turn it off because of some kind of regulation.
Okay. Thank you. Sure.
Thank you. As far as the KV cache, and actually doing any special caching of the data coming off the Bluefield versus it's all going to storage class memory when it's written, and it'll be re- de- de- gets deduped to QLC or whatever the back end is. Uh, it's not like you're holding that data in storage class memory or anything like that.
We are not. So, you know, in general, KV cache will use the different types of media available, right? So it could land in the memory, the high bandwidth memory, right?
Or it could land in, you know, a local SSD. Uh, this is essentially another tier of KV cache. Um, and for us, yeah, we're, we're gonna keep all that metadata in storage class memory.
But we... You know, there's, there's nothing different about how we have to store the data. Essentially, we've just created a more optimal data path.
Right.