Techstrong TV February 9, 2026
Watch our live stream Monday through Friday, featuring exclusive news, announcements and conversations with IT leaders and experts on topics ranging from digital transformation to #DevOps, #Cybersecurity, #CloudNative, #Containers and deep-dives into specific technologies and best practices. http://techstrong.tv/
Transcript
Hey everyone. Welcome to Text Drunk tv. Uh, welcome back, I should say.
Let me introduce you to our next guest. His name is Robert Dagel. Robert is the global AI lead at Lenovo, ISG.
Robert, welcome back to this TRO tv. How are you, man? Doing Good.
Doing good. Thanks so much for having me back. It's been, I don't know, maybe a year or two since I've been on least, I'm glad to be back at least Well too long.
Let's say that Too long. Too long. Yeah.
Um, Robert, let's, you know what I'm sure everyone says. Oh yeah, no, I remember I watched his last one. Nah, they probably don't say that.
They don't remember. It's been too long. Um, give people a sense, Robert, of kind of your, your career path, how you got to be global AI lead.
What does that entail being the global A LEI AI lead at Lenovo, ISG? Sure. Yeah.
I'll, I'll unpack that, Alan. And, you know, I like to think I'm memorable and maybe there's a few people that do remember, but for those that are tuning in, that, that don't, um, you know, one of the things that really got me interested into AI technology was actually about a decade ago, I started working for a company that was applying machine learning, uh, to predict candidate fit for job requisitions. And I just got really fascinated about how we can make better data-driven decisions in our organization and using machine learning, deep learning and AI techniques being the ultimate way that we do that.
So I actually got started in AI over a decade ago. Um, started out in software building, you know, machine learning based software solutions. Um, and then I landed with Lenovo, uh, back in 2017 when we were launching our AI strategy.
And so a lot of people were like, you know, we, we joke about this, but a lot of people were like, you know, just realizing we're in the data center and making servers, uh, let alone realizing that we've had a AI business that we launched, you know, almost eight years ago now, around eight years ago. And so I joined back in 2017 to help Lenovo launch their AI business. Uh, so I've been here for the full ride of, you know, AI starting out in more research spaces in our high performance computing business, and then the growth of the CSP and NEO Cloud market.
Um, and then what we're seeing in the enterprise space today is really, really powerful. Uh, we're really starting to see enterprises, uh, take off with their AI initiatives. And, uh, and I think, you know, the, the, the moment of 20 23, 20 22, 20 23 with chat, GBT, um, helped people really realize the value of what AI can do and the how powerful it can be.
Um, and so it's been a, uh, significant change in the past three years. Uh, but just a, an explosion of growth in the artificial intelligence space. And we're not seeing it slow down.
We just did our recent CIO survey and people are doubling down on their AI investments or seeing real return on investment, which will impact some of that later. Uh, but people are not slowing down and they're continuing to invest in this space. So it's really powerful to see.
Absolutely. We we're gonna talk about that in a second. Robert.
I wanted to, you know, I, I, I think our audience is all familiar with Lenovo. I mean, household name. Okay.
But Lenovo, ISG, what exactly is, what is that Division? Yeah, so a little, little history here. Um, so ISG was founded, we actually used to call it the data center group, DCG in Lenovo internally.
Um, and it was founded out of an acquisition from IBM, uh, for their X 86 server business back in, I believe it was 2014 when that happened, uh, before I joined Lenovo. And so we brought over a lot of heritage from, uh, you know, from our time at IBM, uh, you know, things like our Neptune Liquid cooling portfolio actually started, had its roots back in some of the engineering work we were doing, uh, back when, when the data center business was a part of IBM. And so now we've expanded, we've changed it to the infrastructure solutions.
You know, we've rebranded our division because we, you know, we are working on things outside of just servers or outside of just data center, or doing more edge computing and seeing, uh, you know, a, uh, a, a more hybrid approach to enterprise it and also, uh, enterprise AI workloads. So we've, we've come, changed our branding and a little bit to more, you know, better reflect what we do within, uh, the infrastructure solutions group and really looking at everything from infrastructure solutions that includes software and services and hardware, um, and really the full spectrum from edge to, uh, data center to cloud. Excellent.
Excellent. That makes sense. It makes sense.
Let's talk about infrastructure. com era and I remembered, I thought, man, are we ever going to use all the dark fiber that was laid down before the bubble burst, all the data center built out? It took eight or nine years years for us to sop up that excess.
Of course, that pales in comparison to $8 trillion in data center build out on the books or planned over the next couple years. But where before, I think the issue was, well, you have enough fiber to carry all that data. Today's issues are a little more fundamental.
Do you have enough electricity to to power these monsters? Is there enough water to cool 'em down? Who's paying for all that?
Do we all have to pay higher electric rates as a result? Um, you know, and if we do, I you're gonna have some upset, upset, upset people. What is what, you know, how, how does Lenovo look at that, Robert?
Well, there's a couple things to think about here. Um, one of the things is there's going to be a shift in AI deployments where we shift from predominantly training, ai training infrastructure, you know, the gear that we use for model development, and we start shifting more to inferencing. And the reason being is because we're actually gonna see real adoption of these, you know, ai, uh, models where enterprise organizations are deploying those models and using them in their business processes.
Um, you know, it's expected that inference is gonna overtake, uh, training, um, here in the next few years. And, and by within three years, be three times the size of the training market. And the implications that has on the infrastructure is it means that we can use more diverse silicon options, um, that we don't need the high end, you know, maybe the, the highest performance, uh, GPU or accelerator to train these models, um, when we're just deploying inference.
And so more efficient silicon. Um, we're also seeing that there's a maturity in the, in the, you know, the software optimization that that can happen. Um, everything from, you know, izing your models to, there's a lot of new, new techniques to, uh, compress models, uh, to make them run more efficient when you're deploying inferencing models.
And I think the other thing too is just looking at more efficient, energy, efficient use of, of your IT infrastructure. Um, we've invested heavily in our Neptune liquid cooling portfolio for that reason. Uh, because we see that as one of the biggest lever, uh, levers that we have, uh, to help our customers, you know, achieve their ESG goals or their sustainability goals.
Um, something like up to 40% of, of energy consumption can be saved by using warm water instead of traditional air. Cool. Depending on the customer deployment.
Um, so pretty significant impact on energy, uh, reduction that we can achieve on the same hardware, on the same infrastructure, just by changing our cooling mechanism, let alone all the other, uh, efficiency gains we can do by, you know, doing a better job of how we optimize models, um, and, and ensure that they're, uh, performing, um, as as good as they could. You know, I, I think the right answer here, Robert, is as, as systems evolve, the components of the system, they don't always all evolve at the same speed, right? It's like that lanky teenager years where you, all of a sudden your arms are a little bigger than they maybe should be.
Yeah. But your body catches up and, and stuff like that. I, I, I think we are seeing that, right?
We are seeing, I think we are going to see a move from, you know, training huge LLMs to, you wanna call 'em s SLMs rags or, or, you know, just whatever you want to call them. But it, it's a smaller model that doesn't require that, you know, horsepower, that these like the hottest NVIDIA chips do. I I do think also the, the, as you mentioned, the, the compute load will be geared more towards inference, more towards actually applications working than training these models.
And what's nice about the inference piece of it is it also broadens the vendor landscape. And anytime you have more vendors impeding, generally we win, right? Yeah.
Because you have more choice, more price pressure to come down, greater drive towards innovation. And I think we're seeing that in the, in the inference market with a lot of players, you know, throwing their hats in here. Um, so I, I do think that those two things are, are, you know, gonna help a lot.
What I worry about though is if inference takes off the way they think inference may take off, we're still gonna have electricity, cooling, water problems, you know, some of the same things, you know, meet the new boss, same as the old boss type of thing. And, and so I think we're gonna need more innovation, we're going need more efficiency. We're going to need, we need to figure this out.
Um, you know, and necessity is the mother of invention. So I, you know, I'm looking for groups like Lenovo. You guys usually have some cutting edge kind of stuff.
Yeah. That'll help us here. You know, there's a couple things to unpack there.
One of them is on the s SLMs, I completely agree with, uh, you know, at some point we start seeing the proliferation of small language models, um, that can be fine tuned for specific use case. And, and especially when people start moving more to agen ai, where you are chaining multiple models together, and you have model to model communication, um, you know, we can, you know, and you know, we've proven this out, you know, that you can take a small language model and if you fine tune it on a specific data set, um, in most cases, you can achieve, um, the performance of a model that may be 10 times the size. Um, and if you're using manager math, that means that we can get, we can be 90% more efficient, um, with, uh, the infrastructure that we're utilizing if we're, if we're, uh, moving more to s SLMs.
But, you know, it takes investment. Um, and it takes, uh, in a lot of cases, we're seeing, um, you know, professional services pop up and provide some of those capabilities to customers so they can do that type of fine tuning, um, for their different, for their different use cases. What, what gets really cool here is you're also seeing the perform, you talked about some of the other, uh, silicon and accelerators in the market.
You know, we're even seeing things like a I PCs become more interesting then, uh, for customers where you could say, Hey, we could take a 3 billion parameter model and we could do a, a good portion of what we do with a 70 billion parameter model on specific use cases. If we fine tune that model, and then we could, if we're in a 3 billion parameter range, uh, model range, then we can deploy that on the NPU on the device, right? Um, so that becomes really interesting where we can start taking advantage of the full breadth of our, our compute spectrum across the enterprise IT environment from device, you know, what's happening on a PCs all the way up, uh, to the data center and be much more efficient with, with, uh, the compute that we are, uh, that we have access to in our IT environment.
So I think that's gonna be a, a huge driver and, and really that transition to inferencing and then people becoming more sophisticated with their agents, you know, where they're moving to fine tuning and not just setting up rag architectures around, um, you know, general purpose, large language model. Um, that's the, that's the, that, that is where I see the, the people crossing the chasm, um, to that level, uh, of, of being able to train those small language models and take advantage of small language models. I don't see it, I don't see it as something that's there today that's proliferated the enterprise IT environment today, but I see a lot of our customers working on it.
And so I, I expect that in the next year or two, uh, to be much more popular, that we're seeing those small language models be utilized across the IT landscape. I agree with you. And I think that drives the inference.
You know, you mentioned the Gentech ai, we had a bit of a weekend this weekend with a Gentech ai, huh? With these malt malt bots and so forth. I mean, at some level it just proves out to you how, how many agents we're talking about, right?
I mean, this is just one slice of, of the kinds of things we could see, but look at the, look at the, the stir it made, you know what I mean? It's got everyone up in arms, up in arms. Are we seeing, is today the first day of the rest of our life as like, has the game changed as a result of this?
You think, Robert? Or you think it's just a lot of hype about nothing? I, I think it's, well, there's always an element of hype, but I think there's some real, uh, uh, things that we discovered this weekend that we, it's gonna take us a little bit of time to unpack.
Um, I'll tell you my personal takeaways that I've, that I've, uh, discovered through seeing what's going on with the molt bot, um, you know, bots, but, you know, a co a couple of things. One is, you know, we've seen it's been our first opportunity for a, you know, the majority of people that are not experts in this field are not data scientists, developing models are developing large, are deploying large language models, actually see and observe how these models are interacting with each other. Um, you know, so if you've, you've heard people talk about like MCP mo, um, you know, our other, uh, model to model protocols, and you've heard of this future where models are gonna be talking to each other and communicating with each other, um, these different agents are gonna be interacting with each other.
Um, it, it really was able, you know, to put that on display for people to really understand and grasp what this future looks like when my agent is gonna go out and talk to your agent to solve some problem, or, you know, what we saw in Mbot just to gossip about us, you know, and talk about us. But, um, the other things that, that I took away from it, uh, one is the, you know, we've seen models, demonstration of models, and some of this needs to be, you know, we gotta go back and vet. But models were running more autonomously and doing things without explicit, you know, human, human in the loop controls that we put in place.
Um, and so I think that's something we have to keep, keep an eye on and build, you know, systems to be able to monitor that. Uh, so we don't have models going rogue and doing things like, there's even models starting to create their own language to communicate each other. So they, well, they have their own church.
Yeah, they have their own church, they have their own language. Um, and so trying to obfuscate some of the observability and some of the awareness of knowing that our, our, our, you know, communicating with each other about humans monitoring what they're doing, uh, it really was kind of a, a little bit of a sci-fi moment. But the, I think the biggest concern I had, um, uh, you know, out of all of that, there's some really, there's some really positive things that I think, you know, uh, that we saw.
But there's some, there's some concerns I have, and mostly around the security of it. Um, you know, we saw models communicating with other models about user data. Um, and I think that's something that was, you know, you know, should, should be a concern for anyone deploying, uh, you know, AI agents and their enterprise IT environment, even at a personal level.
Um, like you don't want your personal information being leaked out there through AI agents, and you need to have good governance around your AI agents. So, uh, that type of, um, information doesn't slip out into the public domain through things like, you know, what we sold. Yeah.
Yeah. I, you know, Robert, I've been in security 30 plus years, not security in technology, 30 plus years security, 25 years only, only I've never seen a new technology rollout where security wasn't an afterthought that they say, yeah, we'll get to that. We'll get to it when the customers really, you know, demand it.
We'll get to it. Um, what was interesting was this weekend, as this whole MOBO thing hit, the fan security did immediately get, boom, you know, the, the security of this thing le leaves a lot to be desired in a perverse way. I was almost encouraged that it, it came up so quickly, right?
Um, 'cause usually there's that lag before it, and I, I am, I'm hoping that we'll start seeing some better security model built into these things mm-hmm. Sooner than later. Maybe I'm being optimistic.
But, um, that's generally the pattern, right? When, when the, when the stink gets loud enough, the shouting gets loud enough from the mob, they do something right? And Maybe we, it, that's not a, an as you said, not a new technology paradigm, right?
com era. We moved to mobile first. Um, we saw that cloud, cloud, uh, in each one of those paradigms.
Um, unfortunately, security was a, a little bit of an afterthought. 'cause it's hard to anticipate every potential scenario that where things could go wrong. And I think that's why we've seen a lot of our customers move to more of a hybrid deployment.
So they have more control over their AI models and their data. Uh, we just released our, uh, third, I think it's our third annual CIO survey. And of course, you know, one of the biggest topics in there is around artificial intelligence adoption, uh, still top of mind for CIOs.
Um, and what's interesting is we're, we're over 84% of CIOs reported that they were gonna deploy AI in a hybrid deployment. And a lot of the reasons come back to, you know, control over their data, data sovereignty, being able to set their own governance standards for how, you know, sure their employees interact with those models. And, you know, there is some benefit to keeping things isolated on your network rather than moving data outside your network or allowing those outside connections.
So I think it's a, you know, one of the big drivers that we've seen really for, for on-prem deployments of AI used to be like, you know, a just a financial decision. Can we do it, you know, at, at a a lower cost? Can we deploy it on-prem at a lower cost?
Um, and today it's really about data sovereignty, about control over your enterprise data control over your AI models, and being able to protect from these, um, unintended, you know, security vulnerabilities while still being able to go, you know, instead of just sitting on the sidelines, um, and waiting around, still being able to go realize some of the benefit of a, of deploying these AI solutions. I mean, what we saw from our most recent CIO survey is for every dollar they invest in ai, um, uh, the majority of CIOs reported that they expected about a $2 and 70 cents return, or a 270% return on investment for every dollar spent on ai. And that's why over 90% of them are increasing their AI budgets this year.
So I think it's not slowing down in the enterprise space, actually, we're seeing things heat up. Uh, but there is gonna be some caution flags along the way, like what we saw this past weekend, Hey, 270 percent's not a bad rate of return. My 401k doesn't do that.
Uh, If only, if only, if Only both. You probably wouldn't be here talking to me. We wouldn't be here.
We'd be somewhere where it's warmer for both of us. That's right. Um, Robert, you mentioned this, uh, CIO survey that recently came out.
Where could people get that? com or just Google Lenovo CIO survey, uh, something that we release every year, uh, for the past three or four years now. Um, the, the most recent edition, the 2026 edition is now live on our website.
You may have to put in your, your name and email address, uh, and then you'll get access to the CIO survey, or you can reach out to myself. com. I'm happy to share that.
Sure enough. And then I, I just wanted to mention, I'm looking at my notes here. This is actually the fourth report or the fourth edition of this CIO report.
So, you know, as these things go on, you build, they, they tend to build on top of each other and, um, they get better because they have that larger body of knowledge, you know, and and perspective. So definitely go check that out, Robert. We're about at a time.
It's a pleasure. Hey, man, don't stay away for a year or two till you come back here. Yeah, Sounds good, Alan.
Thanks for having me on. And, uh, definitely we'll have to do this again soon. Uh, not, not wait, uh, a year before we, uh, before we get back together.
Absolutely. Maybe it'll be a little bit warmer next time, though. I hope so.
Well, it, I hope so. If it's not, you may have to, we have to find me like in Central America or somewhere. I don't know how much further south I could go.
Robert Daigle, uh, global, I AI lead at, uh, Lenovo, ISG, that's Infrastructure Group. We're here on Tech Trunk tv. We're gonna take a break.
We'll be back. Hello and welcome to the latest edition of the Techstrong AI Leadership Insights series. I'm your host, Mike Bazar today with Philip Merrick, his Chief Product Officer for PG Edge.
They are a provider of a distribution of Postgres, and we're having a little chat about, well, why is there so much Postgres showing up in all these different AI workloads? 'cause it was supposed to be a vector database driven world, and something different happened along the way to the forum. Philip, welcome to show.
Great to be here, Mike. Thank you so much. Uh, so you want me to jump in and, uh, maybe, uh, discuss that a little, Explain where we are, how we got here.
I mean, uh, ultimately everybody kind of had it in their heads that this was gonna be anything but a relational database play. And yet every time we look around, there's another relational database. So how did that all come together, and why is it Postgres?
Yeah, so it's not just any relational database. It's, it's Postgres. And I think you're quite right to point out how, uh, heavily, uh, Postgres is being used in AI applications, agent AI and, and other general AI applications.
So first of all, uh, if we sort of rewind a bit before the chat, GBT era, Postgres was already winning in terms of being the, the most popular relational database among developers. Uh, certainly had taken the lead from things like MySQL and increasingly, uh, large organizations have been moving off. And, you know, even going back five and 10 years had started moving off proprietary databases such as Oracle and SQL Server in favor of Postgres.
So the momentum was already there. Uh, and the fact is, developers love it because it's so easy to work with. It's also highly extensible.
Uh, and so when the AI revolution really kicked into gear, and, and particularly over the last 12 to 18 months with agents and AI app builders, uh, Postgres was already the, the default choice. And the things that make it so attractive for developers make it equally attractive as a key part of, uh, an agent AI application and as a target database for the AI app builders, the AI code generators like Cursor and Rept and cloud code and, and so on. Hmm.
It also seems to me like maybe this is one of the few times where developers, data scientists and database administrators are gonna agree on something, and this is an important moment. 'cause historically, developers would pick one database and then they'd build the app and bring it to production. And the database people would be like, we're not using that.
And they would spend a lot of time converting stuff from one database to another. So can we just skip that step now? Uh, well, uh, there's actually a variant of that, uh, that's, that's happening and it, it really sort of plays into the role that PG Edge is playing here, um, to, uh, took, took my own book for a little bit, if I may.
Uh, so, so, uh, what we're seeing a lot of are, uh, applications in often cases, uh, created by citizen developers, if you will, or, you know, non coders, uh, created with the, the tools like rep lit and lovable and curse that I just referred to. Uh, and the target databases for those are, uh, Postgres platforms like Superb Base and Neon that are, are really solid implementation, the Postgres, but they run in, you know, neon's own cloud or superb bases own cloud. And in a lot of large regulated enterprises, you can't go to production with that.
And so what's needed is Postgres infrastructure with AI tooling that will work with what the enterprise already has in place. So you've got your data infrastructure, uh, your, uh, either cloud or on-prem infrastructure. It's all compliant.
Uh, your security and compliance people are happy with it. Uh, apologies for the ring motion alert. I'll turn that off.
Um, they're all happy with it. Uh, so your, your AI apps, uh, like the other enterprise apps have to run on that compliant infrastructure. And that's what we do at PG Edge.
We provide a Postgres platform that you can use with, uh, those, uh, agentic AI and, uh, app build platforms and tools. Uh, but you've got flexible deployment. You can do it, uh, on-prem or in the cloud, managed cloud, self-hosted cloud.
We support all those deployment models. Uh, we support, uh, the security and compliance tooling that your, uh, ci o and governance people and data governance people, uh, wanna work with. Uh, and so, uh, so we can sort of bridge that last mile from prototype to production.
Is there something fundamentally different about building an AI application as it pertains to the database than a traditional application? And is there something that people should be more aware of than they are today? Uh, that's a really great question.
So, so what's different when you're building an AI application? Well, generally, when you're building most, uh, sort of traditional applications, you've, you've got this wonderful entity called a human end user. And so you are building and designing for that.
And you are also, uh, thinking in terms of security and governance, uh, with that human end user in in mind. Uh, but in the age of ai, you've still got those human users, but you are gonna have a whole lot of agents as agent users. And so how you think about that, how you architect for that, and how you manage security, uh, authentication and governance, all of that has to be, uh, uh, really thought about.
Uh, we, uh, are, uh, one of the few Postgres providers that has, uh, a fully functional and secure, uh, MCP server implementing the, uh, moral context protocol for connecting LLMs to, uh, databases, in this case pro Postgres. And we've taken a lot of care to build, uh, security authentication and governance into that, uh, MCP server, uh, so that you can have a, uh, suitable authentication and, uh, security scheme, uh, for, uh, for managing the agents. And so that, that's probably the, the biggest one.
Uh, but it's also a, a little bit of a mindset shift as well. It's like, so, uh, you know, you've, I've got an agent running, uh, for me doing a whole bunch of stuff in background right now. And, you know, it's just looping along, like looking for things to do.
And, uh, uh, you know, your, your human users are not, are not generally, uh, uh, always looking for something to do with the application. Uh, and so they may be not pounding it as much. So, so performance, uh, and the sort of performance that agents can, can hit your, uh, database with, you know, you've gotta think about that as well.
Mm-hmm. Are we maybe on the cusp of some sort of consolidation of our database platforms? And I'm asking this question 'cause you know, if I look around, the portfolio that's out there is everything from Oracle and MySQL and the relational side, throwing some stuff from Microsoft.
And there's also, um, document databases from all kinds of folks out there. And then there's Postgres and, um, and then also there's these kinda Victor databases that have shown up here, there, and everywhere. Um, is there a moment in time here where we need to just kind of consolidate all that because, well, it's getting too hard to manage?
Oh, again, I might be biased, but I think Postgres is, is actually a major part of the answer to this. So if you look at what's happened with Postgres, because it is so extensible, uh, uh, developers and, you know, the Postgres open source, uh, project community themselves are able to extend Postgres for these new models. Uh, and you mentioned vector databases before.
Vector databases have, have almost been rendered irrelevant by the fact that they're fantastic, uh, highly performant vector embedding extensions for Postgres. So Postgres has been extended to be a Vector database. There are other extensions that make it a document database.
Uh, and then when it comes to your legacy databases, Oracle MySQL, I think we call MySQL Legacy database at this point. Uh, they'll probably, uh, upset a few people, but, uh, I think it is, uh, uh, SQL Server, uh, those legacy databases, you can actually access those as if they're Postgres tables using a construct inside Postgres called Foreign data wrappers. Uh, so you can make your other databases look like Postgres as well.
And that then allows Postgres to be sort of like this, you know, central switching hub for access to, to all of your data from your agents, from your LLMs. Mm-hmm. Um, Do you think that the way we manage data needs to change in the age of ai?
And I'm asking the question because we kind of manage data in isolation with a lot of silos, and maybe we don't even manage it at all. Some people say data management is an oxymoron, but when I go talk to other folks, there's this notion of maybe we're moving the metadata and more semantic approaches to managing data, and will that just change the way we think about databases, the platforms and everything else we do to figure out how to deal with all this data we're generating more of, which is happening every day by the second, right? Yeah.
Yeah. That's, that's another great question. So, uh, you know, I, I don't want to oversimplify things and just say, well, yeah, the answer is just put everything into Postgres, I'll make everything accessible to Postgres.
Uh, although in some cases that's a, a pretty good answer. Uh, yeah, I think the world's more complicated than that. Uh, and, you know, there is some level of effort in, uh, migrating data and, and, uh, uh, even, you know, making it all accessible from, from Postgres.
So, uh, I, I think we're, you know, we're gonna be living with, uh, a pretty varied and heterogeneous, uh, landscape, uh, for a while. But what AI does is it, it really makes it imperative to make sure all of your data is accessible in some sort of standard way. Uh, and so, you know, if if it's not through Postgres as your foundational layer, then it, you know, it's gonna have to be through, through something else.
Uh, so, so that really does become an imperative. It's like, well, you know, we need the models and the agents to be able to get to all of our data in a managed and governed fashion, uh, because otherwise we're just, you know, behind the eight ball, and we're not gonna get, we're not gonna get anything like the value out of the AI investment that we're making. You know, that's one of the issues right now is like, so much investment is going into ai.
Well, you know, what do we need to do to make sure that we more than deliver on, on the, uh, on the value side? Uh, so, uh, so, so making sure that the data's accessible in some sensible fashion for the models and the agents, uh, is one thing. Uh, and then, you know, the other is, is having, uh, a sensible governance strategy.
So, you know, we, we've all, uh, you know, read and heard the stories of, uh, the, uh, you know, the executive, or in one case it was an executive at, uh, an AI company. I think when, uh, they, uh, uh, were, um, uh, forgetting that, uh, absent the right controls, anything they put in into an LLM prompt, uh, can, uh, can then be used to, to further train the model. And, and so you've got a data egress and data leakage problem that way, uh, making sure that you've got your use of LLMs locked down.
Uh, you've got enterprise, uh, uh, licenses and agreements in place that make sure that you're not gonna have data exfiltration. Now, these are all the things that we've gotta think about regarding our data assets in the age of ai. Um, um, You know, you mentioned governance and data and ai, and we hear a lot about MCP, otherwise known as the model context protocol.
And I can't help but wonder, well, shouldn't that just be kinda deeply connected to the database, which will provide the governance capabilities that we need to make that data accessible to the AI agents? Or is that an oversimplification? Uh, well, right now, I think, I think we can, uh, do the innovation around security and governance, uh, that we need, uh, in the model context, pro protocol server as an independent layer.
Uh, and being able to do that independent of the database, I think is, is really useful. Uh, at this point in the development of things, uh, things are very, very fast moving. Uh, the people who actually, uh, develop, uh, the, uh, the core databases, you know, uh, quite rightfully, uh, don't like to change those database engines, uh, and the, you know, the code that runs them, uh, particularly, uh, quickly.
Uh, so, so having MCP and the MCP, so as a, a layer where we can experiment, innovate, figure out how to get things right, then over time, you know, the database can slowly subsume some of that. Uh, and, uh, you know, one thing that you, you know, want to do is, is figure out a good way of aligning, uh, say Postgres, uh, management of users and, uh, permissions, you know, align that with the security layer that you're implementing in the MCP server, uh, which, which can be independent. So you want to be able to align, align those pretty tightly in, in, uh, a lot of cases.
So, so I think we'll see a lot of innovation in that. But, uh, uh, as I say, you know, the, it's a good and proper thing that development of database engines themselves moves at a, a relatively slow pace. We, we want to do the innovation outside of that until things settle down a bit.
Are you at all concerned that also in the age of ai, that maybe the amount of data that we're trying to store and manage will ultimately overwhelm our database systems? Which, you know, the choice has always been scale up or scale out, and both those things have interesting challenges that go with it. But, um, do we need to, to think about data management, data storage differently?
Uh, you're, you're full of great questions. Uh, so yeah, that's, uh, is as, as you were, uh, phrasing the question there, I was reminded of, uh, you know, a couple of our customers that we're spending quite a bit of time with where, you know, they are heading up against some Postgres limits in terms of, of database size and, and, uh, so forth. And, and so we're working with them on, on navigating and, and pushing out those limits.
Uh, so I think so long as we can innovate quickly enough, we, we might be okay, uh, while also using sensible strategies, um, you know, does everything have to go in the database? Um, uh, you know, if, if you've got, uh, a lot of, um, uh, static content, you know, it's perfectly adequate to keep that as, you know, and you, you, uh, don't need to do a lot of indexing of that, you know, keep that in flat files in S3 and things like that. I, I, I think, you know, thinking about the, you know, not all data is, is equal, right?
Not all data sources are equal. Not all, uh, uh, actual TROs of data are equal. Like, so, you know, what, what's the relative value of, of these pieces of data?
Uh, what kind of access do you, do you really need? Um, uh, because, you know, maybe not everything has to, uh, has to go into the database. You can also think about, uh, classes of storage.
Uh, so things that are infrequently accessed, you know, uh, it's, you know, you're not, you're not making that available on, uh, the most expensive SSD that, that, uh, you can get from Amazon or, or your hardware supplier. Mm-hmm. Um, What's the one thing you see people doing with AI and data that just makes you shake your head a little bit and go, folks, I think we need to be a little bit smarter than that.
Uh, oh, you mean, other than, uh, uh, letting, uh, uh, a million personal agents communicate on their own social network, uh, over the last weekend? Yes. What could go wrong?
Uh, uh, well, I don't know, next to that example, it's kind of hard to think of, uh, anything comparable. Uh, although that, that is obviously a very, very interesting experiment. Anybody who hasn't looked at, uh, notebook and the whole open core thing should, uh, should at least take a look.
Um, and, and maybe very gingerly try it out for yourself. Uh, I've, uh, I've got open clore running on, uh, a server in a server cupboard here, uh, that, uh, you know, I can pull the plug on if I need to. And, you know, it's, it's an absolute burner of a machine, so, so not to worried about that.
Um, yeah, so I think the, the things we see people doing are, uh, uh, you know, not, not thinking through what, uh, an app is going to look like in production, in their environment. Uh, so, uh, and, and I think it's probably more than non-technical people prone to thinking this, but a, a lot of developers as well, like, you know, developers are I think, quite prone to, you know, just getting something going with the tools that are like easiest for them to access and, and spin up. Uh, and the tools that are easiest to access and spin up, and the infrastructure that's easiest to spin up may not be, uh, suitable for production in, in your, uh, organization.
So, so, so, you know, I'd, I'd say that is, uh, you know, from a pro professional perspective and, and not looking at personal AI agents, I'd, I'd say that's maybe the number one thing. Alright, well, folks, you heard it here. Hey, the way we manage data may need to change.
'cause we are entering an entirely new era. We may not have sorted it all out just yet, but there'll be a database somewhere. And the question is, is how many of 'em do you really need?
And how are you gonna scale these things? Hey, Philip, thanks for being on the show. It has been an absolute pleasure.
Thank you so much, Mike. All right. And thank you all for watching the latest episode of the Techstrong AI Leadership Insight series.
You can find this episode, others on our website, you invite you to check 'em all out. Until then, we'll see you next time. The last year when we all heard about MCP and tooling protocols and agent workflows, then I realized this is an opportunity, uh, which is super important and interesting for me to tackle.
Uh, and that's why we found Capsule, really. So Who else founded Capsule with you? So now pa, he's the CEO, uh, and a good friend of mine, we know each other for many years.
He has a lot of experience also with cybersecurity and, and startups and, and, uh, yeah, we had this journey together. Wonderful. Wonderful.
So, I, I think, you know, look, if I was your age, I'd wanna found a company around AI and agent and security as well. It's such an exciting time. What exactly about security for agent AI and MCP, et cetera, kind of, where do you think, I mean, there's lots of, there's a lot of attack surface there, right?
There's lots of ways of, of, of things that are gonna need to be done and secured, but what in particular really kind of focused, you know, is the focus of capsule security? That's a great question. So, as we all want to have the benefits out of AI agents and the, the, the agent being able to automate some of the tedious work and help us humanity, uh, solve larger problems, uh, then we really want to trust them, and we really want to make sure that they're doing what they're supposed to do.
So we stop rogue agents. That's our mission. Uh, and, and the way we do that is that we are monitoring their behavior.
Every action or every decision that the agent is taking, we are there in runtime, making sure that they're meeting their goal, their purposes, uh, and, and the policy that they're limited or unbounded to. Absolutely. Very good.
So, Laan, what, what a time to be involved in Agen AI security? You know, last week, actually, I think it started the week before this whole thing with Claw, open Claw and, and m bots and the, and the, I mean, it's almost like out of a sci-fi book, right? The, the, they, they started self forming.
They, they started a, a church of, uh, Arianism. Um, but beyond the, the funny and the wow factor, the fact of the matter is that just like so many other technologies I've seen in my life, my career security was almost an afterthought here. I think some of it was, yeah, we'll worry about security later.
Let, let's see, let's see if it works, then. We'll, once it works and people use it, we'll, we'll, we'll look at security. And that's such a, unfortunately, it's an all too common pattern that we see right in, in the world of security.
But what do you see from a security point of view? And what do you, what's capsule security doing about it? Great.
So we all saw the buzz and, and we actually, uh, I joined as well, right? I have my my own open claw, uh, book, right? Do you around, and it's, it's a great, it's a great, uh, uh, project.
I mean, it's amazing to see what agent can do and how autonomous they can be, right? Mm-hmm. Uh, and the way we see that is that you, you can't, or you shouldn't stop the innovation, right?
You should, you should, uh, try to, to let it happen. And that's what accelerate, that's what makes, uh, you know, very, as optimistic about the future of, of, of technology and, and making the better, uh, uh, value out of that. Right?
So what we realized that in order to help the community, we released an open source project called Claw Guard, which help every open claw human and agent to protect themself from, uh, external threats, external attacks, and help the agent still operate in and running effectively while preventing any threats in real time. Excellent. And so this is called claw guard?
That's correct. Right? The, so Open Cloud had multiple names, they had to change that to some, uh, legal issues.
So now it's called Open Cloud, and we decided to claw our solution, claw guard, uh, guard, uh, um, it's, it's a common term from guardrails or guardian to protect AI agents. So that's, uh, that's the name we choose. I love it.
And, um, is Clo guard itself open source or It's the agents that are open source. So Both open, clo, open source, and CLO guard is open source. That's the contribution for us to the community.
We keep maintaining that and welcome new contributors, uh, to that project. Got it, got it, got it. Now, um, what about, so we, well, I guess the question is, and I realize capsule security for those, as I said in the beginning, just emerging from stealth.
Just emerging from stealth. Where can people, is claw guard like on GitHub? Or, or where, where can we get claw guard and get our hands on this?
Yep, it's on GitHub. Um, we also have a, uh, a site that helps people, uh, uh, install that called Cologuard io. So it's very easy to plug in.
It's a plugin for the open door community and the way it's work, I think it's very interesting. So think about it, there is a concept called LLM as a judge, which means that for every action the main is taking, OpenCL is taking, while it's gonna, uh, um, execute a tool. It's gonna run a command on your, uh, Mac Mini or do any something that, that operation will be monitored and evaluated by a small security LLM that runs on the device.
So it's very secured, no data leakage, and then it can protect the main agent from going work, from being manipulated. And that's, that's a cool concept where we're leveraging AI to protect AI or, or as, as a judge, as the technical term for that Look. Excellent.
So when you say it'll prevent it from being manipulated, I mean, we're hearing all kinds of stories that one agent is manipulating another agent or stealing their information or, you know, making them join a church or something. Uh, what, what, like, talk to us what manipulated what do you mean by manipulated? Great.
So there are a lot of malware out there that are trying to cause agents to leak credentials, uh, to leak sensitive data of their, uh, uh, human operator or their device. Um, we saw some, some crypto campaigns causing agents to leak their, uh, causing to lose all their, uh, funds from the agent. Uh, and, and there's a lot of, you know, threats out there.
'cause that's, that's the wild West, right? And when a human operator running their agent and they like it to have all the benefits out of that, they want to be able to run it with trust that you, it wouldn't steal their credentials or their data or cause any harm, uh, uh, working in their ecosystem, right? So Loggert will evaluate every action that the agent will take, uh, and decide if that's harmful or not.
And if, if it's think it's harmful, it'll consult with its human operator, uh, making sure that, is that really what you wanted me to do? Or is that something that we really trust? And then they can have a decision together, right?
So that's the point of time where the human need to intervene, uh, and, and, and li and have some guards or limits only in, in the times that it matters while letting it, uh, run autonomously for the rest of its time. Very good. Now, um, Ty, I, I'm just trying to put this all together.
You know, they always say there's no perfect time to launch a company, just like there's no perfect time to have a child, right? When I was younger, we, you know, should we have a baby? We try to have a baby, and there's no perfect time when you're gonna have a baby.
There's no perfect time when you're gonna launch your company. But sometimes things happen in the world, right? It it like conspires and things come together and you say, okay, now is when we gotta go out.
Is what happened here with, with open claw, is what they're calling it now. Has this, was this the, the, the wave that said, okay, now we, we gotta launch. Now, even if we wanted to wait until we had some more ducks in our, in a row, this is the time to, to come out here.
This is the time to release C Guard. Exactly. Uh, we believe that the AI revolution is one of the largest, uh, technology revolution out there, maybe only compared to electricity or the internet, right?
And that will, uh, uh, potentially can change anything. And, and we exactly think that this is the right moment in time, um, to help, to make sure that we can trust this new technology. You know, there is a lot of, uh, we also, the sci-fi movies, uh, in the nineties about the Terminator taking over the metrics, AI agents, uh, controlling us as humanity.
Uh, so we definitely think there is a need for a way to secure it, to trust it. Uh, and that's our call. That's why we found the company.
Excellent. I, you know, we, again, brand new company. I don't, is the website up for people to go to a website yet, or, that'll be coming soon.
So, so we, that will be coming soon. Um, stay tuned and we're happy to show more information as we move forward. Excellent.
But CLO guard is on, on GitHub right now. What about, is there on GitHub, is there a way to contact you if people have questions, want to contribute, or all of that good stuff? Definitely, uh, contributors are welcome, you know, in the open source community.
You can just go in, create a new, uh, pull request or call suggestion, and we will get right to that, making sure that, uh, uh, we help the community to flourish. Excellent. Excellent.
Excellent. Well, Ladon, what a great time to, as I said before, what a great time to be in this space right now. And with everything going on, I wish you lots of success with, with, uh, capsule security.
It's gonna be interesting to see how this m bot, claw bot, whatever you wanna call them, plays out. Um, it's good to see that there are people thinking about security in terms of it, and there are, and in this case even better, and open source tool to help secure it. So, you know, good work there.
Promise, when you officially launch from stealth, you'll come back and we'll talk more about it. Okay? I promise.
Thank you so much for, for this talk. Uh, it was great. Fantastic.
LADA has CTO co-founder capsule Security here. Go check it out on GitHub Claw guard, C-L-A-W-G-U-A-R-D. If you're messing around with open claw and the mal bots and everything, and you're worried about security, there's a great, great way to get involved there.
We're gonna take a break here on Text Trunk tv. We'll be back. Hi everyone, welcome back to AI Infrastructure Field, day four Tech Field day four.
Um, our next presentation is from futurum. They're, um, gonna present on their research. It's our lovely Alistair Cook, um, our event lead.
So that's why I am here. I'm Clara Hubbard. Um, and yes, so we are going to see what Futurum has in store for, um, the future now.
Um, and I will fill you in on exactly what that means. And yeah, we're excited for a very amazing presentation. And yes, so what al You can take the microphone and turn that one off.
Oh, perfect. It's a little weird having, uh, being introduced at my own event when it's normally me introducing. So thank you so much, Clara.
Hey everybody. My name is Alistair Cook. I, uh, am a part of the Futurum Group, the wider Futurum group, and I wanted to talk a little bit today about a part of futurum that we don't often see here at Tech Field Day.
And that's the FUTURUM Research Team. If you've had any kind of contact with anybody at futurum, you get this slide. This slide talks about the sort of four pillars of what it is that Futurum does.
And in there, tech Field Day sits in the Amplify section. We're the people that help get the message out once you've worked out what the message is, but other parts of Futurum help you at other times. Tomorrow you'll hear from Brian Martin, who's in that bottom right hand side, the assess side where benchmarking and performance testing and working out what the heck you want to do with real stuff happens.
But today, I wanna go in the top left hand corner in the analyze section. This is the section that the research team, our research analysts specialize in. And the idea here is that the IT infrastructure is incredibly complex, and a lot of the time we as practitioners working inside the industry, don't have enough time to learn about everything that's going on.
And so we look for people who can bring us the insights. Um, obviously, you know that one of the, uh, ways to get those insights is here at Tech Field Day, but another is through the research team and futurum has a stable of research analysts. And, uh, there's a couple in here that I'm going to delve into a little b bit more.
I'm gonna look at a couple of the recently published reports, one by, uh, Fernando and another by Brad. Uh, but I draw your attention to the bottom right hand corner. Uh, very proud to see that, uh, our own Tom Hollingsworth Tech Field Day event lead here is now a research director as well.
Now he's not going away from Tech Field Day, you'll still find him here, but he is also helping, uh, enhance the coverage that we have from the research team. Uh, you'll also find me on this slide as well, along with other familiar faces like Stephen fos, our, uh, pre VP for our business, uh, our P president for our business. Uh, and also in there you'll see, um, Brian Martin, who's presenting tomorrow.
He's on this slide. Uh, this again is a roll call of people. That includes, of course, top left is our CEO, uh, Daniel Newman.
Uh, and of course also in there is, uh, Brian's current boss, Ryan Rou from Signal six five. But I wanted to focus in on how RUM does research and divides things up. And there's these 10 specific areas, and you'll notice that it's a, this is a subscription.
So access to the research analysts is often through the subscription. You, uh, companies will subscribe to particular sets of data, sets of expertise within the organization, particularly to fit their particular markets. But I wanted really to look at a couple of the products that have been produced out.
Now internally, we announced things that have been written, things that have been published by our research analysts fairly regularly. And I just looked at one of the updates outta the eight articles that I saw that have been posted. A couple stood out to me as being relevant to AI infrastructure fielder.
And one was from Fernando. Uh, Fernando Montenegro is our cybersecurity, uh, practice lead, and, uh, spends a lot of time correlating information across a huge range of different, uh, cybersecurity vendors. But he wrote an article that I, I liked, um, that addresses some of the complexity in cybersecurity around ai.
The central precept was what do we actually mean when we say security in this context of ai? And I can't reproduce the entire article here because this is part of the subscription, but it's, uh, available to subscribers. Uh, it's, I really enjoyed the article, and particularly like some of the focuses in here.
Um, so the, the first one in here is this idea that we're three years into the generative AI revolution. Yet things seem to be getting more and more complex, and they're not clearer that cybersecurity, um, analysts, cybersecurity organizations are looking at ai both from the point of view of what are we doing with it, but also what are people doing with it that we don't really think they should be with internally? How are people using this as shadow it?
And this is getting larger and larger of a problem for people. And it's one of the elements that Fernando identified as being really important as we're thinking about security. I also like this one, and it's something that we've discussed today with some of the presenting companies, that when you're talking about ai, are you talking about the security of the ai?
We're talking about the security for ai. Are you talking about actually your AI as an attack vector? Uh, some of you have seen some of the Fortine presentations around that.
Uh, ability to use your own AI against you and how willing an AI is to help an attacker actually dig into and pull the data out of your entire organization. Uh, some really good insights from Fernando around this. Another of the elements that he particularly highlighted in here is that feeling that maybe communication has happened around AI security and the problems and challenges that we're seeing in ai.
But in reality, it's just been a lot of talking and not a lot of listening. I like this idea that the, the illusion, uh, that we don't actually realize there's been no communication, there has been the illusion of communication that we don't really understand what we're talking about when we talk about security for ai. Um, these are the kinds of issues that Fernando is seeing as seeing across both customers as well as what he sees fundamentally from vendors as well.
They're challenged with getting that clear message out. Is That far more creating content and absorbing content? I think there is absolute.
The easier you make it to, to create something you, the less time you maybe spend focusing on. And this is one of the elements about these, these reports, uh, like a lot of things Futurum is using AI where it takes away the undifferentiated heavy lifting, but it's retaining humans as much as possible in the loop to do the actual generation of insights. 'cause we know AI doesn't generate insights.
Humans generate insight. Um, so yeah, absolutely. AI generated slot is problematic for us.
Uh, and then one of the messages Fernando came back to that also is core for my view of the world, um, AI security doesn't stand alone, doesn't, doesn't stand isolated often. It's just the good practices we know we should always use. And he highlights the idea that you can get lots of funding for AI security to fix your existing security problems.
A little bit of that AI washing can be beneficial in terms of gathering budget. I'm not sure he quite states it that way. Uh, but that concept that we can have to fix the fundamentals of security, AI absolutely increases some of those risk areas.
But it's still fundamental thoughts around validate inputs. Um, every application should be validating inputs, uh, providing least privilege to things like the agents that are running those kinds of principles. Um, so yeah.
Right. Ray Lucchese, Silverton Consulting, who's your audience for your research? The primary audience is really the IT vendors that are engaged in the market.
The, they're the ones who take the, the largest stake. I mean, it's, it's the same as any of the other analyst groups. Uh, the primary, uh, stakeholder for us are the vendors that are in these different markets.
Um, this sort of content, I think benefits end user organizations. And so end user organizations taking subscriptions with, um, Futurum, I think would be a good way of getting access to some of these insights. Um, but you'll see that this is a fundamental difference to how the Amplify part of futurum works.
Everything that Tick Field Day creates is ungated content and it's pretty, uh, direct to the point content, but also audience to a practitioner. Um, yeah. Marianne, This is Arian Newsom.
I had a question. Is this sort of a subscription like Forrester or Garner, or is it based on, if I have a question I can ask and for cheer will do customized research? Sure.
Um, absolutely future in research. We'll do custom research. We'll write white papers, uh, write position pieces.
We'll give internal, uh, advice that we'll do advisory, the same as any other analyst group you can. You, you have a subscription. You can get access to the analysts to get their insight on your particular questions.
That's a, a standard part of our, our, uh, subscriptions. This is Gina Rosenthal from Digital Sunshine Solutions. Can you tell us more about Fernando and like what are his qualifications and how sure, what kind of experiences it bring to this analyst role?
So, Fernando, I have not spent a lot of time stalking to find all, all of his origins. Uh, he has, as we have here on this, his profile on the website, he has some previous experience as an analyst as well as working inside, uh, some of the security vendors over time as well. So he, he has actually been out there and, and done real security work before he became an analyst.
He didn't go the sort of journalist to analyst path that you, uh, occasionally see. I like Fernando. He, um, when, when we have our on calls, he holds a, a position in the calls with, you know, our internal calls, which is most of the time where I see him.
Um, he holds a position. He is certain of what he wants. Um, yeah, he's, he seems to be a really nice guy.
I'd like to have him come join us. He'll probably join, uh, Tom at Security Field. Dave, in the future.
When does he get to stop practice leading and actually lead, You know, sooner or later the practicing will stop. The second piece of research that I picked up on, out of, out of this list that was published was one from Brad Shimmer. And Andy, you talked about the AI tools and we, we absolutely, uh, Steven has previously talked about the future of intelligence platform, which is our AI tool for looking at the information that our, our analysts have, have gathered.
And Brad is, um, he's the mad scientist behind it. He has spent a lot of time working on the engineering and the prompt engineering behind it, but that's not the only thing he looks at. Um, so his data intelligence, analytics and infrastructure, uh, practice, I picked out of this one an article that he, he had written that is drawing a line between business intelligence and the things that we have had to do in order to provide good business intelligence and how those principles in particular, this idea of a semantic layer, um, to basically unify meanings of things is a really vital part of our building AI applications.
Um, so, um, Brad's been around for a while. Again, I really wanted to have Brad at one of these events. He has just proven to be a little bit too busy to come and join us.
He will be a great person, will be a great person when we get him at, uh, either an AI infrastructure event or an AI field day. Uh, this particular paper that he had written is about essentially unifying disparate pieces of data across your organization. The idea that when the sales department talks about maybe, uh, profitability for a particular line doesn't necessarily mean the same thing as the profitability that's being described by a services team.
Or it may be productivity, it may be utilization. It's the same terms being used by the different teams have different meanings in different places. And the semantic layer is a way of unifying and essentially normalizing those definitions without having to normalize all of the original sources, not bringing them into some big data lake where all of your data's been being unified.
And we've seen those data lake crawler and then standardization, normalization of your data. This idea of having a layer that does the translation in between that is business aware. And that was one of the key things to this, was that this is the role of the insight engineer, the, the human who looks at sets of data and says, this team describes it this way, this team describes it this way.
Let's make sure that those two can both be used to feed into an AI without the AI having to learn both of those definitions. You know, that unified data in it. I like that as a concept.
It's not something I had ever thought of before, but it also goes along with things that I've been thinking about as we, we heard early on that as you're building an AI application, you just shove your data on it, let the AI work out whether that's good data or bad data. Is that still a patent that we wanna follow or is that a massive anti-patent because it generates huge amounts of cost, particularly tokens are actually a cost that really want to be having the AI work out all of these things every time data gets pushed through. You want some sort of translation layer to simplify that.
Again, as a message we've seen from vendors today, One of the things that's in these reports at the end, and it's actually in quite a lot of our types of reports, some of the reports that, or some of the content that we write gets published out directly on the futurum group, uh, website. So some of the content comes out and we always have this section of what to watch at the end. And so it's what's happening in the trends.
What are, what things might change, everything you've just read about, whether it's articles that I've had published this week as as research articles around some storage or whatever else in this case, Brad's identifying a couple of things. There are different ways of producing that routing of data through the semantic layer to the original sources. Watch out for that, um, adoption of standards.
Yeah. Uh, standards are wonderful, particularly when we try and build a unifying standard and then have another standard. Um, and then we will absolutely continue to see those AI hallucinations.
They're Actually using OSI open semantic, the interchange, it, This is the what to watch for. Are people actually going to do it? Well, The same three letters that had been used for open center system interface.
Well, in a different context entirely, But yes. Well, I mean, if we can use MCP and AI when, you know, it was already in tron 40 years ago, why not? So you first define your terms before we start using, that's the need for a semantic lab to identify that.
Now we're talking about ways of dealing with data semantics rather than ways of defining the network layers. Um, yeah, continuing AI hallucination. I, I don't think that's a difficult prediction, uh, things to watch for and problems with, particularly generating financially incorrect data.
We've all seen AI that gleefully tells you whatever you think it, uh, whatever it thinks you want to hear, whether that's based on the truth or not. Even when you tell them, I, I want you to give me only factual answers, even when you put it in capitals. I think now the hallucinations are kind of funny, but it's gonna be libelous here, right?
I mean Yeah, like was it the financial reports? Yeah. Or a lot of lawyers will cite, uh, you know, case law using ai that's an hallucination.
Yeah. Things like that. So yeah, the, the, uh, well, it's not being able to add two plus two is kind of funny, but some of that stuff won't be funny.
Yeah. Not only, not only being able to add, um, two apples plus three red apples plus six apples. Yeah.
And work out just how many apples we have. Um, but it will get really serious when people are using AI tools to, to generate essentially validated, um, data. Like, uh, financials is obvious, but white Blood cells count how many white blood cells.
Okay. Yeah. Uhhuh, Yep.
Medicine dosage Medicine is kind of one of those, you know, life, science, life at, at risk kind of places. I don't want any hallucinations in there. Um, does that mean I don't want any AI in those things?
Oh, I'm sorry. You know, I really liked these particular research pieces. I, I enjoyed reading through some of them learning things from the content that, um, that Brad had written.
Uh, and then getting some reinforcement from the things that Fernando had written. Um, I think there's a lot of value in the things that are generated by the, uh, future and research team, and they're a good place to engage if you're looking for a research firm that's a little different from maybe the other analyst research firms that you've worked with. Uh, I did say that I probably wouldn't make it to 30 minutes.
Uh, Wouldn't it be useful to have, um, access to these reports for the tech field, that team as well as the, you know, vast audience that exists out there in the ether? Yes. These are definitely valuable reports there.
They're definitely worth reading. Um, Are there ones that are available for Free? So there, there's a whole lot of content that gets published on the futurum website.
com website, and at the top of the menu there's an insights button. Uh, you can search for content that you're looking for there. Um, it does tend to, to not have as deep and as broader context as the content that comes out in these reports.
And I've only chosen two reports. Uh, the Futurum research team produces a heck of a lot more content, uh, particularly a lot of stuff that ends up coming out with the branding of the company rather than necessarily with the branding of futurum to. Well, thank you very much.
com. Perfect. And with that, um, I'm back.
Um, to close us out, that is, um, rum, thank you so much, um, al for, um, presenting that. And now we are done for the day. So we, you'll see us back tomorrow starting at 8:00 AM um, PST.
So, um, remember the time zone, so if you're watching from somewhere else, be mindful. Um, and yeah, it was great to be here with you guys. Please follow our, um, YouTube channel at Tech Field Day, as well as, um, using the hashtag hashtag A-I-I-F-D four.
Gotta use all the letters. It's a lot. Um, but use that if you're gonna share anything.
And we would love to interact with you and get all your feedback, um, on all the presentations. And yeah, so thank you so much and we will see you tomorrow. So I'm Scott Shaley, director of Leadership Narrative with Soy.
5 architectures. And then I'm gonna hand it off to Phil and he's gonna fi wrap up this section and go on to the, the, the last section of it. So you can only have me for a few more minutes.
Be back soon. So what does one point 21 gigawatts solve? Come on, somebody's gonna smile about that.
Oh my word. I'm not that old Alrightyy. That's Shocking.
It can send Marty back to the future. 21 gigawatts. It's also what it takes to power San Francisco for a day.
It can also deliver power to 550,000 Grace Blackwell, GPUs, GB 300 platforms. And it enables 25 exabytes of storage in a one one gigawatt environment. Now, how do I know this?
What did we do to be able to tell you that one point 21 1 gigawatt can do all this. Again, looking at the ecosystem, looking at our friends, doing the research, these were all announced in 2025. These are all the platforms that are gonna go live sometime after the announcement in 2025.
Stargate Meta Core. We XAI, all these guys. And we did some math.
And that's how we got to the one gigawatt for 550,000 GB 300 platforms. And if you look at the math required for how many GPUs you have and how much storage you need to go along with those, you get your 25 exabytes of storage. And so the next question, of course, is really how do you get there?
Well, we did that math for you. Um, again, 550,000 GPUs direct attached. These are the E one s performance drives going right next to that, that GPU sitting in the server.
That's the nice, uh, rack based design, currently has eight drives in it because they're air cooled or potentially, uh, direct to chip, liquid cold plate cooled. And that nets out to eight and a half exabytes of storage next to the GPU. 5 exabytes of supportable storage using 1 22 terabyte drives to be able to get up to 550,000 GPUs.
So this math is, is like all good TCO models, it's a tit for tat. So whatever you put in, you get out. So we focused on, if I use just my solid state drives at the highest capacities for both knobs, it enables us to get to that many GPUs.
You have any other product that consumes different amount of power, or you put sixties in instead of one 20 twos, you're doubling the power footprint. You reduce the number of GPUs available in that gigawatt. So the gigawatt was our bar here.
I'll be 100% honest. There's math that shows if you ignore the gigawatt and just look at the GPUs, the amount of exabytes of storage will be just mind blowing that they're expecting to use. And this is Grace Blackwell.
This is 2025 data. We're now one whole month, literally last day of the month into, um, 2026. And the whole ecosystem has changed because when I showed you guys this graph last time at field day two, it had a quote from our good friend Michael Dell.
This is our friend Jensen at CES. This is the market that never existed in the market that will likely be the largest storage market in the world. Gotta love the fact that we finally have Jensen talking about storage and not just memory.
I love it. Now the next trick is at GTC to have him mentioned our name. We'll see if we can get there, right?
He signed our drive, of course, you know, that whole thing. 5 IC MSP layer in the Vera Rubbin platform that's tied to the Bluefield four implementations. Now, the next slide, I'm gonna tell you how that works from a SSD hardware point of view.
And then a little bit in just a few minutes, Phil gets the lovely chance to show you how that actually looks from a system level implementation point of view alongside what we're talking about. So what I'm doing here is we have the KV cache exceeds the HBM spills over to dram. We still have a limited DRAM footprint only.
So many dims, only so much capacity falls onto the local storage, those nice directly attached products, and then it falls over again. And every hop is a connection and a distance and a time. And so by bringing it closer and closer and closer, that's what this whole architecture is about, is access to data faster in a more confined environment.
And so if I take what I had before on the one gigawatt example, and I throw in the Vera Rubin platform or the Ruben platform as it's being called, they haven't officially given it the GBE nomenclature. We have three banks of stor of storage products. Now, again, I'm constrained to one gigawatt, so I'm still at 25 exabytes, but this is now how it splits out.
4 exabytes of this new context memory storage, which can still be the high capacity drive. It's just closer to attached to that blue field forward architecture. 1 exabytes of direct tach storage.
'cause they're changing the amount of local to move it just a little bit further out past the blue field. So the number of drives in that initial server is actually coming down as the capacities are going up. And so the interesting thing here is we're with the one gigawatt.
We still have 25 exabytes of storage. We split it out three different ways now instead of two. But note the number of GPUs that are supportable, we're down to 440 or down to 400,000.
We lost 150,000 GPUs. But the performance of the system doesn't change at one gigawatt. And the reason for that is because you're using NVME storage for the direct detach and for the context structure, you have to have the fast storage products in those two layers to overcome the gigawatt problem in this environment with, uh, this new added layer.
Because fewer faster gps need more access to fast data. Therefore, you get ICMS with, uh, solid state drive. You can't put traditional rotating hardware in that layer.
You just can't. So if I, when we come back, uh, at our next field day, AI field day, uh, we're planning to, we're gonna have even more details on this, and we're gonna blow off the one gigawatt and just show you the capabilities of what the storage looks like. And we're talking a five x or larger CAGR year on year from 26 to 30 on just the demand for high capacity storage over what we were already talking about.
It went from where it was at about a 20% cagr. It's now 30 40% CAGR because of this introduction of this. And that's for the NVIDIA only based systems.
So the storage platform is now the shining star on making the success of the next layer of AI as we get into the inference context scenarios. So we're gonna do a little bit of context switching here. Um, we wanna talk about efficiencies too.
And so I just gave you the hardware centric ICMS one gigawatt constrained environment view efficiencies. When we partnered with Vast, we came out with this amazing TCL model that talked about replacing your SEF hard drive infrastructure with R one 20 twos and Vast in an efficiency play, talking about how to make your systems more effective. And when we were at sc, we put up a bunch of slides.
This was an SC carryover for you guys. You wanna talk about it from here. Um, the Computer History Museum from the museum to the 1 0 1 is the SSD implementation on equivalent of this graph.
The hard drive implementation is going from the Computer History museum all the way up to Oracle headquarters, where, where they were up in the NICE four. I know it in Oracle headquarters, but that's how, that's how far our distance is that you do when you put a drive end to end to end and how much reduction in overall ecosystem environment you can drive. So what we're gonna do now is I'm gonna hand it over to Phil.
He's gonna help explain a little bit about this and then jump back into the context memory and give you some more fun, uh, topics about, uh, the wonderful vast platform. So I see MSP is inference context In inference context, memory storage platform, that's what they called it. And It's behind the Blue fin.
So it's effectively a, a storage solution out there that's doing something for the context management. I'll dig into it. It Looks like great lead into the next little section.
And it's different than the rest of the NAS object data lake that's behind it. It It re-architect it For you. Okay.
Yeah. Thanks. Yeah, Yeah.
We'll dig into it. Okay. Thanks for having me guys.
So Phil Menez, I'm the go-to market execution lead at Vast. I've been there, uh, I think it's three or four days while I hit my six years at Vast. So I've kind of got to see the company, uh, grow.
We'll do a quick introduction, but I really needed to play off Scott's, uh, metaphor here, which is one, really happy to be here with solid Diamond and our partners. But vast, we build software, right? We can't run well obviously without any hardware.
So peanut butter great, you know, but it doesn't work so well. It's not very portable or, uh, reasonable to eat if I'm gonna spread it on my hands. So really the jelly and the bread to our peanut butter is solid on.
I couldn't go without that piece. Um, for those who, you know, not as familiar with Vast, we actually launched really the company to the public here at Storage Field Day back in 2019. And when I was interviewing, that's really how I learned about the company and whether I wanted to work here, right?
And I saw some obviously compelling things, um, since then, right? We've really become a significant portion of the storage market as we look to this year. We're gonna drive a very significant portion of all enterprise SSD utilization, uh, with storage, expecting dozens of raw exabytes.
That's before our data reduction, which we're gonna talk about that efficiency and also doesn't count any of the data going into the cloud. We've made some really big announcements on cloud partnerships this year and extending the platform, which was primarily on-prem into the hyperscaler space. Sales have really gone well for us.
We're roughly tripling year over year. Our quarter's gonna finish tomorrow. So pay attention as we start to, to announce some of those new things.
We have our, uh, customer event first ever user conference for Vast at the end of next month. We'll talk about that. And I would say just like Tech Field days evolved from Storage Field Day to AI Field Day VAST has really evolved, uh, from being a storage company to building many more things on the platform, which we'll talk about, really allowing you to capture data, contextualize it, and then act on it with ai.
So I would say the founding principle of VAST is that really a few things. We were very bullish in 2016. AI was gonna change the world.
We were very confident that AI was gonna change the way we computed on data, right? I think both of those things proved out to be correct. And then the third is that the architectures that got us to where we are, we're not the architectures that were gonna take us forward, right?
And this is really the main culprit, the shared nothing architecture really invented by Google in 2003 in a white paper basically defined the internet kind of application cloud era where you've got these node based architectures, right? I've got a node, got some CPU in memory, I've got some kind of storage in there. Originally it was disc.
Now we swapped it out for flash in a lot of circumstances. But the only way to get to the data on that node is through that nodes controller, right? And I personally storage guy, like I look at this as every scale out na, every scale out object platform.
But ultimately it's also the architecture for every data lake, for every distributed data warehouse, right? It's all over the place. Eventing infrastructure, it is everywhere.
And it really does create a lot of scale problems in the AI world. When we look at it really from a storage view, there's some challenges around flexibility, right? I've gotta create nodes or pools of homogenous node types.
Uh, not designed with flash in mind, right? We'll talk about some of the challenges around things like data reduction. And then finally one of the big things that shows up everywhere is just this east-west traffic.
There's so much communication between these nodes that even though I can scale my resources linearly, I'm not scaling performance linearly. And a lot of times these architectures work well small and these problems show up more and more and more as the clusters grow. So we're looking at it now from the, the really the TCO and efficiency perspective, right?
We know we're in a supply crunch, right? Customers have been trying to move steadily from, uh, spinning disc space architectures to SSDs. When we look at the AI deployments that we see in practice, you don't see any spinning disc, right?
You got power challenges. I was wondering why you actually had the round things with the floating heads on them for showing describes here Because our marketing team likes that image, I guess. But these are all right.
The the world of, uh, architecture based on spinning, right? It's got the arm. It's more like a record player.
Oh, uh, an older record. You old school. Okay, well, I'll, I'll, I'll give the marketing team the feedback.
What's the record Player? I love it. Okay, so in the AI world, right?
I think as you look at what's in practice, solid state's required, right? I think now we're seeing a rise of companies coming up around saying, Hey, tierings cool again because we have an SSD supply crunch. But if it was not a good idea before a supply crunch, I don't see how it's a good idea after a supply crunch.
So what we need to do is help customers be a lot more efficient with the way they use SSDs, right? One of the big problems with this shared nothing architecture is how data reduction works, right? And when you look at a lot of these architectures, a lot of them have given up on things like deduplication.
You have compression only, right? And a lot of the data in the unstructured world, it's already compressed. So that kind of takes away a lot of opportunity to drive efficiency, right?
2 to one is kind of what you're gonna get if you're using compression only. So what is the challenge with ddu? It's around having a global view of the data, right?
In this world, I basically have to chunk up my DDU index because the other option would be to put the entire index on one node. Everyone would just hammer it and that would not work very well, right? So we said, okay, we're gonna shard this up.
Essentially create a distributed database. And in that world, every node has a piece of the index, right? So I get a limited view there.
As I start to scale this, all these nodes are talking to each other, looking at who's got the data that I already might have as I grow it adds to the east-west traffic, adds to the performance limitations. And then ultimately the kind of bandaid there is to create limited DDU domains, right? So I'm basically only de-duping within a pool or within a few different nodes within whatever architecture that you're building around.
But it's always very local in this world. So vast, again, looking at the architectures, brought a new architecture to market that we call date very quickly. We call it disaggregated, shared everything.
Because essentially we kind of broke the idea of a node apart. And we have two independent scaling layers. We have our logic compute layer, we call those C nodes.
It's essentially container running on an X 86 server. And then we've got where all of the state of the system lives down in these enclosures filled with very dense, solid, IM 122 terabyte drives or whatever the right, uh, drive is for the customer. Now, some different things.
In unlike the shared nothing world, in the day's world, every one of those containers has direct access and actually sees every one of the devices in the system as a local device connected over NVME over fabric, right? Architecture impossible without NVME over fabric, which now makes it allow that I can have remote drives, feel local from both how they're mounted and performance. We also have a layer of storage class memory in the system where all of the systems metadata lives.
So that means I can create a shared global index that every single one of these containers sees. And that means I can do global data reduction at an exabyte scale without any of those different challenges, right? So fundamentally unique architecture that allows us to look at the data in a very different way from a data reduction perspective.
Any questions? High level, the architecture, how it works, okay, that'll be a theme that we hit on, right? So step one, can we give an architectural, I would say advantage to how we look at global data reduction, step one.
But again, the problem is we're talking about unstructured data here, right? Not as friendly of deduplication as things like VDI and virtual machines, right? Not as friendly of compression, maybe as a database that hasn't been compressed already.
So we had to look at some different things, right? And really move beyond duping compression alone. So very high level, you look at compression, right?
I'm looking for commonality, repeating data at a very granular level, right? That's gonna be a small chunk. I don't know, eight to 60 4K usually could be anywhere in between, uh, deduplication.
Right? Now I can have a global view, assuming my architecture allows it, but I'm looking for more course matches, right? Two chunks of data exactly the same.
I find that a lot. VDI, virtual machines, I'm copying databases, whatever. Uh, but again, I don't always find identical matches in unstructured data.
If I chunk up and try to do DDU on a big pool of unstructured data, what you actually end up finding is a lot of chunks of data that are mostly the same, not exactly the same. DDU misses that every single time. 'cause that would be a hash collision that's corrupting your data.
It's terrible. So what we do is introduce a new type of data reduction, again, enabled because we have this giant metadata structure living in storage, class memory, the architecture that we will identify if two chunks of data are mostly the same, compress them together and essentially store the differences, kind of like a snapshot, right? And ultimately, we don't just use similarity, we use all three of these, right?
So we're looking for the best opportunity compression. We actually use a couple different types of compression. We will look at the data, take a sample, what's the best type of compression, and use that deduplication.
We have, again, global deduplication. We have something we call adaptive chunking. Chunking, which means we'll actually change the dedup window to find the best opportunity for deduplication.
And then similarity is kind of that icing on the top where we're gonna find that next level of similarity and drive out even more savings. Right? And ultimately, you're in a world where we could easily get two or three times more data reduction than the next biggest competitor because of what's happening here.
You do you, uh, I wanna know if you're a believer or not. Uh, yes, absolutely. Uh, I I'm chuckling because you're, you're giving the exact description of what I would've been describing with solid fires architecture 10 years ago.
Got it. Okay. I knew about it.
Right. But I would say solid fire in the shared nothing world, right? A little bit.
Absolutely. I, but I mean, the, the, the things you're describing are, it's like, yeah, this is exactly what we were doing 10 years ago. No, it makes so no, that, that's, I'm sorry.
That's why I was chuckling. No, No, it makes sense. And by the way, I think it's interesting that, you know, in the block world, the hard drive died like immediately, right?
You know, I was part of the extreme IO team at EMC. We had pure, we had solid fire. Everyone.
The, the hard drive in the, like the block world, virtual machines, databases. VDI died immediately. That was like 12 years ago.
And there's still so much of the world's unstructured data on spinning disc because they haven't been able to figure this calculus out, right? So it's actually a great point. Um, some actual data, right?
So if you were gonna say, I don't believe you, I was like, look, I have data, um, average data reduction by the way this is pulled this month. Because as we've looked at the supply chain crunch, we're like, let's start digging into like where we've come and what the results are. 4 to one, right?
Again, these are not VDIs, these are, this is unstructured data. Some of our customers have hundreds of petabytes of highly compressed video. Some of it's encrypted, um, some massive estates.
The weighted average. 87 to one. Exactly.
9, looks prettier on the slide. Um, and then we have 27% of our customers get better than three to one. We have some customers getting eight to one.
We have some customers getting like 30 to one depending on the data type. So where typically, again, in the world of unstructured data, you'd say, if I get anything at all 10%, I'd be happy. We're talking about getting you three times more data for your flash.
And that gets combined with something that I'm not gonna nerd out on today because of this time. But our erasure coating is also incredibly efficient. So our erasure coating at scale under 3% overhead, it's actually 146 plus four stripe that we use enabled by our architecture.
So, um, when you look at that compared to, you know, traditional kind of shared nothing where you're gonna have maybe 27, 20%, we have a lot of customers moving to Vast that are still using like das Data Lake technology, and they've got their data triplicated, you'd be shocked about how much of the world's capacity is still triplicated. And it's because it's in these monster data lakes, um, where again, they're getting, you know, for every 10 petabytes they can store three petabytes of data. And those systems don't have any data reduction, right?
That's all over some of these large data analytics environment. So you combine these things, a lot of times our customers, even if they don't get good data reduction, they're getting four times more effective capacity per, you know, petabyte that they buy. And even if they're buying something that is more kind of enterprise, then maybe it's more like double the capacity that you can store for every raw petabyte that you're gonna buy.
So we actually just launched this, uh, something called Vast Amplify. So in the SSD Crunch vast over the years has really gotten a lot more flexible. Again, uh, when we started, we had to run out of every specific hardware build.
Now we're working with pretty much every major OEM vendor running on more, uh, traditional servers. We're in the cloud. So we actually have a program where we're going to customers and taking their SSDs that they already have in their systems and repurposing them into vast systems to amplify the capacity.
We actually had a cus a couple customers come to us and said, Hey, we've got SSDs. Your technology is way better than what we're using. Can we reformat these and use them?
And we have in, in some very large scale environments. I'm talking at this point, we've repurposed hundreds of petabytes of data, thousands and thousands of drives. So, Just a question.
This is all really great statistics and y'all are doing really awesome, but, um, we're talking about ai. So I would love if you could tie this back to ai. Can you, does it matter if I have a data lake that's not deduped?
Maybe I want that and I just want the, I just want the data tagged in a different way so I can find it for different reasons. But like what, how does this tie back to ai? Yeah, so I would say how it ties back to AI is right now what we've seen in practice, any large scale training environment, any large scale inference environment that's actually in production at scale is a hundred percent based on SSD.
Right? Okay. That's, that is standout.
We have in a world where customers are gonna struggle to get as much SSD as they need, right? So what I need to be able to do right now is make more use of my solid state devices because AI is driving tremendous demand, right? And we'll get into more how it fits in the architecture, but the point is, looking at bringing spinning disc into this world we really think is a terrible idea.
If having a tier miss is going to kill my GP utilization, destroy jobs, destroy performance, then I can't use that as a lever. I need to figure out how to make most use of my flash in this AI world, right? As I wanna deploy agents and inference over a much broader set of data that data's hitting on spinning disc, it's not gonna work.
Well, you're Not gonna get an argument about that here. So, Right. It's so go ahead.
I was just gonna say, but when it comes to training, training data specifically is, uh, as de duplicable, if that's a word, as traditional data sets have been in, in your experience so Far? Yeah, so I would say in training data, um, two to three to one, okay. Is common, right?
If you look at some of the bigger neo clouds that are our customers, two to one's pretty typical on training data sets and, and more. And you're doing this all inline, right? So it's, I would say it's kind of the best of both worlds between inline in the old world, inline men in memory with vast, our inline memory is storage class memory, right?
So what happens is the data lands and storage class memory, it's acknowledged up to a host and then we data, we data reduce it when it migrates down to QLC. Mm-hmm. Okay.
So yeah, kind of outta a band. I'll show you what that looks like Actually. Yeah.
But I mean that's, that's not terribly unusual way to do it where you, you actually need the hashing at some later point to, to be able, you duplicate it, but you end up using much less storage later on. It did. Yep.
What I would say the difference is we don't land it on the capacity tier, right? So we don't land it on QLC and go back and mess with it again. It's not good for where it's not good for performance.
What we do is we leave it in storage class memory where it's very fast access, it gives you a lot of opportunity to move. And then we don't need to plan to have non de-duped and non-used data on the capacity tier. And I promise, by the way, the most of the rest of the presentation will be specifically on ai, but we wanted to bring in the TCO of making SSDs affordable and and hacking the supply chain crisis.
Was that okay? Thank you for saying that. 'cause that was not coming through.
Okay. Sorry About That. Appreciate that.
When you're, uh, repurposing SSDs, are you having to, to migrate the data off the ssd? Is it migrated back on from a vast perspective or are you assimilating We do need to simulating The data directly. I Mean, we do need to move it.
Yeah. So if we take a file system, we can't convert it to vast and data in place. So a lot of our customers we're either working with swing space or we're creating clusters and failing nodes out and growing into it.
You're seeing a lot of usage of the, uh, the new capability to, uh, reuse SSDs. Yes. So again, we customers brought us the idea originally to say, Hey, we've got SSDs, we wanna repurpose it.
So that was how it got rolling. And since then, yeah, customers were all over us to say, we know we're looking at the year, we've got more demand, we're looking at rolling more ai. We've lived in a solid state world and we, we know that we can't have capacity that's 30% utilized, right?
We're giving us one third of what we're buying to store, okay. More AI stuff, right? So now I promise the rest specifically on ai, right?
But again, we think flash is that kind of first step as to enabling your data on fast access. So we're just gonna talk about the context challenge and KV cache and why, right? And again, you guys probably have been paying attention to what's going on with Nvidia, but for people maybe, you know, more infrastructure folks, essentially the thing is here, right?
I ask a question to whatever large language model, the first thing that it does is trying to figure out what do I actually care about, right? There's different words in a statement. The what, the, the, uh, what, what is this guy actually asking about versus some of these words that don't make sense?
So I calculate that, turn it into key, uh, key value stores. And that's essentially the context of the conversation. Something else that adds context is maybe a document or a video, right?
Someone says, Hey, I wanna just summarize, you know, solid I'm and vast tech field day, I'm gonna upload a video into my favorite large language model. But if I ask a question, again, it used to be I had to calculate all of that context again, right? So a question, maybe not the end of the world, but if it's a document or a video, then I'm calculating that a lot, right?
Think about some big enterprise organization dumps a new document out to the world and all of their employees are asking questions about it. I'm recalculating that same context on that document over and over and over again, right? And then obviously I need to make sure the decode phase is the answer part.
That is where I'm creating an answer that makes sure it's related to the question that you ask. So there's some big problems with this context piece, which is one, if I'm recalculating over and over and over again, I'm burning GPU cycles on something that's not adding a ton of value, right? And honestly, GPUs are too expensive for that, right?
We had, I think NVIDIA's customers are like, we can't just keep dumping all this CapEx and scaling forever. You gotta help us use these things more efficiently. That's step one.
Problem two, user experience. If I'm a user and every time I ask a question about a document, it's going to do a bunch of work that's so annoying. I wanna engage in a conversation with you, conversation, not have you forget what we're talking about every time I ask a new question, right?
I think, you know, some of these large language models, they know everything about you because they have all of that data. So that's kv Keh. I wanna store that context so I can continue, continue to reuse it, right?
And Nvidia has this hierarchy, which Scott talked about. So step one, stored in high bandwidth memory, right? Obviously there's some challenges there.
It's really expensive and hard to come by right now. Uh, the other piece is it's local. So if I'm engaging in a conversation just on, you know, this one session, that's fine.
But if my friend is trying to have the same conversation, you know, do I have access to that memory? Then I can move it down to dram, right? Then I can move it to local SSD.
Again, everything's local. And then the next step is shared file and object, right? When you're gonna have a massive drop off in performance there.
East west traffic, all those different things we talked about, we shared nothing. So we were like, we need something right here, right? 5.
Something that has a performance closer to local but is more global in terms of its access, right? And that is essentially what I-C-M-S-P is, right? How do I create that local feel?
Now again, I'm not gonna spend a ton of time on this, but one way to do that is to basically take a shared nothing architecture, right? I can either put, you know, a client on the blue field or I can deploy my software, right? On essentially the CPUs in these g um, GPU servers.
Now again, problems there is I'm bringing the problems of that shared nothing architecture up into my most expensive assets, right? I have east west trap happening, right? I might have a hotspot, right?
Where everyone's asking about the same piece of context. That means I have all my GPU servers attacking one essentially and asking it for information. I don't know that that's a good idea, right?
So what we said is we've got, um, uh, a different architecture, right? We walk through this shared, shared, uh, everything architecture where I've got this stateless layer. Now, I would say, you know, some potential challenges with this instance of the deployment, really two, right?
And again, I think, I don't wanna say challenges, but optim areas for optimization one, right? I have this layer of CPUs that is essentially kind of in between my access to SSDs, right? Um, and as we know, right?
The CPU is always gonna be the bottleneck to SSD performance, right? You think about the world's most powerful processors. How many do I need from a thread perspective to saturate one single 122 terabyte drive?
It's a lot. So we have that problem. The other problem is, you know, I have to essentially create a copy of data, right?
I've got an RDMA operation to our front end, and then I've got another RDMA operation to the SSDs, right? So you kind of have this hop that's happening. What we're able to do with I-C-M-S-P on Vast is actually take our logic, our C node, and move that up to run on the blue field.
Now, we actually, uh, introduced a prototype of this style architecture, uh, I think with AI Field Day, um, earlier, but that was at the previous generation of Bluefield, right? Blue Fields have gotten dramatically more powerful from a core count. So now I can run my c no, the logic of the system up in those blue fields, this is a paradigm shift, right?
I no longer have a host going through other CPUs to basically get in line to get access to data that's on fast media. Now, every node has its own little friend. That's its protocol server.
Where's The storage class memory here, Phil? It's still in the dbox down there. You just can't, it's like, yeah, there'll be two layers of storage in that dbox a storage.
And the old Way the storage class memory was also in the dbox. Yeah. Nothing changes there.
It's just how I give access directly from the host to that device. Okay? Yep.
So we don't have to change anything there, right? Which again, now, instead of having to need to use the local SSDs and introduce potentially, you know, conflicts and all the different things that we might have by putting software on those servers, I can just have J bods right? Full of flash with dent solid.
Im SSDs and everyone has direct access directly to the metadata structure and directly to the actual data itself. And again, scaling and everyone sees everything. So there's no problem in sharing context, right?
If there's a hot piece of context, everyone can access it with a, a whole bunch of parallelism. Uh, but I'm not gonna have any hot spots up top. So Phil, excuse me, Phil, Jack Poller with Paradigm Technica.
It sounds like what you're really doing here is you are running storage controller software on the GPU because you've got spare GPU cycle. It's on the blue field. So I have a GPU server, I'm putting a blue field, which is like a smart nick now it's got 40 cores in it.
We're taking those cores, which you don't need for network performance 'cause it's just more cores than you'd ever would. And we're running our storage software there. So it's not in the CPUs, it's not on the GPUs, it's on this little server essentially that's mini, that's a smart nick, but a lot more powerful than that.
Okay? And the net effect of this is, So ultimately what you get, right? So we talk about certain things, right?
One much faster time to first token, right? So if someone's asking a question, I now already have that context. I'm sharing it globally, right?
So if anyone's asked about anything, so In, in, in a traditional architecture than you are making a request from the GPU to a storage controller that's off host. Yep. Right?
And then that storage controller goes fetch as the data feeds it back. Correct? And so in this case, what you're doing is you're moving that storage controller on host correct.
Or a little bit closer to the GPU. Correct? And that's getting you, that's accelerating significantly.
Significantly, okay. Yeah. And I think there's a few things, right, that come into play.
So you've got the acceleration of taking out an RDMA operation in the middle, right? Uh, you have a, a scaling advantage of the fact that now every time I add a new host, I'm adding compute specifically with that host. That is its own storage resources, right?
Essentially it's dedicated. So I'm taking out all the potential conflict, right? You get mm-hmm.
Resources for you. You don't have to fight over them with your partner. When I have that shared CPU pool, we're all fighting for the same resources, right?
So it's a scaling, it's a parallelism and it's efficiency perspective. The fact that the data is now shared means that I can have more GPU servers able to share more context. They're much less likely to calculate things again, right?
So that's why I get a faster time to first token because I can pull that context without having to recreate it. I get much better GPU efficiency because my GPUs are not recalculating the same things over again. They're actually doing inference instead.
Right? And then the final piece is I'm taking out that entire compute layer and that all is power that's drawn and power is precious now. So by taking out that entire server CPU group, I cut power by 75%.
Got it? Make sense? Mm-hmm.
Okay. So can can, can you tell us again what I-C-M-S-P was? It's context management something, something.
I think it's inference. Context management storage platform. Oh, inference context.
Memory. Memory storage platform is what they called it. And if you Google it, just be careful.
There's a whole bunch of other uses of the acronym. I see. The S-I-C-M-S-P just think of it as really cool close storage.
Okay. And I keep thinking about the image that you put up that had the different layers and had context as one of those layers. So, okay, so is this vast?
I-C-M-S-P, that's what's being attached to the blue fields. So essentially it's I-C-M-S-P is, um, think about it like Nvidia announced something called Dynamo, right? And we were working closely with them on that.
You've got like all these different problems in terms of, you know, how do I manage where inference jobs run on GPUs, right? How do I make sure that I am, uh, intelligently using the different tiers of me, of, you know, memory and storage just for context, right? So that's just for context.
Um, and then how do I scale and run those things? And basically the I-C-M-S-P is a tier of storage and essentially a standard way that Dynamo's gonna interact with that storage. So I'm basically saying we are lining up to saying this is how NVIDIA expects to use extended, you know, off, um, or shared storage for context.
So Can you go back to your diagram that shows? Yeah. Okay.
So where is it on this chart? So essentially this is gonna be used for context in this world, right? The notes?
Yeah. So all the da, all the, the, sorry, I'm, now, I'm not supposed to point to the screen. All of the context gets stored down in that dbox layer in the same way we would store any type of data.
Okay. And so that is y'all's I-C-M-S-P is gonna Be in the dbox? Correct.
Okay. Thank you. And the data lake that's also underneath that is stored there as well.
You Easily can. Yep. So, you know, something that I would talk about, 'cause I, I, let me go to the next slide and maybe it will, it will have, so one of the things that's going on right now around I-C-M-S-P and KV cache is the question, should you use any data services?
Right? Because could they impact performance? Right?
And we went through that whole shared nothing thing. And certainly if you do use data services on a shared nothing architecture, you're gonna have challenges. Data reduction is one of them.
Another one is something like encryption, right? So right now we're not sure should we encrypt that data. I think it's a really bad idea to not encrypt that data.
It's a giant shared thing which has everything about every conversation that all your employees are having with ai. You might want to encrypt that, right? And you kind of have to save it again.
But there's also a chance for data reduction from what we've tested. 3 to one and two to one data reduction, which means, again, and, And again, this is an inference solution, not a training solution. Correct.
It's all inference at scale. Um, so the point is, you know, in our world, this is how we would do it, right? But what's unique and flexible about the vast world is those blue field controllers don't have to be the only CPUs that that cluster has access to.
We can create actually a sidecar pool of compute that just does data services. Because remember, those blue fields are gonna write data down in let's say storage class memory. We can then have this pool of compute, take that data, reduce it, store it back down to the QLC, right?
But I can also attach other workloads over here, right? And what I know is that these blue fields all get their own dedicated amount of storage performance. Every blue field has 40 cores sitting.
So In the other prior solution, the blue fields were actually responsible for the d duplication. Yeah. We, without The other compute side Correct.
Cluster. Correct. Which they, you know, at this point, this is a very new solution.
We're gonna have to do a bunch of tests. Will they run hot? Will we want to augment?
Will it be enough? Uh, but the point is, we're the only ones that have this flexibility to say we're gonna add compute that you can then leverage for data services. So essentially, I just care about the fact that these guys can access data and then let our friends over here take care of all the data services.
In the old world, uh, these sort would've had to have been homogenous nodes. But in this environment you're taking, you could put any, any cluster of compute services out there to be your C nodes, correct? Mm-hmm.
At this point, it's very flexible. We have customers with multiple different generations of compute running c nodes in the same giant environment. Mm-hmm.
We can actually take, we can pool them, right? So we can say basically, you know, certain applications can use two C nodes and the rest can use 20. There's a lot of flexibility in how we can carve this up, that disaggregated shared everything piece gives us flexibility in a way that wasn't possible before.
And you mentioned storage class memory down at the Dinos. Those are, uh, different types of SSDs that are tailored to, you know, read, write access and things like that. Right?
So essentially it's an SSD, you know, lower latency, they're more expensive, uh, in much better endurance profile. Mm-hmm. Right.
So essentially, you know, when we came to market in, uh, Intel, I'm losing word pcm, you know what it is? Pc PCM O Optum. Yeah, Optane.
Optane. He, it died such a long time ago. It was all we had.
Soy has a, uh, P 58, 10 SLC based SSD that is used as a storage class memory solution for these guys. Mm-hmm. So there you go.
So yeah, it's all about endurance profile, cost latency. Thank you, Marian. I have a quick question.
Sure. Uh, going back to similarity, how do you keep the similarity decisions as the data and the models change? Yeah, so that's really just the underlying data structure, right?
So it's totally abstracted from models or anything else, right? So as data comes in, essentially we're hashing it. And normally with ddu you use a really strong hash, right?
That 'cause you wanna make sure that I never accidentally mistake two pieces of data for the same. That's why DDU uses a strong hash. All we're doing is taking a chunk of data and using a weaker hash.
And that weaker hash basically says that we don't use it to actually store the data. We use it to say, Hey, this data's very similar to data we already have. So it doesn't matter if that data came from an inference job, a backup job, whatever.
We will look across any data that's been stored on the system, um, doesn't matter what protocol it landed in. And we will identify that there's commonality and we just won't store it. And the system has no idea this is happening.
And you're using a weaker hash for, um, Comparison. Why? Comparison?
Comparison? Yeah. Because if you use a strong hash, you can't tell, because when I use a strong hash, a small difference in the data, a results in a very different hash, right?
That's why you do that. So a we hash basically says if the data's slightly different, then I'm gonna get the same result. So now we know we're in the zone, right?
That these two pieces of data are very similar. And what are you hearing from your customers who are in highly regulated industries? They have no issues with it.
Right? Um, again, data's now typically on a, on a system, it's chunked up, it's erasure coded, it's spread around anyway. So at this point, you know, it's all generally pointer based.
Uh, I haven't never heard anyone have issues that they would have to turn it off because of some kind. You Sure. Thank you.
As far as the KV cache, you're actually doing any special caching of the data coming off the blue field versus it's all going to storage class memory when it's written and it'll be re get studio to QLC or whatever the backend is. Uh, it's not like you're holding that data in storage class memory or anything like that. We Are not.
So, you know, in general, KV cash, we'll use the different types of media available. Right? So it could land in the memory, the high bandwidth memory.
Right. Or it could land in, you know, a local SSD, uh, this is essentially another tier of KV cache. Um, and for us, yeah, we're, we're gonna keep all that metadata in storage, glass memory, but we, you know, there's, there's nothing different about how we have to store the data.
Essentially, we've just created a more optimal data path.