Techstrong TV February 4, 2026
Watch our live stream Monday through Friday, featuring exclusive news, announcements and conversations with IT leaders and experts on topics ranging from digital transformation to #DevOps, #Cybersecurity, #CloudNative, #Containers and deep-dives into specific technologies and best practices. http://techstrong.tv/
Transcript
Hey everyone. Welcome back here to Tech Drunk tv. I'm, uh, happy to have our next guest on, he's the founder and CEO of a company named Etle.
It's his first time here on Text tv. Let's welcome Christian Ramming. Ramming.
Is that correct, Christian? Yes. That that, that's right.
Yeah. Thanks. Thanks a thanks so much for having me.
Alright. Got it. Right on The first try.
It's a good day. Christian, welcome. Thanks for being here.
As I mentioned, you're the founder and CEO of that leap. We're going to get into that leap in a second, but let's, let's start with you, you know, where, where was, what, what drove you to found Etle and kinda where have you been and you know, that brought you here? Yeah, yeah.
I, uh, I'm an engineer by training. Um, I've, I've been know, been, uh, living the silicon battle life for the past couple of decades now. And, uh, I, uh, I started atle about, uh, a decade ago.
Um, uh, I was CTO at an ATech company before that. And, uh, that's really, you know, there, there's, there's essentially an, an itch from there that I'm scratching with atle. It's, um, um, you know, we, uh, I was, uh, you know, heading up engineering and analytics and, uh, I hired, had a old little engineering team just dedicated to building data pipelines.
You know, we needed to make our analysts and scientists productive. And, uh, that just ended up being a tremendous amount of engineering work. So, um, so that's a, you know, the, the short version of this story is that, uh, I wanted to take all that engineering work out of, um, you know, something that companies needed to do.
And, and so that's really what we're doing at atle. We're building software to, uh, take all the data engineering work out of, uh, data pipelines. Interesting.
And you say you started this a decade ago, so that was before, you know, chat, GPT, generative AI and all of this stuff that's going on today. Did you ever think a decade ago that what you, what you were setting out to do would one day perhaps be enhanced or, or augmented by something like Gen ai? The, the world is so the short answer or no, the world is a crazy place.
I mean, I used to, when I started, I used to have to explain to people what a data warehouse was, you know, so it's not, yeah, I, I, there's also true, I, I know logo has to do that, so yeah, that's, it's come a long way since then. Absolutely. So, you know, for those of us who don't know at Leap give us the last 10 years in a nutshell, Christian, We have, uh, a software product, um, or we've been offering a software product that works great for, uh, data warehouses, uh, in the cloud, right?
Redshift, snowflake, Databricks, and so on. And so what it does is it ingests data from, uh, a bunch of different types of sources, right? Databases, applications, streams and files, uh, brings it all into the warehouse and then makes it easy for data teams to sort of model, uh, data to, to turn it into data products, um, within, within the data warehouse.
And so, uh, so yeah, so we've been on a mission to, to, uh, reduce that data engineering effort and also reduce the time that it takes to, to get all this set up right. Um, that's, that's another big part of it. You want to, you know, get your, your analysts and scientists producted, right?
Uh, as, as quickly as you can, and, um, you know, have your engineers spend time on more, more interesting things than, uh, than data plumbing. Um, and I, I think something that's, that sort of set us apart in the, in the marketplace is we, we work with, um, companies in regulated industries, right? Financial services companies like Morningstar, and in healthcare, right?
Companies like Moderna, um, they need to, uh, keep their data within their own firewalls, essentially, right? Like within their virtual private clouds. And, and so that's, that's something that our, our, uh, our software does.
Excellent, excellent. Um, you know what, before we jump into Iceberg Pipeline platform and all this, let me close the loop on that leap for people who maybe say, sounds interesting, I'd like to learn more. What's the best on-RAMP forum?
Yeah, I mean, uh, uh, it's, it's pretty straightforward. com, uh, sign, sign up for a demo. We have a, we have a great team that's, uh, uh, you know, ready to, to, to, you know, show you exactly what you need to, what you need to see.
Excellent. Cool stuff. Alright, let's move over.
Now, you guys recently, uh, introduced a, a new, uh, platform, the Lib Iceberg Pipeline platform. Yeah. Uh, that, that's right.
So, um, I guess maybe a little bit of, a little bit of background. First, we, you know, we, we, something we've, we've heard a lot over the past few years is, you know, these data warehouses are great, but, um, you know, for, for different reasons, maybe for data sovereignty reasons, or because people are adopting these new AI use cases, um, they, they want to, um, take on Apache Iceberg, basically use data, uh, use Apache Iceberg as the data foundation. And, uh, for, for, for data pipelines, there's a bunch of consequences.
When you, you, you want to run pipelines with Iceberg, there are things that really should work differently. So what we're, what we just launched is, as you said, the, the Iceberg Pipeline platform. Um, it's basically a, a pipeline there that's built from the ground up to support, uh, to support Apache Iceberg and everything that's different about it, right?
So it's, you know, still doing ingestion, still doing modeling, but, uh, but also tying that all together, right? The, the dependencies between the different parts of the system are, are become really important. Um, use cases like streaming, which I'm sure we'll talk more about, are, um, uh, you know, are now possible.
And, uh, and also there's a lot more to do, right? You need to do table maintenance and, uh, um, you know, snapshot removal and, and, and that sort of thing. And so, uh, the Iceberg Pipeline platform is essentially a and sort of an all in one system that ties all of that together, uh, and can, and can run inside of your own virtual private cloud.
Excellent. Now, of course, this is based on Apache Iceberg. That's right.
Let's say someone out there says, iceberg Titanic. I never heard of Apache Iceberg. What, what exactly are we dealing with?
Tell us, you know, if you wouldn't mind, give us a little Apache iceberg background. So, Apachi Iceberg has been around for, for several years now. Um, you know, it's an open table format that's particularly well suited for, um, uh, you know, analytics use cases.
And, um, you know, it's not the only OpenTable format out there. Um, it has, uh, that there are, there are other ones too. But I think that the thing that's special about Iceberg is that, uh, a few years ago, the, you know, the, the, the massive vendors in this space, right?
Snowflake, Databricks, AWS Google all essentially decided to throw their weight behind this one, right? And so they said, look, we, we, we think this, this technology shows a great promise. We're gonna make our products work really well with this specific format.
And so that, that really helped, uh, future proof it, right? Because now it became, uh, much easier for enterprise to say, yes, okay, well, the, the, the big guys are behind it, we can adopt it too, right? And so that's, um, uh, you know, so that, that's, that's a, a, a big part of the reason why it's now seen as a default choice, right?
So, um, so, you know, the more concretely right, you can, uh, store all of your data within, uh, the iceberg, uh, uh, table format, um, which is, you know, built on top of Parquet, right? Parquet files that live in S3 with some metadata layers. And then you can, uh, you can use tools that you already use, right?
Snowflake, Databricks, Athena, uh, BigQuery, right? Um, these engines work really well with the data that's stored in Iceberg, right? So you can, uh, uh, you know, you can get, you can still get your fast queries, you can still support your, your AI use cases, right?
If you're using Bedrock, for example, that plugs in very, very nicely. So you, you know, you own your own data and you continue, can continue to use the tools that you already use in your stack. Love it.
Excellent. And you know, the nice thing is we, we see it with the foundations like Apache, like Linx Foundation. Eclipse Foundation is, it, it fosters this, um, cooperation where companies that normally wouldn't necessarily cooperate Snowflake and Databricks, right?
You don't see them around the campfire singing Kumbaya too often, but, you know, they all, they're all behind and they're all supporters of Iceberg, right? As, as something that transcends individual companies like that. Um, is, is the Iceberg Pipeline platform available now?
IPlan platform is available. We have, uh, customers using it successfully in production, and, um, uh, as I mentioned, you can, you can run it inside of your, your VPC so that, you know, no, no data's flowing out. Uh, you, you know, while, while your pipeline's running.
Very cool. Um, you know, while Apache Iceberg is free and open, I'm sure, uh, the, the, the, uh, excuse me, the, uh, pipeline platform, is it open source, not open sourced SaaS kind of thing? How do, how do you do this here?
So there are, um, it's, it's, uh, proprietary software. Um, we, uh, integrate very tightly with, uh, you know, other, on top other vendor, vendor space. Yes, exactly right.
So it's, it's very tightly integrated with AWS, right? So if you have, you know, if you have a glue catalog, right? If you have, um, you know, you use Athena, um, if you have, you know, CloudWatch for monitoring, right?
It, it sort of plugs right into that. Excellent. And then did we mention, I don't know if we, did I forget, did we mention the, uh, ATLE website?
com? That's right. com.
Yeah, lots of, lots of information about the, uh, uh, the iceberg integration there as well. Um, so you can kind of see, see how it's different. You know, I think that, uh, one of the things that sort of sets this apart from, um, from sort of traditional ETL, right?
You could argue, oh, isn't this just, isn't just just, you know, ETL in a new, uh, you know, for the new, with a new sort of destination. And, and I think this is, uh, you know, it's really, this really is something else. You know, ETL tools move data, but, um, but there's sort of a fundamental difference with operating pipelines end to end, right?
There needs to be, uh, you know, tying together ingestion and modeling. There are, uh, you know, catalogs involved, and also these downstream tools that need to keep their own metadata, uh, fresh about the table. So in order to kind of get a, a, a system that sort of hums end to end, you need, you need, you know, whether it's atle or something else, you need something more integrated than a, than a sort of a traditional ETL tool that will sort of dump your data into iceberg.
But, but that's, then that's sort of it. Fair enough. Christian, we're about outta time.
I want to thank you for popping in here, man, the time goes very quickly, as you could tell. Um, before, before we end anything else, If you're gonna, if you're gonna remember one thing about Etle, I guess, um, you know, we, we, uh, handle ingestion modeling. Uh, we operate your, your pipelines end to end, and we run inside your VPC.
Thanks. Uh, excellent. Thank you.
A really appreciate it. I appreciate it. Christian Rumming, founder, CEO of at Leap here on Techstar tv.
We're gonna take a break. We'll be back with more. Stay tuned.
Hey guys, thanks for the throwaway here with Peter Lee's, the CEO of sim space, and we're talking about, well, cyber ranges as it applies to cybersecurity and training, and maybe just modernizing this whole process. Peter, welcome to the show, Mike. Thanks for having us.
You know, I just saw recently somebody going through some cybersecurity exercises, and it was a tabletop, kind of a version of a tabletop game, and everybody had their role to play. And, you know, as I was watching, and I couldn't help being reminded of, you know, a clue, you know, Colonel Mustard did it in, uh, a library with the candlestick. And as fun as it was, it didn't seem very realistic.
So, um, what should we be thinking about here when it comes to training and modern attacks and resiliency? Is there some way to think about all, all this that's just gonna be maybe better than what we have been doing, because, well, I just don't think the bad guys are sitting around playing table talk games to attack us. Yeah, you're absolutely right.
Well, first of all, I think the whole notion of training has to get redefined as humans plus AI working hand in hand. It's no longer just humans. And I think, um, I've been shocked at the weaponization of ai.
I mean, it is unbelievable today. There's a lot of skeptics, still people who believe it's at the early stages, and that may be true, but we are moving very rapidly into an age of autonomous agent attacks at scale. You know, they're gonna do reconnaissance, vulnerability, exploitation, moving laterally, you know, the whole full, uh, kil the full, um, full value kill chain.
And so I think it has to be humans plus ai. And I think that that notion of training is critical. And I think a cyber range is the cornerstone of rebalancing the fundamental asymmetry of offense, cyber offense, and defense.
Right? It's always cheaper, faster to attack than to defend the attacker chooses when and where the defender has to prepare everywhere. Yeah.
Um, so that's, by the way, that's half the story, Mike, and we can get into the other half, which is testing of tools, agents, capabilities. So yeah, as one old boss of mine said, you know, it's a lot easier to throw grenades than it is to catch 'em. But, um, when you think about this for a minute, uh, I'm not sure everybody knows exactly what a cyber range is.
They're familiar with range as in like the military has an ex exercise. They go out to the range and blow stuff up, and they learn how to do things. But how does that concept apply to cybersecurity?
Well, I think we should define the cyber range. First of all. Um, a realistic and intelligent cyber range is going to be one that enables you to create a realistic replica of your production environment.
So this isn't gonna be some pre-canned laboratory. It's going to be a very rich, um, replica with full network topology. It's going to have integrated attack and activity emulation with your actual security tools and user behavior, which is chained in some ways, and interactive with attack behavior.
You're gonna have a very significant, um, we don't like to use the word digital twin because of that implies, um, a layer of cost and complexity, which is diminishing marginal returns for what I think a cyber range, um, can provide. But yeah, absolutely tabletops, simple labs, individual training, these things are obsolete. They're not gonna move the needle.
You absolutely need to have a realistic replica of your production environment. And I'd say, as I opened with my comment, not just for humans, but increasingly certainly what we see in our business is that, um, vendors are using it to train AI models, to test AI agents and to validate agentic capabilities. And you can't do that in an environment that doesn't closely, um, mirror your own enterprise complexity.
It's gotta be like any other type of agentic or, um, AI modeling. Gotta be very, very, um, tailored to your specifics. And how easy is it to do that these days?
I think people think that there's a level of complexity involved that's too challenging. And then how do I test those AI agents once I set that thing up? Because, well, every AI agent I've seen so far does things that are unpredictable.
Yeah, those are very good questions. Uh, so on the first one, I think that, um, there the answer really depends on the degree of fidelity that you want for your, for your environment. So we can have, um, realistic replicas of production environment set up in a matter of a couple days, or it could take several weeks depending on the level of fidelity.
And that'll depend on the use case as well. Um, in terms of the testing of agents, it's really important to step back and zoom out. One of the things that's really, um, critical in terms is the ability to get clean labeled data to be able to actually run full kill chain threat attacks in varying user environments with every single conceivable permutation of your network topology.
And the different tools you use, be able to collect and synthesize that log data and figure out, did this agent actually work? Thumbs up, thumbs down. And you can imagine doing those training runs thousands and thousands of times in order to begin to bring that capability into bear.
And of course, you could think about dual, you know, traditional data science measures precision and recall, am I really tuning it so that it absolutely is gonna detect every single type of threat, but it may throw off a degree of noise, or am I gonna reduce the noise, but I'm gonna let some of these threats through. So there's, I think there's some significant capabilities and, um, tailoring and customization that are capable in the cyber range, um, uh, uh, platform. But that's going to be user and use case dependent.
Mm-hmm. Have we gotten to a point where we can't quite just put up a standard red team because the number of vectors that could be exploited are just too broad and there's too many things that could possibly happen, and who knows, maybe we'll use AI agents themselves to create the attacks. But as one security expert said to me once, if you can imagine it, somebody's probably trying it.
So how do we deal with all the exponential possibilities? Well, I think v on your original question, can we, uh, are red team still useful, let's call it, or, or rely on? Absolutely.
I think that, you know, again, I, I would go back to my, my touchstone of humans plus ai. There are things, certainly, if I can speak on behalf of our red team, that they're capable of doing that are extremely sophisticated and would, um, would not be, uh, available for agent training. They haven't, the data doesn't exist, the penetration methods, the tooling that they've developed.
That being said, I think what we see is we have actually, um, a fairly expansive set of ecosystem partners. There are a number of red team agent businesses that are beginning to, um, develop their capabilities in our cyber range. And what they wanna do is they wanna partner with us.
So as we run these realistic, um, training exercises, team training exercises, they're going to be able to, um, have a menu. You could think of them as a, to offer capabilities against the, uh, for the enterprise clients to test, test their teams and tools. And so I do think that that is going to be the future where you're gonna see a variety of different, um, scenarios.
And I think the, um, the cyber teams that are really focused not on compliance and check marks, I'm gonna say not on individual training, but on realistic team training that actually helps to outsmart adversaries in any cyber terrain. I think you'll see a tremendous amount of, um, interest in both red team and blue team agents. We're seeing both, um, both parts of the ecosystem get attracted to our, our cyber range, um, and of tremendous interest to clients who want to do team training at, at, uh, at realistic scale.
Some folks are concerned that perhaps we're a little bit over our skis when it comes to AI agents, and we're deploying these things without thinking through the security issues. Um, are you at all concerned that maybe we're just waiting for some sort of catastrophic event before everybody gets serious about AI training and testing? Mike, it's a great, it's a great question.
I think one of the other elements of that, what we are actually doing is to help, um, validate the, uh, AI agents themselves. Anytime you introduce a new capability, you open up a new attack surface. And the reality is that, um, bringing on a blue team agent to help with your SOC preparations, but opening up new attack vectors without being prepared for it would be a terrible mistake.
And so we are actually creating kind of an agent reading system or validation system, and we're, um, taking a hard look at the, um, at the vulnerabilities that the agents present themselves is a fantastic question. Mm-hmm. Do you think on the other end of this, that auditors might start using these platforms to test and validate the environments that they're being asked to vouch for?
You know, that's a fascinating question. We actually see, um, it hasn't yet proceeded to the auditor level, but certainly the cyber insurance providers are taking an active interest. And we've had several of our enterprise clients let us know that we've been instrumental in reducing cyber insurance premiums because they can actually demonstrate progress on a longitudinal basis against the varying stage and progression of threat actors.
So I do think that the insurance market is going to be the leading edge, but I would not be surprised if we see, um, multiple other, um, uh, constituents take advantage of that. No doubt. So as you look at it and you think about testing and training and everything that goes with that, what's that one thing that kind of just makes you shake your head a little bit and go, folks, we need to be a little bit better than that?
Oh, yeah, sure. That's an easy question to answer. I think if you look at our large enterprise clients or large government clients, they are literally using hundreds of tools.
They are never going to be optimally configured. They are in a continuous, what I call process of selection and optimization. Many of these tools have overlapping vectors.
Um, they're continuously evaluating and adopting new tools, and there isn't a programmatic, um, value based, efficacy based way to actually, um, frame decision support around how do I configure and optimize and rationalize my tools in order to improve my security posture. And so I think that to me is just a continuous, um, cyber range use case that we're, we're really excited to, um, be engaged with our clients around. All right.
Hey, folks, you heard in here if your training and testing is rooted in the last century, you're probably not gonna get the outcome you're hoping for. Peter, thanks for being on the show, Mike. Thank you for having us.
All right. And back to you guys in the studio. Now, as you continue to move to your new infrastructure, um, you've got, um, you, you have production nodes and, uh, all embedded on the new, the, the new operating system, the new management platform.
What are you doing differently now, carrying forward? You know, of course we've talked about digital twin, uh, to a certain extent, but, uh, you know, you've talked about integration with CICD processes and so forth. Yes.
What do, what do day two operations look like for the whole team now on the new platform and infrastructure? Well, currently, you know, um, on the new infrastructure we have, uh, the team, now it's becoming more, um, uh, you know, date, the date, day two plus is becoming more, more routine kind of changes, uh, as we are of course moving the, migrating the data centers and, and building our existing data centers on. But, but also from an architecture perspective, we are now trying to do more in terms of, you know, implementing more AI functionality.
Mm-hmm. To, to now to tell us more about the operational, you know, if something goes wrong, if an incident happens, uh, don't give me row row information, even if it's correlated. Even a correlate, give, gimme an RCA, gimme an RCA, and point me to what, what could be the problem and what could be the impacts.
So may maybe something has happened, but actually there is no real impact to the network. So a link might fail. Mm-hmm.
But there's enough in redundancy that we shouldn't worry about. Sure. Tell, tell me.
I don't have to worry. So we're trying to build this more of an intelligent intelligence that there is some entity that is entity, uh, that is aware of our topology. Sure.
Our topology Yeah. And our network that can actually give us useful information about incidents and the results. But the main, the main thing that's interesting also in the operational aspect is I, I, I talked to the operational team and, you know, they, they, they have a huge reduction in, in incidents, you know, like, like 80% to what we had before.
Yeah. Not, not only, not only the tool is part of the aspect, but the, I think the, i, the, the NetOps and tracking and, uh, this idea of doing things before you implement them helps a lot in avoiding, uh, avoiding problems. So, so the, the operations team are saying that, most are saying to me that most of the time they, they don't find the same problems that with what they found before, the fabric is very stable, very consistent.
Stable, consistent is as, as expected. There is no, nothing that surprises them. Most of the things they're dealing with is maybe access to, to servers, uh, low balancers, firewalls, physical, physical stuff that is going on in the data center, but not from a, a design or, or a routing or, or this kind of, uh, perspective.
So it shifts, it shifts their, their, their efforts, you know, away from debugging, you know, lower level Sure. Details that the fabric should, should take care of. So it's a, it's a different, different type of challenges, but much, much less than before, much less before.
So that's a great example of the tooling decisions that you make, really having an impact on your operations team. Right. And you're dropping 80% of your tickets.
Um, that's, um, and, and I know, um, you know, your leadership has, has cited the same statistic, you know, 80% reduction in tickets. That's, that's huge. That's certainly impactful.
Didn't you also make some decisions on how to, um, implement EA, like you could have gone with a local install for EA, but you opted to go with the SaaS option. What were, what were some of the thoughts you had to, to weigh there and why did you ultimately choose the, you know, I a SaaS version? Okay.
So, so the E SaaS, really, uh, we're, we as a network and, you know, network team, we're not in the business of running edda, you know, or, or we want to make use of ed. Sure. We want to architect around it.
So, so we offloaded that, that effort to the EDAs team where they, they can implement it in the cloud, whatever, what, whatever cloud they want, they want to go to, sure. To GKE. They want to go Google or AWS whatever they want, right?
We don't care. We just want the connectivity to it. We want, you know, the upgrades to happen.
We want them to monitor the health of IDA itself, you know, uh, maybe their, their, I comes with many extra add-ons, you know, with RCA, um, you know, correlation. Uh, they do some kind of AI into, into giving you, in, you know, analysis of anything that happens. So, so that's, that's also a plus that comes with it.
We really didn't want to, to think about IDA too much. We wanted to rather use IDA architect around IDA and, you know, think about our, our architecture more and building more enhancements to the solution itself. So EA SAS in that perspective, uh, saved us a lot of effort.
And, and the team there was very helpful for us. So we, we liked that. We, we, we were, we were glad we went that way.
Not in the business of actually running edda. You just want it to work, right. So that's a great Yes.
Encapsulation of that. Very good. Yes.
Well, what else is in store in the future here? You know, you've, you've laid some really important foundations here for, you know, making real improvement for the resiliency of the network and the lives of the operators. Um, what, what's coming down the pike?
What, what are you gonna be able to do in 2026 that you couldn't do in 2025? I think, I think now we're con concentrating on enhancing, enhancing the solution using maybe more AI that is native, you know, native to either, either is coming up with lots of capabilities that mm-hmm. Are AI related, that we can chat, we can chat with either, we can talk to it, we can tell it what went wrong, what, what's happening.
So that part, you know, we used to actually, we build in our solution, but a way or around either to, to, to fill in that gap. But I think in, in the next release, we're gonna be merging some of our tools with IDA and maybe merging our automation into IDA itself. IDA has the capability to write your own apps.
Sure. We would, I, I think, think that's something we would like to do is, you know, bring in our, our automation and pipelining to, to, to be within ida, to be within, to be ada IDA native. So the user doesn't really have to think, I'm dealing with Git, I'm dealing with ida.
Maybe we can merge something together there. And also more on the operations aspect to get, to make operational RCA more intelligent, which either will add, and we'll see what it adds and try and fill in the gaps there. As well as, as you know, we're still Nokia's growing.
Mm. We've got, we've got acquisitions and stuff like that. So there will be more data centers.
And, and as I said, our, our architecture now is very modular. It's not flat. It's not like, it doesn't look like a data center.
It looks like many small mini data centers connected together. Sure, sure. So we are able to absorb, absorb those changes.
And that's, I think will be, uh, main task force next year. Ahed, thanks so much for talking with us today. It's been a pleasure.
I look forward to hearing how things progress in 2026. Thank you very much for having me. It is very exciting for me, and I'm very passionate to, to, to talk about our work.
'cause it's, it really is exciting and interesting. It comes through in every conversation I've had with you. So thank you.
Thank you for sharing that passion and turning it into real concrete. You know, this is what you're actually experiencing in the new, uh, the new network infrastructure in, um, Nokia Enterprise. It.
We are super excited to be here, uh, presenting to this delegation. This is actually our first time, uh, at Tech Field Day. And, uh, my name is Ish Reker.
I am the chief marketing Officer here at, uh, fabric Study ai. And, uh, before this, been in the Valley for almost 30 plus years. Uh, one company in particular was Waca, uh, before this, where I was the head of AI and Strategic Alliances, worked with guys like Fredrik and others, right?
So, understand the infrastructure side of things for AI extremely well. And there's a lot of action happening today. But what I'm going to talk about is something on top of the AI infrastructure today, right?
So this is an agent platform, uh, which I and a colleague Rashid will be talking. So we'll have it, uh, kind of divided into two sections. So I'll be giving the company and the product introduction as well as, uh, you know, Rashid will go do a deep dive and then get to the demo.
Okay? So with that, this is kind of the time allocation we have. We have a section for q and a, but feel free to ask us questions as we go along.
If there are pressing questions. Uh, this is, uh, extremely new for all of us. Okay?
So we all learning as we go along. So who is fabric? So we, uh, used to be cloud fabrics, uh, dot com, right?
For last, uh, eight years or so, from and go to market perspective are very well recognized in the AI ops space. We were an observability and AI ops player to begin with. And then, um, starting last, uh, beginning of the last year, right, January, we rebranded ourselves to fabric start ai, right?
For a reason. Obviously, the use cases had changed with Agent Tech coming on the horizon. Our customers were asking us to kind of get new functionality, right?
And that was really the genesis to fabrics ai. So again, um, we are headquarter here in the East Bay, uh, with, uh, you know, r and d office in India, satellite offices across the globe. Uh, well, very well known in the analyst community, right?
And obviously, as you guys know, agent is all new. Nevertheless, we did get featured into two prominent Gartner publications. One, we just actually came out this, this week with Arun.
Okay? So we'll talk about that as we go along. So, fabrics, uh, the founding team, right?
So we have our CEO and the CP in the back, I Rashid here, but this is kind of our fourth startup, the founding teams, right? So we have had successful exits before, three of them, uh, right? So we've been in this business of actually kind of building companies, uh, and startups, which are kind of morphed into larger business units, servings, tens, and thousands of customers.
Okay? So this is not our first rodeo, okay? So again, uh, for fabrics, uh, uh, in the AIOps days, we had some real large market customers that are also now good amount of large customers, uh, who are actually in production.
Some of them are actually doing POC, okay? So you can see, uh, fortune 500 customers, uh, right, MSPs and CSPs, communication service providers. So obviously, because of our pedigree, we, we tend to work with, uh, the likes of Cisco and IBMA lot, right?
Uh, and that's where the, the partnerships come from. And then expanding to other verticals beyond IT ops, right? So our classics, uh, kind of, uh, sweet spot, was it operations, knock ops, right?
And AI ops, right? But because of agent ticket kind of enables us to kind of cross the domain across in Bops and SecOps and so on and so forth, right? Because we are essentially collapsing at the data level and the application level, okay?
And we'll talk about that as we, we go along. Okay? So let's kind of walk, Hey, why agent take, when, when, you know, we are an AI ops player.
So on the left hand side, you'll see this, right? So the classic workflow, uh, in an AI ops environment. So you have an application and infrastructure to, uh, stack, which is sending signals, uh, to the AI ops kind of, uh, stack, right?
Or, or, uh, basically the backbone, which is essentially leveraging machine learning and to some extent, uh, you know, deep learning to basically do correlation, root cause analysis, right? But the process for remediation is, is primarily manual, right? So now what you are doing is what I call the last mile problem, is you are basically exposing dashboards, you're exposing incidents, and it's up to the domain experts to interpret that and kind of formulate and remediation path, right?
So heavy kind of, uh, intervention with, with, uh, you know, uh, with humans in there, right? So from there, as, uh, you know, the customer started asking, and, and again, this, what drove us actually, uh, to the agent ops was actually a customer environment, right? So we had a telco, uh, which AC suffered, you know, an eight hour of outage just because somebody inadvertently did ACL changes, right?
Access control list changes. And it, it took them almost eight hours to find the problem to ate the problem, right? And these are perfect use cases, which can be remediated with agent ops, agent K ops, right?
So what's the new stack? Like, we still have that application stack we interface with, we have signals coming in right now, the LLM is at the center of reasoning or the decision making, right? And we'll talk about the, the good and the bad part of that as well, right?
Because it doesn't come free, right? I mean, LLMs, at the end of the day, they are non-deterministic processes, right? But what they enable us to do is, now you can have workflows which are semi-autonomous or autonomous, right?
Or why call semi-autonomous is you typically customers till they build trust, they will have human in the loop, right? HITL, what we call, uh, the, the, the LLMs themselves can do deep research, right? So you don't have to depend on a domain expert saying that, Hey, have you gotten my networking logs, or have you gotten my metrics from the application, right?
The LLMs are good in doing what we call chain of thoughts or graph of thoughts, right? Uh, kind of reasoning, okay? Uh, and then, uh, it's cross domain.
It can talk to a multiple, and especially it's true with us. Uh, this is one of the major prerequisite for an a successful agent platform that you have to be able to connect to, you know, a number of different endpoints, right? Because in IT ops particularly, you have storage your network, you have applications, you have a PM tools, what have you, right?
And then, uh, the in-place data access is also something which is very unique to us, and we'll kind of talk about it, uh, uh, as we go along. But each of these tools, what we interface with are domain specific tools, right? And they don't want to integrate with each other.
But by agent taking what we have in terms of what we call dynamic tooling, we are able to get to those endpoints and data sources, right? Directly, right? And we'll talk about how we do that, okay?
So, uh, that's kind of the setting the stage on how we are transitioning from AIOps to agent TK iops, right? Um, now having kind of shown the promise of agent take, right? There are also challenges to agent because the primary kind of LLM, what you're depending on, right?
The reasoning, uh, LLM is actually, uh, a non-deterministic LLM, right? And, uh, it, it's what we call stochastic processes, right? But in addition to that, there are other elements also which contribute to the, to the hallucination, right?
So particularly there are three main kind of, uh, you know, reasons what, what we kind of, um, articulate as, uh, the reasons for hallucination. So first and foremost, context and data management, right? So context is the king, as they say, right?
So the output tokens, which an LLM pro kind of puts out is directly proportional to the veracity of the input tokens, right? That means simply put that context is how much curated context you are providing to the LLM really determines your accuracy, right? So that's the most important thing.
And you'll see our key IP being, uh, that middleware, which comprises of context engine and what we call dynamic tooling, right? Also data management at scale, right? So we are not talking about, you know, an experimental demo.
It's very easy. It's one thing to do a demo, right? Where you have five devices, it's another thing when you are in a real life production environment, you are getting 10,000 rows of SQL, right?
You're getting large logs, you're getting metrics, right? You're interfacing to almost hundreds of devices, right? Uh, how do you manage all of that, right?
So that's where the data management part comes in. How, I mean, for, for MCP, as you guys know, right? The model context protocol is very transformative, right?
You don't have now EFLs then in your code anymore, right? You're essentially handing over a manual to the LLM, right? Hey, these are my tool sets, and out of this, select what you feel appropriate, right?
But at the same time, it also comes at a cost because every conversation has to go to an LLM and back and forth, right? So how we have simplified that, uh, is also, uh, our key ip, lack of connectivity. As I said, we connect to almost 1700 plus data sources, right?
And even beyond that. So that's very crucial when it comes to IT ops and AIOps, uh, use cases, right? And finally, agent tops is where, you know, hey, having a platform is one thing.
Operationalizing is, is another thing, right? And what I mean by operationalizing is observability security, right? Explainability, right?
Trust governance. How do you bring all of this, uh, where the customers you're building in the trust with the customer, okay? Uh, and that's what really, you know, all these studies, what you are seeing about MIT or Cisco talking about, hey, there's a lot of noise, but there's only 5% of the value, right?
And I call that as agentic value gap, right? So you're doing a lot of experimentation and demos, but it's really, you know, very little of enterprise use, uh, coming out of these agent take use cases. Making sense so far, folks?
Any questions, comments? I'm waiting for you to get to it because I have a ton of questions, but I'm waiting for you to kind of lay it out a little bit More, okay? Sure.
Yeah. So, we'll, we'll get to the, to the gist of it, right? So again, if you see the progression with agent take applications, right?
So just like with the three tier application, we introduced a middleware layer, right? Uh, for complex application scenarios. Similarly, with agent take, what we are seeing is most of the applications today are the, the, the two tier one where the AI agents and LLMs are directly talking to MCP endpoints, right?
So MCP has to be unable for the endpoints because that is where you get to a structure which the LLM can understand. Now, what happens when there is data sources, which are not MCP enabled, right? Which is pri primarily telemetry.
You are, you are reading raw data or you are reading from legacy devices, right? So that's a pitfall. That's another reason why, you know, what leads to hallucination on the right hand side is our solution.
So what we have introduced is what we call this fabrics. Middleware, right? And there are two main components to this middleware.
So one is what we call the context engine, and the second one is what we call universal tooling, right? And the context engine in simpl, in simplistic term, what it does is it provides curated data to the LLM only gives it what it needs. It's not dumping everything, right?
Because that's where the, the LLMs you, you are overriding the context. You are corrupting the context and the LLMs hallucinate, right? And the second part, the universal tooling is taking care of how do you dynamically connect to this disparate data sources, right?
You get that data schema and you normalize in a way such that the LLM understands, right? So tho that's the key ip, which we'll be talking about. And what does that do?
It allows us to talk to obviously any MCP data source, right? If you have an MCP data source, that's great, right? And we'll provide you still all that context, um, capabilities and so on.
But also any API, right? So any API, uh, based data source we have, we will actually create that, uh, MCP wrapper on top of that through the middleware, right? So, which is huge as well as raw data and legacy systems, right?
So just to give you an example of this, right? What I mean by this, right? So say my intent, right?
Uh, I am providing an intent to the platform saying that, Hey, I want to get metrics from this application servers. Now, it so happens that those application servers are monitored by Dynatrace as an a PM tool right now, uh, my LLM and my, my system goes and checks that, hey, uh, there is no tools MCP tools available for, uh, getting the metrics from Dynatrace, right? At that point, what we'll do is we'll launch and tool handler a tool creator, which will actually scrap the internet.
It'll look for publicly available Dynatrace APIs, right? And it'll formulate the tools, four or five tools, right? Which, uh, essentially are, are are there to talk to the di Dynatrace controller, get that metrics right?
And we have demo, um, demonstrating that, right? So this is what I mean by dynamically creating, um, you know, um, the, the ability to go and talk to data sources, which are not MCP enabled. Making sense?
So you guys sit at the data layer, the data acquisition layer of an AI process, right? Above the data acquisition, the middleware sits on the above the data acquisition. There is a data fabric which integrates with the data sources, Okay?
Yeah. And what y'all do is you take, uh, lemme see if I get this right. So you, you somehow provide the data or clean the data, get it ready for the MCP, but you also work with data sources that aren't from, that aren't involved with MCP to get them prepared to go into the data set.
That'll be used to train a model. Yes, absolutely. Right?
So in short, yeah. And we will also, uh, integrate with that. That's what we call the dynamic MCP, right?
And you're correlating the data from all those sources so that I understand, you know, what the information in one system means to, to the other, and how it's giving you the full picture. Yeah. That's, that's further down, right?
So now we are sending it all to the LLM. Mm-hmm. Right?
And which is actually correlating and doing the root cause analysis and so on, right? And we provided a template, which we'll talk about, right? We'll give it some examples that, hey, this is how you should go about doing this, right?
Isn't it Ray Esei, Silverton Consulting, isn't there a challenge here with normalization of the data coming from different data sources and trying to understand the range and, and those sorts of characteristics of the data you're getting? Yeah, Absolutely. And that's where our, our value prop comes in.
So the middleware actually takes care of that normalization, right? It normalizes it in a way where the LLM understands our structure. Because if you give LLM unstructured a str, which doesn't have structure, is bound to ate.
So that's, that's really the value prop. You're, I'm also interested in, in how your context engine manages the context. 'cause context, uh, growth is a major challenge in most of these things.
Yeah, Absolutely. So we'll be going actually really deep in, in, um, rashid's section on that context engine. But just to clarify, right?
So this, I mean, again, we come from that, uh, world, right? So what we are talking about is not the size of the context, right? There is a lot of, uh, uh, innovation happening with KV cash and all of that, right?
Mm-hmm. We are talking about the purity of the context, right? So it's not about, obviously size matters as well, right?
But now a lot of these LLMs are coming out with a lot of bigger, uh, context memory, but really it's, uh, uh, it essentially boils down to how pure your context is, right? Because that's how the, the output is going to be determined. Oh, pure, yeah.
Purity, well, not pure storage, but, um, alright, so this is, so what are we talking about, right? So at a system level, this is our platform. It interfaces with number of other capabilities, right?
It could be application devices, it is, uh, uh, IT and business tools. We work with a number of different data platforms, right? And we have, uh, uh, data, um, platforms in a bundled with the product.
Like we have, uh, uh, graph db, we have, uh, you know, and, um, uh, OO open search and so on. But we also can work with customers external tools, right? That's another value proposition.
What we bring, we work with a number of different LLM models, right? And pretty much all the latest ones, right? Like GP five two is out now we are already testing it, right?
It is been out only three weeks or so, right? Uh, cloud, right? We support, uh, others as well.
Okay? Then we interface with output tools, which are essentially ITSM tools, right? They could be automation tools and so on and so forth, okay?
And on the, the, the user interface to the platform is, uh, through two ways, right? So there is a studio, our own studio, not the Microsoft, so what we call it a copilot, right? Which is mostly conversational queries, right?
So you can ask conversational queries, you can test your hypothesis and then convert that conversation into an agent right there. The other user interface is what we call an agent studio, right? So it allows you to build your own agent.
So the platform, uh, out of the box, we provide about 50 agents or so, right? Across different categories. So we have AIOps observability, SecOps, bops, what have you, okay?
And then the real power of the platform comes for partners and end users to build their own, own, um, own agents for bespoke use cases, right? And that's where now we are kind of, you know, broadening the use cases beyond IT ops into SecOps or be ops and so on and so forth. Yep.
So before you go on, this is, uh, Jack Poller with Paradigm Technica. Yeah. Um, on your, your architecture diagram there, you've got a couple of benefits listed.
Faster deployment, MTDR reduction, alert noise reduction. Can you explain a little bit about what sort of, what are we talking about there? In what context?
Yeah, So, uh, in terms of, uh, the consolidation of the tool, right? So now you have, you know, individual domain specific tools, right? Right.
Like you have, uh, Dynatrace, right? You have item tools or a PM tools or what have you, right? Or you have infrastructure tools.
So we have the ability to now go directly go and talk to the devices ourselves, right? Okay. So in some scenarios, you may, you may, over a period of time, you may see that, hey, uh, this platform can actually pro, uh, cater to those use cases.
So, so those use cases would be monitoring your AI infrastructure, uh, Well, yes. Yeah. It's not just AI infrastructure, it's actually the customer deployment, right?
It's application stack. The infrastructure stack stack, okay. Yeah.
Yeah. And so then the MTGR, Yeah. So the next one is alert noise reduction, right?
So that's the standard use case in the AIOps world, right? Okay. You're getting a lot of alert fatigue and you're doing, correlating that, understanding the context, right?
Then, uh, the other one is faster deployment. So the platform is, is what we call it a data driven platform, right? So what that means is there's very little, it's a low-code platform.
So we go and we, we, uh, basically do data discovery, right? Uh, we understand the schema, right? And then we create, uh, you know, uh, declarative templates, right?
So that's another value proposition of the platform where you can, it's composable, right? So that's a faster deployment. And MTTR reduction is literally, uh, you know, one use case, which we'll talk about, right?
Uh, where, uh, in a standard way it was, is it was taking a long time for them to get to a remediation, okay? Right? That was a problem.
State making sense so far. Yeah. Josh here with another follow up question to that.
Um, how much data ingest are you doing, right? You're talking about interfacing when the data is a source and building a template around that, but are you also building a data lake underneath, like, or dependent on the platform, because right, all of my data sources are not gonna be equal, right? If I, I'm, if I'm assuming that the customer doesn't already have a data lake and they've gone through that bronze, silver, gold process, right?
To, to purify things, yeah. You're saying I can still interface directly with another product like ServiceNow or, or whatever's up here like Splunk and not worry about a data lake, or are you building a data lake? No.
So you would have a data platform. You would have a data platform, but a lot of these things, actually, the correlation and so on, it happens in the context engine. Mm-hmm.
Right? So that's where, but we do have persistent layer, right? You could have Splunk, you could have open search, uh, you could have min io, right?
What have you, and then ServiceNow and all those are the output layers, right? Mm-hmm. Where you're basically updating a ticket, uh, or so on.
Okay, that makes sense. Yeah. Yeah.
I gotcha. Yeah. Yeah.
So, uh, real quick, right? Everything I'll have to quickly go through, um, yeah. So, uh, again, uh, capability wise, right?
So what we say how we are different is basically it's a full stack platform, right? That means there are a lot of agent platforms who only focus on the agent layer, right? So we take care of all the way from the data to the agent layer, to the automation layer, right?
So that's what we mean by full stack, uh, purpose built, right? Because a lot of, uh, you know, there's a lot of hype around this. Frameworks like crew and auto gen and so on in the IT world, you're dealing with realtime data, right?
You're dealing with alerts, incidents, time series data. Many of those frameworks don't even support realtime data like crew. Uh, AI doesn't support, right?
So this is a platform which is build grounds up for this particular real time use case, okay? And it's enterprise grade, right? So this is how he'll, um, Russia, when, when he walks, uh, over in the demo, you'll see this is our catalog.
This is the first point of contact for an end customer, right? Where you have this, uh, you know, 40 or 50 AI agents where you are basically, uh, you know, clicking them and demonstrating them. Okay?
So I think this, and then a couple more slides will actually, uh, establish, right? So this is, we are now kind of getting, uh, under the hood on what is our unique value proposition, right? So the middleware is really where, you know, we differentiate against all the other, uh, uh, other platforms, right?
Which are like first generation, if you will, right? And it comprises of the context engine and the universal tooling. And again, as I said, the LLMs, they don't have state, right?
So that's why the context becomes important here. We maintain the state, right? And now when you have all this, you know, large number of, uh, data sets, uh, SQL rows coming to you, or, um, large data, um, large logs, you're coming to you, we keep, uh, it, we basically summarize it and keep a state, uh, in the context engine, right?
So now once you have that state, when you are having the MCP servers, uh, say server one talking to server two, it doesn't have to go all the way to the LLM for every transaction, right? That state enables it to kind of converse it between the two, two different MCP servers. Anyway.
So the three main features of the, of the platform is the agent tops model, right? So this is where we actually operationalize, right? The use cases where you are essentially able to bring in, uh, you know, uh, uh, those, what we call scaffolding, uh, of the, uh, of the elements, right?
Where your trust governance and, uh, observ and so on. Then there is a middleware, which is where, you know, we are doing shared context, intelligent caching. And the tooling, tooling is where we are able to dynamically go and talk to data sources, uh, using, uh, either MCP or raw, raw, all this, right?
And then finally, this is all built on top of the data fabric, which allows you to actually connect to this 1700 plus, 1800 plus data sources. Okay? Uh, yeah, I, I'm not going to walk you through this, right?
But essentially the capabilities we have, like say for example, how do you build trust with a stochastic model, right? So this is where, uh, you know, we have this, what we call prompt templates, right? Which are, are basically give some examples to the LLMs that, hey, this is how you'll go about root cause analysis, right?
This is how you would go about looking at the, uh, CBE, right? And then what we do is we use the LLM to choose those instructions, right? That's what we call dynamic instruction.
So that's kind of a trust part. Everything flows through with a persona, right? So we have guardrails around that, that's very important.
So, and level three engineer will see only certain tools and data sets, which are different than an info IT manager, right? Who will see a different set of data sets and d different tools, right? Same thing with governance.
Uh, there is finops model. So you can leverage look at, hey, what's your cost based on a project, right? So you can create a sandbox and assign certain personas to it, right?
Uh, security, right? So we have this concept of list agency, right? So every agent should only get access to a data sets and tool set, which he's entitled for.
Yep. Reliability and performance observability, right? So again, this observability is very different than the infrastructure observability we are all used to, right?
So we are talking about observability at the agent tick layer. Mm-hmm. So, which is important, right?
And what we do is, uh, not just provide you, uh, all the audit lock trails, but also we have, uh, flows, flow maps, right? Mm-hmm. Which is in real time, that means, say here is a user prompt, which went to the MCP client, it went to the MCP server.
This is what, what was the tool set which was accessed. Mm-hmm. So once you're looking at it, it's very little room for hallucination.
Yeah. Uh, this is how we will actually show this in the demo itself, right? So we operationalize it, like, uh, looking at the persona, looking at the tools, and so on.
Uh, this is actually a positioning slide. Some of you might find this useful. So there are other players as well, right?
And we acknowledge that. So there are observability players. The traditional one, right?
There are AI observative players, like right res and so on. Then there are orchestration frameworks. There are SRE vendors, and that there is platform place.
So Gartner has actually placed us both in the, in the platform, uh, uh, this as well as in the AI SRE quadrant, right? And that arrow indicates that some of these traditional observative players who lack, uh, agent, um, uh, framework on top of that, we are able to compliment that, right? And they have been our partners for a long time, right?
Guys like, you know, uh, Splunk and, and, uh, so on and so forth. Okay. So any questions on this?
On the positioning? No. Yeah.
Uh, again, so this is some examples of how we work. Like we have Splunk agents, uh, we have I om agents, right? We have CRM agents.
Okay? Uh, so this is the agent to clear complementing the existing observability and I om tool. Yeah.
And so I, I actually, if you have a question Yeah. Um, could you kind of give us a, an abstracted description of how one of your customers is actually using this? Yeah, yeah.
So, exactly right. So, uh, I'll skip this. Let me get to this example, right?
So this is an example of en large e energy company, right? So, uh, you have an application and infrastructure stack, okay? Uh, uh, you have a number of different tools, right?
There is observatory tools, there is data, data tools. There are, uh, I tools and so on and so forth, okay? We are basically bringing this in all okay?
Through, uh, our, our, that dynamic MCP tooling, right? And this is where, uh, you know, the, the engine basically is able to curate all of that telemetry to provide it to the LLM, okay? Okay.
For root cause analysis or SecOps and so on. So, so Is basically an MCP run in for MQQT, Uh, well, M-Q-T-Q-T could be in the IOT scenario, right? We don't have your MQQT, right?
These are more enterprise applications, okay? Yeah. Yeah.
But yeah, I mean, if the woulda been, that would be another endpoint. So, so Ray Ese Silver Tank Consulting, my, my question is, you know, a lot of these solutions servers, VMs, networking storage, GPUs, they're moving more and more to MCP server solutions themselves. So they're supplying a lot of the MCP endpoints for what your universal tooling is doing.
I understand from a legacy perspective, a lot of legacy solutions out there may never get to there, but, but, uh, from more current infrastructure perspective, they're all moving to that, in that direction. Yeah. The other question I have is, and, and so the LLMs and the agent tech workload flows are also moving towards more context, sophistication, management, compression, those sorts of things.
It's, it's, it's, it's, I guess my question really is do you have a, a, a, a solution that, that will, you think will last in this environment? Well, Yeah. I mean, again, uh, that's a, there's a good, good point, right?
So again, there are, most of the enterprise solutions do have MCP, right? But there is also a large number of, uh, you know, legacy or raw data particularly, right? Is the same thing with open telemetry, right?
We talk about open telemetry. There's 80% of the data sources still don't support it open telemetry, right? So that's the problem.
Certain, coming to the second part of it, as I said, that the caching, what the infrastructure layer caching we are talking about is, is different. That's good. Good announcement.
But it's really the purity of the, of the context what matters here, right? Which is what we are focusing. It's The second time you mentioned that.
Yeah. Yeah, exactly. I mean, you'll hear that all the time, by the way, I'm going to, AI isn't just for software.
As companies like Flux AI are leveraging generative AI to act as a designer to bring your ideas into the real world, they enable a user to design electronics, creating the schematics needed for a contract manufacturer to produce it Flux AI is also a community of users for support and ideas. Will AI enable on-demand manufacturing of just about anything we can dream up? That's the question on this episode of utilizing AI with Olivier Brashard and Mathias Wagner.
Welcome to utilizing ai, the podcast focused on practical applications of artificial intelligence from the TUM Group. Every Wednesday, we explore news and use cases of the way in which AI is transforming enterprise IT and the industries it serves. I'm your host Steven FoST, president of the Tech Field Day business unit here at the Futurum Group, and today we're talking about bringing AI into the real world.
But before we dive into this discussion, let's meet who's on the panel today. Hi, I'm Oli Blanchard. I'm a research director at the Futurum Group, and my primary focus is intelligent devices.
So anything that has AI embedded in it, from rings and other wearables to PCs, to the iot, to robotics, to cars, uh, is, uh, is the, the stuff that I focus on. Hi, I'm Matthias Wagner. I'm the founder, CEO at Flux.
At Flux, we aim to take the heart out of hardware, and we're building the first AI hardware engineer to help go from a, from a prompt to any physical product you can imagine. Excellent. And, uh, again, I'm Steven Foskett, the organizer of Tech Field a and the host of, uh, utilizing ai.
So I am an iot enthusiast and all around nerd, as you can see from the background, if you're watching the video. Um, and of course, it's pretty exciting to think that somebody like me without a hardware design experience could develop a specialized iot device just by talking and yelling at my computer. Uh, but I think, uh, Mattias, that's pretty much what Flux is doing, right?
You wanna tell me a little bit more about what it is? Yeah, totally. Um, yeah, I mean, I think you're spot on, right?
Uh, you can come to Flux and, you know, just like you, you can enter a prompt, you get a poem or some marketing copy outta that with Flux, right? You can enter a description of a device, you want design and manufacture, and then Flux will go and design that for you, right? It'll source all the components, figure out the architecture, you know, ask follow up questions to nail down the details you haven't thought of, uh, and then go ahead and design it for you, you know, and get that manufacturing ready.
And that can be really like, I mean, um, pretty much anything, uh, uh, I'd say like, you know, there's like simple devices like, you know, think controller for farming irrigation systems or so to like the kind of controller boards that in vending machines all the way to like, you know, uh, uh, the kind of, uh, electronics that go into a satellite that go to space, right? So that's kind of like the range of applications we've seen so far. Yeah.
So one of the questions I have for you actually, mti, is, is what, what level of proficiency in, uh, PCB design, uh, is, is this really for, can, can someone say a, a, a tech savvy, remotely tech savvy farmer, for instance, uh, who wants to build, uh, their own drones or their own equipment or their own like, sort of sensor technology to do very specific things on their farm or in, in their ag agricultural, uh, uh, facility? Could they just come to this and, and use Flux to go from zero to complete design? Um, or do they need to have some semblance of, of training already or experience in that field?
Yeah, great question. I think you're spot on there, uh, with the user segments, right? So that's exactly the kind of that come to us today.
We do have Indeed Farmers, so building their own farm automation equipment. So I'd say like anyone with like a somewhat technical background can gr this today. Of course, what we're working towards too, right?
It's like these, you know, if you think about how easy it's become for like 12-year-old kids to build iPhone apps, we wanna enable them to build their own iPhone down the road or an iPhone equivalent device, right? And that's what we're aiming for. But today, yeah, it's mostly technical crowd.
These are not, uh, uh, uh, uh, users who have necessarily designed A PCB before, but they have like, maybe like a mechanical engineering background, uh, or industrial design background, right? Or, or software engineering background, right? Um, and then, you know, we, we make this as easy as we can here for them to design a board and look, in some simple cases, you can really go from a single prompt to the fully finished thing.
And in other cases, you know, take a little bit back on forth, you know, on being more specific about what you want, um, and steering it in that direction. So I think it's not unlike your typical like, you know, chat typical session, but like sometimes the first prompt you put in, you get exactly what you wanted, and other times you have to nudge it a little bit into the right direction. Okay.
That's pretty good. So what type of, um, what type of equipment do you need for this? Is this just like, just any laptop?
Do you have to have minimum specs to run this? Uh, and, and I also assume that, um, just from looking at the website, right there is, there's certain types of files. So you have to be able to run some kind of design software on this, uh, like CAD or something like that, or is it all, uh, cloud based and, and included in, uh, in Flux, it's batteries included, right?
Uh, it's included, it, it, uh, flux runs in the browser. You need nothing else. It's like, you know, I always say you can go here from like the idea to the fully finished product, uh, uh, all within flux, right?
So, no, you don't need anything else. Uh, in terms of specs. I mean, look, I think like any kind of machine made in the last five years or so, we'll do, we'll do fine an older one probably too.
I think it depends a little bit on the size of project you have. Like a larger, more complex project is gonna take one memory, and then short, maybe you want a more highend machine. But I think, you know, any average machine made in last five to six years will do fine.
Cool. Okay. So, um, following on this, this line, we'll, we'll switch in a little bit, but I'm just, I'm interested because I might actually be interested in using this.
Uh, and I don't have necessarily the practical skills required to do this, um, even though I have the theoretical skills as, as an analyst. Um, but I, I saw on your website that, um, there's a, there's a human support element to this where, uh, users of flux can call on, uh, an active builder community, I think is how you phrased it, and also engineers who might be on call. How does that work?
Is that sort of like a, a, a free floating community that can just use Flux as, as sort of a, a meeting point? Or do you have staff that helps customers? What's, what's the model there?
Yeah, great question. I mean, to start with, right? Flux is a platform's much designed, like in GitHub, right?
GitHub is this like pillar in software engineering, right? For engineers to work together to collaborate together, to learn from each other and to use each other's stuff, right? All these like breads and butter kind of things, you need to make software, whether it's like an encryption library, uh, server, uh, library, whatever.
You can find those in GitHub and use those and remix them and all that kind of stuff. And so we've built the equivalent of that for electronics, right? If you think about electronics, you need the, the, the, the digital representation of a semiconductor or so, right?
And somebody needs to put that in there. And so that's either comes from the semiconductor companies themselves that are represented on flux, right? Or that comes from users who do that, or, or semi distributors, right?
So there's a big community aspect to that to create this, these like nuts and bolts that you need to build your stuff there. Um, so that's under like the first layer. And then of course, yeah, we have like a big Slack community, uh, where people post jobs, you know, people look for jobs if they have questions or they brainstorm ideas.
And you know, I think this idea that everything is better if you do it with others. Um, and that's a providing there. And so there's other users there, there's power users there, and there's of course also like, uh, our team members are there.
I'm, I'm out there, you know? Um, and it goes all the way from like, you know, reporting a and getting that fixed to explaining to how something works to, like learning what to build next for us, right? What's missing, where are the gaps where people get stuck?
Um, and so, yeah, I think, I think, you know, you zoom out here, when we started a company, it was clear about what we were doing was really big and ambitious, and it was gonna take a long time, and it wasn't gonna work if we wouldn't like work directly with the users from the, from day one, right? And so I remember, like, we started a company six years ago, and we shipped the first, I don't even wanna call it a beta, the first thing we shipped three months in. And it was embarrassing to demo that two users, but, you know, but we learned from that, right?
And we moved the whole thing into the right direction and kept focus on that. And so that's, that spirit lives on, right? To be close to users and get feedback and, and, and do better.
So once you've made the design in flux ai, um, AI can't manufacture things. So I assume that you, uh, produce a, uh, production ready, uh, schematic that can be then sent on to a contract manufacturer, of which there are many now who would be happy to manufacture, um, you know, small or even large volumes of these things. Um, can you talk a little bit about how, uh, it goes from ideas to designs to actual physical product?
I mean, you held something up in your hand just a minute ago, I'm assuming you didn't solder that together. Uh, who did I try to avoid soldering these things these days? Because it's, it's tedious.
Um, it's so tiny, you can't even see that here. But you can know it's like tiny, tiny components here. Um, so no, what's the process look like?
Yeah. So today how this works is like, yeah, we, when we, when you're done in flux, we give you the manufacturing files. These are called Gerber files, there some other formats, but, and then any manufacturer in the world of which is 10,000 right?
Can take these files and make this for you. What's interesting in, in PCB boards is like that making PCB boards has been automated for a very long time, right? The whole form factor here exists because it's automateable.
That's the whole innovation here, right? Um, and so in that sense, can AI make this, I mean, it kind of does, right? A robot will actually assemble this, right?
Um, and what we're essentially providing you is the, the files to program these robots on the assembly line to do this for you. It's funny, um, I actually, to that point, I actually have, uh, the name tag in the front of my office at, uh, I'm, I'm at home, but at, at the office here, um, is a PCB that was, uh, created as a one-off, and it, and, and it just spells out my name in the traces. And, um, so you're, you're absolutely right.
I've definitely seen this in action. And of course I've got friends as well who have, uh, had, uh, contract manufacturers. But of course, um, what I hear from folks who've been involved in this is that there's often a, a bit of back and forth and fine tuning, you know, all that.
Um, is, is flux helping with that whole process as well? Yeah, it's a great question. Um, so yeah, look, ideally there isn't back and forth.
Your design is straightforward. They just make it, they say that you have it on your doorstep, right? Ready to go.
Um, but you're right. In some cases, uh, the manufacturer has some questions or feedback. Um, today, what you can just do, you can just copy.
When they email you, you just copy paste that into flux, and the flux will fix that for you, right? Um, and, uh, you know, but, but of course, what we're working towards too is like to integrating that too, right? Because that feedback is of course, an important step of like the learning loop for these agents, right?
To learn what mistakes they made or what cost fiction and manufacturing, and to not do that again in the future, right? And I think that's one of these big, uh, compounding effecting here network effects you have in ai, right? The humans of course, learn too, but it's much more lossy process, right?
Whereas like these c models, when they learn, they tend to not forget that. Um, and so it's very effective. And so you, you reach, quickly hear like a, a high situation of like, knowledge that prevents them from making mistakes in the future, at least for like that category of problems you work on, Right?
So my, my next question, if I can, if I can get in there real quick, is, uh, I just got back from CES, uh, in January. And, um, surprisingly enough, the, the, the big theme at CES wasn't necessarily AI or A IPC or a devices, it was robotics. Uh, even though I think we're still very, very early, early in the game, every single vendor that I talked to, even the keynotes from, uh, the big semiconductor companies from Qualcomm, Intel, a MD, Nvidia, uh, arm, nxp, Texas Instruments, everybody was about robotics.
Whether it's robotic arms or delivery robots. Obviously humanoid robots are always the flashy thing that's shown on stage. Um, but it feels like there's, we're an inflection point with robotics in general, all form factors.
And I, I, I, I look at your company, I look at what you're, the service that you're providing, uh, for, for designers and for, you know, small businesses, large businesses, and, and the, the power that you're sort of democratizing in terms of designing, uh, uh, even FPGAs. And I, I was wondering if you could talk a little bit about where you think you fits into, um, I think this, this transition into sort of mainstream robotics and how your company, how flux specifically with this, uh, this prompt based model can help accelerate that, scale that, and make it more available for, for a lot more companies, whether they're startups or established companies that just wanna build their own robotics, uh, departments. Yeah, good question.
Where to start? Um, nothing with robotics, but yeah, you're totally right. It's an exciting time because, you know, we have, like, for the first time now, these like reasoning engines, right?
Uh, uh, and, you know, for robotics, like the hardware we've had for a while, if you look at robust dynamics been putting out for the last decades, right? It's incredible stuff. But we haven't had, we didn't have these reasoning engines, right?
And so now we have those, and so it's like a new wave of innovation down robotics. And at the same time, by, you've seen it, but all the demos you see, yes, these robots can reason now, but they reason very slowly, right? It's like a watching, like a tour, you know, either a leave of, of, of, of salad or so, um, and I think that will get figured out, right?
I think people were very early and we're probably still a couple of years out from seeing like a truly humanoid, you know, latency kind of capability, uh, uh, robot. But I think it's gonna happen. Um, I think in the meantime, there's of course a huge opportunity for these like domain specific robots, like think about, trying to think of the company's name, but there's like a company they make like robot vacuums, but with AI now, right?
And that's a huge leap forward towards what we happen is like Roomba generation of vacuum stuff, they're kind of like, they work, but they're kind of dumb, right? They'll eat a USB cable, you know, every time. Um, and so I think data sector, you're gonna see a lot here now, starting already now.
And I think flux is kind of in the same category, right? We're like very domain specific. It's an AI model, an AI agent or robot, if you wanna call it that, to make PCB boards, right?
And it doesn't make that by being a human humanized robot that you tell to, and it sits there and sold us this for you all day, right? It does that by being very special on the design and then being able to utilize the existing manufacturing capacity of the world, of the world economy, right? To then turn these designs into, into real products, right?
Um, and I think that in this domain, yeah, I mean, the thing is probably, probably a couple more decades until you have human robots who could replace that capability. Human robots who could replace that capability. Um, but yeah, but, but it's in the end, it's, you know, flex is also just a giant robot.
Yeah, yeah, yeah. No, it's, it's, it's cool. Like I, I, I feel like you're, you're the right company at the right time, especially with, uh, you know, how much easier you, you make it.
Uh, just before Steven asks you, uh, his next question, I'm wondering if, uh, how well you scale, let's say that everybody discovers flux, right? Worst problems to have, uh, and everybody wants to start using it and building their own boards and developing their own robots and their own systems, uh, or improving on existing ones. Um, can you scale well, uh, well enough to, to meet that kind of demand?
Or, um, are you currently structured where it might be, it might turn into a first come, first served, uh, camp service? Yeah, great question. Um, no, I think the way we see this is this, right?
This, today there's a, a market of tens of millions of users in the world ready to adopt flux. And that's mostly people who've technical fields before, again, they've, before just used tools, right? And then just bought the electronic, someone om or contacted that out.
I think those, that market is up for grab here now for a company like ours, right? Um, and then, but if you go all the way here, right? Then, look, if I can type in a prompt and I can get any product manufacturer, I can type in any idea, and I, and that's, and get this spec within like 10 days manufactured ready to go.
Um, and we do this for electronics today, but, you know, why will we not also deliver the enclosure for this, right? The enclosure for this is pretty simple. We know where the plugs are, you know, where the heat sink has to go.
We can make the enclosure for you. And most enclosures aren't like, you know, a piece of jewelry like your iPhone is. Most enclosures are gray p vvc boxes, right?
Especially in industrial applications. So that's, you know, that's not so difficult to do, but if you can do that right, then why would you still go spend two hours on Amazon looking for the product you were looking for, if you could just describe it and it got made for you. Yeah.
Right. Yeah. And so I think on that level, right?
Yeah. Then there's a market of billions of users, right? Uh, and that's, look, that's a long road.
Uh, but that's the road we're on. Yeah. What, yeah, no, that makes a lot of sense.
What about that? So let's take that to the next, uh, the next level. So you're talking about designing PCBs here, but, um, what about everything else?
I mean, it's gotta have an enclosure. Um, you know, what about batteries? What about, uh, you're building your giant robot, uh, you've got, uh, motors and wiring and you know, joints and, and whatever else, you know, what about screens?
What about, um, everything else? I, is that all part of the flux world as well? Or is that something that's handled differently?
Yeah, that's a great question. So, yeah, no, today we're hyper focus just on everything that goes to make this board right? And look, and, and if a display is mounted on the sport, that'll work, right?
Um, and, and we do this simply because, like I said earlier, right? This is, this is a form factor that was designed to be automated. It's meant to be fully automated for a hundred years, and then we're finally here.
Now, we can actually fully automate this, right? And so that's kinda like to be chat here, you know, that we've built out for ourselves now, uh, and we're doubling down on that, uh, for the, for the, for the coming months. But then, yeah, right?
I think from here we go to enclosures from, from there. We're gonna go to like, you know, something that has a piece of an enclosure and maybe motors or display or the batteries or just sensors that are not on the board themselves, right? Um, and yeah, from there, we're gonna go to full robots or smart toasters or, you know, you name your category, right?
Um, because there is like a huge, there's huge infrastructure in the world to build these things if you can deliver them the design files, alright? Like the bottleneck isn't in the manufacturing capability. There's thousand of manufacturers just in China, right?
Who would love more business and who don't care whether they make five units for you one unit or 5 million units, right? They're just looking for more business. Um, and so, and so that's kind of like here, the, the, the need, you know, we're serving, right?
And, uh, in the short term, Yeah. I have a question about software real quick. Um, so, and, and the answer might be very, very short.
Um, but essentially, right now, obviously this is really good, but this is also fairly new. I'm wondering if you like, where the limitations of, of your software, uh, and especially the, the sort of like prompt based, um, you know, design generation, um, that, that you have the AI working in the background, what are some of the limitations that you're working on improving mm-hmm. For 2026?
Where yeah. Where you're not as good as you'd like to be yet in, in some way. And if the answer is no, we're great, we can, you know, everything is perfect.
That's, that's a fine answer too. Um, but I expect that you probably have some things that you wanna improve on. I'm just wondering what those are.
Yeah, great question. Uh, no, everything's perfect. I sleep wonderfully at night every night, you know?
Um, no, no. It's a, it's a bucket, but a lot of holes, uh, that's the valid of it, right? Uh, that's, that's the fun part.
Um, yeah. Where falling short, I think, you know, the big issues here, uh, that we're working on is just recall, right? So, look, it's not good enough if the model gets it right, eight outta 10 times, right?
Especially along of a chain of like thousands of, of decisions that have been made, right? If you, along every step, only get it eight out 10 times, right? And at the end you got it all wrong, right?
And so that's something we're working on. Uh, um, and then the other thing is like speed, like doing that faster, right? I mean, yes, you could argue it's, it's a miracle that the, the model can do something in an hour that would've taken an expert weeks, right?
Um, but if you really wanna wield this like a guitar, which is like, think about a guitar, right? It's an incredible tool. It's like so intuitive and responsive, right?
And playful, and you can discover, right? And like the feedback loop is really fast. Um, we think that that's kind of like what really amplifies, you know, human innovation and creativity, like fast feedback loops, right?
Try something, get a response, try something new. And I think that's what we're trying to get to here, but flux too, right? To be really like a aspiring partner.
And, but you can really go quick and forth and play with idea and iterate, right? Because, you know, we've all like made things, whatever, we've written something or built something, it's all about getting into like an, an, an effective loop here, you know, of like making something, getting feedback, whether it's good or not, changing a little bit, trying it again, right? And that's what we're trying to get to.
Um, and then, you know, I, even if you wanna make an actual iPhone, right? An actual iPhone is like a thousand funnel components, 12 layers. It's like a very high density design, and you need to support like a bunch of manufacturing processes to make a board like that.
And we don't have them fully covered yet. Or like at a level where you could actually effortlessly make something like an iPhone, right? Um, and so that's what we're working on, Right?
No, those are good answers. That's what keeps me up at night. Yeah.
Yeah. Thanks for being candid about that. 'cause you could have just basically said, no, everything's good.
Um, uh, a a question about that. So in, in a past life, I was a, a product manager. And, and so I, we, I helped develop, uh, new products and, and upgrade old ones.
And, and I remember, I think it was an IDO, uh, tenet, uh, IDO the design firm, right? From, from, uh, the, the nineties, um, that was, uh, fail as fast as you can or something like fail fast, you know, uh, succeed faster. And this seems to me you talked about something that would've taken a designer weeks, just a mere hours with this.
And that's how you're, you're shrinking that, that time envelope. Um, I wasn't thinking necessarily, when I'm listening to you talk about this, about the finished product, I'm thinking about as, as this former product manager, the prototyping process of all these different iterations. Like, I have an idea, I wanna build this, um, I wanna create a working prototype as quickly as I can so I can test it and then see if that's really what I need, or if I wanna add some features, make some changes, and then take what I learned, what works, what doesn't iterate, iterate, iterate, iterate until after two or three, uh, different iterations or maybe 15, I finally arrive at a manufacturer, uh, tested prototype that works and, and that I can fairly quickly turn into a production, uh, production ready product.
So to me, in, in, in all of this, one thing that we haven't talked about really is, is this process of either failing quickly or learning quickly, uh, and getting there faster through these prototype iterations. Do you think, um, that is still an old school way of doing this? Do you think that your tool helps get to that finished product with fewer iterations?
Or is it just, you know, faster iterations, uh, to get there? Not less necessarily, but just getting there faster? It's a great question.
I think if you look at like AI coding tools, which I would say is probably like the bleeding edge of AI right now, the process has to make, software has drastically changed in the last 12 months, right? It has changed more than a change in the last 40 years, over the last 12 months. 12 months ago, if you wanted to build a feature in like, some piece of software that you would probably start with like a conversation, a whiteboard exercise, your designer would for some mockups together, you then you maybe do some user testing with that.
Eventually you would get somebody to implement that test that, right? It's like a lot of steps in the process. And if you look at like software teams like us, like work now is, we skip all that, right?
Somebody has an idea at 2:00 AM in the morning, right? They typed the idea into one of these coding agents, and the coding agent would just build the thing into our actual product, right? And on Monday morning, they just show off that working thing in the product that's like ready to be shipped or really close to be ready to be shipped, right?
And I think that's kind of what we're gonna see happening here, and hardware too, what we're working towards to right? Select, remove all these in between steps. Yeah.
Yeah. And make it much more responsive and intuitive, right? Again, like going back to this guitar example, right?
Just like a guitar, if I like pulled a string and then had to wait five days to get like the sound back, you know, that would suck. You know, like a guitar would not be as popular as it's today if that was the case, right? It is so popular and, you know, adopt and, and fun because it's so immediate, right?
And it's immediacy, you know, I think we, we've seen in software happening now, right? And we're gonna see that and happen two now over the next 12 months. Yeah.
No, this could, this could become addictive, um, for Yeah. For tankers. Yeah.
It's fun, right? It's, look, it's fun because you can try so many things right before, right? In this like linear sequential world we were in, I had to have an idea.
I to like kind of spec it out, try it out. I could only do one thing at a time. But that's another thing with these agents, I can now try 20 ideas at the same time as these ideas come, I just drop the prompt and then an hour later I check back in where, where it ended and then look, maybe it went nowhere.
That's fine. No, I don't care. It's just a random idea.
Maybe it's actually amazing, but then let's double down on that, right? And that's like the power you get here now. And I think in Harvard it's gonna be amazing because you know how expensive it's to run one process.
Like you can't try 20 different things, right? Because there isn't enough time and money in the world, right? To do that.
Even like on at large companies with lots of resources, like you think of Apple here or meta making hardware, sure. They maybe have like four or five different teams working on different versions of the same product or idea, right? Uh, with ai, they could be working on like hundreds, thousands of different approaches to the same idea and then pick the winners or mix and merge the winners, right?
Um, and that, and that mind bog speed too. And so that's kind of like the elasticity and you know, IM, we're working on, Yeah, this is, this is very Tony Stark in a little bit, right? And just kind of Oh, For sure.
You know? Yeah. No, no, this is exactly, I have a screenshot of that in our pitch deck for investors, right?
No, no, you're spot on, right? Yeah. Really, I think this is exactly like the kind of future that's gonna be real within a year.
Yeah, no, it's great. And it's gonna be available to everybody, right? That's cinema thing.
Is this gonna be like, like Tony in the Marvel universe, only Tony Stark has this, and that's why the has all these things. But the future of a building, everybody's gonna have access to this, Right? Which means also that you're, you're sort of flattening a lot of the moats that keep some companies with all this expertise and all this investment, uh, sort of in the lead, you're, you're making it much more accessible to startups or individuals or, or even universities, uh, who want to really innovate, uh, yeah.
In, In that space. You know, people always ask me who we're competing with and, and they think of the legacy tools here, and I tell them like, no, the legacy tools are not what we're competing with. That's actually a really small market, the existing PCB design software market, right?
Who we're competing with is the OEMs, right? I have like, my, my favorite example here is like, we have a customer, they make vending machines, like for snacks and beverages, and they used to, before they started using Flux for a single vending machine, it took like four or five individual PCP boards that they had to like, buy from a distributor who bought it from a distributor, who bought it from a distributor, who bought it from an O, right? Had to integrate those on the, into a single machine on the assembly line, configured this wired up, right?
Then in the field, one of these machines breaks down. You've gotta bring replacements for every single board, figure out which ones broken, replace that, right? Um, it's workable, it's cumbersome.
Now, fast forward that to them using Flux now, right? They have a single custom board made with Flux that does exactly what they want, no more, no less, right? Um, the board custom pennies on the dollar because they're just paying materials now.
There's no distributors or middleman that get a cut right on the assembly line. You put one board into the machine, cable in, it's working, the machine stops working in the field, there's one board to replace, right? So the total cost of ownership suddenly is like hundredth of what it was before, right?
Plus, you can make it exactly what you want. You're not dependent on somebody else to build the thing you want. You can make it that, right?
And that's the power, that's the opportunity here. Yeah, no, it is. I, I, I couldn't have said it better myself.
So I think this is all the time that we have, so I really appreciate it, Mathias. Um, this has been really interesting. Uh, and now I'm not just interested in Flux as a, as an analyst.
I'm actually interested in Flux as a, as a potential practitioner. Like, I feel like I might actually be able to, uh, uh, to start dipping my, my toe in, uh, in this PCB design, uh, space. Um, which I think is the point of this.
Like everything that we just talked about, I mean, throughout this episode, obviously this podcast, but also really in the last five minutes has been, uh, sort of encapsulated this opportunity of this intersection between, um, ai, browser-based AI that can, can basically allow you to describe what you wanna build and as a companion, as an expert, it helps you build it and get from basically idea to manufacturing. Um, and, and I really wanna thank you for your time today and explaining this and also, uh, for your vision in building this, because this is the kind of application of AI in the real world, um, that it, that's, that's a lot more real than a lot of the other, uh, AI applications that, that we've heard about. Uh, and also it's ready now.
It's not something that we have to wait for six months, or two years, or three years. Um, it's available now, uh, right now. So I'll, uh, thank you again.
I'm gonna turn it to Steven. Um, thanks. Uh, and I, I hope that we can continue this conversation soon.
Yeah. Um, on that note, um, before we go, um, uh, Mathias, um, where can people continue the conversation with you? Where can they learn more about Flux AI and, um, will we be seeing you in the physical world anytime in the, in the coming months?
Yeah, great question. So look, you can find us at Flex ai. That's a good starting point.
If you wanna start building and start, you know, exploring what this, what this, what this new world looks like. Yeah. Come there.
Sign up. Um, where you can find me. Look, you can find me on Twitter or LinkedIn, you know, just the me if you wanna get in contact with me.
Uh, I'm out there. Uh, in terms of physical world, great timing. You know, we've been remote for the last six years.
We're moving into an office, right? So we have an office in San Francisco and downtown on Second Street. Um, so, you know, we'll, we're aiming here, like as we're moving into also host a bunch of events here, you know, uh, and turn us into like, kind of like a, a maker space if of the new age, right?
Uh, where we together here figure out how to, how to build this. And, and I think that's a big thing. You look, we're extremely early in August.
This is day one, right? And if you're excited here about exploring this, definitely come and try our product. But also, like, if you wanna get involved here, get involved with a community.
Look, we're hiring any role you can imagine. We're hiring for it, right? You also wanna be party and help us build this.
There's a lot of stuff to be figured out. Um, we're really excited about it and we think now's the time, so let's do it. Excellent.
Yeah. Well, hopefully we can see you. Uh, we are out in, uh, California for Tech Field Day, fairly often in the Bay Area, so maybe we'll see you at one of our future, uh, tech Field Day events.
Um, That'll be awesome. Olivier, uh, what are you working on these days? Uh, what can we look forward to hearing from you?
What am I not working on? Uh, robots apparently, uh, but any kind of physical ai. So we have reports coming out fairly, uh, uh, fairly regularly.
Um, whether it's, uh, a market forecast for devices, PCs, et cetera, or, um, uh, semi-annual IT decision maker surveys that, that look at how the enterprise is looking at, at investing in, uh, in AI and AI devices. Uh, I attend pretty much all the conferences, or at least all the major ones. So you're bound to see me at anything that has AI and, and devices sort of like converging.
Obviously we just did cs. I'll probably be at Mobile World Congress in Barcelona. Uh, and then, uh, many more beyond that.
Where you can find me is on x, uh, formerly Twitter. Uh, just look for Olivia Blanchard. You'll find me fairly easily.
I'm on LinkedIn, and obviously if, uh, all else fails, you can Google me or look for me at, uh, the rum group com website, uh, where I publish semi daily articles and commentary on tech. Excellent. Thanks a lot.
And, uh, as for me, you'll find me as s FoST on most social media. Uh, and of course, uh, as I mentioned, I am, uh, uh, a nerd. And so you'll find me, uh, dabbling with, uh, projects like Home Assistant and ESP Home, which I'm sure, uh, Mathias is familiar with as well.
Um, and thank you everyone for listening to the utilizing AI podcast. Uh, if you enjoyed this discussion, please do subscribe on YouTube or your favorite podcast application and consider giving us a rating and review. This podcast is brought to you by the analysts and experts at the Futurum Group, where insights meet ai.
For show notes and more episodes, head over to Text Strong ai, the utilizing AI YouTube channel or the Techstrong TV app. Thanks for listening, and we will see you next Wednesday. So I am Scott Shaley, director of Leadership Narrative with soy.
5 architectures. And then I'm gonna hand it off to Phil and he's gonna fi wrap up this section and go on to the, the, the last section of it. So you can only have me for a few more minutes.
Be back soon. So what does one point 21 gigawatts solve? Come on, somebody's gonna smile about that.
Oh my word, I'm not that old alrightyy. That's Shocking. It can send Marty back to the future.
21 gigawatts. It's also what it takes to power San Francisco for a day. It can also deliver power to 550,000 Grace Blackwell, GPUs, GB 300 platforms, and it enables 25 exabytes of storage in a one one gigawatt environment.
Now, how do I know this? What did we do to be able to tell you that one point 21 1 gigawatt can do all this? Again, looking at the ecosystem, looking at our friends, doing the research, these were all announced in 2025.
These are all the platforms that are gonna go live sometime after the announcement in 2025, Stargate Meta Core Weave, XAI, all these guys. And we did some math and that's how we got to the one gigawatt for 550,000 GB 300 platforms. And if you look at the math required for how many GPUs you have and how much storage you need to go along with those, you get your 25 exabytes of storage.
And so the next question, of course, is really how do you get there? Well, we did that math for you. Um, again, 550,000 GPUs direct attached.
These are the E one s performance drives going right next to that, that GPU sitting in the server, that's the nice, uh, rack based design currently has eight drives in it because they're air cooled or potentially, uh, direct to chip, liquid cold plate cooled. And that nets out to eight and a half exabytes of storage next to the GPU. 5 exabytes of supportable storage using 1 22 terabyte drives to be able to get up to 550,000 GPUs.
So this math is, is like all good TCO models, it's a tit for tat. So whatever you put in, you get out. So we focused on, if I use just my solid state drives at the highest capacities for both knobs, it enables us to get to that many GPUs.
You have any other product that consumes different amount of power or you put sixties in instead of one 20 twos, you're doubling the power footprint. You reduce the number of GPUs available in that gigawatt. So the gigawatt was our bar here.
I'll be 100% honest. There's math that shows if you ignore the gigawatt and just look at the GPUs, the amount of exabytes of storage will be just mind blowing that they're expecting to use. And this is Grace Blackwell.
This is 2025 data. We're now one whole month, literally last day of the month into um, 2026. And the whole ecosystem has changed because when I showed you guys this graph last time at field day two, it had a quote from our good friend Michael Dell.
This is our friend Jensen at CES. This is the market that never existed in the market that will likely be the largest storage market in the world. Gotta love the fact that we finally have Jensen talking about storage and not just memory.
I love it. Now the next trick is at GTC to have him mention our name. We'll see if we can get there, right?
He signed our drive, of course, you know, that whole thing. 5 IIC MSP layer in the Vera Rubbin platform that's tied to the Bluefield four implementations. Now the next slide, I'm gonna tell you how that works from a SSD hardware point of view.
And then a little bit in just a few minutes, Phil gets the lovely chance to show you how that actually looks from a system level implementation point of view alongside what we're talking about. So what I'm doing here is we have the KV C exceeds the HBM spills over to dram. We still have a limited DRAM footprint only.
So many dims, only so much capacity falls onto the local storage, those nice directly attached products, and then it falls over again. And every hop is a connection and a distance and a time. And so by bringing it closer and closer and closer, that's what this whole architecture is about, is access to data faster in a more confined environment.
And so if I take what I had before on the one gigawatt example, and I throw in the VE Ruben platform or the Ruben platform as it's being called, they haven't officially given it the GB nomenclature. We now have three banks of stor of storage products. Now, again, I'm constrained to one gigawatt, so I'm still at 25 exabytes, but this is now how it splits out.
4 exabytes of this new context memory storage, which can still be the high capacity drive. It's just closer to attached to that blue field forward architecture. 1 exabytes of direct attached storage.
'cause they're changing the amount of local to move it just a little bit further out past the blue field. So the number of drives in that initial server is actually coming down as the capacities are going up. And so the interesting thing here is we're with the one gigawatt.
We still have 25 exabytes of storage. We split it out three different ways now instead of two. But note the number of GPUs that are supportable, we're down to 440 or down to 400,000.
We lost 150,000 GPUs. But the performance of the system doesn't change at one gigawatt. And the reason for that is because you're using NVME storage for the direct detach and for the context structure, you have to have the fast storage products in those two layers to overcome the gigawatt problem in this environment with, uh, this new added layer.
Because fewer faster GPS need more access to fast data. Therefore you get ICMS with, uh, solid state drive. You can't put traditional rotating hardware in that layer.
You just can't. So if I, when we come back, uh, at our next field day AI field day, uh, we're planning to, we're gonna have even more details on this and we're gonna blow off the one gigawatt and just show you the capabilities of what the storage looks like. And we're talking a five x or larger CAGR year on year from 26 to 30 on just the demand for high capacity storage over what we were already talking about.
It went from where it was at about a 20% cagr. It's now 30 40% CAGR because of this introduction of this. And that's for the NVIDIA only based systems.
So the storage platform is now the shining star on making the success of the next layer of AI as we get into the inference context scenarios. So we're gonna do a little bit of context switching here. Um, we wanna talk about efficiencies too.
And so I just gave you the hardware centric ICMS one gigawatt constrained environment view efficiencies. When we partnered with Vast, we came out with this amazing TCL model that talked about replacing your SEF hard drive infrastructure with R one 20 twos and Vast in an efficiency play, talking about how to make your systems more effective. And when we were at sc, we put up a bunch of slides.
This was an SC carryover for you guys. You wanna talk about it from here. Um, the Computer history Museum from the museum to the 1 0 1 is the SSD implementation on equivalent of this graph.
The hard drive implementation is going from the Computer History museum all the way up to Oracle headquarters, where, where they were up in the NICE four. I know it in Oracle headquarters, but that's how, that's how far our distance is that you do when you put a drive end to end to end and how much reduction in overall ecosystem environment you can drive. So what we're gonna do now is I'm gonna hand it over to Phil.
He's gonna help explain a little bit about this and then jump back into the context memory and give you some more fun, uh, topics about, uh, the wonderful VAs platform. So I-C-M-S-P is inference context, Infe Inference, context, memory Storage Platform, that's what they called it. And It's behind the Blue fin.
So it's effectively a, a storage solution out there that's doing something for the context management. I'll dig into it. It Looks like great lead into the next little section.
And it's different than the rest of the NAS object data lake that's behind it. It It re-architect for You. Okay.
Yeah. Thanks. Yep.
Yeah, we'll dig into it. Okay. Thanks for having me guys.
So Phil Menez, I'm the go-to market execution lead at Vast. I've been there, uh, I think it's three or four days while I hit my six years at Vast. So I've kind of got to see the company, uh, grow, we'll do a quick introduction, but I really needed to play off Scott's, uh, metaphor here, which is one I'm really happy to be here with solid Diamond and our partners, but vast, we build software, right?
We can't run well obviously without any hardware. So peanut butter great, you know, but it doesn't work so well. It's not very portable or, uh, reasonable to eat if I'm gonna spread it on my hand.
So really the jelly and the bread to our peanut butter is solid on. I couldn't go without that piece. Um, for those who, you know, not as familiar with Vast, we actually launched really the company to the public here at Storage Field Day back in 2019.
And when I was interviewing, that's really how I learned about the company and whether I wanted to work here. Right? And I saw some obviously compelling things, um, since then, right?
We've really become a significant portion of the storage market. As we look to this year, we're gonna drive a very significant portion of all enterprise SSD utilization, uh, with storage, expecting dozens of raw exabytes. That's before our data reduction, which we're gonna talk about that efficiency, and also doesn't count any of the data going into the cloud.
We've made some really big announcements on cloud partnerships this year, and extending the platform, which was primarily on-prem into the hyperscaler space. Sales have really gone well for us. We're roughly tripling year over year.
Our quarter is gonna finish tomorrow, so pay attention as we start to, to announce some of those new things. We have our, uh, customer event first ever user conference for Vast at the end of next month. We'll talk about that.
And I would say just like Tech Field Day evolved from Storage Field Day to AI Field Day, vast has really a evolved, uh, from being a storage company to building many more things on the platform, which we'll talk about, really allowing you to capture data, contextualize it, and then act on it with ai. So I would say the founding principle of VAST is that really a few things. We were very bullish in 2016.
AI was gonna change the world. We were very confident that AI was gonna change the way we computed on data, right? I think both of those things proved out to be correct.
And then the third is that the architectures that got us to where we are, were not the architectures that were gonna take us forward, right? And this is really the main culprit, the shared nothing architecture really invented by Google in 2003 in a white paper, basically defined the internet kind of application cloud era, where you've got these node based architectures, right? I've got a node, got some CPU in memory, I've got some kind of storage in there.
Originally it was disc. Now we swapped it out for flash in a lot of circumstances. But the only way to get to the data on that node is through that node's controller, right?
And I personally storage guy, like I look at this as every scale out na, every scale out object platform. But ultimately it's also the architecture for every data lake, for every distributed data warehouse, right? It's all over the place.
Eventing infrastructure, it is everywhere. And it really does create a lot of scale problems in the AI world. When we look at it really from a storage view, there's some challenges around flexibility, right?
I've gotta create nodes or pools of homogenous node types. Uh, not designed with flash in mind, right? We'll talk about some of the challenges around things like data reduction.
And then finally, one of the big things that shows up everywhere is just this east west traffic. There's so much communication between these nodes that even though I can scale my resources linearly, I'm not scaling performance linearly. And a lot of times these architectures work well small, and these problems show up more and more and more as the clusters grow.
So we're looking at it now from the, the really the TCO and efficiency perspective, right? We know we're in a supply crunch, right? Customers have been trying to move steadily from, uh, spinning disk space architectures to SSDs.
When we look at the AI deployments that we see in practice, you don't see any spinning disc, right? You got power challenges. I Was wondering why you actually had the round things with the floating heads on them for showing disk describes here Because our marketing team likes that image, I guess.
But these are all right. The, the world of, uh, architecture based on spinning, right? It's got the arm.
It's more like a record player. Oh, uh, an older record who? Old school.
Okay, well, I'll, I'll, I'll give the marketing team the feedback. What's the Record player? I love it.
Okay. So in the AI world, right? I think as you look at what's in practice, solid state's required, right?
I think now we're seeing a rise of companies coming up around saying, Hey, tierings cool again because we have an SSD supply crunch. But if it was not a good idea before a supply crunch, I don't see how it's a good idea after a supply crunch. So what we need to do is help customers be a lot more efficient with the way they use SSDs, right?
One of the big problems with this shared nothing architecture is how data reduction works, right? And when you look at a lot of these architectures, a lot of them have given up on things like deduplication. You have compression only, right?
And a lot of the data in the unstructured world, it's already compressed. So that kind of takes away a lot of opportunity to drive efficiency, right? 2 to one is kind of what you're gonna get if you're using compression only.
So what is the challenge with ddu? It's around having a global view of the data, right? In this world, I basically have to chunk up my DDU index because the other option would be to put the entire index on one node.
Everyone would just hammer it, and that would not work very well, right? So we said, okay, we're gonna shard this up. Essentially create a distributed database.
And in that world, every node has a piece of the index, right? So I get a limited view there. As I start to scale this, all these nodes are talking to each other, looking at who's got the data that I already might have as I grow it adds to the east west traffic, adds to the performance limitations.
And then ultimately, the kind of bandaid there is to create limited DDU domains, right? So I'm basically only de-duping within a pool or within a few different nodes within whatever architecture that you're building around. But it's always very local in this world.
So vast, again, looking at the architectures, brought a new architecture to market that we call date very quickly. We call it disaggregated, shared everything. Because essentially we kind of broke the idea of a node apart.
And we have two independent scaling layers. We have our logic compute layer, we call those C nodes. It's essentially container running on an X 86 server.
And then we've got where all of the state of the system lives down in these enclosures filled with very dense, solid, IM 122 terabyte drives or whatever the right, uh, drive is for the customer. Now, some different things. In unlike the shared nothing world, in the day's world, every one of those containers has direct access and actually sees every one of the devices in the system as a local device connected over NVME over fabric, right?
Architecture impossible without NVME over fabric, which now makes it allow that I can have remote drives, feel local from both how they're mounted and performance. We also have a layer of storage class memory in the system where all of the systems metadata lives. So that means I can create a shared global index that every single one of these containers sees.
And that means I can do global data reduction at an exabyte scale without any of those different challenges, right? So fundamentally unique architecture that allows us to look at the data in a very different way from a data reduction perspective. Any questions?
High level, the architecture, how it works, okay. That'll be a theme that we hit on, right? So step one, can we give an architectural, I would say, advantage to how we look at global data reduction, step one.
But again, the problem is we're talking about unstructured data here, right? Not as friendly of deduplication as things like VDI and virtual machines, right? Not as friendly of compression, maybe as a database that hasn't been compressed already.
So we have to look at some different things, right? And really move beyond duping compression alone. So very high level, you look at compression, right?
I'm looking for commonality, repeating data at a very granular level, right? That's gonna be a small chunk. I don't know, eight to 60 4K usually could be anywhere in between, uh, deduplication.
Right? Now I can have a global view, assuming my architecture allows it, but I'm looking for more course matches, right? Two chunks of data, exactly the same.
I find that a lot. VDI, virtual machines, I'm copying databases, whatever. Uh, but again, I don't always find identical matches in unstructured data.
If I chunk up and try to do DDU on a big pool of unstructured data, what you actually end up finding is a lot of chunks of data that are mostly the same, not exactly the same. DDU misses that every single time. 'cause that would be a hash collision that's corrupting your data.
It's terrible. So what we do is introduce a new type of data reduction, again, enabled, because we have this giant metadata structure living in storage, class memory, the architecture that we will identify if two chunks of data are mostly the same, compress them together and essentially store the differences, kind of like a snapshot, right? And ultimately, we don't just use similarity, we use all three of these, right?
So we're looking for the best opportunity compression. We actually use a couple different types of compression. We will look at the data, take a sample, what's the best type of compression, and use that deduplication.
We have, again, global deduplication. We have something we call adaptive chunking. Chunking, which means we'll actually change the DDU window to find the best opportunity for deduplication.
And then similarity is kind of that icing on the top where we're gonna find that next level of similarity and drive out even more savings. Right? And ultimately, you're in a world where we could easily get two or three times more data reduction than the next biggest competitor because of what's happening here.
Do you, uh, I wanna know if you're a believer or not. Uh, yes, absolutely. Uh, I, I'm chuckling because you're, you're giving the exact description of what I would've been describing with solid fires architecture 10 years ago.
Got it. Okay. I knew about it.
Right. But I would say solid fire in the shared nothing world, right? A little bit.
Absolutely. I, but I mean, the, the, the things you're describing are, it's like, yeah, this is exactly what we were doing 10 years ago. No, it makes so no, that, that's, I'm sorry.
That's why I was chuckling. No, no, it makes sense. And by the way, I think it's interesting that, you know, in the block world, the hard drive died like immediately, right?
You know, I was part of the extreme IO team at EMC. We had pure, we had solid fire. Everyone.
The, the hard drive in the, like the block world, virtual machines, V databases, VDI died immediately. That was like 12 years ago. And there's still so much of the world's unstructured data on spinning this because they haven't been able to figure this calculus out, right?
So it's actually a great point. Um, some actual data, right? So if you were gonna say, I don't believe you, I was like, look, I have data, um, average data reduction by the way this is pulled this month.
Because as we've looked at the supply crane crunch, we're like, let's start digging into like where we've come and what the results are. 4 to one, right? Again, these are not VDIs, these are, this is unstructured data.
Some of our customers have hundreds of petabytes of highly compressed video. Some of it's encrypted, um, some massive estates. The weighted average.
87 to one. Exactly. 9, looks prettier on the slide.
Um, and then we have 27% of our customers get better than three to one. We have some customers getting eight to one. We have some customers getting like 30 to one depending on the data type.
So where typically, again, in the world of unstructured data, you'd say, if I get anything at all, 10%, I'd be happy. We're talking about getting you three times more data for your flash. And that gets combined with something that I'm not gonna nerd out on today because of this time.
But our erasure coating is also incredibly efficient. So our erasure coating at scale under 3% overhead, it's actually 146 plus four stripe that we use enabled by our architecture. So, um, when you look at that compared to, you know, traditional kind of shared nothing where you're gonna have maybe 27, 20%, we have a lot of customers moving to Vast that are still using like das Data Lake technology, and they've got their data triplicated, you'd be shocked about how much of the world's capacity is still triplicated.
And it's because it's in these monster data lakes, um, where again, they're getting, you know, for every 10 petabytes they can store three petabytes of data. And those systems don't have any data reduction, right? That's all over some of these large data analytics environments.
So you combine these things, a lot of times our customers, even if they don't get good data reduction, they're getting four times more effective capacity per, you know, petabyte that they buy. And even if they're buying something that is more kind of enterprise, then maybe it's more like double the capacity that you can store for every raw petabyte that you're gonna buy. So we actually just launched this, uh, something called Vast Amplify.
So in the SSD Crunch vast over the years has really gotten a lot more flexible. Again, uh, when we started, we had to run out of every specific hardware build. Now we're working with pretty much every major OEM vendor running on more, uh, traditional servers.
We're in the cloud. So we actually have a program where we're going to customers and taking their SSDs that they already have in their systems and repurposing them into vast systems to amplify the capacity. We actually had a cus couple customers come to us and said, Hey, we've got SSDs.
Your technology is way better than what we're using. Can we reformat these and use them? And we have, and in some very large scale environments, I'm talking at this point, we've repurposed hundreds of petabytes of data, thousands and thousands of drives.
So, Just a question. This is all really great statistics and y'all are doing really awesome, but, um, we're talking about ai. So I would love if you could tie this back to ai.
Can you, does it matter if I have a data lake that's not deduped? Maybe I want that, and I just want the, I just want the data tagged in a different way so I can find it for different reasons. But like, what, how does this tie back to ai?
Yeah. So I would say how it ties back to AI is, right now what we've seen in practice, any large scale training environment, any large scale inference environment that's actually in production at scale is a hundred percent based on SSD. Right?
Okay. That's, that is standout. We have in a world where customers are gonna struggle to get as much SSD as they need, right?
So what I need to be able to do right now is make more use of my solid state devices because AI is driving tremendous demand, right? And we'll get into more how it fits in the architecture, but the point is, looking at bringing spinning disc into this world, we really think is a terrible idea. If having a tier miss is going to kill my GP utilization, destroy jobs, destroy performance, then I can't use that as a lever.
I need to figure out how to make most use of my flash in this AI world, right? As I wanna deploy agents and inference over a much broader set of data that data's hitting on spinning disc, it's not gonna work. Well, you're Not gonna get an argument about that here.
So, right. It, so go ahead. I was just gonna say, but when it comes to training, training data specifically is, uh, as de duplicable, if that's a word, as traditional data sets have been in, in your experience so far?
Yeah, so I would say in training data, um, two to three to one, okay. Is common, right? If you look at some of the bigger neo clouds that are our customers, two to one's pretty typical on training data sets and, and more.
And You're doing this all inline, right? So it's, I would say it's kind of the best of both worlds between inline in the old world, inline men and memory with vast, our inline memory is storage class memory, right? So what happens is the data lands in storage class memory, it's acknowledged up to a host, and then we data, we data reduce it when it migrates down to QLC.
Mm-hmm. Okay. So yeah, kind of outta a band.
I'll show you what that looks like, Actually. Yeah. I mean, that's, that's not terribly unusual way to do it where you, you actually need to do the hashing at some later point to, to be able to de duplicate it, but you end up using much less storage later on.
It did. Yep. What I would say the difference is we don't land it on the capacity tier, right?
So we don't land it on QLC and go back and mess with it again. It's not good for where it's not good for performance. What we do is we leave it in storage class memory where it's very fast access, it gives you a lot of opportunity to move.
And then we don't need to plan to have non de-duped and non-used data on the capacity tier. And I promise, by the way, the most of the rest of the presentation will be specifically on ai, but we wanted to bring in the TCO of making SSDs affordable and and hacking the supply chain crisis. Was that okay?
Thank you for saying that. 'cause that was not coming through. Okay.
Sorry about That. Appreciate that. When you're, uh, repurposing SSDs, are you having to, to migrate the data off the SSDs and migrate back on from a vast perspective?
Or are you assimilating We do need to Simulating the data directly. I mean, We do need to move it. Yeah.
So if we take a file system, we can't convert it to vast and data in place. So a lot of our customers we're either working with swing space or we're creating clusters and failing nodes out and growing into it. You're seeing a lot of usage of the, uh, the new capability to, uh, reuse SSDs.
Yes. So again, we customers brought us the idea originally to say, Hey, we've got SSDs, we wanna repurpose it. So that was how it got rolling.
And since then, yeah, customers are all over us to say, we know we're looking at the year, we've got more demand, we're looking at rolling more ai. We've lived in a solid state world, and we, we know that we can't have capacity that's 30% utilized, right? We're giving us one third of what we're buying to store, okay.
More AI stuff, right? So now I promise the rest specifically on ai, right? But again, we think flash is that kind of first step as to enabling your data on fast access.
So we're just gonna talk about the context challenge and KV C and why, right? And again, you guys probably have been paying attention to what's going on with Nvidia, but for people maybe, you know, more infrastructure folks, essentially the thing is here, right? I ask a question to whatever large language model, the first thing that it does is trying to figure out what do I actually care about, right?
There's different words in a statement. The what, the, the, uh, what, what is this guy actually asking about versus some of these words that don't make sense? So I calculate that, turn it into key, uh, key value stores.
And that's essentially the context of the conversation. Something else that adds context is maybe a document or a video, right? Someone says, Hey, I wanna just summarize, you know, solid I'm and vast tech field day, I'm gonna upload the video into my favorite large language model.
But if I ask a question, again, it used to be I had to calculate all of that context again, right? So a question, maybe not the end of the world, but if it's a document or a video, then I'm calculating that a lot, right? Think about some big enterprise organization dumps a new document out to the world and all of their employees are asking questions about it.
I'm recalculating that same context on that document over and over and over again, right? And then obviously I need to make sure the decode phase is the answer part. That is where I'm creating an answer that makes sure it's related to the question that you asked.
So there's some big problems with this context piece, which is one, if I'm recalculating over and over and over again, I'm burning GPU cycles on something that's not adding a ton of value, right? And honestly, GPUs are too expensive for that, right? We had, I think NVIDIA's customers are like, we can't just keep dumping all this CapEx and scaling forever.
You gotta help us use these things more efficiently. That's step one, problem two, user experience. If I'm a user and every time I ask a question about a document, it's going to do a bunch of work that's so annoying.
I wanna engage in a conversation with you, not have you forget what we're talking about every time I ask a new question, right? I think, you know, some of these large language models, they know everything about you because they have all of that data. So that's KV cash.
I wanna store that context so I can continue, continue to reuse it, right? And Nvidia has this hierarchy, which Scott talked about. So step one, stored in high bandwidth memory, right?
Obviously there's some challenges there. It's really expensive and hard to come by right now. Uh, the other piece is it's local.
So if I'm engaging in a conversation just on, you know, this one session, that's fine. But if my friend is trying to have the same conversation, you know, do I have access to that memory? Then I can move it down to dram, right?
Then I can move it to local SSD. Again, everything's local. And then the next step is shared file and object, right?
When you're gonna have a massive drop off in performance there, EastWest traffic, all those different things we talked about was shared nothing. So we were like, we need something right here, right? 5, something that has a performance closer to local, but is more global in terms of its access, right?
And that is essentially what I-C-M-S-P is, right? How do I create that local feel? Now again, I'm not gonna spend a ton of time on this, but one way to do that is to basically take a shared nothing architecture, right?
I can either put, you know, a client on the blue field, or I can deploy my software, right? On essentially the CPUs in these g um, GPU servers. Now, again, problems there is I'm bringing the problems of that shared nothing architecture up into my most expensive assets, right?
I have EastWest trap happening, right? I might have a hotspot, right? Where everyone's asking about the same piece of context.
That means I have all my GPU servers attacking one, essentially and asking it for information. I don't know. That's, that's a good idea, right?
So what we said is we've got, um, a different architecture, right? We walked through this shared, shared, uh, everything architecture where I've got this stateless layer. Now, I would say, you know, some potential challenges with this instance of the deployment, really two, right?
And again, I think, I don't wanna say challenges, but optim areas for optimization one, right? I have this layer of CPUs that is essentially kind of in between my access to SSDs, right? Um, and as we know, right?
The CPU is always gonna be the bottleneck to SSD performance, right? You think about the world's most powerful processors. How many do I need from a thread perspective to saturate one single 122 terabyte drive?
It's a lot. So we have that problem. The other problem is, you know, I have to essentially create a copy of data, right?
I've got an RDMA operation to our front end, and then I've got another RDMA operation to the SSDs, right? So you kind of have this hop that's happening. What we're able to do with I-C-M-S-P on Vast is actually take our logic, our C node, and move that up to run on the blue field.
Now, we actually, uh, introduced a prototype of this style architecture, uh, I think with AI Field Day, um, earlier, but that was with the previous generation of Blue Field, right? Blue Fields have gotten dramatically more powerful from a core count. So now I can run my C node, the logic of the system up in those blue fields.
This is a paradigm shift, right? I no longer have a host going through other CPUs to basically get in line to get access to data that's on fast media. Now, every node has its own little friend.
That's its protocol server. There's A storage class memory here, Phil. It's still in the dbox down there.
You just can't, it's like, yeah, there'll be two layers of storage in that dbox a storage. And the old Way, the storage class memory was also in the dbox. Yeah.
Nothing changes there. It's just how I give access directly from the host to that device. Okay?
Yep. So we don't have to change anything there, right? Which again, now, instead of having to need to use the local SSDs and introduce potentially, you know, conflicts and all the different things that we might have by putting software on those servers, I can just have JBoss, right?
Full of flash with dense, solid, IM sds and everyone has direct access directly to the metadata structure and directly to the actual data itself. And again, scaling and everyone sees everything. So there's no problem in sharing context, right?
If there's a hot piece of context, everyone can access it with a, a whole bunch of parallelism. Uh, but I'm not gonna have any hotspots up top. So Phil, excuse me, Phil, Jack Poller with Paradigm Technica.
It sounds like what you're really doing here is you are running storage controller software on the GPU because you've got spare GPU cycle. It's on the blue field. So I have a GPU server, I'm putting a blue field, which is like a smart nick now it's got 40 cores in it.
We're taking those cores, which you don't need for network performance 'cause it's just more cores than you'd ever would. And we're running our storage software there. So it's not in the CPUs, it's not on the GPUs, it's on this little server essentially that's mini, that's let's smart Nick, but a lot more powerful than that.
Okay? And the net effect of this is, So ultimately what you get, right? So we talk about certain things, right?
One much faster time to first token, right? So if someone's asking a question, I now already have that context, I'm sharing it globally, right? So if anyone's asked about anything, So in, in, in a traditional architecture, then you are making a request from the GPU to a storage controller that's off host.
Yep. Right? And then that storage controller goes fetch as the data feeds it back.
Correct? And so in this case, what you're doing is you're moving that storage controller on host correct. Or a little bit closer to the GPU.
Correct? And that's getting you, that's accelerating significantly. Significantly, okay.
Yeah. And I think there's a few things, right, that come into play. So you've got the acceleration of taking out an RDMA operation in the middle, right?
Uh, you have a, a scaling advantage of the fact that now every time I add a new host, I'm adding compute specifically with that host. That is its own storage resources, right? Essentially it's dedicated.
So I'm taking out all the potential conflict, right? You get resources for you. You don't have to fight over them with your partner.
When I have that shared CPU pool, we're all fighting for the same resources, right? So it's a scaling, it's a parallelism and it's efficiency perspective. The fact that the data is now shared means that I can have more GPU servers able to share more context.
They're much less likely to calculate things again, right? So that's why I get a faster time to first token because I can pull that context without having to recreate it. I get much better GPU efficiency because my GPUs are not recalculating the same things over again.
They're actually doing inference instead. Right? And then the final piece is I'm taking out that entire compute layer, and that all is power that's drawn and power is precious now.
So by taking out that entire server CPU group, I cut power by 75%. Got It? Make sense?
Mm-hmm. Okay. So can can, can you tell us again what I-C-M-S-P was?
It's context management, something, something. I think it's inference. Context management storage platform.
Okay. Inference context, memory storage platform is what they called it. And if you Google it, just be careful.
There's a whole bunch of other uses of the acronym. I see. The S-I-C-M-S-P just think of it as really cool close storage.
Okay. And I keep thinking about the image that you put up that had the different layers and had context as one of those layers. So, okay, so is this vast I-C-M-S-P, that's what's being attached to the blue fields.
So essentially it's I-C-M-S-P is, um, think about it like Nvidia announced something called Dynamo, right? And we were working closely with them on that. You've got like all these different problems in terms of, you know, how do I manage where inference jobs run on GPUs, right?
How do I make sure that I am, uh, intelligently using the different tiers of me of, you know, memory and storage just for context, right? So that's just for context. Um, and then how do I scale and run those things?
And basically the I-C-M-S-P is a tier of storage and essentially a standard way that Dynamo's gonna interact with that storage. So I'm basically saying we're lining up to saying this is how NVIDIA expects to use extended, you know, off, um, or shared storage for context. So Can you go back to your diagram that shows?
Yeah. Okay. So where is it on this chart?
So Essentially this is gonna be used for context in this world, right? The denotes Yeah. So all the da, all the, the, sorry, I'm, now, I'm not supposed to point to the screen.
All of the context gets stored down in that DBOX layer in the same way we would store any type of data. Okay. And so that is y'all's I-C-M-S-P is gonna be in the dbox.
Correct. Okay. Thank you.
And the data lake that's also underneath that is stored there as well. You Easily can. Yep.
So, you know, something that I would talk about, 'cause I, I, let me go to the next slide and maybe it will, it will help. So one of the things that's going on right now around I-C-M-S-P and KV cache is the question, should you use any data services? Right?
Because could they impact performance? Right? And we went through that whole shared nothing thing.
And certainly if you do use data services on a shared nothing architecture, you're gonna have challenges. Data reduction is one of them. Another one is something like encryption, right?
So right now we're not sure should we encrypt that data. I think it's a really bad idea to not encrypt that data. It's a giant shared thing which has everything about every conversation that all your employees are having with ai.
You might want to encrypt that, right? And you kind of have to save it again. But there's also a chance for data reduction from what we've tested.
3 to one and two to one data reduction, which means, again, And, and again, this is an inference solution, not a training Solution. Correct. It's all inference at scale.
Um, so the point is, you know, in our world, this is how we would do it, right? But what's unique and flexible about the vast world is those blue Bluefield controllers don't have to be the only CPUs that that cluster has access to. We can create actually a sidecar pool of compute that just does data services.
Because remember, those blue fields are gonna write data down and let's say storage class memory, we can then have this pool of compute, take that data, reduce it, store it back down to the QLC, right? But I can also attach other workloads over here, right? And what I know is that these blue fields all get their own dedicated amount of storage performance.
Every blue field has 40 cores sitting. So In the other prior solution, the blue fields were actually responsible for the D deduplication. Yeah.
Yeah. We, without The other compute side correct. Cluster.
Correct. Which they, you know, at this point, this is a very new solution. We're gonna have to do a bunch of tests.
Will they run hot? Will we want to augment? Will it be enough?
Uh, but the point is, we're the only ones that have this flexibility to say we're gonna add compute that you can then leverage for data services. So essentially, I just care about the fact that these guys can access data and then let our friends over here take care of all the data services In the old world that these sort would've had to have been homogenous nodes. But in this environment you're taking, you could put any, any cluster of compute services out there to be your C nodes, correct?
Mm-hmm. At this point, it's very flexible. We have customers with multiple different generations of compute running c nodes in the same giant environment.
Mm-hmm. We can actually take, we can pool them, right? So we can say basically, you know, certain applications can use two C nodes and the rest can use 20.
There's a lot of flexibility in how we can carve this up, that disaggregated shared everything piece gives us flexibility in a way that wasn't possible before. And you mentioned storage class memory down at the D nodes. Those are, uh, different types of SSDs that are tailored to, you know, read, write access and things like that, right?
So essentially it's an SSD, you know, lower latency, they're more expensive, uh, much better endurance profile. Mm-hmm. Right?
So essentially, you know, when we came to market in, uh, Intel, I'm losing word PM you know what it is? PM Optum. Yeah.
Optum. Optane, Optane. See, it died such a long time ago.
It was all we had. Soy has a, uh, P 58, 10 SLC based SSD that is used as a storage class memory solution for these guys. Mm-hmm.
So there you go. So yeah, it's all about endurance profile, cost latency. Thank you, Marian.
I have a quick question. Sure. Uh, going back to similarity, how do you keep the similarity decisions as the data and the models change?
Yeah, so that's really just the underlying data structure, right? So it's totally abstracted from models or anything else, right? So as data comes in, essentially we're hashing it.
And normally with ddu we use a really strong hash, right? That, 'cause you wanna make sure that I never accidentally mistake two pieces of data for the same. That's why DDU uses a strong hash.
All we're doing is taking a chunk of data and using a weaker hash. And that weaker hash basically says that we don't use it to actually store the data. We use it to say, Hey, this data's very similar to data we already have.
So it doesn't matter if that data came from an inference job, a backup job, whatever. We will look across any data that's been stored on the system, um, doesn't matter what protocol it landed in. And we will identify that there's commonality and we just won't store it.
And the system has no idea this is happening. And you're using a weaker hash for, um, Comparison. Comparison.
Comparison, yeah. Because if you use a strong hash, you can't tell. 'cause when I use a strong hash, a small difference in the data, a results in a very different hash, right?
That's why you do that. So a weak hash basically says, if the data's slightly different, then I'm gonna get the same result. So now we know we're in the zone, right?
That these two pieces of data are very similar. And what are you hearing from your customers who are in highly regulated industries? They have no issues with it.
Right? Um, again, data's now typically on a, on a system, it's chunked up, it's erasure coded, it's spread around anyway. So at this point, you know, it's all generally pointer based.
Uh, I haven't never heard anyone have issues that they would have to turn it off because of some kind Ation. Okay. Thank you.
Sure. Thank you. As far as the KV cache and actually doing any special caching of the data coming off the blue field versus it's all going to storage class memory when it's written, and it'll be re d gets stooped to QLC or whatever the backend is.
Uh, it's not like you're holding that data in storage class memory or anything like that. We Are not. So, you know, in general, KV cash, we'll use the different types of media available, right?
So it could land in the memory, the high bandwidth memory, right? Or it could land in, you know, a local SSD, uh, this is essentially another tier of KV cache. Um, and for us, yeah, we're, we're gonna keep all that metadata in storage, glass memory, but we, you know, there's, there's nothing different about how we have to store the data.
Essentially. We've just created a more optimal data path. Hey everyone, welcome back here to Text Drunk tv.
I'm, uh, happy to have our next guest on, he's the founder and CEO of a company named Etle. It's his first time here on Text Drunk tv. Let's welcome Christian ramming.
Ramming. Is that correct, Christian? Yes.
That that, that's right. Yeah. Thanks.
Thanks, thanks so much for having me. Right? I got it right On the first try.
It's a good day. Christian, welcome. Thanks for being here.
As I mentioned, you're the founder and CEO of that leap. We're going to get into that leap in a second. But let's, let's start with, you start with, you know, where, where was, what, what drove you to found Edle and kinda where have you been and you know, that brought you here?
Yeah. Yeah. I, uh, I'm an engineer by training.
Um, I've, I've been, uh, been, uh, living the silicon battle life for the past couple of decades now. And, uh, I, uh, I started atle about, uh, a decade ago. Um, uh, I was CTO at an adtech company before that.
And, uh, that's really, you know, there, there's, there's essentially an, an itch from there that I'm scratching with atle. It's, um, um, you know, we, uh, I was, uh, you know, heading up engineering and analytics and, uh, I hired, had a whole little engineering team just dedicated to building data pipelines. You know, we needed to make our analysts and scientists productive.
And, uh, that just ended up being a tremendous amount of engineering work. So, um, so that's the, you know, the, the short question of the story is that, uh, I wanted to take all that engineering work out of, um, you know, something that companies needed to do. And, and so that's really what we're doing at atle.
We're building software to, uh, take all the data engineering work out of, uh, data pipelines. Interesting. And you say you, you started this a decade ago, so that was before, you know, chat, GPT, generative AI and all of this stuff that's going on today.
Did you ever think a decade ago that what you, what you were setting out to do would one day perhaps be enhanced or, or augmented by something like Gen ai? The, the world is So the short answer, no. The world is a crazy place.
I mean, I used to, when I started, I used to have to explain to people what a data warehouse was, you know, so it's, yeah, I, I, There's also True, I, I know longer have to do that, so, yeah. That It's Come a long way since then. Absolutely.
So, you know, for those of us who don't know at Leap give us the last 10 years in a nutshell, Christian, We have, uh, a software product, um, or we've been offering a software product that works great for, uh, data warehouses, uh, in the cloud, right? Redshift, snowflake, Databricks, and so on. And so what it does is it ingests data from, uh, a bunch of different types of sources, right?
Databases, applications, streams and files, uh, brings it all into the warehouse and then makes it easy for data teams to sort of model, uh, data to, to turn it into data products, um, within, within the data warehouse. And so, uh, so yeah, so we've been on a mission to, to, uh, reduce that data engineering effort and also reduce the time that it takes to, to get all this set up right. Um, that's, that's another big part of it.
You want to, you know, get your, your analysts and scientists productive, right? Uh, as, as quickly as you can. And, um, you know, have your engineers spend time on more, more interesting things than, uh, than data plumbing.
Um, and I, I think something that's, that sort of set us apart in the, in the marketplace is we, we work with, um, companies in regulated industries, right? Financial services companies like Morningstar, and in healthcare, right? Companies like Moderna, um, they need to, uh, keep their data within their own firewalls, essentially, right?
Like within their virtual product clouds. And, and so that's, that's something that our, our, uh, our software does. Excellent.
Excellent. Um, you know what, before we jump into Iceberg Pipeline platform and all this, let me close the loop on that leap for people who may be, say, sounds interesting, I'd like to learn more. What's the best on-Ramp forum?
Yeah, I mean, uh, uh, it's, it's pretty straightforward. com, uh, sign, sign up for a demo. We have a, we have a great team that's, uh, uh, you know, ready to to, to, you know, show you exactly what you need to, what you need to see.
Excellent. Cool stuff. Alright, let's move over.
Now, you guys recently, uh, introduced a, a new, uh, platform, the Lib Iceberg Pipeline platform. Yeah. Oh, that, that's right.
So, um, I guess maybe a little bit of, a little bit of background. First, we, you know, we, we, something we've, we've heard a lot over the past few years is, you know, these data warehouses are great, but, um, you know, for, for different reasons, maybe for data sovereignty reasons, or because people are adopting these new AI use cases, um, they, they want to, um, take on Apache Iceberg, basically use data, uh, use Apache Iceberg as the data foundation. And, uh, for, for, for data pipelines, there's a bunch of consequences.
When you, you, you want to run pipelines with Iceberg, there are things that really should work differently. So what we're, what we just launched is, as you said, the, the Iceberg Pipeline platform. Um, it's basically a, a pipeline there that's built from the ground up to support, uh, to support Apache Iceberg and everything that's different about it, right?
So it's, you know, still doing ingestions, still doing modeling, but, uh, but also tying that all together, right? The, the dependencies between the different parts of the system are, are become really important. Um, use cases like streaming, which I'm sure we'll talk more about, are, um, uh, you know, are now possible.
And, uh, and also there's a lot more to do, right? You need to do table maintenance and, uh, um, you know, snapshot removal and, and, and that sort of thing. And so, uh, the Iceberg Pipeline platform is essentially a, an sort of an all in one system that ties all of that together, uh, and can, and can run inside of your own virtual private cloud.
Excellent. Now, of course, this is based on Pache Iceberg. That's right.
Let's say someone out there says, iceberg Titanic, I never heard of Apache Iceberg. What, what exactly are we dealing with? Tell us, you know, if you wouldn't mind, give us a little Apache iceberg background.
So Apache Iceberg has been around for, for several years now. Um, you know, it's an open table format that's particularly well suited for, um, uh, you know, analytics use cases. And, um, you know, it's not the only open table format out there.
Um, it has, uh, that there, there are other ones too. But I think that the thing that's special about Iceberg is that, uh, a few years ago, the, you know, the, the, the massive vendors in the space, right? Snowflake, Databricks, AWS Google all essentially decided to throw their weight behind this one, right?
And so they said, look, we, we, we think this, this technology shows great promise. We're gonna make our products work really well with this specific format. And so that, that really helped, uh, future proof it, right?
Because now it became, uh, much easier for enterprise to say, yes, okay, well, the big the, the big guys are behind it, we can adopt it too, right? And so that's, um, uh, you know, so that, that's, that's, uh, uh, uh, a big part of the reason why it's now seen as a default choice, right? So, um, so, you know, the more concretely right, you can, uh, store all of your data within, uh, the iceberg, uh, uh, table format, um, which is, you know, built on top of Parquet, right?
Parquet files that live in S3 with some metadata layers. And then you can, uh, you can use tools that you already use, right? Snowflake, Databricks, Athena, uh, BigQuery, right?
Um, these engines work really well with the data that's stored in Iceberg, right? So you can, uh, uh, you know, you can get, you can still get your fast queries, you can still support your, your AI use cases, right? If you're using Bedrock, for example, that plugs in very, very nicely.
So you, you know, you own your own data and you continue, can continue to use the tools that you already used in your stack. I love it. Excellent.
And you know, the nice thing is we, we see it with the foundations like Apache, like Linux Foundation, eclipse Foundation is it, it fosters this, um, cooperation where companies that normally wouldn't necessarily cooperate Snowflake and Databricks, right? You don't see them around the campfire singing Kumbaya too often, but you know, they all, they're all behind and they're all supporters of Iceberg, right? As, as something that transcends individual companies like that.
Um, is, is the Iceberg Pipeline platform available now? Pipeline platform is available. We have, uh, customers using it, successful in production, and, um, uh, as I mentioned, you can, you can run it inside of your, your VPC so that, you know, no, no data's flowing out, uh, you know, while, while your pipeline's running.
Very cool. Um, you know, while Apache Iceberg is free and open, I'm sure, uh, the, the, the, uh, excuse me, the, uh, pipeline platform, is it open source, not open source SaaS kind of thing? How do you, how do you do this here?
Well, there are, um, it's, it's, uh, proprietary software. Um, we, uh, integrate very tightly with, uh, you know, other sits on top vendor, vendor space. Yes, exactly.
So it's, it's very tightly integrated with AWS, right? So if you have, you know, if you have a glue catalog, right? If you have, um, you know, you use Athena, um, if you have, you know, CloudWatch for monitoring, right?
It, it sort of plugs right into that. Excellent. And then did we mention, I don't know if we, did I forget, did we mention the, uh, ATLE website?
com? That's right. com.
Yeah, lots of, lots of information about the, uh, uh, the iceberg integration there as well. Um, so you can kinda see, see how it's different. You know, I think that, uh, one of the things that sort of sets this apart from, um, from sort of traditional ETL, right?
You could argue, oh, isn't this just, isn't it just, you know, ETL and a new, uh, you know, for the new, with a new sort of destination? And, and I think this is, uh, you know, it's really, this really is something else. You know, ETL tools move data, but, um, but there's sort of a fundamental difference with operating pipelines end to end, right?
It needs to be, uh, you know, tying together ingestion and modeling. There are, uh, you know, catalogs involved and also these downstream tools that need to keep their own metadata, uh, fresh about the table. So in order to kind of get a, uh, a system that sort of hums end to end, you need, you need, you know, whether it's athlete or something else, you need something more integrated than, uh, than a sort of a traditional ETL tool that will sort of dump your data into iceberg.
But, but that's, but then that's sort of it. Fair enough. Christian, we're about outta time.
I want to thank you for popping in here, man, the time goes very quickly, as you could tell. Um, before, before we end anything else, If you gonna, if you're gonna remember one thing about atle, I guess, um, you know, we, we, uh, handle ingestion modeling. Uh, we operate your, your pipelines end to end and we run inside your VPC.
Thanks. Uh, excellent. Thank you.
A really appreciate it. I appreciate it. Christian roaming founder, CEO of Atle here on Techstar tv.
We're gonna take a break. We'll be back with more. Stay tuned.
Hey guys, thanks for the throw. We're here with Peter leaves the CEO of sim space, and we're talking about, well, cyber ranges as it applies to cybersecurity and training, and maybe just modernizing this whole process. Peter, welcome the show, Mike.
Thanks for having us. You know, I just saw recently somebody going through some cybersecurity exercises and it was a tabletop, kind of a version of a tabletop game, and everybody had their role to play. And, you know, as I was watching it, I couldn't help being reminded of, you know, uh, clue, you know, Colonel Mustard did it in, uh, a library with the candlestick.
And as fun as it was, it didn't seem very realistic. So, um, what should we be thinking about here when it comes to training and modern attacks and resiliency? Is there some way to think about all, all this that's just gonna be maybe better than what we have been doing?
'cause Well, I just don't think the bad guys are sitting around playing table talk games to attack us. Yeah, you're absolutely right. Well, first of all, I think the whole notion of training has to get redefined as humans plus AI working hand in hand.
It's no longer just humans. And I think, um, I've been shocked at the weaponization of ai. I mean, it is unbelievable today.
There's a lot of skeptics, still people who believe it's at the early stages. And that may be true, but we are moving very rapidly into an age of autonomous agent attacks at scale. You know, they're gonna do reconnaissance, vulnerability, exploitation, moving laterally, you know, the whole full, uh, kil the full, um, full value kill chain.
And so I think it has to be humans plus ai. And I think that that notion of training is critical. And I think a cyber range is the cornerstone of rebalancing the fundamental asymmetry of offense, cyber offense, and defense, right?
It's always cheaper, faster to attack than to defend the attacker chooses when and where the defender has to prepare everywhere. Yeah. Um, so that's, by the way, that's half the story, Mike, and we can get into the other half, which is testing of tools, agents, capabilities.
So yeah, as one old boss of mine said, you know, it's a lot easier to throw grenades than it is to catch 'em. But, um, when you think about this for a minute, uh, I'm not sure everybody knows exactly what a cyber range is. They're familiar with range as in like the military as an ex exercise.
They go out to the range and blow stuff up and they learn how to do things. But how does that concept apply to cybersecurity? Well, I think we should define the cyber range.
First of all. Um, a realistic and intelligent cyber range is going to be one that enables you to create a realistic replica of your production environment. So this isn't gonna be some pre-canned laboratory.
It's going to be a very rich, um, replica with full network topology. It's going to have integrated attack and activity emulation with your actual security tools and user behavior, which is chained in some ways and interactive with attack behavior. You're gonna have a very significant, um, we don't like to use the word digital twin because of that implies, um, a layer of cost and complexity, which is diminishing marginal returns for what I think a cyber range, um, can provide.
But yeah, absolutely tabletops, simple labs, individual training, these things are obsolete. They're not gonna move the needle. You absolutely need to have a realistic replica of your production environment.
And I'd say as I opened with my comment, not just for humans, but increasingly certainly what we see in our business is that, um, vendors are using it to train AI models, to test AI agents and to validate agentic capabilities. And you can't do that in an environment that doesn't closely, um, mirror your own enterprise complexity. It's gotta be like any other type of agentic or, um, AI modeling.
Gotta be very, very, um, tailored to your specifics. And how easy is it to do that these days? I think people think that there's a level of complexity involved that's too challenging.
Yeah. And then how do I test those AI agents once I set that thing up? Because, well, every AI agent I've seen so far does things that are unpredictable.
Yeah, those are very good questions. Uh, so on the first one, I think that, um, there the answer really depends on the degree of fidelity that you want for your, for your environment. So we can have, um, realistic replicas of production environment set up in a matter of a couple days, or it could take several weeks depending on the level of fidelity.
And that'll depend on the use case as well. Um, in terms of the testing of agents, it's really important to step back and zoom out. One of the things that's really, um, critical in terms is the ability to get clean labeled data to be able to actually run full kill chain threat attacks in varying user environments with every single conceivable permutation of your network topology and the different tools you use be able to collect and synthesize that log data and figure out, did this agent actually work?
Thumbs up, thumbs down. And you can imagine doing those training runs thousands and thousands of times in order to begin to bring that capability into bear. And of course, you could think about dual, you know, traditional data science measures precision and recall.
Am I really tuning it so that it absolutely is gonna detect every single type of threat, but it may throw off a degree of noise, or am I gonna reduce the noise, but I'm gonna let some of these threats through. So there's, I think there's some significant capabilities and, um, tailoring and customization that are capable in the cyber range, um, uh, uh, platform, but that's going to be user and use case to bend it. Mm-hmm.
Have we gotten to a point where we can't quite just put up a standard red team because the number of vectors that could be exploited are just too broad and there's too many things that could possibly happen. And who knows, maybe we'll use AI agents themselves to create the attacks. But as one security expert said to me once, if you can imagine it, somebody's probably trying it.
So how do we deal with all the exponential possibilities? Well, I think v on your original question, can we, uh, are red team still useful, let's call it, or, or rely on? Absolutely.
I think that, you know, again, I, I would go back to my, my touchstone of humans plus ai. There are things, certainly, if I can speak on behalf of our red team, that they're capable of doing that are extremely sophisticated and would, um, would not be, uh, available for agent training. They haven't, the data doesn't exist, the penetration methods, the tooling that they've developed.
That being said, I think what we see is we have actually, um, a fairly expansive set of ecosystem partners. There are a number of red team agent businesses that are beginning to, um, develop their capabilities in our cyber range. And what they wanna do is they wanna partner with us.
So as we run these realistic, um, training exercises, team training exercises, they're going to be able to, um, have a menu. You could think of them as a, to offer capabilities against the, uh, for the enterprise clients to test, test their teams and tools. So I do think that that is going to be the future where you're gonna see a variety of different, um, scenarios.
And I think the, um, the cyber teams that are really focused not on compliance and check marks, I'm gonna say not on individual training, but on realistic team training that actually helps to outsmart adversaries in any cyber terrain. I think you'll see a tremendous amount of, um, interest in both red team and blue team agents. We're seeing both, um, both parts of the ecosystem get attracted to our, our cyber range, um, and of tremendous interest to clients who want to do team training at, at, uh, at realistic scale.
Some folks are concerned that perhaps we're a little bit over our skis when it comes to AI agents and we're deploying these things without thinking through the security issues. Um, are you at all concerned that maybe we're just waiting for some sort of catastrophic event before everybody gets serious about AI training and testing? Mike?
It's a great, it's a great question. I think one of the other elements of that, what we are actually doing is to help, um, validate the, uh, AI agents themselves. Anytime you introduce a new capability, you open up a new attack surface.
And the reality is that, um, bringing on a blue team agent to help with your SOC preparations, but opening up new attack vectors without being prepared for it would be a terrible mistake. And so we are actually creating kind of an agent reading system or validation system, and we're, um, taking a hard look at the, um, at the vulnerabilities that the agents present themselves. It's a fantastic question.
Mm-hmm. Do you think on the other end of this, that auditors might start using these platforms to test and validate the environments that they're being asked vouch for? You know, that's a fascinating question.
We actually see, um, it hasn't yet proceeded to the auditor level, but certainly the cyber insurance providers are taking an active interest. And we've had several of our enterprise clients let us know that we've been instrumental in reducing cyber insurance premiums because they can actually demonstrate progress on a longitudinal basis against the varying stage and progression of threat actors. So I do think that the insurance market is going to be the leading edge, but I would not be surprised if we see, um, multiple other, um, uh, constituents take advantage of that.
No doubt. So as you look at it and you think about testing and training and everything that goes with that, what's that one thing that kind of just makes you shake your head a little bit and go, folks, we need to be a little bit better than that? Oh, yeah, sure.
That's an easy question to answer. I think if you look at our large enterprise clients or large government clients, they are literally using hundreds of tools. They are never going to be optimally configured.
They are in a continuous, what I call process of selection and optimization. Many of these tools have overlapping vectors. Um, they're continuously evaluating and adopting new tools, and there isn't a programmatic, um, value-based, efficacy based way to actually, um, frame decision support around how do I configure and optimize and rationalize my tools in order to improve my security posture.
And so I think that to me is just a continuous, um, cyber range use case that we're, we're really excited to, um, be engaged with our clients. Ryan. All right.
Hey, folks, you're heard in here. If your training and testing is rooted in the last century, probably not gonna get the outcome you're hoping for. Peter, thanks for being on the show, Mike.
Thank you for having us. All right. And back to you guys in the studio.
Now, as you continue to move to your new infrastructure, um, you've got, um, you, you have production nodes and, uh, all embedded on the new, the, the new operating system, the new management platform. What are you doing differently now, carrying forward? You know, of course we've talked about digital twin, uh, to a certain extent, but, uh, you know, you've talked about integration with CICD processes and so forth.
Yes. What do, what do day two operations look like for the whole team now on the new platform and infrastructure? Well, uh, currently, you know, um, on the new infrastructure we have, uh, the team, now, it's becoming more, um, uh, you know, date, the date, day two plus, it's becoming more, more routine kind of changes, uh, as we are of course moving the, migrating the data centers and, and building our existing data centers on.
But, but also from an architecture perspective, we are now trying to do more in terms of, you know, implementing more AI functionality. Mm-hmm. To, to now to tell us more about the operational, you know, if something goes wrong, if an incident happens, uh, don't give Mero raw information, even if it's correlated.
Even it's correlated. Gimme, gimme an RCA, gimme an rca and point me to what, what could be the problem and what could be the impacts. So may maybe something has happened, but actually there is no real impact to the network.
So a link might fail. Mm-hmm. But there's enough in redundancy that we shouldn't worry about.
Sure. Tell, tell me. I don't have to worry.
So we're trying to build this more of an intelligent intelligence that there is some entity that is entity, uh, that is aware of our topology. Sure. Art topology.
Yeah. And our network that can actually give us useful information about incidents and the results. But the main, the main thing that's interesting also in the operational aspect is I, I, I talked to the operational team and, you know, they, they, they have a huge reduction in, in incidents, you know, like, like 80% to what we had before.
Yeah. Not, not only, not only the tool is part of the aspect, but the, I think the, i, the NetOps and tracking and this idea of doing things before you implement them helps a lot in avoiding, avoiding problems. So, so the, the operations team are saying that, most are saying to me that most of the time they, they don't find the same problems that with what they found before, the fabric is very stable, very consistent.
Stable, consistent is as, as expected. There is not nothing that surprises them. Most of the things they're dealing with is maybe access to, to servers, uh, low balancers, firewalls, physical, physical stuff that is going on in the data center, but not from a, a design or, or a routing or, or this kind of, uh, perspective.
So it shifts, it shifts their, their, their efforts, you know, away from debugging, you know, lower level Sure. Details that the fabric should, should take care of. So it's a, it's a different, different type of challenges, but much, much less than before, much less before.
So that's a great example of the tooling decisions that you make, really having an impact on your operations team. Right. And you're dropping 80% of your tickets.
Um, that's, um, and, and I know, um, you know, your leadership has, has cited the same statistic, you know, 80% reduction in tickets. That's, that's huge. That's certainly impactful.
Didn't you also make some decisions on how to, um, implement EA, like you could have gone with a local install for EA, but you opted to go with the SaaS option. What were, what were some of the thoughts you had to, to weigh there and why did you ultimately choose the, you know, EA SaaS version? Okay, so, so the IDA SaaS really, uh, we, we as a network and, you know, network team, we're not in the business of running ida, you know, or, or we want to make use of it.
Sure. We want to architect around it. So, so we offloaded that, that effort to the IDaaS team where they, they can implement it in the cloud wherever what, whatever cloud they want, they want to go to, sure.
To GKE. They want to go Google or AWS whatever they want, right? We don't care.
We just want the connectivity to it. We want, you know, the upgrades to happen. We want them to monitor the health of IDA itself.
You know, uh, maybe their, their e comes with many extra add-ons, you know, with RCA, um, you know, correlation. Uh, they do some kind of AI into, into giving you, in, you know, analysis of anything that happens. So, so that's, that's also a plus that comes with it.
We really didn't want to, to think about either too much. We wanted to rather use IDA architect around IDA and, you know, think about our, our architecture more and building more enhancements to the solution itself. So IDA SaaS in that perspective, uh, saved us a lot of effort.
And, and the team there was very helpful for us. So we, we like that. We, we, we were, we are glad we went that way.
Not in the business of actually running ida. You just want it to work. Right.
So that's a great encapsulation of that. Very good. Yes.
Well, what else is in store in the future here? You know, you've, you've laid some really important foundations here for, you know, making real improvement for the resiliency of the network and the lives of the operators. Um, what, what's coming down the pike?
What, what are you gonna be able to do in 2026 that you couldn't do in 2025? I think, I think now we're con concentrating on enhancing, enhancing the solution using maybe more AI that is native, you know, native to either, either is coming up with lots of capabilities that are AI related, that we can chat, we can chat with either, we can talk to it, we can tell it what went wrong, what, what's happening. So that part, you know, we used to actually, we build in our solution, but a way or around either to, to, to fill in that gap.
But I think in, in the next release, we're gonna be merging some of our tools with IDA and maybe merging our automation into IDA itself. IDA has the capability to write your own apps. Sure.
We would, I, I think that's something we would like to do, is, you know, bring in our, our automation and pipelining to, to, to be within ida, to be within, to be IDA native. So the user doesn't really have to think, I'm dealing with Git, I'm dealing with ida. Maybe we can merge something together there.
And also more on the operations aspect to get, to make operational RCA more intelligent, which IDA will add, and we'll see what it adds and try and fill in the gaps there. As well as, as you know, we're still Nokia's growing. Mm-hmm.
We've got, we've got acquisitions and stuff like that. So there will be more data centers. And, and as I said, our, our architecture now is very modular.
It's not flat. It's not like it doesn't look like a data center. It looks like many small mini data centers connected together.
Sure, sure. So we are able to absorb, absorb those changes. And that's, I think will be a main task for us next year.
Ahed, thanks so much for talking with us today. It's been a pleasure. I look forward to hearing how things progress in 2026.
Thank you very much for having me. It is very exciting for me, and I'm very passionate to, to, to talk about our work. 'cause it's, it really is exciting and interesting.
It comes through in every conversation I've had with you. So thank you. Thank you for sharing that passion and turning it into real concrete.
You know, this is what you're actually experiencing in the new, uh, the new network infrastructure in, um, Nokia Enterprise. It. Thank you very much.
Thank you for having me. AI isn't just for software. As companies like Flux AI are leveraging generative AI to act as a designer to bring your ideas into the real world, they enable a user to design electronics, creating the schematics needed for a contract manufacturer to produce it Flux AI is also a community of users for support and ideas.
Will AI enable on-demand manufacturing of just about anything we can dream up? That's the question on this episode of utilizing AI with Olivier Blanchard and Mathias Wagner. Welcome to utilizing ai, the podcast focused on practical applications of artificial intelligence from the Futurum Group.
Every Wednesday, we explore news and use cases of the way in which AI is transforming enterprise IT, and the industries it serves. I'm your host, Steven Foskett, president of the Tech Field, a business unit here at the Futurum Group, and today we're talking about bringing AI into the real world. But before we dive into this discussion, let's meet who's on the panel today.
Hi, I'm Olivia Blanchard. I'm a research director at the Futurum Group, and my primary focus is intelligent devices. So anything that has AI embedded in it, from rings and other wearables to PCs, to the IO ot, to robotics, to cars, uh, is, uh, is the, the stuff that I focus on.
Hi, I'm Matthias Wagner. I'm the founder, CEO at Flux. At Flux, we aim to take the heart out of hardware, and we're building the first AI hardware engineer to help go from a, from a prompt to any physical product you can imagine.
Excellent. And, uh, again, I'm Steven Foskett, the organizer of Tech Field a and the host of, uh, utilizing ai. So I am an iot enthusiast and all around nerd, as you can see from the background, if you're watching the video.
Um, and of course, it's pretty exciting to think that somebody like me without a hardware design experience could develop a specialized iot device just by talking, yelling at my computer. Uh, but I think, uh, Mathias, that's pretty much what Flux is doing, right? You wanna tell me a little bit more about what it is?
Yeah, totally. Um, yeah, I mean, I think you're spot on, right? Uh, you can come to Flux and, you know, just like you think pt, you can enter a prompt, you get a poem or some marketing copy out of that with Flux, right?
You can enter a description of a device you want design and manufacture, and then Flux will go and design that for you. But it'll source all the components, figure out the architecture, you know, ask follow up questions to nail down the details you haven't thought of, uh, and then go ahead and design that for you and know, and get that manufacturing ready. And that can be really like, I mean, um, pretty much anything, uh, uh, I'd say like, you know, there's like simple devices like, you know, think controller for farming irrigation systems or so to like the kind of controller that in vending machines all the way to like, you know, uh, the kind electronics that go into a satellite that go to space, right?
So that's kind of like a range of applications we've seen so far. Yeah. So one of the questions I have for you actually is, is what, what level of proficiency in, uh, PCB design, uh, is, is this really for, can, can someone say a, a, a tech savvy, remotely tech savvy farmer, for instance, uh, who wants to build, uh, their own drones or their own equipment or their own like, sort of sensor technology to do very specific things on their farm or in, in their ag agricultural, uh, uh, facility?
Could they just come to this and, and use flux to go from zero to complete design? Um, or do they need to have some semblance of, of training already or experience in that field? Yeah, great question.
Well, I think you're spot on there, uh, with the user segments, right? So that's exactly the kind of users that come to us today. And we do have, indeed Farmers who are building their own farm automation equipment.
And so I'd say like anyone with like a somewhat technical background can gr us today, of course, what we're working towards to, right? It's like these, you know, if you think about how easy it's become for like 12-year-old kids to build iPhone apps, we wanna enable them to build their own iPhone down the road or an iPhone equivalent device, right? And that's what we're aiming for.
But today, yeah, it's mostly technical crowd. These are not, uh, uh, uh, uh, users who have necessarily designed A PCB before, but they have like, maybe like a mechanical engineering background, uh, or industrial design background, right? Or, or software engineering background, right?
Um, and then, you know, we, we make this as easy as we can here for them to design a board and look, in some simple cases, you can really go from a single prompt to the fully finished thing. And in other cases, you know, take a little bit back on forth, you know, on being more specific about what you want, um, and seeing it in that direction. So I think it's not unlike your typical like, you know, chat tip session, but like sometimes the first prompt you put in, you get exactly what you wanted, and other times you have to nudge it a little bit into the right direction.
Okay. That's pretty good. So what type of, um, what type of equipment do you need for this?
Is this just like, just any laptop? Do you have to have minimum specs to run this? Uh, and, and I also assume that, um, just from looking at the website, right, there's, there's certain types of files.
So you have to be able to run some kind of design software on this, uh, like CAD or something like that, or is it all, uh, cloud-based and, and included in, uh, in flux, It's batteries included, right? Uh, it's included, it, it, uh, flux runs in the browser. You need nothing else.
It's like an, you know, I always say you can go from like the idea to the fully finished product, uh, uh, all within flux, right? So no, you don't need anything else, uh, in terms of specs. I mean, look, I think like any kind of machine made in the last five years or so, we'll do fine an older one probably too.
I think it depends a little bit on the size of project you have, like a larger, more complex project, take more memory, and then sure, maybe you want high-end machine, but I think, you know, any average machine made in last five to six years will do fine. Cool. Okay.
So, um, following on this, this line, we'll, we'll switch in a little bit, but I'm just, I'm interested because I might actually be interested in using this. Uh, and I don't have necessarily the practical skills required to do this, um, even though I have the theoretical skills as, as an analyst. Um, but I, I saw on your website that, um, there's a, there's a human support element to this where, uh, users of flux can call on, uh, an active builder community, I think is how you phrased it.
Mm-hmm. And also engineers who might be on call. How does that work?
Is that sort of like a, a, a free floating community that can just use Flux as, as sort of a, a meeting point? Or do you have staff that helps customers? What's, what's the model there?
Yeah, great question. I mean, to start with, right? Flux is a platform that's much designed like in GitHub, right?
GitHub is this like pillar in software engineering, right? For engineers to work together to collaborate together, to learn from each other and to use each other's stuff, right? All these like breads and butter kind of things, we need to make software, whether it's like an encryption library, uh, a server, uh, library, whatever.
You can find those in GitHub and use those and remix them and all that kind of stuff. And so we've built the equivalent of that for electronics, right? If you think about electronics, you need the, the, the, the digital representation of a semiconductor or so, right?
And somebody needs to put that in there. And so that's either comes from the semiconductor companies themselves that are represented on flux, right? Or that comes from users who do that, or, or semiconductor distributors, right?
So there's a big community aspect to that to create this, these like nuts and bolts that you need to build your stuff there. Um, so that's on like the first layer. And then of course, yeah, we have like a big Slack community, uh, where people post jobs, you know, where people look for jobs if they have questions or they brainstorm ideas, and, you know, I think that everything is better if you do with others.
Um, and that's a providing there. And so there's other users there, there's power users there, and there's of course, also like, uh, uh, our team members are there. I'm, I'm out there, you know?
Um, and it goes all the way from like, you know, reporting a buck and getting that fixed to explaining to how something works to, like, learning what to build next for us, right? What's missing here, where the gaps, gaps, where people get stuck. Um, and so, yeah, I think, I think, you know, you zoom out here, when we started a company, it was clear about what we were doing was really big and ambitious, and it was gonna take a long time, and it wasn't gonna work if we wouldn't like work directly with the users from the, from day one, right?
And so I remember, like, we started a company six years ago, and we shipped the first, I don't even wanna call it a beta, the first thing we shipped three months in. And it was embarrassing to demo that two users, but, you know, but we learned from that, right? And we moved the whole thing into the right direction and kept focus on that.
And so that's like that spirit lives on, right? To be close to users and get feedback and, and, and do better. So once you've made the design in flux ai, um, AI can't manufacture things.
So I assume that you, uh, produce a, uh, production ready, uh, schematic that can be then sent on to a contract manufacturer, of which there are many now who would be happy to manufacture, um, you know, small or even large volumes of these things. Um, can you talk a little bit about how, uh, it goes from ideas to designs to actual physical product? I mean, you held something up in your hand just a minute ago, I'm assuming you didn't solder that together.
Uh, who did I try to avoid soldering these things these days? Because it's, it's tedious, um, because it's so tiny, you can't even see that here. But they, it's like tiny, tiny components here.
Um, so no, what's the process look like? Yeah. So today how this works is like, yeah, we, when we, when you're done in flux, we give you the manufacturing files.
These are Gerber files, there's some other format, and then any manufacturer in the world of which has 10 thousands right? Can take these files and make this for you. What's interesting in, in PCB boards is like that making PCB boards has been automated for a very long time, right?
The whole form factor here exists because it's automatable. That's the whole innovation here, right? Um, and so in that sense, can AI make this, I mean, it kind of does, right?
A robot will actually assemble this, right? Um, and what we're essentially providing you is the, the files to program these robots on the assembly line to do this for you. It's funny, um, I actually, to that point, I actually have, uh, the name tag in the front of my office at, uh, I'm, I'm at home, but at, at the office here, um, is a PCB that was, uh, created as a one-off, and it, and, and it just spells out my name in the traces.
And, um, so you're, you're absolutely right. I've definitely seen this in action. And of course, I've got friends as well who have, uh, had, uh, contract manufacturers.
But of course, um, what I hear from folks who've been involved in this is that there's often a, a bit of back and forth and fine tuning, you know, all that. Um, is, is flux helping with that whole process as well? Yeah, that's a great question.
Um, so yeah, look, ideally there isn't back and forth. Your design is straightforward. They just make it, they say that you have it on your doorstep, right?
Ready to go. Um, but you're right. In some cases, uh, the manufacturer has some questions or feedback.
Um, today, what you can just do, you can just copy. When they email you, you just copy paste that into flux, and the flux will fix that for you, right? Um, and, uh, you know, but, but of course, what we're working towards too is like to integrating that too, right?
Because that feedback is of course, an important step of the learning loop for these agents, right? To learn what mistakes they made or what cost fiction and manufacturing, and to not do that again in the future, right? And I think that's one of these big, uh, compounding effect things, your network effects you have in ai, right?
That humans of course learn too, but it's much more lossy process, right? Whereas like these CI models, when they learn, they tend to not forget that. Um, and so it's very effective.
And so you, you reach, quickly hear like a, a high situation of like, knowledge that prevents them from making mistakes in the future, at least for like that category of problems you work on, Right? So my, my next question, if I can, if I can get in there real quick, is, uh, I just got back from CES, uh, in January. And, um, surprisingly enough, the, the, the big theme at CES wasn't necessarily AI or A IPC or AI devices, it was robotics.
Uh, even though I think we're still very, very early, early in the game, every single vendor that I talked to, even the keynotes from, uh, the big semiconductor companies from Qualcomm, Intel, A MD, Nvidia, uh, arm and xp, Texas Instruments, everybody was about robotics. Whether it's robotic arms or delivery robots. Obviously humanoid robots are always the flashy thing that's shown on stage.
Um, but it feels like there's, we're an inflection point with robotics in general, all form factors. And I, I, I, I look at your company, I'd look at what you're, the service that you're providing, uh, for, for designers and for, you know, small businesses, large businesses, and, and the, the, the power that you're sort of democratizing in terms of designing, uh, uh, even FPGAs. And I, I was wondering if you could talk a little bit about where you think you fits into, um, I think this, this transition into sort of mainstream robotics and how your company, how flux specifically with this, uh, this prompt based model can help accelerate that scale that, and make it more available for, for a lot more companies, whether they're startups or established companies that just wanna build their own robotics, uh, departments.
Yeah, good question. Where to start? Um, I think with robotics, but yeah, you're totally right.
It's an exciting time because, you know, we have, like, for the first time now, these like reasoning engines, right? Uh, uh, and, you know, for robotics, like the hardware we've had for a while, if you look at what Boston Dynamics we're putting out for the last decades, right? It's incredible stuff.
But we haven't had, we didn't have these reasoning engines, right? And so now we have those, and so it's like a new wave of innovation down robotics. And at the same time, by, you've seen it, but all the demos you see, yes, these robots can reason now, but they reason very slowly, right?
It's like a watching, like a turro, you know, either a leave of, of, of, of salad or so, um, and I think that will get figured out, right? I think people are very early and we're probably still a couple of years out from seeing like a truly humanoid, you know, latency kind of capability, uh, uh, robot. But I think it's gonna happen.
Um, I think in the meantime, there's of course a huge opportunity for these like domain specific robots, like think about, trying to think of the company's name, but there's like a company that make like robot vacuums, but with the AI now, right? And it's a huge leap forward towards what we happen as like Roomba generation of vacuums. They're kind of like, they work, but they're kind of dumb, right?
They'll UB cable, you know, every time. Um, and so I think data sector, you're gonna see a lot here now, starting already now. And I think flux is kind of like in the same category, right?
We're like very domain specific. It's an AI model, an AI agent or robot, if you wanna call it that, to make PCB boards, right? And it doesn't make that by being a human who robot that you tell to, and it sits there and sold us this for you all day, right?
It does that by being very special on the design and then being able to utilize the existing manufacturing capacity of the world, of the world economy, right? To then turn these designs into, into real products, right? Um, and I think that this domain, yeah, I mean, the thing is probably co probably probably a couple more decades until you have human robots who could replace that capability.
Human robots who could replace that capability. Um, but yeah. But, but it's in the end, it's, you know, flex is also just a giant robot.
Yeah, yeah, yeah. No, it's, it's, it's cool. Like I, I, I feel like you're, you're the right company at the right time, especially with, uh, you know, how much easier you, you make it.
Uh, just before Steven asks you, uh, his next question, I'm wondering if, uh, how well you scale, let's say that everybody discovers flux, right? Worst problems to have, uh, and everybody wants to start using it and building their own boards and developing their own robots and their own systems, uh, or improving on existing ones. Um, can you scale well, uh, well enough to, to meet that kind of demand?
Or, um, are you currently structured where it might be, it might turn into a first come, first served, uh, camp service? Yeah, great question. Um, no, I think the way we see this is this, right?
This, today there's a, a market of tens of millions of users in the world ready to adopt flux. And that's like mostly people who've like worked in technical fields before where that, again, they have before just used mechanical cat tools, right? And then just bought the electronic someone om or contacted that out.
I think those is, that market is up for grab here now for a company like ours, right? Um, and then, but if you go all the way here, right? Then, look, if I can type in a prompt and I can get any product manufacturer, I can type in any idea and I, and that, and get this back within like 10 days manufactured ready to go.
Um, and we do this for electronics today, but, you know, why will we not also deliver the enclosure for this? But the enclosure for this is pretty simple. We know where the plugs are, you know, where the heat sink has to go.
We can make the enclosure for you. And most enclosures aren't like, you know, a piece of jewelry like your iPhone is. Most enclosures are gray p vvc boxes, right?
Especially in industrial applications. So that's, you know, that's not so difficult to do, but if you can do that right, then why would you still go spend two hours on Amazon looking for the product you were looking for, if you could just describe it and it got made for you. Yeah.
Right. Yeah. And so I think on that level, right?
Yeah. Then there's a market of billions of users, right? Users, uh, and that's, look, that's a long road.
Uh, but that's the road we're on. Yeah. What, yeah, no, that makes a lot of sense.
What about that? So let's take that to the next, uh, the next level. So you're talking about designing PCBs here, but, um, what about everything else?
I mean, it, it's gotta have an enclosure. Um, you know, what about batteries? What about, uh, you're building your giant robot, uh, you've got, uh, motors and wiring and you know, joints and, and whatever else.
You know, what about screens? What about, um, everything else? Is that all part of the flux world as well?
Or is that something that's handled differently? Yeah, that's a great question. So, yeah, no, today we're hyper focus just on everything that goes to make this board, right?
And look, and, and if this display is mounted on this board, that'll work, right? Um, and, and we do this simply because, like I said earlier, right? This is, this is a form factor that was designed to be automated.
It's meant to be fully automated for a hundred years, and then we're finally here now where we can actually fully automate this, right? And so that's kinda like the be chat here, you know, that we've built out for ourselves now, uh, and we're doubling down on that, uh, for the, for the, for the coming months. But then, yeah, right?
I think from here we go to enclosures from, from there. We're gonna go to like, you know, something that has a p of put an enclosure and maybe motors or a display or batteries or bunch of sensors that are not on the board themselves, right? Um, and yeah, from there, we're gonna go to full robots or smart toasters, or, you know, you name your category, right?
Um, because there is like a huge, there's huge infrastructure in the world to build these things if you can deliver them the design files, right? Like the bottleneck isn't in the manufacturing capability. There's thousands of, of, of manufacturers just in China, right?
Who would love more business and who don't care whether they make five units for you one unit or 5 million units, right? They're just looking for more business. Um, and so, and so that's kind of like here, the, the, the need, you know, we're serving, right?
And, uh, in the short term, Yeah. I have a question about software real quick. Um, so, and, and the answer might be very, very short.
Um, but essentially, right now, obviously this is really good, but this is also fairly new. I'm wondering if you like, where the limitations of, of your software, uh, and especially the, the sort of like prompt based, um, you know, design generation, um, that, that you have the AI working in the background, what are some of the limitations that you're working on improving mm-hmm. For 2026, where yeah.
Where you're not as good as you'd like to be yet in, in some way. And if the answer is no, we're great. We can, you know, everything is perfect.
That's, that's a fine answer too. Um, but I expect that you probably have some things that you wanna improve on. I'm just wondering what those are.
Yeah, great question. Uh, no, everything's perfect. I sleep wonderfully at night every night, you know?
Um, no, no, it's a, it's a bucket with a lot of holes. Uh, that's the valid of it, right? Uh, that's, that's the fun part.
Um, yeah, where's the falling short? I think, you know, the big issues here, uh, that we're working on is just recall, right? So, look, it's not good enough if the model gets it right, eight outta 10 times, right?
Especially along of a chain of like thousands of, of decisions that need be made, right? If you, along every step only get it eight outta 10 times, right? And at the end, you got it all wrong, right?
And so that's something we're working on. Uh, um, and then the other thing is like speed, like doing that faster, right? I mean, yes, you could argue, look, it's, it's a miracle that the, the model can do something in an hour that would've taken an expert weeks, right?
Um, but if you really wanna wield this like a guitar, which is like, think about a guitar, right? It's an incredible tool. It's like so intuitive and responsive, right?
And playful, and you can discover, right? And it's like that the feedback loop is really fast. Um, we think that that's kind of like what really amplifies, you know, human innovation and creativity, like fast feedback loops, right?
Try something, get a response, try something new. And I think that's what we're trying to get to here with Flux too, right? To be really like a, a sparring partner and where you can really go quick and forth and play with idea and iterate, right?
Because, you know, we've all like made things, whatever, we've written something or built something, it's all about getting into like an, an, an effective loop here, you know, of like making something, getting feedback, whether that's good or not, changing a little bit, trying it again, right? And that's what we're trying to get to. Um, and then, you know, I mean, with the third side may be much more boring, is to just support much more manufacturing processes, right?
If you think about PCP boards, like if you wanna make, even if you wanna make an actual iPhone, right? An actual iPhone is like a thousand phone components, 12 layers. It's like a very high density design, and you need to support like a bunch of manufacturing processes to make a board like that.
And we don't have them fully covered yet. Or like at the level where you could actually effortlessly make something like an iPhone, right? Um, and so that's what we're working on, Right?
No, those are good answers. That's, That's what keeps me up at night. Yeah.
Yeah. Thanks for being candid about that. 'cause he could have just basically said, no, everything's good.
Um, uh, a a question about that. So in, in a past life, I was a, a product manager. And, and so I, we, I helped develop, uh, new products and, and upgrade old ones.
And, and I remember, I think it was an ideo, uh, tenants, uh, idea of the design firm, right? From, from, uh, the, the nineties, um, that was, uh, fail as fast as you can or something like fail fast, you know, uh, succeed faster. And this seems to me you talked about something that would've taken a designer weeks, just a mere hours with this.
And that's how you're, you're shrinking that, that time envelope. Um, I wasn't thinking necessarily, when I'm listening to you talk about this, about the finished product, I'm thinking about as, as this former product manager, the prototyping process of all these different iterations. Like, I have an idea, I wanna build this, um, I wanna create a working prototype as quickly as I can so I can test it and then see if that's really what I need, or if I wanna add some features, make some changes, and then take what I learned, what works, what doesn't iterate, iterate, iterate, iterate until after two or three, uh, different iterations or maybe 15, I finally arrive at a manufacturer, uh, tested prototype that works and, and that I can fairly quickly turn into a production, uh, a production ready product.
So to me, in, in, in all of this, one thing that we haven't talked about really is, is this process of either failing quickly or learning quickly, uh, and getting there faster through these prototype iterations. Do you think, um, that is still an old school way of doing this? Do you think that your tool helps get to that finished product with fewer iterations?
Or is it just, you know, faster iterations, uh, to get there? Not less necessarily, but just getting there faster? That's a great question.
I think if you look at like AI coding tools, which I would say is probably like the bleeding edge of AI right now, the process has to make, software has drastically changed in the last 12, right? It has changed more than it changed in the last 40 years, over the last 12 months. Like 12 months ago, if you wanted to build a feature, uh, in like, some piece of software, right?
It would probably start with like a conversation, a whiteboard exercise, your designer would for some mockups together, then you maybe do some user testing with that. Eventually you would get somebody to implement that test that, right? It's like a lot of steps in the process.
And if you look at like software teams like us, like work now is, we skip all that, right? Somebody has an idea at 2:00 AM in the morning, right? They typed the idea into one of these coding agents, and the coding agent will just build the thing into our actual product, right?
And on Monday morning, they just show off that working thing in the product that's like ready to be shipped or really close to be ready to be shipped, right? And I think that's kind of what we're gonna see happening here. And too, it's what we're working towards to right.
Select all these in between steps. Yeah. Right?
Yeah. And make it much more responsive and intuitive, right? Again, like going back to this guitar example, right?
Like a guitar I like pulled string and then had to, to wait five days to get like the sound back, you know, that would suck. You know, like a guitar would not be as popular as it's today if that was the case, right? It's so popular and, you know, adopted and, and fun because it's so immediate, right?
And it's immediacy, you know, I think we we're seen in software happening now, right? And we're gonna see that inha too now over the next 12 months. Yeah.
No, this could, this could become addictive, um, for Yeah. For tinkerers, right? Yeah.
Some, right. It's, look, it's fun because you can try so many things right before, right? In this like linear sequential world we were in, I had to have an idea.
I to like kind of spec it out, try it out. I could only do one thing at a time. But that's another thing with DC agents, I can now try 20 ideas at the same time as these ideas come, I just drop the prompt and then an hour later I check back in where, where it end it, and then look maybe went nowhere.
That's fine. No, I don't care. It's just a random idea.
Maybe it's actually amazing, but then double down on that, right? And that's like the power you get here now. And I think in habit it's gonna be amazing because you know how expensive it's to run one process.
Yeah. Like, you can't try 20 different things, right? Because isn't enough time and money in the world right.
To do that. Even like on at large companies with lots of resources, like you think of Apple here or meta making hardware. Sure.
They maybe have like four or five different teams working on different versions of the same product or idea, right? Uh, with ai, they could be working on like hundreds, so thousands of different approaches to the same idea, and then pick the winners or mix and merge the winners, right? Um, and at and at mind boggling speed too.
And so that's kind of like the elasticity and, you know, immediacy we're working on. Yeah. No, this is, this is very Tony Stark in a little bit, right?
And just kind of Oh, For sure. You know? Yeah.
No, this is exactly, I have a screenshot of that in our pitch deck for investors, right? No, no, you're spot on, right? Yeah.
Really, I think this is exactly like the kind of future that's gonna be real within a year. Yeah. No, that's great.
And it's gonna be available to everybody, right? That's another thing gonna be like, like Tony, like in the, in the Marvel universe, only Tony Stark has this, and that's why he's the Ironman and has all these things. But like in the, in the, the, the future of a building, everybody's gonna have access to this, Right?
Which means also that you're, you're sort of flattening a lot of the moats that keep some companies with all this expertise and all this investment, uh, sort of in the lead, you're, you're making it much more accessible to startups or individuals or, or even universities, uh, who want to really innovate, uh, yeah. In, In that space. You know, people always ask me who we're competing with, and, and they think of the legacy tools here, and I tell them like, no, the legacy tools are not where we're competing with.
That's actually a really small market. The existing PCP designs software market, right? Who we are competing with is the OEMs, right?
I have, like, my, my favorite example here is that we have a customer, they make vending machines, like for snacks and beverages, and they used to, before they started using Flux for a single vending machine, it took like four or five individual PCP boards that they had to like, buy from a distributor who bought it from a distributor, who bought it from a distributor, who bought it from an OM somewhere, right? Had to integrate those under into a single machine on the assembly line. Configurators wired up, right?
Then in the field, one of these machines breaks down. You've gotta bring replacements for every single board, figure out which one's broken, replace that, right? Um, it's workable, but it's cumbersome.
Now, fast forward that to them using Flux now, right? They have a single custom board made with Flux that does exactly what they want, no more, no less, right? Um, the board custom pennies on the dollar because they're just paying materials now.
There's no distributors or middlemen that get a cut right on the assembly line. You put one board into the machine, cable in, it's working, the machine stops working in the field, there's one board to replace, right? So the total cost of ownership suddenlys like 100 of what it was before, right?
Plus, you can make it exactly what you want. You're not dependent on somebody else to build the thing you want. You can make it that, right?
And that's the power, that's the opportunity here. Yeah. No, it's, I, I, I couldn't have said it better myself.
So I think this is all the time that we have, so I really appreciate it. Mattias. Um, this has been really interesting.
Uh, and now I'm not just interested in Flux as a, as an analyst. I'm actually interested in Flux as a, as a potential practitioner. Like, I feel like I might actually be able to, uh, uh, to start dipping my, my toe in, uh, in this PCB design, uh, space, um, which I think is the point of this.
Like everything that we just talked about, I mean, throughout this episode, obviously this podcast, but also really in the last five minutes has been, uh, sort of encapsulated this opportunity of this intersection between, um, ai, browser-based AI that can, can basically allow you to describe what you want to build, and as a companion, as an expert, it helps you build it and get from basically idea to manufacturing. Um, and, and I really wanna thank you for your time today and explaining this and also, uh, for your vision in building this, because this is the kind of application of AI in the real world, um, that it, that's, that's a lot more real than a lot of the other, uh, AI applications that, that we've heard about. Uh, and also it's ready now.
It's not something that we have to wait for six months, or two years, or three years. Um, it's available now, uh, right now. So I'll, uh, thank you again.
I'm gonna turn it to Stephen. Um, thanks. Uh, and I, I hope that we can continue this conversation soon.
Yeah. Um, on that note, um, before we go, um, uh, Mathias, um, where can people continue the conversation with you? Where can they learn more about Flux AI and, um, will we be seeing you in the physical world anytime in the, in the coming months?
Yeah, great question. So, look, you can find us at Flux ai. That's a good starting point.
If you wanna start building and start, you know, exploring what this, what this, what this new world looks like. Yeah. Come there.
Sign up. Um, where you can find me. Look, you can find me on Twitter or LinkedIn, you know, just to me, if you wanna get in contact with me, uh, I'm out there.
Uh, in terms of physical world, great timing. You know, we've been remote for the last six years. We're moving into an office, right?
So we have an office in San Francisco in downtown on Second Street. Um, so, you know, we'll, we're aiming here, like as we're moving in to also host a bunch of events here, you know, uh, and turn this into like, kinda like a, a, a maker space if you'll, of the new age, right? Uh, where we together here figure out how to, how to build this.
And, and I think that's the big thing. You look, we extremely early in August. This is day one, right?
And if you're excited here about exploring this, definitely come and try our product. But also, like, if you wanna get involved here, get involved with the community. Look, we're hiring any role you can imagine.
We're hiring for it, right? You also wanna be party and help us build this. There's a lot of stuff to be figured out.
Um, we're really excited about it, and if we think now's the time, so let's do it. Excellent. Yeah.
Well, hopefully we can see you. Uh, we are out in, uh, California for Tech Field Day, fairly often in the Bay Area, so maybe we'll see you at one of our future, uh, tech Field Day events. Um, That'll be awesome.
Olivier, uh, what are you working on these days? Uh, what can we look forward to hearing from you? What am I not working on?
Uh, robots apparently, uh, but any kind of physical ai. So we have reports coming out fairly, uh, uh, fairly regularly. Um, whether it's, uh, a market forecast for devices, PCs, et cetera, or, um, uh, semi-annual IT decision maker surveys that, that look at how the enterprise is looking at, at investing in, uh, in AI and AI devices.
Uh, I attend pretty much all the conferences, or at least all the major ones. So you're bound to see me at anything that has AI and, and devices sort of like converging. Obviously, we just did cs.
I'll probably be at Mobile World Congress in Barcelona. Uh, and then, uh, many more beyond that. Where you can find me is on x, uh, formerly Twitter.
Uh, just look for Olivia Blanchard. You'll find me fairly easily. com website, uh, where I publish semi-daily articles and commentary on tech.
Excellent. Thanks a lot. And, uh, as for me, you'll find me as, as FoST on most social media.
Uh, and of course, uh, as I mentioned, I am, uh, uh, a nerd. And so you'll find me, uh, dabbling with, uh, projects like Home Assistant and ESP Home, which I'm sure, uh, Mathias is familiar with as well. Um, and thank you everyone for listening to the Utilizing AI podcast.
Uh, if you enjoyed this discussion, please do subscribe on YouTube or your favorite podcast application, and consider giving us a rating and a review. This podcast is brought to you by the analysts and experts at the Futurum Group, where insights meet ai. For show notes and more episodes, head over to Techstrong ai, the utilizing AI YouTube channel, or the Techstrong TV app.
Thanks for listening, and we'll see you next Wednesday. We super excited to be here, uh, presenting to this delegation. This is actually our first time, uh, at Tech Field Day.
And, uh, my name is Ish Ker. I'm the, uh, chief marketing Officer here at, uh, fabric Study ai. And, uh, before this, been in the Valley for almost 30 plus years.
Uh, one company in particular was Waca, uh, before this, where I was the head of AI and Strategic Alliances, worked with guys like Fred and others, right? So understand the infrastructure side of things for AI extremely well, and there's a lot of action happening today. But what I'm going to talk about is something on top of the AI infrastructure today, right?
So this is an agent platform, uh, which I and a colleague Rashid will be talking. So we'll have it, uh, kind of divided into two sections. So I'll be giving the company and the product introduction as well as, uh, you know, Rashid will go do a deep dive and then get to the demo.
Okay? So with that, this is kind of the time allocation we have. We have a section for q and a, but feel free to ask us questions as we go along.
If there are pressing questions. Uh, this is, uh, extremely new for all of us. Okay?
So we are all learning as we go along. So who is Fabric? com, right?
For last, uh, eight years or so, from a go to market perspective, a very well recognized in the AI ops space. We were an observability and AI ops player to begin with. And then, um, starting last, uh, beginning of the last year, right, January, we rebranded ourselves to fabric start ai, right?
For a reason. Obviously, the use cases had changed with Agent Tech coming on the horizon. Our customers were asking us to kind of get new functionality, right?
And that was really the genesis to Fabric ai. So again, um, we are headquarter here in the East Bay, uh, with, uh, you know, r and d office in India, satellite offices across the globe. Uh, well, very well known in the analyst community, right?
And obviously, as you guys know, agentic is all new. Nevertheless, we did get featured into two prominent Gartner publications. One, we just actually came out this, this week with Arun.
Okay? So we'll talk about that as we go along. So, fabrics, uh, the founding team, right?
So we have our CEO and the CPO in the back, I Rashid here, but this is kind of our fourth startup, the founding teams, right? So we have had successful exits before, three of them, uh, right? So we've been in this business of actually kind of building companies, uh, and startups, which are kind of morphed into larger business units, servings, tens, and thousands of customers.
Okay? So this is not our first rodeo, okay? So again, uh, for fabrics, uh, uh, in the AI ops days, we had some real large market customers.
There are also now a good amount of large customers, uh, who are actually in production. Some of them are actually doing POC. Okay?
So you can see, uh, fortune 500 customers, uh, right, MSPs and CSPs, communication service providers. So obviously, because of our pedigree, we, we tend to work with, uh, the likes of Cisco and IBMA lot, right? Uh, and that's where the, the partnerships come from.
And then expanding to other verticals beyond IT ops, right? So our classics, uh, kind of, uh, uh, sweet spot, was it operations, knock ops, right? And AI ops, right?
But because of agent ticket kind of enables us to kind of cross the domain across in bops and SecOps and so on and so forth, right? Because we are essentially collapsing at the data level and the application level, okay? And we'll talk about that as we, we go along.
Okay? So let's kind of walk. Hey, why agent take when, when, you know, we are an AIOps player?
So on the left hand side, you'll see this, right? So the classic workflow, uh, in an AI ops environment. So you have an application and infrastructure, uh, stack, which is sending signals, uh, to the AI ops kind of, uh, stack, right?
Or, or, uh, basically the backbone, which is essentially leveraging machine learning and to some extent, uh, you know, deep learning to basically do correlation, root cause analysis, right? But the process for remediation is, is primarily manual, right? So now what you are doing is what I call the last mile problem is you are basically exposing dashboards, you're exposing incidents, and it's up to the domain experts to interpret that and kind of formulate and remediation path, right?
So heavy kind of, uh, intervention with, with, uh, you know, uh, with humans in there, right? So from there, as, uh, you know, the customers started asking, and, and again, this, what drove us actually, uh, to the agent ops was actually a customer environment, right? So we had a telco, uh, which I suffered, you know, an eight hour of outage just because somebody inadvertently did ACL changes, right?
Access control list changes. And it, it took them almost eight hours to find the problem to remediate the problem, right? And these are perfect use cases, which can be remedi remediated with agent take ops, agent ops, right?
So what's the new stack? Like, we still have that application stack we interface with, we have signals coming in right now, the LLM is at the center of reasoning or the decision making, right? And we'll talk about the, the good and the bad part of that as well, right?
Because it doesn't come free, right? I mean, LLMs, at the end of the day, they are non-deterministic processes, right? But what they enable us to do is now you can have workflows which are semi-autonomous or autonomous, right?
Or why call semi-autonomous is you typically customers still, they build trust. They will have human in the loop, right? HITL, what we call, uh, the, the, the LMS themselves can do deep research, right?
So you don't have to depend on a domain expert saying that, Hey, have you gotten my networking logs or have you gotten my metrics from the application? Right? The LLMs are good in doing what we call chain of thoughts or graph of thoughts, right?
Uh, kind of reasoning, okay? Uh, and then, uh, it's cross domain. It can talk to a multiple, and especially it's true with us.
Uh, this is one of the major prerequisite for an a successful agent platform that you have to be able to connect to, you know, a number of different endpoints, right? Because in IT ops particularly, you have storage your network, you have applications, you have a PM tools, what have you, right? And then, uh, the inplace data access is also something which is very unique to us, and we'll kind of talk about it, uh, uh, as we go along.
But each of these tools, what we interface with are domain specific tools, right? And they don't want to integrate with each other, but by agent take and what we have in terms of what we call dynamic tooling, we are able to get to those endpoints and data sources, right? Directly, right?
And we'll talk about how we do that. Okay? So, uh, that's kind of the setting the stage on how we are transitioning from AIOps to agent TKI, right?
Um, now having kind of shown the promise of agent take, right? There are also challenges with agent because the primary kind of LLM or what you're depending on, right? The reasoning, uh, LLM is actually, uh, a non-deterministic LLM, right?
And, uh, it, it's what we call stochastic processes, right? But in addition to that, there are other elements also which contribute to the, to the hallucination, right? So particularly there are three main kind of, uh, you know, reasons what, what we kind of, um, articulate as the reasons for hallucination.
So first and foremost, context and data management, right? So context is the king, as I say, right? So the output tokens, which an LLM pro kind of puts out is directly proportional to the veracity of the input tokens, right?
That means simply put that context is how much curated context you are providing to the LLM really determines your accuracy, right? So that's the most important thing. And you'll see our key IP being that middleware, which comprises of context engine and what we call dynamic tooling, right?
Also data management at scale, right? So we are not talking about, you know, an experimental demo. It's very easy.
It's one thing to do a demo, right? Where you have five devices, it's another thing when you are in a real life production environment, you are getting 10,000 rows of SQL, right? You're getting large logs, you're getting metrics, right?
You're interfacing to almost hundreds of devices, right? Uh, how do you manage all of that, right? So that's where the data management part comes in.
How, I mean, for, for MCP, as you guys know, right? The model context protocol is very transformative, right? You don't have now EFLs then in your code anymore, right?
You're essentially handing over a manual to the LLM, right? Hey, these are my tool sets, and out of this, select what you feel appropriate, right? But at the same time, it also comes at a cost because every conversation has to go to an LLM and back and forth, right?
So how we have simplified that, uh, is also, uh, our key ip lack of connectivity. As I said, we connect to almost 1700 plus data sources, right? And even beyond that.
So that's very crucial when it comes to IT ops and AIOps, uh, use cases, right? And finally, agent tops is where, you know, hey, having a platform is one thing. Operationalizing is, is another thing, right?
And what I mean by operationalizing is observability security, right? Explainability, right? Trust governance.
How do you bring all of this, uh, where the customers you're building in the trust with the customer, okay? Uh, and that's what really, you know, all these studies, what you're seeing about MIT or Cisco talking about, hey, there's a lot of noise, but there's only 5% of the value, right? And I call that as agentic value gap, right?
So you're doing a lot of experimentation and demos, but it's really, you know, very little of enterprise use, uh, coming out of this agentic use cases. Making sense so far, folks? Any questions, comments?
I'm waiting for you to get to it 'cause I have a ton of questions, but I'm waiting for you to kind of lay it out a little bit more, Okay? Sure. Yeah.
So, we'll, we'll get to the, to the gist of it, right? So again, if you see the progression with agent take applications, right? So just like with the three tier application, we introduce a middleware layer, right?
Uh, for complex application scenarios. Similarly with agent take, what we are seeing is most of the applications today are the, the, the two tier one where the AI agents and LLMs are directly talking to MCP endpoints, right? So MCP has to be unable for the endpoints because they, that is where you get to a structure with the LLM can understand.
Now, what happens when there is data sources, which are not MCP enabled, right? Which is pri primarily telemetry. You are, you are reading raw data or you are reading from legacy devices, right?
So that's a pitfall. That's another reason why, you know, what leads to hallucination on the right hand side is our solution. So what we have introduced is what we call this fabrics.
Middleware, right? And there are two main components to this middleware. So one is what we call the context engine.
And the second one is what we call universal tooling, right? And the context engine in simpl, in simplistic term, what it does is it provides curated data to the LLM only gives it what it needs. It's not dumping everything, right?
Because that's where the, the LLMs you, you are overriding the context. You are corrupting the context and the LMS hallucinate, right? And the second part, the universal tooling is taking care of how do you dynamically connect to this disparate data sources, right?
You get that data schema and you normalize in a way such that the LLM understands, right? So tho that's the key ip, which we'll be talking about. And what does that do?
It allows us to talk to obviously any MCP data source, right? If you have an MCP data source, that's great, right? And we'll provide you still all that context, um, capabilities and so on.
But also any API, right? So any API, uh, based data source we have, we will actually create that, uh, MCP wrapper on top of that through the middleware, right? So, which is huge as well as raw data and legacy systems, right?
So just to give you an example of this, right? What I mean by this, right? So say my intent, right?
Uh, uh, I am providing an intent to the platform saying that, Hey, I want to get metrics from this application servers. Now, it so happens that those application servers are monitored by, uh, Dynatrace as an a PM tool right now, uh, my LLM and my, my system goes and checks that, hey, uh, there is no tools, uh, MCP tools available for, uh, getting the metrics from Dynatrace, right? At that point, what we'll do is we'll launch and tool handler a tool creator, which will actually scrap the internet.
It'll look for publicly available Dynatrace APIs, right? And it'll formulate the tools, four or five tools, right? Which, uh, essentially are, uh, are there to talk to the di Dynatrace controller, get that metrics right?
And we have a demo, um, demonstrating that, right? So this is what I mean by dynamically creating, um, you know, um, the, the ability to go and talk to data sources, which are not MCP enabled. Making sense?
So you guys sit at the data la the data acquisition layer of an AI process, right? About the data acquisition, the middleware sits on the above the data acquisition. There is a data fabric which integrates with the data sources, Okay?
And what y'all do is you take, uh, lemme see if I get this right. So you, you somehow provide the data or clean the data, get it ready for the MCP, but you also work with data sources that aren't from, that aren't involved with MCP to get them prepared to go into the data set. That'll be used to train a model.
Yes, absolutely. Right? So in short, yeah.
And we will also, uh, integrate with that. That's what we call the dynamic MCP, Right? And you're correlating the data from all those sources so that I understand, you know, what the information in one system means to, to the other, and how it's giving you the full picture.
Yeah. That's, that's further down, right? So now we are sending it all to the LLM.
Mm-hmm. Right? And which is actually correlating and doing the root cause analysis and so on, right?
And we provided a template, which we'll talk about, right? We'll give it some examples that, hey, this is how you should go about doing this, right? Isn't it Ray Esei, Silverton Consulting, isn't there a challenge here with normalization of the data coming from different data sources and trying to understand the range and those sorts of characteristics of the data you're getting?
Yeah, absolutely. And that's where our, our value prop comes in. So the middleware actually takes care of that normalization, right?
It normalizes in a way where the LLM understands our structure. Because if you give LLM unstructured, which doesn't have structure, is bound to ate. So that's, that's really the value prop.
You're, I'm also interested in, in how your context engine manages the context. 'cause context, uh, growth is a major challenge in most of these things. Yeah, absolutely.
So we'll be going actually really deep in, in, um, rashid's section on that context engine. But just to clarify, right? So this, I mean, again, we come from that, uh, world, right?
So what we're talking about is not the size of the context, right? There is a lot of, uh, uh, innovation happening with KV cash and all of that, right? We are talking about the purity of the context, right?
So it's not about, obviously size matters as well, right? But now a lot of these LLMs are coming out with a lot of bigger, uh, context memory. But really it's, uh, uh, it essentially boils down to how pure your context is, right?
Because that's how the, the output is going to be determined. Yeah. How pure, Yeah.
Purity, well, not pure storage, but, um, alright, so this is, so what are we talking about, right? So at a system level, this is our platform. It interfaces with number of other capabilities, right?
It could be application devices, it is, uh, uh, IT and business tools. We work with a number of different data platforms, right? And we have, uh, da, uh, data, um, platforms in our bundled with the product.
Like we have, uh, uh, graph db, we have, uh, you know, and, um, uh, OO open search and so on. But we also can work with customers external tools, right? That's another value proposition.
What we bring, we work with a number of different LLM models, right? And pretty much all the latest ones, right? Like GPD five, DO two is out.
Now we are already testing it, right? It is been out only three weeks or so, right? Uh, cloud, right?
We support, uh, uh, others as well. Okay? Then we interface with output tools, which are essentially ITSM tools, right?
They could be automation tools and so on and so forth, okay? And on the, the, the user interface to the platform is, uh, through two ways, right? So there is a studio, our own studio, not the Microsoft, so what we call it a copilot, right?
Which is mostly conversational queries, right? So you can ask conversational queries, you can test your hypothesis and then convert that conversation into an agent right there. The other user interface is what we call an agent studio, right?
So it allows you to build your own agent. So the platform, uh, out of the box, we provide about 50 agents or so, right? Across different categories.
So we have AIOps observability, SecOps, bops, what have you, okay? And then the real power of the platform comes for partners and end users to build their own, own, um, own agents for bespoke use cases, right? And that's where now we are kind of, you know, broadening the use cases beyond IT ops into SecOps or bops and so on and so forth.
Yep. So before you go on, this is, uh, Jack Poller with Paradigm Technica. Yeah.
Um, on your, your architecture diagram there, you've got a couple of benefits listed. Faster deployment, MTDR reduction, alert noise reduction. Can you explain a little bit about what sort of, what are we talking about there?
In what context? Yeah, so, uh, in terms of, uh, the consolidation of the tool, right? So now you have, you know, individual domain specific tools, right?
Right. Like you have, uh, Dynatrace, right? You have, uh, item tools or a PM tools or what have you, right?
Or you have infrastructure tools. So we have the ability to now go directly, uh, go and talk to the devices ourselves, right? Okay.
So in some scenarios, you may, you may, over a period of time, you may see that, hey, uh, this platform can actually pro, uh, uh, cater to those use cases. So, so those use cases would be monitoring your AI infrastructure, Uh, well, yes. Yeah.
It's not just AI infrastructure, it's actually the customer deployment, right? It's application stack. The infrastructure stack stack, Okay.
Yeah. Okay. Yeah.
And so then the MTTR, yeah. So the next one is alert noise reduction, right? So that's the standard use case in the AIOps world, right?
Okay. You're getting a lot of alert fatigue and you're doing, correlating that, understanding the context, right? Then, uh, the other one is faster deployment.
So the platform is, is what we call it a data driven platform, right? So what that means is there's very little, it's a low-code platform. So we go and we, we, uh, basically do data discovery, right?
Uh, we understand the schema, right? And then we create, uh, you know, our declarative templates, right? So that's another value proposition of the platform where you can, it's composable, right?
So that's faster deployment. And MTTR reduction is literally, uh, you know, one use case, which we'll talk about, right? Uh, where, uh, in a standard way it was, is it was taking a long time for them to get to remediation, okay?
Right? That was a problem statement. Making sense so far?
Yeah. Josh here with another follow-up question to that. Um, how much data ingest are you doing, right?
You're talking about interfacing with the data is a source and building a template around that, but are you also building a data lake underneath, like, or dependent on the platform, because right, all of my data sources are not gonna be equal, right? If I, I'm, if I'm assuming that the customer doesn't already have a data lake and they've gone through that bronze, silver, gold process, right? To, to purify things, yeah.
You're saying I can still interface directly with another product like ServiceNow or, or whatever that's up here like Splunk and not worry about a data lake? Or are you building a data lake? No.
So you would have a data platform. You would have a data platform, but a lot of these things, actually, the correlation and so on, it happens in the context engine. Mm-hmm.
Right? So that's where, but we do have persistent layer, right? You could have Splunk, you could have o open search, uh, you could have min, right, what have you, and then ServiceNow and all those are the output layers, right?
Mm-hmm. Where you are basically updating a ticket, uh, or so on. Okay, that makes sense.
Yeah. Yeah. I gotcha.
Yeah. Yeah. So, uh, real quick, right?
Everything I'll have to quickly go through, um, yeah. So, uh, again, uh, capability wise, right? So what we say how we are different is basically it's a full stack platform, right?
That means there are a lot of agent platforms who only focus on the agent layer, right? So we take care of all the way from the data to the agent layer, to the automation layer, right? So that's what we mean by full, full stack, uh, purpose built, right?
Because a lot of, uh, you know, there's a lot of hype around these frameworks like crew and auto gen and so on. In the IT world, you're dealing with real time data, right? You're dealing with alerts, incidents, time series data.
Many of those frameworks don't even support real time data. Like CR AI doesn't support, right? So this is a platform which is build grounds up for this particular real time use case, okay?
And it's enterprise grade, right? So this is how, uh, Russia, when, when he walks, uh, over in the demo, you'll see this is our catalog. This is the first point of contact for an end customer, right?
Where you have this, uh, you know, 40 or 50 AI agents where you, you are basically, uh, you know, clicking them and demonstrating them, okay? So I think this, and then a couple more slides will actually, uh, establish, right? So this is, we are now kind of getting, uh, under the hood on what is our unique value proposition, right?
So the middleware is really where, uh, you know, we differentiate against all the other, uh, uh, other platforms, right? Which are like first generation, if you will, right? And it comprises of the context engine and the universal tooling.
And again, as I said, the LLMs, they don't have state, right? So that's why the context becomes important here. We maintain the state, right?
And now when you have all this, you know, large number of, uh, data sets, uh, sql, a rows coming to you, or, um, large data, um, large logs you coming to you, we keep, uh, it, we basically summarize it and keep a state, uh, in the context engine, right? So now, once you have that state, when you are having the MCP servers, uh, say server one talking to server two, it doesn't have to go all the way to the LLM for every transaction, right? That state enables it to kind of converse it between the two, two different MCP servers.
Anyway. So the three main features of the, of the platform is the agent tops model, right? So this is where we actually are operationalize, right?
The use cases where you are essentially able to bring in, uh, you know, uh, those, what we call scaffolding, uh, o of the, uh, of the elements, right? Where you trust governance and, uh, observatory and so on. Then there is a middleware, which is where, you know, we are doing shared context intelligent caching.
And the tooling, tooling is where we are able to dynamically go and talk to data sources, uh, using, uh, either MCP or raw, raw, all this, right? And then finally, this is all built on top of the data fabric, which allows you to actually connect to this 1700 plus, 1800 plus data sources, okay? Uh, yeah, I, I'm not going to walk you through this, right?
But essentially the capabilities we have, like say for example, how do you build trust with a stochastic model, right? So this is where, uh, you know, we have this, what we call prompt templates, right? Which are, are basically give some examples to the LLMs that, hey, this is how you'll go about root cause analysis, right?
This is how you would go about looking at the, uh, CBE, right? And then what we do is we use the LLM to choose those instructions, right? That's what we call dynamic instructions.
So that's kind of a trust part. Everything flows through with a persona, right? So we have guard rails around that, that's very important.
So, and level three engineer will see only certain tools and data sets, which are different than an infa, uh, IT manager, right? Who will see a different set of data sets and d different tools, right? Same thing with governance.
Uh, there is finops model. So you can leverage look at, hey, what's your cost based on a project, right? So you can create a sandbox and assign certain personas to it, right?
Uh, security, right? So, uh, we have this concept of list agency, right? So every agent should only get access to a data sets and tool set, which he's entitled for, okay?
Uh, reliability and performance observability, right? So again, this observability is very different than the infrastructure observability we are all used to, right? So we are talking about observability at the agent layer, right?
Mm-hmm. So, which is important, right? And what we do is, uh, not just provide you, uh, all the audit locked trails, but also we have, uh, flows, flow maps, right?
Mm-hmm. Which is in real time, that means, say here is a user prompt, which went to the MCP client, it went to the MCP server. This is what, what was the tool set which was access.
Mm-hmm. So once you're looking at it, it's very little room for hallucination. Yeah.
Uh, this is how we actually show this in the demo itself, right? So we operationalize it, like, uh, looking at the persona, looking at the tools and so on. Uh, this is actually a positioning slide.
Some of you might find this useful. So there are other players as well, right? And we acknowledge that.
So there are observative players, the traditional one, right? There are AI players like right, are reason and so on. Then there are orchestration frameworks.
There are SRE vendors, and that there is platform place. So Gartner has actually placed us both in the, in the platform, uh, uh, this as well as in the AI SRE partner, right? And that arrow indicates that some of these traditional observative players who lack, uh, agent, um, uh, framework on top of that, we are able to compliment that, right?
And they have been our partners for a long time, right? Guys like, you know, uh, Splunk and, and, uh, so on and so forth, okay? Mm-hmm.
So any questions on this? On the positioning? No.
Yeah. Uh, again, so this is some examples of how we work. Like we have Splunk agents, uh, we have I om agents, right?
We have CRM agents, okay? Uh, so this is the agent to clear complementing the existing observability and I om tool. Yeah.
And so I, I actually do have a question. Yeah. Um, could you kind of give us a, an abstracted description of how one of your customers is actually using this?
Yeah, yeah. So, exactly right. So, uh, I'll skip this.
Let me get to this example, right? So this is an example of en large e energy company, right? So, uh, you have an application and infrastructure stack, okay?
Uh, uh, you have a number of different tools, right? There is observatory tools, there is data tools, there are, uh, it om tools and so on and so forth, okay? We are basically bringing this in all okay?
Through, uh, our, our, that dynamic MCP tooling, right? And this is where, uh, you know, the, the engine basically is able to curate all of that telemetry to provide it to the LLM, okay? Okay.
For root cause analysis or SecOps and so on. So Is basically an MCP run in for MQQT, Uh, well, MQQT could be in the IOT scenario, right? We don't have your MQQT, right?
These are more enterprise applications, okay? Yeah. Yeah.
But yeah, I mean, if there would've been a, that would be another end point, right? So, so great Luc Silverton Consulting. My, my question is, you know, a lot of these solutions servers, VMs, networking storage, GPUs, they're moving more and more to MCP server solutions themselves.
So they're supplying a lot of the MCP endpoints for what your universal tooling is doing. I, I understand from a legacy perspective, a lot of legacy solutions out there may never get to there, but, but, uh, from more current infrastructure perspective, they're all moving to that, in that direction. Yeah.
The other question I have is, and, and so the LLMs and the agent tech workload flows are also moving towards more context, sophistication, management, compression, those sorts of things. It's, it's, it's, it's, I guess my question really is do you have a, a, a, a solution that, that will, you think will last in this environment? Well, Yeah.
I mean, again, uh, that's a, there's a good, good point, right? So again, there are, most of the enterprise solutions do have MCP, right? But there is also a large number of, uh, you know, a legacy or raw data particularly, right?
It is the same thing with open telemetry, right? We talk about open telemetry. There's 80% of the data sources still don't support it open telemetry, right?
So that's the problem. Certain, coming to the second part of it, as I said, that the caching, what the infrastructure caching we are talking about is, is different. That's good.
Good announcement. But it's really the purity of the, of the context what matters here, right? Which is what we are Focusing.
It's the second time you mentioned that. Yeah. Yeah, exactly.
I mean, you'll hear that all the time. By the way, I'm going to, So I'm Scott Shaley, director of leadership narrative with solid. 5 architectures.
And then I'm gonna hand it off to Phil and he's gonna fi wrap up this section and go on to the, the last section of it. So you can only have me for a few more minutes, be back soon. So what does one point 21 gigawatts solve?
Come on, somebody's gonna smile about that. Oh my word, I'm not that old alrightyy. That's shocking.
It can send Marty back to the future. 21 gigawatts. It's also what it takes to power San Francisco for a day.
It can also deliver power to 550,000 Grace Blackwell, GPUs, GB 300 platforms, and it enables 25 exabytes of storage in a one one gigawatt environment. Now, how do I know this? What did we do to be able to tell you that one point 21 1 gigawatt can do all this?
Again, looking at the ecosystem, looking at our friends, doing the research, these were all announced in 2025. These are all the platforms that are gonna go live sometime after the announcement in 2025, Stargate Meta Core Weave, XAI, all these guys. And we did some math, and that's how we got to the one gigawatt for 550,000 GB 300 platforms.
And if you look at the math required for how many GPUs you have and how much storage you need to go along with those, you get your 25 exabytes of storage. And so the next question, of course, is really how do you get there? Well, we did that math for you.
Um, again, 550,000 GPUs direct attached. These are the E one s performance drives going right next to that, that GPU sitting in the server. That's the nice, uh, rack based design, currently has eight drives in it because they're air cooled or potentially, uh, direct to chip, liquid cold plate cooled.
And that nets out to eight and a half exabytes of storage next to the GPU. 5 exabytes of supportable storage using 1 22 terabyte drives to be able to get up to 550,000 GPUs. So this math is, is like all good TCO models, it's a tit for tat.
So whatever you put in, you get out. So we focused on, if I use just my solid state drives at the highest capacities for both knobs, it enables us to get to that many GPUs. You have any other product that consumes different amount of power, or you put sixties in instead of one 20 twos, you're doubling the power footprint.
You reduce the number of GPUs available in that gigawatt. So the gigawatt was our bar here. I'll be 100% honest.
There's math that shows if you ignore the gigawatt and just look at the GPUs, the amount of exabytes of storage will be just mind blowing that they're expecting to use. And this is Grace Blackwell. This is 2025 data.
We're now one whole month, literally last day of the month into, um, 2026. And the whole ecosystem has changed because when I showed you guys this graph last time at field day two, it had a quote from our good friend Michael Dell. This is our friend Jensen at CES.
This is the market that never existed in the market that will likely be the largest storage market in the world. Gotta love the fact that we finally have Jensen talking about storage and not just memory. I love it.
Now the next trick is that GTC to have mentioned our name, we'll see if we can get there, right? He signed our drive, of course, you know, that whole thing. 5 IIC MSP layer in the VE rubbin platform that's tied to the Bluefield four implementations.
Now, the next slide, I'm gonna tell you how that works from a SSD hardware point of view. And then a little bit in just a few minutes, Phil gets the lovely chance to show you how that actually looks from a system level implementation point of view alongside what we're talking about. So what I'm doing here is we have the KV cache exceeds the HBM spills over to dram.
We still have a limited DRAM footprint only. So many dims, only so much capacity falls onto the local storage, those nice directly attached products, and then it falls over again. And every hop is a connection and a distance and a time.
And so by bringing it closer and closer and closer, that's what this whole architecture is about, is access to data faster in a more confined environment. And so, if I take what I had before on the one gigawatt example, and I throw in the Vera Ruben platform, or the Ruben platform as it's being called, they haven't officially given it the GB nomenclature. We now have three banks of stor of storage products.
Now, again, I'm constrained to one gigawatt, so I'm still at 25 exabytes, but this is now how it splits out. 4 exabytes of this new context memory storage, which can still be the high capacity drive. It's just closer to attached to that blue field forward architecture.
1 exabytes of direct attached storage. 'cause they're changing the amount of local to move it just a little bit further out past the blue field. So the number of drives in that initial server is actually coming down as the capacities are going up.
And so the interesting thing here is we're with the one gigawatt. We still have 25 exabytes of storage. We split it out three different ways now instead of two.
But note the number of GPUs that are supportable, we're down to 440 or down to 400,000. We lost 150,000 GPUs. But the performance of the system doesn't change at one gigawatt.
And the reason for that is because you're using NVME storage for the direct detach and for the contact structure, you have to have the fast storage products in those two layers to overcome the gigawatt problem in this environment with, uh, this new added layer. Because fewer faster GP PS need more access to fast data. Therefore, you get ICMS with, uh, solid state drive.
You can't put traditional rotating hardware in that layer. You just can't. So if I, when we come back, uh, at our next field day, AI field day, uh, we're planning to, we're gonna have even more details on this, and we're gonna blow off the one gigawatt and just show you the capabilities of what the storage looks like.
And we're talking a five x or larger CAGR year on year from 26 to 30 on just the demand for high capacity storage over what we were already talking about. It went from where it was at about a 20% cagr. It's now 30 40% CAGR because of this introduction of this.
And that's for the NVIDIA only based systems. So the storage platform is now the shining star on making the success of the next layer of AI as we get into the inference context scenarios. So we're gonna do a little bit of context switching here.
Um, we wanna talk about efficiencies too. And so I just gave you the hardware centric ICMS one gigawatt constrained environment view efficiencies. When we partnered with Vast, we came out with this amazing TCL model that talked about replacing your CEF hard drive infrastructure with R one 20 twos and Vast in an efficiency play, talking about how to make your systems more effective.
And when we were at sc, we put up a bunch of slides. This was an SC carryover for you guys. You wanna talk about it from here.
Um, the Computer History Museum from the museum to the 1 0 1 is the SSD implementation on equivalent of this graph. The hard drive implementation is going from the Computer History museum all the way up to Oracle headquarters, where, where they were up in the NICE four. I know it in Oracle headquarters, but that's how, that's how far our distance is that you do when you put a drive end to end to end and how much reduction in overall ecosystem environment you can drive.
So what we're gonna do now is I'm gonna hand it over to Phil. He's gonna help explain a little bit about this and then jump back into the context memory and give you some more fun, uh, topics about, uh, the wonderful vast platform. So I-C-M-S-P is inference context In inference context, memory storage platform, that's what they called it.
And It's behind the bluefin. So it's effectively a, a storage solution out there that's doing something for the context management. I'll dig into it.
It Looks like great lead into the next little section. And it's different than the rest of the NAS object data lake that's behind it. It it re-architect it For you.
Okay. Yeah. Thanks.
Yeah, Yeah. We'll dig into it. Okay.
Thanks for having me guys. So Philman na, I'm the go to market execution lead at Vast. I've been there, uh, I think it's three or four days.
We'll hit my six years at Vast. So I've kind of got to see the company, uh, grow. We'll do a quick introduction, but I really needed to play off Scott's, uh, metaphor here, which is one, really happy to be here with solid I and our partners, but vast, we build software, right?
We can't run well obviously without any hardware. So peanut butter great, you know, but it doesn't work so well. It's not very portable or, uh, reasonable to eat if I'm gonna spread it on my hands.
So really the jelly and the bread to our peanut butter is solid on. I couldn't go without that piece. Um, for those who, you know, not as familiar with Vast, we actually launched really the company to the public here at Storage Field Day back in 2019.
And when I was interviewing, that's really how I learned about the company and whether I wanted to work here, right? And I saw some obviously compelling things, um, since then, right? We've really become a significant portion of the storage market as we look to this year.
We're gonna drive a very significant portion of all enterprise SSD utilization, uh, with storage, expecting dozens of raw exabytes. That's before our data reduction, which we're gonna talk about that efficiency and also doesn't count any of the data going into the cloud. We've made some really big announcements on cloud partnerships this year and extending the platform, which was primarily on-prem into the hyperscaler space.
Sales have really gone well for us. We're roughly tripling year over year. Our quarter's gonna finish tomorrow.
So pay attention as we start to, to announce some of those new things. We have our, uh, customer event first ever user conference for Vast at the end of next month. We'll talk about that.
And I would say just like tech Field days evolved from Storage Field Day to AI Field Day, vast has really a vault, uh, from being a storage company to building many more things on the platform, which we'll talk about, really allowing you to capture data, contextualize it, and then act on it with ai. So I would say the founding principle of VAST is that really a few things. We were very bullish in 2016.
AI was gonna change the world. We were very confident that AI was gonna change the way we computed on data, right? I think both of those things proved out to be correct.
And then the third is that the architectures that got us to where we are, were not the architectures that were gonna take us forward, right? And this is really the main culprit, the shared nothing architecture really invented by Google in 2003 in a white paper basically defined the internet kind of application cloud era where you've got these node based architectures, right? I've got a node, got some CPU in memory, I've got some kind of storage in there.
Originally it was disc. Now we swapped it out for flash in a lot of circumstances. But the only way to get to the data on that node is through that node's controller, right?
And I personally storage guy, like I look at this as every scale out na, every scale out object platform. But ultimately it's also the architecture for every data lake, for every distributed data warehouse, right? It's all over the place.
Eventing infrastructure, it is everywhere. And it really does create a lot of scale problems in the AI world. When we look at it really from a storage view, there's some challenges around flexibility, right?
I've gotta create nodes or pools of homogenous node types. Uh, not designed with flash in mind, right? We'll talk about some of the challenges around things like data reduction.
And then finally one of the big things that shows up everywhere is just this east west traffic. There's so much communication between these nodes that even though I can scale my resources linearly, I'm not scaling performance linearly. And a lot of times these architectures work well small and these problems show up more and more and more as the clusters grow.
So we're looking at it now from the, the really the TCO and efficiency perspective, right? We know we're in a supply crunch, right? Customers have been trying to move steadily from, uh, spinning disk space architectures to SSDs.
When we look at the AI deployments that we see in practice, you don't see any spinning disc, right? You got power challenges. I was wondering why you actually had the round things with the floating heads on them for showing disc describes here Because our marketing team likes that image, I guess.
But these are all right. The, the world of, uh, architecture based on spinning, right? It's got the arm.
It's more like a record player. Oh, uh, an older reference who? Old school.
Okay, well I'll, I'll, I'll give the marketing team the feedback. What's the record Player? I love it.
Okay, so in the AI world, right? I think as you look at what's in practice, solid state's required, right? I think now we're seeing a rise of companies coming up around saying, Hey, tierings cool again because we have an SSD supply crunch.
But if it was not a good idea before a supply crunch, I don't see how it's a good idea after a supply crunch. So what we need to do is help customers be a lot more efficient with the way they use SSDs, right? One of the big problems with this shared nothing architecture is how data reduction works, right?
And when you look at a lot of these architectures, a lot of them have given up on things like deduplication. You have compression only, right? And a lot of the data in the unstructured world, it's already compressed.
So that kind of takes away a lot of opportunity to drive efficiency. 2 to one is kind of what you're gonna get if you're using compression only. So what is the challenge with ddu?
It's around having a global view of the data, right? In this world, I basically have to chunk up my DDU index because the other option would be to put the entire index on one node. Everyone would just hammer it and that would not work very well, right?
So we said, okay, we're gonna shard this up. Essentially create a distributed database. And in that world, every node has a piece of the index, right?
So I get a limited view there As I start to scale this, all these nodes are talking to each other, looking at who's got the data that I already might have as I grow it adds to the east-west traffic, adds to the performance limitations. And then ultimately the kind of bandaid there is to create limited DDU domains, right? So I'm basically only de-duping within a pool or within a few different nodes within whatever architecture that you're building around.
But it's always very local in this world. So vast, again, looking at the architectures, brought a new architecture to market that we call date very quickly. We call it disaggregated, shared everything.
Because essentially we kind of broke the idea of a node apart. And we have two independent scaling layers. We have our logic compute layer, we call those C nodes.
It's essentially container running on an X 86 server. And then we've got where all of the state of the system lives down in these enclosures filled with very dense, solid, IM 122 terabyte drives or whatever the right uh, drive is for the customer. Now, some different things.
And unlike the shared nothing world, in the day's world, every one of those containers has direct access and actually sees every one of the devices in the system as a local device connected over NVME over fabric, right? Architecture impossible without NVME over fabric, which now makes it allow that I can have remote drives, feel local from both how they're mounted and performance. We also have a layer of storage class memory in the system where all of the systems metadata lives.
So that means I can create a shared global index that every single one of these containers sees. And that means I can do global data reduction at an exabyte scale without any of those different challenges, right? So fundamentally unique architecture that allows us to look at the data in a very different way from a data reduction perspective.
Any questions? High level, the architecture, how it works, okay, that'll be a theme that we hit on, right? So step one, can we give an architectural, I would say advantage to how we look at global data reduction, step one.
But again, the problem is we're talking about unstructured data here, right? Not as friendly of deduplication as things like VDI and virtual machines, right? Not as friendly of compression, maybe as a database that hasn't been compressed already.
So we had to look at some different things, right? And really move beyond duping compression alone. So very high level, you look at compression, right?
I'm looking for commonality, repeating data at a very granular level, right? That's gonna be a small chunk, I don't know, eight to 60 4K usually could be anywhere in between, uh, deduplication. Right?
Now I can have a global view, assuming my architecture allows it, but I'm looking for more course matches, right? Two chunks of data exactly the same. I find that a lot.
VDI, virtual machines, I'm copying databases, whatever. Uh, but again, I don't always find identical matches in unstructured data. If I chunk up and try to do de-dupe on a big pool of unstructured data, what you actually end up finding is a lot of chunks of data that are mostly the same, not exactly the same.
DDU misses that every single time. 'cause that would be a hash collision that's corrupting your data. It's terrible.
So what we do is introduce a new type of data reduction, again, enabled because we have this giant metadata structure living in storage, class memory, the architecture that we will identify if two chunks of data are mostly the same, compress them together and essentially store the differences, kind of like a snapshot, right? And ultimately, we don't just use similarity, we use all three of these, right? So we're looking for the best opportunity compression.
We actually use a couple different types of compression. We will look at the data, take a sample, what's the best type of compression, and use that deduplication. We have, again, global deduplication.
We have something we call adaptive chunking. Chunking, which means we'll actually change the dedup window to find the best opportunity for deduplication. And then similarity is kind of that icing on the top where we're gonna find that next level of similarity and drive out even more savings.
Right? And ultimately, you're in a world where we could easily get two or three times more data reduction than the next biggest competitor because of what's happening here. Do you, uh, I wanna know if you're a believer or not.
Uh, yes, absolutely. Uh, I, I'm chuckling because you're, you're giving the exact description of what I would've been describing with solid fires architecture 10 years ago. Got it.
Okay. I knew about it. Right.
But I would say solid fire in the shared nothing world, right? A little bit. Absolutely.
I, but I mean, the, the, the things you're describing are, it's like, yeah, this is exactly what we were doing 10 years ago. No, it makes so no, that, that's, I'm sorry. That's why I was chuckling.
No, No, it makes sense. And by the way, I think it's interesting that, you know, in the block world, the hard drive died like immediately, right? You know, I was part of the extreme IO team at EMC.
We had pure, we had solid fire. Everyone. The, the hard drive in the, like the block world, virtual machines, databases, VDI died immediately.
That was like 12 years ago. And there's still so much of the world's unstructured data on spinning disc because they haven't been able to figure this calculus out, right? So it's actually a great point.
Um, some actual data, right? So if you were gonna say, I don't believe you, I was like, look, I have data, um, average data reduction by the way this is pulled this month. Because as we've looked at the supply crane crunch, we're like, let's start digging into like where we've come and what the results are.
4 to one, right? Again, these are not VDIs, these are, this is unstructured data. Some of our customers have hundreds of petabytes of highly compressed video.
Some of it's encrypted, um, some massive estates. The weighted average. 87 to one.
Exactly. 9, looks prettier on the slide. Um, and then we have 27% of our customers get better than three to one.
We have some customers getting eight to one. We have some customers getting like 30 to one depending on the data type. So where typically, again, in the world of unstructured data, you'd say, if I get anything at all 10%, I'd be happy.
We're talking about getting you three times more data for your flash. And that gets combined with something that I'm not gonna nerd out on today because of this time. But our erasure coating is also incredibly efficient.
So our erasure coating at scale under 3% overhead, it's actually 146 plus four stripe that we use enabled by our architecture. So, um, when you look at that compared to, you know, traditional kind of shared nothing where you're gonna have maybe 27, 20%, we have a lot of customers moving to Vast that are still using like das Data Lake technology and they've got their data triplicated, you'd be shocked about how much of the world's capacity is still triplicated. And it's because it's in these monster data lakes, um, where again, they're getting, you know, for every 10 petabytes they can store three petabytes of data and those systems don't have any data reduction, right?
That's all over some of these large data analytics environments. So you combine these things, a lot of times our customers, even if they don't get good data reduction, they're getting four times more effective capacity per, you know, petabyte that they buy. And even if they're buying something that is more kind of enterprise, then maybe it's more like double the capacity that you can store for every raw petabyte that you're gonna buy.
So we actually just launched this, uh, something called Vast Amplify. So in the SSD Crunch vast over the years has really gotten a lot more flexible. Again, uh, when we started, we had to run out of every specific hardware build.
Now we're working with pretty much every major OEM vendor running on more, uh, traditional servers. We're in the cloud. So we actually have a program where we're going to customers and taking their SSDs that they already have in their systems and repurposing them into vast systems to amplify the capacity.
We actually had a cus couple customers come to us and said, Hey, we've got SSDs. Your technology is way better than what we're using. Can we reformat these and use them?
And we have in, in some very large scale environments. I'm talking at this point, we've repurposed hundreds of petabytes of data, thousands and thousands of drives. So just a question.
This is all really great statistics and y'all are doing really awesome, but, um, we're talking about ai. So I would love if you could tie this back to ai. Can you, does it matter if I have a data lake that's not deduped?
Maybe I want that and I just want the, I just want the data tagged in a different way so I can find it for different reasons. But like what, how does this tie back to ai? Yeah.
So I would say how it ties back to AI is right now what we've seen in practice, any large scale training environment, any large scale inference environment that's actually in production at scale is a hundred percent based on SSD. Right? Okay.
That's, that is standout. We have in a world where customers are gonna struggle to get as much SSD as they need, right? So what I need to be able to do right now is make more use of my solid state devices because AI is driving tremendous demand, right?
And we'll get into more how it fits in the architecture, but the point is, looking at bringing spinning disc into this world we really think is a terrible idea. If having a tear miss is going to kill my GP utilization, destroy jobs, destroy performance, then I can't use that as a lever. I need to figure out how to make most use of my flash in this AI world, right?
As I wanna deploy agents and inference over a much broader set of data that data's hitting on spinning disc, it's not gonna work. Well, you're Not gonna get an argument about that here. Yeah.
So, yeah. Right. It so go ahead.
I was just gonna say, but when it comes to draining training data specifically is, uh, as de duplicable, if that's a word, as traditional data sets have been in, in guys' experience so far? Yeah. So I would say in training data, um, two to three to one, okay.
Is common, right? If you look at some of the bigger neo clouds that are our customers, two to one's pretty typical on training data sets and, and more, and you're doing This all inline, Right? So it's, I would say it's kind of the best of both worlds between inline and the old world inline men in memory with vast, our inline memory is storage class memory, right?
So what happens is the data lands in storage class memory, it's acknowledged up to a host, and then we data, we data reduce it when it migrates down to QLC. Mm-hmm. Okay.
So yeah, kind of outta a band. I'll show you what that looks like actually. Yeah.
But I mean that's, that's not terribly unusual way to do it where you, you actually need the hashing at some later point to, to be able, you duplicate it, but you end up using much less storage later on. It did. Yep.
What I would say the difference is we don't land it on the capacity tier, right? So we don't land it on QLC and go back and mess with it again. It's not good for where it's not good for performance.
What we do is we leave it in storage class memory where it's very fast access. It gives you a lot of opportunity to move. And then we don't need to plan to have non de-duped and non-used data on the capacity tier.
And I promise, by the way, the most of the rest of the presentation will be specifically on ai, but we wanted to bring in the TCO of making SSDs affordable and and hacking the supply chain crisis. Was that okay? Thank you for saying that.
'cause that was not coming through. Okay. Sorry about that.
I Appreciate that. When you're, uh, repurposing SSDs, are you having to, to migrate the data off the SSDs and migrate back on from a vast perspective? Or are you as simulating We do need to simulate The data directly.
I mean, We do need to move it. Yeah. So if we take a file system, we can't convert it to vast and data in place.
So a lot of our customers we're either working with swing space or we're creating clusters and failing nodes out and growing into it. You're seeing a lot of usage of the, uh, the new capability to, uh, reuse SSDs. Yes.
So again, we customers brought us the idea originally to say, Hey, we've got SSDs, we wanna repurpose it. So that was how it got rolling. And since then, yeah, customers are all over us to say, we know we're looking at the year, we've got more demand, we're looking at rolling more ai.
We've lived in a solid state world and we, we know that we can't have capacity that's 30% utilized, right? We're giving us one third of what we're buying to store. Okay.
More AI stuff, right? So now I promise the rest specifically on ai, right? But again, we think flash is that kind of first step is to enabling your data on fast access.
So we're just gonna talk about the context challenge and KV cache and why, right? And again, you guys probably have been paying attention to what's going on with Nvidia, but for people maybe, you know, more infrastructure folks, essentially the thing is here, right? I ask a question to whatever large language model, the first thing that it does is trying to figure out what do I actually care about, right?
There's different words in a statement. The what, the, the, uh, what, what is this guy actually asking about versus some of these words that don't make sense? So I calculate that, turn it into key, uh, key value stores.
And that's essentially the context of the conversation. Something else that adds context is maybe a document or a video, right? Someone says, Hey, I wanna just summarize, you know, solid, I'm in vast tech field day, I'm gonna upload the video into my favorite large language model.
But if I ask a question, again, it used to be I had to calculate all of that context again, right? So a question, maybe not the end of the world, but if it's a document or a video, then I'm calculating that a lot, right? Think about some big enterprise organization dumps a new document out to the world and all of their employees are asking questions about it.
I'm recalculating that same context on that document over and over and over again, right? And then obviously I need to make sure the decode phase is the answer part. That is where I'm creating an answer that makes sure it's related to the question that you asked.
So there's some big problems with this context piece, which is one, if I'm recalculating over and over and over again, I'm burning GPU cycles on something that's not adding a ton of value, right? And honestly, GPUs are too expensive for that, right? We had, I think NVIDIA's customers are like, we can't just keep dumping all this CapEx and scaling forever.
You gotta help us use these things more efficiently. That's step one. Problem two, user experience.
If I'm a user and every time I ask a question about a document, it's going to do a bunch of work that's so annoying. I wanna engage in a conversation with you, not have you forget what we're talking about every time I ask a new question, right? I think, you know, some of these large language models, they know everything about you because they have all of that data.
So that's KV cash. I wanna store that context so I can continue, continue to reuse it, right? And Nvidia has this hierarchy, which Scott talked about.
So step one, stored in high bandwidth memory, right? Obviously there's some challenges there. It's really expensive and hard to come by right now.
Uh, the other piece is it's local. So if I'm engaging in a conversation just on, you know, this one session, that's fine. But if my friend is trying to have the same conversation, you know, do I have access to that memory?
Then I can move it down to dram, right? Then I can move it to local SSD. Again, everything's local.
And then the next step is shared file and object, right? When you're gonna have a massive drop off in performance there. East west traffic, all those different things we talked about, we shared nothing.
So we were like, we need something right here, right? 5. Something that has a performance closer to local but is more global in terms of its access, right?
And that is essentially what I-C-M-S-P is, right? How do I create that local feel? Now again, I'm not gonna spend a ton of time on this, but one way to do that is to basically take a shared nothing architecture, right?
I can either put, you know, a client on the blue field or I can deploy my software, right? On essentially the CPUs in these g um, GPU servers. Now again, problems there is I'm bringing the problems of that shared nothing architecture up into my most expensive assets, right?
I have EastWest trap happening, right? I might have a hotspot, right? Where everyone's asking about the same piece of context.
That means I have all my GPU servers attacking one essentially and asking it for information. I don't know that that's a good idea, right? So what we said is we've got, um, a different architecture, right?
We walk through this shared, shared, uh, everything architecture where I've got this stateless layer. Now, I would say, you know, some potential challenges with this instance of the deployment, really two, right? And again, I think, I don't wanna say challenges, but optim areas for optimization one, right?
I have this layer of CPUs that is essentially kind of in between my access to SSDs, right? Um, and as we know, right? The CPU is always gonna be the bottleneck to SSD performance, right?
You think about the world's most powerful processors. How many do I need from a thread perspective to saturate one single 122 terabyte drive? It's a lot.
So we have that problem. The other problem is, you know, I have to essentially create a copy of data, right? I've got an RDMA operation to our front end, and then I've got another RDMA operation to the SSDs, right?
So you kind of have this hop that's happening. What we're able to do with I-C-M-S-P on Vast is actually take our logic, our C node, and move that up to run on the blue field. Now, we actually, uh, introduced a prototype of this style architecture, uh, I think with AI Field Day, um, earlier, but that was with the previous generation of Bluefield, right?
Blue Fields have gotten dramatically more powerful from a core count. So now I can run my c no, the logic of the system up in those blue fields, this is a paradigm shift, right? I no longer have a host going through other CPUs to basically get in line to get access to data that's on fast media.
Now, every node has its own little friend. That's its protocol server. There's A storage class memory here, Phil.
It's still in the dbox down there. You just can't, it's like, yeah, there'd be two layers of storage in that dbox a storage class in the old way, The storage class memory was also in the dbox. Yeah.
Nothing changes there. It's just how I give access directly from the host to that device. Okay?
Yep. So we don't have to change anything there, right? Which again, now, instead of having to need to use the local SSDs and introduce potentially, you know, conflicts and all the different things that we might have by putting software on those servers, I can just have JBoss, right?
Full of flash with dent solid. Im SSDs and everyone has direct access directly to the metadata structure and directly to the actual data itself. And again, scaling and everyone sees everything.
So there's no problem in sharing context, right? If there's a hot piece of context, everyone can access it with a, a whole bunch of parallelism. Uh, but I'm not gonna have any hotspots up top.
So Phil, excuse me, Phil, Jack Poller with Paradigm Technica. It sounds like what you're really doing here is you are running storage controller software on the GPU because you've got spare GPU cycle. It's on the blue field.
So I have a GPU server, I'm putting a blue field, which is like a smart nick now it's got 40 cores in it. We're taking those cores, which you don't need for network performance 'cause it's just more cores than you'd ever would. And we're running our storage software there.
So it's not in the CPUs, it's not on the GPUs, it's on this little server essentially that's mini, that's let's smart Nick, but a lot more powerful than that. Okay? And the net effect of this is, So ultimately what you get, right?
So we talk about certain things, right? One much faster time to first token, right? So if someone's asking a question, I now already have that context, I'm sharing it globally, right?
So if anyone's asked about anything, So in, in, in a traditional architecture, then you are making a request from the GPU to a storage controller that's off host. Yep. Right?
And then that storage controller goes fetch as the data feeds it back. Correct? And so in this case, what you're doing is you're moving that storage controller on host correct.
Or a little bit closer to the GPU. Correct? And that's getting you, that's accelerating significantly.
Significantly, okay. Yeah. And I think there's a few things, right, that come into play.
So you've got the acceleration of taking out an RDMA operation in the middle, right? Uh, you have a, a scaling advantage of the fact that now every time I add a new host, I'm adding compute specifically with that host. That is its own storage resources, right?
Essentially it's dedicated. So I'm taking out all the potential conflict, right? You get resources for you.
You don't have to fight over them with your partner. When I have that shared CPU pool, we're all fighting for the same resources, right? So it's a scaling, it's a parallelism and it's efficiency perspective.
The fact that the data is now shared means that I can have more GPU servers able to share more context. They're much less likely to calculate things again, right? So that's why I get a faster time to first token because I can pull that context without having to recreate it.
I get much better GPU efficiency because my GPUs are not recalculating the same things over again. They're actually doing inference instead. Right?
And then the final piece is I'm taking out that entire compute layer and that all is power that's drawn and powers precious now. So by taking out that entire server CPU group, I cut power by 75%. Got it?
Make sense? Mm-hmm. Okay.
So can can, can you tell us again what I-C-M-S-P was? It's context management something something. I think it's inference.
Context management storage platform. Okay. Inference context.
Memory storage platform is what they called it. And if you Google it, just be careful. There's a whole bunch of other uses of the acronym ICMS and I-C-M-S-P.
Just think of it as really cool close storage. Okay. And I keep thinking about the image that you put up that had the different layers and had context as one of those layers.
So, okay, so is this vast I-C-M-S-P, that's what's being attached to the blue fields. So essentially it's I-C-M-S-P is, um, think about it like NVIDIA announced something called Dynamo, right? And we were working closely with them on that.
You've got like all these different problems in terms of, you know, how do I manage where inference jobs run on GPUs, right? How do I make sure that I am, uh, intelligently using the different tiers of me of you know, memory and storage just for context, right? So that's just for context.
Um, and then how do I scale and run those things? And basically the I-C-M-S-P is a tier of storage and essentially a standard way that Dynamo's gonna interact with that storage. So I'm basically saying we are lining up to saying this is how NVIDIA expects to use extended, you know, off, um, or shared storage for context.
So Can you go back to your diagram that shows? Yeah. Okay.
So where is it on this chart? So Essentially this is gonna be used for context in this world, right? The notes?
Yeah. So all the da, all the, the, sorry I'm now, I'm not supposed to point to the screen. All of the context gets stored down in that DBOX layer in the same way we would store any type of data.
Okay. And so that is y'all's I-C-M-S-P is gonna be in the dbox. Correct.
Okay. Thank you. And the data lake that's also underneath that is stored there as well.
You easily can. Yep. So, you know, something that I would talk about, 'cause I, I, let me go to the next slide and maybe it will, it will help.
So one of the things that's going on right now around I-C-M-S-P and KV cache is the question, should you use any data services? Right? Because could they impact performance?
Right? And we went through that whole shared nothing thing. And certainly if you do use data services on a shared nothing architecture, you're gonna have challenges.
Data reduction is one of them. Another one is something like encryption, right? So right now we're not sure should we encrypt that data.
I think it's a really bad idea to not encrypt that data. That's a giant shared thing, which has everything about every conversation that all your employees are having with ai. You might want to encrypt that, right?
And you kind of have to save it again. But there's also a chance for data reduction from what we've tested. 3 to one and two to one data reduction, which means, again, And again, this is an inference solution, not a training solution.
Correct. It's all inference at scale. Um, so the point is, you know, in our world, this is how we would do it, right?
But what's unique and flexible about the vast world is those blue field controllers don't have to be the only CPUs that that cluster has access to. We can create actually a sidecar pool of compute that just does data services. Because remember, those blue fields are gonna write data down in let's say storage class memory.
We can then have this pool of compute, take that data, reduce it, store it back down to the QLC, right? But I can also attach other workloads over here, right? And what I know is that these blue fields all get their own dedicated amount of storage performance.
Every blue field has 40 cores sitting. So in The other prior solution, the blue fields were actually responsible for the D Yeah. We without The other compute side Correct.
Cluster. Correct. Which they, you know, at this point, this is a very new solution.
We're gonna have to do a bunch of tests. Will they run hot? Will we want to augment?
Will it be enough? Uh, but the point is, we're the only ones that have this flexibility to say we're gonna add compute that you can then leverage for data services. So essentially I just care about the fact that these guys can access data and then let our friends over here take care of all the data services In the old world that these sort would've had to have been homogenous nodes.
But this environment you're taking, you could put any, any cluster of compute services out there to be your C nodes, correct? Mm-hmm. At this point, it's very flexible.
We have customers with multiple different generations of compute running c nodes in the same giant environment. Mm-hmm. We can actually take, we can pool them, right?
So we can say basically, you know, certain applications can use two C nodes and the rest can use 20. There's a lot of flexibility in how we can carve this up, that disaggregated shared everything piece gives us flexibility in a way that wasn't possible before. And you mentioned storage class memory down at the D nodes.
Those are, uh, different types of SSDs that are tailored to, you know, read, write access and things like that, Right? So essentially it's an SSD, you know, lower latency, they're more expensive, uh, much better endurance profile. Mm-hmm.
Right? So essentially, you know, when we came to market in, uh, Intel, I'm losing word pcm, you know what it is? PPM Optum.
Yeah. Optane. See it died such a long time ago.
It was all we had. Soy has a, uh, P 58, 10 SLC based SSD that is used as a storage class memory solution for these guys. Mm-hmm.
So there you go. So yeah, it's all about endurance profile, cost latency. Marian, I have a quick question.
Sure. Uh, going back to similarity, how do you keep the similarity decisions as the data and the models change? Yeah, so that's really just the underlying data structure, right?
So it's totally abstracted from models or anything else, right? So as data comes in, essentially we're hashing it. And normally with ddu we use a really strong hash, right?
That, 'cause you wanna make sure that I never accidentally mistake two pieces of data for the same. That's why DDU uses a strong hash. All we're doing is taking a chunk of data and using a weaker hash.
And that weaker hash basically says that we don't use it to actually store the data. We use it to say, Hey, this data is very similar to data we already have. So it doesn't matter if that data came from an inference job, a backup job, whatever.
We will look across any data that's been stored on the system, um, doesn't matter what protocol it landed in. And we will identify that there's commonality and we just won't store it. And the system has no idea that is happening.
And you're gonna use a weaker hash for, um, Comparison. Why? Comparison?
Comparison? Yeah. Because if you use a strong hash, you can't tell, because when I use a strong hash, a small difference in the data, a results in a very different hash, right?
That's why you do that. So a weak hash basically says if the data's slightly different, then I'm gonna get the same result. So now we know we're in the zone, right?
That these two pieces of data are very similar. And what are you hearing from your customers who are in highly regulated industries? They have no issues with it.
Right? Um, again, data's now typically on a, on a system, it's chunked up, it's erasure coded, it's spread around anyway. So at this point, you know, it's all generally pointer based.
Uh, I haven't never heard anyone have issues that they would have to turn it off because of some kind of, Okay, sure. Thank you. As far as the KV cache and actually doing any special caching of the data coming off the blue field versus it's all going to storage class memory when it's written and it'll be re gets to QLC or whatever the backend is, uh, it's not like you're holding that data in storage class memory or anything like that?
We Are not. So, you know, in general, KV cache, we'll use the different types of media available, right? So it could land in the memory, the high bandwidth memory, right?
Or it could land in, you know, a local SSD, uh, this is essentially another tier of KV cache. Um, and for us, yeah, we're, we're gonna keep all that metadata in storage, glass memory, but we, you know, there's, there's nothing different about how we have to store the data. Essentially, we've just created a more optimal data path.