The Power of a True Enterprise AI Orchestration with Kamiwaza
Kamiwaza is a groundbreaking platform designed to enable enterprise AI at scale. Its Distributed Inference Mesh and Locality-Aware Data Engine deliver unmatched performance across diverse environments, including cloud, on-premises, and edge locations, while remaining independent of specific silicon technologies. A key focus is inference optimization for the operational aspect of AI, addressing the challenges of increasing inference loads associated with complex applications like retrieval-augmented generation (RAG) and autonomous agents. The platform tackles the “data gravity” problem, which significantly hinders enterprise AI adoption, by processing data locally wherever it resides to minimize data transfer and maintain data sovereignty.
The Kamiwaza platform distinguishes itself through four key differentiators. First, it provides a complete, opinionated yet loosely coupled enterprise AI stack delivered within Docker containers, allowing for rapid deployment and easy customization. Components like vector databases can be swapped with minimal code changes. Second, the global inference mesh and locality-aware data engine enable distributed inference across multiple locations, intelligently routing requests based on data location and available resources. This approach drastically reduces data movement while maintaining performance and compliance requirements. A global catalog service tracks metadata across all locations, facilitating efficient data access and processing.
Finally, Kamiwaza’s architecture is completely silicon-agnostic, enabling deployment on diverse hardware in various environments. It offers a unified API across all locations, fostering seamless integration with existing enterprise security and authentication systems. This approach, combined with its ability to integrate third-party and in-house Agentic applications, positions Kamiwaza as a “Docker for Enterprise AI,” simplifying the deployment and management of large-scale generative AI solutions across complex, distributed enterprise infrastructures.
Presented by Luke Norris, CEO, Kamiwaza. Recorded live in San Jose, California on January 30, 2025 as part of AI Field Day 6. Watch the entire presentation at https://techfieldday.com/appearance/kamiwaza-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://Kamiwaza.ai for more information.
Transcript
So, um, Kaza, I'm not trying to win the name, uh, the coolest name award for a company. Um, it actually has sort of a deep meaning, uh, in, uh, Japanese. It like a lot of, uh, uh, awesome Japanese words, it's hard to translate into English.
It can mean anything from divine skill God-like, uh, and then in business context, it's, it's typically super human. Uh, so that's obviously what we're sort of grabbing onto. So, uh, not only is, uh, it a cool name, it was in a great book called The ous Deception.
Uh, they talk about sort of, uh, living and being, uh, Kaza. Um, and it just also happens to be the name of one of my favorite games growing up. Kaza, art of the Thief.
I love throwing that one in. So, uh, Matt and I, uh, co-founded it. Matt's on Zoom here.
He's gonna be, uh, showing some of the live demos. Um, yep, there he is. Um, and, and, uh, we sort of have a history of, uh, long history together.
Um, this is our second company. It was CTO at my previous company. Uh, so we've really hit the ground running and we're pretty excited, uh, what we can show off here today.
So, um, I I, I watched yesterday, there was a lot of talk from several of the companies, uh, about training and inferencing. Uh, the focus today and the focus of our technology is really gonna be on inferencing or the operations of generative ai. Uh, so I always like to set the stage a little bit simple chatbots.
It's typically about two to three inferences per request where it actually has a scan back and forth over the model to actually get a response. This is your typical chat GPT one shot. You're just sort of throwing something into a chat interface, then you escalate pretty quickly to chat with rag retrieval, augmented generation, it immediately sort of jumps up from an inference, from a load perspective pretty darn quick.
So you go from about two to three to now five to 30 inferences to actually cover the retrieval of the data, the re-ranking of the data, the ingestion of the data, the processing of it, and then actually pull that back in. Now, the reason I'm building this up is because what we're gonna focus on today is agents. That's really sort of the buzzword, but it's actually what sort of our customers and typical outcomes are being built around.
So an agent takes this to a whole nother stigma. Now, for a simple agent request, uh, that you sort of tell it to go autonomously accomplish something, the inference load goes through the roof. Now it's not just sort of two or three or five or 10, you're talking a minimum typically of a hundred inferences.
Real agents that run real autonomous and are doing real data processing can get into the 10,000 plus inferences just to accomplish a simple sort of request that you would give an agent to go accomplish. Add onto the fact that each one of those inference requests might have a data request associated through the rag process. You might be looking anywhere between a thousand to 10,000 to a hundred thousand data pools actually pulling that data, having the model ingested.
So a simple ag agentic request can kick off massive amounts of compute and masses, and I mean masses amounts of data, ingest, pulling, copying, reading, et cetera. And this is why we come up with the concept of data has gravity. I love this quote from, uh, my good friend Michael here.
Uh, AI inference and data gravity is gonna drive sort of the idea of repatriation the load that it's gonna require to run these ENT workloads needs because of the speed of latency, the actual amount of data that they're ingesting be processed next to that data. This is actually one of the biggest inhibitors in enterprise's adoption of gen ai. Think about it.
If their data spread across cloud, prem edge, multiple clouds, multiple databases in the clouds, multiple data services like Snowflake and Databricks, and a simple agentic service to get an answer has to touch multiple data pools in multiple locations, it's nearly impossible for the enterprise to actually stitch all that together or to move all of that data into one location where the inference engine can actually attach to it to actually run that many pools. And through our research that we ran for about a year and a half, we saw this as the number one sort of impediment for the enterprise adoption in general, generative ai, especially in agen resources, let alone the security, let alone all of the access controls to all that data when it's spread across all those locations, let alone the pipes, you know, just not having enough bandwidth to actually service 'em and to connect. So how do you sort of a, a attack, what has been a general problem in all of IT infrastructure for the last, you know, 30 plus years?
Well, we went about it in what we believe are four unique ways in the market. And these four unique ways in the market is why the Fortune 500 and large government organizations have just been flocking to us. And literally we're having a hard time keeping up with the ingest of net new customers.
So let's go through sort of our quote unquote four major differentiators. Now, once again, we're only talking inference. We're only talking the ability to process and run models, not train models.
First, we deliver a full stack for the enterprise in a docker container. We're not gonna talk Kubernetes today in a docker container. Those set of docker containers is the entire opinionated stack for the enterprise.
On the next slide, we'll do a deep dive on all of the constructs of that. I'll pass it to Mike. Uh, my, uh, Matt.
So to explain the stack in a little bit more detail, we call this though our opinionated but loosely coupled stack. And this is also a game changer in the market. So within 15 minutes, the docker containers can be installed, but on any hardware with a large enough CPU or GPUs to actually run inferencing, it's loosely coupled because the com WAA middleware allows you to swap out all major components of the stack with about one line of code.
So if the enterprise doesn't like the vis vector database that comes with the opinionated stack and they wanna use enterprise DBE vector db, it's one line of code to swap that out. Of course, that's true for, huh? There's a lot more to Installing certain vector databases.
Like MongoDB would take a whole lot more work than say, VIS versus anyway. No, if they already have that one line of code, if no, if they've already had it, if they've already started down their journey and they already have a vector they Already implemented. So instead of using our opinionated one, it's one line of code in their middleware to swap out the other one.
Okay. Uh, Matt, do you actually just want to cover that on how we, I Mean, there's a lot of credentials if they use different. Nope.
Matt, do you want just swap on how we abstract the vector database services? Uh, the short answer is we abstract as much as we can. Yeah, there you go.
Sometimes if we're, if we're, if we're talking about things that have like super diversion things, there is more commonality for some than others, and that's true. Okay. Um, and of course some things like Mongo or PG sql, you know, pg v support might contain metadata that you'd want to use, uh, that might not be universal, right?
That's, so we're normalizing things like collection names and, uh, the interface for, uh, inserting and retrieving factors, but you're right. Uh, but we try to make it relatively easy to kinda hook in under the scenes. And like a good example is that our default vis client, you know, when you're instantiating the economy was the one, the native vis client is just not client in our object, right?
So we're not trying to like obfuscate things, we're trying to just kind of help you along. I'm just saying, oh, one saying you can do something. One line is a dangerous in front of our technicians.
I understand. Yep. Um, second, the loosely coupled architecture also allows us to separate sort of a lot of the dependencies.
Um, and the ability to just update those individual dependencies also allows it to be sort of a much lighter load across it all. Once again, fully containerized. Second, and I'm gonna co combine these two together.
It's what we call our inference mesh, our global inference mesh, and our locality aware data engine. There are two key concepts, and I can take you through the paint by numbers here. So first off, you see the Comisa stack installed in cloud.
This cloud is a private cloud. For this particular instance, you need access to the underlying GPUs. And this isn't a SaaS service.
Think about, you know, Amazon, Microsoft, Google's private side, uh, of the cloud. That cloud could be swapped out for any location. It could be any data center.
Cloud's just a representation here. So you have the full stack of containers installed on A GPU based cluster in the cloud. Then you install the docker containers on a GPU based server in the data center.
You ride the enterprise networking using docker swarm networking. So they connect right over the enterprise network, effectively VXLAN encrypted so that the two are now aware of each other. Now, once the two are aware of each other, you get the concept of the inference mesh, the ability for you to move inference load based off of the underlying hardware that it's running on the models that the individual stacks are running on.
And then you add the third component, which is the data component. So as those stacks through their initial rag or continuous rag processing where they're actually building the underlying embeddings and putting those embeddings into the vector database, we add layers of metadata and put those into our own global catalog service. The global catalog service is then shared between the two clusters.
So the individual clusters know not only what data they can actually process in the rag, but they have the affinity of the location tied to the actual docker container that did the processing. Now, let's go through the paint by numbers on why that's so important. As a request to our stack happens substantiated from an LLM enabled app or user service, it hits our stack.
The LLM on our stack says I need to rag data to resolve this particular request. But in this case, the data for the RAG requires data from both the cloud and data that resides in the data center. So it kicks off at through the paint by numbers here.
Now number three, and is running the rag process in the cloud with the data that resides in the cloud while simultaneously kicking off an inference request. So the data stack that's running in the data center, the stack in the data center realizes then it needs to run the rag process with the data in the data center. And then it only sends the inferred result tokens back to the initiating stack in the cloud there, it concatenates the result and gives a single response back to the user or the application that actually hit it.
In doing this, we've effectively split the data gravity paradigm because we've actually processed all the data local and only sent back the actual result tokens between the two. These tokens could have been sent over a dial up line, a 28 K modem, literally could have moved them. We're talking very, very finite little amounts of data.
And in processing this, we've gone through because that rag process isn't as simple as just a process. As I pointed out with the agent service, it could be 5,000, 10,000, it could just constantly be processing, reprocessing, re-ranking, and rerunning that data in that location. But once it gets a result, it actually sends it back that way.
Once again, we haven't had to move the data beyond the firewalls of the enterprise. We haven't actually decided to move too much data. Our customers are typically in the petabyte exabyte scale arena.
And this allows the inferencing now take it to the next level. Multiple clouds, multiple edge locations, multiple data centers across the world with GDPR and hipaa. This is a way to process and run generative AI solutions through all of those major paradigms.
Uh, you think splitting up the, the inferencing GPU side of things versus the rag thing is going to add more, uh, effectively or or less data transfer? I'm, I'm trying to understand why you did this. I think mostly data sovereignty.
Yeah. Yeah. I mean, Yes.
I mean, but the tokens are still being passed into the cloud, right? And, and with the tokens you could potentially reconstruct, uh, some portion of the data with appropriate LLM. I mean, No, we're talking processing million genomic sequences locally and then only sending back the results saying five genomes.
We're actually having this particular marker would be the tokens that actually pass back nothing in the actual genome sequence. I'm just giving you a, a concept of an idea, right? Right.
You can further have ossific rules where you don't passify data, you don't pass, you know, seven digit numbers between those Two stacks. There's multiple ways to work around that. Two, the data is a lot of those, a lot of the functions run LLMs feel to me like MapReduce operations.
And so, you know, Luke and I in past lives have had customers that had as much as 750 petabytes of data in their, call it data lake, although it's number one data lake that size. Right? Um, and so we were definitely thinking about how you deal with that and the concept of, you know, 5 22 global data centers.
And I wanna say, you know, find me every customer service record, you know, from a customer service call log where XY, Z happens. And in every event you're essentially gonna pull every single customer call transcript. You're gonna have many, many inference passes against that, right?
'cause you can't just dump 'em all into a model and go, tell me this thing. Right? You need to do some level of like entity extraction or organization on the transcript.
And then each of those results in multiple calls now. Yeah. But don't gravity starts to help there.
But then also of course, when data sovereignty comes into play and you're saying, well, the things from Germany we have to keep in Germany, that's the second big reason. But, but if there, it's all coming from the rag, right? Like how, I mean then do you have to have sovereignty per rag or do you have to have bifurcation for rags for different vector databases and then you get it to detect vector database sprawl?
It's, thank you. Yeah. It's, um, well of course at very large scale you're using many vector databases anyways.
Um, 'cause there's limit to how many vectors you're gonna store. I think 750 petabytes of data. I don't think anybody's gonna vectorize all of that.
No, no, there won't. Yeah. You have massive, massive amounts of vectors.
Nonetheless. Um, in this case, uh, what we're doing at our layer to be a little bit more like in the weeds about, you know, how it works is, um, we have a catalog layer that comes with kawa. We'll see it in some of the demos.
Um, when, when people do global clusters, we can interlink the catalog. 'cause we're only really tracking metadata, right? We don't keep data itself.
Um, and even if you put data, data in the vector database that's local to whatever vector database you're targeting, but with the metadata and our awareness of location, we can essentially say, well, the person wants to tech use all the catalogs that match this pattern, and therefore we can essentially remote across the interlink, uh, the inter clusters communication coming on to let's, we're gonna run this here and there, um, to try to draw an analogy. It looks a bit like a, a a dag execution, right? You're running, um, you're distributing functions basically to do this in every location.
Each location retrieves the data that's associated with its local version of the catalog, you know, where to send those requests because the metadata about the catalog is globalized, but you're sending it to each location, doing the rag in each location, getting some level of local inferencing. Although you can move data, by the way, it's not a hard and fast rule, right? But the main idea of the feature is you're gonna do it locally, but then you can coalesce all the results back.
Um, and I think of it a little bit like, to do another analogy. I think everybody here is probably familiar with like Starburst for, for example, right? I, Presto and Trino and Starburst are all essentially like distributed hive SQL type query engines.
Um, the trick about Starburst is they can actually be multi-site and like distribute, um, the query in many places and do like predicate push down to kind of optimize some of that, right? So we're doing something very similar with the, the catalog. So Matt just, just hi, hi by the way.
But, um, hey, the, um, so what you're saying is sort of a data gravity for rags for sure. Okay. All right.
That, that makes more sense RA rag per site because of data gravity. Yeah. Yeah.
Got it. Yeah. Got it.
No. Okay. All right.
Makes sense. Uh, but to challenge to some extent, agent you, you mentioned this earlier, is the agent workloads are doing literally thousands of interface, uh, inferences, which is Each inference can have hundreds. Yeah.
Each of those is potentially another rag operation. I mean, it's a lot of interaction between these two clusters that are going on in order to support this, uh, age agent workload. Is it They, they, there's, they're Bifurcated.
Think, think about, um, site A goes out to site B and C and, and site A is doing it itself. And it says to each one, Hey, for catalog items like this that you have locally, go retrieve them and then pass them into this inference process. Right?
The inference process is that over Ray using distributed remote functions, right? So it's a little bit like executing a remote lamb, uh, that we just make a little bit easier, uh, in each site. It is locally responsible, and this has all happened simultaneously for retrieving the data in that site processing with the language models deployed in that site.
And then you can re all the kind of final answers, if you will, at the bottom of that map, reduce process back to the middle. Does that make more sense? The, The true key is the processing's local Between the data source and the L LM in general.
Yeah. Still not clear to me that you're, you're, you're, you're, you're buying much by segregating the, the rag pros from the inferencing itself, but The, the, the inferencing is local for the rag and is running. It's not separated.
It's running there. It's You're passing the rag. We're not passing the rag information, Not passing the rag.
Only the result of all of the inference and all of the rag processing. Only the, Oh, that was obvious. I thought we were actually No, no, no.
Passing the tokenized rag. No, I'm trying to like beat this India only that result. That's why it's so important.
You're not moving that data. You're not exposing it. It's only that result after all of that is run locally.
And, and I think in practice of course, like this tends to be less about performance and data gravity just 'cause of how heavy weight L one processes are in general, much more about sovereignty. And then you can actually implement things like guardrails. So it's very common practice now for production scale rag or production level rag to be doing things before, you know, responses come out like crosschecking with another model and another process for things like, um, getting PHI and PII when you don't want it, right?
And so you can do that per site, which helps deal with compliance challenges. How are you dealing with a lot of changes in RAG right now? You got cag, you got graphic Rag, you got Agent Rag, and it seems like it could get a little complicated for you guys approach that You have that that is a great question.
And um, so what I'd say is the methodology for doing the RAG is decoupled from the theory of the inference mesh and the distributed data engine. So it, you know, that pattern. But we do actually supply some middleware that kind of helps with rag.
And we are in the middle of a pretty substantial revamp. Um, because, you know, we, we wanna actually make well of the point of the middleware that we have, you know, beyond like what common was WASA does operationally is to make development of these things that the app layer much easier. And so we're definitely pulling a lot of those things in around, um, graph rag, um, you know, enabling things like DOC ETL in pipelines, um, if those are fantastic paper about using models to basically enrich the metadata that use for embeddings that anthropic published, right?
Allowing people to easily turn up like model LM based, um, enrichment, uh, pipelines, things like that. So yeah, it's, it's a, it's a pretty big deal, but it will fit into the inference mesh pretty. Okay, cool.
And, and I'll say that's part of the loosely coupled aspect of this. If, uh, there had a graph rag service that they've already wanted or they want to inject a graph rag service, once again, our middleware will just point over to that and they can take it over. But Not one line of code Several at that point.
Yes. Obviously. Sorry.
Um, no, no bother. But um, but our opinionated Yeah. Is, is sort of No, I, and then our next Rev, like Matt's talking about, we're actually gonna add that feature into our opinionated one on the next.
You Had me, Lucy opinionated that, that's really good architecture. So, So, um, you know, I know, I know Kaas is, uh, a, a relatively new company. There's obviously the big ones folks would've heard of like, you know, uh, you know, could, you know, could say Snowflake, you could say Databricks.
Um, but then there was also their nstitute, like Starburst also a fairly, you know, young company. Um, you could probably say click house, uh, you could probably say, uh, uh, uh, one house. All these other kinds of hybrid data lake kind of approaches.
Uh, you, you separate, there are certainly these hybrid data data lake companies where where is the kaza differentiator that you see as compared to say like someone that's shipping a, a solution set for the hybrid data lake problem. Mm-hmm. Versus the one that you're referring to, which is kind of when ag agentic happens, you are sending all your fastest ships, um, but only bring the results back that matter.
Yeah. Don't bring the entire landmass of America once you've described, For example. Absolutely.
Yep. Okay. Uh, but keep in mind, uh, a lot of those you just talked about are structured data services.
Mm-hmm. Or any data service. So you're, We connect to anything raw, PDFs, word documents, file services, BMP files, we're gonna talk about those on weather mapping services.
Anything that AI can attach to, You can do the wins, you can do words, images, numbers, sounds, it's, It's, it's getting incredible. Okay? And you don't have to put it into structured format services.
Of course you can, you can have it in sql, you could have it in parquet tables. You can mix and match all of that. And that is what our customers are, are, are, are starting to utilize the fact that they could be in Azure, install our stack, connect to a snowflake instance, connect to the Databricks instance, also have their on-prem data that's sitting on their file service.
All in the rag process. All in the agentic flow is why you get outcomes that are outsized for the market. If you are just have, uh, an AI service that is just running in a data lake service like a snowflake or a Databricks, but you need data that's outside of that to complete even five or 10% of that request, that means you have five or 10% information missing.
'cause those AI agents don't reach into it and you actually can't run that process. You don't get the value, you don't get the actual output that's needed for that. Okay.
And, and that, I know you said loose, you know, loosely couple and you've talked about swapping, but I'm also hearing you say that you can compose differently as to what you're trying to achieve. So if it is, you know, cloud service A, cloud service B, and then some on-premises enterprise instantiation of a OLA or whatever it might have been, you can have those, uh, catalogs being globally to, I stumble for the outcome you're trying to achieve. Our inference mesh understands the global catalog of all of the data.
You, you mentioned lab, you would probably swap that out for the KAZA stack. Okay. Where the enterprise, There's a, if there's a, if I, if I paint an enterprise architectural diagram and I'm gonna rip something out and I'm gonna replace it with something your position would be is if there's an OLA appearing on that diagram, Kaza would attack that OLA space.
Uh, OLAP. Yeah. Ola.
Matt, do you wanna cover that? Because I was thinking of llama. Yeah, I mean honestly the use case hasn't come up.
Uh, and it's, for me, I, I think you probably have to tell me Jay, like how are you picturing like, um, the LLM models being involved in the OLA processing, like we do a lot of things. A common customer use case when it comes to like anything that looks like data warehousing or general analytics is they tend to get an LLMs involved a lot in um, kind of the ETL pipeline, ETL pipe, ELT pipeline, right? One way or the other, right?
They're kind of, it's a sing of data, transforming it. And one of those things that's really marvelous, and I'll pontificate for a second, right? 'cause I think this is valuable for everybody to hear, is that that sort of non-deterministic nature of LLMs makes them fantastic, right?
If I told you to write code and I said I want it to correct typos, you, you would have to laugh at me 'cause you'd be picturing like hundreds of thousands of lines of code and it would still be buggy and not work, right? And yeah, language models just get to do that, right? That's the kind of power of the fuzziness of neural networks and transformers and self attention in general is they deal with those abstractions, which makes them great for data cleaning and sort of recognition of things.
Um, and the other thing, and I actually do have a demo. Um, it's the only one I'm not gonna do live 'cause I just don't have, it's hard to do the environment, but I'll, I'll run it for you. Um, which is actually LLM driven data analytics where you kind of either take a data set and you let a language model kind of construct queries and stuff, right?
And our demo, we'll show you, we actually, um, you know, we have a query running with duck TB against backend per K files. Yep. And we give it like a natural language prompt and it, and based on that, the model actually reconstructs both the SQL query and it rewrites the react component to render it to the instructions, which is pretty cool, right?
Um, so that's a good idea maybe of how you're thinking about ola, but some of that stuff ranges from like baby steps to industrial. And so I think we'd have to know a lot more about the scale and the use case to really answer the question. Yeah.
That, so that, that's helpful because um, I was thinking it was displacement, but now what I'm hearing is the OLA might just be simply one or the other go off and run this query Augmentations And the computation comes back. Okay. Thank you.
So the last bullet, the Last bullet or code to run a query, right? I mean, a funny thing is we are certainly living in this world where everybody just can and will throw a more computer at everything. This is one of the reasons why obviously NVIDIA's being hugely successful, but, you know, it's very easy and you'll get a hint from this, from our other demos, which are completely, I mean, some of these are a little off the wall, like how, not how wild agents get, but the fact that you can just go and ask for data and it'll build a whole pipeline for you.
I mean, in fact, I can't show this, right? 'cause it was done for a customer. But, um, you can go see this, um, discussion.
This panel where the chief meteorologist of the DHS was talking about the work we did to extract weather data from some horrible like kind of ancient format called genack and converted to par K. And that ultimately was done by an agent. So I can't show you that work, but I will show you the agent working today.
And, and when you start to realize, wow, this thing can like actually just go and autonomously convert things, you start to get ideas about like, well, what can I do really? Right? And you combine that with like models like tab PFN, which recently got released that actually uses a transformer to do the equivalent of like, um, logistic regression models.
But except you don't train them, you just pass it the entire, what would've been your training set essentially as data. And it instantly infers a correct answer. And it, it performs better than models that are trained for four hours.
So I mean there, yes, you can tune it at some point, but it's com it's wild that you can do that now, neural networks being applied directly to um, you know, analytics questions, right? And that kind analytics thing used to be much more the classic predictive analytics models and not the LLM domain. And now you see this like transformer architecture or troop bridging those worlds.
So That fourth bullet on the other slide of our differentiation is the fact that this is entirely silicon agnostic. So as we're gonna show in a lot of our demos and services today, we will be doing some of the processing on a MD, some of the processing on Nvidia. It could also run on a Mac.
So our community edition or free developer level edition works on any Mac metal and above this laptop here is a lunar leg processor. It runs just fine on that. The new strict points from a MD, uh, maybe you want to get into more, uh, discreet level stuff like Qualcomm processors.
The point of that silicon agnostic approach means that at the edge we can be put on the right hardware at the data center, we can be put on the right hardware, and then the cloud we can be put on the right hardware, the new arm chips and inferencing, uh, technology coming from Intel, sorry, coming from uh, uh, Amazon, not Intel obviously. Um, the inferior chips, the uh, uh, uh, tensor chips over at Google. This agnostic approach allows us to literally run where it needs to on the right price, on the right concept, on the right scale, cloud, prem and edge.
So that opens up sort of that paradigm, but then you have the ability, once again, this is the full stack view. Um, it's, it's quite large. It's 150 plus packages.
But I just want to take the one step back on where this really goes and what the tagline, uh, over the slide says. If we connect to all of the data of the enterprise in all of its locations, in all of its forms, if we are running across cloud, prem and Edge, we are connected to all of the security constructs of the enterprise, the SSO and semi authentication structures for the L LLMs and the users that are accessing that data. And we've federated that and we're running that dispersed across the entire state of the enterprise.
It's not just our apps that customers build and run on common waza, it's also all of the third party and ecosystem apps that you simply plug right on top of common Waza. We then become, and it's a concept I literally built for this room, sort of the docker for generative ai. 'cause now you're talking all of these third party apps simply plug into us and get that access.
Just so I understand this correctly, right? So you're both, uh, a rag inference aggregator for gravity and data, which locality all that stuff. But you're also, are you also helping build the rags as well?
Uh, Matt's gonna, I mean, I just wanted know, Matt's gonna show you like DataWorks would Process through the Python. Okay. That's what saying.
But you don't have to do that once you lose couples. You can pick whatever, but yes, I know be honestly do that. But But the point being like, you look a company like DataWorks.
Yeah, right? That's what, that's their argument that they can take anything everywhere and ragg it. Mm-hmm.
So is that, that's, you're in that business as well? We're we're in the business of enabling that. Okay.
And doing it if you want. We're very, we're very much not SaaS, right? No, I know that are certainly, there's certainly tons of products that are sort of like use case specific, right?
And of course, I mean one of the problems is they're actually almost all delivered as SaaS. And so if you have any, anything you wanna do, and by the way, I mean this, this is not foreign. This is everybody.
If you wanna do something with private secure data that you don't want to put out in the SaaS API and I mean I'm, you know, it's really timely, right? Like just this morning we saw this news about deep seek getting breached, a bunch of their prompt and user data sure is just like completely leaked out in the public through some backdoor, right? I expect that's gonna have fallout, you know, in this direction, right?
Because it's kind of first major service leak. I think I've seen like that from a, a model provider in that way. And it's a great example of like why I think enterprises in a ceases at them.
It's not that they're necessarily negative about generative ai, but they're negative about shoving data into random generative AI services. Especially 'cause to your point about young companies earlier, you know, a lot of young companies, um, and you know, keeping that stuff in the four walls can be, um, very reassuring. So the, the catalog becomes all important in all this discussion.
'cause that's where the enterprise data resides. And so I'm trying to understand, so it's almost like you're parsing out the igenic workflow into various interface rag inter interactions and, and, and divvying 'em out to wherever those interface nodes in your interface mesh mm-hmm. Actually have access to that data.
Is that what you're doing? Yeah, yeah, Absolutely. Locally it accesses that data using the enterprise security it Has to, but how do you know, you just gotta pull that in to actually inference.
How do you get to a point where the agenda, uh, workflow is parsed out and to a point where you know actually what data is required for a particular request? Well, yeah, go ahead. Okay.
So, so if you, if you start thinking about agents, like really agents are almost like programs, right? They're kind of comprised of widgets and, and then you can have multiples and they can interact through interfaces where they're not aware of each other. People can build them on top of frameworks.
Like a G two is a good example, right? For formerly known as Auto Gen. And they're often doing, you know, a, a new version commercially.
Now, um, if you think about those things, every one of those agents basically tends to have some mix of prompts, input tools, um, those sorts of things. So with us, you can give them access to go read the Commonwealth's catalog, but there's the question of what kind of credentials they have or what kind of access they have, right? So you, through our system, can give an agent access to as much data as you have access to right?
Through a role. You can also create roles that have less access and give it to the agent. And now, you know, critically we're mindful of the fact that some things are deployed to act on your behalf, Mr.
User A, but also some things are things that you deploy to act on behalf of other users where it should actually have their permissions, right? And this is actually a really critical aspect of things that we're doing because we've built basically local auth, um, and J two B and t enforcement at the edge and awareness of these things that are at a, um, at a catalog layer so that you can actually take, whether it's local identity, OAuth identity, SAML identity, right? Identity is one layer and then the kind of access controls flow from there.
And then anybody can kind of pass them to applications and agents that leverage our stack. Um, and then we, we, we put a blog up entry on one of our, our GitHub doc site about two weeks ago that shows about 20 or 30 lines of code you can use both from all almost front end. We actually integrate an app directly with Commonwealth's authentication and with our model repository, right?
So you can actually build a completely private gen AI powered application, but then actually benefits from our implementation of, you know, the identity enforcement and we pass that to the catalog. Um, so, And Matt's gonna show this, but uh, and I didn't cover it. The stack is forward presenting everything via an API.
We have our own SDK as well. We have one API that is literally across the entire data state for this power, for this inferencing.