From Storage to Enterprise Intelligence, Unlock AI Value from Private Unstructured Data with CTERA
Discover the obstacles that hinder AI adoption. What matters most? Data quality or quantity? Understand the strategy CTERA uses for curating data to create trustworthy, AI-ready datasets that overcome silos and security challenges, translating raw data into meaningful and actionable insights.
CTERA’s presentation at AI Infrastructure Field Day 3 focused on the transition from traditional storage solutions to “enterprise intelligence,” highlighting the potential of AI to unlock value from unstructured data. While enterprise GenAI represents a massive market opportunity, with projections reaching $401 billion annually by 2028, the speaker, Aron Brand, emphasized that current adoption is hindered by the poor quality of data being fed into AI models. Brand argued that simply pointing AI tools at existing data leads to “convincing nonsense,” as organizations often lack understanding of their own data, resulting in inaccurate and potentially harmful outputs. He identified three main “quality killers”: messy data, data silos, and compliance/security concerns.
To overcome these obstacles, CTERA proposes a strategy centered on data curation, involving several key steps. These include collecting data from various storage silos, unifying data formats, enriching metadata, filtering data based on rules and policies, and finally, vectorizing and indexing the data. CTERA aims to provide a platform that enables users to create high-quality datasets, enforce permissions and guardrails, and deliver precise context to AI tools. The platform is powered by an MCP server for orchestration and an MCP client for invoking external tools, facilitating an open and extensible system.
CTERA’s vision extends to “virtual employees” or subject matter experts created by users to automate tasks and improve efficiency. The system respects existing access controls and provides verifiable answers grounded in source data. The presented examples demonstrated the potential of the platform in various use cases, including legal research, news analysis, and medical diagnostics. The presentation emphasized that the goal is not to replace human workers but to augment their capabilities with AI-powered assistants that can access and analyze sensitive data in a secure and compliant manner.
Presented by Aron Brand, CTO, CTERA. Recorded live on September 11, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/ctera-presents-at-ai-infrastructure-field-day-3/ or visit https://www.ctera.com/products/ctera-data-intelligence/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
For people who just joined, uh, my name is Aaron Brand. I'm CTO of Citer. Uh, so thank you for joining today.
And, uh, now I'd like to, uh, uh, talk about how sitter is thinking about, uh, the transition from storage to enterprise intelligence. So, uh, sitter has started from the storage space. We've seen, uh, um, our customers optimizing their storage footprint, reducing costs, uh, providing access to their data from any location.
Uh, but, uh, but over the time we have learned that organizations, after they consolidate their data estate and create this unified, uh, uh, data layer, they really want to unlock the, the value of their data, uh, for ai, for analytics or other kinds of processing. And this is, uh, what we call enterprise intelligence, uh, providing better decisions, better decision making, and better ability to extract the value from this data estate. And, and enterprise Gen AI is, is really big business according to Morgan Stanley.
Uh, they predict that by 2028, which is not so far in the future, it'll be about $401 billion market per year. Uh, and that's, that's about 23 2%. So that's almost a quarter of the total software spend.
So that's an amazing shift. Um, so why aren't we seeing enterprises doing this yet, right? It, it seems like it, right?
They they're doing it, but why aren't we seeing it in production? That $401 billion represents a net loss of 400 million. 4 billion, Billion.
Yes. So that's, that's the point. Uh, so, um, and this is a, this is, uh, something that was released just, uh, last month, uh, by MI MIT or MIT research, and they found as part of a survey of hundreds of, uh, enterprise companies that 95% of JAI pilots fail, right?
So that's, uh, right. Only one in 20 of these actually produces any ROI right now. So, and, and so why is that?
And so we, we've been talking to many, many customers. Uh, our customers are in, um, particularly very large and, uh, regulated industries, military, federal, uh, healthcare, and we're asking them, you know, why are your gen ai, uh, pilots failing? And what we recognize is that, you know, there's a perception about using gen AI on your data, which is, you know, we'll just point, uh, all these AI tools into all our data, right?
Uh, and we vectorize everything. We, we add a rag layer, we put, uh, plug in the smartest model that we have, right? The GPT five, and then we sit back and, and everything works like magic.
And unfortunately, you know, ai, it, it feels like magic. And I I, every day, I, I'm amazed by the capabilities of, uh, the new, um, models and tools that are created, but they're not magic. And, uh, and, and the, the reality is a bit more like this, right?
Uh, when, when you dump all your data, uh, as it is into these Gen AI models, all you get is more convincing nonsense, right? And the stronger the, the gen AI model is, it's more convincing, but it's not more true. And, um, and, and why is that?
Because organization don't really don't know what they have. They, they have no data from, uh, uh, decades, uh, old. And it's, uh, if you provide the gen ai, certainly you, you ask a question such as, you know, what does the regulation say about this action?
Right? And then it, it pulls some, uh, chunks of documents from different years, and you get some, you, you get wrong answers very confidently. And then what happens is, then what happens is that the, the person who is really an expert on this field sends a reply based on that, and then that it becomes the truth because he's, he's the authority, right?
So it becomes an authority because the person in authority trusted this incorrect answer. Uh, so that's, that's, that's a really big problem. And, and actually, I we're say, uh, I like to say that bad data is worse than no data, right?
It, it leads you to the wrong in the, into the wrong direction. So the next frontier, and what really is required in order to make enterprise gen AI success is not more data, it's quality, quality data. Now, um, and, and we analyze this, and we, we found, uh, that there's three main, uh, quality killers for data.
The first is that the data is messy. The data is messy. It means that you have irrelevant data, outdated, uh, information intermixed with good information, uh, to different levels of quality.
Um, and in order to solve this problem, what you need to do is to curate your data, to classify the data that you could use gen AI or AI tools to, to do this classification and to enrich, to do metadata enrichment. We, we mentioned that in previous, uh, parts of this discussion, uh, to enrich the, your files with additional metadata that allows these, these, uh, AI pipelines or rag pipelines to pull in the correct information, uh, and the most, uh, the highest quality information needed for any specific, uh, task. The second problem is the problem of silos.
And by silos, I mean, I don't only mean that the data is split into different, uh, systems, right? This could be split into, right? You could have some of your data in SharePoint, some of your data in Confluence, some of your, of your data in the global file system.
Uh, but it's also, uh, so it's dispersed across, uh, geographies, right? You can have different data stored, right? In the, in the us different data stored in Europe.
There might be, uh, um, data penetration or rules that require you to source certain, certain data in certain places, or it just due to historical reasons. Uh, and the data is also not stored in a unified format that allows gen ai, it's not ready for gen ai, right? Because if you have, you have a video or you have a PDF that's scanned, or a PDF, that's, uh, um, or a word file, right?
You really need to bring all these things into a unified place and format. You don't think vectorization Does the unification, Uh, the vectorization needs to be done after you, after, after the unification, right? 'cause, uh, um, a a a vector database or a embedding model is not able to look at a binary PDF line, right?
It's just, uh, you have to convert this to something that is, uh, either textual or you can use, if you have a multimodal, um, uh, uh, embedding model, you can feed in images, and then it will be able to answer different kinds of questions, right? But, um, so you really have to, uh, convert to some kind of a format that, that the embedding model will understand. Um, and finally, and this is the, uh, probably the, the, the worst quality killer or the, the, the biggest killer of POCs is compliance and security, right?
Right. Because, uh, the compliance problems, especially if you're a regulated industry, are, are enormous. I mean, the, the, the problem that will happen for a company, if for some reason their gen AI assistant will start spitting out some PII or, or health information or, uh, or information about payroll, it could be disaster.
It could be a disaster for compliance, and it could get, it could get fined, and it would be very embarrassing. And it's very hard, right, uh, to predict what will happen. You have this, uh, uh, black box.
People ask questions, they try to trick it, and you have no idea what was picked back unless you know what you fed into it. So it's very, it's very crucial to have guardrails around the data that is ingested into the system. So at the ingestion level and at the output level, right?
So to protect the system, uh, to look at, uh, the output, look at the re uh, all the different steps of the pipeline, and to enforce guardrails about what is allowed within our, uh, organization. And, and, and finally, you need to, you must then, this is, uh, uh, you must preserve the access controls of the original files, right? Because if you assume that users have access, the, the, the, uh, there are many systems that allow you today to make sure that the access controls are correct, right?
And, and almost any big regulated industry has these tools, uh, like, uh, Coronas and, uh, other, uh, data management tools that allow you to set, to make sure that the acls of the files are correct based on what's in the files. Um, so we, we want to preserve this when we add the, the JI assistance, so it'll enforce the same permissions that files have today and not lose them as part of copying your data to another unmanaged this. Uh, so these, these are the, Yeah, we've just been kind of discussing this in the back channel.
All of these things still don't guarantee that you'll get the right answer. Absolutely. There's no, there's no guarantee of a right answer.
But these are the basic requirements before you could even start, right? So there are a lot of steps after that. And, and this is, this is exactly, I mean, nothing proves Yeah, yeah.
Nothing will, uh, allow to guarantee correct answers. Um, and, but I'll show in a moment what what we are doing. And, and that is to provide, uh, clear the ability for the users to verify that, Well, it's what would require getting the right answers that artificial intelligence actually had some intelligence, which it does not.
Uh, yeah. That, that's a philosophical question. Whether artificial intelligence has intelligence and whether, whether, and I wanna question ask your your question.
People have intelligence that that's also a, it's Not philosophical. It, it's artificial intelligence today is just echoing back garbage that it was given. Um, and so, are we right?
You could have said that. You could have said that. Um, well done, sir.
Yeah, I, I could talk to you, you know, hours about that. Okay. I'm about to derail your presentation Not too late.
We, um, the, the ai Yeah, you're, you're, you're absolutely right. It's, uh, it's very, uh, it, it's a, it, you give it bad data, which is bad answers. And, uh, we are trying to do our best to get, to be, to allow the organization to curate the best data, most better, best quality data, and to allow the users to verify we're not, we're not in a state today where we can allow the AI right to, to trust it entirely.
Um, and I don't know if we ever will. Um, so, so the step, the steps of curating data for gen AI involves, uh, first, uh, collecting them from the different storage silos. So it could be from NFSS, MB S3, SharePoint, OneDrive, uh, whatever you have.
Um, and the ingestion should be timely, right? You, you don't want to have this collection happen on a monthly basis or on a yearly basis because you really want to get relevant answer, right? You wanna ask, what are my competitors doing on this matter?
You? And if they announce something yesterday, you want to know that, uh, the second, and then you have the format unification that we mentioned, where you have to bring all these different, uh, archive. Some could be scan PDFs, some can be video, some can be audio meeting transcripts, uh, word files, and build some kind of a unified format that, uh, that is more suitable for gen ai, uh, enrichment of the metadata.
So, uh, you can, uh, take, take these, uh, documents and enrich, if you have a doc, uh, a repository of contract, you could have, you could say, what is the signing date? Who signed it? Who's the customer who's, okay?
So you, you add this, uh, semi-structured layer over the unstructured data that helps the, the rag engine to pull in the right document. Uh, and then, uh, you do filtering of the data. So you would instruct the system to drop in, uh, information based on certain rules.
So you would say, okay, if it, if it contains personal health information, drop it, drop the file or, or redact this specific part before you, before you, uh, provide it, uh, insert it into the index. Uh, and only after you do all these, you can do the vectorization and indic. So where does Sentara play in this flow?
That's a good question. I'll show you. Um, I have a question too.
Yeah. But, uh, before we get, can you stay on that slide real quick? Yeah, of course.
So, um, the timely ingestion and format unification, um, so what happens, like, let's take your example of a competitor announcing something. What happens if that's on a longer document that's a PowerPoint or a document A doc? Um, it gets pulled out of that document, but what if I make an update to the document?
Do you guys update the format? Unified? Yeah.
Five. Yeah. So, so, um, uh, so I'll show it in a minute, but, uh, uh, what our data intelligence system essentially does, is to connect to all these data sources.
The, the main one would be our global file system, and it, it, uh, works using our notification bus. Uh, it wakes up with it wherever the file is modified, and then it re, it runs the pipeline again and replaces the, the chunks of the file with the nuance. So, yeah, that, that's exactly what it does.
And so you have, so you can be sure that you have the latest version. And, and typically this happens, takes can take seconds, right? I, I can drop a file into the system within seconds.
It's already available for question answering. So that means I'm keeping two versions of the file. I'm keeping the file that's human readable, that's in my format that I wanted, as well as this, this reformatted for, Yeah, that's a semantic layer.
So we, the system builds this semantic, uh, version of, of all your documents, right? And it, it stores a textual version. It stores some, uh, descriptors that contains some metadata that was were extracted from the file.
And this is where it actually is used by the gen AI and not the original, uh, source documents that you can even move them to archive, right? You can move them to, uh, to glacier. You don't have to pull them in in order to answer the question.
Okay? Thank you. Um, so our, our vision, uh, of data intelligence is to, to provide us a platform that allows you to move all the way from curating your data, right?
So creating these high quality data sets, um, carving them out of your data and, uh, and tagging them and, and enforcing, uh, the, the permissions and, and the guardrails, uh, creating guardrails over what is ingested, uh, a retrieval layer, uh, that allows you to, uh, deliver precise, uh, uh, context that you can provide to your Gen AI tools. And then on top of all that, our vision is what we call virtual employees. Virtual employees are like subject matter experts, uh, created by the users themselves, okay?
So the system is, is designed in such a way that, uh, a lawyer or a, a secretary or anyone within the organization is able to, uh, automate the most annoying parts of their work, or the parts where an agent would help them do their work faster and more efficiently, and they can encode their knowledge in a very simple way into virtual employees that we call experts. And, and finally, um, we build the system to respect existing acls or other per permissions from the source data. The users will get only answers based on something that they're allowed to access anyway.
Uh, and we built it, uh, we built it in a way that, uh, is, it provides trustworthy answers that you can, ver you can verify and validate because they're, they're grounded, and it shows you the source of the, of, uh, every answer. And, and the vision is to extend your, your workforce with a virtual, a virtual, uh, uh, team of assistance that every user can add for themselves. And, and so it's not replacing users, it's just making their job more efficient and, and more enjoyable.
And, And this, uh, platform is powered by M-C-P-M-C-P is the underpinning of everything that we do here. So it has both an MCP server, so you don't have to use our user interface in order to inter interact with these experts. Uh, you don't have to use our chat interface.
We have, we provide one in case you want to turnkey solution, but in fact, most, uh, we, we believe that most of the usage would be through MCP allowing you to use to access this wherever your employees are present, right? Wherever they do their work. So if they work in teams, in co-pilot, in chat, GPT or anywhere else, your, their virtual assistants are there, help them.
Uh, and the other side of this is the MCP client, because the system also allows invoking external tools, right? We don't want the system to be a closed system, want to be open, and MCP client is the best way to do it. So you, uh, you can tie in any external tools using that support, the MCP protocol, it could be you can, you can have the system, uh, you know, send an email, uh, create an a document in Confluence, uh, or, you know, uh, whatever system that you have, either using one of the thousands of, uh, MCP servers that are, are available today, or you can develop your own.
So it's built to be an open an open system. And now I'd like to show, show you c Teradata intelligence in action. Okay?
So, uh, I'll, I'll, I'll be switching back and forth between the administrator interface and the end user interface. But the, and I want you to remember that the end user interface is only an example, and you could do everything through, uh, MCP or through other AI tools. You don't have to use ours.
Um, so the most important part in c Teradata intelligence are the experts. The experts are the virtual employees. Um, and, uh, users can define these own experts.
So you could see we have three experts that are defined here, a paralegal, the news, news analyst, and the lab assistant. Okay? These are user defined, and they're, uh, we have, uh, this collaboration capabilities.
So these are my experts, and I have a shared with me tab, uh, that chose experts that will share with me, okay? So this is the self, self-provisioning aspect, uh, of our platform, uh, where, uh, we don't, we don't think that people from, it should create experts for their users, right? They should do their own.
Uh, and here you see, uh, a Navy, Navy, uh, maintenance expert is shared with me. Um, and now, um, LLMs, the system is built to support any LLM. So you could use flaw, deep seek, uh, open ai, whatever you want.
It's built in an open platform. And even private LLMs like lama, uh, it's really not dependent on any specific LLM, uh, it's built in a generic way. And, um, as you can see here, uh, each of these experts has, uh, it has some documents or some, some dataset that's associated with it.
Uh, the dataset can be from different sources. So you can have ctra portal or Global File system as a data source, or you can have a website, or you can have Confluence. So, and we're working to add more and more data sources to this system, uh, as connectors.
So it's built to pull in your data from wherever you, uh, your data is. And, and then after that, uh, you have data sets. So data datasets are curated, high quality, uh, subsets of data selected from your data sources.
Uh, and as you can see here, I have legal documents, medical documents, and these are also self-curated by, by your employees. Now, let's go into the chat interface. Um, and let's look at the paralegal experts.
So, and I'll ask it, uh, do we have petroleum contracts? Okay, so it's, now, it's using MCP to run a cera search. And here it provides me three, three contracts.
And on the right, uh, you can see, I could see the, the text of the original file, and I could verify this response is correct. Now, I, I'd say perform comprehensive research, uh, and I wanted to, uh, to compare three contracts. As you can see, it's running quite a few, um, tools, just doing a deep research and providing me a table comparing three contracts.
Now, this, this will be a very time consuming job for a person, right? Uh, you have to read the contract, compare them side by side, and it's done it in seconds. Now, let's look at, uh, a news analyst expert.
This one is an, uh, expert I built based on Gemini approach, a very, very powerful model. Um, and the, the role of this expert is to use, uh, MCP in order to pull in, uh, an RSS feed, uh, read the documents, um, referenced by this news news source, and create a, an analysis that is suitable for me and tailored to me as a as. And since it, a system knows my role, and it knows who I am, it works, it's integrated with active directory, it knows my role description and my groups, it's able to actually provide an analysis that is tailored to me, right?
So, um, here's an example for the MCP servers that I connected, and it's this expert is using the URL Fetcher, uh, expert that in order to bring in the RSS feed, and then bring in the files that are, or the documents that are referenced by, uh, this, this, uh, news source. Um, okay. And, and as you can see here, as the user provides plain text instructions, how they, how do they want this expert to operate?
What kind of, uh, answer it needs to keep, what, what, what is its tone, uh, what it wants to avoid, uh, and, and it doesn't require any programming. So let's try this. Uh, so I just paste in an RSS feed from, uh, Berkeley, uh, the Berkeley, uh, AI blog, i, i plus president.
And, uh, let's see what happens. Okay? So it says, I, I will produce an analysis of this blog tuned to our interest as CTO.
You see, it's tailored for me. And, and now it's, uh, it's you calling the MCP tool. It's fetching, uh, fetching the RSS, uh, feed, and then it's fetching all the articles, uh, that are mentioned within this feed.
And, and again, I, I didn't develop, you know, single line of code in order order to do this. And it provides me this report tailored for Aaron Brand, CTO, uh, based on this. So it's, it could save me a huge amount of time to understand what is relevant for me within this, uh, news source, right?
So, uh, uh, and all of us, you know, how, how much content we have to, uh, stay read every day in order to stay updated. Think about getting this kind of report every day that summarizes everything while knowing your interest. So I think it's, it's, uh, amazing.
Now, let's jump into a medical use case. So the system is built for very sensitive, excuse me, Aaron, Before you continue. All of these agent activities are running where, Um, so, uh, the, the data intelligence system incorporates, um, um, a chat interface, right?
Uh, which you can use if you want. Mm-hmm. And, and in this case, the chat interface has a backend that orchestrates the calling of tools, right?
But at, at the same, at the same way, you can also use Microsoft copilot, which already has also the ability to call MCP tools, right? And you could do use it from there, right? So it doesn't really, it doesn't really matter as, as long as the front end supports things.
No, I understand that. But so you're, you've built a data intelligence solution into the Ciera storage system. Uh, it's, uh, so this, this is, uh, an add-on, right?
Uh, it's an add-on. It's not, uh, something that all our customers have right now. It, you have to buy another part of the product and, uh, and it can work with, with the, uh, information that you have in the global file system or other data sources.
Okay? Makes sense. But, and so, but so all, so I mean, like, this is running on the SATERA processing engine, the, the storage solution, right?
Yeah. Okay. Um, Okay.
The add on is the MCP client, MCP server. Is that how you package it? Is that kind of the way it's So, yeah.
Uh, the MCP server, uh, it's contained in MCP server that allows you to invoke these experts from tools that you have, like, uh, workflows like NA 10 or from copilot or from any, any other tool that supports, uh, MCP. And it has an MCP client that allows it to, uh, allows you to deploy tools within this ent, uh, flows. So like, uh, fetch a URL, send an email, read my mailba.
You can do all these things, and you're not limited only to operate on the u on the, on this private data state. You can also, it gives the, the eye, uh, you know, the arms and the, and the eyes for this, uh, AI model. Um, now the system also knows how to classify and augment the metadata of the files.
And this is what we, uh, Simon, uh, was talking about earlier, where, where I could say, I want you to extract certain, uh, certain, uh, uh, pieces of information from every file in this dataset. And I do that by, uh, specifying adjacent schema. Uh, in this example, we are, I'm showing a medical use case.
So there, there would be examination date, examination type, uh, and so on the different fields that are extracted, and we use, uh, uh, gen AI models in order to take this, uh, unstructured document and convert it into this kind of a semi-structured format, uh, that's, that's appended to each document. And we can use a combination of this metadata together with the, the semantic search in order to come up with the best, uh, answers for each question. And you can even use this information for analytic use cases.
So you could say, how many documents do I have that have examination types that to blood tests, right? So it will, it will give you a numeric answer because it's able to, uh, do like an SQL query on this data. So datas will be 42, 5 42, still 42 Yes.
Answer for, for everything. Mm-hmm. Um, Well, you think this is the, maybe I missed it.
Sorry. This is the metadata enhancement definition. Correct.
So we're in the meta metadata enhancement phase, right? There's something very human readable about that. Is there a way that, that you are, so the, the closer you can get, um, to the, the, the, the expert mm-hmm.
In this case, you know, a, a medical data analyst type person who defining the schema without requiring them to understand schema formatting, right? The better, the better off you are, right? Yeah.
Um, So, so, But this doesn't strike me as the medical data analyst interface of choice. Um, so, so you're asking about why, why you, are we providing it at JSON? Um, Well, sort of that's the technical question.
It's more like, uh, and maybe it's an unfair question. How are you enabling the, the, the, let's say the stakeholder, the user stakeholder to influence this without Yeah. A layer of, without an additional layer of, yeah.
Well, if, if it's a simple classification where you have certain fields, um, um, so we, we, we have some simple classifiers, so such as the keyword classifier, where all you do is enter a set of keywords and, and then it'll help tag. So this is not the only way to do the classification. You could do it without the LLM.
Um, we started with the power, we started with this, what I showed you is the most powerful classifier. Uh, and it's conceivable to create a nice user interface about this, right? Sure.
Uh, to, to edit this. Uh, we haven't done it because it's so powerful. It allows, 'cause the json schema allows things like lists.
It allows you to do things like, uh, you know, things that are not one dimensional. It could be nested. So it's pretty hard to do this with user interface.
We might do that in the, in the future to have like a builder, right? Yeah. But we, but, but you know, if, if you're a user and you don't know how to write JSON schema, all you need to do is in pt, you know, write me a js o schema for this.
You need to know to ask that, but yeah, fair enough. But, okay. But, uh, yeah.
But yeah, I agree. It's a usability, uh, good usability enhancement that we're considering doing. Um, so, uh, just show you a bit the results of the classification.
So I have some medical reports. I click here on this, uh, document, and I see here the metadata. And you could see down here, medical document metadata, the name of the doctor, some review of what's in this document, the examination date, and the examination type.
So this is so powerful for every, uh, enterprise that has no idea what they have in their file system. This is like hand levels up with what they have today. They can really understand what they have in their data.
Now, let's see this in action. So find all kidney function tests in our database, right? So the system in thinking about and providing me a list, uh, of the, the tests that I requested, and with a summary of the findings, and now create a table of the tests with anomalies.
Okay? So it created, it's, uh, providing me the, the result. Please explain each anomaly.
And now here, the AI is using its own, uh, knowledge in order to explain what this means and create a presentation using reveal J. So now, right now what it did, now it created code, JavaScript code, uh, create a presentation, you know, within seconds from the data. So yeah, I think that's, uh, you know, very, very impressive.
What, what is a, what the AI is capable of doing today. So how many HIPAA violations are in That question? Is there a way, what you've shown us is, uh, what I would assume would come up if you've got the data and it's completely raw.
Yeah. Um, there's a lot of that, uh, of what you've just shown us that is really objectionable from a compliance point of view. Of Course.
Uh, address that for me. Yeah. So first of all, this is not real information, obviously what I, what I showed you here, and this is not something that, uh, uh, will be available in the open, but, uh, uh, and now the system has the ability to, uh, to redact specific fields.
So if, if there's a name of the patient right here, for example, we didn't, we didn't enable that this was just, uh, made up names. Mm-hmm. But, uh, you typically have it an anonymized, when you, as part of the feeding process, we'll just remove the names of the people.
And is that a, is that a manual process? Is that something that you can automate? How, talk to me about that, because I'm envisioning a healthcare provider that might have hundreds of thousands of these, these, you know, lab reports.
And yes, I understand the scientific value of being able to say, okay, give me all of, you know, the times that we've run this particular exam exam or this particular lab in our facility, I wanna, and I wanna have some statistic analysis on that. But the flip side of that is that in order to be able to have that data, it has to be anonymized in some fashion. So either somebody's gotta scrub it manually, uh, or there's gotta be some kind of way that you can assure, um, you know, compliance authorities that, that, you know, your AI bot is not just pulling up all of this hipaa, you know, hip, hipaa sensitive material of course.
And, you know, the security folks are just basically going ape s**t. Absolutely. Absolutely.
Yeah. So I, I, I agree this, this might not be a very realistic example, what I showed now, it was more to show the power of the system. Uh, the system does have the ability to redact PII, uh, and anything that looks objectionable, it, it, it use, it removes it on the source.
Um, and, and I, and I do expect that things like this will be done on data is that is already cleared by the right committees, uh, that is free anonymized. And, and our guardrails will just be another, another step to, uh, to add another level of assurance. So manual, then, it's a manual process at this point in There.
There, it has this abilities. But, uh, you know, the LLMs, I wouldn't say that, you know, they, if they can't be, uh, tricked or can, they can make mistakes. So this kind of a use case, uh, would typically run, uh, on data that is already pre anonymized.
That that, right. This is an example for a research use case. And you, you wouldn't just throw, uh, information right from the patient, uh, charts directly into this.
Uh, so I, I take it as, as feedback and, you know, you're, you're, you're right that, uh, this might be a little bit alarming to people from, from the industry, and I'll, I'll make sure, I'll make sure to correctly For sure, for Sure. Um, I think it's, I think it's alarming because I mean, just that, that last line there of like pro provide a prescription for the first patient, right? The alarming thing for me would be that could be sold into an industry that doesn't understand this is not a source of truth, right?
Mm-hmm. This is the statistical generation of, of something. Mm-hmm.
And so it could turn into something that we're all afraid of, right? Um, yeah. So we, we we're not able to solve the, the, you know, the, the problems, the, the fundamental problems of AI and our trust for AI and, uh, uh, but we, we, we are providing the tools, uh, we're, we're providing the, uh, the way to deploy the agents.
And so, uh, but, but it is designed to be used on sensitive data, right? So the system is built, uh, and, and it's, uh, you know, there are many different, uh, AI frameworks and, uh, and SaaS, uh, rag platforms and so on in the market. But we, we are really try to, we try to make it, and this is our focus, to connect your private data state and the sensitive whatever, you don't want your doctors to post it, paste into chat GPT.
Okay? So, um, there will be a lot of permissions around this. There will be a lot of, a lot of, uh, guardrails around this and, and it'll be used.
Uh, uh, it has to be used, but in a responsible manner. Absolutely. Yeah.
I mean, because it's, it is for, and so I think, uh, one thing that I'd like to kind of, uh, focus in on, because, you know, we, I came down on, on this particular example, um, because it does have objectionable information in it, in the, in the context of the example. But there are use cases where, you know, I'm the healthcare provider, I'm providing healthcare to Ms. Smith, and I need to pull up something from Ms.
Smith's file, and it's, and it's information that I'm searching for. That's very cus that's very, um, patient specific, but it's also something that is appropriate. Mm-hmm.
And so I think one of the takeaways that I'm getting is that, is that before we can even feed this data to an an AI or to an L-L-L-L-M, excuse me, or before we could even have a, an AI bot access this data, we really need to be very careful how we curate the data ahead of time. And I think that, um, that is the thing where, where we need to focus more time on mm-hmm. Is, is how many, how many different ways can we come up with bad outcomes because we haven't carefully curate, Come back and join us for AI ethics field day.
That's, that's, that's a thorny field. Yes, absolutely. I, I, yeah.
So I, I, I totally understand, right? This, and, and, and this is really our approach, not to throw all the data into ai. And, and even worse than that, we we're seeing, uh, you know, customers are talking to us and they're saying, today we have some kind of an AI tool.
And the way it works is that, uh, our employees take files, they drop them into this other system, and then, uh, and then they can share with, with each other, and they get, they get, uh, responses, and then they have another unmanaged shadow data copy of all these sensitive files in some kind of a SaaS system that nobody knows, and you know, what's there. So this is really the, the core here. We believe that we need this system for the, for the more security sensitive part of the business for, for the, uh, for these.
Uh, but, but it, it can provide you real, uh, business advantages, right? If, if you're a legal, let's say that you want to, you're a company that provides legal, uh, medical advice, right? A legal advice in a loss, a lawsuit, right?
You, you might have 300, uh, 300 or 500 documents related to this lawsuit, right? Now, you need, that's lawsuit. You want to save, you want to save the cost.
You, I mean, you don't want to have a doctor. Uh, you know, that's what they do today. They, a doctor can work a week, they have to analyze this thing, and then you, you build a patient for $20,000 right now.
Now what, what if we can do that, uh, on a comparable level within a minute, right? Uh, that, that would be a game changer for healthcare or, right? Or it will make our health insurance cheaper.
So we have to find a way to make it secure, um, that, that's what, that's what we believe in. And, and it's, it's up to the, up to our customers to find what is, what they can do within their regulation. But we will provide them all the tools to, to do this in a secure fashion.
Um, now, uh, here, here's an example. Uh, you can see here, I asked, um, uh, I, I created a guardrail that says you're not allowed to provide direct medical advice. You say, when I said provide, I, it answers questions, but I don't wanna say, provide a prescription for the first patient.
And it says, right, your request violated corporate policy. Okay? So, so these are the examples of the, the tools that we're provide and our customers, you know, um, should do it to the best, you know, to, to do, to make it as secure as possible while respecting their, uh, regulations.
Um, finally, I, I have, uh, one more example, uh, actually two, uh, this example is using the system through Microsoft, uh, uh, uh, copilot. Okay? So I just wanted to show you that I could, I could ask questions through chat, GPT, or sorry, through, uh, Microsoft copilot, and it will, uh, use MCP to, uh, invoke the system and do the same incompliant thing that you just, uh, complained about.
But, uh, yeah. Uh, so you could do it. Uh, you're not restricted to do that through, uh, our interface.
You can do it, do it through wherever your employees are already active and, uh, improve their efficiency. Um, and the final thing I wanted to share you is, uh, an example for, um, main, uh, maintenance and engineering use case. Uh, so in this case, we have, um, uh, uh, a naval Maintenance and engineering, uh, officer that's trained on very, very old, uh, documents.
So what we, what we actually, I fed here is some scanned, uh, documents from the 1960s about maintaining helicopters and, uh, all kinds. This is a very technical and complicated dataset, and also the quality is not very good. But, uh, we're actually using, uh, VLN or a vision language model as, uh, a state of the art, uh, approach for OCR and for question answering on, uh, image data.
In this case, I'm asking, I'm, I'm pasting a, a, a question, a mission debrief, and asking what is the corrective procedure in case we had an over RPM incident on our engine and this, uh, and it's running the search and providing me the required procedure by looking at these documents, uh, while providing me these citations, I can click on and I can check that this is really the correct, uh, procedure. Now, you might have, uh, you know, as a, as a naval officer, you might have documents, you have might have a, you know, a stack of documents that, you know, smell like engine oil up to the up to the ceiling, right? And you, and might take you hours to find the right answer.
And here you can find it within, within seconds. So, Well, you can, you can actually find it within seconds, uh, if you actually just scan it and do an OCR of the documents and then use a general purpose search engine on it. Um, so it's, uh, the OCR will is, um, this is a pretty much state of the art OCR.
So I've never seen any OCR tool as good as, as what this VLM is able, including understanding charts, uh, tables, uh, it's just amazing. Uh, it's much better than anything I've seen before. Uh, so yeah, it's, it's a new form of OCR.
Um, and, and now the, the, and, and the the next part is answering, answering the question rather than looking by the keyword, right? 'cause there was nothing in my question that was actually a keyword that existed within the document. So I, I could search for maybe over RPM and I would find over RPM for this, for this helicopter, and for that one and for, but, um, so yes, it can be, it can be done using, uh, previous technology, but when we actually showed it to customers, they said, oh my God, I mean, no LCR is able to handle these documents.
Um, okay, yeah. But of course, and, uh, the system, um, you can use, uh, ful tech search and uh, uh, I would say, uh, search engines exist for, for decades. And, uh, but, but the reality is that they are these, uh, very technical and complex, uh, uh, cases.
And, and, and the reasoning, right? The reasoning, uh, capabilities of the system don't exist in the search engine. So, uh, because it's able, it's able to do multiple, multiple step questions, right?
And do multiple searches within the same, uh, the same query and combine information, And you're somehow implying that there's a reasoning capability in today's AI LLMs, and that is definitely not the case. Um, okay. That, that's again, the argument whether we have good reasoning capabilities, and I think that, uh, I I, I, I'm sure that they, um, you know, some, it's the AI is not, not worse than the worst people in that, in that capability That that's a very, very low bar.
Um, Yeah, I the point, Yeah, the, the, the proof that we, if it's useful, right? If it, if, if, if people feel that it's useful and they'll use it, uh, and, and, and if the accuracy is good, then it, it does what it needs to do. But when I, when I say reasoning, what I mean is that the answer is not necessarily within one page of one document, right?
The answer could be, uh, in order to, uh, the system knows that, like the question like, uh, I showed of compare these three, uh, three documents, right? So there's no answer to that in, in one of the documents. You need to do multiple steps in order to, uh, to answer this question, and it's just doesn't much faster than a human.
Um, but, uh, yeah, I, I, I, your skepticism is, uh, I, I, you know, um, that's, that's, uh, very understandable. And, uh, but I think, uh, you know, we're, we're in a very fast evolving field, and even if it's not good enough today, I'm sure that in two years, um, no, that will be a thing of the past. Do you actually supply a Vector database and, uh, embedding solutions and things of that vectorization?
Or are you using other tools for that sort of stuff? Um, good question. So we, um, the embedding model, uh, we, we currently, um, allow, we allow you to select the model.
Like we allow, we allow you to select the, Or whatever who, so whoever wants to do the vectorization. Yeah. But our, we actually recommend, uh, the, the open source models that we prefer over the cloud solutions because, uh, one of the concepts here of the system is that we, we don't want your information to move to the, to the, the raw information to be moved to the cloud.
That's very expensive and very large. So the operation of the, all the initial operation of indexing and preparing your data, we prefer it to up happen on premises, uh, within, within your GPU infrastructure. And, and actually the embedding is a relatively lightweight part of this.
Uh, Yeah. Yeah. The vector database is yours.
Yeah. Or is it somebody else's? Uh, some, it's, uh, we, yeah, so we, we work with, uh, open source, uh, open source, um, vector database, and, um, um, yeah, so it's a pretty much a standard, standard solution.
Right? Anything else? The interface there that you provide?
So what's, what's the licensing or sales model on this? Is this is, yeah. How are you, are you selling it?
Are you adding it? Are you, what are you doing? Um, so the system is, is, uh, we marketed today as that, uh, first of all to our existing, uh, client base.
And when we sell to our existing client base today, we mainly pay by the capacity they pay by the capacity, yeah. So we have it as an add-on to, to their system and, and it's based on the capacity of their global file system. Now we have other customers that are coming from the AI side of things where, where the problem is not actually related to storage, but it's related to, okay, we have this data set and we want to, uh, and that's very sensitive and private, and we want to have this rag engine and, and answer engine.
Uh, so in those cases, uh, it's, it's a bit complicated 'cause it's possible that you'll have a very small capacity like, you know, um, terabyte, right? One terabyte and, and it will be worth a lot business wise. So this, this, these are, uh, a bit more complicated.
So we have some kind of a minimum price for this solution. We don't really want to, uh, go down to one terabyte, so there's like a minimum. Um, so yeah.
But, um, so That, which you would take under management, you would tell you pull that one terabyte or, or one petabyte into te even if the rest of what they're, because they're, this is a, this is a greenfield for you, and maybe for them, They, they can use, uh, they can either move it into our global file system or they can keep it at, on their source system. And, and we, they only use our semantic layer. So, Ah, okay.
So, so if I get this right, so you, you, this is an add-on for existing Satara customers because hey, they're gonna bring more data in, they're gonna pay for that data within satara, and this is an add-on to that. Yes. That's encouraging more and more uses of more and more data within the CITA system, that that makes sense.
But then if you're coming in from a different angle on it, um, and you're not putting everything in Cera, at least not yet, then you're gonna pay for that. Exactly. Okay.
Makes sense.