AI Transformation: Moving from Pilot to Production with Aaron Fulkersom at AIE 2024
As the Business Track leader, Aaron Fulkerson will kick off the discussions by addressing one of the most pressing challenges in today’s enterprise landscape: Transitioning AI from experimental pilot projects to full-scale, secure production. This session will explore how businesses can navigate the complexities of deploying AI while ensuring the security of sensitive data and adherence to data sovereignty. With a focus on strategic insights, the talk will provide practical, actionable solutions for AI and business leaders to harness the full potential of AI without compromising on privacy or operational efficiency.
Transcript
Thank you all for tuning in. I'm really excited to be here representing Techstrong and the artificially intelligent enterprise on this, uh, seminal event. And I am excited to speak with you about AI transformation moving from pilot to production, which and 2024 have been the years of a massive breakout in ai, but the enterprise still remains languishing and pilot or sub prod, and they're really struggling to get into production.
So it's an exciting topic I'm happy to be speaking about today. Now, if we look at the last couple of years, it's really clear that chat g PT was the catalyst for what will undoubtedly be the most significant technology supercycle in human history. And I say the catalyst because this is what we commonly see with these super cycles.
There's a technology that sparks the imagination of people at exactly the same time when the effort and expense of adopting a new technology becomes dramatically lower. So for example, if we look at the internet, it was Netscape, the Mosaic web browser, and Netscape that sparked everybody's imagination. If we look at mobile, for example, it was the iPhone.
In this scenario, it was definitely chatt PT that sparked everybody's imagination about what can be accomplished with AI coinciding at a time when the accessibility of these technologies just got much, much better. So for example, in 2023, the global AI market size was valued at almost 200 billion with a 37% compounded annual growth rate. So that's a pretty big market, and it's growing really fast.
However, while AI is now an existential business imperative, these projects are getting blocked and they're getting blocked because of challenges related not to technology, but to data. So what do I mean by an existential business imperative? What I mean is all businesses have to reinvent themselves with ai, and if they don't, they're going to get disrupted.
When I say that, the challenge is data related. The fact is the most valuable data for AI projects is the most sensitive data, and that's what's creating the challenge. Specifically, there's three areas that we commonly see as the top challenges in getting AI projects into production.
First, data's fragmented across a multiplicity of silos and the enterprise and organizations really struggle to bridge these silos, not because of a technology challenge, but because of a governance problem. How do they ensure that their data is kept private, secure and they retain their sovereignty? Well, that's the challenge with these multiplicity of silos, particularly when many times the governance policies across these silos vary and are incompatible.
To make matters worse, we're seeing a rapidly changing regulatory landscape. There's new regulations coming out all the time in Europe, in the us, even across individual states within the United States. That makes it very difficult to understand how to operationalize your data.
So that's a pretty well known challenge. One that's lesser known and emerging is the threat landscape. Specifically through the use of ai, it is easier than ever for hackers to gain access to data that's being operationalized inside the enterprise.
It is far more efficient than ever for bad actors using generative AI to be able to hack and gain access. Furthermore, there are new AI techniques that make it possible to unmask data that's being used in these AI projects. What do I mean by unmasking?
The most common way organizations are operationalizing their data is through data masking techniques. We're gonna talk more about that in a second. But because of these new threats that are emerging through the weaponization of ai, those tried and true legacy techniques of masking data are now obsolete.
So what that means is it's kind of like everybody built walled castles and now there's gunpowder. So those castle walls are entirely obsolete. What I mean by that is every organization and enterprise has created access control ring fencing around each of their data silos to ensure that privacy, security, and sovereignty of the data within that silo.
But then when they go to operationalize the data, they're using a variety of anonymization techniques or masking tokenization, et cetera, and then operationalizing it, which makes irrelevant and obsolete the access control and manual audit par audit procedures that they put in place to enforce those policies. Now, as a result, it's true that AI success is business success, but most of the projects are stalling in their transition to production. For example, in 2023, less than 10% of AI projects reached production.
I commonly see estimates floating around three and 4% of that $200 billion spend actually went into production and began to generate value, which is pretty shocking. Furthermore, of those projects that went into production, less than half actually met the expectations of the organization that deployed it. And the reason why is entirely related to data challenges, the same data challenges we just covered.
So if we look at how organizations have operationalized data, the only thing that they've had available to them to date is data masking techniques. Now, the problem with this is they're laborious in manual. They're complex, they're time consuming, they're costly, they're air prone, and furthermore, they actually hurt the quality of insights.
So when I say that they're laborious manual complex, like this is a a simple fact. In fact, companies that I've worked with will report in most cases that it takes weeks or months, commonly months to prep the data for use operationally. So partly it's because there's data preparation.
Partly it's that there's a human element of trying to bridge those silos policies to get agreement around are we actually adhering to each of their policies? All these data masking techniques are really inefficient and costly and hurt the quality of insight when you run your AI workloads because you've anonymized or hashed the data in a way that prevents you from getting the full value out of it. But to make matters worse, these legacy masking techniques create risk within the organization for the reason I already explained, which is there are new AI techniques, new advances, new advancements in AI that make it easier for ev, for, for than ever before for bad actors to actually invert these masking techniques and expose the data.
So not only are they costly but actually ineffective, let's do a deeper dive into it. So if you look at the common ones, you have everything from synthetic data to masking to tokenization, and we've analyzed it here across five dimensions, versatility, ease of use, security, compute, efficiency and insight quality. So versatility is like what kinds of jobs can you run on this data and what kind of insights can you get out of it?
Ease of use is pretty obvious. Security, same thing, compute efficiency. How efficient is it for you to run your AI workloads on this data?
And then insight quality. So as I've already established, the insight quality, the ability for you to build business insights from data that's been masked or synthetic data that's been created is severely damaged. It's much lower.
And then as I've established, the risk related to security is very, or sorry, the the security is very low, meaning the risk is very hot. Now with synthetic data, you'd be like, it's synthetic data, like what's the security risk? Actually, even with synthetic data, the trends are in the data and AI techniques can pull out those trends.
It is more secure than just masking or tokenizing your existing sensitive PII or proprietary information, but there's still risk associated with it. So these are the only techniques that have been available to date because the only technologies that we've been able to deploy for encrypting the data have been at rest and in transit. So that means the holy grail of providing the cure data processing lies in the ability to do encryption data in use.
This is the holy grail of being able to unblock a, a AI and unlock the value of the data is, well, gosh, let's just encrypt it. Let's encrypt the data all the way through, including during compute. And obviously most people would say, well, that's impossible.
It's actually not. In fact, there's multiple technologies for doing this, and they have their pros and cons. The first is the cure, multi-party computation, and then the other one's homomorphic encryption.
These are mathematical cryptographic approaches to processing data while encrypted. The challenge with these though is that they're really hard to use and you need some arcane wizardry to be able to even get it up and running. Furthermore, the capabilities of these technologies, the versatility, it's low because you don't have the full capabilities of being able to run Python machine learning frameworks that you're already using Spark, et cetera.
But the security is very high. You have encrypted data all the way through end-to-end, including in processing. Um, and your insight quality is much higher than if you're using a masking technique.
But given the cryptographic, um, technique used on the data, the compute efficiency is very low. So you could have a job in Python or Spark that may take milliseconds, that could literally take weeks, sometimes even a month or more with some of these technologies. So they're not very versatile, pure, and they're great in certain use cases.
Now, I've become aware of all of these techniques because I joined OPIC last year in 2023. Prior to that, I was at ServiceNow where I led a business unit. Most of my career has been predominantly in business software, enterprise software, but business application.
So I was attracted to Opaque because this team who founded Opaque happens to be a group of Berkeley researchers who also created Spark and founded Databricks. They also invented Ray, which then turned into any scale. And then more recently, they built an open source project called mc two, which became opaque.
Now, the research for Opaque began nine years ago at the uc, Berkeley Rise lab. Today, it's called the Sky Lab. They rebrand every five years or so.
So this group of researchers who worked on all of these foundational AI technologies worked in collaboration with Intel, who is pioneering a new kind of hardware, a new kind of hardware feature in the CPU. They called hardware enclaves. And what this allowed the researchers to do is to run compute on encrypted data.
Now, that was nine years ago that hardware was not commercially available. Fast forward, this is what we call confidential computing. It's commercially available on every hyperscaler, every CPU, and even the NVIDIA's new H one hundreds supports the capability to do confidential computing.
So let's take a look at confidential computing. This allows you to run secure AI on sensitive data at cloud scale. But when you do a tear down of the five dimensions that we're analyzing these technologies on, here's what you'll find.
The versatility is, uh, pretty good. Um, but the ease of use is pretty low. It's pretty difficult to get an environment set up where you can ensure the applications and workloads that you're building on confidential computing are running at cloud scale.
Um, and it's auto scaling. There's a lot of complexity associated with setting up confidential computing environments. However, when you have it set up, it is incredibly secure.
The compute efficiency is incredibly high. In fact, the performance hit you receive from running AI workloads or just any compute workload on encrypted data using confidential computing is it's negligible. You get a single digit percentage performance hit, so it's not even noticeable in most scenarios.
And then most of all, the insight quality is very high, high. So when I say the versatility is medium, why is that? Because the kinds of workloads and frameworks, et cetera, that you can run on confidential computing at cloud scale depends on your ability to set up the environment so that it's auto-scaling, and how do you set up policies?
How do you set up workflows? How do you set up collaboration across multiple people's data sets? Um, and then obviously the ease of use and governance, self-explanatory.
So the same people who helped pioneer the space of confidential computing, which I strongly encourage you to investigate, there's actually a summit called the Confidential Computing Summit, where anybody who's interested in securing AI and confidential data or sensitive data, they attend the summit. Uh, in fact, it's June 5th and sixth of this year. Um, and you have the CTO of Azure, the Chief Security officer of nvidia, the CEO of Accenture.
It's kind of a who's who of anybody who cares about secure AI sensitive data. Um, anyway, the opaque founders helped pioneer this space and built a software stack on top of confidential computing. That's really a business application.
It's a business solution for anybody to take their existing AI workloads, whether that's analytics using Python, spark or machine learning frameworks, and deploy this on encrypted data at cloud scale. But I'm gonna switch away from opaque and talk more generally about confidential computing. So what was introduced with confidential computing is the ability to run data in use at scale.
Now, if you look at this applied to the hottest topic in technology right now, large language models and generative ai, here's where it can be applied. It's across every aspect of generative ai. One, you can deploy confidential computing to protect the queries.
So you're processing encrypted data like the context and providing data governance rules at the prompt. Level. Two, you can secure data in the model, the model's just data really, right?
So you can secure the model. And we see this often at opaque where let's say it's the New York Stock Exchange or ServiceNow, they want to train the model or they wanna secure the model. They can run that in opaques environment or build their own confidential compute to protect it, um, both the data and the model.
So let's do a deep dive into what this would look like in a generative AI implementation. But I wanna make clear this isn't just about generative ai. I'm just doing a deep dive on generative AI for the sake of it's the hot topic right now.
So let's take a look at, this is the opaque gateway, but you could easily, well, there's some complexity to it, but you could set up your own confidential compute environment for this. As I'm showing you here. So let's say in this scenario, I'm a tax advisor supporting Tracy Garcia, who has a question seeking tax advice.
I write a prompt. That prompt goes through rag. When I say rag, that means that what I wrote is enriched from multiple enterprise data sources with additional context relevant to the prompt I wrote.
So for example, it could include Tracy's address because that has tax implications. It would include the investment playbook that we're running with Tracy. It would include her existing accounts, what has she purchased historically, because that would give an indication of her behaviors and the propensity to buy and other personal and proprietary information.
Now, this data then, if you're using a confidential compute environment or opaque gets encrypted, that data while encrypted then is able to be processed. So what kind of processing might you do? Well, what we commonly see, like for example, with a high performance auto manufacturer from Germany, uh, we help them deploy specific data governance rules.
For example, starting with the basic stuff you can redact or tokenize PII data, um, you can take dictionaries of data like supplier names, parks, et cetera, and redact or tokenize those as well. Now, moving into the more sophisticated in most of these generative AI implementations in the enterprise, they're calling multiple data sources and they can't apply user group policies on those data sources without re-architecting their entire data landscape. So in this scenario, what we commonly will see is metadata's passed to the gateway about who is the user and what group do they belong to.
Oh, it's Aaron Fulkersom. Uh, he works at the dealership in San Diego. Well, I don't want him to see any of the dealership from Austin Texas's data because that's competitive.
So data gets redacted, or it's Aaron Fulkersom the customer, or Aaron Fulkersom the supplier. We can apply these data governance rules before sending it to the LLM. Now, in this scenario, I'm gonna show the data, get decrypted before sending to the LLM, but I wanna make a key point here.
You can run the large language model in a confidential compute environment. You can even run GPUs in a confidential compute environment because the Nvidia H 100 support the capabilities, as I've already mentioned. So in this scenario, I'm showing it being decrypted.
It gets redacted. Tokenized data governance policies applied. Oh, one more thing that that we often do here, prompt compression.
So you can run an LLM as I've just established inside that gateway, and then do prompt compression to improve the performance of your LLM in terms of cost and speed. So in this scenario, as I've said, we're showing it decrypted because oftentimes this is going to chat GPT and people can't host their own chat GPT, but if you can host your own LLM, you could actually run this encrypted end to end, including at the inference time. So in this scenario, you can see that the LM comes back with a response.
The response gets encrypted again, back inside the opaque gateway, and then it gets de anonymized and post-processing occurs. Now, some post post-processing that you might consider is, we've seen scenarios where the, when the l LMS been fine tuned or trained with data, sometimes you need to clean up the training data. And I'll give you a infamous example that I've seen in the real world.
Large software company, fine tune some Gemini LMS or some Gemini models with customer data. They roll it out to an early access program. Then within a week of the early access program, customer one is able to see customer two's data coming out of the LLM, the the fireworks were completely unintentional.
That was not intended. If anything, it should be sad trombone. So in this scenario, uh, it's easy to solve because you can run the training data inside the gateway and clean up any of the training data again before sending it back to the customer.
And one key point that I should highlight here, this report, the way that we run this at Opaque is every operation on data that flows through our platform, every operation on data, every policy gets logged so that you have a complete audit trail digitally signed by the C-P-R-G-P-U manufacturer proving that your data was kept private, secure, and your sovereignty was retained. This is possible because of confidential computing, because deep down within the dies of the CP or GPUs, the data is being decrypted where nobody can see it and processed, where nobody can see it, not opaque, not the Hyperscaler Azure or uh, AWS or GCP. Nobody can see the data and you can prove it because of confidential computing with a cryptographic signature provided by the CPU or GPU.
So that's the basic implementation of confidential computing in a generative AI environment. But I want you to be aware that it's, this is not limited to, uh, large language models. Confidential compute can be applied to any AI workload, um, or any compute workload.
In the case of opaque, what we provide is a confidential workspace for data and MLOps. So we have a whole end-to-end user experience with notebook and no code. But in some of our large enterprises customers, they'll just use us as the compute plane inside their existing environment.
Domino Data Labs or Databricks or whatever they're using today, they'll continue to use. But Opaque provides the compute plane to that confidential workspace. Another application of this is, uh, confidential training and inferencing.
So here's a perfect scenario. Well, lemme give you a scenario on confidential workspace. Uh, where we see this commonly is inside the enterprise where they're taking data from multiple silos.
Let's say it's Workday HR data data from Snowflake or some other, or, or Databricks, some, some data warehouse or like house, right? And then, uh, combining it with CRM data to run reports related to their salaries or employee performance. Why would they use confidential computing or opaque in the scenario?
Because it ensures that you can process the data data without seeing the underlying data. So for example, maybe you don't wanna see gender, race, age. Well, you can do that.
Um, or for example, maybe you wanna look at trends without exposing highly sensitive SEC data, you can do that. So that's the confidential workspace. Um, oh, I'll give you another example.
The European Union uses opaque, so they use it across their EU member countries where they can share data without showing each other citizens' data to identify cyber threat, emerging cyber threats like ransomware attacks, et cetera, that operate across country lines. Okay? Confidential training and inferencing.
So I, I showed you inferencing with LLMs, but you can think about this more broadly around classic machine learning, which is in, uh, some cases, let's say New York Stock Exchange a customer. What they wanna do is they want to train their machine learning models from partner data like their market makers or customer data like we do with some enterprise SA vendors. Customers want the updated trained ML model with their data, but they don't wanna share their data because it's private.
They wanna keep it secure, they wanna retain sovereignty, they don't wanna give it away. So you can dynamically train preexisting models autonomously from encrypted dataset. And then lastly, the opaque gateway around large language models that we've already covered.
And all of these scenarios, think about this as a confidential data pipeline and compute plane with a hardware root of trust. And I wanna make available to everybody here, the securing generative AI in the enterprise White paper. It covers every one of the techniques known today and how they, these techniques can be used for, for securing AI and their pros and cons and trade-offs.
It goes, this is written by two PhD, won a full-time Pro Professor re Luka at Apapa who world renowned Grace Hopper Award winner, uh, professor at uc, Berkeley, and co-director in the Sky Lab at uc, Berkeley. And then the other one is Dr. Rashad Poddar, who, uh, was her, did his doctoral, uh, um, did his doctorate under Euca and Eon Stoica, who's also a founder of Opaque Eon Stoica is the creator of Spark and Ray, and, um, another founder of Opaque.
Anyway, great paper. Um, broadly applicable. I'm sure it'll be really valuable to all of you.
And I just wanna thank you for, uh, your time tuning in. There's, there's a, the business track is packed for the rest of the day. There's a bunch of really fantastic talks that are being given.
I'm excited for all of you to tune in and see the other speakers because there's just such a wealth of experience that has been assembled here by the artificially intelligent enterprise and Tech strong. So thank you so much for your attention. And, uh, feel free to reach out directly to us at Opaque.
Um, we'll make available the link, uh, right here. Okay? Thank you so much.