Accelerating the AI Transformation in our Organizations with Marco Palladino at AIE 2024
AI has taken the world by storm, generating a paradigm shift in how we build applications and make them intelligent. Our teams are thinking of ways to integrate GenAI inside of their applications to create better end-user experiences and gain competitive advantage in the market. But as they do so, it is essential for the organization to develop an AI playbook that will both improve developer productivity while at the same time enforcing governance, compliance, security and end-to-end observability as our applications are building new AI-driven use cases.
Transcript
Hello everyone. My name is Marco Paladino, and I am the CTO and Co-founder of cog. Today we're going to be talking about how we can accelerate the AI adoption and AI opportunity in our organizations by providing our developers with the right platform and the right infrastructure that is going to be making them more productive at scale, while at the same time still having control, security, and governance on the AI usage that our teams are performing.
First and foremost, uh, we really have to look at the adoption of AI engine AI specifically as one of the latest trends that's generating more traffic and more consumption of APIs inside of our organizations. API usage grows when our developers build digital applications to cater to new digital use cases. And so of course, mobile microservices were some very big trends that increased the adoption of APIs, but also gene AI increases the adoption of APIs.
With gene ai, uh, we can, uh, consume ai. We can train AI and AI can interact with our own systems. And for each one of these different, uh, actions, there is going to be an API that enables this operation.
So the more gen AI adoption and the more API consumption there is going to be in our organizations as our developers are building AI applications, they may also consider using, uh, cloud or self-hosted models, uh, that are available to them, uh, from the industry. And something that's quite interesting to see is how quickly open source and, um, open models really were able to catch up when it comes to intelligence to the, uh, initial, uh, clo, uh, closed source and proprietary models that we're all very familiar with. This is now opening a whole world of opportunities for organizations to now determine if they want to choose cloud dusted models or private models in their consumption of gene ai.
And as we're going to be seeing later in this presentation, sometimes it is a mix of both. We may still want to use a self TED model to optimize latency and even cost while at the same time leveraging a cloud model or a cloud provider, uh, for, uh, failovers in case our self TED model doesn't go so well. So there is lots of different orchestration that's currently being built, but it's good for the industry that we have the optionality to choose which provider and which model bar suits our use cases.
So there is lots of activity that's happening right now, and it's very exciting as a space. Chances are that today in your organization, you have a small subset of developers that are very progressive, and they are, um, spearheading gen AI innovation in your company. They are, uh, a small group of progressive developers that are building, um, AI applications either by creating vanilla integrations with ai.
They're building AI agents using frameworks like Long Chain or LAMA Index. Uh, they are leveraging, um, AI driven chat bots or AI driven productivity applications like a GitHub copilot. And the point is that today, not only they are building AI applications that drive an outcome and a use case for the end users and the end customers, which may be internal or external, depending if we're building, uh, uh, something that's, uh, hooked into our public products or not, but they're also building loss of, of, uh, crosscutting requirements that each one of these different applications needs to have.
Today. When we look at the journey of AI adoption, it is very similar to the early days of API adoption. Back in the days, an API developer was building APIs and also was building lots of cross-cutting requirements like authentication, uh, rate limiting traffic control analytics that every API at the end of the day needs to have.
Their job was harder because in order to build an API, they also had to build the infrastructure for that API usage. And for ai, we are in the early days of ai, and the same thing is happening right now. AI developers that are building these new, um, uh, AI applications are also building AI infrastructure to be successful in their production lifecycle for ai.
So there is lots of crosscutting requirements that every AI application or integration needs to have that today. They need to build The key to unlock AI innovation in your organization. It is to make sure that these underlying infrastructure tops being something that they build and starts becoming something that they use.
We don't want them to build the underlying infra to run AI or APIs. We want them to focus on the end use case and everything else is being given to them by the underlying infrastructure that powers these use cases. And so whether they're using cloud-based lms, whether they're using self-hosted l lms, whether they are, uh, generating embeddings and then storing them in a vector database, there is going to be a series of crosscutting requirements that we can remove from their table and we can put in the underlying infra and make it available to them by default out of the box.
This will make the developers more productive when building AI applications. It'll also enable the AI opportunity to a larger population of developers in your organization. Likewise, back in the days every developer became an API developer because today the internet is APIs.
In the future, every developer will be an AI developer as well. And the key to triggers this change, it is to provide the infrastructure that allows AI adoption to scale across every developer in the company. And to do that, we need to provide an infra, and this infra can be pro, uh, given to them in the form of an AI gateway, an AI gateway that essentially covers the security, uh, the governance, uh, the, uh, uh, semantic capabilities of the underlying AI consumption and makes it available out of the box to every developer.
Really, the goal here is to turn ai, uh, in the AI infrastructure as a platform, core service that can support the whole organization. By doing so, not only we're making the developers that are building the AI applications more productive, but also the company now is in full control of establishing an AI playbook that ensures security and governance of all of these AI consumption. Of course, we're all very well aware that there is a huge risk also involved in the usage of ai.
And the risk is the, the risk of leaking customer data or sensitive credentials in the models that can be then extracted somewhere else by enforcing an AI playbook, using AI infrastructure that becomes part of the core product of our, the core platform of our company. We can now, uh, establish, uh, procedures, processes and controls in place to make sure that that doesn't happen. So not only we're making the developers productive, but quite frankly, we're getting peace of mind that the AI usage that's been generated across all the applications, it is proper and it is, uh, responsible usage of ai.
Also, by the way, by knowing what is all the AI adoption that we're generating, we can also, uh, monitor that adoption and perhaps even find opportunities to, uh, reduce costs for, uh, all of these new gene AI use cases. And at scale gene AI is very expensive to run. So this becomes a problem real quick that we need to address.
Of course, Kong, and Im the city of Kong, you know, we work with enterprise organizations that are today, uh, capturing the AI opportunity, and they understood the need of AI infrastructure. And so at con one of our offerings, it is an AI gateway that can run pretty much anywhere in the cloud on Kubernetes. Uh, it's very portable and provides L seven AI capabilities for a series of LLMs as well as vector databases that can be used right away in the organization to accelerate the AI adoption from a developer standpoint.
But when we look from a holistic standpoint, you know, uh, how does AI gateway fit within the broader context of our connectivity truly becomes an additional use case on top of API management on top of ingers controllers that we, we may be using service meshes that we may be using to create a network overlay across all the microservices. And AI gateway happens to be an extension of our connectivity, um, uh, platform that pretty much every enterprise organization has today, um, to enable the developers to now also take advantage of this new gen ai, uh, technology prime. Now, of course, again, at Kong, we provide the full series of end-to-end, uh, uh, products for connectivity, and we also provide a unified control plane that allows to manage all of this from one place.
But today, uh, I'd like to focus on the use cases of AI and, and how an AI gateway can simplify, um, the adoption of ai, therefore making it more successful. When we look at, uh, what an AI gateway does, truly, we're looking at a specialized component that has the, uh, capability of fully understanding the underlying AI traffic that we are generating. So this is not just, uh, adding, uh, an LLM or the API that gives access to an LLM on top of an API management platform.
Uh, this is really about having deep introspection on what that AI consumption is to provide more interesting capabilities on top of that. So first and foremost, an AI gateway, uh, including Kong's AI gateway. One of the things that they would do is to provide one API that allows us to consume, uh, one or more LLM.
These improves developer productivity significantly. We may be using different LLM technologies, whether they're cloud or self-hosted, uh, because they fundamentally have different foundation models that we can use to fine tune our own models, and they all come with our pros and cons. We may use different LMS for different use cases.
The world is multi LLM. Uh, we may also use multi LLM adoption to, uh, leverage, let's say a cloud LLM that's more expensive. Well, at the same time fine tuning a self-hosted model over time so that over time we can rely more and more on the self-hosted model versus the cloud models.
So there is lots of orchestration that may be happening between one LLM and another, and it is very hard for developers to figure this out on their own if they have to integrate specifically for each LLM instead. It is much better if the infrastructure can provide one endpoint that allows to consume every LLM the same way, way, therefore, a developer may be building an API that let's say uses, uh, open ai, and at the flip of a switch, they could be using that same code base that they're writing the same integration to consume Misra, for example. This allows for much better productivity and much, uh, faster time to market for the gen AI use cases.
Then on top of these AI traffic, we need to centralize how we secure communication to the LLMs. And so in the AI gateway itself, we can store the credentials and rotate those credentials to the LLMs that we're using without having to share them with the, uh, developers that are making the AI consumption. This is useful because to the developers in our teams, we can give them a proxy credential, let's say a job token that at one time it will be replaced with the actual LLM credential to consume the gen AI capability.
If that credential for the LLM needs to be rotated, we can do it, uh, on the gateway itself without having to tell the teams to update their credential because it's a different one. At the same time, we can implement more advanced load balancing capabilities to load balance traffic between one and element to another based on availability or even based on cost. And, um, we can also implement failover in such a way that we could be running all traffic on our own self-hosted model.
And if the latency increases above a certain threshold, then we could be leveraging a different provider in the cloud, for example, to keep the uptime of the service always up and running for the end user and doing all of that at one time. On the gateway itself, of course, comes out of the box as part of the infrastructure of AI that we're deploying. And then finally, the observability of AI being able to capture, um, the requests that we're making, but also, quite frankly, L seven AI information like the providers and the models that we're using, how many input tokens, output tokens are being consumed.
This gives us an opportunity to fully understand how the adoption of AI is happening across the board, and it also gives us an opportunity to improve the cost and the span of AI by seeing how much we're spending across all these different models. 1 A or what an AI gateway should do. More interesting.
AI gateway and infrastructure technologies are going to be actually doing deep introspection on the, uh, AI usage that we're doing, and they're going to be providing more advanced capabilities. For example, capabilities like the ability to firewall the prompts by essentially determining what are the, uh, operations that are allowed and not allowed by the developers. These allows us to, for example, ensure compliance, ensure that there is no restricted topic that's being, uh, discussed using the prompts.
It allows us to have some level of control. We can then decorate the prompts with context that we want to standardize across every, every request. Uh, whenever, uh, a developer integrates with ai, part of the job is to make sure that the, uh, user input can't, for example, jailbreak, um, our rules.
Um, it's called prompt jailbreaking. Uh, it allows us to make sure that we cannot, uh, accept profanity as an input or an output in our models. It allows us to determine what is the role of engagement of ai.
These rules will change over time and every time they do change. We don't want our developers to first build all of this, and we don't want them to then update over time all of this. Instead, we can push all of these enforcement inside of the AI infrastructure, and the core team can then apply these roles for every team or select teams and select applications without having to change the implementations that developers are building.
And then the ability to also ensure a playbook for, um, prompt lifecycle. You see the prompts. It is what we are using to access gene AI and prompts need to have a lifecycle in the way they're being introduced, in the way they're being versioned, in the way they're being decommissioned, in what are the prompts that we want to allow or disallow in the organization.
By providing a templatized, prompt, uh, management system, we can enforce these AI playbook inside of the AI gateway in such a way that we, from an organizational standpoint, from an an architectural standpoint, uh, we always know what is the usage of AI that's being done. But at the same time, um, the developers can think in advance about what are the prompts that they want to ask, follow a process that the organization will have to implement to then load those prompts inside of AI gateway and then start consuming these AI gateway. This is, again, like the early days of APIs.
Back in the days with APIs, everybody was creating APIs ad hoc for specific use cases. And over time, that creates the part of a, of a mess of APIs in the organization, lots of duplication and so on. We don't have to repeat the same mistakes again with ai.
With ai, we know already that that is going to happen. And by leveraging AI gateway and AI infrastructure, we can ensure a life cycle earlier on in our adoption journey so that we're not gonna end up with a big mess down the road in a few years from now. And then finally, the ability to be able to, uh, interject, um, the, to inject our AI in a existing either API traffic or AI traffic on the request or response lifecycle.
Like I, I usually talk about these use cases as no code use cases because the AI gateway itself becomes the client for ai. For example, we may have a client that's either submitting existing API traffic today or new AI traffic to an API or an LLM. And before that request reaches the upstream LLM or API, we may want to do, uh, AI driven, uh, uh, uh, implement AI driven functionalities on top of that lifecycle.
For example, we may, we may want to take the request to train a model. We may want to enrich that request before it is being sent to the upstream API. We may want to, uh, convert the actual format or the language of that, uh, content in, uh, you know, from one thing to another, and we could be using an LLM to do that on the request lifecycle.
And then likewise, we can also do the same on a response lifecycle. The AI gateway can inject, can intercept the response, and before it is being sent to the client, uh, it can then, um, reduce the hallucinations. It can implement rag, it can implement sensitization in such a way that, for example, if we detect any personal information or if we detect something that should not be shared, it'll strip it out before sending it to the original client.
AI infrastructure fundamentally allows us to inject AI behaviors inside the organization, even without having the developers building it for us. These use cases I'm talking about right now, these ones that don't even require a developer to build an AI integration, we can inter inject AI today on the existing API traffic that we have today without even having to build anything simply by using AI gateway and then configuring these capabilities more interestingly, there is a whole series of capabilities we can enable for semantically, uh, manage, uh, AI requests, for example, semantec caching. It is a use case that allows us to semantically understand what are the prompts that we are being receiving from the applications.
And if we find diff two different prompts that have the same semantic meaning, then we can return the same response. By doing so, we can improve the latency, uh, of AI, because now we're returning a cash response. And we can also reduce cost of consuming AI because the cash response will never generate an actual LLM request.
The prompts are never going to be always the same, but they can be semantically the same. It means they can have the same meaning. For example, a very down to earth example.
If we ask LLM, how long does it take to cook pasta, the LLM will respond 10 minutes. If we ask LLM in a different request, how long does it take to cook spaghetti semantically? I'm asking the same question because spaghetti is faster.
Therefore, we could be returning a cash response instead of generating another LLM query, then we can route semantically to different models. The value of AI is being driven in the way we fine tune our models and, um, we are going to be fine tuning models for different use cases. For example, let's say that we're building a chat bot for customer support, and also we're building a chat bot for new sales conversations.
The models may be different because they're fine tuned on different data sets when the developer is building, uh, an application that needs to consume the models that we have, and over time, we're going to be having hundreds of models, so it becomes a big routing problem. We can semantically route the, the request to the prompt that has been fine tuned for that specific purpose. So the AI gateway can find out that the incoming request is a sales question, and with it without the developer knowing which model in the organization is very suited for answering a sales question, the AI gateway knows that, and therefore it can automatically route the request to the right model.
Over time, we can replace the models and improve the lifecycle of our models without having to tell the developers what to change in their applications because it's all being done by the underlying infrastructure. Without this, the developers will have to constantly update their applications to keep using new versions of the models, and they will have to release new versions of their applications, whether they're mobile or web. Every time something like this changes, it's lots of work that will drastically reduce their productivity.
Semantic routing fixes that. And likewise for semantic firewalling in, in a, a traditional API request or in a traditional web request, we may want to block requests based on a series of predefined rules. Uh, and if the request matches those rules, uh, that are very predefined, then we may want to block that request.
Think of SQL injection in ai. We may want to firewall the request based on the semantic meaning of what is being asked. By doing so, uh, we can now essentially, uh, ensure that the behavior of our AI usage in the organizations is always consistent, but it's being done on the actual meaning of the request.
Therefore, we don't have to think in advance about all the different ways that a user may want to, um, you know, jail break our prompt. We could be having a more generic rule that determines that if there is an attempt to jailbreaking whatever that attempt is to block that request because it may be not proper. Semantic caching, semantic routing, semantic firewalling, these are all part of great AI infrastructure that can simplify enormously the productivity of the developers and the governance and compliance of the organization.
For all the semantic capabilities, there is usually going to be a, um, a few steps, um, that pretty much involve creating embeddings for the prompts that we're receiving and then storing the embeddings in our vector database. And then at one time, uh, being able to, with high performance, ideally non-blocking performance, being able to then perform all of these functions. So to wrap it up, it's important that we think about how we are turning every, developing the organization into an AI developer, because that is going to be unlocking a whole new generation of applications that our organization could create.
And for that we need modern AI infrastructure. It'll improve the developer productivity, it'll make, uh, the organization safer. Uh, it'll enforce compliance and governance on the AI usage.
It'll also allow us to observe and potentially optimize costs of AI that we're performing across the board. But this is only the beginning of the AI opportunity in the world. When we think of, uh, API platforms in general, we are today thinking of fundamentally, um, you know, uh, creating APIs that developers can use, uh, to harness the data and the functionality in our applications.
But we do have an opportunity in the the coming years to build almost like a new layer in the network stack. And that is an intelligence layer that allows developers to not build against the individual APIs and instead build against that intelligence layer first and foremost. And that layer will be in charge of orchestrating requests across all the underlying APIs.
To fully understand this, we have to look at the evolution of API platforms. When we look at modern API platforms, you know, we built APIs so we can extract data and services from our silos, APIs break down the silos in turn our company into a platform, our products into a platform. Now, developers inside and outside of the org, they can build on top of our API platform, um, and and harness our data and our services.
In the current generation of API platforms, we are fundamentally not changing the underlying architecture of how this works, but we are enriching our APIs with LMS and Gen AI to make some of these operations smarter over time. Uh, and that's the middle picture you're seeing here. So the API platforms are being enriched with AI usage.
This is where we are right now in the industry. And the evolution of this. It is to remove the concept of APIs altogether from the developers that are implementing against our platform, and instead provide a large action model that can orchestrate the intended, uh, operation that we want to build to the underlying APIs that developers are creating.
Let's say that I am building, um, an application that wants to, uh, order, uh, the most popular, uh, food among the most popular restaurants in my area using, let's say Uber Eats. How many API calls would that involve? First I have to find where I am, then I have to find what the most popular restaurants are.
Then I have to find what's the most popular food, uh, is for each one of these different restaurants. Then once I found that out, I have to go ahead and order that, uh, make that order for that specific food in that specific restaurant using the Uber Eats APIs. There's lots of API orchestration, but with a large action model, I can ask the large action model, find the top food among the most popular restaurants in the area, and order it already with Uber Eats.
And the large action model will be in charge of making the orchestration requests to the underlying APIs. So the developers in our org, they still build the APIs. It is the developers that are using our APIs that are not using our APIs directly anymore, but they're using, um, our intelligence, organizational intelligence to make sure that these, um, these, these happens at Kong.
We're actually working towards building the future of API platforms. And this, it is also an area, uh, that we're investing in. You can see the, uh, clam name that is the con large action model.
So I'm very bullish about what, um, it's happening in Indu industry right now when it comes to ai. I am very bullish, um, into, uh, the role of AI infrastructure in the future of API platforms when it comes to AI that we're going to be building in our organizations. Ultimately, I think that, um, everything that you're seeing here, uh, quite frankly, it's going to be inevitable.
Uh, and, uh, I am personally, uh, excited to be working with you to, uh, unlock this opportunity in your respective businesses. So thank you so much. I hope you enjoyed this presentation.
com. Uh, and if you have any questions, uh, you can find me on, on social media subnet, mark on Twitter, uh, I'll be happy to, I'll be happy to have a conversation with you. Thank you so much.