Discover AI Innovation at Techstrong with LangChain 101 | AI in Action 2023
Dive into the Future of AI With the Austin LangChain Users Group at Techstrong!
Join us for an exciting lab session at Techstrong Group, presented by the Austin LangChain Users Group. We believe that collaborative learning “in the open” is key to integrating AI into our businesses, communities and everyday lives.
What We’re Bringing to Techstrong:
Empowering Insights: Gain a comprehensive understanding of LangChain — an open source project that’s reshaping the AI landscape.
Practical Tools: Learn how to use LangChain with Google Colab and OpenAI’s API, making AI more accessible and hands-on.
Create and Innovate: Build your own LangChain application and showcase it through Streamlit, crafting a simple yet effective web interface.
Cost-Effective Learning: Manageable expenses with a paid OpenAI API account, tailored for budget-friendly exploration.
Expand Your Horizons: Access a variety of learning labs, offering diverse experiences to enhance your AI skill set.
This session is perfect for anyone eager to explore AI, from beginners to seasoned enthusiasts. We’re committed to creating an engaging, open learning environment, fostering innovation and community growth through AI.
Join us at Techstrong and be part of a movement that’s bringing AI closer to everyone. Let’s learn, innovate and grow together with LangChain!
Transcript
So, uh, my name's Colin McNamara. I want to invite you to joining Austin Link Chain user groups. A couple of members here, uh, for our link Chain 1 0 1, uh, presentation and labs.
So a little bit about Austin Link Chain. Hey, hey Tip. Hello from Florida.
Uh, little bit about, uh, Austin Link Chain. So we're a users group. We're enthusiasts from, uh, Austin, Texas, the capital of Texas.
Um, you can find us, um, on Discord, on our GitHub. Um, there's a pinned, uh, some pin links, as well as a lot of QR codes that you can find us. Uh, we meet once a month in person as well as we meet virtually.
Uh, so people that aren't, uh, aren't able to meet, make the in-person meetups, they can go catch up. Um, we are early adopters inside of Lane Chain. This is really an exercise in learning in the open.
So anyone who's been doing any development within, uh, with AI may see that, uh, things are moving really, really fast. And so what we are doing is coming together to teach each other, uh, things as we've learning. And so we're sharing this.
We have the opportunity to share this, uh, with y'all here. Um, joining me, uh, is Kareem Ani and, uh, Ricky. Oh, God, I'm gonna murder your last name.
Ricky. Ricky Perio. Did I get, did I, did I nail it?
Yeah, let's, right. Sweet. I'm actually switch to the camera right here so everyone can see us.
So, joining me is, is Kareem Ani and Ricky Perio. Um, Kareem, um, we'll go into Zach later. Uh, Kareem works in the, the office of the Governor, uh, doing data science, uh, for the state of Texas.
And Ricky is a, uh, manufacturing engineer doing, uh, basically full stack development, uh, for, uh, chips head supply chain. So this is just a smattering, uh, sampling of folks we have in Austin, lane Chain user group at the end of the day. There's a lot of manufacturing, there's a lot of industry, there's a lot of development here, and it's really interesting to see us all to come together, whether that's from the Artis program, from Nvidia, um, from Always Cool where I work at, and so many others.
So, let's get into it, and we'd love to see you connect with us. Where is the next slide, bud? Okay.
So a little bit about me. Um, I live in Austin, Z side. Um, the cool side, if you ask anyone, uh, the, I I am, uh, the managing partner for engineering and finance at Always Cool Brands.
Um, we are a, uh, product design consultancy and supply chain integration company. Uh, my background is in hyperscale cloud engineering operations. Uh, had my hands in, uh, many global platforms that, um, I would say all of us on the call have, have used or do use some of these things I've had a chance to work on.
Uh, my open source experience, uh, is started 25 years ago with Linux, um, made its way to OpenStack open daylight, a bunch of really cool projects. Um, I'm using Lang Chain for business operations. com.
That's a a, a great start to everything. Um, you can find me on LinkedIn under the same name and on x occasionally. So I'll ask you, um, how safe is your next flight?
And we all get on planes. You know, I, I have, I fly so much on Southwest. I got a companion pass.
Um, in this case, uh, we, there's a distributor, um, out of London called A OG NICs. Uh, what they did to make money, and this is this, this is actually a classic scam in the supply chain. Uh, this distributor went ahead and, uh, fake certificates for bogus engine parts on 126 engines.
These made their way into airplanes all across the fleet, and this was in the Airbus fleet. Um, what happened with this specific fraud, and this is probably about, uh, three months ago when this happened, it exposed weakness in, in the, uh, in the airspace parts market, right? So what happened?
The airlines had to ground, I think, 90 air buses and take their engines off, take take out all their parts and go through a complete Indian supply chain audit. Now, I've had to do this work in a hyperscaler. It's horrible, right?
Uh, but the key thing here is distributors, manufacturers, brokers, there's so many good ones out there. Um, but there are so many bad ones too. And they really try to play with the provenance or the history of what you get, get into your product, whether it's a jet engine or it's a food that you eat.
So at, at always cool brands, you know, we design products from, for retailers, for grocery, for club stores, and for influencer agencies. Uh, we do formulations. We design packaging, uh, we arrange manufacturing as well as, uh, bulk material sourcing.
So the entire chain, when something comes out of the ground, when something ends up in the customer's hands, we handle that and we build those. Uh, primarily we do it to, for people's own brands, helping them build their brands to give their customers. We operate under a lot of regulatory controls.
Um, whether that's ISO frameworks, whether I-S-A-I-W 42 for sustainable procurement, or 9,000 for manufacturing, um, safe quality foods. So a lot of our work is for ingest ing. Inject ingestibles is something you drink or eat, uh, needless to say, when you're doing that, and you have to maintain things to the highest standard.
We have to maintain Hass ups. We have to follow the, uh, cap, the FDA Kappa program, a corrective and preventive action program, which is anytime you're dealing with food or drugs, uh, or pharmaceuticals. And then, uh, at the times when we're helping helping brands, uh, sell part of those brands to scale manufacturing, we actually fall under SEC regulations.
So, for us, Lang chain has been a godsend. Um, our first, our first integration interaction with Lang Chain was, uh, in, uh, putting together a rag solution for due diligence, right? So as we are building up these supplier, uh, supplier and manufacturing relationships, um, throughout the world, which result in products on shelves, uh, there are a lot of fast, uh, a lot of fast talkers out there, a lot of companies like a OG Technics that are trying to pass, uh, something bad or something good.
Um, we went ahead and implemented a, a rag solution where we took in all the documentation, all the emails that we have with our suppliers, and then we created transcripts out of our interviews and had to build a simple document like every business, it runs on documents at some point. And in our code, we told it to very be very specific and say, I don't know if I don't know. And this is the simplest use case for, for, for a lane chain application.
What it showed us was holes in the story, right? What a a of, of people that are trying to scam and very specifically allowed us to continue building intellectual property to be effectively build a, a fancy form that showed when there's holes in our manufacturer stories, when there's holes in the products, and allows us to go ahead and, um, protect ourselves from the issues. And effectively, effectively, the SEC.
It'll throw you in a, create a whole bunch of, of legal problems if you don't do the due diligence like we do. Uh, beyond that, so beyond protecting us from risk, uh, these rag solutions are really interesting. As, as we are con putting all the information from our supply chains, the point of manufacturer, um, our distribution networks, the methods of distributions, we capture a lot of information that is really hard to get if we don't do that specifically around sustainability.
And Scope three. So these rag solutions using lane chain allow us to, uh, much more easily report on, um, on sustainability metrics that are required to sell products within the eu, um, and actually in America in, in the next probably 18 months. So we love Lang Chain.
Uh, what is Rag, uh, very simply, uh, it allows you to retrieve a document, right? In this case, uh, throw in the, in the case I have up here, retrieve some PDFs, right? So first thing we did was we just printed everything PDF f threw it in the Google drive, sucked it in, right, chunk those documents up into little, into little bits using some code.
Um, and, and then throw that in what's called a Vector Store. So if any of you use Google Drive, um, if you've used NFS file share, if you've, um, uh, if you've used any sort, sort of, um, place, you can stick your documents. What a Vector Store fundamentally is, is just like any other document store.
But what it does is it translates these files coming in into effectively like, like a map through space and time and rep that in some ways represents a neural network that we have in our brain. So, as we dumped all of our data into a vector store, and then wrote little agents that went ahead and queried that vector store with simple questions, each simple question was a question about our supply chain, right? Each simple question was a question about distribution.
Each simple question led to a verification of statements by statements by the manufacturer or distributor. And the, the followup documentation that came, you know, days, weeks, months later, is a really powerful tool. So this query goes into this vector store, retrieves the relevant splits that we threw into it in that first stage, consolidates that altogether, throws it through the large language model, spits out an answer, and in our case creates the whole document.
Um, when this document, when that, when that code base and that document is checked into Git, we have, what we have is a control process. We can define what data we put in, what decisions we used to, to what decisions were used to populate the document, who made the decision, and when they made the decision. So this is a simple use case that we used OpenAI to add a lot of value in our business.
Um, and I'm gonna tell you more about the project. So let's move forward. So what is Link Chain?
It's an open source library, uh, for building large language model or LM applications. It is extraordinarily fast moving. Uh, there's a lot of things that, uh, can confuse people in link chain.
The documentation, uh, can be very unstructured at times. Uh, it's currently in transition between effectively a stable and experimental. So many people, so many platforms, uh, have been, uh, contributing their integrations to the project.
It's really, really exciting. dev, or I would highly, highly, um, highly suggest following, uh, Harrison Chase on X who founded the project, um, you'll see a lot of smart people doing a lot of really cool hackathons, doing a lot of great things in the community, and, uh, pushing the state of the art forward within the ai, but more specifically, um, within the Python and TypeScript packages, taking the code that you can write natively to OpenAI's, uh, API or Anthropics API or a llama locally, and delivering that to you as a template, effectively in a simple, reusable piece of code. So you, if you're like me and you know, you're a business owner, has some engineering background, not a great programmer, that you can take this take, you can use this project and add value to yourself, your company, whatever.
Let's talk about, uh, Lang chain. Uh, the key concepts we want you to understand. So, first things first, there is at the base of this right, LCEL, uh, link chain expression language, this is a layer of abstraction above effectively programmer specific code.
So if you are a business user, if you, you're a power user for Excel, if you can write a BB script to make Excel, do amazing things, you will love this. If you can write complicated formulas, you're gonna feel comfortable with this after taking some time to understand the basics. Lane chain expression language, um, allows you to put something simple and have it exp expanded underneath.
On top of this, there's connections to, uh, models for, uh, model io, right? So integration components to models. They, uh, whether this is OpenAI hosted, whether this is, uh, a model that's you're running in hosted inference at hugging space, whether it's a model that you're running locally on your laptop, or maybe you're running inside of a container in your enterprise infrastructure, there's retrieval integrations.
And I talk about how our use case where we just dumped all of our documents inside of our, our user data folders in, in, in a Google drive, and then, um, wrote a little out that said, suck in this data, suck in the vector store and make some sense outta it. When we ask the agent what's going on with our customer base and our business, um, there's retrievers, there's document letters, there's all sorts of fun things that stand inside this retrieval bucket. There's tooling for agents, and we'll get into what agents are a little, uh, a little further, uh, down the line, but there's tools that allow the agents to like, go to the internet, maybe log onto Jira, uh, and pull some ticket statuses.
Uh, maybe you want to, uh, maybe those agents want to do something with a local file system. Right? On top of that, we have our chains, right?
So that's our application layer. These are chains, agents, and executors. Um, we'll get into a little deeper, but chains allow us to chain prompts together to change some of our logic together, but more importantly, allows us to have reusability and composability in our AI applications.
These expand into templates, which, uh, are part of the recently released link serve application. Uh, there's a hub out there that allows, uh, any of us with a simple command line to go ahead and download and edit templates and post templates to the projects. Uh, this can be used both internally.
You can post back to your own repos, you can post back to the public reposts really, really exciting times. Uh, this came out about 30 days ago, and, uh, we've been playing with it here in awesome, uh, link chain users group. Um, and we'll get into it kind of our, in our later sessions.
Uh, Lang Serve itself, which is a way to take those templates and run little AI agents within, behind a web services layer. And, um, and go ahead and effectively connect to them in, in a way that you can either run 'em on Google Cloud, run internally, and then Lang Smith, which is a tool, uh, that we use to go ahead and see what our applications are going. So we're gonna take a little bit of a deeper look, gonna re uh, refer back to some of these concepts, and then we're gonna swing in to getting some labs, and we can get some stick time playing with this stuff.
Okay. So today we're gonna be working with chats. So there's two types, uh, of large language models out there.
Two kind of switches of how they function. It's really how they're trained. Uh, the first type is a chat model.
When I ask, when I ask an el something simple, it infers a, uh, it, it, it, it infers and expand. So basically, I'm gonna get a, uh, a longer response out of it because it's, it's expecting me to chat with it. One-on-one, I think some human is interacting with it.
Now, there's other modes and other models you can connect to where when you write your code, it, it sends a very specific response. Both are good for both for different situations. Today, we're gonna be chatting with our applications.
It's gonna have a complex output for a simple input. Now, models themselves, what does Lang Chain connect to? About everything.
About everything. So there, if you, Lang chain has 20 plus integrations, as well as about 10 different text embedding models. Uh, we don't, we won't go into it too different, uh, too specifically, but the, the embeddings are the ways that we can store, or kind of the format that we store our files in, in maps between space and time in our vector store.
It's incredibly extensible and provides a great level of layers of abstraction, which is really good whenever you're creating an application. Next concept we talked about is a fundamental component of building a chain. Now, if any of you are like me, you spend a lot of time in chat, GPT, all of us are, I was using it like five times this morning to, to structure discord announcements and Twitter announcements and stuff like that, right?
So one of the first things that we learned in our prompt engineering is we want to tell the AI agent, like, what should we do? What should you do for me, right? So in this case, we have a prompt, right?
And this is, I took a little, a, a little snippet out of one prompt, and I want you to pay attention to the purple parts of it. The, the target, the bolded parts. So prompts allow us to effectively store our system message.
So instead of blogging in each time and saying, you are a helpful ai, that's an expert in software development, specifically Python and fixing my books. And then I paste them, my bugs, and I see, see, see what comes out with, you can go ahead with Lang chain. You can basically store that prompt, and you can say, in this case, you're a helpful assistant.
That rewrites, rewrites the user's text to sound more upbeat. Now, that is up here, I think you can see my, my cursor. Now, if we screw, if we look down at the bottom, and we don't have to go through all the code right now, but where it says text equals, I don't like eating tasty things, that's a message that we passed it up to the, to the prompt template, and it's gonna go ahead and respond with the AI message there.
It says, I absolutely love indulging in delicious treats. This is one simple example. Uh, there's some examples that I've been using recently, which is to take input and create a full formed message that is good for, um, uh, that, that that's a fully formed ticket, right?
These are simple use cases that we use to build our chains, right? So speaking of chains, I think my thing is causing a flashing on my camera here. I'll turn that off.
Okay, so taking our prompt templates, right? We chain them together, right? So there's a bunch of different types of chains that we can, that we will see inside the link chain project, right?
There's a simple chain, which is, if, if the simplest chain is one link, which is the prompt I told you, put one input, you tell it to do something and output comes out, well, we can link these prompts together. We can not only link these prompts together, we can have each chain use a different LLM, use different models, use different sources for their data. Um, we can chain chains together, right?
So if we have a sim, a simple chain, which is a chain of chains, right? We have a sequential chain, which allows us to split our input into two chains that come together in a final sequence and output. Those are really handy.
And then we have a really fun one, which is a router chain, which is basically we define the personalities of our AI agents. And then the input chain actually, uh, directs the re requests based on looking at the, the, the data that we put into it. What is the best, what is the most capable AI agent that can service this response?
So when you think about it, if you're writing up a, uh, a customer service bot internally, that goes off your internal documentation, and then you're also writing a, uh, a distribution calculation bot, which takes in, um, the vehicles that are transporting your goods, the percent, the weight, the percentage weight of the trailer, um, the distance tran, the distance transported, and like if it's diesel battery or train, uh, ship or whatnot, and compute your, your carbon impact, right? These are type of things where you can specify your chains to, to give it lots of, lots of features and functions, lots of specif specificity, and even to pop out to things that aren't good at math, like, well, for mouth or whatnot. Super fun.
We talked a little bit about retrieval, augment generation in the always cool brands use case. Um, for that, there's 50 plus document loaders right now. So whatever the document type, there's a lot of stuff that can pulled in.
And there's also been most, in, in recent times, a really exciting, um, uh, a really exciting growth and what's called multimodal rag, which is the ability for this, these, this code and these LLMs to understand images natively as text. And same thing with, uh, with audio as well, both to understand it, but also transform it back out. So really, really interesting.
Um, there's 10 different splitters as we talked about. We talked, it connects about 10 different, uh, vector stores. So whether, whether you're using pine cone, whether you're using chroma, whether you're using, uh, Mongo, you can use Redis.
Uh, there, there's, there's a huge list that you can go through. Um, there's a lot of really, really good ones out there. Uh, and I encourage you to play.
And the next thing we talked about a little bit is the tools, right? So tools are generally functions that you expose to an agent. An agent is, is just an AI agent.
You type something and you tell it to do something, and it uses kind of, uh, it uses a framework to figure out what to do. Um, common tools are internet search, right? So you might have a DuckDuckGo tool that you expose out to your, your AI agent.
Um, you may, uh, connect in through a Google search. A API, it works pretty good. I use it myself.
Um, you can also do a really fun thing in your tools, is you can expose multiple vector stores to a tool and start to apply some sort of rules-based access control to your agents as well as your people. 'cause humans are agents too, right? Um, there's 50 plus tools that are in the repo.
It is really cool modular way of building these AI applications. And really, I kind of think 'em as robots, um, a lot. So speaking of agents, uh, agents are independent entities.
They generally have access to cools. Uh, the, the, the common, the common thing you're gonna see with an agent is it reacts, uh, the react framework. You'll see a lot in LMSs, and that is it reasons.
Uh, so it takes a look at your, your, your query, your prompt into it, and it, it does some reasoning. So how can I best, uh, achieve this goal? Um, and sometimes, and then it acts and it'll kind of loop through it, you know, it's, it's not uncommon to see maybe 127 inter iterations on something complicated.
Um, so at the end of the day, and this is part, part of how we tune our applications, how we, uh, build our prompts to be able to minimize this expense, minimize this time, but also this is an int inte a intelligent digital worker that is working for us, right? Really, really cool. So, again, these are a combination of, uh, large language models, the code we write or use, um, memory and tools, both short-term and long-term memory, right?
And we'll get into a little bit of that in our labs. Uh, linker itself allows you to take AI agents or AI templates and effectively, uh, expose them behind a tool called fast, API as a web service, uh, a self-documenting web service, uh, that allows for extraordinarily easily templating, as well as, uh, it puts a little playground, uh, that you can work with it really, really neat. Um, and as you progress in your lane change journey, you're going to be using a lot of this Lang Smith.
Our final component is, uh, it's a web interface. So this is a tool, you don't have to use it. You can, you can take logs locally, um, and do like a call, uh, a callback handler, uh, that, uh, pulls information going into lm.
The challenge is, and that is great if you're popping this into lowkey or if you're popping this into some of your log management systems. However, for a, uh, an easy user interface that allows you to tweak your tweak, your prompts, allows you to create different data sets, it allows you to visualize latency. Um, the link chain team has a web-based tool called Link Smith.
It's really, really cool. Um, and it's easily added to any of your link chain code. You basically export a variable of the endpoint, uh, your key, and, uh, you get a web interface that is super duper easy to see what's going on and improve these things.
Um, it is in, uh, a beta right now. So if anyone needs, uh, access to a, a beta key, pop onto our discord, um, say hi, and we'll get you one again. Lane chain itself is a bunch of components.
The documentation can be, uh, can be complex as a user group. Um, we are coming together and teaching ourselves in our community here in Austin, um, how, how to use this code and sharing, sharing our code, sharing our use cases, sharing our learning. Um, link chain itself is, uh, a combination of, uh, a bunch of components, an application layer, templates, links, or, uh, templates, a service layer, and then a debugging tool chain.
Okay? So let's get to labs. Okay?
So first thing I want you to do, and I'm sharing slides here, I'm gonna share my screen. I, I, I'm gonna share my entire screen, uh, window. One moment here.
Let me, uh, cancel this, and I'll get this over here, and I'll share my screen tire screen up. There we go. Cool.
Okay. So the first thing that I want you to do is go ahead and let me make sure I'm looking at the lab slides here. Go ahead and get your API key.
com/api keys, you'll get to this, right? So for those, let's create a new secret key, and we'll call this Lang Chain one oh one by the lab, create a secret key, copy it. Now, this is live, so I'm just gonna go ahead and delete this one.
Delete revoke. Now we're gonna do next is come into our repo. Okay?
So that is where is my screen here, go to slides. We're gonna do next is go to our repo. So if we either scan this QR code, go to this repo, and specifically we're gonna go to this notebook, right?
So for those that's in the repo, there's, I'm gonna give you a little tool here. So the first thing is we have a resources. Uh, first thing is read me.
So this is a little bit what, what we are resources and presentations. So this has all the presentations that we, that we're given. We also have, um, our notes for all our meetings.
You can see how something like this is built. We have, um, our lab guide and the presentation. So you can open up either of these.
Now, anything that's in resources, uh, it is under the, uh, creative Commons license. So what this allows you to do is basically, you can use it for your own internal uses. You can use it for corporate usage.
We don't care. Just give some attribution, throw some, throw some shouts to us in lane chain where you found it. Next, we have our labs.
All the code under here is, uh, licensed under Apache two, because this basically means you can use it for private, you can use it for public, you can use for business, you use it for whatever. He just can't sue us for copyright. Um, sort of extraordinarily extensible.
The goal that we've had for us is that within Austin, but also the global community, is that as we are teaching ourselves how to use these AI middleware tools, that each of us can go and teach it at work. We can teach it at school, we can do whatever we can share. So I wanna start with everyone going to the Lang Chain introduction.
Okay? So what you should see here is a page that says Link chain introduction. This is a Jupyter Notebook.
Um, Jupyter Notebooks are effectiveness webpages that can run Python. Now, as we're looking at here, there's no place to press play, right? So let's go ahead and click on open in CoLab, and then we'll see if it connects.
So what's Google CoLab? Um, anytime you're playing with, uh, AI code, uh, not anytime. Many times when you're playing an AI code or you're doing any sort of data science or machine learning, a lot of your courses are gonna be using, uh, Google CoLab.
Uh, what this allows you to do is host Python applications or host the Jupyter Notebook notebook applications, effectively Python, um, and run them up in Google's platform. Uh, there's a free tier, which we're using. Uh, I have pro in here, but we're not using any of the pro features for this.
So the first thing I want you to do is file and save a copy and drive, and it's gonna allow you to go ahead and if you need to make any edits, you'll go ahead and save it. So any of our labs here will just be in your Google Drive. You don't have to worry about getting down to our repo, although as you get a little more comfortable with this stuff, you'll, you'll show up.
So the, let's go ahead and introduce, first thing that we're gonna do is we're gonna execute some codes. So in these black, um, these, these, uh, black sections, this is code that's actually gonna be executed. Now, it's not gonna be executed on your laptop.
This is safe. Um, it's actually executed up at Google inside of a containerized environment. So the first thing that we're gonna do is we're gonna use a Python package manager called PIP, to quietly install the Lang chain packages, the OpenAI packages, and then we're gonna throw a coherent tick token.
Um, because this is a fast moving project, it'll complain if you don't have these, uh, installed with it, um, even though we're not gonna be using 'em today. So, as we press play, and it says connecting, there we go. Okay.
As you as it, as we press play, we did a quiet install. So we don't have a bunch of stuff crowding up our screen, but what it's actually doing is executing code is downloading from the internet, a bunch of different packages that allow us to connect back, um, connect to our stuff, right? Connect back to OpenAI, but also a L chain, which is effectively an application, a set of applications, set of modules that we can use to, uh, avoid writing code.
Next one, imports and packages. So for those people that are, um, not heavy-duty, Python programmers, importing a package is effectively loading software into this little Python robot that we're building here. First thing that we're gonna import is a base callback handler.
This is effectively, um, helping us un helping our application understand how to connect to, uh, different models. We're installing that from the Lang Chain project in the callbacks folder and base. Right?
Next, we are importing chat OpenAI. So I talked about the, the type of models that we use, right? So chat versus a standard, a large language model.
We're saying, okay, we're gonna import some code that allows us to chat to utilize OpenAI as chat model, right? We're installing that from Lang chain under chat models. And next we're in, we're, we're installing, uh, or importing the, uh, chat message application.
What this allows us to do is get a sense of state. So each time that we interact with the machine natively, um, it'll forget about history. But what this does is it does, uh, it gives us history in our messages, okay?
And that's, that's from Lang Chain's schema. So we press play on that, and that works. So, green checkbox means all good.
So next, what I want you to do is to go and, um, put a secret in. So on your left hand side of Google collab, there's this little key. There are a couple things on here, right?
So you can, you, you can access your files and your folders that's inside of your instance kind of cool. So if you're creating files, you're creating outputs, your learning, you can dump stuff in here and download it. Uh, you can also get a terminal if you pay for pro.
Um, but for our use cases, it creates a place where you can save your secrets. Um, it, it makes it so you don't have to put your, your keys in your code. It's kind of handy.
So in this case, you can see that I have OpenAI API key. So what I want you to do is go to the bottom, click add new secret, and then we're gonna type in OpenAI underscore API underscore key, right? We're gonna paste it in.
Now, this is the key that I just deleted to you, so this isn't gonna work, right? So I'm just gonna ahead and delete this, but you'll see, you'll be able to save yours. Where is it?
Oh, there, it's, you'll see that I have an OpenAI, API key. Now I can do two things. I can manually turn it on just by pressing this button.
Now, I talked a little bit about, uh, Lang Smith and, and our ability to debug applications. If I wanted to, uh, turn on Lang Smith, right? Um, I could go ahead and turn on the API key the endpoint in the project, and then turn on tracing, and I'll be sending all this application up to my debugging application.
Super, that simple. So hopefully by now you put your OpenAI API key, um, to find it in this specific wording and then paste it into this application. And now we go to back to our code, and we want to go ahead and pull that a, that API key from Google Cloud.
So from this interface, so from Google collab, we're gonna import user data, right? So this is the Google Cloud, um, code. Uh, we're going to pull the OpenAI, uh, a p key from the user data, and we define, which we define in uppercase really loudly.
And we're just gonna go ahead and print it if it's in, in this Python application. So now we get, we have a return that says, yes, the proper API key has been found in, in our cloud secrets. Next, what we're gonna do is we're actually gonna create an object.
So, um, in Lang chain, uh, this Python lang chain, uh, we create create objects. So what we're gonna do here is we're gonna pass this open AI API key. We're gonna pass that API key to the software chat OpenAI that we installed from Lang Chain.
And we're gonna instantiate that in object called LLM. So we don't have to type all this complicated stuff all the time. We can just go ahead and interact with the LLM object, and it'll have chat OpenAI with a proper key press play next.
So since we, now that we've defined our, our large language model, what we wanna do is give it some memory, right? So there's many ways within, uh, uh, our, our lmm application LinkedIn applications that we can give memory many different types of, of memories. There's short-term memory, there's long-term memory.
We can keep memory, um, in in chat message like we're doing here. We can dump it to a JSO file, uh, which, which some of the A GI applications do. We can throw it in a vector store.
And there's many ways that we can have memory, both in our agent, but also the notion of shared memory in this, in this case, we're using chat message and think of it as just a file. It's an object, really. Um, so we're gonna insert a message using chat message, and we're gonna say the assistant, or we're gonna say this, this message is from the assistant.
So, um, we're the lmm itself. We're gonna tell the, the lmm to prompt to, the first thing is telling that large language model, that chat chat model to say, how may I help you? This is actually what it's telling us.
We're bundling this inside of messages. Now, next we need to give some sort of human input. So I want you to do is press play here.
And what this is going to do is gonna open up a box. You can ask your l your large language model, any question. So, let's see.
Say, um, I am developing AI applications. Uh, what can I do to minimize risk of an AI application? Putting my computer now deleting files.
Can I do anything with containers? Okay, so we've stored this prompt. Say, I'm, I'm an AI application developer.
Uh, I'm curious, um, what can I do to minimize risk? And does this run in a container? So we're gonna take this input, the prompt that we had, and we're gonna say, this is from the user.
We're gonna pass it through chat message, and we're gonna pin it to our messages object. Next thing we're gonna do, we're gonna press play and print our messages, and we can see the first message we passed in, which was to the lmm, to tell it to be effectively be helpful. And then we have our user prompt, which is, you know, I'm gonna developing these AppSec, so what can I do to make it more secure?
And let's take this me message right here. Let's, which is represented by this object, pass it to the lmm, which is the object that represents the API key in our chat message and store it in a response. Let's press play.
And you can see on the bottom here, let's make, it's doing a lot. It's rap it, it's wrapping it, it's handling all the responses, it's passing it up to lmm. Now let's go ahead and view the response.
There's a lot of information, by the way. So you see the AI message inside of here, and then there's some other stuff. What we can do is we can actually filter the content of the response back.
So let's press play, make it a little bit easier to read. So in this case, to minimize the risk of AI application, deleting files. So you can take several precautionary member measures, including the use of containers, um, slash this is kind of a, uh, how Python responds.
Um, but this slash n slash n one just means new line. So says, for containerization, you can use docks or Kubernetes, Kubernetes isolate AppSec. Uh, in sandbox, you can limit permissions.
You can say, you know, make sure that each a app application runs with at least the least possible permissions necessary. Um, you can backup, use, backup and recovery with, with this, and then you can test and verify. Okay, that's a really cool message.
So next, we're gonna take this response and we're gonna pin it to our messages. So what we've gone through here is we've gone through one run of, of having a conversation with, in a, with the ai, but wait, what if you actually want to have a conversation versus just one, one message. So let's scroll back up to number six.
I want you to press play here and let's continue this application. Be like, okay, Wow. Um, oh, that is not what I want Here.
Um, that's good info. Mm. Well, my auditors feel, um, when they see my AI app, locations, open containers, should I, once I see how they feel, okay.
And that's in our messages, viewer messages. Now, what we can see here, the first interaction, the second interaction, the response, the response, and then the new message we brought in. So one thing, when you're doing, when you're playing with these AI applications, they have a window.
So there's only so much data you can stick in. And even when you have a very large window, like I think some of the latest ones have 128,000 tokens, um, they'll kind of loose fidelity depending where you are. But, so one thing to remember is when we're, we're writing our simple A app applications, um, a lot of the game is figuring out what we're going to pass up to these models to control cost and, uh, to control, to ensure functionality.
So let's run this new stack of messages, um, through the LLM, and we can see on the bottom that's executing. Again, this is, this is handling, this is all up in, uh, inside o uh, Google collab. Let's take the response and let's make it a little more readable.
So auditors generally view the use of containers, clouds of light. Well, that's cool, um, 'cause of isolation. Um, reproducibility reduce potential security breach, uh, and version control.
Um, okay, cool. And then we'll press play because we want to, again, append this to our messages. And then let's ask like an architectural decision, like, okay, like pleasing my auditors.
Uh, what will they think about me setting up a pipeline where my business users commit to a gi to expand simple AI services? Can I do anything to make this more robust dinner? We're gonna go ahead and, uh, uh, pin this to our messages.
We're gonna display again, all of our messages. So we always have a chain of thought here. We're giving sort of a personality.
Now, one of the fun things that you can do is as you're inter, as you're building your applications in playing with them, you can go ahead and use some of these to structure a more complete and, uh, functional application. So let's pass this new message, stack up to the lmm, and store it in the response object. As you can see on the bottom, it's executing.
Check our response. Let's filter it a little bit so it's easier to read. We have a lot here.
Okay? Setting up a pipeline where business users commit, um, can be a beneficial approach. It promotes collaboration agility.
However, there are few considerations to make it robust and aligned with auditing requirements. Access controls. We need to implement proper access controls, uh, code review.
So implement code review process. Ensure the made changes made by business users, uh, meet the required quality and security standards. Uh, versioning tagging.
So use proper versions and release automated testing. Uh, good CID cd, CI/CD practices, um, and documentation. We need to make sure that these things are documented.
Uh, by implementing these measures, we can create a robust pipeline. I honestly think this, uh, in my experience in past lives, I think I had to do 38, uh, audits a year when I was a service owner at OCI, stuff like this would actually fly. Um, so it's kind of interesting.
Okay, let's pin this response. So as you can see here, we have the ability you just wrote or interacted with, and you can change a lot of stuff in here. You built your first working AI application, running inside a Google Notebook.
Now we're gonna do, um, further on in this, in, in, in our labs here. We're gonna show you how to create a simple AI application and build it inside a simple web interface called, uh, with streamli, uh, that you'll actually be able to access through a browser. We're gonna do some hacks, so it actually runs on Google Cloud.
Uh, then we're gonna see how to do, how to make streamlet pretty and, and how to make it a, a fully functional interface. And then last thing we're gonna see, we're gonna see in our labs is how we can actually use aama to, uh, make it so your applications don't leak any data. Don't talk, don't share any data externally.
So for those of you that are in military or government applications or in high security application, or maybe you just want to be able to run an application that runs on your own laptop lab at home, or your own servers. So I'm going to go ahead and introduce, let's get back to slides here. Introduce, uh, Ricky Pirruccio to take the stage.
Hello everyone. Um, my name is, uh, Ricky. Um, I'm from Austin, Texas.
Uh, wanna thank, uh, our host for having us here, and Colin, for putting this together. Um, and, uh, so telling you a little bit about my background, um, I, uh, have a bachelor's in mechanical engineering. I've been working in manufacturing for the last six years.
Um, I started out on aerospace and, and then pivoted in, um, semiconductor equipment manufacturing in my current role at Applied Materials, working as a manufacturing engineer. Um, I do a lot of work in supply chain, um, and that involves a lot of, uh, uh, data-driven, uh, business decision making. Uh, that's kind of how I use my, uh, tech skills, uh, in the manufacturing sector.
Applied. Um, my tech journey started a couple years ago with a hack reactor galvanized. Um, I took a full stack, uh, JavaScript engineering course.
Um, and I learned, uh, not just JavaScript, I learned a whole lot about software development in general, debugging testing, uh, running containerized services with Docker, uh, deploying code, um, interacting with databases like sql, no sql. Um, and then later, uh, picked up some Python, Python. Um, so my, my current interest in using Lang Chain, um, comes from, uh, educational, uh, sort of learning perspective, but also, um, trying to understand how to, uh, integrate AI and Lane Chain into my current work.
Um, supply chain is filled with gaps, many gaps, um, which stem from, uh, uh, sort of a, a knowledge and, uh, also lack of, uh, sort of acquiring data in the way that you need it. Uh, so I believe integrating lane chain ai, um, into supply chain and applied, we could dramatically, um, improve these gaps and, um, resolve these issues that we currently, we face all the time. Um, at the bottom, you'll see some of my links as to my social media platforms, uh, LinkedIn, GitHub.
The QR code here will take you to my LinkedIn profile. Uh, please connect with me, um, be more than happy to. So we will start, um, we will start with, uh, a brief tutorial with, uh, access and chat model via the API and a web interface using streamlet.
Um, so little bit about streamlet. Um, it's a powerful bionic frontend framework that allows us to build, uh, beautiful interactive web AppSec for link chain projects, uh, with absolute ease. Um, so for those engaged with Link Chain, uh, streamline is a game changer.
It allows us to focus on our core projects without being bogged down with the complexities of frontend development. It's intuitive, it's onic, and it seamlessly integrates with the tools we already use in our link chain projects. Um, so we're gonna dive into this lab.
I'm gonna share my screen. There we go. Okay, so here is our lab, and you can find this in our, in our repo.
Um, um, it was on the QR code beforehand. Um, probably go back to the slide, so you can see that It's also in the, uh, the presenter's. It's in the handouts too.
So if you go, if you go into your handouts, there are links to the repo. There you go. Yeah.
So here it is, uh, going back to my screen. So here's our lab. Um, this is in CoLab as, uh, Colin just showed.
Um, so, uh, we can, we first start off with just installing, um, packages, which I already done. Um, and again, this is, uh, this lab is for just demonstrating how we can use streamlet, uh, with OpenAI, uh, in order to build a chat bot. So, uh, the goal here is to just kind of show just the, how easy it is, uh, to do something like this.
Um, and Streamlet is really the facilitator here, so already installing the packages beforehand. Um, and, uh, the way this works, you're gonna reconnect here. Okay?
So you gotta connect to your runtime and collab, uh, to, uh, have access to the Python environment. So here is our script on this, uh, second cell. Um, it has our entire application to make, uh, an interactive chat bot with OpenAI 40 lines of code.
Um, so in order to, uh, save the script in our runtime storage, we will run this, it's gonna save it to the streaming app dot pi file. And, uh, then we will serve our streaming app. So we're going to press this command, and it's going to run our streamline application on our runtime.
Um, this, uh, and, um, greater than operator, uh, it's used to, uh, redirect standard output, standard error to a file name logs text, um, in the content directory. And so we can use it for logging purposes. Uh, that's all it does, and the end operator, uh, basically, uh, tele streamlet to run in the background.
So allowing us to execute for commands. So, uh, running this gives us an IP address, which we will need for, uh, this command right here. Uh, we're using NPX, um, which is a Joss Script package manager to run local tunnel, uh, which essentially, um, we can use to host our application to the internet, um, very simply for just demonstration purposes.
Um, and we're exposing port 85 0 1 because that's where streamlet runs on, on our local, uh, runtime. So, uh, the port will be exposed and we will have access to streamlet on the internet. So it runs, it opens up this, uh, window.
I'll add these side by side, so you can look at the code and, uh, uh, render application. And here we just gotta paste in our IP address so we can access our application. It's kind of a security measure of, uh, local tunnel.
And there you go. So the way this works, um, users' first input or OpenAI, API key in the sidebar, um, which you used for, um, interfacing with, um, data ai. So I already have my API key here.
I'll just paste it in this box and we can start using our application. So I'll just say hello, there you go. Link chain.
And it gives us some, uh, information here. It doesn't provide information. L chain is a blockchain, so it would add some information on Lang chain.
Um, Well, this is, remember that our model, it, this model was trained before Lang Chain was released. So, you know, this is a great example, uh, of hallucinations from a model. Uh, although when we're interacting with our agents, we get 'em access to recency data and access to the internet.
It can Kind of fill those. There you go. Um, so yeah, this is, uh, just a brief demonstration, how to use, uh, uh, just how to build this like chat bot, um, 40 lines of code.
Um, if we dive in just a little deeper, um, streamlet, um, all it's doing is, uh, we have, we're, we're, we're saving these messages on the session state. Uh, and then we're basically using a stream handler, um, as a callback, um, to update the chat window, um, as the AI generates responses. Um, so at this point, I want to move on to our second lab.
And, uh, going back to the slides, here it is. So here you can, um, if you have to, if you, uh, use this QR code, uh, it will take you to our repo. So in that repo, um, it will have the link to collab, um, and then it will, it will be this, uh, notebook right here.
So, um, this is just a introduction of streamlet, how to use, uh, streamli, um, in order to build these, uh, web interfaces with Python. Um, just wanted to show how easy it is, um, and how you can start using it. So, um, we first start with installing packages.
Um, here we're using PIP with just install streamlet. That's all we're using. Um, already have it installed.
Um, and then we can move on to creating a simple streamline application. So 19 lines of code. This is a very, um, simple application.
There's no widgets. Um, it's basically just statically, uh, rendering, um, some texts and numbers. Um, title header, um, just to kind of show the ease of doing this with stream.
So We will first save the script to our Streamline app dot pi, and this works very similarly as the previous lab. Uh, we will run streamlet again on the background. Uh, here.
We will generate our IP address, uh, that we then use for, uh, local tone and host our application to the internet. So here we go. We have a side by side, we get the IP address, paste it in this box, and here is our render application.
And, uh, here's the code side by side. Um, so yeah, we got the title, we got the header, uh, we got the texts, uh, we got some numbers. Um, St right is, uh, kind of the way, just like a Swiss Army knife of, uh, rendering anything, um, on the screen.
Uh, that's kind of what Streamline calls it. Um, and what's cool also is that you can literally just have literals in your code, um, and streamline it will wrap it the st dot, right? Um, all on its own.
So you really don't even need to use st dot, right? You can just literally have literals on your, on the, on your code and it will get rendered. And now I'm just kind of looking at this, uh, ui.
Um, I mean, it looks good. Um, it's already formatted. You didn't really have to put a lot of thought into it.
It's kind of our goal with using Streamlet. You know, we, we wanna focus more on building these AI applications. We don't want to, um, you know, spend a lot of time with, with front end things.
Um, and I just wanna kind of show too that you can also get a dark mode just like that. Very neat, very cool. Um, so moving on, um, kind of wanna talk to you about making the application more interactive.
Um, to do that we will use Session State. Okay? So, uh, session State basically, um, is, um, mechanism, uh, to share variables between reruns of each user session.
Uh, so this allows storing a persistent state. Um, it allows, uh, for manipulating state also using callbacks. Uh, and when we refer to state, we mean the information is streamline.
App keeps track of while it's running. Um, so we will re-save our streamline app pi file. We will, we then first, okay, so that got that file got written, uh, we then have to rerun our stream lab in the background rerun this NPX local tunnel command.
Open this in a new window and paste in our IP address. So going back to the code, okay, so state is really easy to initialize and streamlet. Um, you do it just like a, an object.
So the session state object is just an object, um, called session state. Um, you can pass in any key, um, like counter for example here, you can name it wherever you want. You can use bracket notation, you can use dot notation.
You can also, um, use, uh, with widgets. You can have this like key property, um, argument that you can, um, you can let initialize state with. So you have several options here to do.
So. Uh, what's neat about state also is that you can, um, use callbacks to, um, update it and maintain it. Um, so here in this application you will see this object right here.
That is the session state object that is in this line right here, 54. And uh, is literally just displaying the state in real time. So we can enter our name and you can see the session state object being updated.
We can do the same thing here with some texts. Uh, we have to interact with this button widget in order to do that. And you can see the text being updated.
We also have an increment counter button here that will just increment this counter, that's initialized at zero. So you see clicking that goes to 1, 2, 3. And then we also have, um, another button here called Delete State.
So that just, uh, resets the state. Um, and you can see that the state has reset when we click it. Um, really neat.
Again, very simple. Um, not a lot to it. The state is updated with these callbacks Um, and, uh, it makes it for a very neat, uh, ui.
Now, the last thing I wanna talk about is caching. I just wanna show you how simple it is to do this, a streamlet. So, doing this again, we're gonna rewrite our, uh, streamline app pi file.
We're gonna rerun our streamline app, restart our local tunnel. And there you go. Now notice it's, it took some time there.
Um, so here's what's happening in Streamli. You can really easy cash. re cash, uh, underscore resource, uh, decorator.
The streamlet makes available. Uh, so you can just wrap this with any expensive computation and the results will get cashed every time. Uh, so now in this expensive computation, we're just, uh, taking the square root of a number, um, sorry, squaring a number.
And, uh, we're, we have a fleet timer here for three seconds in order to simulate the extensive computation, taking a long time to run. So if we do four, for example, press enter, you'll see the timer running, meaning the computation is running. Um, and after three seconds, this gives us our result.
16. So if we go back, let's do another one now, like, uh, five, you can see the expensive computation running again, taking three minutes, three seconds. Um, if we go back to four, we don't have the timer anymore because the result is coming straight from the cash.
Um, and same thing with five coming straight from the cash. If you have another number, the expensive computation runs again, because it hasn't been running yet. So it's is a really neat feature.
Um, you know, working, uh, with, uh, lane chain, uh, we certainly have many expensive computations that we would certainly could use caching, uh, for. So really neat feature, really simple. Um, and, uh, with that, I will kick it off to Kareem and he will show you how to run a streamline application.
But instead of using OpenAI is going to use a local model that lives in your CPU. We cannot hear your audio. Kareem.
Nope. Can you hear Ricky? Nope.
Okay. Sorry. You, you cut off for a little bit.
I can see you and hear you guys now. Kareem, can you speak? How about now?
Hey, we can hear you loud and clear. Welcome back. Awesome.
Uh, okay, so, Uh, what's the demo with our technical issues, right? Yeah. Okay.
Uh, so, um, hello everyone. I'm, I'm excited to be here. My name is Kareem Ani.
Um, I am, um, currently resident of, uh, leaner Texas. I'm a software engineer with the Office of the Governor. Um, over my career, I've worked with, uh, a lot of different, uh, technologies and stacks.
Um, and, uh, my current, um, interest in long chain is, um, essentially, um, uh, you know, ways to explore how I could use large language models in use cases where it might not be feasible or viable, uh, to rely on large language, uh, model vendors like, uh, OpenAI, andro, et cetera. And, uh, this may be due to a variety of reasons. Um, not nec not limited to concerns around data privacy, legal compliance, um, alignment issues, uh, or, you know, even trust in stability and future of, uh, you know, set vendors.
Um, it's just a way to, you know, or maybe just a way to hedge betts against, uh, any, you know, uh, issues that might arise in the future. Um, the other thing that I'm really interested in is, um, looking at deploying, um, you know, local large language models, um, uh, you know, in a way that, uh, would scale for small to medium sized use cases. Um, um, my, uh, socials are, uh, up on the screen.
And, um, with that, um, let's jump straight to the, uh, to the lab. Um, uh, this lab builds on top of what you've seen so far, um, with, um, Collin's, um, introduction to Lang chain, uh, in a notebook, um, without any user interface to, uh, Ricky's, uh, lab, which puts a streamlet wrapper around that. And lets you, um, you know, communicate with the lang, you know, with the, uh, large language model, um, in a chat format.
Um, I, uh, like I mentioned, you know, in my case, I'm mostly in, uh, interested in, in exploring, uh, the possibility of bringing in your own, uh, you know, hosting your own models or, um, you know, uh, doing it in a way where, which would allow you to easily swap out, uh, models, uh, you know, from a, uh, one for a model with another. And long-chain really makes it, um, um, uh, re very easy, uh, to do that. Um, for this, uh, for this lab, um, you can use the QR code to, uh, take you to the, uh, exercise.
The exercise is also available in the handouts. Um, so let me go ahead and share my screen and jump right into the lab. Okay.
Um, cool. So as with, uh, and we, we try to do this for, um, all the labs that we, um, make available, um, uh, in the, in the repo. Um, have a, a nice little utility function that, you know, uh, a button that lets you open it straight into Google collab, uh, so you can start, you know, messing around with it.
Um, in the interest of time, because there are certain, um, uh, things here, certain commands here that do take, uh, a little, uh, a little time to prime to load up. I went there and I, um, ran them, uh, beforehand. Uh, hopefully, um, it'll, it'll work just as we, uh, plant.
So, let's see. Okay. Um, and I'll, I'll walk through, uh, this code in just a minute.
Okay. It is priming at the moment. So let me go switch.
Let me switch back to the, uh, to the code itself. Uh, this code is going to look very familiar, uh, to what you've, you've seen so far. Uh, we start by installing, uh, our, uh, you know, packages.
In this case you'll notice a, uh, an omission, uh, uh, OpenAI is not included in here. We've, we are sticking with just, uh, l chain and streamlet. Um, uh, the Streamlet application that we build is also, um, going to look very familiar with maybe only a few, uh, adjustments.
Um, one notably is we have a call here, uh, you know, a cash resource call, uh, back to, you know, uh, what Ricky just shared a few minutes ago. Uh, what this function does it, it queries the, um, ulama API server running in the background, uh, to check to see what available models, uh, might be there that we might be able to use. Um, we have a streamli handler, uh, object here.
Um, and this is again, uh, to, uh, that gives us, uh, the illusion of, uh, uh, the, uh, large language model typing. Well, the, uh, LLM is run in, is being run in streaming mode. So it is truly sending us stuff one token at a time.
And this handler allows us to, uh, essentially render that in that way. And it gives us, you know, it sort of helps us complete the illusion of you actually chatting with, uh, uh, a language model. Um, we no longer need the, uh, API key in, in its place.
I've, uh, added, um, you know, a, a placeholder for you to provide your, um, Alama API server, um, and, uh, a drop down here, which shows you a list of models, uh, that might have been downloaded, pre-download on your machine. Um, rest of the code follows a pretty, uh, similar, um, uh, path. Um, with the exception here, instead of chat OpenAI, we are using Chat alama.
This is our streamli app. Um, uh, so far, um, then the, uh, you know, the next step of downloading and running, um, and, and setting up alama, um, for this exercise, you know, we are just downloading it straight from Llama's website, uh, making, you know, setting it to be executable, uh, launching the a p service in the background by, with the llama serve command. And, um, so just so that we have a model to work with, um, I went there and I pre-download, um, mytral and Llama too.
Um, and, um, once that was done, I went ahead and started, um, uh, the Streamlet application and, uh, opened up a tunnel, provided the IP address. And here, here's our application, uh, because we are running this, um, service locally. So one thing that, um, uh, people got a little comfortable with, uh, uh, with the chat, GPT is the immediate, you know, uh, it's, it responds back immediately, right quickly.
Um, however, um, the way large language models work is you have, um, the, uh, the model file, which is in many gigabytes, uh, which has to be loaded up into the GPU. That task alone takes, um, some time. And what you don't see with our GPT is, uh, you know, uh, that delay because OpenAI, pre-primed their servers with those weights, which is why, um, in this case, I had to go ahead and I, um, uh, prime this, uh, before the demo.
Um, in the case of, uh, uh, Google collab, um, it, it takes a few seconds. It doesn't take minutes, but it, uh, no, nonetheless, I mean, that is something, uh, to keep in mind, but once we have it, uh, this will now work very similar to what, um, you saw in Ricky's demo. Uh, so say, um, let me go ahead and clear out the history now that it's Prime say, okay.
Um, I am interested in learning about RLHF. What can you tell me about it? Oh, because I took a little, because I let it sit for a while.
Um, Google collab said, well, you're not using it, so I'm gonna shut it down. So let me, let me go ahead and, um, but yeah, in, in the meantime, while it's priming up, um, let, let's take in any questions. Um, um, you, you have, uh, some use cases, uh, Colin that, you know, uh, you mentioned about, uh, that you use, um, you know, long chain and, um, you know, uh, an interface like this for your day-to-Day.
Um, what, what are some interesting use cases that you can think of, um, for using a, uh, a self-hosted, uh, language model instead of OpenAI? Um, well, many use cases. If I think back to, uh, my first responder background, you know, uh, having to deal with, uh, information disconnected areas, right?
Um, if you are running oil and gas or energy exploration, um, having, and, and you're somewhere in, like somewhere in South America, I mean, access to a model, uh, that you can preload with your data, uh, is incredibly valuable. Um, if I think of, uh, any war fighting capability, if I think of a, say your consumer brand and you want to constrain your costs, and you don't want to risk, um, any external vendor maybe training on your data, um, maybe you want to, uh, create a lo adapter out, uh, and, and apply that on top of this. And, and maybe Enri, uh, uh, further enrich it with, with a rag, with, uh, uh, a rag infrastructure.
You know, there, there's many cases where you want to be able to control the data, say you're a nerd just like me, or just like us, right? And you want to make sure that you have the ability to reason using these advanced tools and not be constrained by Sam Altman getting fired and rehiring you. Like we all saw, we all saw open AIS performance go, go to hell in a hand basket.
Um, having our own resources, being able to switch resources around on the back end. And, you know, some of the great things about Lang Chain, um, is this allows us to stay in business when our providers aren't. Oh, absolutely.
That, that, those are, uh, those are great use cases, uh, and, and, um, all, all valid reasons why you might want to consider, uh, at least checking out, um, you know, the, um, the LLM space outside of, um, uh, you know, uh, the, uh, hosted ones. Uh, let's see, what can you tell me about it? Okay, so, uh, it looks like we are back in business, uh, and, uh, like you can see it's, it's already, you know, yeah.
Um, in this case, um, Alama is giving me the capability, not unlike, um, open ais, uh, where instead of having to interact with a, you know, and another way to interact with a local model would be to code against it or to use something like Lama CPP, but then you run into scaling issues. But with something like a lama, uh, you, you have, uh, an inference engine that is exposing, uh, an API, which is, uh, opening AI compatible, which means that you could just flip a switch and point your applications to, uh, to, to alama. Um, you also have the ability to run multiple models, like in this case, uh, I'm showing you, uh, you know, ulama and, uh, LAMA two, uh, 7 billion per, uh, model and Mistral 7 billion model.
But there's, um, thousands of, um, you know, fine tunes and open source models available, you know, uh, that are either derived from, uh, Lama or Are, am I still there? Yes, you're still here. We got clear.
Okay. Perfect. Okay.
So, um, with all Lama, you know, all you need to do is identify what model works best for you. Um, uh, think of all, so what alama does, um, is it brings you Docker like capabilities to language models. Uh, in fact, the, uh, the command structure for alama is, is very similar to Dockers.
You know, you're doing a doc, uh, a alama pull, um, you know, uh, and you know, you can even create variations on the LA language model. So say, you know, you want to, uh, uh, preset a system prompt. You want to set the parameters of the language model, you, uh, you want to prime it with some data you can create, um, you know, a model file, which is very similar to a Docker file.
Uh, you can say, okay, I want to start with Mytral as my base models. You do base, you know, you, you, you say from mytral, and then you start giving it all of these details. And now you can, what will happen is, um, Ulama will create a custom model based on those instructions, uh, that now you can use for your use case.
So now, um, you can take your prompt engineering game to another level, you know, try out different things, and once you, once you're satisfied with it, now you can have it hosted as an API and, um, and, you know, have your applications take advantage of that. Um, uh, you can have, um, you know, uh, the, the OMA team is current, is constantly working on enhancing, uh, the product. Um, uh, you can, because it's an API, you can put it behind, you can put a bunch of instances behind, um, uh, uh, reverse proxy and, and have, uh, uh, you know, uh, high availability.
Each instance will work with, uh, it'll, it'll, uh, queue up requests. So if I'm, if I'm, um, if I have multiple AppSec con, uh, connected to the same, uh, Ola instance, uh, it'll queue up the requests as they're coming in and, you know, uh, uh, and respond to them one at a time. But, you know, um, you, you still have that capability and they're looking into approaches like, um, uh, you know, uh, pro, uh, batching, which should make it even more, um, robust in handling multiple parallel requests at the same time, um, like I mentioned, because you're hosting it yourself, you run into the classical problem of, uh, cold starts, uh, you know, uh, which is something that you have to, uh, you know, account for.
But for when you're running smaller to mediums as applications, it, it, you're looking at seconds, uh, not minutes. Um, it's only when you're running very large models in the, um, you know, uh, upwards of, uh, you know, uh, tens of 20, you know, uh, um, of gigabytes, you know, maybe, uh, you're looking at a 70 gigabyte model, uh, where you need a hard, you need hardware that's capable of that, but at the same time, those are the types of models that will take a while to sort of get load up and prime up. Um, the other thing that, uh, to keep in mind is, uh, to, to sort of, uh, make a note of is with, so now you have two applications.
You have your, um, alama, API server, and you have your streamlet application, and you can already see how you can create a, a deployment story around it. Uh, you have a lightweight streamli application, uh, that uses L Chain, uh, to create interesting AI applications, um, that you can, uh, you know, uh, deploy on, uh, commodity hardware, uh, don't, they don't require, um, you know, uh, heavy CPU or memory requirements. Um, and then you can have a few instances of, um, alama, uh, like servers running, uh, on with, uh, on instances with GPUs.
And you can start to see a robust, um, AI infrastructure beginning to sort of take shape, uh, where you can write multiple applications that you know, that you can build, uh, to, uh, to run within your, um, uh, you know, uh, uh, organization or, or your communities, um, that can communicate with, uh, you know, with these, um, uh, language models that are being served up by, uh, locally, uh, through all LAMA APIs. Um, that is, um, that is it, uh, from my demo, uh, back, uh, back to you, um, uh, Colin. Cool.
Thank you so much, Kareem, and thank you so much, Ricky. Um, I am consistently impressed by, uh, how you embrace learning in the open with these new project, uh, projects, uh, impressed by the contributions and your interaction with, I just checked. We have, uh, over 130 members of our user group right now.
So our, uh, interactions with the larger community, and you are a great representation of people love learning just like all of us here. So, to wrap, wrap things back up. So, uh, Ricky and Kareem showed us how we can build simple web interfaces and change how they interact with our users and interact with their language models.
Shifting from externally hosted models, OpenAI to locally hosted models by, um, using aama. Um, we showed how we can enrich our web interfaces very simply, and how we can use caching so simple, um, simple but effective ways. We can think of how we build our AI applications in ways that we can, um, that we can deploy out to, uh, out to our end users.
We ask our own, we ask for our large language models, uh, uh, some questions, uh, about how we could possibly containerize 'em fit, fit our, fit our AI applications into a, a format that is deployable as manageable. It's repeatable that will pass audits, um, and that won't, uh, that will minimize the risk of us downloading some code that blows up our machine. Uh, we all know those pranks are out there.
Let's protect ourselves. Uh, Lang Chain itself is an open source project that is publicly hosted. I'm loving the feedback.
Thank you so much. Um, that, uh, wraps in an expression, language wraps in connection to all sorts of models, ability to retrieve data, the ability to define tools to present to agents, to build agents out of prompts and chains, and little, little robots that we can play with. So exciting to be able to both download and edit templates of other people who have done this using Lanes serve, but also to contribute, contribute templates to, in our private repos in our business or clubs or in public re in backup to the public re, uh, repo, uh, open source repo, that is link chain.
We talked a little bit about Link Serve, which is an abstraction layer that allows us to create simple web services APIs. There's some later labs, uh, that you can play with if you wanna play with a repo or maybe you wanna join us at Austin Link chain, both in person and Austin, Texas, but also virtually, um, in our off sessions for people Cannot apply or can't, cannot work, uh, can't show up in person. And we, and last, uh, there's Link Smith, which is a simple web interface.
Uh, if you need a, um, if you need a, um, a troubleshooting interface, it's a little usable, uh, it's really, really good. We have a beta code pop onto our discord and we can help you out. So again, uh, I wanna thank each and every one of you for joining us in learning in the open.
Uh, this is new code. Uh, one way that I heard it described, uh, someone mentioned to me when I was showing, uh, in my early application, it's like, yo, you're like that dude who figured out what HTML was in 19 nine in, in, in 1997. And I replied, I actually was that guy in the small farm town who figured it out, who HTML was in 1997.
Um, we're in this, the greatest transition that, that I've seen in my 25 career, uh, 25 year career. Uh, we are, uh, taking back these applications in these open source projects. We are taking back ai, um, into, into the users, into the community, and we, uh, in Austin Lang Chain, uh, hope you can join us, uh, either virtually or physically connect with our users, connect together.
But again, the code in the repo are, this presentation is Creative Commons attribution. You can use this for whatever you want. It's good.
Give a shout out back to Austin lha. The labs themselves are Apache too. You can use this code if you want to teach this yourself in your own school, in your own organization.
You wanna teach it at work, you wanna stand on a corner and write some code, you can do this. Please contribute back. We'd like to join the fun.
So again, thank you so much. Thank you to our host, uh, at Techstrong. Thank you to everyone who's joining us, and I encourage you to have an amazing day and have some fun with ai.





