Unlocking Value With Generative AI: What Every Business Leader Needs to Know | AI in Action 2023
As business leaders stand at the cusp of the generative AI revolution, it’s crucial to understand the transformative potential this technology holds. This keynote address will dissect the intricacies of generative AI, illustrating its capacity to not only streamline operations but to also engender new forms of creativity and efficiency within the business ecosystem.
Historically, technological advancements have been the precursors to periods of significant prosperity, spurring job creation and opening new markets. Similarly, generative AI is poised to become the harbinger of an unprecedented wave of economic and social growth. We will explore this phenomenon, drawing parallels with past technological shifts, and providing a forward-looking analysis of the prosperity that generative AI is likely to bring.
The talk will further illuminate how upskilling and workforce development are integral to unlocking productivity gains with generative AI. By investing in human capital, businesses can foster a symbiotic relationship between technology and talent, ensuring that the workforce evolves in tandem with these advancements.
Leaders will gain strategic insights into leveraging generative AI for optimizing operations, enhancing customer engagement and driving innovation. This session will also delve into the ethical framework that must underpin the deployment of generative AI, to ensure that the benefits of this technology are equitably distributed.
Attendees will leave with a blueprint for integrating generative AI into their strategic roadmap, ready to propel their organizations into a future where technology and human ingenuity converge to create new jobs, markets, and opportunities for growth.
Join us to chart the path forward in this new era, where generative AI not only shapes the future of business but also heralds a new age of prosperity and opportunity.
Transcript
Welcome to Unlocking Value with generative ai, with every business leader needs to know about artificial intelligence. My name is Mark Hinkle. I run an AI consultancy called Ity Labs.
I've spent 25 years as an executive in emerging technology, and today I help businesses understand and use artificial intelligence. And I published the what I learned in a weekly newsletter called the Artificially Intelligent Enterprise. If you'd like a copy of this presentation, I'll make it available there.
So let's get started. Let's look at artificial intelligence by the numbers. And what we start with is the economic value that generative AI is likely to add to the world economy.
7 trillion. That's, that's an immense amount of value that generative AI has to add to the economy in the next 10 years. Um, on top of that, let's look at two other numbers.
85 and 97 million. That's 85 million jobs that will likely go away due to generative ai. The bright side is that generative AI will likely create 97 million jobs.
So that's a net of 12 million jobs by 2025 that could be created above what would be taken away by generative ai. Mm-Hmm. Another bright spot is that 97% of business owners believe that chat, GAPT, by open AI, will help their business.
So businesses are optimistic about ai. That's the, that's the, uh, world we live today. Let's go in and dive deeper into these numbers.
Why generative AI is ready for prime time. It's here for the long haul, and it's ready for prime time because of three really, really distinct factors. Number one, enhanced accessibility.
In the past, generative ai, or there was no generative ai. It was AI that analyzed data. Today we have generative AI that works via natural language processing.
So instead of the way we interact with ai, um, being code like Python or go, we're actually using our, our regular natural language and English language for us and virtually any language around the world. The second reason that it's ready for primetime is just the massive efficiency gains, and we'll talk about them in a little bit later, that you get from generative ai. Um, it not only makes you more effective and more, um, allow you to do more work, but it frees you up to do tasks that before you couldn't get to, and the higher value work, not the redundant things that take up 80% of our time today.
And finally, um, we're in a boom around ai, AI and machine learning advancements. So we are seeing technology that is moving faster and being implemented in ways that we've never seen before. And these advancements are enabling us to do amazing things that we couldn't have thought of 10, 15 years ago.
So what kind of things are they? The, the biggest thing, and this is why we call it generative ai, and not just ai, is that it can generate unique content. So that's, that's a big difference.
Before we might've used algorithms to analyze data, now we're using it to analyze data, make an inference, and generate original content. It's also able to automate complex tasks, so we can take large amounts of data, and before where you would've had a business analyst, um, looking at that data and manually making comparisons, now you can run it through generative AI to, um, make comparisons, draw inferences, understand the sentiment of the data, whether it's positive or negative. And those tasks are highly automatable and much more efficient than using the human to do.
And on top of that, and the thing that's driving all of this is the ability to do these previous two things allows AI to create massive economic value. And that academic value is, um, what we all want to capture in our businesses. Probably one of the biggest, um, concerns we have is how artificial intelligence is going to impact the workforce.
And those impacts are gonna be very, very large. It doesn't mean that that workforce will necessarily, um, be removed. It will actually mean that we will find, um, different uses for our human power, um, alongside ai ai.
But what that means in the near future is, according to Goldman Sachs, 300 million jobs could be affected. Um, in research conducted by the University of Pennsylvania, um, and open ai, 80% of the US workforce could have at least 10% of their work tasks affected by the introduction of large language models like GPT-4 that powers, uh, chat GBT or Gemini that powers Google Bard. Um, an additional 19% of workers may see 50% of their tasks impacted.
Um, but we don't know exactly when this could happen. Um, but it will have some impact on our daily workload. Let's look at what share of businesses today are looking to adopt AI technologies.
Um, as of this year, um, AI and big data skills, the skills needed to implement this technology is ranked 15 as, uh, core skill for mass employment. Um, it's also going forward is the number three priority for company training strategies between now and 2027. And for companies with over 50,000 employees, it's the number one priority.
And why is it such a disruptive and big priority, um, and why is it gonna be so impactful? Um, the reason that AI and specifically gen generative AI is gonna be impactful is that 40% of the working hours across industries can be impacted by this kind of technology because language tasks count for about 62% of total time worked in the US and the overall share of language tasks in general. Um, 65% of them have high potential to be automated or augmented.
So not necessarily replacing our workers, but taking that work that prevents us from getting to the deeper, more thoughtful work where we can apply critical thinking and creativity is being eaten up by things like scheduling, um, sorting our email, going through our, um, paperwork that could be easily automated so that we can do more deeper work. Now, let's take a look at what kind of, uh, uses for Gen AI and 2 20 23. There are at the very top marketing and sales.
Um, these, this is the number one area where we see people automating today. So, um, out doing outreach, it is a grind. It is very repetitive.
Um, you have to spend time crafting emails or, um, pre presentation decks. You have to do ads. All these things are very much a function of using our resources to create content that could be easily automated, not necessarily deciding what content to create, which is where the human comes in, but actually creating those emails, creating slide decks, creating ads, all of that thing requires a lot of human interaction and that human interaction.
And that point is not as high value as coming up with the kinds of pitches that we want them to pursue or the kinds of assets that we want to create. The other place that we see a lot of automation today is in the development of software. So all of these knowledge worker tasks are very easily augmented with generative ai.
And let's just look at what kind of results people that are using that. Again, what we've seen from different research, uh, Microsoft, which makes a, uh, um, coding assistant called copilot for their GitHub. Um, they're seeing developers have a 55% time savings when they're developing code.
They accomplish tasks more quickly, and they are more successful at completing the tasks assigned to them. Um, another place that we see huge improvements is reduction in human answered customer support requests. So, uh, a lot of the chat bots that we see on, um, websites for the, the cursory answers, the, um, how to change my password, how to, um, interact with billing, all of those things that, um, are somewhat automated or even more automated because we can interact with a large language model that can make better decisions about how to, um, route customer requests so that they don't need to talk to a live human being.
And the great thing about AI is they can work 24 7. Um, another place we see great, great results is in video editing. Uh, a company called Runway, um, has been, uh, in the news recently because they work, um, part of the, uh, um, last year's Academy award-winning, uh, everything all everywhere, all at once.
Um, but using that kind of AI technology is saving video editors 90% of their time. And finally, uh, a more qualitative answer. Um, in medical circles, they're seeing that AI chat was rated 70 times, 79% higher in quality for their responses over that of physicians that are doing, um, you know, email or chat, um, telehealth.
So we talked about what the impact is on our workforce. Um, let's talk about what the positive impact is and how we're gonna get there. So let's start with some of the things that our, our workforce will probably need to be doing over the next couple of years, and that's, that's re-skilling.
So, um, ironically, this, this report from the World Economic Forum, uh, looked at, uh, the future of jobs in 2023. And the number one skill that, um, companies needed to look at was analytical thinking. So this is one place where artificial intelligence is far behind humans.
Um, they, they really are looking at the ability to be critical of the results that are generated from generative ai, which that means that we'll have to be a better, um, editor and less of a creator. Um, after analytical thinking, the next skill that people are looking at for, um, upcoming years is AI and big data. So that's the technical skills to interact with these, um, new AI advancements.
So all of this so far sounds a little dire, but, um, let's go back and take a little look at history and let's look at how major technology improvements have affected, um, jobs in the past. So, um, after every major, um, technology breakthrough, so in this case, we're looking at, um, data from, uh, the US Department of Labor that was analyzed by Goldman Sachs. You can see that when the electric motor came about and when the personal computer came about, um, shortly thereafter, there was a boom in productivity.
So that's the graph on the right. So you see the trough where, where it goes down a little bit, and then there's a big trough that peaks above the, um, or a big, um, peak above that trough of where we start to generate, um, more, um, more work or more productivity. And just like the electric motor had changed the way that, um, we live, we also, and the personal computer, I believe that AI will do the thing.
Same thing. And this is what happens after those productivity booms, is it leads to new inno um, uh, new innovations lead to new, new occupations. So, um, you can see here how, uh, after 1940, 60% of the jobs today, um, that exist, they didn't exist before 1940.
Uh, one of the biggest areas of growth was professionals. So people that work in offices versus working on farms, working in agrarian economy, working, um, in, in areas that were hard on our bodies and didn't allow us to enjoy, um, as much leisure as we got later in life. So there is, there is a bright side to ai, and it's not coming for our jobs as much as it should be, making our life better.
So we talked about how we, how jobs and work life is gonna be affected from these technology booms. Let's go back to, um, one of the things that comes up a lot in conversation today, and that's the legal and the ethical concerns of ai. So, you know, there's a lot of things that we've seen in sci-fi movies like the Terminator and Matrix of AI taking over.
Um, it's a fear. I don't know how realistic that fear is, but, um, there's a number of things that we need to think about now so that we never get, so that we never enter the matrix, I should say. So the big, the big one that we, we look at, um, and is preventing a lot of companies from moving forward is, um, concerns about privacy and surveillance.
So, um, they wanna make, we wanna make sure that, um, AI is, um, respecting our privacy, that when we put our, our information into the system, that that system is using the, uh, information responsibly. Just like when we go to a medical provider and hope that our healthcare is, is, um, being used response, our health information is kept private. We want the same from our ai.
Um, the other big concern is bias and discrimination. So if the language model is not trained to be fair, it could discriminate based on certain, um, ethnic age and other, um, factors. When it does things like help us approve, uh, home loans or, uh, underwr insurance.
So we need to make sure that these systems are acting in the same fair manner that we, uh, would expect from a human being. Um, the other thing beyond that is inclusivity and accessibility. So, um, we want to make sure that these, these tools are available to everyone and not just those of means or in certain groups.
So, um, we want to, just like the internet was, uh, available to everyone, we wanna make sure that AI tools are available to everyone, whether it's the large language models software that is, um, used to run these chatbots or just to access to the tools that allow us to, um, be more productive without having any, uh, um, artificial walls for them to take advantage of it. And finally, um, there's the moral and legal responsibility around these. So, for example, in the healthcare SES sector, um, ethical issues will be around for like informed consent to use the data, um, the allocation of responsibility of who's in charge of that data, um, and making sure that we adhere to regulatory things like hipaa.
So, uh, we want to make sure we're, um, our AI is guided in the way that it, um, plays within the same rules as our other technology. So what's the risk around this? I think the biggest risk, and this comes from ai, which is a global insurer.
The, the risk is that, um, in surveys that they are experts in risk and they see the future risk of AI as ranked number 17 as future risks for their customers while their customers see, um, the risk is 49th among their current risks. The, the, the takeaway here is just this is if you are not responsibly using these tools, if you are not understanding how these tools work, that they could provide a risk to your business, um, the risk of not using them, the risk of using them improperly. So, um, that, that is a trend of those who aren't educated on AI right now, is they're not sure how powerful and how fast it's moving.
So that's, that would be the big takeaway from ai if you dive into to other risks and what those risks, um, are specifically are, are really, um, there's a big distribution, but the top three risks are using AI and getting inaccurate data. So if you were using financial models and AI didn't give a good financial model, you could actually be, uh, at risk of, um, making bad business decisions. Uh, the second one, which is not unique to ai, but is, uh, a new attack vector is, uh, cybersecurity.
So are these, these artificial intelligence tools, the large language models, are they, um, hardened against attacks from, uh, bad actors? And the third is IP infringement. So understanding what the, um, IP that inform these models versus what IP is generated by these models so that you're not infringing on someone else's copyright.
Uh, right now there was a big, um, case just settled that Sarah Silvers and then, uh, brought against, uh, some of the AI creators that she said that they infringed on her ip, that case was, was thrown out. But these kinds of issues are things that we need to understand, uh, US copyright, um, office is giving guidance on what types of things we need to care about when we're talking about ip. So, um, so if everyone has access to these new tools and we all have acce, they're all accessible to all of us, um, how do you gain an, uh, competitive advantage?
Um, the way you do that is via data. So the saying today is data is the new oil, and data is what fuels machine learning, which powers ai. So these, um, the data that we have is the information we feed into a large language model to train it, to make inferences about new data.
And one of the things that we are really astounded by is things like chat GBT, which means trained by lots and lots of public data, but the thing that it doesn't have access to is your data for your enterprise. So it's either your, your logistics customer data, your product data, your in industry knowledge, all of these things that live in data, um, warehouses and data lakes behind your firewall are the kinds of things that you will be leveraging for, uh, the future. Um, on top of that, 90% of the data that we have is unstructured.
So we've used it once and then probably left it sit in these data warehouses. So, um, uh, 90% of that unstructured data gets of, of that 90% of the data that you have about, um, 58% of it has only been used once. And the reason it's only been used once is 'cause it takes a lot of, um, human power to analyze that data and make it useful.
But if we can automate that, we can extract even more value out of the data that we have today. So what we'll use that data to do is to improve, uh, large language models and be training them. That's probably the biggest and, um, change that you'll see.
Um, because this, these large language models are very good at data analysis and it allows, uh, business people to interact with them from via large, um, natural language. It'll be much easier to do data analytics than it's ever been before. Um, and we'll be able to use that data to personalize experiences, um, because it is more, more of the data is accessible and we can automate a lot of those experiences.
Um, and finally for some of us, we can actually monetize that data. So just like grocery stores make money from selling groceries, they also make money from selling data back to, um, the people that put products on the shelves. Um, we will have that same opportunity to, um, monetize data that is non-differentiating by selling it to, um, AI providers.
So here's my takeaways. The impact in general of AI is massive. It is going to change the economics of many, many businesses, hopefully for the better.
And it's gonna change those economics through improved productivity and not just the improvement of productivity of things we do today, but creating more time for us to work on things that we haven't gotten to before. Unfortunately, it will be a little bit of a bumpy road because we'll require a workforce to adapt to this new ai. And finally, the one, the companies that are, um, taking effect, uh, taking action and leveraging AI today, um, are gonna have a huge competitive advantage of those who are magnets.
So, um, that is my overview on sort of where we are today. What I'm gonna do now is I'm gonna hand this over to John Willis. John Willis is one of the godfathers of DevOps.
He's the author of, uh, great new book, um, one Edward Steming, profound Knowledge. Um, so he's gonna dive deeper into the anatomy of an LOM and talk about some of the data points that I made earlier. So take it away, John.
Thanks, mark. That was great. Uh, really plays well into, um, the presentation we've been talking about and, and working together on something I call the anatomy of an LLM.
You know, DA Mark said, data is a new oil, right? That's been, and it's been it, we've been saying that for quite a while, but now it's, it's, it's more true than it's ever been. So I wanted to walk you through, you know, how do you take the, your own data?
Mark talked about, you know, the common data you get from the models, um, you know, like the GT before and, and the, the, all the foundational models as you will. Um, but how do you take your corporate data? And, and so I'm gonna show you the techniques through the Learn way.
I learned how to do this by using a book I wrote recently and Mark s uh, Deming's journey, profound Knowledge. Um, so I'm gonna give you a little bit of an introduction. I just wanna tell you just something about Edwards Demming because, uh, as spent the last 10 years thinking about this guy.
Um, and then we'll talk about the whole data strategy idea of like what, you know, how, you know, again, what I'm gonna show you is I took the PDF of my book and I talk, walked through the choices you make as a sort of a data engineering process. And, and, uh, we, we will call it this, we'll talk about this, we call 'em, uh, retrieval augmentation or rags, or specific use of a rag, uh, a vector database. And so, just a little bit about me real quick.
I've done like 12 books, 12 startups over the years, uh, worked for, you know, sold the company, Dell sold the company to, uh, Docker early in it chef. Uh, currently I'm, uh, uh, technical, um, doing, uh, consulting and, uh, what I call fractional evangelism for certain companies. And the one I'm doing most work right now is, uh, MongoDB, particularly on their new vector database, the Atlas Vector data search.
Uh, very exciting stuff. Um, I, I happen to be a little opinionated, but I do think it's one of the best choices, um, out there for a large enterprise of the type of work that, you know, that I mainly focus on is infrastructure operations, DevOps, DevSecOps, uh, SRE. So here's my book real quickly.
Um, if you're so inclined, you might, you wanna learn about this really fascinating person, uh, you can order it, but that's all I'm gonna say about that. Um, one of the things as I, I've been, um, really playing around with the, um, degenerative effect, it was Mark who got me started in 2020. You know, he asked me if I, he showed me a tool called Jasper, and I was literally about three quarters away down with my book.
And I was like, oh my God, this is gonna be amazing, you know, and, you know, and like I said, mark tell to write something about Dr. Demming. And he wrote like a paragraph and I was like, oh, I went home and I found out this is early GPT-3, right?
5 or chat GPT. And I realized that we weren't really calling hallucinations yet, but the, the, the amount of fact-checking I had to do it was really a terrible research tool. And then three five came out, I rechecked it.
Still not a great research. If you're trying to do research on like Abraham Lincoln or you're wanna, you wanna write a a 10th grade report, you know, on the Constitution it's gonna do a fabulous job. But if you're gonna do something, I'm new and creative, like a book that I'm creating that, that I'm trying to create new ideas on, without feeding the right corpus data, it's going to hallucinate, right?
And that's the, and so, um, so I started learning and playing, but it wasn't until I learned that you could actually use these, you could take your own corpus data, vectorize it really, they call 'em vector databases. And now I can get things like chat, GPT or other foundational models to answer questions correctly based on the data that I, I preload and, and we'll talk all about that. But so early on, I've been following this and I, I wrote an article for Techstrong called the, the Rise of Shadow ai.
You know, if you remember the, the sort of the shadow IT problem, right? And, you know, for as being an infrastructure operations, DevOps, DevSecOps, whatever you wanna call it, you know, I think a lot about like that world. In fact, my running joke over the years is, you know, as we get into these major transitions like cloud or going before that distributed computing and like operations people are like cicadas, right?
We go to sleep for like 17 years. The world gets so screwed up, we wake up and we have to clean up the mess. And so what I'm trying to do now is, is set the stage for can we at least try to make it less messy because like all the things Mark pointed out, but like the underlying infrastructure tax or technical debt, we're gonna pay, you know, you're gonna have all these different implementations and this business unit is gonna do LAMA two.
This business is gonna use bar, this one is gonna use line chain with this, and this one's gonna use three different vector databases. And, you know, the poor infrastructure operations, DevOps, whatever, you know, I'll just call 'em INO people are going to have to support all this. So we've gotta get our head around it much sooner than we did in prior, um, you know, changes like cloud or even going way back, I'm pretty old distributed computing.
So I've written a number of articles so you can find those on text Strong about, just as I'm learning, I'm trying to write articles about what I'm learning. So, uh, but I can tell you, I used my book. What I want to do is show you now what I've learned about how to use retrieval augmentation, specifically something called a vector database.
So I, I took the PDF of my book, um, that I recently we published, and I loaded into a vector database, and I'll show you the whole process here. And so one of the things you have to understand, you know, people talk about how do you train data? And I think it, it, the conversation gets a little skewed because in a sense, when you're using vector databases, you're kind of training it, but you're not training in the way that when you talk about training a foundational model, like GBT four has been trained by open ai, they spent ridiculous amount of computing resources and time and effort and, and just vast amount of generic data from everything that's known on the internet, basically.
And then slam it into this, this, um, model. And it's why it answers so many incredible questions and answers and creates. And, but like, but it also hallucinates, right?
We know this, right? Um, so there are ways to train your model, but like, that's like the old data scientists way, right? You know, like the, like the reason, um, Chet GBT is so, uh, popular is now anybody, my aunt, my uncle, you know, my, your kids can literally use this advanced AI techniques that when they don't have to know anything about math or data science or artificial intelligence.
So there's the foundational models, the, um, again, they're, um, they don't know your data. Uh, they're not up to date, they're point in time. So when they're trained, they become the, the sort of that model is timestamped, basically of one inch train.
And like I said, they can hallucinate. So this idea is this idea of, uh, retrieval augmentation or rags, um, becomes an additional data strategy to get your own data. This is like if I have my own standard procedures data, or I've got risk control, or I have incident data, or you can see my mind thinks about the way infrastructure and operations work, but any sort of corpus of data.
And what you basically wind up doing is take that data and you, it's called, uh, embeddings or vectorizing it. And what you're really doing is saying, Hey, I'm gonna use the GPT-4 foundational model. I'm gonna use LAMA two, whichever one you choose as my sort of foundational model, but I'm gonna then create my own data in the same format as that foundational model.
And I can merge, you know, this, this isn't a scientific or an AI term, but merge 'em together such that when I'm asking my prompts and queries, I'm actually seeing 'em both together. But I'm see, but the questions and the queries, the prompts are going against my data first. So, um, so like, and I'll show you this example of the, in my book, right?
Uh, if I just start asking questions about Dr. Deming against g PT four, I may get one pa like if I ask, gimme a bi, gimme a two page biography of Dr. Deming, the first couple of paragraphs will be reasonably accurate.
The third paragraph will be start going really south, and then it, it'll start saying crazy stuff, right? Because there's not as much known information, but if I prefix it with a biography and, and then, and the biography is reasonably accurate, then uh, then I will get more accuracy from the questions, right? And so the, the techniques you use to get that data in are, and we, we don't have to go too deep here, but I I do do in my longer version, I have 90 minute versions of this out there in different places.
But you literally wind up taking your data and you split it up in a certain way. And that becomes important how you split it up. Some of you have probably played around with something like Claude and has a PDF load and you put the PDF in CLO and it magically throws up a prompt and you start asking questions and like, oh my goodness.
Now you can do that. But if you really want higher efficacy on your data output and your prompts, you're gonna want to have to really think about the data engineering that you're gonna use. So I can just load in my book and I'll show you, or I can take data strategies of how to cut up and split my data, for example.
I might want to split it up by chapters, by subheadings, and that's gonna cluster the information I ask more accurate. And then, um, and, and so this is what they call vector embedding strategies. And then the metadata, this is another really interesting thing you need think about, is not all the data that you're gonna wanna run against is gonna, you, you're gonna want it in the vector or the LLM if you will, right?
The sort of the combination of the foundation model in your vector database. Because in some cases you want that metadata to be separate in maybe a document object or A-J-S-O-N object, you know? So like, for example, I might wanna ask questions against, um, all the incident data that's been collected last month, right?
And so, but what I might wanna do is separate like the, the status, the, the priority. Was it a P one, P two, who was the author? All those things I might wanna just keep in JSON objects.
And now as I search across, this is again, why I like Margo DB so much 'cause it's kind of built in. They actually take their vector data and the JSON objects and they're in the same, same object. So now I can say, give me all the objects that have John Willis that are P ones and P twos and only happen incidents that only happen on weekends, right?
Um, if those, that metadata is actually in the JSON document objects, then I would get back all the, um, the vectors, basically the, the, the sort of LLM data that was associated with that. And that's a much cleaner search than trying to ask that question directly against the LLM because in a vector data is not, like, it's not a binary search, it's a nearest neighbor. And that gets into the complexity of vectors and all.
But, and then, um, you know, how I want to chain the query. And then, um, an interesting thing too is when we talk about concerns of, you know, sort of the, the data, are we giving the right answers? Are we reducing hallucinations?
There's a whole new, um, definition of observability. So things like observability in your classic, um, you know, how you think about, um, you know, performance and, and latency and, you know, CPU and memory and all that, that's still here. But now there's a whole nother section of what is the, um, what they would call the, um, the evaluation.
Can I put a percentage that the data is going to answer these questions, 98% accuracy, and I wanna make sure as I'm updating the data, it doesn't fall below or I want to increase the, the accuracy. And so tracing. Um, and so this idea of a vectorized data is you basically go through a tech splitter and then, um, you literally take that, split a deck, and then you pick an embedding that would be associated with the foundational model you, you want to use.
So that, and there are different types of embeddings for different types of use cases. We'll go through that. And then you have your prompts and your queries, and it's important to know, you know, again, i I, I think there's a fine line between not having to become a data scientist and knowing enough about the, the data science of all this gen ai, um, so that you can at least be sort of effective and intelligent on what you're doing.
Again, um, a vector, you don't have to know the math behind vectorizing data. There are tools that like, that do this very well for you. Lang chain is one, there's many more.
Um, but you do need to understand like how data gets placed in a vector, and it actually gets placed in this three tier structure based on indexes. Just sort of like how you can shard the data. Your chunks are basically locality of words in the vector, and then the tokens are actually how the words get stored.
And so what happens is a vector is a multi-dimensional and sometimes 512 dimensions. So you think about two dimensions, three dimensions, 512 dimensions, right? It's like you don't want to know the math unless you're really a mathematician.
But like, what happens is the words get placed in a locale. Like, like here's an example. Like newspaper and magazine and article are, are relatively close to in this, in this case, um, you know, like Margo DB or Pine Cone or a couple of the defaults are usually 512 dimensions, right?
So, um, the words are close, but you notice that apple is pretty far from newspaper. So when you're writing in a sentence and, and it's, it's a, it's basically an auto complete on steroids. It's finding the next nearest word and the highest probability of the next word, which turns into the ne the next sentence, which turns into the next paragraph.
And it's all about, well, you know, um, in, in data science star analogy, um, approximate near neighbor, right? Which is like how close, it's really the, it's a near go one level deeper. They're actually started floating point numbers and it really just calculates the distance between the floating point numbers that have been, um, vectorized in a, in the, in some cases a 512 dimension.
And so this is just some example code of how you would basically set the embeddings. Um, and I'm sorry, I I take that back. This is how you, you split.
So there's a couple of strategies for splitting data. So what you're doing is taking my PDF and I'm splitting it up. And one is I just want to chunk it every turn to physics characters.
Uh, the reason you do an overlap is so you don't split a sentence in the middle of it. And, and that's the, that's a default. And I always say, if you're getting started with this stuff, take something that you know and just take the default of how to chunk it and then see what the output looks like.
Run some prompts against it, run your own version of a chat GP against the data that you've just loaded. And like I said, I did it in my book, if you wanna get more advanced, there's what they call a recursive chunking strategy. And this is a strategy where, um, where basically you, um, you basically, it, it does the logic of never splitting up a sentence, a paragraph or a page.
So the logic, you can see the code, it's all open source, right? And then one I prefer most is converting my data to markdown. Now, I didn't do that for this particular scenario, but convert it to markdown and then what will happen, it'll split by head of one head or two head of three.
And now you can have a really tight clustering of your data. You know, in my book, if I just load the PDF it, by default, it's just gonna go, I, I think I used, um, a thousand with a 200 overlay, right? Um, that's sort of, I, you get lucky if I wanna increase the efficacy of the prompts or the evaluation of the prompts, I could convert all my, um, chapter headers, subheaders, um, even citations into different markdowns, and it would automatically create that data in locations.
And I would see if I ran the prompts against that version, I have not in this example, um, the, uh, prompt evaluations are much higher in percentage. And this is example of like using a thousand chunk size of 200 valor in my book, you know, again, not fortunately to read my book, but I can look at this quickly and say, Hey, it did a reasonably okay job of splitting out the data. So these are gonna be vectorized by these sort of paragraphs or these chunks.
And so, um, so chances are when I ask questions, it's gonna answer 'em in these sort of clusters of finding the next word, the next sentence. Uh, but if I used like a 500, this would've been terrible. See how tight it chaptered, it was like, there's a story about Buffalo bill there, it would be spread all over the place.
I, uh, the, the, what I would call efficacy or what the data scientists would call evaluations would be very percentage lower. So what my point is, you really want to take your data and you wanna play with this. Um, you know, and then there's a whole discussion about the embedding model you want to use for GPT, uh, GPT-4, GTB three five.
It's, uh, the default is something called text embedding ada. But, um, but again, there's a lot of strategies here for embeddings. You know, whether you wanna do search, focus, cluster focus, recommendation engines, anomaly.
So again, there's a lot here to learn, um, and or classification. And this is just interesting, um, for classification, this is something I did early on when I was learning, is I went ahead and I, if you, there, there's some real power in this. Like if I gave it a paragraph with three people in the paragraph and at the end of the paragraph I said, Jake was a baseball player, Amelia was an artist, and Ryan was a caller.
If I just send that in the prompt. And, and that's now in sort of the memory of the sort of LLM, um, now I can give it another paragraph and notice down here the answer without me telling it anything, it will tell me, Ethan's a football player. Olivia is a doctor, Lucas is a carpenter.
That's, that's, that's just the tip of the iceberg of how powerful this stuff is. Um, I talked about metadata, uh, you know, the, the, the, the, uh, here's some example code. We'll have the slides, how you can load in any PDF or any file, really any text file.
And by the way, it's not just PDFs. I mean there are, if you start using Lang chain, you can load in Slack channel Jira data. I mean it's got connect, it's got, I I know probably 70, 80, probably more than that now.
Different connectors for load, different loader data loading types. And then here's how you can run the query. This is all out on, um, MongoDB website, um, under the Atlas, uh, vector search.
Um, but this is how you extract it. So you could load your data in and you can use it over and over. Um, and then, um, I'll go real quickly here.
We wanted to keep this short, but um, also, you know, think about this as sort of a three stage process. So your ingress is how you get your data in and how you need the data engineer. The, I talk about chunking, embedding strategies.
And then you want to think about how do you process the data from a prompt perspective. Like when you're gonna ask your questions, there are different, um, and it really comes down to do you want speed or accuracy? So sometimes you want higher speed, less accuracy.
Sometimes you want, um, lower speed, high accuracy, right? And so these are different, um, what they call, um, the, the stuff refine or map reduce. And then I just show examples here of what kind of results you're going.
And this is interesting because, um, this is a really good example of like where I'm going for speed. I'm getting terrible accuracy. 'cause I ask about my book.
And, and the, the stuff example is, is is basically an example of I want accuracy. Don't worry about speed. I'm totally overloading the scientific description of stuff, but, and, and it's pretty accurate.
It's like the book is about demming, but then, uh, when I ask for refine or um, actually refine it actually sort of mutates the answer of what my book is really about. Um, and then like I said, what I wanted to tell you last is sort of the output. The third, like if there's a three, there's ingress, there's process, which is sort of the stuff refine how you want to get your prompts, your prompt engineering.
And then finally is sort of observability and tracing. And in this case, um, you know, MongoDB, all the tools have pretty good. Um, the only thing, again, another thing about MongoDB, um, you know, transparency, I said I am working for them being paid as a consultant, but I do think their enterprise, they've got many years experience in the enterprise.
So you'll see right out the bat better monitoring, better graphs just because they've been in this business for quite a while, uh, doing, you know, all the other stuff they do. Great. Um, and then for L Chain, this is relatively new, probably only a couple of months old, something called Lang Smith where they add a tracer.
And one of the things when you're doing prompt and prompt engineering, it's very difficult to find out why you've got the answer that you don't agree with, right? And some of it might be hallucination, but some of it like just, and so, um, if you're losing Lang chain, um, you can turn on this tracing option and you can see exactly very detailed what it's doing, how it's going back, you know, and you really can, and it also has a whole section of evaluations where you can sort of say, I want to benchmark the output. You know, I don't want harmfulness, I don't want, um, I don't want hallucinations or I wanna minimize hallucination.
I wanna make sure that I'm continually getting a certain high percentage of non hallucinations, non-harmful information. And then this is an interesting one too, which is, um, I've, I've told their ex, uh, Cheff people, I, one part, my background was working with Chef, so I have a place my heart for anybody, ex chef, um, but Y labs and they're focusing completely on this new form of observability. Very interesting.
Uh, I know Patrick Debar has been doing some work with him. I've been trying to track, you know, what his, you know, he, he seems, you know, last time I talked to him he seemed to be pretty high on this. Um, and, uh, and that's, that's it.
And I've got a bunch of resources here on examples and, uh, and he goes, I know that was a lot of information, a very short period. I do have a, you know, an hour and a half version of this presentation where I spent a lot more time. But I just wanted to give you a feel for, you know, mark did a great job of sort of presenting the, the, you know, data is the new oil.
Um, but, and, and that we, there's gonna be some bumpy roads and we gotta roll up our sleeves. So even though there's a lot of what we would think is magic here, we gotta decouple the magic to get it, to do the things we want it to do, right? And then we have the best of all worlds.
So anyway, thank you so much. Thank you, mark. Thank you Textron.





