Techstrong Gang – April 30, 2024
Alan Shimel, Amanda Razani and John Willis discuss the rollout of small language models by Microsoft and Apple, as well as the benefits of smaller models to the enterprise, before contemplating the impact AI is having on networking. Additionally, they talk about OpenTelemetry and whether or not it has become a de facto DevOps standard.
Transcript
Hey, hey, happy Tuesday. We've got a little bit of ai, a little bit open telemetry, a little bit of this, a little bit of that. You're watching Textron Gang.
Hey everyone, it's Alan Shimel, as I said in the opening, we gotta some this, some that at a venue today. Um, I'm joined. Well, no one here locally in our tech drunk headquarters in Boca Raton.
Everyone's on the road. I'll be on the road soon too. But we're joined today from, well from their homes.
Let me introduce who we have, first of all, from, uh, Alabama. He's, he's only home for a short time. He's probably about to hit the road too.
Our resident expert on cloud DevOps, ai and everything else in between my friend John Willis. Hey John. Welcome.
Yeah, I guess a jack of all or master at none. Well, You will, well, that depends on your point of view. There you go.
Absolutely. You know, I'll take a little bit of Jack over John. There you go, John.
Then a master of probably a lot of other people. Um, and then joining us from San Angelo, Texas near Abilene, Texas, our own editor for Textron AI and Digital CXO, Amanda Ani. Hey, Amanda, how are you?
Hello. Good. Glad to be Here.
Good to have you on. So, as I I said in the opening, we got a little bit of this and that, as they say, down in, in the islands. Um, we are gonna start off with some AI kind of interesting news.
And, and this comes from an article up on tech strong ai, uh, apple and Microsoft pushing LLMs to the edge. And, and truthfully, it's more s SLMs than LLMs, but, you know, uh, we, we spoke about this on a, an earlier, uh, Textron gang, I think it was last week, about one of the ways of combating the huge energy cost of, of, you know, doing AI in the cloud or in data centers is pushing the AI to the local devices to the edges, if you will. But of course, do the local devices have the horsepower to run the ais?
Seems like Apple and Microsoft are, are trying to enable that. John, this you get rather technical. Why don't, what do you think?
Can we, do we foresee iPhones running AI right on the phone? Yeah. You know, I, I think, um, I mean the iPhone thing is interesting, um, but I, I think more about like the, the real sort of solutions will be edge based, right?
Where we can put models in, um, shipping containers next to stadiums, right? Like that kind of stuff that, you know, where, you know, where you can put shipping containers near, um, you know, near, um, hotspots and, uh, you know, so the reel the edge, you know, obviously everybody's excited about iPhones and, and, uh, sort of some form of a chat GPT on iPhone. And yes, that's but more interesting is, you know, how do you use this technology for all the other things that are going on?
You know, like, you know, there's always talk about sports betting in a game while they're playing baseball. You know, so you can literally bet on a pitch or, right? Like, and so the idea without AI is really just to put sort of shipping containers near the stadium.
Grit, a lot of power that you'd need for 30 or 40,000 people making bets if you wanted to integrate AI and stuff like that. You know, you'd need to put on the edge some models, right? And so, so I think that, you know, the same thing applies like small language models that may work on phones or may work in shipping containers that are unmanaged that don't require, um, you know, like the, the complex management of GPUs, right?
You know, they Put the pitching clock in. I don't wanna see betting on pitches. They're, they're under the pitches are under the gun enough as it is.
I hear you. No, but you know, like, I mean, no, I I, you do this with your kids, I'm sure when you, when when they were younger sporting you'd like, you'd bet. Like, is he gonna, you know, who's gonna hit All baseball curve?
No. Yeah, yeah. All that stuff.
So, you know, Yeah. My son, my older one was a baseball freak, so we would have those discussions in between. Oh, yeah.
No, I, I, I've had not to go too far in this conversation, but I think at one time my youngest son was down about like five grand. I would never would pay. He kept going double nothing.
Yeah, it's easy when you don't have to pay that. Yeah. Yeah.
Well, you don't really have to pay. But, um, but yeah, no, grownups were, and that's just one example, autonomous vehicles. I mean, just so many different things that require a lot of complex logic on the edge.
And I think that that whole idea was small language model. Um, and you know, the one thing I too, like, I, I think, you know, I feel like I'm broken record, but I think people conflate the training versus the inference. We, we talked about this last time, right?
And, and, and, and they're both distinct, but it's inference where, um, you know, where, you know, when a model's built, it's all about inference, right? It's not about training. Um, so, so again, I think the, you know, the, the running these small language models on the edge, wherever hotspots different sort of locations, um, will, you know, not require, you know, large farms of even GPUs for inference.
And then you get into the cooling, you get into the, uh, one, one other thing I think is interesting about, uh, small language models is I think the, you know, there's two ways to augment this. The, we talk about retrieval augmentation generation, right? There's vector databases, there's some other forms.
The, the ilms work really well for specific tasks. So they're a good marriage with retrieval augmentation. 'cause now what you do is you take a very specific data, you combine it with, you know, pretty good, um, you know, inference engines, you know, some of these engines are, you know, a dwarf compared to like GPT-4 or the really paid ones, but they have really big context windows.
They have, you know, lots of parameters, right? Which is sort of the neural network of 'em. So you can, you know, if you're sort of not wanting to ask, you know, questions about Abraham Lincoln and all this stuff, and you want just supply your own corporate data, like a rag, you know, like in vector database, then you can literally get very specific data cleaner, isolated.
You don't have to run it, you know, with like massive GPUs. You can actually, there's a security thing. I, it's easier to secure small language models.
Mm-Hmm. Then I think, you know, this last thing I, I wanna hear Amanda's thoughts too, but, um, the, the, the one last thing is I think, um, mortals will be able to do very soon training models. So right now, you really have to be an ai, you know, click really, really, you know, I've, I've had like my friend Jo Josephine X you know, try to explain it to me like 10 times now about how these things work.
And they're incredibly complex. You literally have to take three or four years of learning and training. But there are tools now that are starting to make training models easier for mere humans and, uh, than fireworks.
Uh, there's A-D-S-P-Y, uh, new model outta Stanford. And so as that keeps evolving where like reasonably smart humans can actually train models, again, SLMs become real powerful, right? Because now the idea of training a very specific model for specific tasks that are sitting on the edge, that may be related to utility work or, anyway, um, yeah, I think the SMS are really interesting.
Yeah, I'm seeing, I'm seeing so many more companies, we've posted a lot of articles recently that are going the way of these small learning models, um, for various reasons. Like you said, specific tasks. They don't need these large language models that they have to hone down to what they actually need as a company.
So, um, I think we're gonna see more of that. And then I do have questions though, on the heat side that you mentioned. I know that that's been mentioned less, less compute power and, um, reduce the heat, but is it really reducing the overall heat of the world?
I mean, you're just taking this large one that's producing tons of heat, you're dividing it up so you have a little less heat, but you're just spreading it out. Is it really, um, leaning itself towards a smaller carbon footprint? Um, I have questions.
Yeah. I think it is though, right? Because like, if, if I can do very specific inference, again, the training and, and the inference are two different things, right?
Somebody's gotta train even the small language model, and that's gonna have, uh, you know, a, a footprint, 'cause any training not as big as a GPT-4. But if, um, for example, if you could run small language models, would sort of vectorize data that's very specific to say, utility work or very, you know, or, or edge healthcare, right? Um, I, I'm not sure, and I, you know, I'll go out on a limb and say, I'm not sure you're gonna need GPUs.
'cause that, that the whole point of one is the points of those small language model is, um, sort of, again, I'll say at inference time, the ideal is that you don't need massive compute power. So, but, but I, I think your point is like, if all this is built and we're spreading it out and we're using compute and electricity in more places, then yeah. Does it really reduce the carbon footprint?
Probably not. Yeah, that was my question, You know? Yeah.
I'll give you an analogy though, too. Even doing this, you know, uh, next strong gang today, the two of you are remote. I'm here in the office, so they record me in studio, and I'm, I'm actually sitting at a table with a green screen in a, one of our studio B here at Textron.
It's a really big kinda room, and it's set up as a, as a recording studio. Well, I thought that was a fancy, uh, uh, desk and all that stuff. No, it's all fake, right?
It's all fake. But you know, that kid's done here and we have a whole bunch of computers and, and masking and everything getting done. The two of you are remote, and we, we basically take your feed, your feed gets uploaded over the net and it gets recorded in the cloud, but it makes for problems, this latency, we gotta make sure your audio matches your video and you, you know, the words match the movement of your lips and all that stuff.
We've been moving now replacing our present recording software to, you know, something that's more custom made for this kind of situation. And in that situation, it actually records you locally on your computer taking the network latency out of the equation and, and records you. Because the fact of the matter is, you know, especially if you're on max with some of these new arms and everything else, you know, the process is you got enough horsepower to do that locally on your computer and lose the latency that we get from the network.
Uh, I, I think somewhat, you know, it doesn't the same thing apply here to, to AI computation if we can somehow run it locally like that. I'm assuming processes continue to, you know, add more not towards lore kind of stuff, but continue to beef up. Yeah, I mean, it, it comes down to like, you know, the Simon Wardley Hass been screaming about cloud from, you know, beginning of cloud time, right?
About like the, the footprint of cloud and, you know, the electricity and the power and all that stuff, right? And, you know, so I mean, and I, I think what, you know, if you take the big, the Microsoft to Google's, the Amazon's, Facebooks, right? And, you know, like there is a lot of like footprint going on there, right?
Like, and then more we move to the cloud, the like, so, so I think at the end of the day, um, I think back to what, what Amanda's talking about, I don't, I don't know that like, I mean, the real power, the real issue right now, like the, you know, the game was the carbon footprint, and I'm not an expert in this anymore, but the carbon footprint like, started to get really interesting with cloud. Then you have blockchain, your crypto, right? And that sort of like really upped the game.
And now you've got, you know, the ai, which has like moved it to a whole new level. So like the things that we do between sort of bandwidth and like, I think dwarfs and compare the real concerns are like, if, you know, we're on this freight train of AI and GPUs, you know, the, um, you know, and, uh, there are people, you know, I've done some podcasts with these sort of outliers who think that there's a different form of, uh, neural networks that don't, aren't gonna require, in fact, they call the GPU processing and intellectual laziness. In other words, in other words, we've driven this, like we were forced down this path of compute and what Was the path of least resistance, let's call It.
Exactly. And, you know, they're like, the whole thing of matrix multiplication, all the things that happen in these sort of neural networks, like, you know, um, these people will stepped back and said, like, could we do this a different way where we weren't forced down what they call an intellectual laziness or, you know, through GPUs and like, everything has to be GPU based. I think that's sort of interesting.
I don't know where that winds up, you know, like, are they swimming upstream or, you know? Well, I, I think that's the way of things, John, right? You got the revolution, which is gen AI coming to this, and now we'll see evolution, right?
Feedback loops continuous, it's almost DevOps in ai, right? Feedback loops, iteration, reiteration, reiterate again, you know, making it more efficient, looking at different alternatives that, that, you know, can get more bang for the buck, but all with under the guise of, of this ai. So, um, yeah, I, I fully expect stuff like that to happen.
I don't know if it'll mean that we're gonna find a radically different approach to neural networks that don't involve GPUs, but frankly, who knows? I mean, you know, yeah. Yeah.
Smart people will be working on That. We forced to do that. We, we might be forced to do that, right?
So, yep. Yeah. Or find, you know, make fusion real and, and then we don't worry about it.
Um, all right. You know, speaking of, of retina work latency though, and whether you could do stuff on the edge on the device, or in the data center in cloud, uh, we're gonna be back in just a minute talking about AI and networking and how AI's changing networking, but we're gonna take a break first here on Tech Drunk Gang. We'll be right back.
Cloud native now is the web's leading resource for the growing cloud native ecosystem. com is your destination for news, thought leadership, features and webinars on cloud native architecture, Kubernetes, serverless, cloud native application development, microservices, service mesh, cloud native security, and more. Stay on the cutting edge of modern application development at Cloud Native.
Now. Hey, we're back here on Textron Gang. Uh, our next block is actually around an article that John wrote on Textron ai, and it's about how AI is changing networking.
You know, John, I couldn't think of a better person to write this article. Given your history and background. Um, how is AI changing networking?
What's the point of the article? Kick it off. Yeah, I think, uh, you know, the, the, there's a bunch of interesting things.
Uh, you know, um, I, um, I, I'm working with a company, so, um, you know, Mike Dekin, who is, uh, the creator of a CI, which basically was Cisco's, uh, SDN, right? He literally did that as a spin out. Uh, you know, the, they, they brought it back into Cisco.
It basically became the core of their SDN solution. He's working with a, a, a, uh, uh, you know, uh, an infrastructure company where they've created this sort of holistic, um, you could call it this idea of composable infrastructure, right? Like one API, because the, the, the problem you have, there's a number of problems that happen here with, with a ai, right?
Uh, particularly if you're gonna run your own infrastructure, you're gonna run your own GPUs. One is you have, uh, this concept of GPUs starving. We can come back to that.
You have heat management, uh, and then like in the, you know, you have complexity in just configuring GPUs and not, and all the sort of like, the, the, like, how do you wanna do memory sharing? So it gets incredibly, it, it's sort of exponentially complex to manage your own infrastructure that is not, not only, uh, you know, composable like that means like ai, AI and API to manage this stuff, right? Instead of like, somebody working on a network, someone working on storage, someone working on the compute, and then somebody trying to DevOps all that, right?
As opposed to can we have, like Mike says, you know, calls it a single API to manage it all right? And, and, but that's even hard with just, um, you know, composable infrastructure in Gen general, then you add GPUs alongside CPUs with shared memory, right? And even the people who configure this stuff, who are like network gurus have been configuring, they'll tell you when you throw the GPUs in the mix, it gets even more complex and then add more complexity.
People are finding the thing running things like Kubernetes or KS three or, or some type of rancher under the covers to manage that is a good idea, right? Because you can use, uh, like the, the whole CRD model to do that. And then to combine that with now it just gets like Crazy.
So, no, I mean, multi-threaded, you know, CIC And then add in security and access and network security, right? Like, like, think about the complexity. So I, like I said, a little shout out to a company and Morgan called Hedgehog, which, you know, Mike de Wilkins part of.
And so I, I've been working with them and, and like, this is crazy too. They actually keep some of their stuff at a service provider that does, uh, liquid immersion where they put the actual, all the componentries in, like vegetable oil. I know.
I'm like, like, and it's like, this is crazy stuff. But, but at the same time, I work with a group that I, that every year I go to Vail and I, I do, I speak to, um, different, um, you know, uh, large cap technology investors, right? It's been something I've been doing for 10 years.
And I was in, um, talking to one of the guys that works there, and he was talking about like the complexities of heat management, heat transfer management, and that, well, he said like one of the more interesting degrees right now that like the big, you know, the Facebooks, the, the Googles they can't hire enough is, is heat transfer engineers. So like an old throwback. And so I did a bunch of research and, and, uh, you know, I found like there's this, all this interesting stuff about GPU starving, um, the, the, the management of heat, um, the GPU starving, you know, is part, like how do you do sort of memory sharing?
So that bringing back like the technologies like PCE and, and, you know, Nvidia has the N NV link and, um, like, so there's just a whole bunch of stuff here. And, you know, sometimes I just wanna learn stuff. So I, you know, and I knew a little about SDN and I was working about, you know, software defined networking Yep.
Or one API. So I just started doing some research on it, and I, you know, I asked Mike, you know, would he be interested in an article? He is like, gimme as many as you got.
And, you know, and it, it kind of fit. 'cause I'm trying to work with this, this vendor to see, you know, if you were gonna build your own, um, infrastructure with GPUs, you know, like you're gonna need something that sort of masks this incredible complexity. And it's more than just saying DevOps in it or, or, you know, um, you know, just using Terraform to do it.
'cause I mean, you can terraform a lot of things, but, but at some point when you're during a very complex infrastructure, like a Kubernetes cluster that's managing network compute and storage through APIs that have GPUs and memory sharing between CP, like, on and on and on, um, it, it's a fascinating, uh, on if, if, if there is a need to run it. This ties in with the small language model stuff I hadn't really thought about that, is if there is a need, and I think there will be like everything else we've seen to have on-prem implementations of these AI models, which I think we all probably agree that's going to happen, uh, for security for all these type of things, then like, this is an interesting space to be tracking. There is, sure.
Is Amanda, wondering if you have any thought, I know this is kind of down in the weeds, but thoughts on this one? Well, uh, my thoughts are, um, you know, we talk about utilizing these GPUs, um, and of, uh, of course they can process a lot more. Um, they have a lot more power, but, um, aren't they gonna also generate a lot more heat as compared to CPUs, which still do a lot of tasks better than GPUs in some instances?
Yeah, no, the GPUs, I mean, you know, the, the, definitely the GPUs up, you know, up the sort of power grid. I mean, this is the, and, and then the bigger problem too is the GPU U starving is that the people, even in the most complex environments, they're getting, uh, low utilization rates on GPUs and they call GPU U starving because the, the sort of the data from CPU to memory to, um, to GPU is just, you know, like the technology is emerging. And that's why these sort of shared memory models are becoming really interesting.
But then you have to manage that, and then you have to have all this sort of, like, so it's back, but, but it is, you know, from my understanding, it is the GPUs that are causing the sort of like the, the, the, so the outcry of the foot, the, the carbon footprint, you know, maybe lower C crisis, you know, maybe capital C crisis. Not sure, but, but, um, but yeah, no, and, and that is, but that's why the, um, that that the thing that I'm, um, you know, uh, it looks like I'm gonna get access to a ton of service to these people, and they actually can run it in a, um, in, in a sort of a liquid immersed environment. So I actually might have a, uh, you know, like, uh, in fact, uh, I'm probably gonna invite you and your team if you wanna come.
The plan is that DevOps State Birmingham, I would love to do a Full of, and the guy who runs it might actually bring one of these immersion tanks, and we'll have it in the lobby, and we'll do a demo with it, like, knock on wood. And then it's gonna be, that'll be a big deal if we can pull it, actually run, um, run a full blown GP structure with some training examples, completely isolated. Um, and, and, and, um, you know, I mean, there are, I, you know, again, I'm not an expert in any of this stuff.
Like I said earlier, Jack of all expert, none. But I do, I, people have told me that sometimes the actual chips start eroding. So there's still sort of, but that's over time, you know, it's like there, so this isn't a completely solved problem, but it's a really interesting problem back Amanda to like, maybe, you know, and they've been doing this, um, since crypto, right?
Like the, the Bitcoin monitoring, a lot of those, um, that actually, I don't know if that's when it started, but that's when it got real popular. The idea that that like, um, small startups that were trying to do, um, Bitcoin mining as a sort of service provider couldn't get the power that they needed. And so therefore they started experiment with this liquid immersion type scenarios.
'cause that Fascinating stuff. That'd be really cool to see. Yeah, I, I like, it'd be a miracle if we pull it off, but like, You always take the vegetable oil when you're done with it and use it as biodiesel.
Yeah. I Mean, or a toilet or a frozen Turkey in it or something, I don't know yet. Interesting.
Keep us posted on that one, John. Yeah. All right.
Hey, we're gonna take another break here on the gang. We're gonna be back, we're moving a little bit away from ai, I think, as we talk about Open telemetry and, uh, observability DevOps, et cetera. Stay tuned.
You're on Textron Gang. Hey, Alan Shimmel back here on Text Trunk Gang. ai articles.
com, talking about, uh, a recent announcement from Company Embrace about adding, uh, support for open telemetry to, uh, mobile applications. But the bigger issue is, has Open Telemetry become the defacto standard for observability in DevOps and, and in DevOps as well? Um, it's sure.
Seems like that, right? Open Telemetry is the second largest open source project for the CNCF behind Kubernetes itself, followed actually in third place by Argo. But, um, you know, there are so many companies that are in the observability space, for instance, that when you look under the covers, they're just using an Open Telemetry agent, you know, porting all that data to open telemetry and then building their interface and so forth from there.
And that's not a new business model, right? We've seen this, I've seen it in security back in the early two thousands when most IDs, ipss actually were running Snort under the covers. And, and many, many of the vulnerabilities, uh, solutions scanning and remediation solutions were running Nessus scans under the covers.
Um, it seems to be, that's the same thing we're seeing here with Open Telemetry. John Wonder what you're saying? Yeah.
I think hotel is, is the deal. I mean, it's, it's the standard, you know, at least, you know what I would call a standard. I mean, I, I think, um, I, I don't think if you're in the observability game, you know, I did talk to an interesting company the other day that, uh, they're, they're trying to do their own sort of AI thing, and they're sort of not using hotel, but that they're, that's a, that's a freak occurrence, right?
But, you know, like hotel is the standard. Um, you know, and I think, you know, like it was, there was, I've been tracking this for quite a while in that even before it became open telemetry, you know, like you had the original, um, you know, Adrian Co No, uh, Adrian Cole, right? He was like down, going back, he was a J Cloud.
So guy, he, when he was a Twitter, he was one of the main committers for, uh, Zipkin Open Zipkin, right? Yep. Which was that whole idea when the service mesh, it was sort of not exactly what the service mesh, but it, the service mesh really exposed the capability to do distributed tracing.
Netflix was doing it. So this idea that you could literally run these sort of, especially at microservices, you could run these traces, like follow the thread, like almost, you know, breadcrumbs of things that were going on in this, did this, this, and, you know, and then, you know, then you had, um, Yeager, which then became a big part of the service mesh Sure did. And then, you know, they combined that in open tracing.
And then at some point when they sort of talked about open telemetry, I think the whole world converged in that, you know, they called the three pillars, like traces, metrics, and logs, right? Um, and so, so now like, okay, this is what we do. We do, uh, tracing, we do metrics, and we do logs.
What else is there? Um, and, and, and it's a standard, it is an API set. So anybody, you know, anybody who's doing anything observability is, you know, uh, unless they believe they have a secret sauce, they're going to support hotel.
Um, and it, you know, I think their secret sauce is like how they store the data, how they visualize the data, you know, like a versus, um, but yeah, I think it, it is the way you do things. And the only other interesting thing is like you thought you were gonna get outta the AI discussion. Now, I'm sorry, but what's interesting is, um, I was working on this article about long calling postmodern, uh, uh, observability, right?
Like, think about the original observability came from instrumentation and, you know, the original, uh, instrument control systems and you know, sort of, and then, you know, um, uh, what, so cherry majors sort of really exposed this idea for monitoring, you know, and, and really sort of coined the modern version of that. But that's infrastructure and applications. And, and then really what's sort of hit it under the cover was, um, data, data ops, and like, how do you do it, you know, like sort of observability of data, which is sort of different.
Yeah. Like I'm, I'm looking at the schemas. I'm looking at, like, other than looking at the database performance, I'm looking at things like what changes in the data and what changes in the schema.
But more recently, a lot of these, and we, I've talked about this before, a lot of these, um, players at NL Ops were using their form of observability, which is where they're monitoring the correctness, the relevance and hallucination management, and they're calling observability, and rightfully so. So you actually really now what I would call, uh, uh, you heard it here for us folks, a postmodern observability, which is really three levels, which is infrastructure, applications, data, MA data, and, you know, um, ML ops or generative AI ops, which all fall under the category observability. And, and by the way, some of the vendors, um, in the, uh, this, you know, sort of gen AI observability, uh, level are using open telemetry, which is really cool.
Yeah. Like, so they're, they're like, Hey, you know, like, put your data here, you can throw it up in Prometheus. You can do all this great stuff because you're following the standard.
You know? So, so I think that's, You brought up Prometheus, John. It, it's interesting, right?
To me, there is, so open telemetry is very much, you know, under the covers the, the underpinnings of a lot of, of, a lot of this. But Prometheus is no slouch either, right? And a lot of times it, it kinda works like this with, with open tell, right?
Um, the Prometheus is, is also I, I think in the top five or six open source projects at CNCF. And then you have about, you mentioned Jager, another one, our friend Jonah Cowell, of course, is one of the maintainers on the Jager project, on the Jager project. Other, uh, uh, these, the Grafana labs people as well, right?
That's a, I mean, to me, the real revolution or evolution here was that what we used to call the a PM application performance monitoring and management market that had, you know, several successful companies, AppDynamics, of course, Cisco paid billions for New Relic Public Company, uh, Dynatrace more. There were, there were many great, and there still are many great a PM, uh, companies out there. And products Datadog at that time too.
They, they shifted under the covers from proprietary models of data collection to relying on, on, on, on triumvirate and more of open source projects led by open telemetry. Yeah. Which, you know, you mentioned the open trace kind of project, and that was married to a few other things to give us what is today o hotel.
I think That's the beauty of all this, right? It is very easy for even like these new AI players to get data into Prometheus or Grafana, right? Because there is sort of a standard now there's a way to do it.
It's, it's a way for me to sort of integrate, you know, uh, Dynatrace and whatever other open source. So like, this is why O OT is so successful. Um, and, and again, I think the, the thing that really caused the conversion in, you know, I'd love to have somebody tell me I'm wrong.
'cause I'm, I'm up, you know, I, I'm wrong about 50% of the time. But hey, um, the, um, you know, I, I think it was the service mesh that like, was a forcing factor. And it was a convergence point where, you know, service mesh, you know, you could, you couldn't have all that complexity in the service mesh that all the things that it does without sort of having something in the middle that allowed you to sort of integrate Prometheus with, with zip you know, whatever zip, zip cans become o part of hotel, all that stuff.
And even traffic management and all that stuff just required, you know, a a new architectural design that required to be sort of standard API based. And so I think that was the thing that was a forcing factor of the success of, in my opinion, hotel as well. I, I, I think also the open tell, and Amanda, maybe you remember there was a company ServiceNow bought.
It wasn't light Tell what were it, it was the four four guys who came outta Google, Ben, um, was it Ben Har, not Ben Horowitz. Um, anyway, I'm at a loss. Amanda, I don't know if you would remember this company either.
It was an acquisition by ServiceNow. These guys really help kind of push open out to what it is today. And I apologize, I'm sometimes I, I remember things and then I hear their guy.
Anyway. Hey, let's move on a little bit here from, from the, uh, open telemetry and DevOps thing, and we got a few minutes. John, what, what do you got going on exciting you want to share that you haven't shared yet?
Yeah, no, the, I think what was interesting is, um, you know, I've been running a lot of hackathons and, and, and I've been learning, just watching what people come up with, right? Is another sort of, you know, for most anybody who's been following me in the last year for me, has just been a total immersion in trying to learn as much as I can about what this AI stuff is, how it relates, uh, my, what I care about is how it's going to impact, you know, sort of the DevOps, our infrastructure or technical debt, all the things that are gonna fall back on us. But, um, but one of the things I fell into is, uh, there were a couple of guys, you know, prototyping this idea and, and of, um, so embeddings are not secure.
So all this stuff about like creating vectors and you can reverse engineer embeddings Mm-Hmm. Right? So all this stuff embeddings are like, you, you take sort of words and you tokenize 'em and you spray 'em, and you put 'em in these high dimensional spaces.
And so, um, so there is some papers and, and, and research and people pointing out that like, it is not impossible. In fact, it is feasible to reverse engineer embeddings, particularly if you have a box with GPUs, right? Um, you know, and, and, and interesting thing about embeddings is, and, and actually GPUs I've learned this recently is, and, and again, I'm, I'm willing to get corrected here, but like GPUs in essence really do matrix multiplication.
That's what they do. They do, they literally, uh, um, you know, um, multiply matrixes, which create another layers, and there layers and layer layers which create these deep neural networks, right? Um, so one of the things we found out is if you take an embedding, let's say you have a chunk of data that is a description of what to do in case of, um, you know, like, you know, let's take a crazy thing.
Like let's say there are, uh, set of steps you have to do to turn down a border router because you might be compromised, or you might have some crazy Al Gore going nuts, right? These are true stories, right? Like your large financial institution, and you realize that, you know, a night capital's happening, and so therefore there are two people that can actually shut off the border, border border routers, right?
And there's a set of procedures, almost like a nucleus sub of what you have to do. These are two stories, right? Like mm-Hmm.
And, and let's say now you've got that in an LLM. Now you might say you're crazy for putting that in a little chat, but let's just say that's really proprietary data, right? And, and so, and, and that proprietary is vectorized, right?
So it's in this highly, you know, maybe through a rag or something like that, which then when I ask, Hey, what do I do emergency? I have to, uh, I need to, uh, turn off the, the, the guy who normally does it is not here and I need to turn off the border routers, right? And it's gonna answer that question again.
I, I'll, I'll take as many hits. I know if you wanna say, John, you're crazy. Nobody would ever do this.
That's fine. I'm just using it as a training sample. And let's say that like we wanted to encrypt that data, but the embeddings are not encrypted, right?
You can't really just encrypt embeddings because embeddings are basically highly dimensional fields or tokens in a vector. And the way that it finds the next word, the next sentence, the next paragraph is to similarities like, like, like angles in the vector. So I know this is gonna, that hopefully this will make sense when I get through it.
What we found is if I apply, um, matrix multiplication against the embedding vector, it will still give the same location and the same word still just be encrypted words. So this is really fascinating that you can literally use the same math that they use to create the embeddings to put like some complex key on top of the embedding, which then encrypts the embedding itself. So now that you there, like it's gonna point to the gibberish words that you've encrypted alongside of that.
But now if I have a key, I can do that question. I get the gibberish paragraph back. But now I can go ahead and run my key and get the answer of what it is.
And, and the beauty is like, it's, it's a way, it's such a simple idea that you can basically turn the embeddings into a secret by not really encrypting it, but by using the same math that uses to solve Multiplication. Because the, the location Is they're gonna be, the problem is once we know that you could undo that, though, what's gonna stop the bad guys from doing it as well? Well, they need the key.
That's true. 'cause otherwise they just, yeah. I mean, you've got a 4,096 key, or Yeah, yeah.
The key that you use to use the matrix multiplication, same key that you use to encrypt the data, you know, and, and, and well, if you're building this stuff on A GPU, you can make the key like incredibly large, right? And so, yeah, no, I, I think it's gonna be an interesting conversation. Um, you know, I like, like again, somebody in a hackathon was playing with this and then, uh, me and Joseph Enox did a bunch of research and like, it's an emerging topic, the whole idea of like securing, uh, embeddings Mm-Hmm.
Um, so I Guess the biggest thing will be safeguarding those keys, Right? Well that surprise I would not use the same key. Probably good practices.
You don't use the same key for the encryption that you would do for the, uh, You don't have to. Yes, sir. That's a good idea.
I mean, You don't have, right? 'cause this way you got, it's a doublet system. Orthogonal, really.
Yeah. One, one is a key that just changes this consistent. Like if, if the distance between this word and this word, if I use matrix multiplication, it's the same distance.
Yep. You know, so, So I don't know how many people out here we lost on that, But if you're interested, write to John about it. It's fascinating stuff.
I, as we continue, Hey guys, we're, we're at 45 minutes in. I, I am afraid I have to pull the plug on this version of the gag. It was great having you both on here remotely.
Hopefully we'll have more people in person later this week. I wanna quickly shout out though that, you know, this is Tuesday. We'll be back out with fresh content on the gang in Text Trump tv.
We'll schedule on Thursday, next week's a big week. We are out at RSA Network Monday through Thursday, live on broadcast all as well as our DevSecOps AI event. We will also be in Las Vegas for the Service Now user conference.
And I don't remember the name of it off the top of my head. Um, also this week we are in Vegas, Mitch, and some of the, uh, team is out there for the Atlassian conference. So stay tuned for live coverage on that.
Another busy, busy week here on Tech Drunk tv. John, Amanda, thanks for joining me. Thank you for watching this.
This is Alan Shemel for Textron. We're out.