Enhancing GPU Usage with Phison Technology’s Michael Wu
In this Techstrong.ai video, Mike Vizard talks to Michael Wu, president and general manager for North America at Phison Technology, about how NAND memory technologies can be used to optimize usage of graphical processor units (GPUs) to improve artificial intelligence (AI) applications.
Transcript
ai video series. I'm your host, Mike Azar, today with Michael Wu, who is president in general manager for Fissan Technology. And we're talking about, well, how to make gen AI more affordable, because as we get into this, people are starting to realize it's expensive, and a lot of that has to do with the underlying infrastructure.
Michael, welcome to show. Hi. Hi, Mike.
Nice to meet you. Glad to be in on show, if You don't mind level set for us for the uninitiated, but from your perspective, what does make gen AI so expensive? Because, well, people are starting to figure out tokens and tokens are starting to kick off processes on backend infrastructure.
There's GPUs, datas, networks, all kinds of good things involved. Yeah, sure. Uh, so first of all, um, I like to say that everybody just love gen AI right now because when the tragedy just come on, everybody is like, wow, this is, there's a free version.
It's easy to use, and overnight it just exploded, right? Everybody become a AI guru, right? So it's amazing how information, uh, is, is transforming the way that we work every day.
But what the, the, what the problem is, the what, the problem we try to solve is that how does small businesses, right, uh, or small medium business take advantage of the gen ai, right? Uh, the problem is you've seen that company across the globe, the forbid people to actually use treasury BT to put in proprietary company information, right? So, what's happening is that, uh, you know, companies in order to enjoy the generative AI and with their own data, they have to build an infrastructure that is contain a very premium, uh, GPUs, uh, which currently is all allocated by a very, very large company that does pre-training, right?
With this Dutch model. So the, the, the allocation of the GPU, the cost of the GPU, it's what really make the small business not, not, you know, not able to kind of get into the game of general AI right now. And some of that also has to do with, um, I don't wanna lose control of my data.
So people are using various techniques to expose their data to a large language model, but, um, there are performance challenges that go with that, I would assume. So, um, walk us through like, what do I need to do exactly to share my data with an LLM in a way that, um, doesn't wind up being too slow to use? Uh, yeah, sure.
So, uh, right now, uh, due to the big infrastructure you have to kind of build, to have your own gen ai, right? To how train your own local data, right? Uh, as I said before, you need to have massive amount of GPU, and the reason why it is so, uh, expensive and it's not easy to get is because, uh, those expensive GPU, uh, are the key to able to train a very large model, right?
There's a memory chip, uh, on every GPU card. It's called, uh, VAM or you know, there's, uh, you, you've heard HBM there. And so all the training data, if you wanna have your own AI train your own data, right?
It needs to have enough of GPU in parallel to actually train your data. So what we are being able to do is using our expertise in SSD right name flash, right? To create a tier of memory that actually expand the GPU memory that allows a massive amount of data instead of trying to buy more GPU, right?
And spend a lot of millions of dollars, right? You actually put, uh, uh, what we call adaptive cash SSD that, uh, that Greg, you know, uh, uh, very fast to kind expand the memory from a matter of, you know, 40 gigabyte all the way to two giga, two terabyte, right? So we unlock the possibility of training a very large model, which is very essential for the organization, small organization, right?
To be able to put all the data into this, uh, this AI training machine, uh, to, to able to make it useful for the company. It sounds like you're increasing the utilization rates of the GPU using caching. Is that a fair assessment?
Uh, we are, uh, ex we are making the host to think that there are actually tons and tons of memory behind the GPU that has, that's part of the GPU, right? So it allows the completion of the training without a certain amount of memory on the GPU card, it won't even complete the training. So you think one of the issues we seem to be incurring here is that, um, not only is the GPU hard to find and it's expensive, but when we do get it, it's not the most efficient thing in the world.
So we need to kind of figure out other ways to, uh, optimize its performance. Well, right now, with the current infrastructure and the way that the GPU memory is progressing, uh, for any company, small million company that want things on-prem that have want to have their data secured, they are able to train probably a 70 billion model se 7 billion model or 13 billion model. Uh, just to give you an idea, 13 billion model, it would require eight GPUs to run, right?
And if we're talking about the latest release, the open source line mastery model, 70 billion, good luck, right? You need 30, 30 GPUs, right? Uh, something like H 100 or 6,000 data to actually train a basic 70 billion model.
What if I can tell you that instead of a dirty GPU that has to run across a rack and you have to hire a, a expert IT person to connect all this together and make sure it run together, I can put all that into a one single workstation that has four GPUs with our adaptive cash and finish the whole 70 billion training. Okay? And, uh, and actually what's uh, more exciting is you heard about NMA three, 7 billion, 70 billion, and you know, very soon you're gonna see 400 billion coming up, right?
So, uh, because of our innovation around this technology, uh, we are able to also decrease the DRAM requirement on the system at the same time to allow beyond 70 even today. So, who is savvy enough in an organization these days to do that math? 'cause a lot of times if I talk to folks, data science teams are not the most infrastructure aware people.
So who is stepping up to kinda have this conversation? Well, so actually, uh, you hit the nail on the head. Nobody, this thing is so new that people say, okay, if even if I have a 40,000, you know, workstation, um, how do I use it?
How do I use it? How do I make my productivity improve with this machine? Right?
So, going through this exercise, what we have found is that there are three key element to the success of integration for gene AI on every organization. Number one, it has to be so easy to use, just like when chat GP just came on board, right? Just like iPhone, it has to be very easy to use.
Number two, very affordable, right? And number three is secure, right? Data fully secure.
So we just solved the, the, the last problem, which is the secure, if we can have a 30,000 or 50,000 infrastructure on a workstation, I think you can afford it, right? Right. So as far as affordability and security, if we have a on-prem workstation that sits on the, the different companies, that part is solved on the ease of use is the most critical one.
And we, as we go through this journey, we found out that a lot of people don't even know they need a generative ai, right? How to use. So this is why in the co upcoming trade shows, uh, you know, FMS around the corner, uh, we are announcing a end-to-end software solution.
We call it the Adaptive Pro suite. So adaptive is the name for our technology, uh, using a hybrid solution, uh, SSD and a specialized, um, adaptive link middleware to make the GPU expand its memory. The pro suite is a software that our expectation is on a single desktop, uh, environment, uh, environment.
You can load in your data, drag and drop, you can train the data, and usually you want to train the data when people are sleeping, right? So, and then you could, on the same gui you could actually start typing inquiries about data from your company, just like a tragedy, bt, but it's gonna be fine tuned with all your company proprietary data, thousands and thousands of data should document and all that, right? So if you could take this suite to anyone in the company, you actually don't need an AI engineer to be in a company to integrate the Gen ai, right?
Right. So that's our whole, uh, we think that there's so much potential for the Gen ai, but if we do not solve the ease of use, it won't, it won't scale. And if the end user, if the end device, if every company doesn't have a gen AI machine everywhere, uh, this AI bubble may not last.
This, this AI momentum, right? May become a bubble because, you know, look, look at how smartphone is really, you know, uh, and, and the, the accessibility of the internet, right? It's really what make the AI really take off, right?
So we believe that with the on-prem affordable solution, uh, you can make this edge devices enabled with ai and in return, actually, it makes more money for the people on the crowd. We hear a lot too lately about small language models, and by extension there'll be medium sized language models. Um, so not everything necessarily needs to be a large language model, but if there's gonna be small language models, it seems like there's gonna be a lot more of them.
And I'm gonna have a lot more training projects, and so on that workstation, I might be training more models than ever. How do you see that might play out? So, uh, the reason why, and, and you hear that out there, right?
The the training is the small model and all that. And, and I can tell you that, uh, a lot of the devices, uh, you know, uh, a lot of people are trying to integrate IPC or AI phone, right? It, it just have no choice.
But having a small model, so first of all, the reason why people are kind of talking about small model is due to the hardware limitation, okay? That you wanna run inference, you wanna run model, it can be, it has to be small 7 billion or, or something small, right? Um, but that's not what we're talking about.
We're talking about enabling businesses. You can have a rack, you can have a workstation that's there, right? You could plug in different things that make it trend a bigger model.
So, but what's the difference? Real tangible difference about the small model, but it's a big model. We've done, uh, university studies, we have actually competition results.
The difference is the quality of the answers are different, right? Uh, you know how people are getting into the small model, uh, they do quantization, they, you know, they, they actually sacrifice the accuracy, right? So if you take a 7 billion model all day long, a, a pre-trained model, right?
70 billion from lama, and you put your data in there and train the data, the data, the accuracy, and also the intuitiveness of the responses when you compare with the 7 billion is night and day, it's different, right? So that's why we think that yes, the end user due to the hardware limitation, small model, whatever it fits, it's okay. But when talk about businesses where, you know, internally, we're actually even enabling the AI to turn a firmware code to a document that we will submit to a automotive, automotive certification.
So you have to have a very high quality model in order to not have to manually change a lot of things after the gif ai for the usual results. Do you think that thanks to the rise of ai, people are suddenly paying more attention to what's going on in their infrastructure that they might have taken for granted all these years? Oh, wow.
Big time. Big time. I can tell you, um, the only way you are gonna get a budget these days is that you are investing in the ai, right?
I even seen the companies, they are cutting some workforce out so that they can buy more h right? Because at the end of the day, the power of the AI has to be unlocked by the GPU and its server architecture, right? That's the machine that's powered, right?
So right now, I can tell you every organization that they're thinking about, you know, uh, additional investment on their server, firstly, they, they're gonna, and, and even when we pitch push this, uh, gen AI training machine, people are saying, well, my company is actually going to buy a inference machine, right? So that they can actually run some model and, and run some, maybe they're doing, um, some, some content creations. Uh, they want to use a large model to expedite some of the process.
So what happened if I already have a budget for the inference, and now you have this training, how does it work? And our answer to them is, well, I want you to train when you're sleeping and when you wake up, that become a really powerful inference machine, right? So, but in general, every AI spends, any hiring spends are all geared towards AI now because people know that this is the biggest, you know, evolution in the recent years, Did we maybe over or underestimate what it would take to accomplish all this last year?
And we had this kind of, you know, rational enthusiasm for all things AI and looks like we're spending this year trying to figure out how to actually turn it into reality and operationalize it. And this may be the year of AI infrastructure, and then next year is when do we see the ROI, Um, you, uh, mentioned the keyboard. Actually, one of the, the speech by one of the executive is the key for this evolution is the ROI, right?
Uh, people are throwing money right now, right? For the infrastructure. This is for the first time they're building a, uh, a big ai uh, infrastructure, right?
But at the end of the day, uh, we know that they're doing this massive infrastructure because they know the ROI will pay at the end of the day, right? Uh, people can charge for the crowd, uh, training, computing and all that. Um, they are going to charge signif, uh, uh, enough so that they make their ROI, right?
But my question is, how about other companies, right? Uh, when they throw in 10 extra hundred, this, how do they make their profit back? Because everybody is doing the same thing.
Everybody is investing, and you know, if we learn from the internet, the times, right? Um, 90% of the company may fail and 10% may win, right? So there's, there's a, the question is, how do you really make a reasonable ROI by investing this AI machine?
Now, if I tell them that, Hey, maybe you don't need a Maserati, I'll give you a Toyota, okay? But you could actually finish the application you're developing using this more affordable gen AI training business appliance. We call it appliance because we want it to be so easy, like a microwave, right?
A appliance, right? Um, what if you have that option, then you may not have to, you could be profitable much faster, uh, versus, you know, try to make back with a lot of big investment right? From the GPUs.
So ultimately, what's your best advice to all those IT leaders that have been kinda watching this whole AI revolution occur? And I think sometimes maybe they felt a little left out of it, or are they back in the game, or where are they on this equation? Yeah, definitely.
As I mentioned, right? My advice to the it, uh, managers is that AI is gonna help you so much and not take away some job, right? That that's the biggest thing, right?
And I think that if you're the IT manager within every company, your goal is to actually make things more efficient, right? Uh, you know, if you don't make things efficient, all you do is keep adding servers, keep adding, um, resources because there are more headcounts coming, right? Uh, what if you can shrink, you know, 20 emails into one check bot, right?
On the just really simple things, right? So have an open kind of attitude on how to build around your a ai, uh, uh, infrastructure, right? Uh, it may sound intimidating at the first, uh, but actually as I kind of look into this AI ecosystem, right?
They, the communities are very, uh, resourceful. The resources are out there and, um, uh, you need to kind of have a strategy around how to kind of bring all this good things of AI and, and, you know, come and track us out, right? I think that, uh, normally the first hesitation of any getting into the AI is the entry tickets are so high.
5 million to build a gene AI machine for our engineering team got rejected. And our CEO just say, guys, figure it out, right? So I don't think we're the only company that gonna have the same issue, right?
So check out our machine. We're gonna have a lot of demo. Uh, actually we are, our goal is to, you know, the process software, we actually are wanted to be ware and let a lot of people experience it because knowing the power of fine tuning a generic model and make it yours, that's the first step.
And you should not afraid of getting to that step, right? Because the first step is what take you to the next step where, okay, I love this fine tuning, I love this custom ai, but I don't have the money to build it, and we have a solution. All right, folks, you heard it here.
Hey, AI models are all the rage, but they're not gonna be worth the whole, I give a lot if you can't figure out how to run 'em. Hey, Michael, thanks for being on the show. Of course.
Thank you for inviting. All right, and thank you all for watching the latest episode of the Techron AI video series. You can check out this episode and others on our website where we have well many other video editors.
Until then, we'll see you next time.