The DRAM Barrier – Why VMware Advanced Memory Tiering is a Data Center Game Changer with VMware
Memory is often the most expensive and restrictive bottleneck in modern datacenters. VMware Memory Tiering (an industry exclusive) solves this by automating data placement across high-performance and cost-optimized memory tiers. This session explores how this unique hypervisor integration drives 40%+ TCO savings, improves VM density, and ensures smarter resource consumption. Learn why VMware is the sole leader in transforming memory from a hardware constraint into a strategic advantage. This innovative feature, “VMware Advanced Memory Tiering with NVMe,” addresses the rapidly escalating cost of DRAM, which now accounts for up to 96% of a server’s bill of materials. Presented as a core component of vSphere, and thus included in VMware Cloud Foundation (VCF) and VMware vSphere Foundation (VVF), this technology aims to overcome the “DRAM barrier” by intelligently managing memory resources.
The core of VMware Memory Tiering involves using less expensive NVMe devices as a secondary memory tier, with DRAM remaining the primary, high-performance tier (Tier 0). From a VM’s perspective, this combination appears as a single, logical memory space, making the underlying tiering transparent. VMware employs a proprietary algorithm that constantly monitors memory page activity, classifying pages as hot, warm, or cold based on recent access patterns. When DRAM utilization reaches a configurable threshold (e.g., 70-75% pressure), cold, inactive pages are proactively moved to the NVMe tier, freeing up DRAM for active workloads. This intelligent, proactive approach differs from reactive measures like swapping or ballooning, enabling customers to achieve over 40% reduction in total cost of ownership by purchasing less physical DRAM, and doubling VM density on existing hardware due to more efficient CPU and memory utilization. The NVMe devices must be directly connected, dedicated solely for this purpose, and meet specific endurance and performance requirements, with hardware RAID support for data mirroring and redundancy.
For operational flexibility, VMware Memory Tiering offers configurable ratios between DRAM and NVMe, starting with a default 1:1 ratio (providing 100% more memory capacity) and scalable up to 1:4 (a 4X increase), with a maximum partition size of 4TB. This allows administrators to adjust capacity based on workload needs without physical hardware changes. The feature seamlessly integrates with existing vSphere functionalities like HA, DRS, and vMotion, as well as various encryption methods (host, VM, vSAN), with vMotion being “tier-aware” to handle VM migrations between hosts with and without memory tiering. However, certain specialized VMs, such as latency-sensitive applications, monster VMs, and security-hardened VMs (e.g., those using TDX or SEV for memory encryption), are not supported as the hypervisor cannot classify their encrypted memory pages. VMware provides extensive documentation, including performance whitepapers, deployment guides, and Hands-on Labs, to aid in understanding and implementing this transformative technology.
Presented by Dave Morera, Staff Technical Marketing Architect, VCF Division, Broadcom. Recorded live at Cloud Field Day in Santa Clara on March 12th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/vmware-by-broadcom-presents-at-cloud-field-day-25/ or visit https://techfieldday.com/event/cfd25/ or https://www.vmware.com/ for more information.
Transcript
Hello, everyone. My name is Dave Moreira. Happy to be here.
First time as a presenter here, but I'm no stranger to Field Day. I've been a delegate a few times in the past, few years ago. Um, so I'm very happy to be here, and, something that, that I've learned during that experience is that you guys like to dive deep right away, and that's exactly what we're gonna do, right?
Um, so today we're gonna talk about one of my favorite features for VMware Cloud Foundation 9, which is, VMware Advanced Memory Tiering with NVMe. You may have heard about it or not, but my goal here is so that you understand what it is, how it works, and how it benefits our customers in a way. Um, I call it the beyond the DRAM barrier.
I wanted to call it the RAMpocalypse, so maybe I should just- ... trademark that from now on. Uh, but you or most of you are aware of what's happening with the prices in RAM, right?
So this is very timely subject, and I hope you get a lot out of it. Yeah, when we're talking about this feature in particular and how it fits within the whole VCF ecosystem, right? We have a lot of, components within VCF, right?
We have the compute component, storage, networking, which is vSAN, NSX, et cetera. So the compute component is what we, you guys know as vSphere, which, which is vCenter and ESX. And yes, we changed the name back to ESX from ESXi, 'cause we like to change things up.
So this feature is part of that core compute component. Say that fast five times. Um, so it is included on the core vSphere area, right?
So whether you're running VCF, VVF, it is already part of, the ecosystem of VCF. Right, so we got that out of the way. So what exactly is memory tiering, right?
So I have this quick diagram here. What we're trying to do is leverage less expensive devices to act as memory, right? Um, so in this case, we're gonna take high-speed NVMe devices and pair that up with DRAM.
DRAM is always gonna be that tier zero, what we call tier zero, the first point of entry for the, memory pages, and underneath that, we're gonna have NVMe devices to kind of back that off, right? And but from a VM perspective, we're creating a logical memory unit per se, and presenting that to VMs, containers, VKS. Whatever you're running on VCF, that's what we're presenting to VM.
So from a VM perspective, they have no idea what is happening. They don't know there's NVMe, or they don't know there's DRAM. They don't know what type it is.
They just know, "Cool, I have a bunch of memory, I can use that. " The way, a- again, they consume that logical memory, and I do want to point out that this is a native solution, so we don't have a separate appliance anywhere. Uh, we're not taking additional resources, a ton of resources anywhere else, so it is part of vSphere, which is our core compute component.
And the goal here is to lower the total cost of ownership for what we're seeing here is the cost of DRAM. So even before all this shift to HBM memory happened for AI, we started working on this a few years ago, and we did see that most of the bill of materials, about 80% of it, was the cost of memory, just, you know, a few years ago. Now, that has shifted to almost 96%.
So when you go buy a server that is probably now $200,000, 95% of that price comes from memory alone, right? And this is what customers are looking at, and they've been keeping me pretty busy lately asking, "Hey, up to 40%, which is what we're doing here," right? So instead of buying one terabyte, I can get away by buying five 12 gigs of memory, for example.
So if we dive a little bit deeper into the NVMe side, you may be asking, "Well, okay, how does it work? Do, kinda, kinda be connected over the wire? " So the answer to that is it needs to be a direct connected NVMe device dedicated only for this purpose.
A lot of customers think, "Oh, I'm wasting space. " No. No.
If you think about it, it's going to affect performance. We don't want anything to affect that I/O path, so we're dedicating a device just for memory. We're creating a partition, dedicated for memory alone, and, you know, we can scale up depending on, the ratios, and I'll talk about the ratios in a second here.
The other question I get is, "Well, cool, that's great, but I know those devices may fail sooner than DRAM," and that's a valid point. 0, we have, support for hardware RAID, so you can buy a controller, even Read Rock is supported as well, and you can have two devices working together. So the memory pages that go into NVMe are gonna be mirror that way, so if you lose one device, you have a backup there.
Any questions so far? I'll throw out one. Sure.
Um, one that came to mind was speaking of the DRAM and, like, the cost and everything, so, like, what percentage DRAM reduction are customers realistically seeing right now? So, I mean, I'll explain how it works later to, it, it will make more sense, but we're seeing about a 40% reduction in cost by doing this. Um, so alone, all the way it comes from the memory alone.
Not only that, but also you double your density- Yeah ... of the VM. So you can bring 2x amount of VMs, and because now you can push the, the DRAM and NVMe together even more.
Yes. Thank you. Good question.
All right. So that's what it was. Uh, but, how does it work?
Where's, where's the magic here, right? Um, so what we're seeing, I'm, I'm gonna put this up, right? A lot of customers are seeing this, and even internally, we're, we're consuming a lot of memory.
Think about those workloads. I used to be an admin back in the day, and my SQL Servers every week kept asking for more RAM. " Like, you know, one terabyte of RAM when active memory was very small.
So what we're seeing here is that we're memory bound on the host level, but also we're seeing that the CPUs are being starved for memory. So that's another problem. Number one problem here, we run out of memory.
Uh, the other problem here is that we're not really using resources efficiently, right? So we're wasting space, we're wasting not only resources, but also power cooling rack space. If I have a server that is fully popular- populated with CPUs, all the DRAM slots are populated, you know, how can I move forward?
I need to buy more servers, and doing that now, it's, it's very, very pricey. So if we dive into that, scenario, right, where we're using most of our memory and little bit of the CPU. We have an algorithm that's proprietary, that is every few cycles it's looking at the activity of the pages, not only if that page is being used, but how often, how recently, what pattern.
So we looked at all these things, and we make decisions based on that, and we classify those pages into hot, warm, cold, very cold, right? Depending on the activity. Every few seconds and, you know, based on that information, and, and then we bring NVMe, and then we make the decision of moving some of those inactive pages down.
" So yes, vSAN, OSA back in the day, very similar, right? We, we pushed down some of those, data that is not being used, push it down to inexpensive devices. So that's exactly what we're doing.
We take NVMe, push the pages that are not being touched down to NVMe, and DRAM is gonna keep those pages that are active. So from a VM perspective, they're always going to talk to DRAM because they want that data fast, they wanna get it fast. When a VM comes down and say, "Okay, I need that page that is cold," we're gonna bring that up to DRAM, and we're gonna read from DRAM.
So we're constantly moving up and down. Is it per VM or is this host-wide or how can that- Yes and no ... get allocated?
Yes, so the way we can configure this is, by cluster, per host. Um, we can disable the VM. So if a VM if, for example, is latency sensitive, it probably doesn't want this, right?
Just in the chance there's a, latency there, so we can disable the, the VM from having memory tiering. Good question. So Dave, this...
And then Jack Palmer from Paradigm Technical. This seems like standard ca- standard memory tiering and caching, and you're just, what's... I, I don't understand what the, the, the, the secret sauce is here other than you're pushing...
I mean, what's on NVMe? What's, what's the storage mechanism that's on the NVMe bus? So you're thinking about, swapping, is that what you're saying?
You, you, you thinking- Well, yeah. I mean, you, you, you've replaced... Y- you've, you've described what is essentially a memory cache, right?
Yeah. And a cache replacement algorithm, you know, classify what's hot, cold, warm- Right ... in the cache.
Yep. And we decide what we evict out of the cache and how we bring it back in. That's what- Yeah.
In a way, yeah ... standard system memory caching, right? Right.
So you're pushing it... You keep saying you're pushing it down to NVMe. What's the backing store on NVMe?
Is it SSD? Is it DRAM, SRAM on, on an NVMe card or? It's, so it's a PCIE NVMe device, so it could be any form factor that way.
Uh, but it would do require certain endurance and performance classes to, to meet our requirement for performance. So it's just a standard backing store. It could be an SSD.
Yeah, yeah. It could be, a DRAM, a slower DRAM or something that's somewhere down in the NVMe side of the world. Yep.
Yep. Okay, thank you. Uh, Denny Cherry from Denny Cherry Associates Consulting.
Um, what's the class... What, what are you considering when you're looking at the, the, the pages? Is it reads?
Is it writes? How, how are you deciding what's hot and what's cold? Uh, both.
So we look at, like, the activity based on how recent it was read or written, how often was it on the pattern. All those things come into play to make the decisions. Okay.
Right? Um, now there's, expand here a little bit more on that. If the DRAM is being used maybe only 50%, we don't want to move pages just for the sake of moving pages, right?
Because you know, that's additional resources, maybe a performance hit. But, so we have to pass a threshold about 70, 75%. When we have memory pressure, then we start proactively moving pages out, right?
So we're not, being reactive here. We're proactively thinking, "Okay, we're getting full. " So just a quick question.
Leon Talera from Infocert. Um, we are enabling DRAM using this method, this, technology. Being able to see larger memory than, larger capacity than you can have with a single RAM, or it is another- Yeah, we'll get to the ratios in a second, but yeah, we're starting with a one-to-one ratio.
Oh, yeah. So you're, doing 100% more- Right ... with all that.
Thank you. Yep. Uh, so again, you know, this is presented to a VM as one logical unit, both, both of these tiers, and once we move pages down to NVMe, make room on DRAM, then we can push more in.
Here's where I talked about the having 2X density on the VMs. 5 more, VMs here, and then we can push both resources at the same time. Because we have memory available, we can actually utilize more of the CPU that was just sitting there idling, right?
So this is, this is what we want to see. Uh, you know, we pay for those resources, we wanna use as much as we, we can. I mean, that's the whole idea of virtualization, right?
Good questions, by the way. Uh, you also may be thinking, and you touched on that, about swapping and ballooning. That's still in place, right?
So these methods here, we still have TPS, which is memory sharing, compression, ballooning, swapping. That stays in play, in place, but these are more reactive, things that we do down the road. If there's a lot of pressure, we're still going to do swapping and ballooning eventually.
Memory tiering is gonna come in right after TPS and help us to be more proactive, gives us more space, manage those more intelligently, and those resources, return those resources back to the host. But yeah, if we have more pressure, if we're pushing overcommitting and doing things like that too much, then we're still gonna have those, reactive measurements to, to gather some of those resources back again. com.
Um, can you talk a little bit more about ballooning? I recognize all those other ones, but that's interesting. Yeah.
B- ballooning, it's, we p- pretty much have a driver within the guest OS- Okay ... that it starts to inflate, so it's kind of reclaiming some of those, pages back. Ah.
Right? Okay. Uh, but what happens there is that we don't, we don't classify those pages.
So because it's reactive, we just ... If it's active or coil, it doesn't matter, we're trying to reclaim those, as many as possible. Ah, okay.
Right? Thanks for the clarity. Versus the tiering where we actually classify the pages and we, we don't want to mess with the hot pages.
Gotcha. Yeah. Thanks.
Good question. Yeah. Thanks for the clarification.
All right, so we talked about what it is, how it works. Uh, now I want to talk about some of the operations, the, uh ... And one of them is the ratio, is how to configure all that stuff.
Um, so in ... So we came up with Memory Tiering in vSphere 80U3, which was tech preview, very limited. And the ratio was on a 4:1, a DRAM to NVMe ratio, meaning that you could only expand 25% with NVMe.
You know, it's small, but 25% more is mo- it's ... Okay. Now, we went GA with VCF 9, and the default now is a 1:1 ratio, so you have 100% more out of the bag, right?
So you have, for example, a host with 256 gigs of DRAM. You buy a drive that's at least the same size, and you get twice as much. You get almost half a terabyte.
What if, you know, we go and we buy a bigger drive? So, if we don't change the ratio, we're still only going to use 256, 'cause we're basing everything off the DRAM. A couple things happen here.
You've got- Number one, you know, with, the NVMe controller, we do ... We're leveling, so we're s- you know, distributing those reads and writes across the drive, even if it's bigger. Uh, so it's going to extend the life of that drive a little longer.
Based on w- on our testing, endurance and performance classes that we require or recommend, those drives should last about five years, so having a bigger drive doesn't hurt to ex- extend it a little more. But also, it prepares you for the future, right? So we know your workloads are great for a 1:1 ratio, but what if the active memory is really low, right?
I can actually push this even more and change my ratio to a 1:2. Uh, so I get 300% or 200% more, giving me almost a terabyte. Uh, and I didn't have to do anything.
So I had the same drive from example two, you know, 512. The partition was already created. All I did was go to vSphere, change my ratio from a 1:1 to a 1:2, and now I have, you know, an extra 256 gigs of memory from that.
Uh, so that is very ... allows you to be very flexible expanding. So just imagine doing this with DRAM alone, right?
You have to take some DIMMs out, put bigger DIMMs in, et cetera. Uh, so it's a little harder. So this is very flexible and allows you to go up and down as necessary.
How far can you push that? So we can go up to 4X. Okay.
Um, so, so there are some workloads that we notice, for example, VDI, right? Some of them are not doing much. They're just a kiosk somewhere, so we can push those up to four, 4X.
The maximum, partition size is four terabytes, for now, so hopefully we can expand that later. I was talking to our customer yesterday. They have huge, infrastructure, so they wanna push that up even more, so, yeah.
Four terabytes maximum partition size and 4X, memory increase here. Uh, and also, last example, you know, it doesn't matter what the ratio is. If you don't have a big enough drive, obviously you can't make space out of nothing.
So I bring this up just to, for those who are watching, you guys, it's important to properly size, think about the future. You know, what can you ... it i- to invest h- invest here.
Um, don't try to save money on getting a smaller drive. You know, you're saving a lot of money already with DRAM alone, so don't try to c- cut corners here. All right, and a lot of ...
Another question I get is does this work with everything else, right? Um, you know, when there's a new feature and we say, "Oh, it doesn't support vMotion," people don't, are not very excited, and, you know, rightfully so. Uh, vSphere's been around for over 20 years and, you know, HA, DRS, vMotion, that's something that is very important for admins.
And, you know, being an admin back in the day, it's ... you know, I can see that. So yes, we support HA, DRS, fast suspend resume, which is what we use when we hot add, hot remove devices.
So yeah, definitely in. I talked about RAID being supported. And vMotion's interesting.
Um, when I say interesting, because, you can configure, an entire cluster, but I also said you can configure per host, which means that you can have, within the cluster, you can have a host that con- configured for Memory Tiering, but some other hosts are not configured for Memory Tiering, right? Why? Because maybe you have VMs that are latency sensitive, and, maybe monster VMs that we don't support there yet, but we wanna put them on those hosts, right?
So there's, there's some flexibility there. Uh, but vMotion knows that, okay, this host has Memory Tiering, this one doesn't. I can still move stuff from DRAM to RIMA- DRAM and, you know, if I have Memory Tiering, I can do the, the tier out capability.
If I don't, then I just keep it on DRAM. So it does know. It's aware of that.
Hm. I have a quick question. Yeah.
You mentioned all the features that are supported. Yeah. But you didn't mention any features that are not supported.
Features? Or, um ... The unsupported stuff is mostly, VM profiles.
So security VMs, so if you think about, TDX, SEV, they encrypt the memory, right? So we can't actually see the memory pages to classify them, so that's one thing, you know, I've been g- been, been getting a lot of questions about lately. Okay.
Uh, those are not supported. Monster VMs, latency VMs because of performance. Uh, so version, this is the first version.
There's about maybe five VM profiles not supported, but next version will be less and less and less. Okay, great. So for those type, for those, my graph, for those types of VMs, do you have to ...
If you had to configure to enable it at the cluster level, would you have to then go disable it on those specific- At the VM level, yep. Okay. Yeah, you can go to the VM level or have dedicated host, for those.
Okay. Yep. So- And ...
Oh, sorry. I was just gonna ask, and then what happens if the NVM tier becomes full? Like, does it fall back to compression, or does it- Yes ...
get swapped? Yep, it goes down the list of, reactive measures there. Yep.
But it, it sounds like this is going to potentially conflict with the guest OS's own concept of memory management. Are there specific use cases that are ... should or should not be used for this, that you should not use this or should use this for?
Um, so some that we've seen so far, like in, in-memory databases, obviously that's, something I wouldn't use for, at least, yet. Um, I, I did mention the latency sensitive. Um, there, there's some.
There, we have a list on our performance guide that I can share. Right. But what's ...
I mean, what is your ... How do you un- how do you understand what the operating system is doing in measuring ... in, in managing its own memory pools versus what you're doing?
It seems like you're in conflict, that if the m- the, the, the guest OS, right, if you're running a Linux system, it's got its own concept of all the memory, the physical memory that it has. It may not be, you know- Mm ... it's virtualized physical memory, right?
Mm-hmm. And now you're virtualizing physical memory again. But that OS has a concept of what it thinks is going to be used, and it's doing memory ca- page replacement based on what it knows is going to be used or not used.
So in a normal situation where it's not just a, you know, an in-memory database is just a pile of memory. But for a normal application running here, you're, you're potentially in conflict with what the OS is doing. Um, yes and no.
I mean, we do it at a higher level, and because that's transparent to the guest OS, pretty much what we're doing, right? So we're trying to be proactive. If you know ...
I mentioned the patterns and all that, those things, right? If we know that page will may be read soon, we may able to move it before it gets called out from the VM. So there's, there's lot of ...
I can't talk about the algorithm itself, but yeah, we, we do it at a higher level and try to be proactive about what the VM is trying to do, if that makes sense. I can share more info with you later. Thank you.
Thanks. Uh, the other question I get is, security, right? It's a, it's a C drive.
I can walk into a data center, pull out a drive, and get some, some memory pages out of that. You know, point taken, so we do support encryption. Uh, we support it at two levels.
" So all those VMs within the host are gonna get their pages encrypted when they move down to NVMe. Also, we can do it at the VM level. Uh, so we can go VM by VM, enable me- encryption for NVMe if you want to.
Now, this also works with all the other encryption stuff that we have, VM encryption, vSan encryption, vMotion encryption, all that stuff. There are separate layers everywhere. Uh, it works, again, at a different layer, and it is supported with the entire ecosystem.
0 and, and forward, right, we're leveraging VMware configuration profiles, which is kind of like host profiles, but for the entire cluster. Uh, so here what we're doing is just, just going in, enabling memory tiering, and, you're passing that out to all the hosts. You can o- obviously get some, host overrides if you don't want to, to do that, but it would be ...
That's just the easiest way you can do it. You can do, PowerCLI, ESXCLI commands as well if you wanna, if you wanna go that route. Uh, so the advantage of using VCP configuration profiles is that you set it, and then it lets it do its thing, right?
0, it does require a maintenance mode reboot, and, but it will roll through all those, based on whether you have vSan or not, what VMs needs to be moved, et cetera. So once you set it, it does all the stuff for you. Okay.
Uh, some failure scenarios, obviously with RAID, if one device fails, it will go to the other device, right? Pretty straightforward. But what if I'm using a single device, which, you know, some people don't like to use in controller.
I understand why it's an additional cost. Now you get a new point of failure there. Uh, you know, some, there are some other cons in there.
But if you have a single device, which we support, the ... an HA event will be triggered on the VMs. And I say VMs because only the VM pages are going to NVMe.
All the p- the pa- memory pages for the kernel, they stay on DRAM always. So if we have a failure on a single device, the host keeps running fine. The VMs may or may not have an HA event, and I say that because they will fail when the VM tries to call down those pages that are inactive, right?
So, as soon as you have miss a, lose a drive, some of those VMs may go down right away, you know, they reboot somewhere else. Or, you know, 10 minutes later you may see a couple more, or an hour later, or never, right? If that VM gets rebooted before any, any of that.
So, yeah, it's not a full crash per se, but it kind of stagger i- in a way, depending on the workload, what the workloads are doing. All right, and I wanted to leave you with this. There's a lot of information here.
The, the white, performance white paper I was talking about that has a ton of information, it's here. Uh, so on the performance, there's a blog and a white paper. Definitely, take a look at that.
I have a series for, a, a blog series for deploying considerations. I talk about what devices to use, why, sizing, what qualifies a workload to be, compliant for this, et cetera. So definitely, a lot of information here.
Hands-on lab, if you wanna kick the tires, configure it, you know, play with it, you know, definitely, take advantage of that. All right, so to recap what we talked about, you know, main driver here is to help our customers lower their TCU, and we're seeing this up to 40%, even ... This was even before the RAMpocalypse, right?
So, we're ... A lot of our customers now are very active, actively looking to deploy this, you know, to save a lot of that, that cost. VM consolidation, again, when we saw that, that diagram, moving pages down to NVMe allows us to free up DRAM, be able to utilize more of that CPU resources so, we have more, more VMs running on the same host, saves us an extra, some extra money there by not having to buy additional servers.
And yes, being able to use all the resources more efficiently, DRAM, CPU, NVMe, all that stuff. Any last questions before I get cut off? Just I wanna make sure I heard this right, that this feature is both in VCF and VVIA?
Yes. It's a, a vSphere component pretty much. Okay.
Yep. It's at the core. Yep.
All right, so with that, I wanna thank you all. Great questions, I appreciate it. And, if you need more information, happy to help.
Thank you.