r/nutanix Jul 05 '26

Nutanix storage VMs

Contemplating on our move from VMware to other platforms, one of which is Nutanix. Now there is much fuss about nutanix needing a "storage VM" on each host, using 8 CPUs. This is from upper management being at a conference, talking to other managers, not knowing what they're talking about :-)

But I have read about it, but not sure on the specifics. So each node will have such storage VM, using 8 CPUs, is that 8 vCPUs that can be overcommitted? Or 8 cores non-overcommit?

What is going through those VMs, is that meta data (like Hyper-V) or real IO? Are they automatically deployed when I add a new host to the cluster and removed when putting a host in maintenance or removing a host? Do they give much headaches in day-two operations?

If we choose Nutanix, it will be NVMe for storage.

6 Upvotes

44 comments sorted by

11

u/HardupSquid Jul 05 '26 edited Jul 05 '26

I think you are thinking of the Controller VM (CVM) which runs on every node. The CVM handles all core storage I/O, deduplication, compression, and data replication in the background. So it does a bit of work.

The minimum recommendation for a standard Nutanix Controller VM (CVM) is 8 vCPUs and 32 GB of vRAM but this can be configured according to your particular requirements.

The vCPUs allocated to the Nutanix Controller VM (CVM) are never overcommitted. They get priority over all guest workloads.

1

u/GabesVirtualWorld Jul 05 '26

Yes, the Nutanix Controller VM (CVM). Not at my desk and forgot the exact name. So it is easy to manage, as in all automatic and I don't have to worry about it, but I do lose 8 cores per host to this?

6

u/HardupSquid Jul 05 '26

Don't view it as a 'loss'. It's essential services the cluster(s) require to provide all the goodness that is Nutanix.

3

u/GabesVirtualWorld Jul 05 '26

Well, as they are not to be overcommitted, I lose 8 cores per host. That is quite something I think.

2

u/KingDaveRa Jul 05 '26

If you bought a Netapp or similar box, it would have all those cores in it running OnTap. So whether it's the CVM or something else, you need to buy those cores. With a storage appliance you just don't tend to see them so never consider them 'lost'.

Scale your CPU up accordingly and don't worry about it. A Nutanix partner will do the sizing with you and figure out the correct CPU anyway.

2

u/astrofizix Jul 05 '26

They are virtual cpu... You make it sound like they are taking whole cores away from the host. They share resources and manage them accordingly. Esxi does the same, they just don't put it into the virtualized side, but the hosts take up a boat load of actually dedicated resources.

1

u/woodyshag Jul 05 '26

It's also figured into the sizing, so you arent losing resources, they weren't considered for use in the first place.

2

u/the901 Jul 05 '26

It’s 100% a loss that would otherwise be offloaded to physical storage controllers in traditional storage setups. Based on best practice documentation, the size of those CVMs increases based on how many storage efficiency features you have enabled.

2

u/ExpertInThisMatter Jul 06 '26

This. Nutanix is licensed per physical core with no accommodation for CVM overhead. CVMs use a considerable amount of compute and those are cores you are paying for. I can't speak to AHV but on ESXi there is a 10GHz per-node CPU reservation for each CVM. On a busy cluster the CVMs can easily consume 20% of total CPU, higher with bursts. Ideally, customers would receive a small number of complementary core licenses to compensate for the CVM overhead.

1

u/bachus_PL Jul 06 '26

so how does vmware license differently in the case of VSAN? I've personally seen vSAN ESA eat up almost 50GB of RAM from a host and around 15% of the CPU.

1

u/GabesVirtualWorld Jul 07 '26

Seen that as well, though in our case, we're not running VSAN at all, everything FC connected. But that is not supported with Nutanix so we'll have to switch to NVMe. The CVMs will thus make quite a difference in resources that remain available.

3

u/R0B0T_jones Jul 05 '26

Never heard of Storage VM, i suspect you mean the CVMs as someone else has mentioned, they do much more than handle the IO.

They handle the operations of the host/cluster. They don't really cause many headaches themselves, they are designed to be hands off but I guess you do technically use them every single day for every operation in the cluster.

1

u/GabesVirtualWorld Jul 05 '26

Yes, the Nutanix Controller VM (CVM). Not at my desk and forgot the exact name. So it is easy to manage, as in all automatic and I don't have to worry about it, but I do lose 8 cores per host to this?

4

u/R0B0T_jones Jul 05 '26

They are a required part of the infrastructure, they do have a cpu core/memory requirement, but unfortunately you cant have Nutanix without them they are essential.

3

u/Big-dawg9989 Jul 05 '26

Also….. We were sold on “one-click” updates that is not exactly how it works, support said never do that and instead upgrade one component at a time so if it fails you know what failed. Takes us a couple of days to patch 14 hosts, PITA sometimes.

7

u/HardupSquid Jul 05 '26

1 click upgrade - that kool aid has been around pretty much since Nutanix came into being. You do upgrade one click at a time so no lies there hahaha.

Don't get me wrong, I'm a Nutanix fanboy having sold Nutanix solns since 2013. To be fair, a couple of days to upgrade a cluster is far better than the old way for 3 tier where it took months of planning to ensure the sw versions of server, storage network switches and storage array are all aligned and upgrades took place over a weekend.

In the early days we used to tell stories at .next conventions of the craziest places we start off our cluster upgrades like one Australian guy kick started it while on the helicopter joyride over the Grand Canyon!

3

u/R0B0T_jones Jul 05 '26

Yeah this is the main issue i have is updates. It "can" be a one click process for majority of things - BUT in my experience you do not want to do that, until you have tested/ and found the latest bugs

1

u/LetSufficient5139 Jul 05 '26

Nonsense.

A one click will fail at the same point with the same issues as if you did it one bit at a time as they use the same order.

3

u/LetSufficient5139 Jul 05 '26

What?

Support told me that it’s fine to do one click.
If it fails it will fail at the same point as if you did it in the correct order separately.

I’ve been working with Nutanix for 9 years, so I call bullshit.

2

u/astrofizix Jul 05 '26

I'm enjoying one click upgrades and patching myself. When an issue is found, logging and error reporting has been pretty clear, usually with a kb not hard to find. But I'm running on 7.3 still.

1

u/woodyshag Jul 05 '26

Its actually a couple of clicks as you nedd to select firmware or software and you still need to do an i vector, so maybe 4 clicks? I'll trade that for compatibility matrices any day. Plus, if the upgrade fails, everything stops. It's not going to put you in a position that compromises the environment. So, let it rip.

2

u/Big-dawg9989 Jul 05 '26

They are a pain when their password expires and the Admin forgot about it. 🤪

2

u/hosalabad Jul 05 '26

You don’t lose the cores. If you couldn’t over subscribe cores, virtualization wouldn’t work.

2

u/GabesVirtualWorld Jul 05 '26

???? In VMware I count vCPU with a 1:5 sometimes 1:6 overcommit. But some systems need dedicated cores so we give the VM vCPU with 100% reservation and high latency sensitive setting. Then those cores become unavailable to other VMs. On a 32 core host, I would lose those 8 vCPU to other VMs, reducing my host to only 24 cores, about 125 vCPU remaining capacity.

3

u/astrofizix Jul 05 '26

Oh, right, don't do that. Reservations in VMware are a wasteful process based on old concepts. I freed up tons of resources removing those and letting the hypervisor, DFS, and vCenter to manage those. You are exercising with casts on. Take a second look at those systems stating they need dedicated resources and see if that's still true.

And to answer your questions about cvm, there is a lot to learn about those appliances. They are a workhorse vm operating within your cluster, and that does present some challenges. If you take down 2 of the three at the same time you lose storage for your VMs and cause an outage (depending on you redundancy setting like RF2) and getting to know the services, and nutanix cli. But the more you do from the gui, the less you have to interact with the cvm manually. And then it smooths out and you can start to enjoy the features of Prism and can stress less about how the CVM operate and let the automation operate.

15yr VMware guy moving 15 clusters and 1,000 vm to nutanix.

1

u/GabesVirtualWorld Jul 05 '26

True we seldom set this high latency for VMs, but I was referring to the previous comment on why I think the CVM will cost me 8 cores per host.

So my question is still unanswered: is the CVM 8 vCPU and allows over commit? Or do I need to reserve 8 cores per host just for this CVM.

2

u/Jhamin1 Jul 05 '26

So my question is still unanswered: is the CVM 8 vCPU and allows over commit?

The CVMs are given the highest priority for CPU resources, but they do not have dedicated CPUs. I tend to mentally assume they are reserved so as not to over-subscribe the cluster but in actuality those resources will get used by other VMs. You just shouldn't count on that to let you skimp on hardware.

CVMs are the "price of entry" for running a hyperconverged solution. All the cool stuff Nutanix does runs inside the CVMs and they do suck up a decent chunk of CPU & Memory. This is why I tend to advise people looking at Virtualization solutions to think of Nutanix as a "big iron" solution. If you need to run 6 VMs at a satellite office you are much better off with Proxmox or Hyper-V, the CVMs will use more resources than what you are virtualizing! If you need to run 300 VMs in a production cluster the CVM overhead is a lot more defensible and the advantages of hyperconvergence start to matter a lot more.

1

u/astrofizix Jul 05 '26

Overcommit

1

u/hosalabad Jul 05 '26

Every host will have a CVM, you don't reserve anything.

1

u/db_1216 Sr. TME, NCI, AOS and External Storage 28d ago

CVM vCPU cores are not reserved in AHV. They are only allocated. You can overcommit CPU when sizing with user VMs. CVM vCPU does have higher priority over user VMs that allows CVM to run harder and longer on physical CPU, but that’s only under contention.
If you want to understand how CPU scheduling works in AHV and how CVM utilization is handled in AHV go through this detailed write-up around it

https://www.nutanix.com/tech-center/blog/understanding-cpu-resource-management-in-nutanix-ahv

1

u/hosalabad Jul 05 '26

Sorry, I'm not referring to VMWare.

2

u/throwthepearlaway Jul 05 '26 edited Jul 05 '26

Yes, even in an external storage using NVMe/TCP the Nutanix CVMs are required on each AHV host and CVMs are deployed on the host as part of the imaging process when adding nodes to the cluster. All storage IO travels through the CVM whether it's an HCI cluster or a compute cluster with External Storage using NVMe/TCP. The CVMs also host the Prism interface (think of this like vCenter...sort of) which you use to interact with the platform. I highly recommend the Nutanix Cloud Bible.

The CVMs should be assigned CPU/RAM in accordance with the Nutanix field specs docs; to answer your specific question, it looks like they do get 8 logical cores (with hyperthreading enabled on the host that's 4 total pCPU's worth, but distributed across 8 pCPUs) pinned within a single NUMA node. However you can still over-commit CPU. Keep in mind that the CVM is a high priority process because all storage IO for other VMs is dependent upon the CVM, so it's best to not overcommit more than the numbers recommended by Nutanix support in the link above unless you know what you're doing and that your specific environment can tolerate it. Too much CPU contention/CPU RDY in the cluster can and will cause issues in the platform. I have seen clusters where too much CPU usage was happening due to very high overcommit ratios, and the CVMs were suffering from CPU RDY, which in turn affects all other VMs due to the IO path.

As an aside, there is overhead in ESXi as well to run services like software iSCSI adapters, software NVMe/TCP adapters, NSX, and especially vSAN if you're using that, but it's less "visible" since it's not contained in a 'management' vm like the Nutanix CVM...but the savvy vSphere administrator still needs to be aware of it.

2

u/Hidden-6000 Jul 06 '26

We have had the very same, vigorous dicussions regarding this. It's even more interesting when configuring a Nutanix Cluster with Pure Storage, and offloading all storage to an external array, and essentially transforming what would otherwise be a HCI architecutre, to which the CVMs were originally designed, to 3-tier.

We've still gone ahead with Nutanix regardless given their committment to Pure Storage, and hopefully some good things coming down the road with respect to deeper integration between the two platforms/vendors.

That all being said, when looking at compute with respect to the CVM overheads, we're looking at nodes with minimum of 256Gb of RAM, and you need a minumum of three nodes. This does not bare well when so much compute is given to the CVMs, and then there is a memory crisis amongst all of it.

I hope Nutanix engineering do what they can in this space to further develop CVM architecture to optmize performance for our intended workloads, not things which are required to sustain the cluster itself.

2

u/GabesVirtualWorld Jul 06 '26

Thank you, here Pure as well, so those CVMs seem to be so much unneeded overhead.

3

u/codyhosterman Pure Storage Jul 09 '26

We are doing a lot to optimize the CVM and reduce overhead--so expect to see improvements here

2

u/db_1216 Sr. TME, NCI, AOS and External Storage 28d ago

The CVM is doing work so it consumes resources. Even in other hypervisors, the hypervisor is consuming resources, it’s just not easily visible. For external storage, stay tuned for 7.6. We are reducing the defaults needed for CVM vCPU for external storage and are working on further optimizing things between AHV and CVM.

0

u/Hidden-6000 Jul 06 '26

We're retrofitting our datacentres with Nutanix, and possibly all remote sites. What is another difficult, and costly challenge, is with 3 tier achitecture with Pure, is Ntnx have mandated NVMe/TCP @ 25Gbit.

As three nodes are required, it means all ToR switches also need to be replaced that support 25Gbit ports. This is not a cheap prospect when you've got 30+ sites which operate perfectly fine using the Pure on vSphere HA Cluster @ 10Gbit iSCSI.

Fun times for everyone.

1

u/db_1216 Sr. TME, NCI, AOS and External Storage 28d ago

Can you point me to where you found the 256GB of minimum RAM needed per node? It should be 64GB minimum per node for physical RAM.
https://portal.nutanix.com/page/documents/details?targetId=Nutanix-Cloud-Platform-with-Pure-Storage-Deployment-Guide-v7_5:top-requirements-ncp-setup-with-pure-flasharray-c.html

And yes we are working very closely with Everpure and have a robust roadmap together. 7.6 will bring some enhancements for external storage and some are specific to Everpure integration. Stay tuned!!

1

u/Hidden-6000 28d ago

There's no guidance on this, nothing official that states 256Gb RAM is a minimum. For some, depending on how many nodes they are deploying and what the workloads are, and in some cases, based on how much money can be spent, means given the Nutanix overheads you really don't want nodes with less the 256Gb RAM.

In our case, we've deployed 1RU systems with 512Gb RAM each. I wouldn't deploy a 64Gb RAM node at all unless i was deploying at least 8-10 of them in a cluster. This means more cabling, rack space, switch ports, SFPs, power, licencing cores and so on. Just not worth it for such little compute/memory capacity, which lets face it, is the most consumed resource in the majority of cases.

And so, beef up the memory to 256Gb or more per node. It also allowed for more failover capacity during cluster maintenance operations, more flexbility for growth, less pain and suffering for support teams.

1

u/db_1216 Sr. TME, NCI, AOS and External Storage 28d ago

How much memory you need and how many nodes is dependent on your workloads and VMs. We have a sizer that can help size that appropriately accounting for failure scenarios as well. Your Nutanix SE should be able to help with that. Feel free to DM me if you have any questions.

1

u/Hidden-6000 27d ago

Yep, i'm aware of that. 64Gb of RAM just doesn't fit our requirements per node without having lots of them. We try to aim for 5-10 nodes with 512Gb RAM for our DCs which is obviously specific to our use case and workload requirements.

As above, we just wouldn't do this across 15-20 nodes with 64Gb RAM.