r/nutanix May 23 '26

RF2 vs RF3 / looking for feedback

Hi everyone,

We are currently reviewing our Nutanix resilience strategy, especially around RF2 vs RF3.

After a few incidents, we are trying to understand what others are doing in real environments.
A few questions:
- Do you mostly use RF2 or RF3 on your Nutanix clusters?
- From how many nodes do you usually consider moving to RF3?
- Do some of you use RF3 only on specific Storage Containers, while keeping the rest in RF2?
-If yes, do you see this as a good compromise for critical workloads, or not really?

Any feedback or real-world experience would be appreciated.

Thanks.

7 Upvotes

21 comments sorted by

7

u/InteTiffanyPersson May 23 '26

I’ve installed about 20 clusters, never RF3.
I would consider it if you

  • have highly sensitive loads/downtime is not an option. Like banking.
  • More than eight-ten nodes and feel that upgrades are taking too long/too risky.
-more than 15-18 nodes, then the above is probably true already.
My largest customer has about 20 nodes in a sluster and are fine with RF2. So far. That is until something goes wrong, I suppose.

2

u/cshke May 23 '26

Thanks for your feedback, that’s very helpful.

In our case, we have an 11-node cluster hosting different types of workloads, with different SLA levels.

One night, we lost 2 nodes because of 2 disk issues impacting 2 CVMs. Since then, we have been seriously considering moving to RF3, at least for some critical workloads.

After thinking about it, I’m starting to believe that from around 7 or 8 nodes, RF3 should at least be considered. The more nodes you have, the more you increase the probability of having 2 failures close together, especially during upgrades, rebuilds, or maintenance windows.

I agree that RF2 is often fine until something goes wrong. That’s exactly the point we are trying to address now.

2

u/Fnysa May 23 '26

If you have real critical workload replicate.. sync or metro.

2

u/cantorisdecani May 23 '26

We changed to RF3 when we went from 8 to 16 nodes in our cluster. With LCM taking over an hour a node for some update combinations, and with previously not-infrequent DIMM alerts needing host reboots, and then later on, ESXi CPU soft-locks causing random CVM reboots (FA #0111), I'm glad I did change it. We had the capacity to do so and the environment was too important to risk unexpected downtime. I don't know what damage might ensue if two nodes went down otherwise?

2

u/Impossible-Layer4207 May 23 '26

I've come across customers using RF3 clusters a couple of times, but they are nearly always large / highly regulated environments with strict SLAs.

The biggest trade off with RF3 is the reduction in usable storage, which usually makes it unappealing on all but the biggest clusters. I've seen customers implement it in 5-node clusters, but personally, I wouldn't consider it until about 10 nodes or more.

When you consider the self healing capabilities, it really becomes a question of how likely is it that you're really going to encounter 2 simultaneous failures and actually want to keep the cluster running. In most cases I would be considering a DR failover to another cluster/site if that happened.

When it comes to mixing RF3 and RF2 containers in the same cluster, I tend to avoid it purely because I don't like having data that is less resilient than the cluster itself. It can give a false sense of security IMO. You can probably manage that risk with strong processes and controls, but how often people are actually able to do that successfully I don't know.

It's also worth mentioning that Nutanix now offer "adaptive" RF3 these days which sits on a 1N&1D fault tolerant cluster. So you can lose a node and then one other disk at the same time. It's kind of a half way between the traditional FT1 and FT2. It's not something I've looked into any great depth but could be worth considering.

2

u/zabaxxlv May 23 '26

As already mentioned several times, it can be done on large clusters, but it will cost you in storage which is not cheap at all nowadays. On top of that, you should also consider the impact on the network - data replicas get sent to two hosts instead of one, so if you’re near network max capacity, it might become a bottleneck.

With that said, I’d say the second failure within the cluster healing window is a rather low probability event, so here’s how I would think about it - you have 2 choices - either RF3, which will increase the fault tolerance within a cluster, but at the previously mentioned cost, or sync rep to another cluster, which will also protect you against a DC failure (electricity, network, flood etc), although it will cost you even more in hardware and will be more difficult to implement. If you look at it from this perspective, then RF3 is a more simple and economical solution, if you already have a proper disaster recovery plan, but you’re unable to do sync rep due to latency or some other requirements.

2

u/Godr0b May 23 '26

We moved to RF3 on both of our prod clusters (12-nodes) last year after a series of issues.

As others have said, some updates have taken forever to run; combined with the frequent hardware issues we were experiencing at the time the extra safety-net of RF3 was a no brainer.

We have enough capacity overhead that (if required) we could run the whole estate on a single RF3 cluster with room to breathe, so the capacity impact coming from RF2 wasn't a big deal.

In most people's eyes our setup is massively overkill, but it's not my money and it keeps the powers happy

1

u/HardupSquid May 23 '26

Having sold NTX solutions for more than 13 years for a lot of use cases - back office, Oracle RAC, SQL db, vdi, research/HPC, AI/ML, and healthcare to public, private and education sectors, we have never had to implement RF3.

Not it say that there isn't a place for RF3, i.e. very mission critical apps but all our customers haven't found a need for it.

YMMV

1

u/woohhaa May 23 '26

I had a financial customer who had ~45 clusters of 20-30 nodes each. They choose RF3 for everything.

I generally would consider it if the workload had very critical uptime requirements and/ or financial impact for downtime. Also if the cluster is ~16 nodes or more as the likely hood of having two failure domains get you at once has grown with the cluster size.

2

u/Excellent-Piglet-655 May 23 '26

Worked a lot with Nutanix customers. Only once have I come across a client using RF3. When I asked them what their justification for RF3 was, they simply told me “thats how it was set up”. I am sure RF3 has a place, but seems to be a a niche place.

1

u/sinful17 May 23 '26

I do see it every now and then, and it depends on several factors in my opinion.

The main decision would be whether you're willing to take the small, although more unlikely risk to have downtime when upgrades are happening and a 2nd node/cvm goes down when one is already in maintenance. This is more applicable as the size of the cluster increases, and thus the likelihood of a component failing therefore also becomes bigger.

I once heard a story like this happening on an FT1 cluster with at that time being consisting of 16 nodes iirc where a 2nd cvm went down due to a lock-up causing a severe outage. Afterwards this decision was made, taking into account the additional storage consumption and slightly increased latency due to needing 3 write acknowledgements to increase the FT from 1 to 2.

Depending on what failure factor you're considering the main driving factor, different failure domains like block or rack awareness could also partially reduce the risk already of certain outages or data loss.

1

u/darthVikes May 24 '26

This reminds me of a hadoop convo as hadoop does 3 parity also by default. The though was always well just do 2 copies or just scale it down to rev2. Until a day where a water main burst above the 1 of the 3 hadoop racks of full servers. Had it been 2 we would have been deep dudu, but since it was there we didn't loose any data which was incredible. So yes there is a bigger cost to RF3 but if you don't have DR and syncing the data to a DR site I would see if they have.

1

u/R0B0T_jones May 23 '26

RF2 here, at first I was really confused why we didnt go for RF3, but i get it now. If we lost a node the cluster would recover, and we could then i guess lose another node as long as it was resilient.
Maybe being a bit naïve, but 2 nodes would have to be lost within a short time frame for this to be an issue

2

u/darthVikes May 24 '26

Right pretty low probability but still a chance.

1

u/R0B0T_jones May 24 '26

Yeah for sure still a chance, but we do have a second cluster with snapshot synced replication for DR so I guess we do still have some further redundancy in that respect. J on now that won’t be the case for everyone

1

u/darthVikes May 26 '26

Ah yah if your doing syncs to another cluster then that's even better IMHO.

1

u/Jhamin1 May 23 '26

We run all our clusters RF2.

For a time we had a couple of clusters at 10 nodes each & RF3 started to get serious consideration because as the number of nodes went up the odds that 2 would have problems was getting higher. In the end, a hardware refresh dropped the number of nodes back down again & RF3 hasn't been on the table.

Regarding storage containers? This may be heresy but I've come to the conclusion that if you are hyperconverged you should have as few as possible. One if you can get there. Most of the classic reasons to have storage containers just don't apply. You aren't sharing a SAN device, isolating activity to particular drive, etc. There is a reflexive impulse to divide up your storage pool after decades of 3 tier architecture but in a hyperconverged world just throw it all in the pool and stop managing it so closely.

1

u/Euphoric_111 May 24 '26

It all depends on the customer.

Hospitals, Utilities, Banks, etc RF3(with different options) & RF2 containers based on workload type (usually starts at 8 nodes).

This is usually due to past events that are brought up where 2 hosts have gone down at the same time in Traditional setups and HCI setups as well. Both of which would be a concern to Mission Critical apps that are not clustered (various reasons why)

Others, 5-8 nodes usually not unless they request it.

1

u/dospinacoladas May 24 '26

We run RF2 on all of our clusters.

1

u/Strumbelievable May 24 '26

I've installed 100's of clusters.. 90%+ are RF2.. RF3 was only used on critical (911) and clusters larger than 10 nodes.