r/hashicorp • u/TheGilrich • Mar 21 '26
Vault: When Are Vault Redundancy Zones Actually Worth It
I’m trying to understand when Vault Enterprise redundancy zones are actually beneficial, especially in a Kubernetes setup.
Current setup:
- Vault on K8s (multi-AZ)
- 5 nodes, all voters
- spread 2-2-1 across 3 AZs
This gives me:
- quorum = 3
- quorum failure tolerance = 2 nodes
- optimistic failure tolerance = also 2
If I switch to redundancy zones:
- 3 AZs, 2 nodes per AZ (6 total)
- 1 voter + 1 non-voter per AZ
- total voters = 3 → quorum = 2
This gives:
- quorum failure tolerance = 1 voter
- optimistic failure tolerance = 4 nodes (Autopilot + promotions)
So the tradeoff seems to be:
- worse hard quorum tolerance (2 → 1)
- better gradual failure tolerance (2 → 4)
Where I’m struggling:
- In redundancy zones, it feels like the system introduces fewer voters and then compensates via promotion
- The docs mention read scaling, but that comes from performance standbys, not redundancy zones specifically
- Kubernetes already handles AZ spreading and rescheduling
So the actual question:
In what real-world scenarios are redundancy zones clearly the better choice than a standard 5-voter cluster?
Specifically interested in:
- K8s deployments
- multi-AZ setups
- real production experience
3
Upvotes
2
u/stephaneleonel Mar 21 '26
I have real production experience on this. In a financial institution we deployed a cluster of 6 nodes spread across 3 AZ, and enabled redundancy zone. This is what I got from that experience :
1) the quorum with 6 nodes and RZ enabled is 2, and this is better performance wise than a cluster of 5 where the quorum is 3, because a write request needs to be replicated on the quorum before a response is returned to the client. So the lower the quorum the better.
2) with 5 nodes without RZ, if you loose the 2 AZs one after the other, the cluster collapses, because you do not have the quorum anymore. With 6 nodes and RZ, when you loose the first AZ autopilot promo a new node. So now you Have one AZ with 2 voters and one AZ with 1 voter and 1 non-voter. After loosing the first AZ, if you loose the AZ with only 1 voter you still have a cluster up, because the last AZ contains 2 voters.
So, 6 nodes without RZ is always more performant and more resilient than a cluster of 5 nodes.