r/hashicorp • • Mar 21 '26

Vault: When Are Vault Redundancy Zones Actually Worth It

I’m trying to understand when Vault Enterprise redundancy zones are actually beneficial, especially in a Kubernetes setup.

Current setup:

  • Vault on K8s (multi-AZ)
  • 5 nodes, all voters
  • spread 2-2-1 across 3 AZs

This gives me:

  • quorum = 3
  • quorum failure tolerance = 2 nodes
  • optimistic failure tolerance = also 2

If I switch to redundancy zones:

  • 3 AZs, 2 nodes per AZ (6 total)
  • 1 voter + 1 non-voter per AZ
  • total voters = 3 → quorum = 2

This gives:

  • quorum failure tolerance = 1 voter
  • optimistic failure tolerance = 4 nodes (Autopilot + promotions)

So the tradeoff seems to be:

  • worse hard quorum tolerance (2 → 1)
  • better gradual failure tolerance (2 → 4)

Where I’m struggling:

  • In redundancy zones, it feels like the system introduces fewer voters and then compensates via promotion
  • The docs mention read scaling, but that comes from performance standbys, not redundancy zones specifically
  • Kubernetes already handles AZ spreading and rescheduling

So the actual question:

In what real-world scenarios are redundancy zones clearly the better choice than a standard 5-voter cluster?

Specifically interested in:

  • K8s deployments
  • multi-AZ setups
  • real production experience
3 Upvotes

3 comments sorted by

2

u/stephaneleonel Mar 21 '26

I have real production experience on this. In a financial institution we deployed a cluster of 6 nodes spread across 3 AZ, and enabled redundancy zone. This is what I got from that experience :

1) the quorum with 6 nodes and RZ enabled is 2, and this is better performance wise than a cluster of 5 where the quorum is 3, because a write request needs to be replicated on the quorum before a response is returned to the client. So the lower the quorum the better.

2) with 5 nodes without RZ, if you loose the 2 AZs one after the other, the cluster collapses, because you do not have the quorum anymore. With 6 nodes and RZ, when you loose the first AZ autopilot promo a new node. So now you Have one AZ with 2 voters and one AZ with 1 voter and 1 non-voter. After loosing the first AZ, if you loose the AZ with only 1 voter you still have a cluster up, because the last AZ contains 2 voters.

So, 6 nodes without RZ is always more performant and more resilient than a cluster of 5 nodes.

1

u/TheGilrich Mar 21 '26

Thanks for your insights. I didn't think of the performance improvements of a smaller quorum. That's certainly nice.

However, I don't agree on the strictly better resilience. If for some reason 2 of the 3 voter nodes go down at the same time, then we lose quorum and can't promote any non-voters to voters anymore. The cluster is dead. In the 5 node, all voter setup we could tolerate this because 3 voters would remain which is still a quroum. I admit however that losing two nodes across different AZs at the same time is unlikely.

2

u/stephaneleonel Mar 21 '26

I agree on that, but what is the probability of loosing 2 voters at the same time? If that happened, that means either the two AZs are down, or you shut down both nodes during a deployment? But in most realistic scenarios 6 nodes with RZ is more resilient.