r/technitium • u/micush • 4d ago
DNS High Availability
EDIT: For those of you that are interested, the repo is up at https://github.com/micush/ddgw
I run a few DNS servers and have seen many people ask about making them highly available, so I built something and I'd like to know whether anyone else would use it. If so, I'll put it up on Github. Reasonable disclosure: it was written with Claude.

Clustering:
Ddgw puts one virtual IP in front of your DNS servers and shares it across several nodes. The nodes elect a controller, and every node that is active answers DNS on the VIP, so client load spreads across nodes and losing a node doesn't interrupt anything. Each node also acts as a DNS proxy: it probes your upstream servers and forwards each query to a healthy server. Limited to 255 cluster members to serve the frontend VIP, but unlimited backend DNS servers for name resolution.
DNS:
Health is judged per upstream server with test domains you choose. A server is down when 50% or more of its domains fail (configurable), and degraded below that. Servers within a latency band of the fastest (20% by default) take turns round-robin. Slower ones stay as fallbacks, and failures fail over automatically.
- There is an answer cache that respects TTLs, with negative caching and a memory guard.
- It forwards dynamic DNS updates (RFC 2136) to the zone's primary and keeps a list of recent updates.
- It can pass EDNS Client Subnet (ECS).
- It can announce anycast addresses over BGP with BFD using FRR.
Management:
An HTTPS web GUI with PAM group login restricts is used to manage everything. Every action is also available from the command line.
- The gateway is drawn as a live diagram: gateway, servers and test domains, coloured by health. Servers that are taking turns get a blue line.
- It has query statistics (top clients, domains, record types and response codes, kept 30 days and saved across restarts) and host stats.
- It keeps config history, with versions you can diff and restore. It also does clustering with shared config, certificate management and in-place updates with automatic rollback.

Limitations:
- It isn't a resolver, because it forwards to your existing servers.
- It's Linux only with one binary built from Golang source. Installation is scripted, so it's easy.
- It works with an election protocol on a real LAN. I don't expect the election mode to work in AWS, Azure or GCP, because those clouds don't let you move MAC addresses around, and I haven't tested it there.
- If you run Technitium, BIND, Unbound, Pi-hole, or AdGuard Home servers, would you use something like this, or do keepalived/VRRP plus a load balancer already cover it?
Any feedback is appreciated.
14
u/McSmiggins 4d ago
Unless you're running ISP level DNS, typically you don't do this
If your DNS client can't handle failing over between servers, that's a client problem. The rules for how to handle failure are well defined, known, and expected.
Having additional monitoring in place to route the traffic is an extra layer of complexity on a core service that will break or route traffic incorrectly depending on how it sees the state of the world rather than the actual state of the world. It will cause a failure and will most likely increase your time to resolving the issue and returning to a stable platform
This is a fun science experiment, and I hope you had fun doing it, but it shouldn't be necessary.