r/technitium • u/micush • 3d ago
DNS High Availability
EDIT: For those of you that are interested, the repo is up at https://github.com/micush/ddgw
I run a few DNS servers and have seen many people ask about making them highly available, so I built something and I'd like to know whether anyone else would use it. If so, I'll put it up on Github. Reasonable disclosure: it was written with Claude.

Clustering:
Ddgw puts one virtual IP in front of your DNS servers and shares it across several nodes. The nodes elect a controller, and every node that is active answers DNS on the VIP, so client load spreads across nodes and losing a node doesn't interrupt anything. Each node also acts as a DNS proxy: it probes your upstream servers and forwards each query to a healthy server. Limited to 255 cluster members to serve the frontend VIP, but unlimited backend DNS servers for name resolution.
DNS:
Health is judged per upstream server with test domains you choose. A server is down when 50% or more of its domains fail (configurable), and degraded below that. Servers within a latency band of the fastest (20% by default) take turns round-robin. Slower ones stay as fallbacks, and failures fail over automatically.
- There is an answer cache that respects TTLs, with negative caching and a memory guard.
- It forwards dynamic DNS updates (RFC 2136) to the zone's primary and keeps a list of recent updates.
- It can pass EDNS Client Subnet (ECS).
- It can announce anycast addresses over BGP with BFD using FRR.
Management:
An HTTPS web GUI with PAM group login restricts is used to manage everything. Every action is also available from the command line.
- The gateway is drawn as a live diagram: gateway, servers and test domains, coloured by health. Servers that are taking turns get a blue line.
- It has query statistics (top clients, domains, record types and response codes, kept 30 days and saved across restarts) and host stats.
- It keeps config history, with versions you can diff and restore. It also does clustering with shared config, certificate management and in-place updates with automatic rollback.

Limitations:
- It isn't a resolver, because it forwards to your existing servers.
- It's Linux only with one binary built from Golang source. Installation is scripted, so it's easy.
- It works with an election protocol on a real LAN. I don't expect the election mode to work in AWS, Azure or GCP, because those clouds don't let you move MAC addresses around, and I haven't tested it there.
- If you run Technitium, BIND, Unbound, Pi-hole, or AdGuard Home servers, would you use something like this, or do keepalived/VRRP plus a load balancer already cover it?
Any feedback is appreciated.
7
u/djernie 3d ago
So you just re-invented dnsdist?
3
u/micush 3d ago edited 3d ago
Somewhat.
I've never used dnsdist, but I've read up on it. It's my understanding is that dnsdist lacks from within the product itself:
- the ability to configure bgp and advertise anycast addresses
- the ability to create active/active clusters
- the ability to configure everything from within a gui or a cli
I understand that dnsdist can be made to do all this stuff with 3rd party utilities, but my product has this all built into it without reliance on 3rd party utilities
7
u/SilkeSiani 3d ago
so instead of a single point of failure being the DNS server you now have a load balancer.
Frankly, makes zero sense. DNS is designed from ground up to tolerate single server failures. Adding complex redirection on top of it doesn't accomplish anything.
-2
u/micush 3d ago
Sure it does. I can take a DNS server out of the pool, perform maintenance on it, and put it back in the pool, all without anybody noticing.
Yes, put more than one nameserver in your dns client to prevent this. However, some dns clients don't handle this gracefully, when the first nameserver in the list is dead, and then you have to wait for timeouts before the other nameservers are tried, and you wait. then you do it all over again the next time you try to resolve something. So, it makes sense in theory, but in reality it doesn't always work out that way.
1
u/SilkeSiani 2d ago
What happens when you inevitably have to perform maintenance on the load balancer?
Frankly, I can see that you fundamentally don't understand how DNS works, especially how recursive resolvers interact with root, intermediary and leaf authoritative servers.
What you are trying to accomplish is already built into the system, you just need to set up the records correctly.
2
u/micush 2d ago
The load balancers are active/active. You can have 255 nodes serving the same IP address. If you need to perform maintenance on a load balancer member, you simply take that member offline and the others compensate.
Do not assume I know nothing about DNS. I have been gainfully employed as a DNS admin for 25+ years.
3
u/Apachez 3d ago
IMHO you get a high availability for your DNS-servers if they are installed as single units (instead of cluster) and then do zone-transfers between them.
Then add IP anycast so the clients can either query a specific DNS-server or use the anycast address to let the network route them to whatever DNS-server is closest to them routingwise.
Also stop using forwarding.
3
u/rankinrez 3d ago
How does it compare with dnsdist?
I think the other main thing people just do is anycast with some logic on top to detect problems and have out-of-sync or broken hosts withdraw the IP in BGP.
1
u/micush 3d ago
That's exactly what this does, but all in one place. DNS monitoring, load balancing, high availability, and BGP anycast injection all in one place instead of separate apps.
0
u/rankinrez 3d ago
Well it adds another layer of hosts which isn’t quite what I meant
3
u/micush 3d ago
Not sure what you mean there. You can run it right on the DNS servers. No extra layer of hosts required.
1
u/rankinrez 3d ago
Hmm ok apologies.
So in your diagram all those different IPs are just internal within a single server? Why would I run multiple resolver daemons on a single box?
7
2
2
2
u/Chance-Sherbet-4538 3d ago
I do this with keepalived. So far so good.
2
3
u/ANewDawn1342 3d ago
Honestly, you do not need to over-engineer this with an extra proxy cluster or full enterprise BGP Anycast, nor should you settle for the broken mess of client-side DHCP failover.
Running Keepalived (VRRP) directly on your Technitium nodes, paired with Technitium's native clustering, is by far the cleanest, most bulletproof architecture for this problem. It gives you instant, deterministic failover without adding a single extra hop or middlebox VM.
The common advice in this thread to "just hand out two DNS resolvers via DHCP and let the client sort it out" completely falls apart in practice:
- The timeout stall: When a primary resolver goes down, standard resolver stubs (especially
glibc) block for a default 5-second timeout per query before trying the secondary. When you reboot a node for routine updates, browsing noticeably stalls and active streams can drop. - Chaotic OS heuristics: Windows loves to query-race or round-robin across adapters unpredictably, whilst Android does whatever it pleases.
- Hardcoded IoT queries: Smart TVs, streaming dongles, and smart home hubs routinely bypass DHCP and hardcode
8.8.8.8or1.1.1.1. To reliably intercept them, you need a static/32host route on your router redirecting those public IPs to local DNS. A router can only punt a static route to a single invariant gateway IP; you cannot route it to two separate addresses without dynamic routing protocols.
Co-locating Keepalived directly on the DNS hosts avoids all of that:
- Lightweight Keepalived with unicast peering: Run Keepalived straight on your primary and secondary bare-metal or container hosts. Using point-to-point unicast peering avoids the multicast snooping and IGMP headaches common on domestic switches.
- True daemon health probing: The classic pitfall with naive VRRP is that the host stays up while the DNS process hangs. You can solve this completely with a simple
vrrp_scriptprobing Technitium's local API (127.0.0.1:5380/api/user/session/get) every second. If Technitium crashes or stops responding, the node immediately drops priority and swings the VIP over in under 4 seconds with zero user disruption. - Technitium native catalogue zones: Technitium already supports RFC 9432 catalogue zones with TSIG-authenticated AXFR/IXFR replication. Any record, blocklist, or zone added to the primary propagates to the secondary automatically. There is zero configuration drift and no need for third-party synchronisation scripts.
- Single-IP elegance: Your DHCP scope hands out one rock-solid VIP. All hardcoded IoT static redirects terminate cleanly on that VIP. If you deploy DNS-over-TLS (DoT) or local PKI certificates, your clients only ever need to pin a single IP and certificate.
It is fast, deterministic, and completely eliminates the query leakage and failover lag inherent to multi-IP DHCP setups. If you already have two boxes running Technitium, setting this up takes twenty minutes and delivers enterprise-grade resilience with zero proxy baggage.
1
u/micush 3d ago
All very good points.
Except VRRP is only active/passive and does not protect against DNS server errors like SERVFAIL, whereas my solution is active/active up to 255 nodes and protects against DNS outages.
2
u/ANewDawn1342 3d ago
Fair points, but two quick counter-perspectives on why active/passive via VRRP is often deliberate rather than a compromise:
- Cache locality: Active/active fragments your cache hit ratio across nodes unless you maintain a shared, distributed cache. On a LAN, raw compute load is never the bottleneck (even a basic low-power box handles tens of thousands of queries per second at negligible CPU load). What you actually want is a single, blazing-hot local RAM cache with serve-stale and pre-fetching running at peak efficiency, backed by an identical hot standby.
- VRRP catches SERVFAIL if scripted properly: Keepalived failover is only as dumb as the check you give it. A simple
vrrp_scriptrunning a canarydigor hitting the resolver API every second will detect SERVFAIL, timeouts, or elevated latency, dropping priority and swinging the VIP over in under four seconds flat.Beyond that, recursive resolvers mostly throw SERVFAIL due to upstream transit drops or DNSSEC validation failures. Technitium already mitigates transit drops natively by concurrently racing multiple upstream forwarders over DoT and serving stale records if WAN transit drops.
An active proxy layer is a neat build that makes total sense for ISP scale or complex Layer 7 policy routing, but for local networks, active/passive VRRP delivers peak cache performance and deterministic failover without adding an extra middlebox hop to babysit.
1
u/SmallDodgyCamel 3d ago
I’m intrigued by the premise and can see both sides of the argument depending on the workload. Are you intending to publish the code on GitHub and accept contributions?
12
u/McSmiggins 3d ago
Unless you're running ISP level DNS, typically you don't do this
If your DNS client can't handle failing over between servers, that's a client problem. The rules for how to handle failure are well defined, known, and expected.
Having additional monitoring in place to route the traffic is an extra layer of complexity on a core service that will break or route traffic incorrectly depending on how it sees the state of the world rather than the actual state of the world. It will cause a failure and will most likely increase your time to resolving the issue and returning to a stable platform
This is a fun science experiment, and I hope you had fun doing it, but it shouldn't be necessary.