TinyDNS tinydns.org

How DNS resolves a name, and how the servers that do it are run.

Where it actually breaks

Lame delegation

A parent zone names a nameserver that refuses to answer for the child; resolvers queue up and get nothing.

A monitoring dashboard on a wall screen with one red panel
Every failure in this section is visible from outside the organisation that caused it, and invisible from inside it.Photograph

How a server becomes lame

Delegation is a promise made in two directions. When a registrant adds nameservers to a zone, the registry writes NS records into the parent zone pointing at those servers, and both sides of that handshake are supposed to agree. A server is lame when the parent points at it but the server itself carries no authoritative data for the zone — it has not been configured with the zone, it was decommissioned without the NS records being cleaned up, or the delegation was entered incorrectly in the first place.

The resolver has no way to know any of this before it sends a query. It follows the referral, contacts the server named in the NS record, and receives either SERVFAIL, REFUSED, or — more insidiously — a referral back to the root, because the server is running as a resolver and treats the query as recursive rather than authoritative. RFC 1034 ↗ established the model that separates these roles; a server that blurs them poisons the delegation chain silently.

A hand-annotated zone file printout on a desk beside a keyboard
The serial at the top of the file is the one line a secondary reads before deciding whether to pull anything at all.Photograph

What goes wrong in practice

Lameness comes in several forms, and they fail differently.

The cleanest failure is outright REFUSED. The server exists, it is reachable, and it tells the resolver immediately that it has no authority over the zone. The resolver marks that nameserver as lame, moves to the next one in the NS set, and if all of them are lame, the query dies with SERVFAIL. From the operator's side this is the easiest form to diagnose: dig +norecurse example.com SOA @ns1.example.com returns REFUSED inside milliseconds.

The subtler failure is a server that answers NOERROR with an empty answer section — no SOA, no records — because it is running a minimal recursive configuration and technically responds to any query without ever actually holding zone data. This passes a naive health check (the server responded!) while breaking every real query.

The worst case is a lame server that is also slow or intermittently unreachable. Resolvers will time out waiting for it, retry, eventually give up, and the added latency propagates to end users as inexplicable slowness rather than a clean error. When there are four nameservers in the NS set and two of them are lame, queries resolve — but inconsistently and slowly — and the problem can persist unnoticed for weeks.

Where lameness originates

The commonest source is lifecycle mismatch. A zone migrates from one DNS operator to another: the new authoritative servers are configured and the NS records at the registrar are updated, but nobody removes the old servers from the NS set. The former operator's servers continue to receive referrals for a zone they no longer carry.

A close second is the provisioning race. NS records propagate from parent to resolver caches with whatever TTL the parent zone carries — often a day or more. If an operator adds NS records to the parent before loading the zone onto the named servers, every resolver that picks up the referral during that window hits a lame server. The correct order is: configure the zone on the servers first, verify they answer authoritatively, then add or update the NS delegation.

Third is simple decommissioning failure: a server is retired, its IP is released or reassigned, but the NS record outlives it. An ICANN SSAC advisory ↗ identified lame delegation as a significant source of unnecessary query load on the root and TLD servers, because resolvers following dead referrals eventually re-query upward.

When a registrant adds nameservers to a zone, the registry writes NS records into the parent zone pointing at those servers, and both sides of that handshake are supposed to agree.

Finding and fixing it

The diagnostic is simple: query each nameserver in the NS set directly with recursion disabled and expect an authoritative answer for the zone's SOA. Any server that does not return AA=1 with a SOA is lame. Tools like dig, drill (from NLnet Labs), and check-soa make this a one-liner per nameserver.

The fix is always the same: either load the zone onto the server so it can answer authoritatively, or remove it from the NS set. Neither step alone is sufficient. Removing an NS record fixes the parent's promise; it does not help resolvers whose caches still carry the old referral. Correcting the zone data on the server fixes the server; it does not fix an NS set that still lists five servers when only three are current. Both sides of the original handshake have to be brought back into agreement.

A patch panel with numbered ports and neatly combed fibre looms
Numbered ports are a name’s last hop. Everything above them is a chain of referrals that exists only as records in other people’s zones.Photograph

The diagnostic is simple: query each nameserver in the NS set directly with recursion disabled and expect an authoritative answer for the zone's SOA.

Read next, in this section