DNS debugging: a field guide for when nothing resolves
Updated

DNS failures in a fixed order: local resolver, authoritative answer, then the path between them. With a fictional outage example.
Roughly half of all it works on my machine mysteries are DNS. The name resolves on one laptop and not another, the new record works for a colleague and not for you, the deploy is green and the page is dead. DNS debugging is a fixed order, and skipping steps is what makes it feel like magic.
The order
Ask what the resolver says. Query the configured resolver for the exact name and record type. Note which server answered, the answer section and the TTL. No answer section with a referral means keep walking up, not trying random fixes.
Ask the authoritative server directly. Query the domain's own nameservers, bypassing every cache. If they answer correctly, the record is fine and the problem is caching or the resolver path. If they do not, stop debugging clients and fix the record.
Check the path and the caches. Browser cache, OS cache, resolver cache, then TTL expiry. Flush one layer at a time and re-query; the layer where the answer changes is the layer that lied to you.
Worked example: a fictional launch-day outage
The context below is fictional. Fictional shop BrightCart (fictional) points its checkout at a new payment host on launch morning. Half the team sees the new page, half sees the old error. The deploy log is clean.
Step one: the failing laptops get no answer section from the office resolver, while working ones show the new record with a long TTL. Step two: the authoritative servers answer the new record correctly everywhere. Conclusion: record fine, office resolver cache stale. The office resolver had cached the old negative answer; after the negative TTL expired, every laptop converged without any further change. Total fix: one diagnosis, zero config edits.
Decision table: where the fault lives
| Resolver says | Authoritative says | Fault |
|---|---|---|
| Correct | Correct | Your local cache or app |
| Stale or empty | Correct | Resolver cache or resolver path |
| Anything | Wrong or missing | The record itself; fix it there |
Checklist: a DNS report worth reading
- Exact queried name and record type.
- Which server answered, and the full answer section.
- Authoritative answer for comparison.
- TTL values at each layer, and which cache was stale.
Related reading
- DevOps Foundations covers networking, DNS, HTTP and TLS in Module 2.
- The TLS companion lab: Inspect TLS Failures with a Test CA.
Straight answers
Frequently asked questions
Is it DNS? How do I know fast?
If the IP works but the name fails, it is DNS or TLS. Test the IP directly first; that single check splits the problem in half.
Which tool first: ping, dig or nslookup?
A resolver query tool first. You want to see which server answered and what it said, not whether packets flow.
What about caching?
Assume every layer caches: browser, OS, resolver, authoritative TTL. Check the TTL on the record before assuming your fix did nothing.