DevOps Foundations · Module 2: Networking, DNS, HTTP and TLS
How DNS Resolution Really Works
Half of all cannot connect incidents are the name, not the network. This lesson makes resolution observable so DNS stops being the thing you blame and starts being the thing you check.
10 min reading
Objectives
- Trace a hostname from application to resolver to authoritative answer
- Separate a DNS failure from an application-port failure with two commands
- Explain what TTL controls and what it does not
- Predict the effect of /etc/hosts and nsswitch order before testing
Why this matters
The deploy went out, health checks fail, and three engineers argue about whether the database moved. Dig once and the argument ends: the name resolves to an address nobody owns anymore. DNS incidents feel like network incidents because every tool downstream of a bad answer fails identically. The fix is always upstream of the arguing: resolve the name by hand, compare with what the application uses, and the incident shrinks to one record.
Concepts
Resolution is a chain, and each link can lie independently. The application asks the system resolver, which consults /etc/nsswitch.conf order: files (/etc/hosts) first by default on most Linux systems, then DNS. That means a stale hosts entry beats the entire global DNS, which is why a forgotten line from a debugging session six months ago can blackhole one hostname on one machine while every other machine works. Check hosts before anything clever.
The DNS leg itself walks down: root to TLD to authoritative server, with your configured resolver (/etc/resolv.conf) doing the walking and caching the answer for its TTL. dig +trace shows the walk; a plain dig shows the cached answer plus which server gave it. Two records matter most in incidents: A/AAAA for the address, and CNAME for the alias that adds one more lookup and one more TTL. A CNAME to a deleted target resolves to nothing with an error that reads exactly like a network failure.
TTL controls caching, not propagation speed in the way teams imagine. Lowering a TTL to 60 seconds tells resolvers to re-ask every minute, but resolvers that already hold the old answer keep it until it expires, and some middleboxes ignore low TTLs entirely. Plan record changes around the old TTL, not the new one: the cutover completes old-TTL seconds after you think it does, and during that window two answers coexist legitimately.
Worked example
Service fails to reach db.internal on port 5432. Name or port?
$ getent hosts db.internal 10.20.6.10 $ dig +short db.internal 10.20.9.99 $ ss -tn dst 10.20.6.10 2>/dev/null | head -3 (no connection attempts visible)
Expected reading: the system resolver (files first) says .6.10 while direct DNS says .9.99. Someone changed the record and a hosts file pins the old address, or the reverse. Either way this is a name problem, not a port problem: no packet should be aimed at .6.10 at all, so testing the port there is testing a ghost. Fix the name layer (update or remove the pin, wait out the TTL), re-resolve with both tools until they agree, then test the port exactly once. Lab L04 replays this split until it becomes reflex.
The common wrong move
Changing the application to use the IP address directly. It fixes today and breaks every future move: failovers, migrations and scaling all assume the name is the stable handle and the address is disposable. Hardcoded IPs also defeat TLS hostname verification downstream, manufacturing the next incident in lesson 4. Repair names at the name layer; the application keeps the name.
Lab and next step
Lab L04 gives you a wrong record and a wrong port at once and demands the two-command split before any fix. Next, lesson 3 follows the packet past the address: HTTP semantics and the proxies that stand in between.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Add a fake name to your /etc/hosts pointing at 127.0.0.1, show that getent and dig disagree about it, then remove the line and show agreement. Write the three outputs and one sentence on which source each tool consulted.
Pass criteria
Hosts entry shown, disagreement demonstrated with both tools, agreement restored after removal; the one sentence correctly attributes files vs DNS.