The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Foundations · Module 2: Networking, DNS, HTTP and TLS

How DNS Resolution Really Works

Half of all cannot connect incidents are the name, not the network. This lesson makes resolution observable so DNS stops being the thing you blame and starts being the thing you check.

10 min reading

Objectives

  • Trace a hostname from application to resolver to authoritative answer
  • Separate a DNS failure from an application-port failure with two commands
  • Explain what TTL controls and what it does not
  • Predict the effect of /etc/hosts and nsswitch order before testing

Why this matters

The deploy went out, health checks fail, and three engineers argue about whether the database moved. Dig once and the argument ends: the name resolves to an address nobody owns anymore. DNS incidents feel like network incidents because every tool downstream of a bad answer fails identically. The fix is always upstream of the arguing: resolve the name by hand, compare with what the application uses, and the incident shrinks to one record.

Concepts

Resolution is a chain, and each link can lie independently. The application asks the system resolver, which consults /etc/nsswitch.conf order: files (/etc/hosts) first by default on most Linux systems, then DNS. That means a stale hosts entry beats the entire global DNS, which is why a forgotten line from a debugging session six months ago can blackhole one hostname on one machine while every other machine works. Check hosts before anything clever.

The DNS leg itself walks down: root to TLD to authoritative server, with your configured resolver (/etc/resolv.conf) doing the walking and caching the answer for its TTL. dig +trace shows the walk; a plain dig shows the cached answer plus which server gave it. Two records matter most in incidents: A/AAAA for the address, and CNAME for the alias that adds one more lookup and one more TTL. A CNAME to a deleted target resolves to nothing with an error that reads exactly like a network failure.

TTL controls caching, not propagation speed in the way teams imagine. Lowering a TTL to 60 seconds tells resolvers to re-ask every minute, but resolvers that already hold the old answer keep it until it expires, and some middleboxes ignore low TTLs entirely. Plan record changes around the old TTL, not the new one: the cutover completes old-TTL seconds after you think it does, and during that window two answers coexist legitimately.

Worked example

Service fails to reach db.internal on port 5432. Name or port?

$ getent hosts db.internal 10.20.6.10 $ dig +short db.internal 10.20.9.99 $ ss -tn dst 10.20.6.10 2>/dev/null | head -3 (no connection attempts visible)

Expected reading: the system resolver (files first) says .6.10 while direct DNS says .9.99. Someone changed the record and a hosts file pins the old address, or the reverse. Either way this is a name problem, not a port problem: no packet should be aimed at .6.10 at all, so testing the port there is testing a ghost. Fix the name layer (update or remove the pin, wait out the TTL), re-resolve with both tools until they agree, then test the port exactly once. Lab L04 replays this split until it becomes reflex.

The common wrong move

Changing the application to use the IP address directly. It fixes today and breaks every future move: failovers, migrations and scaling all assume the name is the stable handle and the address is disposable. Hardcoded IPs also defeat TLS hostname verification downstream, manufacturing the next incident in lesson 4. Repair names at the name layer; the application keeps the name.

Lab and next step

Lab L04 gives you a wrong record and a wrong port at once and demands the two-command split before any fix. Next, lesson 3 follows the packet past the address: HTTP semantics and the proxies that stand in between.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

Add a fake name to your /etc/hosts pointing at 127.0.0.1, show that getent and dig disagree about it, then remove the line and show agreement. Write the three outputs and one sentence on which source each tool consulted.

Pass criteria

Hosts entry shown, disagreement demonstrated with both tools, agreement restored after removal; the one sentence correctly attributes files vs DNS.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.