The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Foundations · Module 1: Linux working model

Logs and Systematic Diagnosis

Logs are the machine's statement of what it did. This lesson builds the fixed order in which to read them, so every incident ends with a cause instead of a coincidence.

10 min reading

Objectives

  • Follow a failure from symptom to cause using logs before changing anything
  • Separate what happened from why using timestamps across two sources
  • Write an incident record another engineer can act on
  • Explain what each of the first four diagnostic commands rules out

Why this matters

Two engineers get the same page. One restarts three services, edits a config, clears a cache, and the system recovers; nobody knows which action mattered, so the incident will return. The other reads four sources in a fixed order, finds the single cause, fixes it once, and writes ten lines that prevent the next page. The difference is not talent; it is sequence. Diagnosis is a procedure, and procedures beat memory under pressure.

Concepts

Start with the service's own voice, newest first: journalctl -u svc-app --since "30 min ago" for systemd units, or the application's log file for everything else. Errors cluster at the end, but causes precede them, so read backward from the first error, not forward from the noise. Correlate with the kernel's voice: dmesg -T shows OOM kills, disk errors and network link events with human timestamps. Then add the two numbers from lesson 3 (df -h, df -i, free -h) and the two facts from lesson 2 (is the process alive, as whom). Four commands rule out four worlds: dead process, permission denial, resource exhaustion, kernel-level fault. Only then form a hypothesis, and test it with the cheapest check available before changing anything.

Timestamps are the joining key. A 502 at 11:42:13 in the proxy log means nothing until the upstream log shows what it received at 11:42:13, and the deploy record shows what changed at 11:41. Clock skew across machines invalidates this join silently; NTP-synced clocks are a prerequisite, not a nicety. When two sources disagree about order, distrust the story and re-check the clocks before re-checking the theory.

Write the record as you go, not after: symptom with time, each check with its result, the hypothesis, the fix, the verification. The template is five lines, and the discipline matters more than the format. A record that says "restarted, works now" is a confession that the cause is still in production.

Worked example

API returns 500 since 11:42. The order:

$ journalctl -u svc-api --since "1 hour ago" --no-pager | grep -iE "error|fail|denied" | head -5
Oct 02 11:42:13 api[4182]: ERROR: open /srv/app/config.yaml: permission denied
$ ls -l /srv/app/config.yaml
-rw------- 1 root root 96 Oct 2 11:41 /srv/app/config.yaml
$ journalctl --since "1 hour ago" --no-pager | grep -i deploy | head -3
Oct 02 11:41:02 deploy[9001]: installed config.yaml (owner root, mode 600)

Expected reading: first error at 11:42:13 (permission denied, lesson 1 pattern), file owned by root with mode 600 from a deploy at 11:41:02, one minute before symptoms. Cause: the deploy wrote root-owned config over the service-readable one. Fix: restore group-readable ownership per lesson 1 and verify with a read-as-service-user plus one successful request, then record all five lines. Nothing was restarted, nothing was deleted, and the deploy job gets the follow-up, not the service.

The common wrong move

Changing things to see what helps: restarting, redeploying, clearing caches until the symptom moves. Every change destroys the crime scene a little more, and if the symptom stops, two changes share the credit and the cause stays unknown. The rule is read-only first: logs, process state, resource numbers and deploy history are all free to read. The first write to the system should be the fix, applied to a named cause, with a named verification after it.

Lab and next step

Module labs L01-L03 each demand this record as their deliverable: symptom, checks, cause, fix, verification, with the command choices explained. From here the path continues to M02, where the same method meets the network: when the cause is not on this machine, the logs you need live somewhere else, and the check order crosses DNS, ports and TLS.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

Take any recent failure you fixed (or invent nothing: use a stopped local service). Write the five-line record: symptom with time, each check with result, hypothesis, fix, verification. Then mark which check you skipped in reality and what it could have ruled out.

Pass criteria

Record has all five lines with real timestamps; at least three checks named with results; the skipped check is identified honestly with what it would have ruled out.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.