The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Find the error source: logs say what, metrics say where

Updated

An engineer working with dual monitors in an office

Logs narrate the failure, metrics locate it: start from the error rate spike, then read the lines around it.

Beginners open logs and read. Operators open metrics and aim. The error rate graph tells you when and where; the log lines tell you what. Either alone wastes an hour.

Narrow, then read

Metrics narrow the window. Find the spike minute and the failing endpoint, instance or queue. A flat graph with user complaints means the metric is missing, which is itself a finding.

Logs tell the story. Read the first error in the window, not the loudest. Then the lines around it: what changed just before, what retried just after. Correlate with deploys: most spikes start within minutes of a change.

Worked example: a fictional error spike

The context below is fictional. Fictional checkout BrightCart (fictional) shows errors climbing at 12:04 on one instance only. Logs on that instance show a connection timeout to the payment host starting 12:03, right after a config deploy to that instance alone.

Metrics aimed (one instance, one minute), logs named (timeout, config change). Fix: revert the config on that instance, errors flat by 12:11. Action items: config deploys to all instances together, plus a payment-latency metric that pages.

Checklist: a failure located in minutes

  1. Spike minute and failing unit named from metrics.
  2. First error plus surrounding lines quoted.
  3. Last change before the spike identified.
  4. Missing metric filed as an action item if the graph was blind.

Straight answers

Frequently asked questions

Logs or metrics first?

Metrics first to narrow time and place, then logs for the story. Reading full logs without a window is drowning.

What if there are no metrics?

Start with the narrowest log window around the first user report, then add the missing metric as the incident action item.

How much log is enough?

The first error and the lines around it, plus one confirmed-healthy line for contrast. The rest is noise until proven otherwise.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.