Find the error source: logs say what, metrics say where
Updated

Logs narrate the failure, metrics locate it: start from the error rate spike, then read the lines around it.
Beginners open logs and read. Operators open metrics and aim. The error rate graph tells you when and where; the log lines tell you what. Either alone wastes an hour.
Narrow, then read
Metrics narrow the window. Find the spike minute and the failing endpoint, instance or queue. A flat graph with user complaints means the metric is missing, which is itself a finding.
Logs tell the story. Read the first error in the window, not the loudest. Then the lines around it: what changed just before, what retried just after. Correlate with deploys: most spikes start within minutes of a change.
Worked example: a fictional error spike
The context below is fictional. Fictional checkout BrightCart (fictional) shows errors climbing at 12:04 on one instance only. Logs on that instance show a connection timeout to the payment host starting 12:03, right after a config deploy to that instance alone.
Metrics aimed (one instance, one minute), logs named (timeout, config change). Fix: revert the config on that instance, errors flat by 12:11. Action items: config deploys to all instances together, plus a payment-latency metric that pages.
Checklist: a failure located in minutes
- Spike minute and failing unit named from metrics.
- First error plus surrounding lines quoted.
- Last change before the spike identified.
- Missing metric filed as an action item if the graph was blind.
Related reading
- Hands on: Find Error Source in Logs and Metrics.
- The first-hour playbook: Incident response.
Straight answers
Frequently asked questions
Logs or metrics first?
Metrics first to narrow time and place, then logs for the story. Reading full logs without a window is drowning.
What if there are no metrics?
Start with the narrowest log window around the first user report, then add the missing metric as the incident action item.
How much log is enough?
The first error and the lines around it, plus one confirmed-healthy line for contrast. The rest is noise until proven otherwise.