FDE Foundations · Module 4: Software Craft
Errors, Idempotency, and Retries
Distributed systems fail constantly and that is fine, if every failure lands in one of two buckets: retry later, or park and inspect. Idempotency keys are what make retrying safe.
12 min reading
Objectives
- Classify errors as retryable or permanent
- Use idempotency keys for every write
- Design retry with backoff and a dead-letter path
Two buckets
Classify each failure: transient (network blip, 503, timeout) means retry; permanent (validation failure, 422, missing record) means park with the payload and context. Retrying permanent failures burns quota and hides real problems; parking transient ones loses data. Misclassifying one as the other is the root of most pipeline bugs.
Idempotency keys
Every write operation gets a client-generated idempotency key: order ID plus step, event ID, or a content hash. If the operation re-runs after a timeout, the key makes the second run a no-op instead of a duplicate. This one habit removes an entire class of 3 a.m. incidents: double-processed invoices.
Backoff and limits
Retry with exponential backoff and jitter, capped attempts, and a circuit breaker on the destination. Retrying a struggling service in a tight loop turns its bad day into an outage.
Dead letters
Permanent failures go to a quarantine store with the original payload, the error, and the attempt count. A dead-letter queue nobody inspects is a data-loss queue with better branding; put review of it in the runbook from day one.
Quick check
An optional 2-3 question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Exercise
Add a retry design to a fictional sync job: classify five example failures as retryable or permanent, define the idempotency key, and specify backoff, attempt cap, and the dead-letter record fields.
Pass criteria
Five failures correctly classified with reasons, a concrete key formula, capped backoff with jitter, and dead-letter fields that allow later reprocessing.