Webhook integration debugging: a field guide
Updated

Signatures, retries and silent drops: the ordered checklist for webhook failures at customer sites.
Webhook debugging fails when engineers guess. The field order below replaces guessing with elimination.
The short answer
Verify signature, bound the retries, make redelivery idempotent, log every drop with a reason. Then replay the failure to prove the fix.
The elimination order
| Step | Check | Fix |
|---|---|---|
| 1. Signature | Timestamp skew, secret rotation, body canonicalization | Log the exact mismatch field, never just failed |
| 2. Retries | Attempts per delivery, backoff shape | Cap attempts, add jitter, dead-letter overflow |
| 3. Idempotency | Same delivery twice, edit-after-send | Natural keys plus idempotency keys, replay-safe consumers |
| 4. Silent drops | Rows that vanish without a log line | Every drop writes reason, delivery id and payload hash |
Worked example: the double-charge scare
A fictional store (fictional) reports duplicate charges after a provider outage. Logs show redeliveries with new ids and a consumer without an idempotency key. The fix: key on order id plus version, unique constraint at the store, replay of the outage window showing zero double-applies. The scare becomes the readout entry. The rollback post covers the response shape; the runbook post turns this into a three-symptom entry.
Checklist: webhook health
- Every delivery carries an idempotency key you control.
- Retry budget written down: attempts, backoff, dead-letter rule.
- Drops always log reason plus payload hash.
- Replay path tested quarterly, not invented during incidents.
Related reading
Straight answers
Frequently asked questions
Where do webhook failures hide?
In four places in order: signature mismatch, retry storms, idempotency gaps, and silent drops nobody logged. Check in that order.
Should webhooks retry?
Yes, with a budget: capped attempts per delivery, backoff with jitter, and a dead-letter store after the budget. Unbounded retries turn incidents into outages.
How do you prove a fix?
Replay the exact failed delivery against staging and show accepted plus quarantined counts. Stories do not prove; replays do.