# Runbook: orders-api - elevated checkout errors (TRAINING EXAMPLE)

> Training example. Fictional service (`orders-api`), fictional environment
> (`training`). Numbers below are a teaching scenario, not a real customer
> incident. This document connects to no real service.

## 1. Purpose, scope and business impact

- Purpose: return checkout error rate to its reference level after a release.
- In scope: `orders-api` checkout path in the `training` environment.
- Out of scope: payment settlement, other services, customer data repair.
- Business impact if this fails: training shoppers cannot complete checkout.

## 2. Symptoms and example signals

- Symptom: error rate over the last 5 minutes is 8 percent; reference is 0.5 percent. Rise started after the latest release.
- Example dashboard panel: `training / orders-api / checkout error rate (5m)` showing 8 percent vs 0.5 percent baseline.
- Example log line (synthetic): `ERROR checkout failed order=demo-order-1042 reason=upstream_timeout release=v2.4.1`.

## 3. Preconditions and authority

- Who may intervene (role): on-call engineer for orders-api.
- Required access: read dashboards and logs; permission to halt the rollout.
- Preconditions: confirm you are looking at the `training` environment, not production.

## 4. Diagnosis (expected vs observed)

| Step | Check | Expected result | Observed result |
| --- | --- | --- | --- |
| 1 | Latest change to orders-api | v2.4.0 stable | v2.4.1 deployed 20 min before the rise |
| 2 | Dependency status (inventory service) | healthy | healthy |
| 3 | Error samples | mixed causes | all `upstream_timeout` on checkout |

## 5. Decision

- Related to the latest change? Yes: timing plus uniform error signature.
- Data compatibility confirmed? Yes: v2.4.1 changed no stored schema.
- Approval required? Yes: release captain confirms before switching traffic.

## 6. Intervention

| Step | Action | Risk | Stop condition |
| --- | --- | --- | --- |
| 1 | Halt the v2.4.1 rollout so no new instances receive it | low | halt command fails twice |
| 2 | Switch traffic to verified previous version v2.4.0 | medium: brief latency bump | error rate does not fall within 10 min |

## 7. Rollback (when it does NOT apply, say so)

- Rollback precondition: stored data schema unchanged between v2.4.0 and v2.4.1 (true here).
- Rollback steps: halt rollout, point traffic at v2.4.0, verify.
- Cases where rollback is wrong: if the release had migrated stored data irreversibly, switching
  versions would not repair data and a rollforward fix would be needed instead.

## 8. Verification

- Error-rate check: 5-minute checkout error rate back near 0.5 percent for 15 minutes.
- Functional check: synthetic checkout probe succeeds 10 times in a row.
- Example record check: order `demo-order-1042` is in its expected state, processed exactly once.

## 9. Stop and escalate

- Stop conditions: error rate keeps rising after the switch; any sign of data corruption.
- Escalate to (role): incident commander plus database owner.
- Communication note: post status in the incident channel every 15 minutes until resolved.

## 10. Incident record and follow-up

- Incident log link or ID: training-incident-042.
- What to update after the event: add the `upstream_timeout` signature to this runbook; add a
  pre-release check for the changed timeout value.
- Last tested date: 2026-10-05 (tabletop walkthrough).
- Version: 1.0.
- Review note: re-test after any checkout dependency change.
