SLOs for customer-facing systems: a starter set

Updated

Monitoring dashboards on control room screens

Three SLIs, honest thresholds, and a report someone reads: the smallest SLO set worth keeping.

Most SLO decks die as decoration. Three owned SLOs with a reading habit survive.

The short answer

Three SLIs measured from the user side, thresholds from baseline, weekly report reviewed with the owner, misses reported with fixes.

The starter set

SLIMeasuresExample threshold
AvailabilitySuccessful operations per attempt995 of 1000 lookups succeed monthly
Latencyp95 lookup time on live datap95 under 3 seconds over 7 days
CorrectnessOutcomes without quarantine surprisesFewer than 5 unexplained drops weekly

Worked example: the SLO that caught the drift

A fictional lookup system (fictional) holds availability while p95 latency creeps from 1.8 to 2.9 seconds over three weeks. The weekly report flags the trend before the threshold breaks; index work lands in the next slice. Nobody pages, nobody argues: the line did its job. The constraints post sets the measurement window; the readout post gives the report its audience.

Checklist: SLOs that live

  1. Measured from the user seat, not the server log alone.
  2. Every SLO has an owner, a query and a review date.
  3. Report reviewed with the customer, not just stored.
  4. Misses carry cause, fix and prevention with dates.

Straight answers

Frequently asked questions

How many SLOs should we start with?

Three: availability from the user seat, latency at a stated percentile, and correctness of the outcome that matters. Each with an owner and a review date.

Who sets the thresholds?

Both sides, in the brief, from baseline measurement. Thresholds dictated after shipping get chosen to flatter.

What happens when an SLO misses?

A dated entry: cause, fix, prevention shipped. Misses with plans build trust; hidden misses end contracts.

Bu sayfanın Türkçesi