Reliability review · Q4 FY26 · 8 September

Every service met its objective and the error budget is 71% gone

Six weeks into a thirteen-week quarter. Nothing here is breaching an SLO, and 1 hour 55 minutes of unreliability is left to spend before 31 October. An objective is a floor; a budget is a total.

Remaining
1h 55m
Burn rate
7.1 min/day
Runs out
24 Sep

Error budget spent

71%

278 of 393 minutes, against a 99.7% quarterly objective across all services.

Budget remaining · 1 Aug – 31 Oct · 393 min at full
Budget falls from 393 minutes on 1 August to 115 today in three steps: webhook DNS 12 Aug, thumbnail worker 19 Aug, and the storage failover 4 Sep which alone took 134 minutes. At this burn rate it reaches zero on 24 September, five weeks before the quarter closes.

1 Aug31 Oct

Each step down is an incident. The dashed tail is the projection at today's burn rate: budget gone 24 September, then five weeks of quarter with nothing left to spend.

The thing this board is for

Four services, four objectives met, and seventy-one per cent of the quarter spent in six weeks. Both are true and only one of them is on a status page. The publish pipeline cleared its floor by six hundredths of a point and took two-thirds of the money.

Where it went

Sorted by budget consumed, not by date · four events, 4h 38m total
EventWhenDowntime BudgetShare of spend
Storage failover — publish pipeline 4 Sep2h 14m 134 min 48%
Webhook delivery — DNS in eu-west 12 Aug1h 06m 66 min 24%
Thumbnail worker rollout 19 Aug41m 41 min 15%
Sub-minute blips · 17 events ongoing37m 37 min 13%
Total6 weeks4h 38m 278 min71% of budget

What the overspend has already cost

  1. 01

    The publish pipeline has been frozen since 5 September

    Three releases held, including the thumbnail rework — which needs the pipeline it is not allowed to touch. The freeze lifts when the budget clears 40%.

  2. 02

    On-call is doubled until the circuit breaker lands

    The storage failover has no automatic recovery, so a second responder covers publish out of hours until 26 September.

  3. 03

    The latency objective is unfundable this quarter

    Tightening p95 to 300 ms would spend budget we no longer have. It moves to Q1 — recorded, not quietly dropped.

By service · last 90 days, one bar per day

Publish pipeline objective 99.70% · actual 99.76%

90 days · read each row left to right

1 degraded day · 1 partial outage — 63% of the quarter's spend Met, barely

Webhooks objective 99.50% · actual 99.95%

90 days · read each row left to right

1 degraded day · 24% of the spendMet

API objective 99.90% · actual 99.99%

90 days · read each row left to right

1 degraded day · 13% of the spendMet

Web app objective 99.90% · actual 100%

90 days · read each row left to right

No incidents · none of the spendMet

Operational — full bar Degraded — two-thirds bar Partial outage — short bar Height and colour both encode the day; counts are written beside each strip.