Incident postmortem SEV-1
Incident
INC-SYNTHETIC-OPERATIONS-REGION-REVIEW-EXTENDED-IDENTIFIER
Status
Resolved
Authors
D. Kaur, R. Idowu
Reviewed
18 Sep 2026

Synthetic extended title for an unusually complex cross-functional operating review, including multiple regions, jointly accountable teams, a revised delivery window and the full descriptive name without any ellipsis

Impact

Duration
3h 47m

09:14 – 13:01 UTC

Customers
12,400

18% of active accounts

Failed checkouts
41,900

39,600 recovered by retry

Data loss
None

No double charges

A dependency bump raised the payment-provider timeout from two seconds to thirty. When the provider slowed, every stuck request held its connection fifteen times longer and the pool saturated — three lines inside a forty-one file pull request.

Checkout failure rate · 0–60% · 09:00–14:00 UTC
Checkout failure rate rose from under 1% at 09:00 to a peak of 52% at 11:30, then fell below 2% by 13:15. Detection at 09:47 came when the rate was about 6%.

Gridlines 0/30/60%. Dashed: detected 09:47 at 6%, mitigated 13:01. Peak 52% — nine times worse than when the alert fired.

Timeline · 12 September, UTC

  1. 09:02

    Release 2026.9.3 rolls out

    A dependency bump changes the payment-provider request timeout from 2s to 30s. All six regions by 09:11, no alarms.

  2. 09:14

    Impact beginsImpact start

    Provider p99 latency rises from 240ms to 4.8s. Each slow request now holds a connection for up to 30s and the 200-connection pool begins to queue in eu-west.

  3. 09:47

    Detected33 min to detect

    The success-rate alert fires after five minutes below 95%. No latency alert ever fired: the SLO is measured at p50, which never moved.

  4. 10:26

    Scaling up makes it worse

    Per the runbook, on-call scales pods from 12 to 40. Each opens connections against the same provider-side ceiling; the failure rate climbs from 28% to 41%.

  5. 13:01

    Mitigated3h 47m total

    A diff against 2026.9.2 surfaced the timeout constant at 11:38 and the rollback completes in the last region. Success rate returns above 99%; a retry job recovers 39,600 of 41,900 checkouts by 13:41.

Root cause

A timeout raised from 2s to 30s turned a slow dependency into an outage.

A pool's throughput is capacity divided by hold time. At two seconds a stuck request freed its connection fast enough that a slow provider cost latency and nothing more; at thirty it holds fifteen times longer, so the pool saturates at a fifteenth of the load it was sized for. The slowdown itself was ordinary — twice last quarter, neither an incident. What changed was our tolerance.

Contributing factors

  1. 01

    The latency SLO is measured at p50

    A p99 blowout raises no alert. Thirty-three minutes passed before a success-rate alert caught it indirectly.

  2. 02

    A behaviour change hid inside a dependency bump

    Three lines in a forty-one file pull request, no note in the description, approved in under four minutes.

  3. 03

    The runbook's first move is the wrong move here

    "Scale up" is right for CPU saturation and harmful for pool exhaustion. It cost 72 minutes and raised the failure rate.

What went well

  • Idempotency keys held. No double charges across 41,900 failures; the retry job recovered 94% within 40 minutes.
  • The rollback was clean. Six regions, 83 minutes, no manual steps.
Action items · 3 open, 1 complete
IDActionOwner DueStatus
AI-1 Alert on p99 checkout latency and on success rate, not p50 alone. M. Halvorsen19 Sep Done
AI-2 Add a circuit breaker to the provider client, opening at 25% errors over 30 seconds. R. Idowu26 Sep In progress
AI-3 Treat timeout, retry and pool constants as behaviour changes in review — CODEOWNERS plus a required note. Platform3 Oct Not started
AI-4 Rewrite the checkout runbook to diagnose saturation before scaling, pool metrics first. D. Kaur10 Oct Not started