- Incident
- INC-2026-041
- Status
- Resolved
- Authors
- D. Kaur, R. Idowu
- Reviewed
- 18 Sep 2026
Checkout was unavailable for 3 hours 47 minutes on 12 September
Impact
- Duration
- 3h 47m
- Customers
- 12,400
- Failed checkouts
- 41,900
- Data loss
- None
09:14 – 13:01 UTC
18% of active accounts
39,600 recovered by retry
No double charges
A dependency bump raised the payment-provider timeout from two seconds to thirty. When the provider slowed, every stuck request held its connection fifteen times longer and the pool saturated — three lines inside a forty-one file pull request.
Gridlines 0/30/60%. Dashed: detected 09:47 at 6%, mitigated 13:01. Peak 52% — nine times worse than when the alert fired.
Timeline · 12 September, UTC
- 09:02
Release 2026.9.3 rolls out
A dependency bump changes the payment-provider request timeout from 2s to 30s. All six regions by 09:11, no alarms.
- 09:14
Impact beginsImpact start
Provider p99 latency rises from 240ms to 4.8s. Each slow request now holds a connection for up to 30s and the 200-connection pool begins to queue in eu-west.
- 09:47
Detected33 min to detect
The success-rate alert fires after five minutes below 95%. No latency alert ever fired: the SLO is measured at p50, which never moved.
- 10:26
Scaling up makes it worse
Per the runbook, on-call scales pods from 12 to 40. Each opens connections against the same provider-side ceiling; the failure rate climbs from 28% to 41%.
- 13:01
Mitigated3h 47m total
A diff against 2026.9.2 surfaced the timeout constant at 11:38 and the rollback completes in the last region. Success rate returns above 99%; a retry job recovers 39,600 of 41,900 checkouts by 13:41.
Root cause
A timeout raised from 2s to 30s turned a slow
dependency into an outage.
A pool's throughput is capacity divided by hold time. At two seconds a stuck request freed its connection fast enough that a slow provider cost latency and nothing more; at thirty it holds fifteen times longer, so the pool saturates at a fifteenth of the load it was sized for. The slowdown itself was ordinary — twice last quarter, neither an incident. What changed was our tolerance.
Contributing factors
- 01
The latency SLO is measured at p50
A p99 blowout raises no alert. Thirty-three minutes passed before a success-rate alert caught it indirectly.
- 02
A behaviour change hid inside a dependency bump
Three lines in a forty-one file pull request, no note in the description, approved in under four minutes.
- 03
The runbook's first move is the wrong move here
"Scale up" is right for CPU saturation and harmful for pool exhaustion. It cost 72 minutes and raised the failure rate.
What went well
- Idempotency keys held. No double charges across 41,900 failures; the retry job recovered 94% within 40 minutes.
- The rollback was clean. Six regions, 83 minutes, no manual steps.
| ID | Action | Owner | Due | Status |
|---|---|---|---|---|
| AI-1 | Alert on p99 checkout latency and on success rate, not p50 alone. | M. Halvorsen | 19 Sep | Done |
| AI-2 | Add a circuit breaker to the provider client, opening at 25% errors over 30 seconds. | R. Idowu | 26 Sep | In progress |
| AI-3 | Treat timeout, retry and pool constants as behaviour changes in review — CODEOWNERS plus a required note. | Platform | 3 Oct | Not started |
| AI-4 | Rewrite the checkout runbook to diagnose saturation before scaling, pool metrics first. | D. Kaur | 10 Oct | Not started |