A dependency bump raised the payment-provider timeout from two seconds to thirty. When the provider slowed, every stuck request held its connection fifteen times longer and the pool saturated — three lines inside a forty-one file pull request.
09:14 – 13:01 UTC
18% of active accounts
39,600 recovered by retry
No double charges
Gridlines 0/30/60%. Dashed: detected 09:47 at 6%, mitigated 13:01. Peak 52% — nine times worse than when the alert fired.
A dependency bump changes the payment-provider request timeout from 2s to 30s. All six regions by 09:11, no alarms.
Provider p99 latency rises from 240ms to 4.8s. Each slow request now holds a connection for up to 30s and the 200-connection pool begins to queue in eu-west.
The success-rate alert fires after five minutes below 95%. No latency alert ever fired: the SLO is measured at p50, which never moved.
Per the runbook, on-call scales pods from 12 to 40. Each opens connections against the same provider-side ceiling; the failure rate climbs from 28% to 41%.
A diff against 2026.9.2 surfaced the timeout constant at 11:38 and the rollback completes in the last region. Success rate returns above 99%; a retry job recovers 39,600 of 41,900 checkouts by 13:41.
A timeout raised from 2s to 30s turned a slow
dependency into an outage.
A pool's throughput is capacity divided by hold time. At two seconds a stuck request freed its connection fast enough that a slow provider cost latency and nothing more; at thirty it holds fifteen times longer, so the pool saturates at a fifteenth of the load it was sized for. The slowdown itself was ordinary — twice last quarter, neither an incident. What changed was our tolerance.
A p99 blowout raises no alert. Thirty-three minutes passed before a success-rate alert caught it indirectly.
Three lines in a forty-one file pull request, no note in the description, approved in under four minutes.
"Scale up" is right for CPU saturation and harmful for pool exhaustion. It cost 72 minutes and raised the failure rate.
| ID | Action | Owner | Due | Status |
|---|---|---|---|---|
| AI-1 | Alert on p99 checkout latency and on success rate, not p50 alone. | M. Halvorsen | 19 Sep | Done |
| AI-2 | Add a circuit breaker to the provider client, opening at 25% errors over 30 seconds. | R. Idowu | 26 Sep | In progress |
| AI-3 | Treat timeout, retry and pool constants as behaviour changes in review — CODEOWNERS plus a required note. | Platform | 3 Oct | Not started |
| AI-4 | Rewrite the checkout runbook to diagnose saturation before scaling, pool metrics first. | D. Kaur | 10 Oct | Not started |