INC-1187: payments-gateway authorization timeouts after connection pool change¶
Fictional example written against the postmortem template. Services and numbers are from the Incident Ops demo domain.
| Severity | Sev1 |
| Service | payments-gateway (Tier1), impact on checkout |
| Date | 2026-09-17 |
| Duration | detection → mitigation 0:45; → resolution 3:33 |
| SLA | ack Met (9 min of 15), resolve Met (3 h 33 min of 4 h) |
| Incident commander | Ana Ribeiro |
| Postmortem owner | Daniel Okafor |
| Status | Published |
Summary¶
A configuration change deployed at 14:05 capped outbound HTTP connections from payments-gateway to the card processor at 10 per replica. Under the midday peak, processor latency rose from a p50 of 300 ms to 1.2 s and requests queued for a connection until they hit the 10 s client timeout. 31% of authorization requests failed for 50 minutes. Rolling back to the previous Container Apps revision restored service at 16:52. The root cause was a pool size chosen without a capacity calculation, in a change classified as low risk.
Impact¶
- Authorization requests failed: 109 400 of 352 800 (31%) between 16:02 and 16:52.
- Orders that failed checkout after client retries: 24 870.
- Customers saw "payment could not be processed"; no duplicate charges (idempotency keys held).
payments-gatewayavailability SLO (99.9% monthly): about 70% of the monthly error budget consumed (budget ≈ 155 000 failed requests at 60 req/s average).
Timeline (UTC)¶
| Time | Event |
|---|---|
| 14:05 | PR "chore: consolidate HTTP client settings" deployed; sets MaxConnectionsPerServer = 10 on the processor client (previously unbounded). Canary at 10% for 15 min passes at ~25 req/s total |
| 14:20 | Revision promoted to 100% |
| 15:40 | Traffic ramps toward peak; p95 authorization latency climbs from 0.9 s to 3.5 s. No alert (threshold 5 s) |
| 16:02 | First timeouts (TaskCanceledException after 10 s) |
| 16:07 | Alert "payments-gateway 5xx > 5% for 5 min" fires; INC-1187 opened automatically as Sev2 |
| 16:16 | Primary on call acknowledges (9 min) |
| 16:20 | Raised to Sev1 (checkout failing for > 25% of requests); IC Ana Ribeiro, comms lead and scribe assigned; bridge opened |
| 16:24 | First stakeholder update sent |
| 16:31 | Dependency telemetry shows processor latency up but processor 5xx flat; requests spend most time before the first byte is sent |
| 16:38 | Scribe links the 14:05 deployment from the production environment history; IC decides to roll back |
| 16:48 | Traffic shifted to previous revision |
| 16:52 | Error rate under 0.5%; incident marked Mitigated |
| 17:22 | 30 min stable; second stakeholder update with mitigation |
| 18:30 | Load test in staging reproduces queueing at 10 connections and 120 req/s with 1.2 s upstream latency |
| 19:40 | Root cause recorded; incident Resolved |
Time to detect: 5 min from first error, 2 h 2 min from the change. Time to acknowledge: 9 min. Time to mitigate: 45 min from alert.
Root cause¶
Concurrent connections needed = arrival rate × latency (Little's law). At peak, 120 req/s × 1.2 s = 144 concurrent requests across 6 replicas, or 24 per replica. The cap of 10 per replica allowed 60 in total. The remaining requests waited in the handler's connection queue; queue wait counted against the 10 s timeout, so under load most queued requests timed out without ever reaching the processor.
The cap was introduced to protect the processor's documented limit of 200 concurrent connections per merchant. The value 10 was copied from another client's settings and not derived from traffic or latency.
Contributing factors¶
- Risk classification. The change was labeled
choreand reviewed as a refactor. A connection limit on a Tier1 dependency is a capacity change. - Canary at low traffic. The 15-minute canary ran at ~25 req/s, where 60 connections were sufficient. Canary analysis did not look at the latency trend at peak.
- No queueing telemetry. There was no metric for time waiting for a connection; dependency duration hid it inside the total call time.
- Timeout budget inverted.
payments-gatewaywaits 10 s on the processor whilecheckoutgives up onpayments-gatewayafter 8 s and retries once. Retries added load to the queue during the incident. - Latency alert too loose. p95 at 3.5 s for 20 minutes did not alert (threshold 5 s); the first signal was errors.
Detection¶
The 5xx alert worked as designed and fired 5 minutes after the first error. A p95 latency alert at 2 s would have fired around 15:50, about 12 minutes before customer-facing errors. The incident was opened as Sev2 by the alert rule and raised manually; the rule's severity mapping did not account for checkout impact.
Response¶
What went well:
- Acknowledged in 9 minutes; roles assigned within 4 minutes of Sev1 declaration.
- The deployment history on the
productionenvironment made the suspect change visible in one click. - Rollback to the previous revision took 10 minutes end to end and needed no schema change.
- Stakeholder updates went out on the 30-minute cadence without being asked.
What was hard:
- Two people investigated the processor's status page for 10 minutes in parallel, unaware of each other; the scribe's timeline was ahead of the bridge.
- The runbook for high error rate did not mention recent configuration changes, only code deploys.
Where we got lucky:
- Idempotency keys prevented duplicate charges when
checkoutretried. - The peak was the regular midday peak, not a promotion day at 3× traffic.
Action items¶
| # | Action | Type | Owner | Due | Work item |
|---|---|---|---|---|---|
| 1 | Size the processor connection limit from Little's law at 2× peak (≥ 50 per replica, ≤ 200 total) and make it configuration per environment | prevent | Daniel Okafor | 2026-09-24 | AB#2311 |
| 2 | Invert timeout budget: processor client 5 s, checkout → payments-gateway 8 s, retries only on idempotent calls with jitter |
prevent | Priya Nair | 2026-10-01 | AB#2312 |
| 3 | Add connection-queue wait metric and a dashboard tile for the processor client | detect | Daniel Okafor | 2026-10-01 | AB#2313 |
| 4 | Alert on payments-gateway p95 > 2 s for 5 min (Sev2) |
detect | Ana Ribeiro | 2026-09-24 | AB#2314 |
| 5 | Map "checkout success rate < 90%" alert directly to Sev1 | detect | Ana Ribeiro | 2026-09-24 | AB#2315 |
| 6 | Release checklist: changes to timeouts, pool sizes, retry or rate limits are capacity changes and need a load test result in the PR | process | Marcelo Roman | 2026-09-24 | AB#2316 |
| 7 | Add "recent configuration changes" step to the high error rate runbook | process | Ana Ribeiro | 2026-09-22 | AB#2317 |
| 8 | Canary analysis compares p95 latency at matching traffic level, not only error rate | mitigate | Priya Nair | 2026-10-15 | AB#2318 |
Lessons¶
Any setting that limits concurrency (pool sizes, semaphores, rate limits, timeouts) is a capacity decision and must be derived from arrival rate and latency at peak, not copied. Canary analysis at low traffic does not validate capacity changes.