Operations¶
How production is supported: who responds, how fast, how incidents are run, how alerts become incidents, and what is measured.
| Page | Covers |
|---|---|
| Support model | L1/L2/L3, hours 10:30–18:30 ET, on-call 1 week in 6, handoffs, on-call health |
| Severity and SLA | severity definitions with examples, ack/resolve targets, slaState |
| Alerting | Prometheus rules, Alertmanager routing, Azure Monitor, dedupe, severity mapping, alert hygiene |
| Incident response | roles, triage flow, escalation paths, stakeholder comms cadence and templates |
| Postmortem template | blameless format; example Sev1 |
| Runbooks | API error rate, dead letters, escalation not firing, with KQL |
| KTLO metrics | MTTA, MTTR, SLA compliance, recurring rate, toil, KTLO split, and how Insights computes them |
| SLOs | SLIs, objectives, burn-rate alerts, error budget policy |
flowchart LR
A["Alert or report"] --> I["Incident<br/>Triggered"] --> K["Acknowledged<br/>(SLA ack)"] --> M["Mitigated<br/>(service restored)"] --> R["Resolved<br/>(root cause)"] --> P["Postmortem<br/>Sev1/Sev2"] --> F["Follow-up actions<br/>tracked to closure"]
I -. "not acknowledged" .-> E["Escalation<br/>primary → secondary → lead"] -.-> K