Skip to content

Operations

How production is supported: who responds, how fast, how incidents are run, how alerts become incidents, and what is measured.

Page Covers
Support model L1/L2/L3, hours 10:30–18:30 ET, on-call 1 week in 6, handoffs, on-call health
Severity and SLA severity definitions with examples, ack/resolve targets, slaState
Alerting Prometheus rules, Alertmanager routing, Azure Monitor, dedupe, severity mapping, alert hygiene
Incident response roles, triage flow, escalation paths, stakeholder comms cadence and templates
Postmortem template blameless format; example Sev1
Runbooks API error rate, dead letters, escalation not firing, with KQL
KTLO metrics MTTA, MTTR, SLA compliance, recurring rate, toil, KTLO split, and how Insights computes them
SLOs SLIs, objectives, burn-rate alerts, error budget policy
flowchart LR
    A["Alert or report"] --> I["Incident<br/>Triggered"] --> K["Acknowledged<br/>(SLA ack)"] --> M["Mitigated<br/>(service restored)"] --> R["Resolved<br/>(root cause)"] --> P["Postmortem<br/>Sev1/Sev2"] --> F["Follow-up actions<br/>tracked to closure"]
    I -. "not acknowledged" .-> E["Escalation<br/>primary → secondary → lead"] -.-> K