Incident response¶
How an incident is run from detection to closure.
Roles¶
Roles are filled for every Sev1 and Sev2. For Sev3/Sev4 the assignee holds all three.
| Role | Default | Responsibilities | Does not |
|---|---|---|---|
| Incident commander (IC) | L2 primary until handed off; engineering lead for Sev1 longer than 1 h | owns severity, decisions, priorities, role assignments, when to escalate, when to declare mitigated | debug hands-on |
| Comms lead | L1 lead or secondary on call | stakeholder updates on cadence, status page, single point of contact for questions | make technical decisions |
| Scribe | any engineer not debugging | timeline in the incident (POST /notes): actions, findings, decisions with timestamps |
filter what goes in |
| Subject matter experts | owning team (L3) | investigate and mitigate as directed by the IC | change severity or communicate externally |
Handing off IC is explicit: "I am handing IC to Ana, Ana confirm" in the bridge and a note on the timeline.
Triage flow¶
flowchart TD
A[Alert, report or page] --> B{Real and current?}
B -- no --> B1[Close as noise<br/>tag alert for tuning]
B -- yes --> C{Existing incident<br/>for this symptom?}
C -- yes --> C1[Add note, link, raise severity if needed]
C -- no --> D[Open incident<br/>pick severity by impact]
D --> E{Sev1 or Sev2?}
E -- yes --> F[Acknowledge, assign IC,<br/>open bridge, comms lead]
E -- no --> G[Acknowledge, assign owner,<br/>work in business hours]
F --> H[Mitigate first:<br/>rollback, failover, scale, feature flag]
H --> I{Impact stopped?}
I -- no --> J[Escalate: L3, lead, vendor]
J --> H
I -- yes --> K[Mark Mitigated,<br/>monitor 30 min]
K --> L[Root cause, fix, Resolve]
G --> L
L --> M[Postmortem for Sev1/Sev2<br/>within 5 business days]
Mitigation before diagnosis: restore service with the cheapest reversible action, then find the cause.
First 15 minutes¶
For a Sev1 or Sev2, the person acknowledging:
- Acknowledge in the console (
POST /acknowledge) — stops the escalation timer. - Confirm severity against severity definitions.
- Open the bridge (Teams meeting linked in the incident note).
- Assign IC (self by default), comms lead and scribe.
- Post the initial update (template).
- Check what changed: deploys in the last 24 h (GitHub deployments on
production), config changes, Azure Service Health, dependency status. - Pick the matching runbook: runbooks.
Escalation paths¶
| Trigger | Escalate to | How |
|---|---|---|
| Not acknowledged within ack window | next on-call level (automatic) | CheckAcknowledgementSla → POST /escalate |
| Sev1 not mitigated in 60 min, Sev2 in 2 h | engineering lead + owning team lead | IC pages directly |
| Needs code change or deep component knowledge | owning team (L3) | page the team's on-call or lead |
| Azure platform issue suspected | Azure support, severity A for Sev1 | support request; record case number in notes |
| Data loss, data exposure or security | security on call + engineering lead, immediately | page; Sev1 regardless of scope |
| Customer contractual impact | account/product owner | comms lead |
| Sev1 lasting > 2 h | head of engineering | IC |
Escalating early is never a mistake in review. Escalating late is the common one.
Stakeholder communication¶
| Severity | Audience | Channel | First update | Cadence | Resolution |
|---|---|---|---|---|---|
Sev1 |
engineering leadership, product owners, support, account owners | incident channel + email list + status page | 15 min after declaration | every 30 min, even if nothing changed | within 1 h of resolution; postmortem link in 5 business days |
Sev2 |
product owners of affected service, support | incident channel + email | 30 min | every 60 min | within 2 h |
Sev3 |
owning team, support | incident channel | 4 h | daily | on resolution |
Sev4 |
owning team | the work item | — | — | on resolution |
Rules: say what is known, what is not, and when the next update is. No speculation about cause before it is confirmed. Times in UTC with EST in parentheses.
Templates¶
Initial update¶
[INC-{number}] {Sev} — {title}
Status: Investigating
Impact: {who is affected and how; % of requests or users if known}
Started: {time UTC} ({time EST})
Services: {serviceId}
IC: {name} | Comms: {name}
What we are doing: {current action}
Next update: {time UTC} or sooner if status changes
Progress update¶
[INC-{number}] {Sev} — {title}
Status: Investigating | Identified | Mitigating | Monitoring
Impact: {current impact; changed since last update?}
Since last update: {findings and actions, 2–4 bullets}
Next steps: {actions and owners}
Next update: {time UTC}
Mitigated¶
[INC-{number}] {Sev} — Mitigated
Impact ended: {time UTC}; duration {h:mm}
Mitigation: {what restored service}
Residual risk: {anything still degraded or temporary}
Next: root cause and permanent fix; postmortem by {date}
Resolved¶
[INC-{number}] {Sev} — Resolved
Root cause (summary): {one or two sentences}
Customer impact: {scope, duration}
Follow-up: postmortem {link} on {date}
Executive summary (Sev1, within 24 h)¶
What happened: {one paragraph, plain language}
Impact: {customers, transactions, revenue if known, SLA breached?}
Duration: detection {t}, mitigation {t}, resolution {t}
Cause: {confirmed / under investigation}
What we are doing to prevent recurrence: {top 3 actions with dates}
Closing an incident¶
Mitigatedwhen user impact has stopped and has stayed stopped for 30 minutes.ResolvedwithrootCausefilled: one or two sentences naming the cause, not the symptom.- Sev1/Sev2: postmortem scheduled before resolving; owner assigned on the incident (template).
- Temporary mitigations still in place become work items tagged
incident-followupbefore resolving.
Anti-patterns¶
- Debugging in a direct message instead of the bridge and the timeline.
- IC debugging hands-on and losing track of the whole picture.
- Holding updates until there is good news.
- Lowering severity to make the SLA look better.
- Resolving without root cause because the symptom went away.