Architecture overview¶
Incident Ops is the system an operations team uses to open, acknowledge, mitigate and resolve incidents against a severity-based SLA, to turn alerts into incidents, to get paged when nobody acknowledges, and to measure KTLO load from the history. Every interface is defined in the contract; these pages explain how the pieces fit and why.
| Page | Covers |
|---|---|
| This page | goals, system context, containers, failure modes, capacity |
| Context map | bounded contexts and how they relate |
| Incident lifecycle | status state machine, write/read paths, event flow per transition |
| SLA timers and escalation | scheduled-message timer loop, idempotency guard, on-call rotation |
| Real-time updates | SignalR hub, negotiation, reconnect |
| Deployment | Azure topology, DNS, security, delivery pipeline |
| Observability | Application Insights vs Prometheus, correlation, KQL and PromQL |
| Code structure | layers, aggregates and boundaries per module, and how they are enforced |
| Alerting | alert → incident → SLA → escalation → notification |
Design goals¶
In priority order:
- Escalation must fire even when the API is idle and must never double-escalate.
- An alert becomes exactly one incident, however many times it repeats.
- The console reflects changes from other operators within 2 seconds.
- Each module is deployable on its own, against a published contract.
- Runs on consumption SKUs: near-zero cost when idle.
System context¶
C4 level 1: who uses the system and what it talks to.
flowchart TB
classDef person fill:#08427b,stroke:#052e56,color:#fff
classDef system fill:#1168bd,stroke:#0b4884,color:#fff
classDef external fill:#999,stroke:#6b6b6b,color:#fff
operator(["Operator / on-call engineer<br/>[Person]<br/>Triages and runs incidents"]):::person
lead(["Engineering lead / Product Owner<br/>[Person]<br/>Reads KTLO metrics and RCAs"]):::person
system["Incident Ops<br/>[Software system]<br/>Incidents, SLA clock, escalation,<br/>alert ingestion, KTLO analytics"]:::system
prom["Prometheus + Alertmanager<br/>[External system]<br/>Metric alerts"]:::external
azmon["Azure Monitor<br/>[External system]<br/>Alert rules, action group"]:::external
chat["Teams / Slack / email<br/>[External system]<br/>Notifications"]:::external
aoai["Azure OpenAI<br/>[External system]<br/>RCA drafts"]:::external
operator -- "opens, acknowledges, resolves<br/>HTTPS" --> system
lead -- "views KPIs, recurring issues<br/>HTTPS" --> system
prom -- "webhook v4" --> system
azmon -- "common alert schema webhook" --> system
system -- "pages on-call" --> chat
chat -- "page" --> operator
system -- "chat completions" --> aoai
Containers¶
C4 level 2: deployable units and how they communicate.
flowchart LR
classDef container fill:#438dd5,stroke:#2e6295,color:#fff
classDef store fill:#438dd5,stroke:#2e6295,color:#fff
classDef external fill:#999,stroke:#6b6b6b,color:#fff
web["Console<br/>[Container: React, TypeScript]<br/>Static Web App"]:::container
api["Incidents API<br/>[Container: .NET 8 Minimal APIs]<br/>domain, SLA, alerts, hub"]:::container
sql[("Database<br/>[Azure SQL Basic]<br/>incidents, timeline, rotation")]:::store
signalr["Azure SignalR Service<br/>[Container: managed]"]:::container
bus{{"Service Bus<br/>[topic incident-events,<br/>queue sla-checks]"}}:::store
fn["Functions<br/>[Container: .NET 8 isolated]<br/>ScheduleSlaCheck, CheckAcknowledgementSla,<br/>NotifyOnCall"]:::container
logic["Logic App<br/>[Container: Consumption]<br/>notification fan-out"]:::container
insights["Insights API<br/>[Container: Python FastAPI]<br/>KPIs, recurring, anomalies, RCA"]:::container
aoai["Azure OpenAI"]:::external
alerts["Alertmanager /<br/>Azure Monitor"]:::external
appi["Application Insights"]:::external
web -- "REST" --> api
web -- "REST" --> insights
web <-- "WebSocket" --> signalr
api -- "hub messages" --> signalr
api -- "EF Core" --> sql
api -- "CloudEvents" --> bus
bus --> fn
fn -- "scheduled messages" --> bus
fn -- "POST /escalate" --> api
fn -- "HTTP" --> logic
insights -- "GET /export" --> api
insights --> aoai
alerts -- "webhooks" --> api
api -.-> appi
fn -.-> appi
insights -.-> appi
| Container | Responsibility | Runtime | State |
|---|---|---|---|
Console (apps/web) |
incident list, detail, timeline, on-call, KPIs | Static Web App | none |
Incidents API (services/api) |
system of record; transitions; SLA deadlines and slaState; alert ingestion and dedupe; events; SignalR hub; /metrics |
Container App | Azure SQL |
Functions (services/functions) |
SLA timers, escalation on missed acknowledgement, paging | Function App, Flex Consumption | Service Bus |
| Logic App | fan-out to Teams/Slack webhook and email | Logic App, Consumption | none |
Insights (services/insights) |
MTTA/MTTR/SLA KPIs, recurring clusters, volume anomalies, RCA drafts | Container App | none, reads the export endpoint |
Infrastructure (infra) |
Bicep, Azure DevOps samples; deployed by .github/workflows/infra.yml |
GitHub Actions | Azure is the state |
Boundaries follow ownership of data: only the API writes incidents. Each container is one bounded context; the context map shows how they relate. Functions and alert sources change incident state exclusively through API endpoints, so transition rules exist in one place, the domain layer.
Analytics and AI¶
Insights pulls GET /api/incidents/export?from=&to= and computes with Pandas. It holds no database: the dataset is a few thousand rows, and recomputing per request with a short in-memory cache is cheaper than keeping a second store consistent. Recurring issues use TF-IDF over title and description, latent semantic analysis and k-means with the cluster count chosen by silhouette; anomalies flag weeks whose volume per service is at least 3 standard deviations above the previous 8 weeks (rolling z-score). RCA drafts send the incident and its timeline to the Azure OpenAI deployment rca-drafts (GPT-5 family, gpt-5.4-mini by default, model parameterized in Bicep) and return a structured draft for a human to edit. Rationale: ADR 0007. Metric definitions: KTLO metrics.
Failure modes¶
| Failure | Effect | Mitigation |
|---|---|---|
| Service Bus unavailable | Events wait in the transactional outbox; timers and pages start late | API keeps serving; the outbox dispatcher retries with capped backoff and delivers when Service Bus recovers. Runbook: escalation not firing |
| Function throws on check | Message retried, then dead-lettered after 10 deliveries | Dead-letter alert. Runbook: dead letters |
| API unavailable when check fires | Escalation call fails, message retried | Retries cover a ~10 minute outage; beyond that dead letters are replayed |
| Duplicate or late check message | None | Level-matching guard makes the check idempotent |
| Alert webhook repeated or retried | None | Fingerprint dedupe appends a timeline entry instead of a new incident |
| API down while alerts fire | Alert sources retry; incidents created late | Alertmanager retries the webhook; ApiDown itself routes to on-call outside the API path (alerting) |
| SignalR unavailable | Console stops live-updating | Client refetches every 30 s |
| Azure OpenAI throttled or down | RCA draft unavailable | Insights returns 502 problem details instead of a mislabelled fallback draft; KPIs unaffected |
| Demo environment powered off | API unreachable; the console shows that the environment is paused | Started on request with the power workflow; Service Bus and the database are re-created and the API re-seeds on first start |
Capacity and cost¶
Sized for a team of 6 to 40 engineers and up to a few hundred incidents a month. All compute is consumption-billed; the API keeps one warm replica while the environment is on, and Service Bus Standard and the Basic database bill a fixed hourly charge while they exist, which is only while the environment is on. Log Analytics ingestion is billed in both states. Ingestion is the cost to watch: sampling stays off for exceptions and Service Bus dependencies, and is set to 25% for successful GET requests.