Skip to content

Service level objectives

SLIs and SLOs for the Incidents API (services/api, role incident-ops-api), measured from Application Insights, and the policy for spending the error budget.

SLIs and SLOs

Window: rolling 28 days. Excludes /health/* and requests rejected by rate limiting (429).

SLI Definition SLO
Availability requests with resultCode < 500 ÷ all requests 99.5%
Read latency GET /api/* with duration ≤ 300 ms ÷ all GET /api/* 95%
Write latency POST /api/* with duration ≤ 800 ms ÷ all POST /api/* 95%
Real-time freshness IncidentChanged sent within 2 s of the commit 99%
Escalation timeliness escalations performed within 60 s of ackDueAt 99%

Why these numbers: the API runs on consumption infrastructure with one warm replica and a Basic database capped at 5 DTU, and the demo environment is powered on only for evaluations, so the SLOs are measured while it is on. 99.5% leaves room for scale-out cold starts and DTU saturation under bursts while still flagging real regressions. Escalation timeliness is held tighter because it is the safety net for every other service.

Error budget

SLO Budget per 28 days
Availability 99.5% 0.5% of requests; at 50 000 requests/28 d, 250 failed requests (≈ 3 h 22 min of full outage)
Escalation timeliness 99% 1 late escalation in 100

Queries

Availability SLI over the window:

requests
| where timestamp > ago(28d)
| where cloud_RoleName == "incident-ops-api"
| where name !startswith "GET /health" and resultCode != "429"
| summarize total = count(), good = countif(toint(resultCode) < 500)
| extend sli = round(100.0 * good / total, 3), budgetUsedPct = round(100.0 * (total - good) / (total * 0.005), 1)

Latency SLIs:

requests
| where timestamp > ago(28d)
| where cloud_RoleName == "incident-ops-api"
| where name startswith "GET /api" or name startswith "POST /api"
| extend method = tostring(split(name, " ")[0])
| extend threshold = iff(method == "GET", 300.0, 800.0)
| summarize total = count(), good = countif(duration <= threshold) by method
| extend sli = round(100.0 * good / total, 2)

Escalation timeliness (Functions log the delay between ackDueAt and execution as customDimensions.delaySeconds):

traces
| where timestamp > ago(28d)
| where cloud_RoleName == "func-incident-ops"
| where operation_Name == "CheckAcknowledgementSla" and tostring(customDimensions.decision) == "Escalated"
| summarize total = count(), onTime = countif(todouble(customDimensions.delaySeconds) <= 60)
| extend sli = round(100.0 * onTime / total, 2)

Alerting on burn rate

Multi-window burn-rate alerts on availability (burn rate 1 = spending the budget exactly over 28 days):

Alert Long window Short window Burn rate Budget spent at trigger Severity
Fast burn 1 h 5 min 14.4 2% Sev2, page
Slow burn 6 h 30 min 6 5% Sev3, ticket
let slo = 0.995;
let burn = (window: timespan) {
    requests
    | where timestamp > ago(window)
    | where cloud_RoleName == "incident-ops-api"
    | where name !startswith "GET /health" and resultCode != "429"
    | summarize errorRatio = 1.0 * countif(toint(resultCode) >= 500) / count()
    | extend burnRate = errorRatio / (1 - slo)
};
union (burn(1h) | extend window = "1h"), (burn(5m) | extend window = "5m")

Fast burn fires when both windows exceed 14.4. Low traffic can make one failed request look like a high burn; the alert rule requires at least 20 requests in the short window.

Error budget policy

Budget remaining (28 d) Policy
> 50% Normal. Releases follow the release readiness checklist.
25–50% Releases that touch the failing path need a second reviewer and a canary at 10% for 30 minutes.
0–25% Feature releases to the API pause unless the Product Owner and engineering lead accept the risk in writing. Reliability work moves to the top of the sprint.
Exhausted Freeze on feature releases until the budget recovers above 0 on the rolling window. Only reliability fixes and security patches ship. A postmortem is required if a single incident consumed > 20% of the budget.

The policy is agreed with Product in advance so it is a rule applied, not a negotiation during an incident. Disagreements escalate to the head of engineering and the product lead, with the decision recorded in the weekly stakeholder update.

Review

Monthly, in the ops review: SLI per objective, budget consumed and by which incidents, whether objectives are still the right ones. Changing an SLO is a PR to this file approved by the engineering lead and the Product Owner.