Skip to content

Runbooks

One runbook per alert or recurring failure. Each has: symptoms, impact, triage steps with KQL, mitigation, verification and escalation. KQL runs in the Application Insights resource (Logs blade) shared by every module; cloud_RoleName identifies the module.

Runbook Alerts that link to it Default severity Tier
api-high-error-rate Prometheus ApiHighErrorRate, ApiHighLatencyP95, ApiDown; Azure Monitor API failed requests, availability, response time Sev2 (Sev1 for ApiDown) L2
service-bus-dead-letters Azure Monitor dead-lettered messages, Function failures Sev2 L2
sla-breach-escalation-not-firing Prometheus SlaBreachesOpen; incident Triggered past ackDueAt + 5 min without Escalated entry Sev2 L2

How alerts reach these runbooks: alerting.

Conventions used in the queries

Item Value
Role names incident-ops-api (API), func-incident-ops (Functions), incident-ops-insights (Insights)
Incident id in logs customDimensions.incidentId
Event type in logs customDimensions.eventType
Function invocations requests with name = function name
Correlation operation_Id (W3C trace id), propagated over HTTP and Service Bus

Writing a new runbook

Copy the structure of an existing one. Every query must have been run against real data at least once before the runbook merges. Mark each step L1 or L2 so the support model tiers know where they stop.