Observability¶
Two telemetry paths with different jobs. Application Insights is the system of investigation: traces, logs, exceptions, dependencies, correlated across modules. Prometheus is the system of alerting on the API's metrics in the local stack and in any Kubernetes or self-hosted deployment; in Azure, Azure Monitor alert rules over Application Insights do the same job.
What goes where¶
| Signal | Application Insights | Prometheus |
|---|---|---|
| HTTP requests (rate, errors, latency) | requests table, every module |
http_server_request_duration_seconds histogram, API /metrics |
| Dependencies (SQL, Service Bus, HTTP) | dependencies table |
— |
| Exceptions with stack traces | exceptions table |
— |
| Structured logs | traces table with incidentId, eventType, decision |
— |
| Distributed traces | end-to-end transaction view via operation_Id |
— |
Domain metrics (meter IncidentOps.Api) |
customMetrics |
incidentops_incidents_open{severity}, incidentops_sla_breached_open, incidentops_sla_compliance_30d, incidentops_incident_events_total{kind,severity,source} |
| Availability | standard test on /health/live, 15 min, 1 location, while the environment is on |
up{job="incident-ops-api"} |
| Alerting | Azure Monitor alert rules → Action Group | rules → Alertmanager |
| Retention | 90 days (Log Analytics) | 15 days (local Prometheus default) |
Both are fed by OpenTelemetry in the API: the Azure Monitor exporter and the Prometheus exporter run side by side on the same instruments, so a request counted in one is counted in the other.
flowchart LR
subgraph api ["Incidents API (services/api)"]
otel["OpenTelemetry SDK<br/>traces, metrics, logs"]
end
fn["Functions"] --> appi
ins["Insights"] --> appi
otel -- "Azure Monitor exporter" --> appi["Application Insights<br/>+ Log Analytics"]
otel -- "/metrics (Prometheus exporter)" --> prom["Prometheus"]
appi --> azrules["Azure Monitor alert rules"]
azrules --> ag["Action Group"]
prom --> am["Alertmanager"]
ag -- "common alert schema" --> ingest["POST /api/alerts/azure-monitor"]
am -- "webhook v4" --> ingest2["POST /api/alerts/alertmanager"]
appi --> wb["Workbooks, KQL"]
prom --> pq["PromQL, Grafana"]
Correlation¶
- W3C trace context on HTTP (
traceparent), propagated by the OpenTelemetry instrumentation. - On Service Bus, the publisher sets
Diagnostic-Id/traceparentas an application property; the Functions trigger continues the trace, so oneoperation_Idcovers console click → API → topic → function →POST /escalate→ API. - Every log line inside an incident operation carries
incidentIdandincidentNumberas scope properties (customDimensionsin Application Insights).
KQL examples¶
End-to-end trace for one incident:
let incidentId = "<incident id>";
let ops = union traces, requests
| where timestamp > ago(2d)
| where tostring(customDimensions.incidentId) == incidentId
| distinct operation_Id;
union requests, dependencies, traces, exceptions
| where operation_Id in (ops)
| project timestamp, cloud_RoleName, itemType, name, message, success, duration
| order by timestamp asc
p95 latency per route, last 24 hours:
requests
| where timestamp > ago(24h) and cloud_RoleName == "incident-ops-api"
| summarize p95 = percentile(duration, 95), count() by name
| order by p95 desc
Event publish failures:
dependencies
| where timestamp > ago(24h)
| where cloud_RoleName == "incident-ops-api" and type has "Service Bus" and success == false
| summarize failures = count() by bin(timestamp, 15m), resultCode
More queries in the runbooks and SLOs.
PromQL examples¶
Error ratio over 5 minutes (basis of ApiHighErrorRate):
sum(rate(http_server_request_duration_seconds_count{job="incident-ops-api", http_response_status_code=~"5.."}[5m]))
/
sum(rate(http_server_request_duration_seconds_count{job="incident-ops-api"}[5m]))
p95 latency per route (basis of ApiHighLatencyP95):
histogram_quantile(0.95,
sum by (le, http_route) (rate(http_server_request_duration_seconds_bucket{job="incident-ops-api"}[5m])))
Open incidents by severity and breached SLAs (basis of SlaBreachesOpen):
Request rate by status class:
sum by (http_response_status_code) (rate(http_server_request_duration_seconds_count{job="incident-ops-api"}[1m]))
Rules built on these: alerting.
Sampling and cost¶
- Application Insights: adaptive sampling off for exceptions, dependencies to Service Bus and all non-2xx requests; successful
GETrequests sampled at 25%. KQL counts over sampled data usesum(itemCount)instead ofcount()when precision matters. - Prometheus: scrape interval 15 s; no high-cardinality labels (no incident id, no user) on metrics.