Service Bus scheduled messages for SLA timers instead of polling¶
Context and problem statement¶
An incident not acknowledged by ackDueAt must escalate one level, up to level 3. A Sev1 has a 15-minute acknowledgement window, so the timer must be accurate to well under a minute, and it must fire even when the API receives no traffic. How do we implement per-incident timers?
Decision drivers¶
- Accuracy: escalation within 60 s of
ackDueAt. - No double escalation under retries, duplicates or concurrent acknowledgement.
- Zero cost and zero work when no incident is open.
- Survives restarts and scale-to-zero of every compute component.
Considered options¶
- Service Bus scheduled messages on a
sla-checksqueue, one per incident per level - Timer-triggered function polling
GET /api/incidents?open=trueevery minute - Durable Functions durable timers, one orchestration per incident
- Background
IHostedServicein the API
Decision outcome¶
Chosen option: scheduled messages. ScheduleSlaCheck schedules { incidentId, escalationLevel } at ackDueAt on every incident.triggered and incident.escalated. CheckAcknowledgementSla escalates only if the incident is still Triggered at the same level.
Consequences¶
- Good: the timer is durable in the broker; no compute runs while waiting.
- Good: the level guard makes processing idempotent; acknowledgement simply makes the pending message a no-op, so nothing has to be cancelled.
- Good: dead-lettering gives a visible, replayable failure path.
- Bad: timer state is not queryable as a list ("what escalates next?"). Mitigated by
ackDueAton the incident, which the console sorts by. - Bad: depends on Service Bus being available at trigger time; covered by the escalation runbook and an alert on the absence of
Escalatedtimeline entries for breached incidents.
Pros and cons of the options¶
Polling timer function¶
- Good: simplest code; easy to reason about.
- Bad: up to one interval of extra latency; 1 440 executions a day when nothing is open; every poll loads the API and database (and wakes serverless SQL from pause).
- Bad: concurrency between overlapping polls needs a lock.
Durable Functions timers¶
- Good: timers and cancellation are first-class.
- Bad: adds a storage-backed orchestration state store and replay semantics for a single-step workflow; harder to operate and to explain in a runbook.
Hosted service in the API¶
- Good: no extra infrastructure.
- Bad: Container Apps scales the API to zero; timers die with the replica. Multiple replicas need leader election.