KTLO metrics¶
What we measure about keeping the lights on, how each number is computed, where it comes from, and what it triggers.
Metric catalog¶
| Metric | Target | Source | Endpoint / query |
|---|---|---|---|
| MTTA | Sev1 ≤ 10 min, Sev2 ≤ 20 min | incidents | GET /api/metrics/summary (30 d), Insights GET /api/kpis |
| MTTR | Sev1 ≤ 2 h, Sev2 ≤ 6 h | incidents | same |
| SLA compliance | Sev1/Sev2 ≥ 95%, Sev3/Sev4 ≥ 90% | incidents | same, per service and severity |
| Breached open | 0 | incidents | GET /api/metrics/summary breachedOpen |
| Incident volume per service | trend down quarter over quarter | incidents | Insights GET /api/kpis weekly trend, GET /api/anomalies |
| Recurring-issue rate | ≤ 15% | incidents | Insights GET /api/recurring |
| Toil % | ≤ 20% of engineering time | Azure Boards | tag toil |
| KTLO vs feature split | planned 25% KTLO ±5 pp | Azure Boards | tags ktlo, tech-debt |
| Postmortem actions closed on time | ≥ 85% | Azure Boards | tag incident-followup |
Definitions¶
All durations from incident timestamps in UTC, over incidents created in the period.
| Metric | Formula | Notes |
|---|---|---|
| MTTA | mean(acknowledgedAt − createdAt) over acknowledged incidents |
Median and p90 reported alongside; one outlier moves a mean a lot at low volume |
| MTTR | mean(resolvedAt − createdAt) over resolved incidents |
Time to restore is mitigatedAt − createdAt, reported as MTTM; resolve includes root cause |
| SLA compliance | count(resolvedAt ≤ resolveDueAt) ÷ count(resolved) |
Per severity; open incidents past due count as breached in the period they breach |
| Ack compliance | count(acknowledged before first escalation) ÷ count(acknowledged) | escalationLevel = 1 at acknowledgement |
| Incident volume | count per service per ISO week | Normalized per Tier for comparisons |
| Recurring-issue rate | incidents in a cluster of size ≥ 3 ÷ all incidents, over 180 days | Clusters from GET /api/recurring |
| Anomalous week | weekly count per service with robust z-score > 3.5 vs trailing 12 weeks | From GET /api/anomalies |
| Toil % | hours on items tagged toil ÷ total completed hours in the sprint |
Manual, repetitive, automatable, no lasting value |
| KTLO split | points tagged ktlo ÷ completed points; same for tech-debt; rest is feature |
Planned vs actual per sprint |
How Insights computes them¶
Insights pulls GET /api/incidents/export?from=&to= (flat incidents, no timeline) and works on a Pandas DataFrame.
| Endpoint | Computation |
|---|---|
GET /api/kpis?days=90 |
Parse timestamps as UTC; tta = acknowledgedAt − createdAt, ttr = resolvedAt − createdAt; groupby([serviceId, severity]) → mean/median/p90 in minutes and compliance ratio; weekly trend via resample("W-MON", on="createdAt") |
GET /api/recurring?days=180 |
Text = title + " " + rootCause; scikit-learn TfidfVectorizer (English stop words, 1–2 grams, min_df=2); clustering over cosine distance; clusters of size ≥ 3 returned with count, services and sample titles |
GET /api/anomalies?days=180 |
Weekly counts per service; median and MAD over a trailing 12-week window; z = 0.6745 × (x − median) ÷ MAD; flagged when z > 3.5 |
POST /api/rca/draft |
Incident + timeline to Azure OpenAI; output is a draft input to the postmortem, not a metric |
Choices worth knowing:
- Median/MAD instead of mean/standard deviation for anomalies: incident counts are small and spiky; one bad week should not hide the next.
- TF-IDF instead of embeddings for recurrence: deterministic, explainable (top terms per cluster), no model cost, good enough on short titles.
- The API's
/api/metrics/summarycovers the fixed 30-day headline numbers for the console; Insights covers breakdowns and trends. Both use the same formulas; snapshot tests in Insights pin them (quality gates).
Metrics from Azure Boards¶
Toil, KTLO split and follow-up closure come from work items, not incidents. Queries in ADO hygiene. They depend on tagging discipline, which the weekly hygiene sweep checks.
Review cadence¶
| Forum | Frequency | Metrics | Output |
|---|---|---|---|
| Ops review | weekly, 30 min | open breached, last week's Sev1/Sev2, anomalies, on-call load | actions into the sprint |
| Sprint review | every 2 weeks | KTLO split planned vs actual, toil %, follow-up closure | allocation for next sprint |
| Monthly ops report | monthly | MTTA, MTTR, SLA compliance trend, recurring clusters top 5 | stakeholder update section |
| Quarterly planning | quarterly | all, quarter over quarter | KTLO allocation for the quarter (KTLO vs roadmap) |
Triggers:
- Recurring cluster with ≥ 5 incidents in 90 days → problem record (a Feature under the KTLO epic) with an owner.
- Toil > 20% two sprints running → automation work prioritized from the KTLO allocation.
- SLA compliance below target for a service two months running → service review with the owning team and Product Owner.
Pitfalls¶
- MTTR goes down when teams resolve faster with less rigor. Read it with recurring-issue rate.
- Severity inflation makes compliance look worse; deflation makes it look better. Audit a sample of severities monthly.
- Low volume makes percentages swing. Show counts next to every ratio.