Incident metrics
Incident metrics measure the response rather than the failure: how quickly someone picked it up, how long service took to restore, how far the impact spread, and who absorbed the work.
The recipes in this section use ServiceNow for incident records and Jira Service Management for the response-time metrics, because JSM exposes dedicated SLA fields for them. Where those fields are not available, each recipe gives an expression that derives the same value from raw timestamps.
Response metrics
- MTTA: Mean time to acknowledge. The average elapsed time between an incident being created and someone responding to it, across the incidents in the selected timeframe. In Jira Service Management this reads the Time to First Response SLA field. It measures whether anyone is picking incidents up, and it usually degrades before MTTR does.
- MTTR: Mean time to recover. The average elapsed time between an incident being created and being resolved. In Jira Service Management this reads the Time to Resolution SLA field. Without one, it is calculated from the created and resolved timestamps and converted to minutes. MTTR is one of the four DORA metrics.
The two move independently, and that is what makes them worth having side by side. Rising MTTA against flat MTTR points at a saturated rotation. Flat MTTA against rising MTTR points at diagnosis, tooling or ownership rather than responsiveness.
State metrics
- Active major incidents: A count of priority 1 incidents still open. In a calm system this is zero, which makes it the fastest read on the dashboard.
- Major incident status: Derives an operational state from the incident lifecycle, mapping New to Error, In Progress to Warning, and Resolved or Closed to Success.
- Time since last major incident: Elapsed time since the last declared major. Read it as a resilience trend across months rather than as a live signal.
Impact metrics
- Ticket creation rate: Ticket creation bucketed over time. A rising slope means impact is still spreading, and it typically moves before anyone has confirmed the scope.
- Incidents over time: Incident counts per bucket, giving the volume context that a single current-state number cannot.
- Affected services: Child incidents grouped by business service or configuration item, which is the blast radius of the event.
- Assignment group load: How tickets distribute across assignment groups, showing where support pressure is concentrated and whether an incident has spread across teams.
- Unresolved critical issues: Open high-priority issues and bugs, returned from Jira with a JQL query. Standing risk that is still present while new change is being deployed.
What good looks like
The recipes use the following baselines as worked examples. They are a reasonable starting point rather than an industry target.
The MTTR thresholds are deliberately wide, because incident complexity varies and the monitor is meant to catch a sustained increase rather than one long incident.
Set your own from your own history rather than adopting these. A monitor that fires constantly gets ignored, which leaves you worse off than having no monitor at all.
Common pitfalls
- Starting the MTTR clock at acknowledgement: That measures fix time and hides detection delay. Both numbers are useful, but they are not the same metric, and mixing them makes recovery look faster than the customer experienced it.
- Averaging across incident classes: A priority 1 and a priority 4 in the same average produce a number that describes neither. Filter to a class, or split the tile.
- Reading a mean without the spread: A handful of very long incidents will drag an average well past what a typical incident looks like. Pair the average with a count or a distribution before drawing a conclusion from it.