Alert metrics

Alert metrics measure the alert stream itself rather than the systems producing it: how much it fires, when it fires, where it comes from, and how much of it a person actually needed to see. They are the input to reducing on-call load without reducing coverage.

The recipes in this section use the Azure Alerts data stream with a monitor condition of Fired. The same metrics work against any alerting tool that exposes an alert list with a severity and a timestamp.

What each metric measures

  • Alerts over time: Alert volume, as a count of fired alerts per time bucket. This is the baseline the others are read against. The trend matters more than the number, because a stable count is manageable at almost any level and a climbing one is not.
  • Alert volume by service: The same count grouped by the resource or service that raised it. Noise is almost never spread evenly, so this is what tells you which single service to tune first.
  • Alert time distribution: The ratio of alerts firing outside working hours to those firing during them, categorized with a SQL Analytics query on the alert timestamp. This is the metric that maps most directly to how on-call actually feels.
  • Actionable alerts: A severity-based proxy for whether an alert needed a human. Sev1 and Sev2 are treated as likely to require action, Sev3 and Sev4 as noise candidates.
  • Major alert volume: High-severity alerts over the selected timeframe. Read alongside incident data, it separates an isolated failure from broader strain on the estate.

What good looks like

There is no universal target for alert volume, because it scales with the size of the estate. The ratios travel better than the counts.

  • Out-of-hours share: A rotation being woken for a meaningful share of its alerts is carrying noise, not coverage. This is usually the fastest thing to improve.
  • Actionable share: If most volume sits at Sev3 and Sev4, the configuration is reporting rather than alerting.
  • Concentration: Expect a small number of services to account for most of the volume. That list is the tuning backlog.

Common pitfalls

  • Counting every alert state: Each recipe filters to a monitor condition of Fired. Including resolved and acknowledged transitions counts the same alert several times and inflates every number on the dashboard.
  • Reading volume without severity: A fall in total alerts that is entirely Sev4 has not improved anyone’s night. Always read volume next to the actionable split.
  • Treating severity as ground truth: Actionability here is a proxy. Severity is assigned by whoever wrote the alert rule, so the metric is only as trustworthy as your severity discipline. If the split looks wrong, the rules are usually the problem rather than the tile.

Was this article helpful?


Have more questions or facing an issue?