How to build an on-call alert noise dashboard

For many teams, the on-call experience has quietly degraded. Engineers are woken up for alerts that don’t require action, the same services trigger incidents week after week, and no one can confidently answer a simple question: is our on-call setup actually healthy?

What exists instead is an endless stream of alerts, some post-incident metrics, and a growing sense of fatigue. This is where operational intelligence becomes essential.

In this tutorial, we’ll build a dashboard that brings together alert volume, incident data, and response times into one view. Rather than focusing on uptime alone, this dashboard looks at the operational reality of being on call.

The goal isn’t just visibility, but action. At the end of this guide, you'll have built a tool that reduces noise, improves readiness, and protects the people behind the screen.

What is alert noise?

Alert noise is the share of your alert volume that reaches a person and needs nothing from them. It is measurable, and it is what turns an on-call rotation from a safety net into a cause of attrition.

Four properties separate noise from signal, and the dashboard measures each one:

  • Volume: How many alerts fire per day, and whether that is trending up.
  • Timing: How much of it lands outside working hours.
  • Concentration: Which services generate most of it.
  • Actionability: What share is severe enough to justify waking someone.

MTTA and MTTR then show whether the noise is costing you response time.

Data sources to use

Depending on how your teams operate and the technologies they use, you can draw from multiple data sources to build this dashboard.

In this example, we’ll use:

This combination gives us both system-level telemetry and the human response layer in one place.

The Azure plugin ships a prebuilt Alerts dashboard. Build your own when you want the alert volume and the response times from your service desk on the same screen.

Configure the tiles

We’ll build this dashboard one tile at a time, drawing from different elements of the SquaredUp toolkit depending on each use case.

Each tile configuration is documented as a self-contained tutorial. This means you can work through them sequentially to build the full dashboard, or dip into specific tiles as needed and adapt the patterns to your own data sources and use cases.

Alert time distribution

This tile shows how many alerts fire outside working hours versus during the day.

A heavy skew toward evenings and weekends is often the clearest signal of noisy thresholds or low-value alerts. Reducing unnecessary out-of-hours noise is usually the quickest win for improving on-call quality of life without increasing risk.

See how to create an alert time distribution tile for detailed instructions.

Alerts over time

Here we establish a baseline. By tracking total alerts per day over time, you can see the true weight of on-call load.

Rather than focusing on individual alerts, this approach highlights trends in alert volume, making it easy to spot sustained increases, sudden spikes, or periods of relative calm.

See how to create an alerts over time tile for detailed instructions.

Alert volume by service

This tile reveals where the noise lives. By grouping alerts by service, it highlights which systems generate the most interruptions.

Instead of spreading effort thinly, you can focus tuning and investigation where it will make the biggest difference.

See how to create an alert volume by service tile for detailed instructions.

Actionable alerts

Not every alert deserves attention. This tile classifies Azure alert events by severity, creating a practical proxy for actionability. The result is a clearer view of how much of your alert stream genuinely requires human intervention.

See how to create an actionable alerts tile for detailed instructions.

MTTA

This tile tracks how quickly incidents are acknowledged. Viewed alongside alert volume, it shows whether noise is delaying the first human response.

See how to create an MTTA tile for detailed instructions.

MTTR

MTTR (mean time to recover) tracks how quickly it takes you to resolve an incident. Viewed alongside alert volume, they show whether noise is slowing response times or extending incident duration.

See how to create an MTTR tile for detailed instructions.

Next steps

You now have a dashboard that measures the on-call experience rather than uptime: how much alert volume there is, when it lands, where it comes from, and how long response takes.

To get the most value from this dashboard:

  • Review it weekly and tune the noisiest service first.
  • Compare out-of-hours volume before and after a threshold change.
  • Watch MTTA alongside alert volume. Rising acknowledgement time usually means the rotation is saturated.
  • Reassign ownership where the alerts and the team that can act on them have drifted apart.

Alert noise then becomes something you can evidence and reduce, rather than something the rotation absorbs.

Was this article helpful?


Have more questions or facing an issue?