How to build a major incidents dashboard
It’s easy for major incidents to get lost amid noise, competing priorities, and uncertainty. Alerts spike, tickets flood in, and Slack channels light up before anyone can clearly spot that something big is happening.
This is where operational intelligence earns its keep. You don't need more data. You need a single, coherent view that cuts through the noise and turns uncertainty into action.
In this tutorial, we’ll build a dashboard that brings together alert activity, incident data, service impact and ownership into one view.
Instead of reacting to isolated symptoms, you’ll be able to see the scope, direction, and momentum of an unfolding event.
What is a major incident?
A major incident is one with widespread or business-critical impact, declared explicitly so the response is coordinated rather than left to whoever noticed first. In ServiceNow and most other ticketing tools that is the priority 1 class.
Four things decide how a major is handled, and the dashboard answers each one:
- Scale: How many majors are active, and how long since the last one.
- Momentum: Whether ticket and alert volume is still climbing.
- Blast radius: Which business services the child incidents land on.
- Load: Which assignment groups are carrying the response.
Data sources to use
Depending on your environment, you can mix and match different systems to build this dashboard. In this example, we’ll use:
- ServiceNow to capture major incidents, child tickets, ownership, and impact.
- Azure to surface alert volume and infrastructure signals.
This pairing gives us both sides of the story: the human workflow layer (tickets, assignments, escalation) and the system telemetry layer (alerts, spikes, degradation).
Both plugins ship a prebuilt dashboard, ServiceNow with Active Incidents and Azure with Alerts. Build your own when you need both sides of the story on one screen, or when you want to change how a signal is derived.
Configure the tiles
We’ll step through the dashboard tile by tile using different elements of the SquaredUp toolkit depending on the question we’re answering.
Each tile is documented as a self-contained guide. You can follow them independently, adapt the logic to your own data sources, or build the full board sequentially.
Active major incidents
This tile answers the most direct question of "how many major incidents are active right now?".
By filtering for priority 1 incidents that are still active, this block immediately tells you whether you’re in a crisis state. In a calm system, this number is zero. When it isn’t, the rest of the dashboard becomes your war-room.
See how to create an active major incidents tile for detailed instructions.
Time since last major incident
Resilience isn’t just about handling incidents well. It’s also about how often they occur.
This tile tracks the time elapsed since the last declared major incident. A growing duration suggests stability. A short interval between majors can signal systemic fragility.
See how to create a time since last major incident tile for detailed instructions.
Major incident status
Not all major incidents are equal. Some are newly declared. Others are stabilizing. Some are resolved but under observation.
This tile derives a health state from the incident lifecycle and maps it to clear operational signals:
- New: Error
- In Progress: Warning
- Resolved / Closed: Success
See how to create a major incident status tile for detailed instructions.
Ticket creation rate
When a major incident unfolds, ticket volume often tells you how fast impact is spreading.
By bucketing ticket creation over time, this chart shows whether the situation is accelerating, plateauing, or stabilizing. A rising slope suggests expanding impact. A flattening curve indicates containment.
Viewed alongside alert volume, it helps correlate technical failure with user-reported disruption.
See how to create a ticket creation rate tile for detailed instructions.
Major alert volume
Alerts often precede or amplify major incidents. This tile tracks alert counts over the selected timeframe, highlighting spikes that align with service degradation. By focusing on fired alerts and grouping by hour or minute, you gain a clear view of technical pressure building beneath the surface.
It helps answer "Is this incident isolated, or is the system under broader strain?"
See how to create a major alert volume tile for detailed instructions.
Affected services
Major incidents rarely impact a single component.
By grouping child incidents by business service or configuration item, this tile reveals the blast radius. Is everything concentrated in one platform? Or is the issue cascading across dependencies?
This is where root cause often begins to emerge visually.
See how to create an affected services tile for detailed instructions.
Assignment group load
Incidents don’t just impact systems. They impact people.
This tile shows how tickets are distributed across assignment groups during an active major. A tightly contained issue might sit with one team. A spreading incident often spans multiple groups, indicating escalation and cross-functional coordination.
It provides visibility into operational load and highlights where support pressure is concentrated.
See how to create an assignment group load tile for detailed instructions.
Next steps
You now have a dashboard that brings alert activity, incident data, service impact and ownership into one view.
To get the most value from this dashboard:
- Watch whether alert spikes precede ticket growth.
- Track which services the child incidents concentrate on as an incident spreads.
- Check how assignment load redistributes during escalation.
- Use the time since the last major to measure resilience over weeks and months.
During a major, the scale and direction of the event are on one screen rather than spread across the tools that hold the pieces.