Deployment metrics

Deployment metrics measure the delivery pipeline itself: how often changes reach production, how long they take to get there, how often they fail, and how much risk is standing in front of the next release. Three of the four DORA metrics live in this section.

The recipes use Azure DevOps, mostly its Build Runs and Deployments data streams, with Jira supplying incident data where a metric needs both sides. The same metrics work against any pipeline that exposes build or release records with a result and a timestamp.

What each metric measures

  • Deployment frequency: Counts the deployments completing on each day of the timeframe, then averages those daily counts. A DORA metric, and the simplest read on delivery throughput.
  • Deployment frequency over time: Where the average answers how often, this answers when. A point per day, so gaps, clusters and weekend releases stay visible. It is the first thing worth checking against an incident timeline.
  • Lead time: The elapsed time from a change being created to it reaching production. A DORA metric, and the one that reflects the whole pipeline rather than any single stage. Read from Azure DevOps with a WIQL query.
  • Deployment success rate: The share of deployments completing without failure. A decline usually arrives before the trouble does, pointing at pipeline instability, thinning test coverage or environmental drift.
  • Deployment failures: A count of build runs that failed, by day. These never reached production, so it measures pipeline reliability rather than customer impact.
  • Change failure rate: Correlates successful build runs with incident-priority bugs over the same timeframe, so it measures the changes that did ship and then caused problems. A DORA metric, and the one that needs two data sources combined with a SQL Analytics query.
  • Deployment health: Rolls the current state of deployments into a single green, amber or red signal, for when the question is only whether it is safe to ship.
  • Deployment risk: A composite score weighted from other published KPIs. Build it last, because it returns nothing until the tiles feeding it exist.

What good looks like

There is no universal target for deployment frequency, because it scales with team size and release model. The DORA research bands run from fewer than one deployment a month at the low end to several a day at the high end, so the useful comparison is your own trend rather than the band you land in.

  • Direction over level: A stable rate at any frequency is healthier than one drifting downwards. Read the slope, not the number.
  • Failures against change failure rate: Failures caught in the pipeline are the pipeline doing its job. Both rising together is the combination worth acting on.
  • Lead time spread: An average hides the long tail. A handful of changes stuck for weeks moves it without describing anything typical, so read it next to the timeline.

Common pitfalls

  • Counting attempts as deployments: The Build Runs stream includes runs that failed. Every frequency recipe filters Result to Succeeded, and leaving it open inflates throughput with work that never shipped.
  • Reading deployment failures as change failure rate: One counts what the pipeline stopped, the other counts what got through and caused an incident. They move independently and mean different things.
  • Building the composite first: Deployment risk reads from published KPIs, so change failure rate, deployment success rate and unresolved critical issues all have to exist and be published before it returns anything.

Was this article helpful?


Have more questions or facing an issue?