Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

The Metrics That Matter Once You've Outgrown DORA

DORA's four keys, deploy frequency, lead time for changes, change failure rate, and time to restore, are a genuinely good starting point, and most teams that adopt them stop there. That's a mistake, because DORA measures what happens after code ships. It says nothing about why a feature took three weeks to get through review, or why your best engineer is quietly burning out while the dashboard still looks green.

This is a guide to the metrics worth adding once DORA stops telling you anything new, the ones to actively avoid, and how to introduce any of this without a team assuming it's about to be used against them.

Why Start With Pull Request Cycle Time, Not Velocity?

Break cycle time into its parts: time from PR opened to first review, time from first review to approval, time from approval to merge. Most teams that think they have a "slow reviewers" problem actually have a "nobody's assigned" problem, the PR sits for a day before anyone even looks at it. Once you split the metric, you can tell whether the fix is a review SLA, smaller PRs, or more reviewers on a bottlenecked codebase area, instead of a vague mandate to "review faster." Most teams that instrument this for the first time are surprised by which stage is actually the bottleneck; it's rarely the stage the loudest complaint in retro points to.

Break pull request cycle time into these three intervals:

  1. Time from pull request opened to first review, which often reveals an assignment problem rather than slow reviewers.
  2. Time from first review to approval, measured separately from the wait for that first review.
  3. Time from approval to merge, the last stretch before the change actually ships.

The SPACE Framework Covers What DORA Leaves Out

SPACE adds satisfaction, performance, activity, communication, and efficiency to the flow metrics DORA already gives you. In practice this means pairing a quarterly developer experience survey, three or four questions on friction, tooling, and focus time, with your flow data. A team that's shipping fast on DORA's numbers but scoring low on satisfaction is usually burning down technical debt it hasn't told anyone about yet, or living with a build pipeline everyone's learned to route around instead of fix. Keep the survey short enough that people actually finish it; a fifteen-question survey with a five percent completion rate tells you less than a four-question one everyone answers.

Vanity Metrics to Delete From Your Dashboard

Lines of code, commit count, and PRs merged per engineer all reward the wrong behavior: padding diffs, splitting one change into five trivial commits, merging small stuff while a hard problem sits untouched. If a metric can be gamed by an engineer optimizing for the number instead of the outcome, it will be, not out of bad faith, but because that's what any kind of review pressure tied to a metric does to a team over time. Replace headcount-normalized output metrics with team-level flow metrics; individual comparison metrics create exactly the wrong incentives in a codebase where most valuable work is collaborative, and where the hardest, most valuable tickets are also the ones that produce the least visible output per hour spent.

How Do You Tie Metrics to a Question, Not a Dashboard?

"Are we getting faster or slower at shipping safely" is a question. A dashboard with twenty tiles is not. Pick three to five metrics that answer specific questions your leadership team actually argues about, review latency, incident recovery time, and onboarding time to first merged PR are common ones, and build the rest of your reporting around explaining changes in those, not adding new tiles. A metrics program that grows every quarter without anyone removing anything is a sign nobody's actually using most of it, and it's worth a standing review that asks which tiles nobody has referenced in the last quarter.

Reading the Cluster Data Without Misusing It

The highest-performing teams ship several releases a day, while teams further down the curve stretch to weekly or monthly and slower still1. Use that kind of cluster data to set a realistic target for where your team sits today, not as a stick. The fastest way to destroy trust in a metrics program is to use it in a performance review the same quarter you introduced it; give the team a full cycle to see the numbers before anyone's compensation depends on them, and say so explicitly when you launch the program.

Rolling It Out Without Triggering a Trust Problem

Announce what you're measuring and why before the first number appears on a dashboard, and be specific about what it won't be used for, individual comparison, stack ranking, a factor in a layoff decision, since silence on that question gets filled with the worst assumption by default. Start with a pilot team that already trusts leadership rather than rolling it out org-wide on day one; a pilot that goes well gives you a real internal reference when the next team asks skeptical questions, which is more persuasive than any framework citation.

Executive Capability Standard

What Good Looks Like

A good metrics program tracks a small set of flow and satisfaction signals your team already trusts, not a dashboard nobody opens after the first month.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read the SPACE framework paper and pick two dimensions besides raw velocity to start tracking.
2. Do Manually:Ask engineers in retro what's actually slowing them down before you instrument anything new.
3. Delegate:Give one engineering manager ownership of the metric definitions so they don't quietly drift team to team.
4. Automate:Pull deploy frequency, lead time, and review turnaround straight from git and CI history instead of self-reported logs.
5. Buy:Bring in a purpose-built engineering intelligence tool once a spreadsheet can't keep pace with per-team benchmarking.

How to Get Started

Frequently Asked Questions

How often should we actually look at these metrics as a team?

Weekly for flow metrics like cycle time, since they're volatile and a single bad week is noise. Quarterly for the satisfaction survey and any trend analysis, since developer sentiment doesn't move week to week and over-surveying just creates fatigue without adding signal.

Should individual engineers see their own metrics broken out?

Team-level, yes, always. Individual-level flow metrics invite gaming and rarely account for the difference between someone doing deep, hard work and someone shipping small, easy tickets. If you need to evaluate an individual, that's a conversation with their manager grounded in code review and project outcomes, not a dashboard filter.

What's a reasonable first target if we're currently deploying monthly?

Aim for weekly before you aim for daily. The jump from monthly to weekly usually comes from smaller PRs and better test coverage; the jump from weekly to daily usually needs real investment in deployment automation and feature flags, so treat them as separate projects with separate timelines.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides