The Metrics That Matter Once You've Outgrown DORA
DORA's four keys, deploy frequency, lead time for changes, change failure rate, and time to restore, are a genuinely good starting point, and most teams that adopt them stop there. That's a mistake, because DORA measures what happens after code ships. It says nothing about why a feature took three weeks to get through review, or why your best engineer is quietly burning out while the dashboard still looks green.
This is a guide to the metrics worth adding once DORA stops telling you anything new, the ones to actively avoid, and how to introduce any of this without a team assuming it's about to be used against them.
Why Start With Pull Request Cycle Time, Not Velocity?
Break cycle time into its parts: time from PR opened to first review, time from first review to approval, time from approval to merge. Most teams that think they have a "slow reviewers" problem actually have a "nobody's assigned" problem, the PR sits for a day before anyone even looks at it. Once you split the metric, you can tell whether the fix is a review SLA, smaller PRs, or more reviewers on a bottlenecked codebase area, instead of a vague mandate to "review faster." Most teams that instrument this for the first time are surprised by which stage is actually the bottleneck; it's rarely the stage the loudest complaint in retro points to.
Break pull request cycle time into these three intervals:
- Time from pull request opened to first review, which often reveals an assignment problem rather than slow reviewers.
- Time from first review to approval, measured separately from the wait for that first review.
- Time from approval to merge, the last stretch before the change actually ships.
The SPACE Framework Covers What DORA Leaves Out
SPACE adds satisfaction, performance, activity, communication, and efficiency to the flow metrics DORA already gives you. In practice this means pairing a quarterly developer experience survey, three or four questions on friction, tooling, and focus time, with your flow data. A team that's shipping fast on DORA's numbers but scoring low on satisfaction is usually burning down technical debt it hasn't told anyone about yet, or living with a build pipeline everyone's learned to route around instead of fix. Keep the survey short enough that people actually finish it; a fifteen-question survey with a five percent completion rate tells you less than a four-question one everyone answers.
Vanity Metrics to Delete From Your Dashboard
Lines of code, commit count, and PRs merged per engineer all reward the wrong behavior: padding diffs, splitting one change into five trivial commits, merging small stuff while a hard problem sits untouched. If a metric can be gamed by an engineer optimizing for the number instead of the outcome, it will be, not out of bad faith, but because that's what any kind of review pressure tied to a metric does to a team over time. Replace headcount-normalized output metrics with team-level flow metrics; individual comparison metrics create exactly the wrong incentives in a codebase where most valuable work is collaborative, and where the hardest, most valuable tickets are also the ones that produce the least visible output per hour spent.
How Do You Tie Metrics to a Question, Not a Dashboard?
"Are we getting faster or slower at shipping safely" is a question. A dashboard with twenty tiles is not. Pick three to five metrics that answer specific questions your leadership team actually argues about, review latency, incident recovery time, and onboarding time to first merged PR are common ones, and build the rest of your reporting around explaining changes in those, not adding new tiles. A metrics program that grows every quarter without anyone removing anything is a sign nobody's actually using most of it, and it's worth a standing review that asks which tiles nobody has referenced in the last quarter.
Reading the Cluster Data Without Misusing It
The highest-performing teams ship several releases a day, while teams further down the curve stretch to weekly or monthly and slower still1. Use that kind of cluster data to set a realistic target for where your team sits today, not as a stick. The fastest way to destroy trust in a metrics program is to use it in a performance review the same quarter you introduced it; give the team a full cycle to see the numbers before anyone's compensation depends on them, and say so explicitly when you launch the program.
Rolling It Out Without Triggering a Trust Problem
Announce what you're measuring and why before the first number appears on a dashboard, and be specific about what it won't be used for, individual comparison, stack ranking, a factor in a layoff decision, since silence on that question gets filled with the worst assumption by default. Start with a pilot team that already trusts leadership rather than rolling it out org-wide on day one; a pilot that goes well gives you a real internal reference when the next team asks skeptical questions, which is more persuasive than any framework citation.
What Good Looks Like
A good metrics program tracks a small set of flow and satisfaction signals your team already trusts, not a dashboard nobody opens after the first month.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we actually look at these metrics as a team?
Weekly for flow metrics like cycle time, since they're volatile and a single bad week is noise. Quarterly for the satisfaction survey and any trend analysis, since developer sentiment doesn't move week to week and over-surveying just creates fatigue without adding signal.
Should individual engineers see their own metrics broken out?
Team-level, yes, always. Individual-level flow metrics invite gaming and rarely account for the difference between someone doing deep, hard work and someone shipping small, easy tickets. If you need to evaluate an individual, that's a conversation with their manager grounded in code review and project outcomes, not a dashboard filter.
What's a reasonable first target if we're currently deploying monthly?
Aim for weekly before you aim for daily. The jump from monthly to weekly usually comes from smaller PRs and better test coverage; the jump from weekly to daily usually needs real investment in deployment automation and feature flags, so treat them as separate projects with separate timelines.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
What to Track About Engineering Productivity Besides DORA
Why DORA's four metrics don't capture the whole picture of engineering health, and what to measure alongside them without turning metrics into a scoreboard.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
How Long Does It Take a New Engineer to Ship Something Real?
Time to first meaningful commit is a real, measurable signal. Here is how to find where new hires actually get stuck and fix it without a full rebuild.
What "Zero Trust" Actually Means for Device Verification
Zero trust device verification means a device is trusted continuously, based on its current state, not once at login. Here is what that actually requires.
How to Benchmark Your System Before It Has to Scale
A practical runbook for benchmarking throughput and capacity before you actually need the headroom, so scaling decisions are based on data, not guesses.
Finding Your Real Latency Bottleneck Before Customers Do
A practical approach to latency benchmarking: how to define what slow means, set a budget, and find where the time actually goes before users complain.