Engineering Metrics Worth Tracking Beyond DORA
The four DORA metrics, deploy frequency, lead time for changes, change failure rate, and time to restore, are a genuinely good foundation, and also an incomplete picture. They tell you a lot about how safely and quickly code ships. They tell you almost nothing about whether engineers are stuck waiting on reviews, drowning in interrupt work, or burning time on something DORA was never designed to measure.
This is a look at which additional metrics actually add real signal on top of DORA, and which ones sound useful but mostly just invite engineers to game the number instead of improving the thing it's supposed to measure.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Where DORA's Four Metrics Stop Telling You Anything
DORA measures the pipeline from code being written to code being safely in production. It has nothing to say about what happens before code exists, how long an engineer waited for a design decision, how much of their week went to meetings or interrupt work, or how long a pull request sat waiting for review before anyone looked at it.
How your team actually tracks against DORA's own deploy frequency clusters is itself worth knowing before layering more metrics on top: a team stuck releasing only once a month has a very different bottleneck to diagnose than one already deploying daily, and the additional metrics worth prioritizing differ accordingly1.
Pull Request Review Time: A Metric That Actually Helps
Time from a pull request being opened to first review, and separately, time from approval to merge, surfaces a bottleneck DORA doesn't capture at all: work that's finished but stuck waiting on another human. This is usually one of the easiest bottlenecks to fix once it's visible, since it's often solved by a review SLA norm rather than any process overhaul.
Track it at the team level, not the individual level. The moment this becomes an individual scorecard, it stops measuring review bottlenecks and starts measuring who's fastest to rubber-stamp things, which is the opposite of what you actually want.
Interrupt Work: The Metric Most Teams Are Flying Blind On
Ask engineers to roughly tag their time, planned roadmap work versus unplanned interrupts, production incidents, urgent bug fixes, ad hoc requests, for a couple of sprints, and the resulting split is often more revealing than any other metric on this list. A team that believes it's spending most of its time on the roadmap and is actually spending a third of it on interrupts has a very different problem than a team whose estimates are simply optimistic.
This doesn't need precise time tracking to be useful. A rough, honest self-report captured consistently over a few sprints tells you more than a perfectly accurate measurement nobody bothered to take.
The Metrics That Mostly Invite Gaming
Lines of code, commit count, and individual story points completed per engineer all correlate weakly with actual value delivered and strongly with how easy they are to inflate. An engineer who splits one unit of work into ten small commits, or pads a pull request with unnecessary changes, looks more productive on paper without having delivered anything more.
If a metric is easy to move without actually changing the underlying work, engineers under any kind of pressure will eventually find that shortcut, not out of bad faith, but because that's what the metric is rewarding. Avoid tracking anything at the individual level that has an easy, low-effort way to inflate it.
Turning Metrics Into an Actual Conversation, Not a Scoreboard
The value in any of these metrics comes from a team looking at its own trend and asking why, not from comparing engineers or teams against each other on a leaderboard. Review time creeping up might mean the team's grown faster than its review capacity; a rising interrupt-work share might mean a specific system needs investment, not that anyone's underperforming.
Bring these numbers to retrospectives as a discussion prompt, not a performance evaluation input. The moment a metric is tied to individual evaluation, the incentive to report it honestly starts competing with the incentive to look good on it, and honesty usually loses first.
A starting set that adds signal without becoming a scoreboard:
- Track time from pull request opened to first review, and from approval to merge, at the team level rather than per person.
- Have engineers roughly tag time as planned roadmap work or unplanned interrupts for a couple of sprints, and look at the split.
- Review trends in team retrospectives and ask why they moved, instead of comparing engineers or teams on a leaderboard.
- Skip lines of code, commit count and story points per engineer, since they are easy to inflate without doing more real work.
What Good Looks Like
Good engineering metrics practice means tracking a small set of team-level signals, DORA's four plus review time and interrupt-work share, used as a discussion prompt for trends, not as an individual scoreboard or performance input.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Are the four DORA metrics enough to understand engineering productivity?
They're a strong foundation for how safely and quickly code ships, but they don't capture what happens before code exists, like time waiting on design decisions, review bottlenecks, or how much of the week goes to unplanned interrupt work. A few targeted additional metrics fill in that picture without needing a large new tracking system.
Is pull request review time a useful metric to track?
Yes, tracked at the team level rather than per individual. It surfaces work that's finished but stuck waiting on another person, which DORA doesn't measure at all, and it's often one of the easier bottlenecks to fix once it's visible, frequently through a simple review SLA norm.
Why shouldn't we track lines of code or commit count per engineer?
Because they correlate weakly with real value delivered and are easy to inflate without doing more actual work, splitting one change into many small commits or padding a pull request, for example. Any metric that's easy to move without changing the underlying work will eventually get gamed under pressure, whether or not anyone intends to.
How should we actually use these metrics once we're tracking them?
As a discussion prompt in team retrospectives, looking at trends and asking why, not as an individual performance scoreboard. The moment a metric feeds into individual evaluation, the incentive to report it honestly starts competing with the incentive to look good on it, which undermines the metric's usefulness.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Cutting a New Engineer's First Week Down to a Day
A step-by-step way to cut new engineer environment setup from days to hours, including the setup steps teams forget to check when something breaks.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
Load Testing Numbers That Don't Match What Users Actually Feel
Why a clean throughput benchmark often fails to predict real-world scaling behavior, and how to build one around your real traffic mix and first bottleneck.