Beyond DORA: Picking Developer Productivity Metrics Worth Tracking
DORA's four metrics, deploy frequency, lead time, change failure rate, and recovery time, measure how fast and safely you ship. They say almost nothing about whether your engineers are stuck in meetings, blocked on reviews, or drowning in context-switching, which is usually where the real productivity problem lives.
The build-vs-buy question isn't really about dashboards. It's about whether you can define a handful of metrics precise enough to act on, without accidentally building a system that rewards the wrong behavior.
What DORA doesn't measure
The four DORA clusters, top, high, medium, and low, rank teams primarily by two numbers: deploy frequency and change failure rate. On deployment frequency specifically, the top cluster releases on demand, while the low cluster can take up to six months between deploys1.
None of that tells you why a specific pull request sits for three days before anyone reviews it, or whether your best senior engineer is spending half their week in incident calls instead of writing code. Review latency, PR size, and time in focused work fill that gap, but only if you track them per team, not per individual, and only if the team agrees on what they're for before rollout.
None of these additional metrics are useful in isolation either. A short review latency paired with a rising change failure rate usually means reviews are getting rubber-stamped, not that the team got faster at reviewing well. Read the metrics as a set, not one at a time.
Build vs buy: what actually differs
Building your own dashboard from Git and calendar data is genuinely feasible with a weekend of engineering time and a few API queries. What you don't get for free is normalization across repos with different branching strategies, historical benchmarking, or a UI your engineering managers will actually open more than once.
A commercial engineering-metrics platform buys you that normalization and a maintained integration layer, at the cost of another vendor with access to your Git history and a recurring line item. For a small team, an internal dashboard usually wins on cost. Past a certain size, the maintenance burden of keeping a homegrown tool current with every new repo and CI change starts to compete with actual engineering work.
Metrics that reward the wrong behavior
- Lines of code or commit count, which rewards padding diffs and penalizes engineers who ship a clean, small fix.
- Story points completed, which teams learn to game within a quarter once it's tied to anything that matters to them.
- Individual PR count as a standalone number, which pushes people toward small, low-risk changes and away from the large refactors that actually need doing.
- Any metric reported to leadership as an individual scorecard rather than a team-level trend, which turns a diagnostic tool into a performance-review weapon and destroys the honesty of the data within a quarter.
A four-metric starting set that holds up
Deploy frequency and change failure rate from DORA, plus PR review latency (time from open to first review) and PR cycle time (time from open to merge). All four are objective, pulled straight from your Git and CI data, and hard to game without also making the underlying work worse.
Track them as a rolling trend per team, not a single-sprint number, and review them in a retro, not a one-on-one. If a number moves, ask the team what changed before assuming the metric is telling you the truth.
A worked example: what the four metrics would have caught
Say a team's PR cycle time doubles over a month while deploy frequency stays flat. Looked at alone, cycle time is easy to write off as normal variance. Looked at alongside a rising change failure rate on the same team, it's a different story: reviews are probably being rushed to compensate for the slower cycle, and the two numbers moving together is the actual signal, not either one on its own.
This is the case for tracking a small set together rather than picking a single favorite metric and watching it in isolation. A dashboard with one number invites a single-minded fix; a dashboard with four invites the harder, more honest question of what's actually changed in how the team works.
What Good Looks Like
A good productivity metrics setup tracks a small number of team-level, hard-to-game signals, reviewed as a trend and discussed with the team, not handed down as a scorecard.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we buy a developer productivity platform or build our own?
Build if you're a small team and want to start with deploy frequency and PR cycle time pulled straight from your Git API. Buy once the maintenance of a homegrown dashboard across a growing number of repos starts costing more engineering time than the platform would.
Is it safe to track individual engineer metrics?
Track them for your own diagnostic use if you must, but don't report them upward as a scorecard. The moment a metric like PR count becomes something a review depends on, engineers optimize for the number instead of the work, and the data stops being useful within a quarter.
How often should we review these metrics?
Weekly is often too noisy and monthly can be too slow to catch a regression early. A rolling multi-week trend, reviewed every two weeks in a team retro, gives enough signal to spot a real shift without reacting to normal week-to-week variance.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
Build vs. Buy for a New Engineer's First Working Day
Whether to build your own developer environment automation or buy a hosted one, based on how often you actually hire and what your stack demands.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
What to Track About Engineering Productivity Besides DORA
Why DORA's four metrics don't capture the whole picture of engineering health, and what to measure alongside them without turning metrics into a scoreboard.
What 'Zero Trust' Actually Requires From Every Device on Your Network
What zero trust device verification actually requires in practice, beyond the buzzword, and where small teams should start first.