Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

Beyond DORA: Building a Productivity Metric Set Your Engineers Won't Game

DORA's four metrics (deployment frequency, lead time for changes, change failure rate, and time to restore) are a strong, well-validated starting point, but they're deliberately narrow: they measure delivery performance, not the broader friction that slows a team down day to day. Extending beyond them is where most teams either build something genuinely useful or accidentally build a system engineers learn to game. Here's how to do the former.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What does DORA leave out of engineering productivity?

DORA's own research treats deployment frequency and the other three metrics as strong proxies for delivery performance across a wide range of teams, and that's worth taking seriously as a foundation rather than something to immediately move past1. What it doesn't capture well is individual or team-level friction: time lost to flaky tests, time spent waiting on code review, or the cognitive cost of context-switching between too many concurrent priorities.

Extending beyond DORA means adding metrics that fill those specific gaps, not replacing DORA's delivery-focused view with something unrelated.

Build these yourself: metrics tied to your own workflow

Pull request review turnaround time, the time from when a PR is opened to when it receives its first substantive review, is workflow-specific enough that it's usually cheap to build from your existing version control and CI data, and it directly measures a common source of engineer frustration that DORA's metrics don't touch. Time spent on flaky test reruns is similarly straightforward to pull from CI logs and directly quantifies a cost most teams underestimate.

These are worth building yourself because they're simple queries against data you already have, and building them in-house means you control exactly what counts as a review or a flaky failure, rather than accepting a vendor's generic definition.

Consider buying: metrics that need broad, normalized comparison

A platform that benchmarks your metrics against a wide set of similar companies, or that infers deeper signal from git history and calendar data using more sophisticated modeling than a straightforward query, is harder to replicate in-house well. This is where a dedicated engineering intelligence platform earns its cost: not for the metrics you could build yourself, but for the comparative context and the modeling effort that would otherwise take significant engineering time to build and maintain.

Be skeptical of a platform's more subjective, individual-level scoring features specifically. The comparative benchmarking is generally sound; scoring individual engineers from git activity alone is where these platforms are most prone to measuring the wrong thing.

How do you design metrics that engineers won't game?

Any metric that becomes a target changes the behavior it measures, often for the worse. A team measured on deployment frequency alone can hit the number by splitting trivial changes into many small deploys without shipping anything more valuable. A team measured on PR review turnaround alone can hit the number with fast, low-quality rubber-stamp reviews. The fix isn't avoiding metrics; it's pairing every efficiency metric with a corresponding quality metric, so gaming one makes the other visibly worse.

Review the metric set itself periodically for signs of gaming: a metric that's improved sharply without a corresponding process change is worth investigating before celebrating it.

Using the metrics without creating a surveillance culture

Report these metrics at the team level, not the individual level, and use them to start a conversation about process (why is review turnaround slow, what's driving flaky tests), not to rank individual engineers against each other. Metrics used to open a conversation about a systemic bottleneck tend to improve the underlying problem; metrics used to evaluate individuals tend to produce the exact gaming behavior described above, faster than any other cause.

Introducing a new metric without triggering distrust

The first time a new productivity metric shows up on a dashboard, engineers reasonably wonder what it's actually being used for. Head that off directly: explain in plain terms what the metric measures, what decision it's meant to inform, and explicitly what it won't be used for, such as individual performance ranking, before the dashboard goes live rather than after someone asks.

A metric introduced with that kind of context tends to get treated as a genuine improvement tool. One that shows up unexplained tends to get treated as a threat, which is exactly the condition that produces the gaming behavior this article is trying to avoid in the first place.

Before a new metric goes on a dashboard, tell the team:

  • What the metric measures, described in plain terms rather than as a formula.
  • What decision the metric is meant to inform, so nobody has to guess at its purpose.
  • What it will not be used for, explicitly including individual performance ranking.
  • That results are reported at the team level, to start a conversation about process rather than to compare engineers.
Executive Capability Standard

What Good Looks Like

Good productivity measurement means metrics extend DORA's delivery focus with workflow-specific data, pair efficiency metrics with quality counterparts to resist gaming, and report at the team level rather than individual ranking.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read DORA's own published research on what its four metrics do and don't capture, so any extension is deliberate rather than duplicating ground DORA already covers well.
2. Do Manually:Pull PR review turnaround and flaky test rate from your existing CI and version control data as a starting pair of additional metrics.
3. Delegate:Assign a specific engineer or small group to own the metric set's design and review it periodically for signs of gaming.
4. Automate:Automate the collection and team-level reporting of your chosen metrics so the review conversation happens on a regular cadence without manual data pulls each time.
5. Buy:Bring in a dedicated engineering intelligence platform once comparative benchmarking or deeper git-history modeling becomes valuable enough to justify the cost over building it in-house.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Should individual engineers ever see metrics about their own PRs or deploys?

Visibility into your own data is generally fine and can be useful for self-awareness. The risk is comparison and ranking against teammates, which is where the same data starts driving gaming behavior rather than genuine improvement.

How many metrics should a small team track beyond the core DORA set?

Two or three additional metrics, each tied to a specific, known friction point your team actually experiences, is more useful than a large dashboard of metrics nobody consistently reviews. Add a new metric only when it answers a question you're actually asking.

Is it worth building a custom dashboard, or should we start with spreadsheets?

Start with a spreadsheet or a simple scheduled query and a shared document. A custom dashboard is worth building once you've confirmed the specific metrics are genuinely useful and worth the ongoing maintenance, not before.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides