Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Build vs Buy for Measuring Engineering Productivity Honestly

DORA metrics, deploy frequency, lead time, change failure rate, recovery time, are a genuinely useful, well-validated measure of delivery pipeline health. They were never designed to answer a different question leadership often wants answered: is this specific team, or this specific engineer, productive. Bolting that expectation onto DORA metrics is where a lot of well-intentioned measurement programs go wrong.

Measuring engineering output honestly means picking the right metrics for the actual question, and being honest about which questions metrics can't answer at all.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

DORA measures the pipeline, not the person

Deploy frequency and lead time describe how fast and how safely code moves from commit to production across a whole team's system, shaped heavily by tooling, process, and architecture, not any individual's effort or skill. Using DORA metrics to rank individual engineers misapplies a system-level measure to an individual-level question, and it reliably produces the wrong incentives: engineers optimizing for the metric instead of for genuinely useful work.

A common mistake is handing leadership a DORA dashboard with no framing, which invites them to read it as a scoreboard for people. The fix is to state, alongside the numbers, which question they answer (how healthy is the delivery pipeline?) and which they do not (who is productive?). For example, if lead time rises, investigate the pipeline before the people: slow reviews, flaky tests, manual approvals, or a fragile release process are far more likely causes than any one engineer. Framing the metric this way protects the team from misuse and keeps the conversation on things you can actually change.

The metrics that measure individuals tend to measure the wrong thing

Lines of code, commit count, and story points closed all correlate weakly with actual value delivered and strongly with gameable behavior: padding commits, splitting work into artificially small tickets, avoiding the hard, slow-to-close problems in favor of easy, fast-to-close ones. Any individual metric that's easy to measure automatically is usually easy to game, and a team that starts optimizing for the number instead of the underlying work will find a way to move it without actually improving anything.

Build in-house measurement only for what your team genuinely can't buy

Standard DORA metrics are well covered by existing tooling; building a custom pipeline to recompute deploy frequency and lead time from scratch is rarely worth the maintenance cost. Where building in-house makes more sense is a metric specific to your own product or process that no off-the-shelf tool tracks, time from a specific customer-reported bug to a verified fix, say, where the value comes from the custom definition, not from reinventing a standard metric a platform already handles well.

Where a platform earns its cost

Tools like ClickUp can automatically compute standard delivery metrics from existing commit and deployment data without a custom pipeline, and a workflow platform like Process Street can enforce and track the process steps, review, testing, deployment checklist, that actually drive change failure rate down, which is a more actionable lever than the failure rate number itself. Teams in the fastest, on-demand deploy cluster recover from failures quickly precisely because those underlying process steps are consistently followed, not because someone is watching a dashboard1.

Pair every quantitative metric with a qualitative check

A team's DORA numbers can look healthy while morale and code quality are quietly deteriorating, if the metrics are being optimized directly rather than as a side effect of genuinely good practice. Regular, honest qualitative check-ins, what's actually slowing the team down, where technical debt is accumulating, catch problems the quantitative dashboard won't show until they've already become a trend worth worrying about.

A decision rule worth writing down

Use DORA metrics to assess pipeline and process health at the team or system level, never as an individual performance metric. Build custom measurement only for something genuinely specific to your product that no platform already tracks well. Buy platform support for standard metrics and process enforcement, since that's where an off-the-shelf tool's maintained, battle-tested implementation beats a custom one built and maintained in-house.

Before adopting any engineering metric, run it through these checks:

  • Confirm it describes pipeline or team health at the system level, since DORA measures were never designed to rank individual engineers.
  • Ask whether it is easy to game; anything simple to compute automatically usually is, and people will optimize the number instead of the work.
  • Check whether a platform already computes it from existing commit and deployment data before building a custom pipeline to recompute it.
  • Build in-house only when the definition is specific to your product or process, such as time from a customer-reported bug to a verified fix.
  • Pair the number with a regular qualitative check-in on what is slowing the team down and where technical debt is piling up.

A worked example: the metric that got gamed within a month

Say a team starts tracking pull requests merged per week as a productivity signal. Within a month, PR counts rise, but average PR size shrinks sharply and review quality drops, because engineers learned the fastest way to move the number was splitting normal work into smaller, faster-to-approve pieces rather than doing more of it. The metric technically improved. The actual output didn't. That's the pattern worth watching for with any individual metric simple enough to compute automatically: if it's easy to measure, it's usually easy to game.

Executive Capability Standard

What Good Looks Like

Good engineering productivity measurement means DORA metrics used at the team and pipeline level only, no individual metric relied on alone, platform tools handling standard computation, and regular qualitative checks alongside the numbers.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit which metrics your team currently tracks and whether any are being applied to individuals rather than the system as a whole.
2. Do Manually:Run one honest qualitative check-in asking what's actually slowing the team down, separate from whatever the dashboard shows.
3. Delegate:Assign an owner for interpreting metric trends in context, so a number moving isn't treated as automatically good or bad without investigation.
4. Automate:Automate standard DORA metric computation through existing tooling rather than maintaining a custom pipeline for numbers a platform already handles.
5. Buy:Bring in a platform for delivery metrics and process enforcement once manual tracking is consuming more time than the platform would cost.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is it ever appropriate to use DORA metrics for individual performance reviews?

No. DORA metrics describe system and pipeline health, shaped by tooling, architecture, and process far more than any one person's effort. Applying them to individuals misapplies the measure and tends to produce gaming behavior rather than genuinely better outcomes.

What should we measure instead for individual contribution?

Individual contribution is genuinely hard to measure with a single clean number, and most metrics that try, lines of code, commit count, story points, are easily gamed. Qualitative input from peers and managers, alongside the outcomes of the work itself, tends to be more honest than any automatically computed individual metric.

Should we build our own metrics dashboard or buy one?

Buy for standard delivery metrics like deploy frequency and lead time, since platforms already compute these reliably from existing data. Build only for something specific to your own product or process that no off-the-shelf tool tracks, where the custom definition itself is the value, not the computation.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides