Build vs Buy for Measuring Engineering Productivity Honestly
DORA metrics, deploy frequency, lead time, change failure rate, recovery time, are a genuinely useful, well-validated measure of delivery pipeline health. They were never designed to answer a different question leadership often wants answered: is this specific team, or this specific engineer, productive. Bolting that expectation onto DORA metrics is where a lot of well-intentioned measurement programs go wrong.
Measuring engineering output honestly means picking the right metrics for the actual question, and being honest about which questions metrics can't answer at all.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
DORA measures the pipeline, not the person
Deploy frequency and lead time describe how fast and how safely code moves from commit to production across a whole team's system, shaped heavily by tooling, process, and architecture, not any individual's effort or skill. Using DORA metrics to rank individual engineers misapplies a system-level measure to an individual-level question, and it reliably produces the wrong incentives: engineers optimizing for the metric instead of for genuinely useful work.
A common mistake is handing leadership a DORA dashboard with no framing, which invites them to read it as a scoreboard for people. The fix is to state, alongside the numbers, which question they answer (how healthy is the delivery pipeline?) and which they do not (who is productive?). For example, if lead time rises, investigate the pipeline before the people: slow reviews, flaky tests, manual approvals, or a fragile release process are far more likely causes than any one engineer. Framing the metric this way protects the team from misuse and keeps the conversation on things you can actually change.
The metrics that measure individuals tend to measure the wrong thing
Lines of code, commit count, and story points closed all correlate weakly with actual value delivered and strongly with gameable behavior: padding commits, splitting work into artificially small tickets, avoiding the hard, slow-to-close problems in favor of easy, fast-to-close ones. Any individual metric that's easy to measure automatically is usually easy to game, and a team that starts optimizing for the number instead of the underlying work will find a way to move it without actually improving anything.
Build in-house measurement only for what your team genuinely can't buy
Standard DORA metrics are well covered by existing tooling; building a custom pipeline to recompute deploy frequency and lead time from scratch is rarely worth the maintenance cost. Where building in-house makes more sense is a metric specific to your own product or process that no off-the-shelf tool tracks, time from a specific customer-reported bug to a verified fix, say, where the value comes from the custom definition, not from reinventing a standard metric a platform already handles well.
Where a platform earns its cost
Tools like ClickUp can automatically compute standard delivery metrics from existing commit and deployment data without a custom pipeline, and a workflow platform like Process Street can enforce and track the process steps, review, testing, deployment checklist, that actually drive change failure rate down, which is a more actionable lever than the failure rate number itself. Teams in the fastest, on-demand deploy cluster recover from failures quickly precisely because those underlying process steps are consistently followed, not because someone is watching a dashboard1.
Pair every quantitative metric with a qualitative check
A team's DORA numbers can look healthy while morale and code quality are quietly deteriorating, if the metrics are being optimized directly rather than as a side effect of genuinely good practice. Regular, honest qualitative check-ins, what's actually slowing the team down, where technical debt is accumulating, catch problems the quantitative dashboard won't show until they've already become a trend worth worrying about.
A decision rule worth writing down
Use DORA metrics to assess pipeline and process health at the team or system level, never as an individual performance metric. Build custom measurement only for something genuinely specific to your product that no platform already tracks well. Buy platform support for standard metrics and process enforcement, since that's where an off-the-shelf tool's maintained, battle-tested implementation beats a custom one built and maintained in-house.
Before adopting any engineering metric, run it through these checks:
- Confirm it describes pipeline or team health at the system level, since DORA measures were never designed to rank individual engineers.
- Ask whether it is easy to game; anything simple to compute automatically usually is, and people will optimize the number instead of the work.
- Check whether a platform already computes it from existing commit and deployment data before building a custom pipeline to recompute it.
- Build in-house only when the definition is specific to your product or process, such as time from a customer-reported bug to a verified fix.
- Pair the number with a regular qualitative check-in on what is slowing the team down and where technical debt is piling up.
A worked example: the metric that got gamed within a month
Say a team starts tracking pull requests merged per week as a productivity signal. Within a month, PR counts rise, but average PR size shrinks sharply and review quality drops, because engineers learned the fastest way to move the number was splitting normal work into smaller, faster-to-approve pieces rather than doing more of it. The metric technically improved. The actual output didn't. That's the pattern worth watching for with any individual metric simple enough to compute automatically: if it's easy to measure, it's usually easy to game.
What Good Looks Like
Good engineering productivity measurement means DORA metrics used at the team and pipeline level only, no individual metric relied on alone, platform tools handling standard computation, and regular qualitative checks alongside the numbers.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
ClickUp can compute standard delivery metrics automatically from existing commit and deployment data instead of a maintained custom pipeline
Process Street fits when the goal is enforcing and tracking the process steps, review, testing, deployment checklists, that actually drive the metrics you care about
Frequently Asked Questions
Is it ever appropriate to use DORA metrics for individual performance reviews?
No. DORA metrics describe system and pipeline health, shaped by tooling, architecture, and process far more than any one person's effort. Applying them to individuals misapplies the measure and tends to produce gaming behavior rather than genuinely better outcomes.
What should we measure instead for individual contribution?
Individual contribution is genuinely hard to measure with a single clean number, and most metrics that try, lines of code, commit count, story points, are easily gamed. Qualitative input from peers and managers, alongside the outcomes of the work itself, tends to be more honest than any automatically computed individual metric.
Should we build our own metrics dashboard or buy one?
Buy for standard delivery metrics like deploy frequency and lead time, since platforms already compute these reliably from existing data. Build only for something specific to your own product or process that no off-the-shelf tool tracks, where the custom definition itself is the value, not the computation.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
What to Track About Engineering Productivity Besides DORA
Why DORA's four metrics don't capture the whole picture of engineering health, and what to measure alongside them without turning metrics into a scoreboard.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
Build vs Buy for a Working Dev Environment on Day One
Why the real bottleneck in getting a new engineer to their first commit is usually access, not code, and where automated provisioning is worth the cost.
Diagnosing Slow Requests Before You Blame the Database
A step-by-step way to find out whether a slowdown is the network, the app, or the database, before you add caching or upgrade infrastructure to fix it.
Building a Throughput Benchmark You Can Actually Trust
A worksheet approach to benchmarking throughput: what load pattern to test, what to record, and how synthetic benchmarks lie about real capacity.