Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

What to Measure About Engineering Velocity Besides DORA

The four DORA metrics, deployment frequency, lead time for changes, change failure rate, and recovery time, are a genuinely useful, well-validated baseline, which is exactly why some teams stop there and assume they've covered engineering velocity completely. They haven't. DORA measures the delivery pipeline well; it says very little about the work happening before code ever reaches that pipeline.

The gaps below are specific and measurable, not vague calls for "more visibility." Each one is something you can start tracking this quarter alongside your existing DORA reporting.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What DORA doesn't see: time from idea to first commit

A team can look excellent on all four DORA metrics while individual engineers spend a large share of their time stuck before they ever write the first line of code: unclear requirements, waiting on a design decision, or blocked on a dependency from another team. Track time from when work is assigned to when meaningful work actually starts, since a team that's fast once coding begins but slow to start isn't actually as productive as the delivery metrics alone would suggest to someone reading only the dashboard.

What DORA doesn't see: interruption and context-switching cost

DORA's metrics are aggregate and don't capture how fragmented an individual engineer's day actually is. A team hitting good deployment frequency numbers in aggregate can still have engineers whose actual focused work time is heavily fragmented by meetings, interruptions, and context switching between unrelated tasks. Track a rough measure of protected, uninterrupted work time, even something as simple as a periodic survey, since this is often where real capacity is being lost without showing up anywhere in the delivery metrics.

What DORA doesn't see: rework caused by unclear requirements

A feature that ships fast, gets used, and then gets substantially rebuilt within weeks because the original requirements were wrong or incomplete looks fine on every DORA metric right up until the rebuild, which then also looks like normal, healthy delivery activity rather than what it actually is: wasted first-pass effort. Track a rough rework rate, work that gets substantially redone shortly after shipping, as a signal about requirements quality, not engineering execution quality. A rising rework rate alongside flat or improving DORA numbers is a specific, useful signal: the pipeline is healthy, but something upstream of it is sending work through that pipeline before it's actually ready.

What DORA doesn't see: how developers actually feel about the system

Engineers working in a codebase with high cognitive load, a lot of accumulated complexity that makes even small changes feel risky, will often be slower and more error-prone in ways that don't cleanly separate out in the DORA numbers from other causes. A brief, regular developer experience survey, asking specifically about friction points and confidence making changes, surfaces this directly rather than trying to infer it indirectly from delivery metrics that were designed to measure something else entirely.

Combine the two sets of metrics instead of picking one

DORA metrics and these upstream, qualitative signals answer different questions and are both useful together: DORA tells you how well the delivery pipeline performs once work is in flight, while the metrics above tell you about everything that happens before and around that pipeline. A team looking only at DORA can chase pipeline optimizations while missing that the actual bottleneck is upstream, in unclear requirements or interruption load that no amount of deployment tooling will fix.

For example, suppose deployment frequency dips one quarter while lead time and failure rate stay flat. DORA alone offers no explanation. Put a developer experience survey and a rework measure beside it, and a pattern may appear: engineers report more interruptions, and several features shipped early were rebuilt soon after because the requirements changed. A common mistake is to treat that finding as a reason to push the team harder on the pipeline. The better response is to fix the upstream cause, such as tightening requirements before work is assigned or protecting focus time, and then watch whether the DORA numbers recover on their own. Treat the pairing as a diagnostic tool, not a scorecard for ranking teams.

Report both sets together, not in separate conversations

It's common for DORA metrics to show up in an engineering leadership review while developer experience data, if it's collected at all, lives in a separate survey nobody cross-references against delivery numbers. Put both in the same report, side by side, so a leadership team asking why deployment frequency dipped last quarter can see, in the same place, whether a spike in reported interruptions or a rise in rework happened at the same time, instead of treating the two as unrelated stories with no obvious connection between them.

To add upstream visibility to your DORA reporting, track these signals:

  • Time from when work is assigned to when meaningful work starts, which exposes delay from unclear requirements, pending design decisions, or dependencies on other teams.
  • Protected, uninterrupted work time, even if you only measure it through a periodic survey, since fragmentation often hides lost capacity.
  • Rework rate, meaning work substantially redone shortly after shipping, read as a signal about requirements quality rather than engineering execution.
  • A short developer experience survey on friction points and confidence in making changes, which surfaces cognitive load directly.
  • Interruption load and rework in the same report as deployment frequency, so leaders can see whether they moved together.
Executive Capability Standard

What Good Looks Like

A complete view of engineering velocity combines DORA's delivery pipeline metrics with upstream signals: time from assignment to first meaningful work, protected focus time, rework rate, and a specific, behavior-focused developer experience survey, since DORA alone doesn't capture what happens before code reaches the pipeline.

Building The Capability (5-Stage Skill Ladder)

1. Learn:read your current DORA metrics alongside a rough estimate of how much rework happened in the last quarter to see whether the two tell a consistent story
2. Do Manually:run a short, specific developer experience survey by hand this quarter and compare the results against your DORA numbers
3. Delegate:assign an owner for tracking upstream, pre-pipeline metrics separately from whoever owns delivery pipeline tooling and DORA reporting
4. Automate:instrument time-to-first-commit and rework tracking directly in your issue tracker so the data collects continuously rather than through periodic manual review
5. Buy:a work management tool like ClickUp can hold the upstream timing data (assignment to first commit) that DORA doesn't track, and a checklist tool like Process Street is useful for documenting the measurement process itself so results stay comparable across quarters

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Should we replace DORA metrics with these instead?

No, they answer different questions and work best together. DORA metrics are a widely used measure of delivery pipeline health; the metrics above cover the upstream and human factors DORA was never designed to capture. Dropping DORA in favor of these would lose a genuinely useful, comparable baseline.

How do we measure developer experience without it turning into a popularity contest?

Keep the survey specific and behavior-focused, asking about concrete friction points, confidence making a change in a specific area, time lost to interruptions, rather than open-ended satisfaction questions. Specific questions produce more actionable, less easily-gamed signal than general sentiment.

Is rework rate hard to measure accurately?

Rework rate needs a working definition, such as a change to a feature within a set window of its release that goes beyond a normal bug fix. Even a rough, consistently applied definition tracked over time beats not measuring it at all. Consistency matters more than precision: agree on the window and the threshold once, then apply them the same way each quarter.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides