Technology leadershipExplainer3 min readUpdated September 2026

DORA Metrics for a Small Engineering Team, Without the Dashboard Sprawl

DORA metrics are four measures of software delivery: how often you deploy, how long a change takes to reach production, how often changes cause failures, and how fast you recover. A small team can track all four with a spreadsheet and the data already in its repository and incident tickets.

The point isn't a benchmark score. It's to notice when delivery is slowing or getting riskier, early enough to change something. Small teams need a lighter touch than the vendor dashboards assume.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What are the four DORA metrics?

Each metric answers a specific question about your delivery process:

  • Deployment frequency. How often you release to production. Count deploys per week, not per engineer.
  • Lead time for changes. The time from a commit being merged to it running in production. Median, not average, because one stuck change skews the mean.
  • Change failure rate. The share of deploys that cause a problem needing a fix, rollback or hotfix.
  • Time to restore service. How long it takes to recover once a failure reaches users.

The first two describe speed, the last two describe stability. Reading them together is the point: a team that ships daily but breaks things every third deploy isn't fast, it's noisy.

How do you collect the data with tools you already have?

You don't need a new platform to start. Follow these steps:

  1. Pick one definition of a deploy. For most teams it's a successful run of the production deployment job. If you use GitHub Actions, that job's run history already has timestamps and the commit that triggered it.
  2. Record the merge time for each change. The pull request's merge timestamp minus the production deploy time gives lead time.
  3. Tag failures. When a deploy needs a rollback, hotfix or incident, mark it in your incident tracker with the deploy it came from. A simple label is enough.
  4. Log restore times. For each incident, write the time it started and the time service was back.
  5. Roll up weekly. Export to a spreadsheet and record the median lead time, deploy count, failures and median restore time.

When you outgrow the spreadsheet, an observability tool such as Datadog fits when you want to compare deploy times against latency and error trends. Confirm in a demo how it takes in deploy events.

A worked example for a team of six

Say a six-person team deploys eight times in a week, and two of those deploys needed a rollback. Change failure rate that week is two out of eight, one in four. On its own that number is alarming and meaningless, because eight deploys is a tiny sample. Track it over a month or a quarter before drawing a conclusion.

Now suppose lead time has crept from one day to four over two months. That's a more useful signal. Ask what changed: a slower test suite, a new manual approval step, larger pull requests waiting on one reviewer. Fix the largest cause, then watch the trend.

The method is the same each time: pick one metric that moved, ask what changed, make one change and measure again.

Why do DORA metrics mislead small teams?

Small samples make ratios jumpy, so use rolling windows and look at trends instead of single weeks. Other traps are more human:

  • Ranking individuals. These are team metrics. Using them to rate engineers invites gaming, such as splitting one change into ten to inflate deploys.
  • Comparing teams with different systems. A team maintaining a regulated device shouldn't be measured against a marketing site.
  • Counting deploys that don't matter. Documentation or config-only pushes inflate frequency.
  • Ignoring context after a bad week. One large incident doesn't mean the process is broken. Look for the pattern.

Pair these with service-level objectives, so stability is defined by what your customers experience. See the SLO template for small engineering teams for one way to set them.

What should you change first when a number looks bad?

Match the fix to the metric:

  • Low deployment frequency: shrink change sizes, and automate the manual steps between merge and release.
  • Long lead time: look for waiting, not working. Queues for review, slow builds and approval gates are the usual culprits.
  • High change failure rate: improve tests where failures cluster, add staged rollouts and make rollback a one-step action.
  • Slow restore time: write runbooks for your most common failures, and rehearse a rollback so it isn't the first time under pressure.

Pick one metric per quarter to improve. If you're choosing monitoring tools to support this, the Datadog vs New Relic vs Dynatrace comparison can help you decide what fits a team your size.

Executive Capability Standard

What Good Looks Like

The team can state its deploy count, median lead time, failure rate and restore time for the last quarter, and knows which one it is trying to improve.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read the four metric definitions and agree, as a team, what counts as a deploy, a failure and a restored service.
2. Do Manually:Export a month of production deploy history and incident records into a spreadsheet and calculate the four numbers by hand.
3. Delegate:Make one engineer responsible for a weekly metrics note and a short monthly review of the trend.
4. Automate:Emit deploy events from the CI pipeline and tag incidents with the deploy that caused them, so the numbers calculate themselves.
5. Buy:Add an observability platform with deploy markers once manual tracking takes more than an hour a week or the team spans several services.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Are DORA metrics useful for a team of five engineers?

Yes, if you read them as trends over weeks or months. With a small team, single-week numbers are noisy. The metrics still show whether releases are getting slower or riskier, and they give you a shared vocabulary for deciding what to fix.

How do I measure lead time for changes?

Take the time from when a change is merged to when it's running in production, and report the median across a period. If you deploy through a CI pipeline, both timestamps are already recorded. Some teams start the clock at first commit to include review time.

What counts as a change failure?

Any production deploy that causes degraded service and needs a rollback, hotfix or patch. Decide the definition once, write it down and apply it consistently. Trivial issues that never reached users usually don't count.

Should DORA metrics be used in performance reviews?

No. They describe how a team and its delivery system perform, not individual effort. Tying them to reviews encourages gaming, such as splitting changes to inflate deploy counts, and it destroys the honest incident reporting the metrics depend on.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides