Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

The Metrics DORA Doesn't Capture for a RAG Team

DORA metrics still matter for a RAG team, but they miss what is specific to retrieval systems: how often relevance regresses, how long a bad retrieval takes to diagnose, and how much time goes to reactive chunking fixes. A team can score well on deployment cadence while retrieval quality quietly degrades.

A RAG team that only tracks DORA metrics can look high-performing on deployment cadence while its actual retrieval quality quietly degrades underneath.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why don't standard engineering metrics capture RAG-specific work?

DORA metrics measure how fast and reliably you ship code. They say nothing about whether what you shipped actually retrieves the right documents, because relevance quality isn't a deployment-pipeline concern, it's a property of the chunking, embedding, and ranking choices baked into the system. A team can deploy frequently, with a low change failure rate, while a slow but steady decline in retrieval relevance goes completely unmeasured by any of the standard metrics.

What does a relevance regression rate actually measure?

Track how often a change to chunking, retrieval parameters, or the embedding model measurably worsens results on your comparison query set, the same set used for embedding model migrations. This is the RAG-specific equivalent of a change failure rate: it tells you whether the team's changes are net positive for the thing users actually experience, not just whether the deployment pipeline ran cleanly.

Why does time-to-diagnose-a-bad-retrieval matter as its own metric?

When a user reports a wrong or irrelevant answer, how long does it take to determine why, a chunking issue, a stale index, a reranking bug, a genuine gap in the corpus? This is a distinct skill and tooling problem from general incident response, since it requires being able to trace a specific query back through retrieval to see exactly what was retrieved and why. Teams with good tracing answer this in minutes; teams without it can spend a day reconstructing what happened for a single bad answer.

How much of the team's time goes to reactive chunking fixes?

Track the share of engineering time spent on unplanned chunking or retrieval fixes, prompted by a bad answer someone noticed, against planned improvement work. A high reactive share signals accumulating technical debt in the chunking and retrieval logic, since a healthy system's fixes are mostly planned, not constant firefighting in response to user complaints.

Should these metrics replace DORA, or sit alongside it?

Alongside, not instead. DORA still tells you whether your delivery pipeline itself is healthy, and a RAG team benefits from fast, low-risk deployment just like any other team. These additional metrics answer a different question DORA was never built to answer: whether what you're shipping is actually improving retrieval, which is the thing a RAG product is ultimately judged on by its users.

A starter set of RAG-specific productivity metrics

  • Relevance regression rate: how often a change measurably worsens results on your comparison query set
  • Time to diagnose a bad retrieval: from a reported bad answer to a confirmed root cause
  • Reactive versus planned chunking and retrieval work, tracked as a share of total time
  • Comparison query set size and freshness, since a stale or small set makes the first metric unreliable
  • Time from a confirmed relevance issue to a shipped, verified fix

Start with one metric, not the full list at once

Trying to instrument all of these simultaneously usually means none of them get built well. Pick whichever one addresses your team's current biggest blind spot, often time-to-diagnose if debugging bad retrievals currently eats unpredictable chunks of time, get it measured reliably, and add the next one once the first is actually informing decisions rather than sitting in a dashboard nobody checks. A single metric that's trusted and actually used beats a full dashboard of numbers nobody looks at closely enough to act on. Revisit the priority every quarter or two, since the blind spot that mattered most when you started measuring isn't necessarily the one that matters most once the first metric has already improved it.

Executive Capability Standard

What Good Looks Like

Good productivity measurement for a RAG team means tracking relevance regression rate and time-to-diagnose alongside standard delivery metrics, so shipping fast and shipping retrieval quality both get measured, not just the first one.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand why DORA metrics don't capture retrieval quality, since that gap is what these additional metrics are meant to fill.
2. Do Manually:Track relevance regressions against your comparison query set by hand after each significant change, before building any automated measurement.
3. Delegate:Assign one person ownership of maintaining the comparison query set itself, since a stale or too-small set makes every metric built on it unreliable.
4. Automate:Automate running the comparison query set and flagging regressions as part of your deployment pipeline, so relevance regression rate gets measured on every change.
5. Buy:Use a project or workflow tool to track the reactive-versus-planned time split instead of estimating it from memory at the end of each cycle.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Are DORA metrics still useful for a RAG team?

Yes, for what they measure: delivery speed and reliability. They just don't capture whether what's shipped actually improves retrieval quality, which is a separate property of chunking, embedding, and ranking decisions that a deployment pipeline metric has no visibility into.

What's the RAG equivalent of a change failure rate?

A relevance regression rate: how often a change to chunking, retrieval parameters, or the embedding model measurably worsens results on a standing comparison query set. It answers whether changes are net positive for what users actually experience, not just whether the deployment itself succeeded.

Why track time to diagnose a bad retrieval separately from general incident metrics?

Because it requires tracing a specific query through retrieval to see exactly what was retrieved and why, a distinct skill and tooling need from general incident response. Teams with good tracing answer this in minutes; teams without it can spend a day reconstructing a single bad answer.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides