Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Build Your Own One-Page Production Risk Register

Most engineering teams carry a real, working knowledge of where their system is fragile, which dependency has caused trouble before, which part of the stack nobody's confident touching. The problem is that this knowledge usually lives only in a few senior engineers' heads, informally, which means it disappears the moment they're out sick, on vacation, or leave the company.

A one-page production risk register turns that tribal knowledge into something the whole team, including a new hire, can read and act on. Here's how to build one.

Column one: the component or dependency

List each meaningful piece of your architecture as its own row: your primary database, each critical third-party API you depend on, your deployment pipeline, your background job system, and anything else that would cause real damage if it failed. Keep this granular enough to be useful but not so granular that the register itself becomes a maintenance burden nobody keeps updated. A dozen to two dozen rows is usually the right range for a small or mid-sized system.

Column two: the known failure mode

For each row, write the specific, concrete way that component tends to fail, based on what's actually happened before, not a generic worry. 'The database connection pool can exhaust under a traffic spike' is useful. 'Database issues' is not, because it gives a new engineer nothing to actually check when something goes wrong at two in the morning. Pull this straight from past incident reviews where you have them, since the most realistic failure modes are almost always the ones that already happened once, not the hypothetical ones someone imagines from scratch.

Column three: the current mitigation, honestly stated

Write down what's actually in place today to catch or prevent that failure, and be honest about gaps rather than writing down what you intend to build eventually. 'No monitoring on this yet' is a genuinely useful entry, since it tells the next person exactly where the real exposure sits, while a vague or aspirational entry just hides the gap behind a false sense of coverage. This column is the one most tempting to overstate, precisely because writing down a real gap can feel like an admission, but an honest gap is far more useful to the next engineer than a comforting fiction.

For example, a row for the primary database might read: failure mode, connection pool exhausts under a traffic spike; mitigation, an alert exists but nobody has tested the runbook it points to. That entry is more valuable than a confident line saying the database is monitored, because it tells the next engineer which step to rehearse. A useful habit is to mark each mitigation as tested, untested or absent, and to revisit the untested entries first during the quarterly review. Untested mitigations are where a comforting register most often hides a real gap, since a control that exists only on paper may not work when it's needed.

Column four: who should be paged, and what should they check first?

For each row, name the person or team who'd actually get paged for that failure, and the first one or two things worth checking, drawn from real past incidents where possible. This is the single most valuable column during an actual incident, since it turns a register that's merely informative into one that's directly actionable the moment something breaks at an inconvenient hour.

How often should you revisit a one-page risk register?

A register that sprawls into a lengthy document stops getting read during an actual incident, when reading speed matters most. Force yourself to keep it to one page, cutting detail rather than adding more rows, and set a recurring quarterly review where the team checks whether each row still reflects reality, since an architecture changes faster than most documentation naturally keeps up with on its own.

A one-page register comes together in these steps:

  1. List each meaningful component or dependency as its own row, aiming for a dozen to two dozen rows so the register stays maintainable.
  2. For each row, write the specific failure mode that has actually happened before, pulled from past incident reviews where you have them.
  3. State the current mitigation honestly, including gaps such as no monitoring yet, rather than what you intend to build eventually.
  4. Name who gets paged for that failure and the first one or two things to check, drawn from real past incidents.
  5. Cut detail to keep it on one page, give it a named owner, and set a recurring quarterly review.

A worked example: the register that caught a growing gap

Say a quarterly review of the register surfaces a row for a background job system with a mitigation column that still reads 'manually restarted if it stops,' written eighteen months earlier when the team was smaller. In that time, the job system has taken on real, revenue-relevant work nobody explicitly decided it should carry, and the informal manual-restart mitigation quietly stopped being adequate somewhere along the way, without any single moment where that shift was obvious. The register doesn't fix the gap by itself. It's what makes the gap visible enough for someone to actually decide to fix it, rather than the team continuing to rely on tribal knowledge that hasn't kept pace with what the system actually does now.

Executive Capability Standard

What Good Looks Like

A useful production risk register is a single page listing each critical component's known failure mode, its honest current mitigation, and who to page first, reviewed and corrected on a regular cadence rather than written once and forgotten.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand which of your system's components would cause the most damage if they failed, as a starting point for the register.
2. Do Manually:Manually interview your most senior engineers about the failure modes they already carry in their heads and write the first version of the register from that.
3. Delegate:Assign a named owner for the register's quarterly review, separate from whoever happens to be busiest at the time.
4. Automate:Link the register to your incident tracker so a new incident automatically prompts a check of whether it should update an existing row.
5. Buy:Use an existing incident management or runbook platform to host the register rather than a document that's easy to lose track of.

How to Get Started

Frequently Asked Questions

Who should be responsible for keeping the risk register updated?

Give it a named owner, not a shared, unowned responsibility, since anything without a specific owner tends to go stale the first time the team gets busy. The owner doesn't need to know every answer themselves, just needs to run the quarterly review and chase down updates from whoever does know.

Should the register include every possible failure, no matter how unlikely?

No. Focus on failures that are either likely enough to plausibly happen again or severe enough that even a small chance of them is worth planning for. A register trying to capture every conceivable failure becomes too long to actually read during a real incident, which defeats its purpose.

Is this the same thing as a formal disaster recovery plan?

Not quite: a disaster recovery plan covers large, rare events, while this register captures the everyday, specific ways your system tends to break. It is more operational, holding the kind of knowledge a senior engineer already carries informally and a new hire can't learn without asking around.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides