Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

A 30-Minute Audit for Finding Your Costliest Technical Debt

Ask an engineering team where their technical debt is and you'll get a list shaped by whoever complained most recently, not by what's actually expensive. This is a way to find the real answer in about half an hour, using evidence you likely already have instead of a fresh survey.

The point isn't to produce a perfectly ranked list. It's to replace a debate driven by recency and volume of complaint with one grounded in incidents, review friction, and onboarding pain, all of which you can pull from records you already keep.

Minute 0 to 10: pull your incident and hotfix history

Look at your last quarter of incidents and emergency fixes, and tag each one by the system or code path it touched. A pattern usually emerges fast: a small number of systems account for a disproportionate share of incidents, and those are your highest-cost debt, whether or not anyone has been vocal about them. This is a better signal than opinion because it's grounded in things that actually happened, not things people worry might happen. If you don't already tag incidents by system, this first pass will be rougher, but even a quick manual read through the last twenty incident summaries usually surfaces the same handful of names.

Minute 10 to 18: check where code review takes longest

Pull recent pull request review times and comment counts, filtered by which part of the codebase they touch. A system that reliably generates long review threads and multiple rounds of back-and-forth is one where the code's structure is fighting the reviewer, which is a direct, measurable cost even when nothing has broken in production yet. This surfaces debt that hasn't caused an incident but is still slowing every change down, quietly, in a way that rarely gets raised on its own because no single slow review feels worth complaining about.

Minute 18 to 25: ask where new engineers get stuck

Talk to, or pull notes from, whoever onboarded most recently. New engineers hit undocumented, confusing, or fragile parts of the system faster than anyone else, because they don't yet have the tribal knowledge that lets experienced engineers work around the rough edges without noticing them. Their confusion is a leading indicator the rest of the team has stopped seeing, precisely because everyone who could see it clearly has since learned to route around it out of habit.

Minute 25 to 30: rank by blast radius, not by annoyance

Sort what you've found by how many other systems or teams depend on the affected code, not by how irritating it is to work in. A messy but isolated internal tool is a lower priority than a merely inconvenient system that sits in the path of every other team's work, even if the isolated tool generates more complaints. Blast radius, not volume of complaint, is what determines actual cost.

Run the audit in these four passes:

  1. In the first ten minutes, pull the last quarter of incidents and hotfixes and tag each one by the system it touched.
  2. Next, review pull request times and comment counts by area of the codebase to spot code that fights its reviewers.
  3. Then ask the most recent hire where they got stuck, since new engineers hit fragile parts first.
  4. In the final minutes, rank what you found by how many systems or teams depend on the affected code.

What to do with the ranked list

Take the top one or two items and turn them into a scoped remediation plan with a defined outcome, not an open-ended "clean this up" ticket that will lose priority against every feature request that follows it. Debt remediation that competes directly against feature work without a specific scope and owner tends to lose that fight indefinitely, no matter how real the cost is.

A remediation plan is scoped when someone outside the team could tell whether it is finished. For example, instead of a ticket that says clean up the billing service, write that the two modules generating most incidents will be split from the shared code, with the owner, a target date, and the incident count you expect to see fall. Attach the audit evidence so the reason is visible at planning time. A common mistake is bundling several systems into one plan, which makes it easy to defer. Take the top item only, finish it, and let the next audit decide what comes second.

Re-run the audit instead of trusting the first result forever

The systems that show up on this list will shift as the product and the team change, and a debt audit done once a year ago is telling you about a codebase that no longer exists in quite the same shape. Keep the underlying data, incidents tagged by system, review times, onboarding notes, flowing continuously so the next audit takes minutes to run rather than requiring the same half hour of manual digging every time. Share the ranked list with the whole team after each run, including engineers who weren't part of producing it, since seeing the evidence tends to land better than being told the conclusion secondhand.

Executive Capability Standard

What Good Looks Like

Good technical debt prioritization ranks candidates using incident history, review friction, and onboarding pain, not the loudest recent complaint, and turns the highest-ranked items into scoped remediation plans with a defined outcome and an owner.

Building The Capability (5-Stage Skill Ladder)

1. Learn:pull your last quarter of incidents and tag each by the system it touched to see where the pattern actually is
2. Do Manually:run the 30-minute audit above by hand once per quarter and keep the ranked list somewhere visible to planning
3. Delegate:give a rotating owner responsibility for running the audit each quarter, so it isn't dependent on one person remembering
4. Automate:tag incidents and long-running pull requests by system automatically so the audit's first two steps take minutes instead of a manual pull each time
5. Buy:teams whose deployment frequency has slipped into a slower cluster are often carrying exactly this kind of unaddressed debt1, which is as much a reason to invest in remediation as any single incident

How to Get Started

Frequently Asked Questions

How often should this audit be repeated?

Quarterly is a reasonable default for most teams, timed to line up with planning cycles so remediation work has a real chance of being scheduled rather than just documented and forgotten until the next audit.

What if the highest-cost debt is in a system nobody wants to touch?

Treat that reluctance as a finding, because a system everyone avoids usually carries high risk and thin documentation. Both are debt in their own right and often explain why it keeps generating incidents. Scope a small remediation with a named owner instead of leaving it alone.

Should technical debt remediation get its own budget?

It helps. Debt work that has to win an ad hoc argument against every new feature request in every planning cycle tends to lose consistently, even when the data clearly shows its cost, simply because a shipped feature is easier to point to than an incident that didn't happen.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides