A 30-Minute Audit for Finding Your Costliest Technical Debt
Ask an engineering team where their technical debt is and you'll get a list shaped by whoever complained most recently, not by what's actually expensive. This is a way to find the real answer in about half an hour, using evidence you likely already have instead of a fresh survey.
The point isn't to produce a perfectly ranked list. It's to replace a debate driven by recency and volume of complaint with one grounded in incidents, review friction, and onboarding pain, all of which you can pull from records you already keep.
Minute 0 to 10: pull your incident and hotfix history
Look at your last quarter of incidents and emergency fixes, and tag each one by the system or code path it touched. A pattern usually emerges fast: a small number of systems account for a disproportionate share of incidents, and those are your highest-cost debt, whether or not anyone has been vocal about them. This is a better signal than opinion because it's grounded in things that actually happened, not things people worry might happen. If you don't already tag incidents by system, this first pass will be rougher, but even a quick manual read through the last twenty incident summaries usually surfaces the same handful of names.
Minute 10 to 18: check where code review takes longest
Pull recent pull request review times and comment counts, filtered by which part of the codebase they touch. A system that reliably generates long review threads and multiple rounds of back-and-forth is one where the code's structure is fighting the reviewer, which is a direct, measurable cost even when nothing has broken in production yet. This surfaces debt that hasn't caused an incident but is still slowing every change down, quietly, in a way that rarely gets raised on its own because no single slow review feels worth complaining about.
Minute 18 to 25: ask where new engineers get stuck
Talk to, or pull notes from, whoever onboarded most recently. New engineers hit undocumented, confusing, or fragile parts of the system faster than anyone else, because they don't yet have the tribal knowledge that lets experienced engineers work around the rough edges without noticing them. Their confusion is a leading indicator the rest of the team has stopped seeing, precisely because everyone who could see it clearly has since learned to route around it out of habit.
Minute 25 to 30: rank by blast radius, not by annoyance
Sort what you've found by how many other systems or teams depend on the affected code, not by how irritating it is to work in. A messy but isolated internal tool is a lower priority than a merely inconvenient system that sits in the path of every other team's work, even if the isolated tool generates more complaints. Blast radius, not volume of complaint, is what determines actual cost.
Run the audit in these four passes:
- In the first ten minutes, pull the last quarter of incidents and hotfixes and tag each one by the system it touched.
- Next, review pull request times and comment counts by area of the codebase to spot code that fights its reviewers.
- Then ask the most recent hire where they got stuck, since new engineers hit fragile parts first.
- In the final minutes, rank what you found by how many systems or teams depend on the affected code.
What to do with the ranked list
Take the top one or two items and turn them into a scoped remediation plan with a defined outcome, not an open-ended "clean this up" ticket that will lose priority against every feature request that follows it. Debt remediation that competes directly against feature work without a specific scope and owner tends to lose that fight indefinitely, no matter how real the cost is.
A remediation plan is scoped when someone outside the team could tell whether it is finished. For example, instead of a ticket that says clean up the billing service, write that the two modules generating most incidents will be split from the shared code, with the owner, a target date, and the incident count you expect to see fall. Attach the audit evidence so the reason is visible at planning time. A common mistake is bundling several systems into one plan, which makes it easy to defer. Take the top item only, finish it, and let the next audit decide what comes second.
Re-run the audit instead of trusting the first result forever
The systems that show up on this list will shift as the product and the team change, and a debt audit done once a year ago is telling you about a codebase that no longer exists in quite the same shape. Keep the underlying data, incidents tagged by system, review times, onboarding notes, flowing continuously so the next audit takes minutes to run rather than requiring the same half hour of manual digging every time. Share the ranked list with the whole team after each run, including engineers who weren't part of producing it, since seeing the evidence tends to land better than being told the conclusion secondhand.
What Good Looks Like
Good technical debt prioritization ranks candidates using incident history, review friction, and onboarding pain, not the loudest recent complaint, and turns the highest-ranked items into scoped remediation plans with a defined outcome and an owner.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should this audit be repeated?
Quarterly is a reasonable default for most teams, timed to line up with planning cycles so remediation work has a real chance of being scheduled rather than just documented and forgotten until the next audit.
What if the highest-cost debt is in a system nobody wants to touch?
Treat that reluctance as a finding, because a system everyone avoids usually carries high risk and thin documentation. Both are debt in their own right and often explain why it keeps generating incidents. Scope a small remediation with a named owner instead of leaving it alone.
Should technical debt remediation get its own budget?
It helps. Debt work that has to win an ad hoc argument against every new feature request in every planning cycle tends to lose consistently, even when the data clearly shows its cost, simply because a shipped feature is easier to point to than an incident that didn't happen.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
How to Decide Which Technical Debt to Pay Down First
A framework for deciding which technical debt actually deserves engineering time, based on how often it's touched and what it's slowing down.
Auditing Security on Your MCP and Agent Tool Stack
A step-by-step way for a CTO to audit which tools an AI agent can reach, what each one can do, and where the access is broader than it should be.
A 30-Minute Audit for Finding Technical Debt That's Actually Costing You
A focused 30-minute audit for CTOs to find the technical debt that's actually slowing the team down, and the pitfalls that waste remediation effort.