Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

A 30-Minute Check for Blind Spots in Your Observability Setup

Most observability gaps aren't discovered by a careful review, they're discovered at 2 a.m. during an incident when someone realizes there's no metric, log, or trace for the exact thing that's failing. You can find most of these gaps ahead of time with a short, structured check instead of waiting for the incident to find them for you.

This is a 30-minute walkthrough you can run this week. It won't catch everything, but it reliably surfaces the most common blind spots: missing alerts, telemetry that exists but nobody's watching, and dashboards that were built for a system that's since changed shape.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Minutes 0 to 10: how do you spot silent alerts in your history?

Open your alerting tool and filter to the last 90 days. Sort by which alerts have fired the most, and separately, which have never fired at all. An alert that's never fired is either genuinely tracking something that's never gone wrong, or it's misconfigured and silently useless; you often can't tell which without checking the underlying query directly.

Cross-reference against your incident history: for each incident in the last quarter, was there an alert that should have caught it earlier? If the answer is no for more than one or two incidents, that's your biggest gap, and it's worth fixing before anything else on this list.

Minutes 10 to 20: how do you trace one request end to end?

Pick a critical user-facing flow, like checkout or login, and trace a single request through every service it touches. At each hop, ask: is there a metric for latency and error rate here, and would someone get paged if this hop started failing. It's common to find that the entry point and the database are well instrumented but a middle service, often one added later or owned by a different team, has no visibility at all.

Pay particular attention to anything asynchronous: a queue consumer, a background job, a webhook handler. These tend to get instrumented last because they don't sit directly in the request-response path a developer is staring at while building the feature, but a stuck queue consumer can silently back up for hours before anyone notices, precisely because nobody's watching it the way they'd watch a slow API response.

For example, imagine a checkout flow that passes through an API, a payment service, and a queue consumer that sends receipts. During the trace you find latency and error metrics on the first two but nothing on the consumer. The fix is small: add a queue-depth metric and an alert on messages waiting longer than a set age, then confirm the on-call rotation actually receives it. A useful decision rule: if a hop can fail silently for an hour without anyone being paged, instrument it first, ahead of polishing dashboards that already work.

Minutes 20 to 25: check who actually owns each dashboard

A dashboard with no clear owner tends to drift out of date the moment the system it monitors changes shape, a service gets split in two, a queue gets renamed, and nobody updates the panel. Skim your top five dashboards and confirm each one has a name attached, not just a team. If a panel shows a metric that no longer exists or a service that's been decommissioned, that's a signal the whole dashboard hasn't been reviewed in a while.

The same check applies to who actually looks at each dashboard, not just who's nominally responsible for it. A beautifully built panel that nobody opens outside of an incident retrospective isn't really observability, it's decoration. If you can't name the last time someone checked a given dashboard outside of an incident, fold its most important metric into an alert instead of trusting a human to remember to look.

Minutes 25 to 30: confirm downtime budget and alert thresholds actually line up

Your availability target implies a specific downtime budget: a service promising 99.9% uptime can absorb roughly 8.76 hours of downtime across a year, while 99.99% only allows around 52.6 minutes1. Check whether your alert thresholds are actually tuned to that budget. It's common to find alerts set to fire on any error at all, which trains the on-call rotation to ignore them, when the real target should be an error-budget-burn-rate alert tied to how fast you're eating into that annual allowance.

If you can't say off the top of your head how much of this year's downtime budget you've already used, that's itself a finding worth writing down before you move on.

Keep this checklist handy for your next run:

  1. Minutes 0 to 10: filter 90 days of alert history, compare alerts that fire most with those that never fire, and match recent incidents to alerts that should have caught them.
  2. Minutes 10 to 20: trace one critical flow, such as checkout or login, through every service and confirm each hop has latency and error metrics plus a page.
  3. Minutes 20 to 25: confirm each top dashboard has a named owner, and that someone actually looks at it outside of incidents.
  4. Minutes 25 to 30: compare alert thresholds to your downtime budget, and move noisy any-error alerts to error-budget burn-rate alerts.
Executive Capability Standard

What Good Looks Like

Every service on a critical user-facing path has an owned alert tied to its downtime budget, and every dashboard has a named owner who's reviewed it in the last quarter.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Run the 30-minute check above on your most critical user-facing flow and write down every gap you find.
2. Do Manually:Manually trace two more critical flows end to end and add the missing metrics and alerts you find.
3. Delegate:Assign a named owner to every dashboard and alert group, not just a team, so drift gets caught earlier.
4. Automate:Set up automated alert-effectiveness reporting that flags alerts which have never fired or that fire without any follow-up action.
5. Buy:Bring in an observability or SRE consultant for a deeper audit if the 30-minute check surfaces gaps across most of your critical paths.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

What if we find more gaps than we can fix in one sprint?

Prioritize by blast radius, not by how easy the fix is. A missing alert on a payment flow matters more than a missing alert on an internal admin tool, even if the admin tool fix is quicker. Fix the highest-impact gap first and schedule the rest.

How often should we repeat this 30-minute check?

Quarterly is a reasonable cadence for most teams, and immediately after any significant architecture change, like splitting a service or adding a new critical dependency, since that's exactly when new blind spots get introduced.

Should every service have the same level of observability?

No. Tier your services by how much damage a silent failure would cause, and invest instrumentation effort proportionally. A payment or auth service deserves deep tracing; an internal reporting job that runs nightly probably doesn't need the same level of detail.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides