A 30-Minute Check for Blind Spots in Your Observability Setup
Most observability gaps aren't discovered by a careful review, they're discovered at 2 a.m. during an incident when someone realizes there's no metric, log, or trace for the exact thing that's failing. You can find most of these gaps ahead of time with a short, structured check instead of waiting for the incident to find them for you.
This is a 30-minute walkthrough you can run this week. It won't catch everything, but it reliably surfaces the most common blind spots: missing alerts, telemetry that exists but nobody's watching, and dashboards that were built for a system that's since changed shape.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Minutes 0 to 10: how do you spot silent alerts in your history?
Open your alerting tool and filter to the last 90 days. Sort by which alerts have fired the most, and separately, which have never fired at all. An alert that's never fired is either genuinely tracking something that's never gone wrong, or it's misconfigured and silently useless; you often can't tell which without checking the underlying query directly.
Cross-reference against your incident history: for each incident in the last quarter, was there an alert that should have caught it earlier? If the answer is no for more than one or two incidents, that's your biggest gap, and it's worth fixing before anything else on this list.
Minutes 10 to 20: how do you trace one request end to end?
Pick a critical user-facing flow, like checkout or login, and trace a single request through every service it touches. At each hop, ask: is there a metric for latency and error rate here, and would someone get paged if this hop started failing. It's common to find that the entry point and the database are well instrumented but a middle service, often one added later or owned by a different team, has no visibility at all.
Pay particular attention to anything asynchronous: a queue consumer, a background job, a webhook handler. These tend to get instrumented last because they don't sit directly in the request-response path a developer is staring at while building the feature, but a stuck queue consumer can silently back up for hours before anyone notices, precisely because nobody's watching it the way they'd watch a slow API response.
For example, imagine a checkout flow that passes through an API, a payment service, and a queue consumer that sends receipts. During the trace you find latency and error metrics on the first two but nothing on the consumer. The fix is small: add a queue-depth metric and an alert on messages waiting longer than a set age, then confirm the on-call rotation actually receives it. A useful decision rule: if a hop can fail silently for an hour without anyone being paged, instrument it first, ahead of polishing dashboards that already work.
Minutes 20 to 25: check who actually owns each dashboard
A dashboard with no clear owner tends to drift out of date the moment the system it monitors changes shape, a service gets split in two, a queue gets renamed, and nobody updates the panel. Skim your top five dashboards and confirm each one has a name attached, not just a team. If a panel shows a metric that no longer exists or a service that's been decommissioned, that's a signal the whole dashboard hasn't been reviewed in a while.
The same check applies to who actually looks at each dashboard, not just who's nominally responsible for it. A beautifully built panel that nobody opens outside of an incident retrospective isn't really observability, it's decoration. If you can't name the last time someone checked a given dashboard outside of an incident, fold its most important metric into an alert instead of trusting a human to remember to look.
Minutes 25 to 30: confirm downtime budget and alert thresholds actually line up
Your availability target implies a specific downtime budget: a service promising 99.9% uptime can absorb roughly 8.76 hours of downtime across a year, while 99.99% only allows around 52.6 minutes1. Check whether your alert thresholds are actually tuned to that budget. It's common to find alerts set to fire on any error at all, which trains the on-call rotation to ignore them, when the real target should be an error-budget-burn-rate alert tied to how fast you're eating into that annual allowance.
If you can't say off the top of your head how much of this year's downtime budget you've already used, that's itself a finding worth writing down before you move on.
Keep this checklist handy for your next run:
- Minutes 0 to 10: filter 90 days of alert history, compare alerts that fire most with those that never fire, and match recent incidents to alerts that should have caught them.
- Minutes 10 to 20: trace one critical flow, such as checkout or login, through every service and confirm each hop has latency and error metrics plus a page.
- Minutes 20 to 25: confirm each top dashboard has a named owner, and that someone actually looks at it outside of incidents.
- Minutes 25 to 30: compare alert thresholds to your downtime budget, and move noisy any-error alerts to error-budget burn-rate alerts.
What Good Looks Like
Every service on a critical user-facing path has an owned alert tied to its downtime budget, and every dashboard has a named owner who's reviewed it in the last quarter.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Tenable's asset and vulnerability telemetry is worth feeding into the same observability pipeline you use for application metrics, so a security gap shows up in the same place your team already looks.
CrowdStrike's endpoint telemetry belongs in this check too: confirm its alerts are actually routed somewhere your on-call rotation watches, not siloed in a separate console nobody opens.
Frequently Asked Questions
What if we find more gaps than we can fix in one sprint?
Prioritize by blast radius, not by how easy the fix is. A missing alert on a payment flow matters more than a missing alert on an internal admin tool, even if the admin tool fix is quicker. Fix the highest-impact gap first and schedule the rest.
How often should we repeat this 30-minute check?
Quarterly is a reasonable cadence for most teams, and immediately after any significant architecture change, like splitting a service or adding a new critical dependency, since that's exactly when new blind spots get introduced.
Should every service have the same level of observability?
No. Tier your services by how much damage a silent failure would cause, and invest instrumentation effort proportionally. A payment or auth service deserves deep tracing; an internal reporting job that runs nightly probably doesn't need the same level of detail.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
What to Actually Monitor Before You Buy an Observability Tool
Answers to the questions engineering teams actually ask before setting up monitoring: what to track, how many alerts is too many, and when to add tracing.
A Practical Checklist for Observability That Gets Used
A short checklist for building observability that people actually rely on during an incident, instead of dashboards nobody opens and alerts nobody trusts.
What Your Alerts Are Actually Telling You
A practical walkthrough for auditing an observability setup: which alerts you can trust, which ones get ignored, and what telemetry gap to close first.
The Signals That Tell You a RAG Pipeline Is Degrading
Uptime dashboards miss RAG failure modes. Here are the retrieval, drift, and groundedness signals worth instrumenting before quality quietly drops.
Watching What Your Agents Actually Do in Production
Answers to the observability questions a CTO actually has about agentic systems: what to log, what to alert on, and what a normal trace looks like.
What to Actually Alert On in a Zero Trust API Setup
A worksheet for building an alerting matrix for zero trust APIs that catches real problems without burying your team in noise.