Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Why Your SLA Alerts Keep Missing Real Breaches

Polling-based SLA checks miss real breaches because failures at scale are shorter than the polling interval, so the reliable fix is to compute SLA metrics from the event stream itself. A burst of failures that starts and recovers between two checks never appears on a dashboard that only samples on a timer.

The fix is not a bigger dashboard. It is moving SLA detection into the event stream itself, so a breach is caught by something watching continuously rather than something that only looks up periodically.

Understand why polling breaks down first, not last

Polling checks give you a health snapshot at fixed intervals, which is fine when incidents are rare and long. As a pipeline scales, failures get shorter and more frequent, a retry that resolves itself in ten seconds, a partition that briefly falls behind, and a check that runs every sixty seconds simply misses most of them by construction, not by bad luck.

This is usually invisible until someone manually reconstructs an incident timeline from raw logs and finds the SLA dashboard reported green the entire time. That gap is the tell that polling has become the wrong tool for the current scale, not a sign the team is monitoring poorly.

How do you detect SLA breaches from the event stream itself?

Instead of polling an endpoint, compute SLA metrics, like consumer lag, processing latency, and error rate, directly from the event stream as messages flow through, using a stream processor or a metrics pipeline that updates continuously. This turns SLA monitoring from a sampled snapshot into something closer to a real-time count.

This is a real engineering project, not a config toggle, so scope it as one: define the exact metrics that constitute your SLA, build the continuous computation for them, and only then wire up alerting on top.

Should SLA breach thresholds use sustained conditions or single data points?

A single slow message is not a breach. A sustained trend usually is. Define your alerting logic around a window, such as lag exceeding a threshold for a defined number of consecutive minutes, rather than firing on the first data point that crosses a line, which mostly produces noise your team learns to ignore.

Getting this window right takes a couple of iterations. Start conservative, watch what actually pages your team for a few weeks, and tighten or loosen the window based on how many of those pages turned out to be real.

Separate the alert from the automated response

Detecting a breach quickly is only half the job. Decide, for each SLA, what should happen automatically when it fires, such as scaling out consumers or failing over to a backup path, versus what should only page a human. Conflating these two things means either your team gets paged for events that resolve themselves, or a real breach waits for a human to notice and act.

Document this split explicitly per SLA rather than leaving it as an implicit assumption in whoever built the alerting rule, since the next engineer to touch it needs to know which failures are supposed to self-heal.

Questions to settle for each SLA before it goes live:

  • Which metric defines the SLA, such as consumer lag, processing latency, or error rate, and where is that definition written down?
  • How many consecutive minutes above the threshold count as a sustained breach rather than noise?
  • Which breaches trigger an automatic action, such as scaling out consumers or failing over to a backup path?
  • Which breaches only page a human, and who is the named responder for each one?
  • How will you track near-misses that approach the threshold without ever crossing it?

Report on near-misses, not only actual breaches

A pipeline that stays just under its SLA threshold repeatedly is telling you something a binary breach or no-breach view will never show. Track how often and how closely you approach the threshold, not only whether you cross it, so a slow drift toward trouble shows up before it becomes an actual incident.

Review this trend on the same cadence as your capacity planning, since a rising rate of near-misses is often the earliest honest signal that a topic is running out of headroom well before anyone declares an incident.

Plot near-misses over time rather than looking at a single week in isolation. A team that only checks the current dashboard misses the difference between a one-off blip caused by an unrelated deploy and a steady climb that means the underlying capacity has been shrinking relative to demand for months.

Keep the SLA definition itself under change control

The metrics behind an SLA tend to get redefined quietly over time, a threshold gets loosened after a noisy week, a metric gets swapped for one that is easier to compute, and eventually nobody can say with confidence what the SLA actually measures today versus what it measured a year ago. Treat a change to the SLA definition itself as a decision that needs sign-off and a changelog entry, not a config edit an individual engineer makes to quiet an alert.

This matters most when a metric moves in a way that makes the SLA easier to hit rather than harder. That kind of silent loosening is exactly what an executive or a customer relying on the SLA number would want to know about, and it is much cheaper to catch with a changelog than to explain after the fact.

Executive Capability Standard

What Good Looks Like

Reliable SLA enforcement computes its metrics continuously from the event stream rather than from periodic polling, alerts on sustained conditions instead of single data points, has an explicit split between automated response and human paging, and tracks near-misses alongside actual breaches.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Reconstruct a past incident timeline from raw logs and compare it against what your current SLA dashboard reported during that window.
2. Do Manually:Define, in writing, the exact window and threshold that should count as a sustained breach for your top SLA metric.
3. Delegate:Assign an engineer to build continuous, stream-based computation for your highest-priority SLA metrics.
4. Automate:Wire automated responses, like consumer autoscaling, to the SLAs where a fast self-healing action makes sense.
5. Buy:Bring in outside infrastructure help if your team has never built stream-based monitoring and is still relying entirely on periodic polling.

How to Get Started

Frequently Asked Questions

How often should a polling-based SLA check run before it becomes unreliable?

There is no fixed interval that stays safe forever, since it depends on how short your incidents typically are relative to the polling cycle. Once you notice incidents resolving between checks, that is the signal to move toward continuous, stream-based detection rather than just shortening the interval further.

Should every SLA breach trigger an automatic response?

No. Decide per SLA whether the right response is an automated action, like scaling out consumers, or a page to a human. Automating a response that should have human judgment behind it, or paging a human for something that should self-heal, both waste attention in different ways.

What is a near-miss and why does it matter for SLA monitoring?

A near-miss is when a metric like consumer lag approaches your SLA threshold without technically crossing it. Tracking these separately from actual breaches surfaces a slow drift toward trouble earlier, often before it shows up as a real incident anyone would notice.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides