AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

What to Actually Alert On When You Serve Models in Production

Most observability setups for model serving start as a copy of whatever dashboard the team already had for a normal API: request rate, error rate, latency percentiles. Those still matter, but none of them tell you when a model is technically healthy and quietly wrong.

Getting observability and alerting right means adding a second layer of signals on top of the usual ones, tuned to catch quality drift, not just downtime.

The signals a standard API dashboard misses

  • Output length distribution. A sudden shift toward much shorter or much longer responses often shows up before users complain.
  • Refusal or fallback rate. How often the model declines, times out, or falls back to a default response, tracked as a trend, not just a raw count.
  • Token-level cost per request. A spike here without a matching traffic spike usually means something changed in your prompt template or the model's behavior.
  • GPU memory pressure and queue depth, which predict a latency problem before it shows up in your latency graphs.

None of these need a research team to build; they're mostly a few extra fields logged on every request.

Setting alert thresholds without an error budget

An error budget gives you a number to alert against instead of a gut feeling. If your target is 99.9 percent availability, you're working with roughly 8.76 hours of allowed downtime for the whole year1; at 99.95 percent, that budget shrinks to about 4.38 hours.

Once you know your annual budget, break it into a monthly or weekly allowance and alert when you're burning it faster than a steady pace would allow, not just when you hit zero. That catches a slow leak, a service that's degraded but not fully down, before it consumes the whole year's budget in a bad month.

A basic alerting setup that catches quality drift

Start with three tiers. Page immediately on infrastructure failures: the endpoint is down, GPU nodes are unreachable, or queue depth is climbing without bound. Alert during business hours on quality signals: refusal rate above its normal range, output length distribution shifting, or token cost per request rising without a traffic increase to explain it. Log everything else for weekly review: small drifts that aren't urgent but are worth a human look before they compound.

The mistake most teams make is putting all of this in the first tier, which trains everyone to ignore alerts, or leaving it all in the third tier, which means quality drift gets caught a week after users noticed.

Tracing a bad response back to its cause

When something goes wrong, you want to answer one question fast: was this the model, the prompt, or something upstream. That requires tagging every logged request with the model version, prompt template version, and any retrieved context or tool results involved, not just the final output.

Without that tagging, a support ticket about a bad answer turns into an afternoon of guessing. With it, you can filter for every request that used the same model and prompt version and see whether the issue is isolated or systemic within minutes.

What not to bother monitoring

  • Raw GPU temperature, unless you're managing your own hardware; your cloud provider already handles this layer.
  • Every individual prompt and response in a real-time dashboard; sample instead, and log the rest for offline review.
  • Model confidence scores in isolation, when the model doesn't have a well-calibrated confidence signal to begin with; a raw score without calibration is noise dressed up as a metric.

Monitoring everything is the same mistake as monitoring nothing: both leave you unable to tell a real signal from background noise.

A common alerting mistake: thresholds tuned once and forgotten

Traffic patterns and model behavior both drift over time, and a threshold that made sense at launch can become useless within a quarter. A refusal rate alert tuned for a slow initial rollout will stay silent through a real problem once traffic grows, because the raw counts look small against the new baseline even though the underlying rate is unhealthy.

Revisit your alert thresholds on a fixed schedule, not only after an incident forces the review. A short monthly check against the previous month's actual data catches drift in the thresholds themselves before it becomes a blind spot in production.

Executive Capability Standard

What Good Looks Like

Solid observability for model serving means you can answer, within minutes, whether a bad response came from the model, the prompt, or something upstream, and your alerts are tuned to catch quality drift, not just downtime.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Log model version, prompt template version, and output length on every request for one endpoint, even before you build dashboards on top of it.
2. Do Manually:Pull a week of logs and manually chart refusal rate and output length by hand to see whether either one is drifting.
3. Delegate:Assign an engineer to own the alerting thresholds for one endpoint and review them monthly against actual incidents.
4. Automate:Build dashboards and alerts for refusal rate, output length distribution, and token cost per request, tiered by urgency.
5. Buy:Bring in observability tooling or consulting once you're running enough endpoints that manual log review can't keep up.

How to Get Started

Frequently Asked Questions

What's the single most useful alert to set up first for a model-serving endpoint?

Refusal or fallback rate, tracked as a trend rather than a raw count. It catches a model quietly failing more often, a prompt template that stopped working after an upstream change, or a provider outage, all in one signal, usually before customers report anything and often before your latency or error rate graphs show a thing.

How do we set a sensible error budget for a model-serving endpoint?

Pick an availability target that matches how critical the endpoint actually is, then work out its downtime budget for the year and split it into a weekly allowance. Alert when you're burning that allowance faster than a steady pace would, not only when you hit zero. Most customer-facing endpoints do fine around three nines rather than reaching for five.

Why would output length or refusal rate matter more than latency for catching problems?

Because a model can be fast and still be wrong. Latency tells you the system is responding; it says nothing about whether the response is good. Output length shifts and refusal rate spikes tend to show up before a quality problem is visible in error rates or user complaints, which is exactly when you want to catch it.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides