How Much Redundancy Your Model-Serving Stack Actually Needs
High availability for model serving gets sold as an all-or-nothing decision: either you run active-active across two regions, or you're one outage away from a bad week. In practice, most companies need something in between, and the right amount depends on what happens to your product when inference goes down, not on what a vendor's reference architecture recommends.
Start your high availability and disaster recovery planning by defining that failure mode first, not by picking an architecture off a diagram.
What actually fails, in order of likelihood
- A single GPU node dies or gets preempted, the most common failure and the easiest to handle with a few replicas behind a load balancer.
- A model provider's API has an outage, if you depend on a third-party model at all, which is a failure entirely outside your infrastructure.
- A whole region goes down, rare, expensive to protect against fully, and usually not worth solving before the first two.
- A bad model deploy takes down quality, not availability, which redundancy doesn't help with at all; that's a rollback problem, not a failover one.
Most teams over-invest in the third case and under-invest in the first.
Active-active, active-passive, or just more replicas
For the common case, a single node dying, you don't need multi-region anything; you need enough replicas in one region that losing one doesn't drop your capacity below demand, plus health checks that pull a failing node out of rotation automatically.
Active-passive across two regions adds protection against a regional outage, with a standby that can take traffic within minutes, at the cost of running, or quickly provisioning, a second set of GPUs you mostly don't use. Active-active removes that failover delay entirely, serving real traffic from both regions all the time, but it roughly doubles your steady-state GPU spend and adds real complexity to keeping model versions in sync across regions.
Pick active-passive unless your product genuinely can't tolerate a few minutes of failover time.
Handling a third-party model provider outage
If part of your stack depends on an external model API, your own redundancy doesn't protect you from their outage. The realistic options are a fallback to a different provider or a smaller self-hosted model for degraded service, or accepting the outage and communicating it clearly.
A fallback model is worth building when the endpoint is core to your product, even if the fallback is noticeably worse; a degraded answer usually beats no answer. It's not worth building for a feature nobody would notice missing for an hour. Decide which category you're in before an outage forces the decision under pressure.
Write the decision down before you need it: which endpoints get an automatic fallback, and which get a manual switch with a status page update. Deciding this during an actual outage, under pressure and with customers already noticing, produces worse choices than deciding it in a calm planning session.
What your uptime target actually costs you in downtime
Nines get thrown around casually, but the gap between them is stark once you translate it into a real number. Three nines, 99.9 percent availability, allows roughly 8.76 hours of downtime a year1. Four nines cuts that down to about 52.6 minutes, which is a different engineering problem entirely.
Going from three nines to four is usually multi-region active-active plus a lot of operational discipline, not a config change. Most model-serving products for founders and mid-sized companies are well served by three nines; reserve the push for four only when a customer contract genuinely requires it.
Redundancy decisions that aren't worth the cost yet
- Multi-region active-active before you have multi-region customers or a contract that requires it.
- A fully automated failover for a third-party provider outage that happens a few times a year and can be handled with a status page and a manual switch.
- Redundant everything for an internal tool that a handful of employees use, where a short outage costs almost nothing.
Spend the redundancy budget on the failure mode that actually happens to you, not the one that sounds most impressive in a design doc.
What Good Looks Like
The right amount of redundancy means you've named which failure mode actually threatens your product, sized your replica count and failover plan to that failure mode, and can state in plain terms how many hours of downtime your uptime target allows.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need multi-region redundancy for our model-serving stack?
Probably not yet, unless a regional outage would be catastrophic for your specific product or a customer contract requires it. Most availability problems come from a single node failing or a third-party provider outage, neither of which multi-region solves by itself. Fix replica counts and provider fallbacks first; multi-region is usually the last lever, not the first.
What's the difference between active-passive and active-active failover?
Active-passive keeps a standby region ready but idle, taking over within minutes of a failure. Active-active serves real traffic from both regions all the time, removing the failover delay but roughly doubling steady-state GPU cost and adding complexity around keeping model versions in sync. Most teams should start with active-passive.
How much downtime does a three-nines availability target actually allow?
About 8.76 hours a year, or a little over 43 minutes a month if you spread it evenly. That's a reasonable target for most model-serving products aimed at founders and mid-sized companies; pushing higher is a real engineering investment worth reserving for cases where a contract specifically requires it.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
What SOC 2 Actually Expects From a Model-Serving Team
What SOC 2 expects from a team serving AI models: how change, access, patch, and vendor controls apply, and the evidence to have ready.
What a Real Security Audit of Model Serving Should Cover
A practical checklist for auditing AI model serving and inference: endpoint access, weight security, prompt logging, and patch timelines.
Designing Fallback Logic That Doesn't Make Things Worse
How to design retry, fallback, and fail-visibly logic for AI model serving without causing a retry storm during an outage.