Status Page Best Practices for When Your Service Is Down
The best status page practices during an outage are simple: host the page away from your own infrastructure, post within minutes of confirming impact, update on a promised schedule, and describe what customers experience, not what your servers are doing.
The rest of this playbook covers how to set the page up before you need it, and what to write once you do.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Where should the status page live?
A status page that goes down with your product is worse than none. Put it on a different provider, a different domain or subdomain that doesn't share your DNS zone's failure modes, and a different login system from production. Test that you can publish from a phone on cellular data with your main network unreachable.
Link to it from places customers look: the app footer, the support widget, your help center and your email signature. Also make sure your in-app error message tells people where to check. If customers can't find the page during an incident, they'll open tickets instead, and support becomes your status channel.
How to design components customers recognize
Components should map to things customers say, not to your architecture. 'Sign-in', 'Dashboard', 'API', 'Email notifications' and 'Payments' are useful; 'ingest-worker-3' is not. A workable process:
- List the five to ten customer-facing capabilities people would notice missing.
- Assign each one an internal owner who can confirm its state.
- Decide what each status means in your terms: degraded means slow or partial, outage means unusable for most customers.
- Separate regions or tenants only if customers actually experience them separately.
- Add third-party dependencies, such as your payment processor, only if you'll keep them updated.
Resist the urge to publish every internal service. Granularity you can't maintain becomes misleading fast.
What should each update say?
Use a fixed sequence of states, commonly investigating, identified, monitoring and resolved, so readers can tell where things stand. Every update should answer four questions: who is affected, what they'll notice, what you're doing, and when the next update is coming.
Compare two versions of the same message. Vague: 'We're experiencing some issues with our platform.' Better: 'Some customers can't log in. Existing sessions still work. We've identified a failing authentication service and are rolling back a change. Next update in 30 minutes.' The second is what a support agent or an anxious customer can act on. Avoid jargon, avoid guessing at causes, and avoid promising a fix time you can't guarantee. If you don't know the estimate, say so and give the time of the next update instead.
Mistakes that erode trust during an outage
Customers forgive outages more easily than evasiveness. Common errors:
- Waiting for a root cause before posting anything, so the first message arrives an hour after customers noticed.
- Using 'degraded performance' for what is really an outage, which customers read as spin.
- Missing your own promised update time. If nothing changed, post 'no change, still investigating' at the time you said.
- Blaming a vendor in the first message. Explain impact first; discuss vendors in the review if relevant.
- Marking incidents resolved too early and then reopening them, which teaches people to distrust green statuses.
How do you close the loop afterward?
Resolve the incident on the page only after you've watched the metrics recover. Then add a short summary within a day or two: what happened in plain terms, how long it lasted, and what you're changing. It doesn't need to be a full postmortem; the internal review belongs to the blameless postmortem template.
Decide in advance who's allowed to publish, and tie the rule to your severity scheme so the page isn't updated by whoever is nearest. The incident response plan outline has a place for the communications role. If you want status updates published from the same place you coordinate the incident, tools like incident.io include that workflow; the comparison in PagerDuty vs incident.io for B2B SaaS covers how that differs from a paging-first product.
What Good Looks Like
Customers can see the current state of each feature they use, hosted separately from production, with updates on a promised schedule.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Should a status page be hosted separately from your product?
Yes. Use a different provider and separate credentials so an outage of your infrastructure doesn't take the page down with it. Confirm you can publish from a phone when your main environment is unreachable.
How often should you update during an outage?
Commit to a schedule in each message, such as every 30 minutes, and keep it even when nothing has changed. A 'no change' post at the promised time reassures customers more than silence.
When should you post the first update?
As soon as you've confirmed customer impact, even without a cause. Say what customers may notice and when the next update will come. Waiting for a diagnosis leaves customers guessing.
Do you need a public status page if you have few customers?
A simple page is still worth having. It gives support a link to send and cuts repeat tickets. If your customers are all enterprise accounts, a private page or email list can also work.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Writing a Blameless Postmortem: Template and Example
A blameless postmortem outline with section-by-section guidance, example rewrites of blaming language, and how to keep action items from stalling.
Incident Response Plan for a Startup: A Fill-In Outline
An incident response plan outline for small engineering teams: roles, the first 15 minutes, communication steps, a security branch and a review process.
incident.io or PagerDuty: Picking On-Call for B2B SaaS
How B2B SaaS teams should weigh incident.io against PagerDuty for on-call paging, Slack-based triage, and postmortems that hold up with SOC 2 auditors.
Incident Response Where an Outage Can Void a Run
In life sciences and biotech consulting, a system outage can invalidate a lab run, not just annoy a user. Compare incident.io and PagerDuty on that basis.
PagerDuty vs Opsgenie vs incident.io: Incident Platforms Compared
Compare PagerDuty, Opsgenie, and incident.io for on-call routing, automated escalation policies, Slack-native triage, and DORA incident recovery.
Defining Incident Severity Levels: A Four-Tier Template
Define SEV1 to SEV4 by customer impact, with response expectations, who can declare and change severity, and mistakes that cause false alarms.