ObservabilityTemplate3 min readUpdated September 2026

SLOs for a Small Engineering Team: A Starter Worksheet

A service level objective (SLO) is a target for how reliably a user-facing service should work, measured by a service level indicator (SLI) such as the share of requests that succeed. A small team needs only one or two SLOs, an error budget and a rule for what happens when the budget runs out.

Start with the journey your customers would complain about most, not with every service you run. The worksheet below walks through it in order.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you pick what to measure first?

Choose one or two user journeys where failure is obvious to customers, such as signing in, loading the main dashboard, or completing checkout. Avoid starting with internal components like a queue or a cache. Customers experience journeys, and internal metrics only matter because they affect journeys.

For each journey, write one sentence in plain language: "A customer can complete checkout within three seconds." Then choose the SLI that measures it, ideally from a place close to the customer such as your load balancer or gateway logs, not from inside the service that might be broken.

The SLO worksheet: fill in these fields

Copy this outline into a document and complete it for each journey:

  1. Service and journey: what the customer is trying to do.
  2. SLI, availability: the count of good requests divided by valid requests. Decide what counts as good, for example a non-5xx response.
  3. SLI, latency: the share of requests faster than a threshold, such as the 95th percentile under a set number of milliseconds.
  4. Measurement source: where the data comes from and who owns the query.
  5. Target and window: for example, a percentage over a rolling 28 or 30 days.
  6. Error budget: the allowed share of bad events, meaning whatever the target leaves over.
  7. Alerting rule: notify when the budget is burning faster than sustainable, not on every blip.
  8. Owner and review date: one named person and a quarterly check.

Keep exclusions explicit, such as planned maintenance, and don't count health checks or bot traffic as requests.

What target should you set?

Don't pick a number that sounds impressive. Every extra nine costs more than the last, and users of a typical web app can rarely tell adjacent high targets apart. The arithmetic is worth having in front of you. Say you set a 99.9% target: it allows about 8.76 hours of downtime per year. Say you tighten it to 99.99%: that leaves about 52 minutes.

Ask two questions. First, what does your current system actually deliver? Measure for a month before setting a target. Second, what will your customers tolerate and what have you promised in contracts? A target slightly better than today's real performance, and no stricter than your customer commitments, is a good start.

How do you use the error budget?

The error budget turns reliability into a shared decision instead of a debate. Say your SLO is 99.9% over 30 days: you can afford about 43 minutes of failure in that window. Write a short policy, agreed by engineering and product, that says what happens in each state:

  • Plenty of budget left: ship normally and take reasonable risks.
  • Budget burning quickly: slow risky releases, and prioritize the fix for whatever is burning it.
  • Budget exhausted: pause non-essential feature releases until reliability work brings it back, with an explicit way to override for urgent business needs.

The policy only works if leadership honors it once. Otherwise the SLO becomes a dashboard nobody acts on.

How do you avoid the common traps?

Watch for these:

  • Too many SLOs. Five well-owned objectives beat thirty ignored ones.
  • Measuring server-side success while users see failures, such as a CDN or DNS problem that never reaches your logs. Add a synthetic check from outside.
  • Alerting on the SLO threshold itself. Alert on burn rate, so you're paged early enough to act.
  • Setting targets from ambition and never reviewing them. Revisit quarterly with real data.
  • Forgetting dependencies. If your provider only promises a lower availability, your own target can't exceed it without redundancy.

Your monitoring platform, such as Datadog or New Relic, can compute SLIs and burn rate alerts, but a spreadsheet and a query are enough to begin. When you're ready, connect the SLOs to your on-call rotation so someone knows what to do when the budget starts to burn, and track delivery health alongside them with DORA metrics.

Executive Capability Standard

What Good Looks Like

Each critical user journey has a measured SLI, a target based on real performance, an error budget with an agreed policy and a named owner.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand SLIs, SLOs, error budgets and how a target converts to allowed downtime.
2. Do Manually:Pick one journey, pull a month of request data by hand and calculate the current availability.
3. Delegate:Assign a named owner for each SLO and schedule a quarterly review.
4. Automate:Compute SLIs and burn-rate alerts in your monitoring tool and show remaining budget on a dashboard.
5. Buy:Use your monitoring platform's SLO features once spreadsheet calculations become hard to maintain.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Datadog

Fits when your request metrics already live there and you want SLO and burn-rate alerts computed on the same data.

Visit Datadog→
New Relic

Fits when your team already uses it for application monitoring and wants SLO tracking in the same tool.

Visit New Relic→

Frequently Asked Questions

What is the difference between an SLI, SLO and SLA?

An SLI is a measurement, an SLO is your internal target for it, and an SLA is a contractual promise, often with financial consequences. Set your SLO stricter than any SLA so you have a margin.

How many SLOs does a small team need?

One or two for the most important customer journeys. Add more only when you have owners and alerts for the first ones. A short list you actually use beats a long one nobody acts on.

What is an error budget?

It's the amount of unreliability your SLO allows, which is whatever the target leaves over. For example, a 99.9% target over 30 days leaves a budget of roughly 43 minutes of failure. Teams use it to decide how much release risk to take.

Should we aim for four nines of availability?

Only if customers need it and you can afford the engineering. For example, 99.99% allows roughly 52 minutes of downtime a year. Measure your current performance first, then choose the lowest target that meets real customer needs.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides