SLOs for a Small Engineering Team: A Starter Worksheet
A service level objective (SLO) is a target for how reliably a user-facing service should work, measured by a service level indicator (SLI) such as the share of requests that succeed. A small team needs only one or two SLOs, an error budget and a rule for what happens when the budget runs out.
Start with the journey your customers would complain about most, not with every service you run. The worksheet below walks through it in order.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you pick what to measure first?
Choose one or two user journeys where failure is obvious to customers, such as signing in, loading the main dashboard, or completing checkout. Avoid starting with internal components like a queue or a cache. Customers experience journeys, and internal metrics only matter because they affect journeys.
For each journey, write one sentence in plain language: "A customer can complete checkout within three seconds." Then choose the SLI that measures it, ideally from a place close to the customer such as your load balancer or gateway logs, not from inside the service that might be broken.
The SLO worksheet: fill in these fields
Copy this outline into a document and complete it for each journey:
- Service and journey: what the customer is trying to do.
- SLI, availability: the count of good requests divided by valid requests. Decide what counts as good, for example a non-5xx response.
- SLI, latency: the share of requests faster than a threshold, such as the 95th percentile under a set number of milliseconds.
- Measurement source: where the data comes from and who owns the query.
- Target and window: for example, a percentage over a rolling 28 or 30 days.
- Error budget: the allowed share of bad events, meaning whatever the target leaves over.
- Alerting rule: notify when the budget is burning faster than sustainable, not on every blip.
- Owner and review date: one named person and a quarterly check.
Keep exclusions explicit, such as planned maintenance, and don't count health checks or bot traffic as requests.
What target should you set?
Don't pick a number that sounds impressive. Every extra nine costs more than the last, and users of a typical web app can rarely tell adjacent high targets apart. The arithmetic is worth having in front of you. Say you set a 99.9% target: it allows about 8.76 hours of downtime per year. Say you tighten it to 99.99%: that leaves about 52 minutes.
Ask two questions. First, what does your current system actually deliver? Measure for a month before setting a target. Second, what will your customers tolerate and what have you promised in contracts? A target slightly better than today's real performance, and no stricter than your customer commitments, is a good start.
How do you use the error budget?
The error budget turns reliability into a shared decision instead of a debate. Say your SLO is 99.9% over 30 days: you can afford about 43 minutes of failure in that window. Write a short policy, agreed by engineering and product, that says what happens in each state:
- Plenty of budget left: ship normally and take reasonable risks.
- Budget burning quickly: slow risky releases, and prioritize the fix for whatever is burning it.
- Budget exhausted: pause non-essential feature releases until reliability work brings it back, with an explicit way to override for urgent business needs.
The policy only works if leadership honors it once. Otherwise the SLO becomes a dashboard nobody acts on.
How do you avoid the common traps?
Watch for these:
- Too many SLOs. Five well-owned objectives beat thirty ignored ones.
- Measuring server-side success while users see failures, such as a CDN or DNS problem that never reaches your logs. Add a synthetic check from outside.
- Alerting on the SLO threshold itself. Alert on burn rate, so you're paged early enough to act.
- Setting targets from ambition and never reviewing them. Revisit quarterly with real data.
- Forgetting dependencies. If your provider only promises a lower availability, your own target can't exceed it without redundancy.
Your monitoring platform, such as Datadog or New Relic, can compute SLIs and burn rate alerts, but a spreadsheet and a query are enough to begin. When you're ready, connect the SLOs to your on-call rotation so someone knows what to do when the budget starts to burn, and track delivery health alongside them with DORA metrics.
What Good Looks Like
Each critical user journey has a measured SLI, a target based on real performance, an error budget with an agreed policy and a named owner.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
What is the difference between an SLI, SLO and SLA?
An SLI is a measurement, an SLO is your internal target for it, and an SLA is a contractual promise, often with financial consequences. Set your SLO stricter than any SLA so you have a margin.
How many SLOs does a small team need?
One or two for the most important customer journeys. Add more only when you have owners and alerts for the first ones. A short list you actually use beats a long one nobody acts on.
What is an error budget?
It's the amount of unreliability your SLO allows, which is whatever the target leaves over. For example, a 99.9% target over 30 days leaves a budget of roughly 43 minutes of failure. Teams use it to decide how much release risk to take.
Should we aim for four nines of availability?
Only if customers need it and you can afford the engineering. For example, 99.99% allows roughly 52 minutes of downtime a year. Measure your current performance first, then choose the lowest target that meets real customer needs.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
On-Call Rotation for a Small Team: A Worked Schedule
Build a fair on-call rotation for a team of four to six: primary and secondary roles, handoffs, swap rules, time off after pages and escalation.
DORA Metrics for a Small Engineering Team, Without the Dashboard Sprawl
How a team of five to fifteen engineers can track the four DORA metrics, pull the data from tools you already use and avoid the common misreadings.
Moving to the Cloud: A Migration Plan for a Small Business
A phased plan for a small business cloud migration: inventory, choose a strategy per system, build a landing zone, pilot, cut over and retire the old.
Cloud Security Posture Management (CSPM) for a Small Team
What CSPM is, what it catches, how it differs from other cloud security tools and how a small team can adopt it without drowning in alerts.
Do You Need an Internal Developer Portal Under 50 Engineers?
Most teams under 50 engineers can wait on a developer portal. See the signs you're ready, cheaper alternatives and how to start small.
Information Security Policy for a Small Business: Outline and Examples
Write a short information security policy set for a small business: which policies you need, a section-by-section outline and example requirements.