Incident Management & On-Call Operations3 min readUpdated September 2026

On-Call Tools for Agencies Running Client Software

You are on call for an application you built, running in a client's cloud account, alerting into a Slack workspace you do not own. When that pager goes off at two in the morning, the first problem is not diagnosing the bug. It is figuring out which of three tools across three client accounts is actually the one paging you tonight.

That ownership question, more than any feature comparison, decides whether incident.io or PagerDuty fits a development shop better, and it changes depending on whether you are running one long-term retainer or six short engagements at once.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Whose workspace does the incident actually live in

incident.io is built around the assumption that your team runs the response inside its own Slack. If you are responding inside the client's Slack instead, because that is where their stakeholders are, you either need incident.io deployed in their workspace or you are coordinating across two chat systems at once while the outage is live.

PagerDuty's web-based incident record does not care whose Slack anyone uses, which makes it easier to keep one paging setup even when every client's collaboration tooling looks different. For an agency running several client codebases at once, that platform independence can matter more than any single workflow feature either tool offers.

Per-seat pricing gets strange when half your responders are not your staff

Both platforms price primarily by the number of people who can be paged or who touch the incident tooling. For an agency, that number is not stable: it might include your engineers, a client's ops lead who wants visibility, and a subcontractor rotating through for a specific project.

Before signing an annual contract with either vendor, map out how many active seats a typical engagement actually needs, and check whether the vendor's tiering punishes you for adding a client stakeholder as a light viewer rather than a full responder. A contract sized for your busiest quarter can leave you overpaying for the other three.

What happens to the incident history when the engagement ends

When a retainer ends, the client keeps the code, the infrastructure, and usually the monitoring. Whether they keep the incident history depends entirely on whose account it lived in. If you ran incidents through your own incident.io or PagerDuty account, the client is left with no record of past outages unless you export and hand it over deliberately as part of offboarding.

Build that export into your engagement close-out checklist rather than discovering the gap when a client asks for a full history of postmortems after you have already moved your team on to the next project.

An offboarding checklist when an engagement ends:

  • Export incident history and postmortems so the client keeps a record of past outages.
  • Transfer or remove your team's access to the account.
  • Hand over any runbooks embedded in the tool.
  • Confirm the client has an active on-call rotation before your last responder is removed.

Recovery speed still tells you something real

Regardless of whose tool it is, the underlying discipline matters more than the label on the escalation policy. The gap between a team that recovers from a failed deployment in minutes and one that takes weeks usually comes down to whether rollback is a rehearsed, one-command action or something an engineer improvises live while a client watches the incident channel.

For a shop juggling several client codebases, that rehearsal has to happen per client, since a rollback script that works cleanly on one client's infrastructure may not exist at all on another's, especially where a client's own team wrote the original deployment pipeline before you took over support.

A practical default for most engagements

For a single client with a modern stack and a Slack-first culture, incident.io tends to be the faster setup: less configuration, and the client can see exactly what happened without learning a new tool. For a client with an existing PagerDuty contract, multiple legacy integrations, or a preference for a web dashboard over Slack, adopt their existing tool rather than introducing a second one just because your team prefers it.

The right platform for an agency is usually whichever one the client will actually keep using after you leave, not the one your own engineers find most comfortable during the engagement itself.

A worked example: two clients, two tools, one engineer on call

Say one engineer covers overnight support for a client running incident.io and another running PagerDuty, because each adopted its own tool before your agency got involved. That engineer needs two separate paging setups configured correctly on their phone, and a habit of checking which client's system is actually alerting before diagnosing anything, since the fix for one client's stack rarely applies to the other's.

Document which tool belongs to which client somewhere more durable than the engineer's memory, ideally in the same runbook that already lists each client's rollback steps and escalation contacts, so a new team member covering that rotation for the first time is not guessing.

Executive Capability Standard

What Good Looks Like

A well-run development shop can name, for every active client engagement, exactly which tool pages which responder, keeps a rehearsed rollback for that client's stack, and hands over a complete incident history at offboarding without a special request.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List every active client engagement and write down which tool, if any, currently pages your team for that client's production issues.
2. Do Manually:Track on-call coverage per client in a shared calendar and rely on personal phone alerts until a client's incident volume justifies dedicated tooling.
3. Delegate:Assign one engineer per engagement as the incident owner responsible for that client's rollback runbook and postmortem quality.
4. Automate:Deploy incident.io or PagerDuty per client, or adopt the client's existing tool, so escalation and channel setup no longer depend on someone remembering the right steps.
5. Buy:Build offboarding into your contract template, including a guaranteed export of incident history and a transition period where the client's own team shadows your rotation.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Should a development agency standardize on one incident tool across every client, or use whatever each client already has?

Standardize where you can, but defer to a client's existing setup when they already run a mature on-call program. Introducing a second tool just for your convenience usually creates confusion during handoff and rarely survives past the end of the engagement.

What should be included in an offboarding checklist when an engagement using incident.io or PagerDuty ends?

Export incident history and postmortems, transfer or remove your team's access, hand over any runbooks embedded in the tool, and confirm the client has an active on-call rotation of their own before your last responder is removed.

Which account should hold the incident history on a client engagement?

Ideally the client's, when the incident lives in their systems. If you ran incidents through your own account, the client is left with no record of past outages when the retainer ends. Decide whose account holds the history before signing, and check how per-seat pricing treats client staff and subcontractors.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides