Skip to main content

On-Call Handoffs: Transfer Risk, Not Just the Pager

GuideWritten by Burak YReviewed by Burak YLast reviewed 3 September 2026

In short

An on-call handoff transfers current operational risk, not just the pager. The outgoing responder passes on recent changes, active incidents, temporary mitigations, known-noisy alerts and upcoming events; the incoming responder confirms they can receive a page and understands what is fragile right now. A handoff that only says "nothing happened" has moved the rota, not the risk.

Key takeaways

  • The rota moves the pager automatically; the handoff exists to move everything the rota cannot see.
  • Temporary mitigations and silenced alerts are the highest-value items in any handoff, because nothing else in your monitoring will remind you they exist.
  • An asynchronous handoff is not complete until the incoming responder acknowledges it in writing.
  • If an item would not change what the incoming responder does in the next few hours, it belongs in the shift log, not the handoff.
  • A schedule boundary does not end an incident: hand over roles explicitly and have the incoming responder repeat the state back.

At a glance

A handoff is the moment ownership of production moves from one person to another. The schedule does that part automatically at the appointed time. Everything else — what is currently fragile, what is half-fixed, what will page tonight and can safely be ignored — only moves if someone moves it.

Handoff problems are rarely caused by people forgetting to talk. They are caused by teams treating the handoff as a status update rather than a transfer of risk.

What a handoff is actually for

Three things, in this order.

  1. Continuity of context. The incoming responder should not have to reconstruct the last twelve hours from dashboards while an incident is running.
  2. Continuity of ownership. At every moment one person owns production, and both people agree who that is.
  3. A checkpoint on invisible state. Temporary changes — a silenced alert, a paused job, a manual failover, a scaled-out worker pool — do not show up in normal monitoring. The handoff is where they are re-declared, revoked, or given an owner.

A meeting that does not do these three things is a meeting, not a handoff.

What must transfer

ItemWhy the incoming responder needs it
Open incidents and their current stateSo they can continue, not restart, the investigation
Temporary mitigations, with expiry and revert stepsThis state is invisible; if it is not spoken, it is forgotten
Silenced or suppressed alerts, and until whenA silence that outlives the reason for it is a coverage gap
Changes shipped in the last shiftThe first thing to check when the next page arrives
Alerts known to be noisy right nowPrevents the incoming responder chasing a known-benign signal
Scheduled events in the coming shiftMigrations, batch jobs, certificate renewals, marketing sends
Anything degraded but not pagingPartial capacity, a lagging replica, a queue that is draining slowly

Two items matter more than the rest: temporary mitigations and open-ended silences. They are the only things on the list that no dashboard will remind anyone about.

What must not become ritual

The opposite failure is a handoff that recites every alert, every deploy and every ticket from the shift. Long handoffs get skimmed, and skimmed handoffs are worse than short ones because everyone believes the information was transferred.

Use one test on every item: would knowing this change what the incoming responder does in the next few hours? If not, it belongs in the shift log, not the handoff.

Synchronous or asynchronous

Time zones decide the shape; risk decides the depth. Both modes work, and neither is a substitute for a written note.

SituationHandoff mode
Active incident, or a mitigation expiring this shiftLive conversation, plus the written note
Overlapping working hours, ordinary shiftWritten note first, then a short live sync
Follow-the-sun with no overlapWritten note, a named question window, explicit acknowledgement
Quiet shift, no temporary state, same teamWritten note and acknowledgement only

The rule that holds across all four rows: an asynchronous handoff is not complete when the note is posted. It is complete when the incoming responder acknowledges it. A note sent into a channel with no reply is a broadcast.

The handoff note

Keep the same fields every time, so a reader can find what they need without reading the whole thing.

  • Ownership — who is primary and secondary, from when to when, in a named time zone.
  • Open incidents — identifier, severity, current state, next action, who is doing it.
  • Temporary state — each change, why it exists, when it expires, how to revert it.
  • Recent changes — what shipped that could plausibly cause the next page.
  • Noise — what is firing and known-benign, whether it is silenced, and until when.
  • Upcoming — anything scheduled in the next shift, and who owns it.
  • Open questions — anything the outgoing responder could not resolve.

Write the note during the shift, not at the end of it. A note assembled from memory in the last five minutes loses exactly the detail that turns out to matter.

What the incoming responder verifies before accepting

Accepting a handoff is an action, not a default. Before saying yes:

  • The pager reaches you. Right device, right contact method, notifications not muted, correct rotation active. Run a real paging test on the first shift after any change of phone, number or tool.
  • The paging tool shows you as primary. Check the tool, not a calendar invite — calendars and schedules drift apart.
  • Every temporary mitigation has an expiry and a revert procedure you understand.
  • Every silence has an end time. Open-ended silences are inherited debt; give each one an owner today.
  • You can reach secondary and the next escalation step, and you know who is unavailable.
  • You have read the runbook for anything currently degraded, before it pages.

Then say so explicitly. "Acknowledged, I have the pager from 09:00 UTC" is a short sentence that removes a whole class of coverage gap.

Handing over during an active incident

A schedule boundary does not end an incident. Someone stays until the handover is finished.

  1. Hand over roles one at a time, starting with incident lead. Handing over "the incident" as a single object is how details get dropped.
  2. The outgoing responder states five things: what is known, what has been ruled out, what has been changed, what happens next, and what would trigger escalation.
  3. The incoming responder repeats it back in their own words. If they cannot, the handover has not happened yet.
  4. Announce the change of lead in the incident channel with a timestamp, so stakeholders know who to ask.
  5. The outgoing responder stays reachable for a stated period. That is a courtesy, not a plan — the incoming responder should be able to proceed without them.

If the incident is severe and the outgoing responder has been awake for hours, a full handover to a rested person is the safer operational choice. Fatigue is a system risk, not a personal weakness.

Common failure modes

  • "Nothing happened." Usually untrue. Something was silenced, deployed or restarted.
  • The undocumented fix. A restart, a scale-up or a flag flip that lives only in one person's memory.
  • Silences with no expiry. The alert that was noisy in March is still off in September, and nobody knows.
  • Handoff by broadcast. Posted, never acknowledged, occasionally never read.
  • Handoff as management report. Written to look competent rather than to be useful at 03:00.
  • Schedule and calendar disagree. Two people think the other has the pager.
  • No handoff at weekend and holiday boundaries. The longest unobserved stretches get the weakest transfer.

A worked example

The following is an invented example used to show the shape of a good note. It does not describe any real team.

A six-person team splits on-call across two regions with no overlap in working hours. The handoff is asynchronous, at 20:00 in the outgoing responder's time zone.

The note says: one open Sev-3, elevated checkout latency, mitigated at 14:10 by scaling payment workers from six to twelve — revert once the fix ships, do not revert before. One silence on the disk-usage alert for a log node, expiring 06:00. A schema migration scheduled for 02:00, owned by a named engineer who is awake for it. One deploy to the search service at 17:40, not yet exercised at peak load.

The incoming responder verifies the pager reaches them, notes that the silence expires inside their shift, confirms the migration owner is genuinely online, skims the payments runbook, and acknowledges.

At 02:20 an alert fires on search latency. Because the deploy was in the note, the incoming responder checks that change first, finds the regression and rolls it back. The time saved came from four lines of a written note, not from cleverness under pressure.

What to do next

  1. Write down the fields your team's handoff note must contain, and put them in a shared template.
  2. Search your alerting tool for silences with no end time. Give each one an owner and an expiry today.
  3. Add an explicit acknowledgement step to the handoff, and define when the outgoing responder may go offline.
  4. Decide the mode for each boundary — live, written, or both — and record it in your on-call policy.
  5. Agree the active-incident handover rule now, in writing, before you need it at 03:00.
  6. At the next handoff, note what the incoming responder still had to ask for. That question is the field missing from your template.

Frequently asked

How long should an on-call handoff take?
Long enough to cover open incidents, temporary state and upcoming events, and no longer. On a quiet shift with no temporary mitigations, a written note and an acknowledgement may be the whole handoff. If a quiet-shift handoff regularly runs long, it is usually a shift log being read aloud rather than risk being transferred.
Do we still need a handoff if nothing happened during the shift?
Yes, but it should be short. Even a quiet shift can leave behind a silence with no expiry, a deploy that has not yet been exercised under load, or a scheduled job in the next window. The note should also confirm who holds the pager next and that they can receive one.
Who owns an incident that is still open when a shift ends?
The outgoing responder owns it until the incoming responder has explicitly accepted it. Announce the change of incident lead in the incident channel with a timestamp so everyone knows who to ask. If the incident is severe and the outgoing responder is exhausted, a full handover to a rested person is safer than continuing.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Unsubscribe at any time. We do not sell subscriber data.