Skip to main content

How to Write an Incident Handoff That Survives the Postmortem

GuideWritten by oncall.fyi editorialPublication approved by Burak YApproval recorded 11 September 2026
Sources and verification
Source dates
Oldest source check: 10 September 2026.
Technical verification
Separate technical verification has not been recorded.

Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.

In short

Maintain one timestamped incident record that distinguishes observations, hypotheses, decisions and outcomes. Name the technical owner separately from the coordinator, link recovery evidence, and assign each follow-up an owner and due date. At handoff, have the incoming responder explain the current state from the record. Reuse that timeline for the postmortem instead of reconstructing the incident from disconnected chat summaries.

Key takeaways

  • Separate coordination responsibility from technical conclusions.
  • Record uncertainty while it exists instead of rewriting history after recovery.
  • Follow-up actions need owners and verifiable completion conditions.

Keep a single working record

An external service desk can coordinate an incident without having the system context to decide its technical cause. Name a technical investigation owner and a coordinator explicitly. Let each contribute to one record; do not make a polished service-desk summary replace the primary evidence.

Use UTC or another clearly recorded common time basis. Link the incident identifier from the relevant chat and ticket so responders know which document is current.

Copy this handoff structure

text
Incident ID and current severity:
Customer or service impact, with evidence:
Coordinator / technical owner / current responder:
Current state and last verified time:
Timeline:
  time | observation | decision/action | result | evidence
Working hypotheses and what would disprove them:
Actions already attempted and their effects:
Recovery evidence and remaining uncertainty:
Next action, owner and expected update time:
Follow-ups: owner | due date | completion evidence

Keep observations separate from hypotheses. “Errors began at 10:04” and “the new release caused errors” are not interchangeable entries. If a theory is later rejected, preserve it with the evidence that rejected it; this explains why the team took a particular action at the time.

Before the outgoing responder leaves, the incoming responder should state the current impact, what has been ruled out, the next action and who can authorize a risky change. Ask them to open the key evidence with their own account. Missing permissions at handoff are a response gap.

In an illustrative handoff, a rollback reduces errors but one region still fails. Record partial recovery rather than “resolved.” Keep the remaining region's owner and next check explicit.

Convert the record into the review

After stabilization, build the postmortem from the same timeline. Add the impact summary, contributing conditions, detection and response gaps, and actions that reduce recurrence or improve recovery. Google SRE's postmortem chapter (opens in a new tab) supports the blameless learning approach; it does not establish facts about your incident.

Avoid replacing a contributing system condition with a person's name. “The runbook lacked the recovery precondition” produces a more actionable follow-up than an unsupported conclusion that someone should have known it.

Test whether context survived

Give the record to an engineer who was not present. Ask them to identify the impact, the evidence for recovery and unfinished work. Record what they had to ask and how long reconstruction took.

Assign every material follow-up an owner and a checkable result, such as a demonstrated restore or a passing notification test. Reuse the handoff checklist for live transfer and this fuller structure for the review. A completed report without completed actions does not establish that the response improved.

Did this help?

Your answer helps us improve this guide. We save only the page and your choice for 30 days.

No name, email, or incident details are requested.

Frequently asked

Who should write the technical cause?
The accountable technical owner should verify that conclusion against evidence. A coordinator can maintain the record without being the authority on system behavior.
Should rejected hypotheses be deleted?
Keep them clearly marked with the evidence that rejected them. They explain decisions made under uncertainty and can reveal investigation gaps.

Sources

Vendor facts change. Each source below shows the date this page last checked it.

  1. Postmortem culture Google SRE. Checked 10 September 2026.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Follow the field guide

Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.

Subscribe with RSS