How to Write an Incident Handoff That Survives the Postmortem
Sources and verification
- Source dates
- Oldest source check: 10 September 2026.
- Technical verification
- Separate technical verification has not been recorded.
Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.
In short
Maintain one timestamped incident record that distinguishes observations, hypotheses, decisions and outcomes. Name the technical owner separately from the coordinator, link recovery evidence, and assign each follow-up an owner and due date. At handoff, have the incoming responder explain the current state from the record. Reuse that timeline for the postmortem instead of reconstructing the incident from disconnected chat summaries.
Key takeaways
- Separate coordination responsibility from technical conclusions.
- Record uncertainty while it exists instead of rewriting history after recovery.
- Follow-up actions need owners and verifiable completion conditions.
Keep a single working record
An external service desk can coordinate an incident without having the system context to decide its technical cause. Name a technical investigation owner and a coordinator explicitly. Let each contribute to one record; do not make a polished service-desk summary replace the primary evidence.
Use UTC or another clearly recorded common time basis. Link the incident identifier from the relevant chat and ticket so responders know which document is current.
Copy this handoff structure
Incident ID and current severity:
Customer or service impact, with evidence:
Coordinator / technical owner / current responder:
Current state and last verified time:
Timeline:
time | observation | decision/action | result | evidence
Working hypotheses and what would disprove them:
Actions already attempted and their effects:
Recovery evidence and remaining uncertainty:
Next action, owner and expected update time:
Follow-ups: owner | due date | completion evidence
Keep observations separate from hypotheses. “Errors began at 10:04” and “the new release caused errors” are not interchangeable entries. If a theory is later rejected, preserve it with the evidence that rejected it; this explains why the team took a particular action at the time.
Hand over decisions, not just links
Before the outgoing responder leaves, the incoming responder should state the current impact, what has been ruled out, the next action and who can authorize a risky change. Ask them to open the key evidence with their own account. Missing permissions at handoff are a response gap.
In an illustrative handoff, a rollback reduces errors but one region still fails. Record partial recovery rather than “resolved.” Keep the remaining region's owner and next check explicit.
Convert the record into the review
After stabilization, build the postmortem from the same timeline. Add the impact summary, contributing conditions, detection and response gaps, and actions that reduce recurrence or improve recovery. Google SRE's postmortem chapter (opens in a new tab) supports the blameless learning approach; it does not establish facts about your incident.
Avoid replacing a contributing system condition with a person's name. “The runbook lacked the recovery precondition” produces a more actionable follow-up than an unsupported conclusion that someone should have known it.
Test whether context survived
Give the record to an engineer who was not present. Ask them to identify the impact, the evidence for recovery and unfinished work. Record what they had to ask and how long reconstruction took.
Assign every material follow-up an owner and a checkable result, such as a demonstrated restore or a passing notification test. Reuse the handoff checklist for live transfer and this fuller structure for the review. A completed report without completed actions does not establish that the response improved.
Did this help?
Your answer helps us improve this guide. We save only the page and your choice for 30 days.
No name, email, or incident details are requested.
Frequently asked
- Who should write the technical cause?
- The accountable technical owner should verify that conclusion against evidence. A coordinator can maintain the record without being the authority on system behavior.
- Should rejected hypotheses be deleted?
- Keep them clearly marked with the evidence that rejected them. They explain decisions made under uncertainty and can reveal investigation gaps.
Sources
Vendor facts change. Each source below shows the date this page last checked it.
- Postmortem culture — Google SRE. Checked 10 September 2026.
Related
One practical idea, occasionally
The On-Call Brief: short field notes, templates, and operational lessons.
Follow the field guide
Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.
Subscribe with RSS