How to Write and Track Postmortem Actions That Can Ship
Sources and verification
- Source dates
- Oldest source check: 12 September 2026.
- Technical verification
- Separate technical verification has not been recorded.
Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.
In short
Turn each postmortem finding into a bounded change with an accountable owner, a realistic delivery decision and a checkable result. Separate immediate containment from longer architectural work and place both in the normal team planning queue. Review blocked and overdue items with the responsible team, record accepted residual risk, and close an action only when its stated evidence exists rather than when the report is finished.
Key takeaways
- Split immediate mitigation from larger engineering work.
- An owner and due date need a capacity decision behind them.
- Close actions with evidence and keep residual risk visible.
Begin with the failure you want to change
A postmortem action should identify the condition it addresses. “Improve monitoring” leaves the implementer to guess the scope and the reviewer to guess whether it worked. “Detect loss of the payment worker heartbeat within the agreed window and demonstrate backup delivery” supplies a behavior that can be checked.
Attach the action to the incident finding and its evidence. Distinguish prevention, detection, mitigation and recovery work. They can all be valuable, but a faster recovery procedure should not be reported as proof that the failure cannot recur.
Google's postmortem-culture workbook (opens in a new tab) discusses learning and follow-through after incidents. The planning method here is a proposed way for a small team to turn that learning into ordinary, deliverable work.
Separate containment from the architectural project
An incident may expose a large design problem that cannot responsibly be resolved in a few days. Write a short-term containment item, a bounded investigation or design decision, and a separately planned implementation when appropriate. Give each its own evidence and risk statement.
For a fictional queue outage, immediate containment might limit incoming work with an existing supported control. A later project might isolate customer workloads. The temporary limit could reduce impact while leaving capacity and fairness risks unresolved. Record those limits; do not close the architectural finding because the quick mitigation shipped.
| Item | Scope | Completion evidence |
|---|---|---|
| Containment | Add the approved queue limit and runbook step | Demo overload triggers the intended bound |
| Design decision | Compare isolation approaches and choose a pilot | Reviewed decision with constraints and owner |
| Engineering pilot | Isolate one low-risk workload class | Failure exercise stays inside the defined boundary |
These examples describe work structure, not instructions for an arbitrary production queue.
Make ownership and capacity explicit
Name one accountable owner even if several engineers contribute. Ask that owner to accept the scope and identify the team that supplies capacity. A person listed in a document without an agreed allocation is a contact, not a delivery plan.
Set a due date using urgency, dependencies and actual capacity. If the action displaces planned work, have the appropriate planning owner make that tradeoff visible. When the deadline is uncertain, assign a near-term decision date rather than inventing a confident implementation date.
A useful action record contains:
Incident and finding:
Change in behavior we want:
Scope and exclusions:
Accountable owner / delivery team:
Dependencies and capacity decision:
Next milestone and agreed date:
Completion check and evidence location:
Risk that remains until completion:
Avoid making “be more careful” or “train everybody” the only response to a missing control. If training is necessary, specify the task to rehearse and how the system supports a person who makes a mistake.
Make follow-ups visible in the normal queue
Create linked work items in the tool the team already uses for planning. Keep the postmortem as context and link back from each item. Do not maintain an independent spreadsheet whose status drifts from the engineering queue.
Run a short weekly review of unaccepted, blocked and overdue actions. Ask what decision would move each item forward. A blocked dependency may need another team; an oversized task may need a narrower pilot. An action no longer worth implementing needs an explicit risk decision by the responsible owner, not silent disappearance from the report.
Share a compact update that highlights the risk removed, remaining gaps and the next decision. Making every engineer read a long report is less useful than making the actionable part relevant to the services they operate.
Close with a demonstration
Use the stated completion check. For a notification fix, retain a controlled receipt and backup-escalation test. For a restore action, keep the isolated restore result. For a runbook improvement, have an engineer who did not write it perform the approved exercise.
The PagerDuty postmortem template (opens in a new tab) separates analysis and follow-ups, but a filled template is not evidence that work shipped. Keep the deployed change, test result or reviewed decision beside the closed item.
Look for recurring findings
At the next service review, compare new incidents with still-open and recently completed actions. If the same failure recurs, determine whether the action was unfinished, too narrow, incorrectly verified or aimed at a different contributing condition. Record that distinction without turning the review into an individual performance ranking. Completion counts are useful administration; reduced risk requires evidence about how the system now behaves.
Did this help?
Your answer helps us improve this guide. We save only the page and your choice for 30 days.
No name, email, or incident details are requested.
Frequently asked
- Should every action be due within two weeks?
- No. Match the scope, urgency, dependencies and capacity. Use a near-term containment or design milestone when the full change needs longer.
- What if an action is no longer worth implementing?
- Have the responsible owner record that decision and residual risk explicitly, rather than leaving it indefinitely open or deleting it silently.
Sources
Vendor facts change. Each source below shows the date this page last checked it.
- Postmortem culture — Google SRE. Checked 12 September 2026.
- Postmortem template — PagerDuty. Checked 12 September 2026.
Related
One practical idea, occasionally
The On-Call Brief: short field notes, templates, and operational lessons.
Follow the field guide
Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.
Subscribe with RSS