Skip to main content

How to Write and Track Postmortem Actions That Can Ship

GuideWritten by oncall.fyi editorialPublication approved by Burak YApproval recorded 12 September 2026
Sources and verification
Source dates
Oldest source check: 12 September 2026.
Technical verification
Separate technical verification has not been recorded.

Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.

In short

Turn each postmortem finding into a bounded change with an accountable owner, a realistic delivery decision and a checkable result. Separate immediate containment from longer architectural work and place both in the normal team planning queue. Review blocked and overdue items with the responsible team, record accepted residual risk, and close an action only when its stated evidence exists rather than when the report is finished.

Key takeaways

  • Split immediate mitigation from larger engineering work.
  • An owner and due date need a capacity decision behind them.
  • Close actions with evidence and keep residual risk visible.

Begin with the failure you want to change

A postmortem action should identify the condition it addresses. “Improve monitoring” leaves the implementer to guess the scope and the reviewer to guess whether it worked. “Detect loss of the payment worker heartbeat within the agreed window and demonstrate backup delivery” supplies a behavior that can be checked.

Attach the action to the incident finding and its evidence. Distinguish prevention, detection, mitigation and recovery work. They can all be valuable, but a faster recovery procedure should not be reported as proof that the failure cannot recur.

Google's postmortem-culture workbook (opens in a new tab) discusses learning and follow-through after incidents. The planning method here is a proposed way for a small team to turn that learning into ordinary, deliverable work.

Separate containment from the architectural project

An incident may expose a large design problem that cannot responsibly be resolved in a few days. Write a short-term containment item, a bounded investigation or design decision, and a separately planned implementation when appropriate. Give each its own evidence and risk statement.

For a fictional queue outage, immediate containment might limit incoming work with an existing supported control. A later project might isolate customer workloads. The temporary limit could reduce impact while leaving capacity and fairness risks unresolved. Record those limits; do not close the architectural finding because the quick mitigation shipped.

ItemScopeCompletion evidence
ContainmentAdd the approved queue limit and runbook stepDemo overload triggers the intended bound
Design decisionCompare isolation approaches and choose a pilotReviewed decision with constraints and owner
Engineering pilotIsolate one low-risk workload classFailure exercise stays inside the defined boundary

These examples describe work structure, not instructions for an arbitrary production queue.

Make ownership and capacity explicit

Name one accountable owner even if several engineers contribute. Ask that owner to accept the scope and identify the team that supplies capacity. A person listed in a document without an agreed allocation is a contact, not a delivery plan.

Set a due date using urgency, dependencies and actual capacity. If the action displaces planned work, have the appropriate planning owner make that tradeoff visible. When the deadline is uncertain, assign a near-term decision date rather than inventing a confident implementation date.

A useful action record contains:

text
Incident and finding:
Change in behavior we want:
Scope and exclusions:
Accountable owner / delivery team:
Dependencies and capacity decision:
Next milestone and agreed date:
Completion check and evidence location:
Risk that remains until completion:

Avoid making “be more careful” or “train everybody” the only response to a missing control. If training is necessary, specify the task to rehearse and how the system supports a person who makes a mistake.

Make follow-ups visible in the normal queue

Create linked work items in the tool the team already uses for planning. Keep the postmortem as context and link back from each item. Do not maintain an independent spreadsheet whose status drifts from the engineering queue.

Run a short weekly review of unaccepted, blocked and overdue actions. Ask what decision would move each item forward. A blocked dependency may need another team; an oversized task may need a narrower pilot. An action no longer worth implementing needs an explicit risk decision by the responsible owner, not silent disappearance from the report.

Share a compact update that highlights the risk removed, remaining gaps and the next decision. Making every engineer read a long report is less useful than making the actionable part relevant to the services they operate.

Close with a demonstration

Use the stated completion check. For a notification fix, retain a controlled receipt and backup-escalation test. For a restore action, keep the isolated restore result. For a runbook improvement, have an engineer who did not write it perform the approved exercise.

The PagerDuty postmortem template (opens in a new tab) separates analysis and follow-ups, but a filled template is not evidence that work shipped. Keep the deployed change, test result or reviewed decision beside the closed item.

Look for recurring findings

At the next service review, compare new incidents with still-open and recently completed actions. If the same failure recurs, determine whether the action was unfinished, too narrow, incorrectly verified or aimed at a different contributing condition. Record that distinction without turning the review into an individual performance ranking. Completion counts are useful administration; reduced risk requires evidence about how the system now behaves.

Did this help?

Your answer helps us improve this guide. We save only the page and your choice for 30 days.

No name, email, or incident details are requested.

Frequently asked

Should every action be due within two weeks?
No. Match the scope, urgency, dependencies and capacity. Use a near-term containment or design milestone when the full change needs longer.
What if an action is no longer worth implementing?
Have the responsible owner record that decision and residual risk explicitly, rather than leaving it indefinitely open or deleting it silently.

Sources

Vendor facts change. Each source below shows the date this page last checked it.

  1. Postmortem culture Google SRE. Checked 12 September 2026.
  2. Postmortem template PagerDuty. Checked 12 September 2026.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Follow the field guide

Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.

Subscribe with RSS