Skip to main content

How to Build an Incident Timeline While You Respond

GuideWritten by oncall.fyi editorialPublication approved by Burak YApproval recorded 12 September 2026
Sources and verification
Source dates
Oldest source check: 12 September 2026.
Technical verification
Separate technical verification has not been recorded.

Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.

In short

Build the incident timeline as work happens, using a shared time basis and a short record of observations, hypotheses, decisions and results. Link the original evidence and keep event time distinct from the time somebody recorded it. Preserve rejected hypotheses with the evidence that changed your mind. Reconcile the record after stabilization, and mark missing or uncertain details instead of filling gaps with confident guesses.

Key takeaways

  • Record why an action was taken as well as when it happened.
  • Retain rejected hypotheses so the next responder avoids repeating them.
  • Late notes and uncertain timestamps should remain visibly qualified.

Choose one record before conversation splits

Open a shared incident document or the incident system's timeline and link it from the active coordination channel. Name the person maintaining it. On a small incident this can be the responder; when typing competes with recovery work, ask another participant to capture short notes.

Keep the record reachable by the incoming shift and the teams likely to be paged next. Confirm access with an ordinary responder account. A beautifully structured document behind the wrong permissions still forces the next engineer to reconstruct the incident from chat.

Google's incident-management guidance (opens in a new tab) describes maintaining a working incident record. The worksheet below adds a practical distinction between evidence, interpretation and decisions so the record can support later review.

Record facts in a fixed time basis

Use UTC or another explicitly named timezone. Preserve the timestamp supplied by the original event and note when the evidence was captured. If a device's clock is suspect, retain that uncertainty instead of silently aligning it to the narrative.

A late note can be useful: “Recorded at 14:20 UTC; operator recalls beginning the restart around 14:11, exact time not yet confirmed.” That should not become a precise 14:11 event until another source supports it. Comparing deployment, application and chat timestamps may reveal clock offsets rather than the order you initially assumed.

Use a compact timeline table

The following entries describe a fictional incident:

Event time UTCTypeRecordEvidence or result
14:03ObservationCheckout errors increased in region ASaved dashboard interval
14:06HypothesisNew application release may contributeChange began at 14:01
14:08Check resultOld and new instances both show errorsScoped log-query result
14:09DecisionCheck dependency saturation before rollbackRelease-specific hypothesis weakened
14:13ActionApply the documented dependency mitigationAuthorized operator and change record
14:17ObservationError rate recovered in region AFresh user-path check; observe further

Keep entries short enough to write during response. Link to restricted evidence instead of copying tokens, customer data or large raw log extracts into broad chat channels. Use a stable evidence location with retention appropriate to your incident process.

Capture decisions and rejected hypotheses

A decision needs the question being answered, the evidence available, the chosen action and the reason for choosing it. It may also need the condition that would reverse the decision. “Restarted worker” is incomplete when the important question later is why a restart appeared safer than a rollback.

Maintain a small hypothesis list with states such as open, supported, weakened and rejected. Rejected means a particular claim was contradicted by evidence; it does not mean every possible variant has been disproven. Write the scope. “Not isolated to release B” is more accurate than “deployment ruled out” if a shared configuration change remains possible.

At handoff, ask the incoming responder which checks they would repeat. If they repeat a completed check because the evidence is missing, improve the record. If they repeat it because the system has changed since the earlier observation, that can be the correct next action.

Reconcile after stabilization without rewriting history

Compare the working notes with notification logs, deployment history and available application events. Add missing links and correct errors transparently. Keep the original reasoning visible when later evidence changes a conclusion. A postmortem should explain what responders knew at the time, not imply that the final cause was obvious from the start.

The PagerDuty postmortem template (opens in a new tab) offers a reference structure for a timeline and follow-ups. Your record still needs your own evidence; a template or generated summary cannot supply missing observations.

Verify that another engineer can use it

Give the sanitized record to a teammate who did not attend. Ask them to identify current impact, the evidence for recovery, the strongest remaining hypothesis and the next action. Note any questions that require the original responder's memory.

For the next incident, capture those missing fields as work happens. Measure reconstruction time and repeated investigations across a few comparable incidents. The goal is a record that helps response now and preserves reasoning later, without turning the responder into a full-time transcription service.

Did this help?

Your answer helps us improve this guide. We save only the page and your choice for 30 days.

No name, email, or incident details are requested.

Frequently asked

Should we delete a hypothesis after disproving it?
Keep it with the evidence and scope of rejection so later responders understand the decision and do not repeat an obsolete investigation.
Can a generated summary replace the timeline?
Use it as a draft to verify against original records. It cannot establish missing facts, exact timestamps or unrecorded decision reasoning.

Sources

Vendor facts change. Each source below shows the date this page last checked it.

  1. Managing incidents Google SRE. Checked 12 September 2026.
  2. Postmortem template PagerDuty. Checked 12 September 2026.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Follow the field guide

Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.

Subscribe with RSS