Skip to main content

How to Use the First Five Minutes After an Incident Acknowledgement

GuideWritten by oncall.fyi editorialPublication approved by Burak YApproval recorded 12 September 2026
Sources and verification
Source dates
Oldest source check: 12 September 2026.
Technical verification
Separate technical verification has not been recorded.

Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.

In short

After acknowledging an incident, confirm current impact, state who owns the response and open the service runbook with a fixed investigation window. Identify the next useful check and the person who can act on it. Escalate with evidence when access, expertise or response capacity is missing. The first five minutes should produce a shared picture and a bounded next action, not an unsupported root-cause claim.

Key takeaways

  • Acknowledgement starts ownership; it does not prove investigation has begun.
  • Choose the next check from impact and service context.
  • A handoff completes when the receiving responder explicitly accepts it.

First confirm that someone is taking responsibility

An acknowledgement button can stop escalation even when the responder has no working laptop or production access. State ownership in the incident record and immediately report any constraint. If you cannot investigate within the agreed response target, request the backup and keep responsibility explicit until they accept.

This five-minute sequence is an example for an actionable service incident. Adapt it to your service and response policy. It does not replace emergency procedures or authorize a production change, and it is not a universal response-time promise.

Minute one: establish observed impact

Read the alert's service, environment, start time and measured condition. Open a relevant user-facing signal or safe synthetic check. Determine what is failing, who is affected and whether impact is continuing. If monitoring is missing, write “impact unknown” and investigate that uncertainty instead of interpreting an empty graph as recovery.

Use a short statement such as: “Checkout requests in region A have increased errors since 14:03 UTC; region B is currently within its usual range.” This is an invented example. It is more useful than “the platform is broken” because another responder can test its scope.

Minutes two and three: open the same evidence window

Use the runbook's service overview, relevant logs and change history with the same timestamps and cluster context. Prefer read-only investigation first. Choose a bounded question: does the problem affect one release, one dependency, or all instances? Record the answer and a link to the evidence before trying the next question.

Recent changes are leads, not conclusions. A deployment near the start of an incident can be unrelated. Compare affected and unaffected instances and note competing explanations. If a mitigation is already documented, verify its preconditions before proceeding; for example, a rollback may not be safe after an incompatible data migration.

The Google SRE incident-response workbook (opens in a new tab) provides broader context for systematic investigation. The minute-by-minute structure here is a proposed drill, not a claim that every incident can be diagnosed in five minutes.

Minutes four and five: commit to a next action

Choose the next check or approved mitigation, name its owner and specify when they will report back. Record what result would support the hypothesis and what would make you stop. If several responders are present, give them different questions rather than allowing simultaneous uncoordinated changes.

A compact working note might look like this:

text
Observed impact: checkout failures in region A; region B checked separately
Current responder: primary on-call
Established: errors affect both old and new application instances
Uncertain: dependency latency increased near the incident start
Next check: dependency owner compares request latency and saturation
Check owner / update due: named responder / 14:12 UTC
Change authority: service recovery runbook and incident lead

The five minutes end with a useful state even if the cause remains unknown. If coordination is consuming investigation time, assign an incident commander.

Reduce escalation handoff delays

When support, SRE and a development team all participate, keep one incident record and one person responsible for overall progress. Send the next team the observed impact, the specific question requiring their expertise, evidence already gathered and actions already attempted. Ask for an explicit acceptance and update time.

Measure transfer delay from the request timestamp to that acceptance. Measure tool-access delay separately. A team that replies quickly but cannot access the dashboard has a different problem from a request waiting in an unattended queue.

If the receiving team does not accept within the agreed interval, use the next escalation contact. Do not silently remove the current owner simply because a ticket changed queues. Google's incident-management chapter (opens in a new tab) explains why clear coordination and role transfer matter; your exact escalation intervals belong in your own policy.

Rehearse with someone new to the service

Replay a sanitized past incident or create a harmless demo failure. Give a new responder only the alert and the normal runbook. Observe whether they can state impact, locate evidence and identify the next owner without private coaching.

Record broken links, missing access, ambiguous service names and repeated handoffs. Change the alert or runbook to remove the largest delay, then repeat that case. Compare time to a useful next action and quality of the handoff, not just time to click acknowledgement.

Did this help?

Your answer helps us improve this guide. We save only the page and your choice for 30 days.

No name, email, or incident details are requested.

Frequently asked

Do we need the root cause in five minutes?
No. Establish impact, ownership, useful evidence and a bounded next action; preserve uncertainty where the cause is not proven.
When does an escalation handoff finish?
When the receiving responder accepts the specific responsibility and can access the necessary context, not merely when a ticket is reassigned.

Sources

Vendor facts change. Each source below shows the date this page last checked it.

  1. Incident response Google SRE. Checked 12 September 2026.
  2. Managing incidents Google SRE. Checked 12 September 2026.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Follow the field guide

Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.

Subscribe with RSS