How to Audit an On-Call Rotation with Eighty Pages a Week
Sources and verification
- Source dates
- Oldest source check: 12 September 2026.
- Technical verification
- Separate technical verification has not been recorded.
Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.
In short
Audit a complete rotation using notification and incident records, then separate urgent actions, duplicate notifications, nonurgent work and unknown outcomes. Count interruptions as well as unique incidents. Bring the loudest rules to their owners with evidence, a proposed change and a failure test. Prioritize fixes using observed responder time and service risk, and compare the next rotation without assuming fewer pages means better coverage.
Key takeaways
- Eighty notifications are not necessarily eighty different incidents.
- An alert that recovered before action is not automatically a false positive.
- A reliability proposal should include an owner, reserved capacity and a test proving coverage survives.
Define the sample before counting
Use a full rotation with a known start, end and timezone. Include quiet shifts and busy nights instead of selecting only the worst incident. Export notification attempts, incident transitions and responder actions where available. A provider may count an SMS retry as another notification while an incident export contains only one row; document which dataset you have.
Do not turn an anecdotal number such as eighty pages into a benchmark for all teams. Here, eighty is a worked planning example. Record the actual counts for your rotation and distinguish team totals from the interruptions experienced by one person.
The alert log analyzer can summarize a CSV locally in your browser. Its results are limited by the fields you supply: it cannot infer a responder's work or whether a customer was affected from a timestamp alone.
Classify outcomes with the responder
Review the record with the person who handled it and the team that owns the rule. Use “unknown” when the evidence is incomplete.
| Classification | Evidence to look for | Candidate change |
|---|---|---|
| Urgent human action | A decision prevented or limited time-sensitive harm | Keep paging; improve context if needed |
| Same incident, repeat interruption | Same underlying condition and occurrence | Repair grouping or retry behavior |
| Real but nonurgent work | Work can wait within an agreed deadline | Route to an owned ticket queue |
| Recovered without intervention | Condition and recovery visible in telemetry | Examine duration and user impact before tuning |
| Unknown outcome | Missing notes or missing telemetry | Improve evidence before removing coverage |
A latency spike can be a real service failure even if the responder cannot reproduce it after acknowledgement. Conversely, a threshold can be too sensitive without the measurement itself being false. Keep that distinction in the report so the repair matches the failure.
Count interruptions and work separately
For each rule, record delivered pages, unique incidents, after-hours interruptions, acknowledgement delay, active investigation minutes and whether work was possible. Avoid counting two people investigating the same hour as only one person-hour; avoid adding overlapping intervals twice for the same responder.
Illustrative arithmetic: eighty interruptions at an observed average of six active minutes consume eight person-hours. That does not include recovery time between tasks, sleep disruption or long-term impact. Measure those separately if the team chooses to track them; do not invent a universal multiplier to make a proposal look larger.
Show the top few rules by interruptions, then inspect the risk they cover. A rare restore failure may warrant immediate attention even though it does not appear near the top of the volume chart.
Turn page data into scheduled reliability work
Create one concrete change per candidate rule. Include the observed problem, owner, estimated implementation effort, desired behavior, acceptance test and rollback. Pair a paging team representative with the rule's author when those are different groups. An urgent route should not be abandoned while the groups negotiate responsibility.
An example proposal is: “Replace the instantaneous queue-depth page with the agreed job-age condition; replay the known stuck-worker case and the normal traffic peak; keep the old rule on a test receiver for one rotation.” This is easier to schedule and verify than “reduce alert fatigue.”
Reserve implementation time in the next planning cycle and assign someone to review the result. Google's discussion of toil (opens in a new tab) describes repetitive operational work as a capacity issue; your local measurements determine which changes deserve that capacity.
Compare the next rotation fairly
Keep the same counting definition and note traffic, deployments, staffing and holiday differences. Compare urgent-action detection, interruptions and unresolved failures. Run a controlled trigger, recovery and missed-primary test on each changed route. A short quiet window is weak evidence for a condition that historically happens once a month.
Publish a small review record: what changed, what remained noisy, which tests passed, what was not tested and when the owner will review again. Refer to the page, ticket and dashboard criteria before demoting a rule. This is a prioritization method, not a justification for globally muting a difficult rotation.
Did this help?
Your answer helps us improve this guide. We save only the page and your choice for 30 days.
No name, email, or incident details are requested.
Sources
Vendor facts change. Each source below shows the date this page last checked it.
- Being on-call — Google SRE. Checked 12 September 2026.
- Eliminating toil — Google SRE. Checked 12 September 2026.
Related
One practical idea, occasionally
The On-Call Brief: short field notes, templates, and operational lessons.
Follow the field guide
Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.
Subscribe with RSS