Skip to main content

How to Decide Between a Page and a Ticket

GuideWritten by oncall.fyi editorialPublication approved by Burak YApproval recorded 12 September 2026
Sources and verification
Source dates
Oldest source check: 12 September 2026.
Technical verification
Separate technical verification has not been recorded.

Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.

In short

Page when waiting for the normal work queue would create unacceptable harm and an available responder can change the outcome. Use a ticket when work can safely wait, with an owner and a due time. Record the evidence behind that decision, test the replacement route and keep escalation for conditions that become urgent.

Key takeaways

  • Urgency is a deadline for useful action, not a synonym for an unusual metric.
  • A ticket needs an owner, due time and escalation when risk increases.
  • Change delivery only after replaying an incident the alert must still catch.

Write the action deadline first

Take one alert that repeatedly wakes someone who does nothing. Ask what would happen if nobody responded until the next working period. Write the expected harm, how soon it could occur, and the action that changes it. If those answers are unknown, involve the service owner before changing its route.

An illustrative queue warning might give a team several days to add capacity. A queue whose oldest item will miss a customer deadline in twenty minutes needs a different response. The metric name alone cannot decide whether either condition should ring a phone.

Google's alerting on SLOs (opens in a new tab) gives a basis for matching notification urgency to error-budget consumption. Translate that into your actual response windows; copying another service's numbers can produce a misleading policy.

Use a small classification table

ConditionDelivery candidateWhat must be written down
Harm is happening and action cannot waitPageFirst action and escalation
Harm is approaching but the work queue has timeTicketLatest safe completion time
Useful investigation signal without required workDashboard or logWho uses it and why
Intent or owner is unknownOwned review queueTemporary coverage and review deadline

Treat the last row as a short investigation state. It is not permission to abandon an alert that might protect an urgent customer journey. Review incident history and the service contract with someone who understands the dependency.

Make the ticket operational

A different delivery channel does not make an alert useful by itself. Include the service, environment, observed value, expected range, first diagnostic link and proposed due time. Route it to a queue with a named team that actually checks it.

Define repeat behavior. The same condition should update its existing work item or be visibly correlated, rather than creating a fresh ticket every evaluation interval. Recovery can close the signal while the engineering task stays open if the underlying defect still needs repair.

Also specify the upgrade condition. For the queue example, an age threshold or failed customer transaction can activate an urgent rule even while the capacity ticket remains open. Review both conditions together so the change does not create a period with neither useful ticket nor page.

Rehearse the change without doubling pages

Use historical events or a test environment to compare the existing route and proposed classification. Keep only the agreed production path capable of paging while the candidate produces review evidence.

Include these cases in the exercise:

  • A brief deviation with no user impact.
  • A persistent condition that deserves daytime work.
  • A rapidly worsening version of the same condition.
  • Missing measurements or an unreachable data source.
  • Recovery followed by recurrence before the original ticket is complete.

For each case, record the observed destination, expected owner and earliest useful action. A lower page count is not a successful result if an urgent fixture becomes invisible.

Review the outcome after a complete work cycle

Check whether tickets were seen before their deadlines and whether a responder could act when an urgent condition occurred. Reopen the delivery decision if work waits beyond the available safety margin. Include nights, weekends and team holidays when calculating that margin.

Use the alert log analyzer to inspect exported history, then investigate individual classifications with the owner. The tool summarizes uploaded records; it cannot decide customer impact or prove that an unobserved incident would have been detected.

Did this help?

Your answer helps us improve this guide. We save only the page and your choice for 30 days.

No name, email, or incident details are requested.

Sources

Vendor facts change. Each source below shows the date this page last checked it.

  1. Alerting on SLOs Google SRE. Checked 12 September 2026.
  2. Monitoring distributed systems Google SRE. Checked 12 September 2026.

Related

One practical idea, occasionally

The On-Call Brief: short field notes, templates, and operational lessons.

Follow the field guide

Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.

Subscribe with RSS