Skip to main content

What is going wrong with your on-call?

Start with the symptom. Find a concrete first test, a practical guide and an artifact you can use with your team.

A notification never reached me

Locate the failure between event creation, routing, the provider and the phone.

Too many alerts need no action

Inspect an export, find recurring patterns and test changes without hiding outages.

Nobody owns the page

Map service ownership, historical schedule changes and backup responsibilities.

Our monitor can fail with the system

Check independent observation, heartbeats and missing-data conditions.

A cron job fails silently

Distinguish process success, valid output and heartbeat delivery.

We need to leave PagerDuty

Compare a tested operating model and migrate without losing alert coverage.

A Prometheus rule change feels risky

Check parsing, test data, runtime loading and receiver delivery separately.

On-call is wearing the team down

Review coverage, interruption load, handoffs and recovery time.

Guided diagnosis · no signup

A page did not reach you. Where did it stop?

Follow the last proven step. These answers stay in this tab; this walkthrough does not connect to your provider or phone.

Check 1 of 6

Does the paging provider show an incident for the missed event?

Compare the source event time and ID with the provider's ingestion and incident logs.