The practical field guide for on-call teams
Build an on-call system people can trust.
Practical guides, assessments, and templates for alert quality, escalation, fair rotations, handoffs, compensation, and incident response.
Free tools. No signup required. Written for the people who carry the pager.
Detected
API latency high
Routed
Payments primary
No acknowledgement
5 min
Escalated
Platform secondary
Owned by a human
Acknowledged
What brought you here?
Start with the problem. Leave with a test, a working document or a next action.
My phone missed a critical alert
Trace the failure from the event to your phone.
Start here →We are drowning in noisy alerts
Inspect an alert CSV locally and find the biggest patterns.
Start here →Nobody owns the next escalation
Write an explicit primary, backup and fallback path.
Start here →I need to prove paging works
Build a controlled test and record delivery evidence.
Start here →A cron job can fail silently
Check successful work and detect missing check-ins.
Start here →We need a different on-call provider
Compare scope, complete costs and acceptance tests.
Start here →Already working on a service? Continue your saved results and action plan.
Before the incident
On-call fails quietly before it fails loudly.
Coverage gaps, noisy alerts, unclear ownership, and brittle escalation paths usually exist long before the night they become urgent. A good on-call system makes those risks visible while there is still time to fix them.
Coverage gaps
Know who owns each hour, including holidays, leave, and handoffs.
Noisy alerts
Separate signals that require action from notifications that merely create interruption.
Ambiguous escalation
Define what happens when the first responder cannot acknowledge or resolve the issue.
Uneven load
Measure nights, weekends, interruptions, and recovery—not only total shift hours.
Learn it. Plan it. Improve it.
Learn
Clear guides for the decisions behind schedules, escalation, alerting, handoffs, and incident response.
Browse guidesPlan
Score alert fatigue, draft an escalation path, assess readiness, and export practical team artifacts.
Open the toolsImprove
Keep dated results, service documents and follow-up work together. Compare progress after your next change.
Open your workspaceUseful before the pager rings.
Alert Log Analyzer
Inspect a CSV locally for alert volume, acknowledgment timing and actionability, with explicit missing-data checks.
Analyze your alert logOn-Call Readiness Assessment
Score coverage, escalation, alert quality, documentation, fairness, testing, and notification delivery.
Get your scoreAlert Fatigue Score
Estimate how much paging load is actionable, duplicated, auto-resolved, or missing context.
Check alert noiseEscalation Policy Builder
Turn a verbal fallback plan into explicit steps, timeouts, targets, and stop conditions.
Build a policyPager Test Checklist Generator
Generate a controlled end-to-end paging test with expected evidence.
Generate a checklistOn-Call Policy Template
Make coverage, handoffs, compensation, recovery, escalation, and ownership explicit before the pager rings.
Use the template
A practical path from signal to response.
01
Detect
A monitoring system identifies a condition that may need action.
02
Route
Ownership rules choose the responsible service, team, and path.
03
Page
The current responder receives the configured notification.
04
Acknowledge
The system records that a human has taken ownership.
05
Escalate or resolve
The path continues until the incident has a responsible next step.
06
Learn
The team improves alerts, runbooks, or ownership based on what happened.
Reliability includes the responder.
A schedule can be technically complete and still be unfair, exhausting, or impossible to sustain. Strong on-call programs make room for coverage requests, shadowing, recovery after disrupted nights, clear compensation, and escalation without blame.
Moving providers
Migrate the operating model, not only the configuration.
A safe migration includes schedules, escalation rules, contact methods, integrations, ownership, runbooks, audit requirements, and a tested cutover path.
Opsgenie migration checklist
Inventory, parity matrix, dual-run plan, controlled cutover test, and rollback.
Grafana OnCall OSS alternatives
What changed, what still works locally, and how to choose a replacement.
Provider evaluation worksheet
Compare ten providers by required workflows, plan costs and trial evidence.
FROM THE TEAM BEHIND ONCALL.FYI
Need a simpler path from critical signal to notification?
MonoDuty connects webhook alerts, uptime checks and heartbeats with incident ownership, schedules and configured delivery paths. Compare its plan limits and escalation behavior against the work your team needs to do.
MonoDuty and oncall.fyi are built by the same team.
The On-Call Brief
One practical idea for calmer, fairer incident response.
Short field notes, templates, and operational lessons. No daily noise and no incident theater.
Follow the field guide
Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.
Subscribe with RSS