Find Devices That Have Silently Lost Backup Coverage
Sources and verification
- Source dates
- Oldest source check: 12 September 2026.
- Technical verification
- Separate technical verification has not been recorded.
Publication approval and technical verification are recorded separately. Automated link checks establish reachability, not accuracy. A checked example verifies only its stated test cases, not the whole article or your production setup.
In short
Compare the devices that should be protected with active backup assignments and recent usable recovery points. A failed-job dashboard misses devices whose agent stopped or whose plan was removed before a job began. Track inventory freshness, policy exceptions, and onboarding deadlines. Do not mark a new device protected until its first backup and the required restore validation are evidenced.
Key takeaways
- An empty failed-job list does not prove all devices are protected.
- Coverage starts from the expected inventory, not from backup jobs.
- A missing inventory feed must become unknown rather than a false 100 percent.
Start with the devices that ought to have backups
Choose an authoritative inventory: managed endpoints, production servers, or an approved service registry. Give each device a stable identifier that survives display-name changes. Record its owner, protection requirement, first-backup deadline, and any approved exception. A laptop renamed in one system should not become two different coverage records.
Import backup assignments, agent observations, and recovery-point metadata separately. Include when each source was fetched successfully. If the backup API times out, an empty response must not be interpreted as an empty estate or a clean report.
Join the inventory to actual protection
Build the report as a left join from expected devices to backup evidence. This preserves rows for devices that have never produced a job. Match stable IDs where available and put ambiguous matches in a review queue rather than guessing from similar hostnames.
| Evidence for an expected device | Coverage state |
|---|---|
| No assigned plan | Unassigned |
| Plan exists but no first recovery point | Awaiting first backup or overdue |
| Agent stale and recovery point too old | Protection stale |
| Recent recovery point, restore not yet checked | Backup present; restore evidence pending |
| Policy satisfied with required restore evidence | Protected within the declared scope |
| Approved, unexpired exception | Explicit exception |
| Inventory or provider feed stale | Unknown |
These are proposed operational states for your report, not a promise that a particular vendor exposes identical fields. Use the supported reporting API or exports available in the installed product and document where evidence is incomplete.
Measure recovery-point age, not agent optimism
An agent can be online while its plan is disabled, credentials expired, or destination unreachable. A plan can exist while its scope excludes a newly added volume. Compare required data scope against the actual completed backup, and record the recovery point that a restore operator can access.
Give devices different policies when their needs differ. A rarely connected laptop and a continuously changing database should not share an unexplained freshness threshold. Exceptions need a reason, accountable owner, and expiry date; otherwise “temporarily excluded” becomes permanent invisible exposure.
An independent deadline check (opens in a new tab) can watch the reconciliation job itself. The report should also display source ages so someone opening a cached copy can tell whether its conclusions remain current.
Test disappearance without causing data loss
Use a disposable test device. Remove its test policy, stop its test agent, deny access to the test destination, and omit it from one simulated provider export. The device must remain in the expected inventory and move into the correct exception state.
Then simulate an inventory-feed failure. A system that announces full coverage after the feed returns no devices is measuring report success rather than protection. Keep the previous inventory with an explicit stale marker until a complete current inventory can be established.
Make onboarding and retirement explicit
Onboarding closes only after the first appropriate recovery point and required restore check. Retirement requires a recorded decision about retained backups, retention, and ownership before the device disappears from the expected set. Review the unresolved queue by risk and age, with links to device and backup records. The useful daily question is which expected device lacks evidence now, not how many jobs happened to finish overnight.
Did this help?
Your answer helps us improve this guide. We save only the page and your choice for 30 days.
No name, email, or incident details are requested.
Sources
Vendor facts change. Each source below shows the date this page last checked it.
- Configuring deadline and heartbeat checks — Healthchecks. Checked 12 September 2026.
Related
One practical idea, occasionally
The On-Call Brief: short field notes, templates, and operational lessons.
Follow the field guide
Email subscriptions are not open yet. Read new guides in your feed reader — no email address needed.
Subscribe with RSS