What Is Cron Monitoring? Dead Man's Switch Explained
The cron job that didn't run is the one you'll never know about.
Cron monitoring uses a dead man's switch to catch scheduled jobs that fail silently. Learn how heartbeats and grace periods work.
The cron job that didn't run is the one you'll never know about.
Cron monitoring uses a dead man's switch to catch scheduled jobs that fail silently. Learn how heartbeats and grace periods work.
Cron jobs fail silently more often than they fail loudly. A job that crashes mid-run, a crontab line that got commented out during debugging and never restored, a server that ran out of disk so the job couldn't write its output, a dependency that changed and broke the script — none of these send you an alert. You find out weeks later when a report didn't generate, a backup turned out to be stale, or a cache stopped refreshing.
The reason cron jobs fail silently is that cron itself has no concept of "this job should have run by now." Cron runs things on a schedule; it doesn't notify you when it doesn't. Traditional log monitoring doesn't help either — it can tell you a job ran and errored, but it can't tell you a job was supposed to run and didn't. That absence-of-signal problem is what cron monitoring solves.
Cron monitoring (also called heartbeat monitoring or dead man's switch) flips the model. Instead of watching for errors, it watches for the absence of success. Your cron job pings a unique URL at the end of every successful run. The monitoring service expects that ping within a time window — say, every 24 hours, plus a grace period. If the ping doesn't arrive in time, it alerts you.
This catches every failure mode: the job crashed, the server is down, the crontab was deleted, the job ran but exited non-zero, the job ran but took too long. All of them result in no ping, and no ping means an alert. The grace period absorbs normal variance — a job that usually takes 5 minutes but occasionally takes 20 shouldn't page you every time it runs long.
The grace period is the buffer between "the ping was due" and "we alert." Set it generously enough that normal runtime variance doesn't cause false alerts, but tightly enough that you learn about a real failure promptly. For a job that runs hourly and takes 2 minutes, a 10-minute grace period is reasonable. For a nightly backup that takes 30-90 minutes, give it 2 hours.
SurePing's heartbeat monitors let you set the expected interval and a grace period per job. Every plan, including Free, includes heartbeat monitoring — you create a monitor, get a unique ping URL, and add a `curl` call to the end of your cron job. If the ping stops, you're alerted through your configured webhooks.
SurePing includes this monitor type on every plan — including the free tier.