Cron job monitoring uses a heartbeat, also called a dead man's switch: the job requests a unique URL each time it finishes, and the monitor alerts you when no ping arrives by the scheduled time plus a grace period, or when the job reports a failure. It catches jobs that crash, hang, never start or quietly stop running.
Why do cron jobs fail silently?
Most monitoring watches things that answer requests: a website, an API, a port. A scheduled job answers nothing. When it fails, the evidence lands in a log file nobody reads, in the local mail spool that cron's MAILTO setting writes to, or nowhere at all. When it does not run in the first place, there is no evidence of any kind.
The ways a scheduled job goes wrong fall into a handful of patterns, and it is worth knowing them because a good monitor has to catch all of them with one mechanism.
- It never startedThe cron daemon was stopped, the server was rebuilt from an image without the crontab, a typo broke the crontab file, or the job moved to a new host and the old line was deleted before the new one was added.
- It started in the wrong environmentCron runs commands with a minimal environment: a short PATH, no shell profile and a different working directory. A script that works in your terminal fails under cron with "command not found" or a missing credentials file.
- It started and crashedA full disk, an expired API token, a renamed database column or an out-of-memory kill. The job exits with an error, and unless something is watching the exit code, that is the end of the story.
- It started and hungA network call with no timeout, a stuck database query, or a lock held by an earlier run. A hung job never exits, so it never reports success or failure.
- It ran and did nothing usefulThe job exited cleanly, but the backup file is empty or the export wrote zero rows. From the outside it looks like a success.
- Runs overlappedA run took longer than the interval, so the next one started on top of it. The two compete for the same files or locks and one or both fail.
What is a heartbeat (dead man's switch) monitor?
A heartbeat monitor reverses the usual direction of a check. Instead of the monitor calling your service, your job calls the monitor: it requests a unique URL each time it completes. The monitor knows the job's schedule, so it knows when the next ping is due. If that time plus a grace period passes with no ping, it alerts. The name dead man's switch comes from trains and machinery that stop unless the operator keeps actively holding a control. Silence is the alarm.
This one design covers almost every failure above. A job that never started, crashed before the end, or is still stuck all produce the same symptom: no ping by the deadline. The only rule is that the ping must be sent after the work is done and only when it succeeded.
If the ping is the first line of your script, a job that crashes one line later still reports healthy. Put it after the last step that matters, and make it conditional on that step's success.
Late, failed or running too long: three different signals
A heartbeat can tell you three different things, and it helps to design for each one deliberately.
- FailedThe job ran and knows it failed: a non-zero exit code or a caught exception. It should say so with an explicit failure ping, so you are alerted straight away instead of waiting for the grace period to run out.
- Late or missingNo ping by the deadline. The job did not start, crashed without reaching the failure ping, or is still running. You know something is wrong, and the job logs tell you which.
- Running too longSome heartbeat tools accept a separate start ping and alert when the gap between start and finish exceeds a limit. Without a start signal, a run that overruns its grace period shows up as a missed run. That is usually the right outcome: a backup that has not finished by 04:00 and one that never started are both a backup you do not have.
Turn hangs into failures. Wrap the command in a time limit, so a stuck run is killed and reports failure instead of hanging forever: "timeout 2h /usr/local/bin/backup.sh" stops the script after two hours and exits with status 124. Combined with a failure ping, a hang becomes a prompt, specific alert rather than a vague missed run.
Record how long each run took, even when it succeeds. A nightly job that took 20 minutes last month and 55 minutes this week is heading for an overlap with the next run long before it actually fails, and a duration trend shows that weeks in advance.
How long should the grace period be?
The grace period is the extra time after a run is due before a missing ping counts as missed. Because you ping at the end of the job, it has to cover the job's run time as well as any delay in starting it. Too short and you get false alerts on slow nights; too long and you hear about real failures late.
- Start from the longest normal run time, not the average. Look at a few weeks of durations if you have them.
- Add scheduler delay. Hosted schedulers can start jobs late under load; GitHub Actions, for example, can start scheduled workflows several minutes after the scheduled time.
- Add headroom for known slow runs, such as month-end batches or the first run after a large data import.
- Keep it shorter than the gap between runs, so a missed run is flagged before the next one is due.
A worked example: a backup is scheduled for 02:00 and normally finishes between 02:10 and 02:25. A 45-minute grace period puts the deadline at 02:45. A normal night pings well before that; a hung or missing run is reported shortly after 02:45, hours before anyone needs the backup.
How do you add a heartbeat to a crontab?
The pattern is the same for every scheduler: run the job, and only if it succeeds, request the heartbeat URL. Treat the URL like a password and keep it out of source control. In a crontab, you can define it once at the top with a line such as HEARTBEAT_URL=https://YOUR-HEARTBEAT-URL, which standard cron implementations pass to the commands below it.
- Ping only on success: 0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL" > /dev/null
- Report failures too: 0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL" > /dev/null || curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL?status=fail" > /dev/null
- Kill hung runs: 0 2 * * * timeout 2h /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL" > /dev/null
- Prevent overlapping runs: */15 * * * * flock -n /tmp/sync.lock /usr/local/bin/sync.sh && curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL" > /dev/null
The curl flags matter. -f makes curl treat an HTTP error as a failure, -sS keeps it quiet but still prints errors, -m 10 caps the request at ten seconds so a slow network cannot hold up your job, and --retry 3 retries transient errors so a brief network blip does not look like a missed run.
In "job && ping || fail-ping", the failure ping also fires if the success ping itself fails after its retries. For anything beyond a one-liner, use a small wrapper script that runs the job, captures its exit code and duration, sends exactly one ping (success or failure), and exits with the job's own exit code.
Monitoring backups, queues, CI and Kubernetes jobs
- BackupsPing only after checking the result: the file exists, it is not empty, and the upload to off-site storage completed. A backup that has never been restored is a hope, so schedule a periodic restore test as its own job with its own heartbeat.
- Queue workersA long-running worker has no natural schedule, so give it one. Ping from the worker loop every few minutes after a successful poll of the queue, and set the monitor to expect that cadence. A heartbeat proves the worker is alive; it does not prove it is keeping up, so watch queue depth or the age of the oldest message separately.
- ETL and data pipelinesPing at the end of the final stage, and send a failure ping from the pipeline's error handler. One monitor per pipeline keeps alerts specific.
- Scheduled CI workflowsAdd a final step that pings on success and a step that runs only on failure (if: failure() in GitHub Actions) to send the failure ping. Store the URL as a repository secret.
- Kubernetes CronJobsPing after the main command inside the container (the image needs curl), set concurrencyPolicy: Forbid to stop overlapping runs, set timeZone so the schedule is unambiguous, and use activeDeadlineSeconds on the job to kill runs that hang.
Time zones and daylight saving time
A heartbeat monitor reads your cron expression in a time zone, and it has to be the same one your scheduler uses. Many servers run cron in UTC even when the team does not, and a mismatch makes every ping look hours early or late.
Daylight saving time adds two awkward nights a year in zones that observe it. When clocks go forward, a run scheduled in the skipped hour has no real time to run at; when they go back, the repeated hour happens twice. Schedulers handle these differently. The simplest fix is to schedule important jobs in UTC, or outside the 01:00 to 03:00 local window where most changes happen.
How heartbeat monitors work in SutramX
SutramX has a Cron job heartbeat monitor type. Here is exactly how it behaves, so you can set it up with the right expectations.
- ScheduleYou give it your job's 5-field cron expression and the IANA time zone it runs in, such as UTC or Asia/Kolkata, and the form previews the next three runs. To check an expression first, use the free cron expression parser.
- Grace periodOptional, from 0 to 10,080 minutes (7 days). Left empty, it defaults to 10% of the time between runs, never less than 1 minute: 6 minutes for an hourly job, 2 hours 24 minutes for a daily one.
- SignalsA GET, POST or HEAD to the heartbeat URL records a successful run. Adding ?status=fail records a failed run and opens an incident immediately. Adding ?duration_ms= records how long the run took, shown as the check's response time.
- No start signalSutramX tracks completion only. A run that overruns its grace period is reported as missed, with the message "No heartbeat received for the run expected at …". Use a timeout in the job if you want a hang reported as a failure sooner.
- EvaluationHeartbeats have no target to probe, so they are not checked from probe regions. SutramX evaluates them itself every couple of minutes, and multi-region confirmation does not apply. The next successful ping resolves the incident.
- Pauses and maintenancePings to a paused monitor are accepted but ignored, and after resuming SutramX waits for the next scheduled run. A maintenance window holds back the alert for a missed run until the window ends.
- Security and limitsThe token in the URL is the only credential. If it leaks, Rotate URL invalidates the old one at once. Each URL accepts up to 60 pings a minute. Create one monitor per job: several jobs sharing a URL hide each other's failures.
Cron and heartbeat monitors are included on every SutramX plan, Free included, and missed runs alert through the same channels and escalation as any other monitor. Full settings and copy-paste examples for Bash, Python, Node.js, GitHub Actions and Kubernetes are in the heartbeat monitor documentation.
A checklist for every scheduled job
- One heartbeat monitor per job, named after what the job protects ("Nightly database backup", not "cron 3").
- The ping is the last step and only runs on success; failures send an explicit failure ping.
- The job has a time limit, so a hang becomes a failure.
- The grace period covers the longest normal run plus scheduler delay.
- The monitor's time zone matches the scheduler's.
- The heartbeat URL lives in a secret or environment variable, not in the repository.
- Alerts go to a channel someone reads, with escalation for jobs that matter overnight.