Monitors

Heartbeat & cron monitors

Monitor cron jobs, backups and scheduled tasks with a heartbeat URL. Ping format, schedules, grace periods, failure signals and copy-paste examples.

A heartbeat monitor (Scheduled job (cron) heartbeat in the type list) works the other way round from every other monitor type. SutramX does not call your service. Your job calls a unique heartbeat URL each time it finishes. If a run is due and no ping arrives within the grace period, SutramX opens an incident and alerts you. Your job can also report a failure on purpose.

Use heartbeats for anything that runs on a schedule and can fail silently: cron jobs, backups, queue workers, ETL pipelines, scheduled CI workflows and Kubernetes CronJobs.

How it works#

  1. You create a heartbeat monitor and give it your job's cron expression and the time zone it runs in.
  2. SutramX gives you a heartbeat URL like https://api.sutramx.com/heartbeat/<token>.
  3. Your job requests that URL when it succeeds, or adds ?status=fail when it fails.
  4. SutramX works out from the schedule when the next run is due. If that time plus the grace period passes with no ping, the monitor goes Down and an incident opens.
  5. The next successful ping marks the monitor Up again and resolves the incident.

Heartbeat monitors have no target to probe, so they are not checked from probe regions. SutramX evaluates them itself, from one source, every couple of minutes. The multi-region quorum described in Regions & confirmation does not apply to them.

Heartbeat: silence is the alarm
An hourly job with the default 6-minute grace period and Missed runs allowed at 0 (the default). A late ping inside the grace period is fine; when the grace period passes with no ping, the monitor goes Down.

Create a heartbeat monitor#

  1. Go to Monitors → New monitor.
  2. Set What do you want to check? to Scheduled job (cron) heartbeat.
  3. Enter the Cron expression of your job, for example 0 2 * * * for 02:00 every day, and pick the Time zone your server's crontab uses. New monitors start with your profile time zone. The form previews the Next expected runs and shows how often the job Runs.
  4. Enter a Name, for example Nightly database backup.
  5. Optionally set a Grace period (minutes) and Missed runs allowed. Leave them empty to use the defaults.
  6. Optionally choose a Group and add Tags.
  7. Click Create monitor.

The next screen shows the heartbeat URL and a ready-to-paste crontab line. Until the first ping arrives, the monitor's checks say "Waiting for first heartbeat". If the first scheduled run passes its deadline without a ping, the monitor goes Down.

Heartbeat monitors have no check interval, location picker, Test button or advanced options. Those settings only apply to monitors that SutramX calls.

Settings#

FieldWhat it doesDefault / limits
Cron expressionYour job's schedule, used to work out when each ping is dueRequired. 5 fields: minute, hour, day of month, month, day of week
Time zoneThe time zone the cron expression is read in, as an IANA name such as America/New_York or Asia/KolkataUTC
Grace period (minutes)Extra time after a run is due before a missing ping counts as missedOptional, 0–10,080 minutes (7 days). Empty = 10% of the time between runs, at least 1 minute
Missed runs allowedConsecutive missed runs to tolerate before alerting, for jobs that sometimes skip a run0 (alert on the first miss). 0–5
GroupPuts the monitor in a monitor groupNo group
TagsLabels for filteringUp to 20 tags

Through the API, the same settings live in config: cron_expression, timezone (an IANA name; default UTC; fixed offsets such as +05:30 are rejected because they ignore daylight saving), expected_interval_ms (the dashboard fills it in from the cron expression), grace_seconds (0–604,800) and missed_runs_allowed (0–5).

Cron expression syntax#

SutramX uses standard 5-field cron syntax, evaluated in the monitor's time zone:

text
┌───────── minute        (0-59)
│ ┌─────── hour          (0-23)
│ │ ┌───── day of month  (1-31)
│ │ │ ┌─── month         (1-12 or JAN-DEC)
│ │ │ │ ┌─ day of week   (0-7 or SUN-SAT; 0 and 7 are Sunday)
│ │ │ │ │
* * * * *
  • Wildcards (*), lists (1,15), ranges (1-5) and steps (*/15, 10-50/10, 5/15) are supported.
  • Month and weekday names are case-insensitive (MON-FRI, jan).
  • If you restrict both day of month and day of week, a day matches when either matches, like classic Vixie cron.
  • The API also accepts the macros @hourly, @daily, @midnight, @weekly, @monthly, @yearly and @annually. The dashboard form needs the 5-field form.
  • An expression that never matches a date in the next 5 years is rejected.
ScheduleExpression
Every 5 minutes*/5 * * * *
Every hour at :3030 * * * *
Every 6 hours0 */6 * * *
Daily at 02:000 2 * * *
Weekdays at 09:1515 9 * * 1-5
First day of every month0 0 1 * *

Daylight saving time is handled like standard cron:

  • Clocks go forward: a run scheduled in the skipped hour (for example 02:30 in New York on the March change) is expected at the moment the clocks change, 03:00. Several runs inside the gap count as one.
  • Clocks go back: a run in the repeated hour is expected once, at its first occurrence.

Grace period#

The grace period absorbs normal variation in how long a job takes. The default is 10% of the gap between two runs, and never less than 1 minute:

Job runs everyDefault grace
5 minutes1 minute
1 hour6 minutes
6 hours36 minutes
1 day2 hours 24 minutes

Ping at the end of the job and set the grace period to cover its longest normal run time. A job that takes 40 minutes and runs hourly needs a grace period of at least 40 minutes.

A ping that arrives a little early, before its scheduled minute, still counts for that run. Clock skew between your server and SutramX does not cause false alerts.

The heartbeat URL#

text
https://api.sutramx.com/heartbeat/<token>

The token is a long random string. It is the only credential, so no API key or header is needed. Treat the URL like a password.

RequestEffect
GET, POST or HEAD the URLRecords a successful run
Add ?status=failRecords a failed run. The monitor goes Down straight away
Add ?status=okRecords a successful run (same as no parameter)
Add ?duration_ms=12345Records how long the run took, in milliseconds. It shows as the check's response time

Accepted status values:

  • Success: ok, success, succeeded, up, pass, done
  • Failure: fail, failure, failed, down, error, err

Values are case-insensitive. duration_ms must be a number from 0 to 604,800,000 (7 days). Any request body is ignored, so you can safely POST logs or JSON.

Responses#

StatusBodyMeaning
200OKPing recorded. Also returned for a paused monitor, whose pings are ignored
400status must be "fail" or "ok"Unknown status value
400duration_ms must be a number between 0 and 604800000Bad duration_ms
404Unknown heartbeat URLThe token does not exist, for example after the URL was rotated or the monitor deleted
429Too many heartbeats for this monitor; slow downMore than 60 pings a minute for one URL
429Too many heartbeat requestsMore than 600 pings a minute from one IP address
500Heartbeat could not be recorded; retryTemporary problem. Retry the request

A job only needs to ping once per run. Use retries (for example curl --retry 3) so a brief network blip does not look like a missed run.

Failure signals#

There is no separate "start" signal. SutramX tracks completion only:

  • A success ping marks the run as done.
  • A failure ping (?status=fail) opens an incident immediately, even if the run was on time. The monitor stays Down until a later success ping arrives.
  • No ping by the deadline (after the allowed number of missed runs, if you set Missed runs allowed) opens an incident with the message No heartbeat received for the run expected at … (grace …s).

Rotate the URL#

If a heartbeat URL leaks, open the monitor and click Rotate URL in the Heartbeat panel. The old URL stops working immediately and returns 404. Update every job that pings it, or the next run will be reported as missed.

Through the API:

bash
curl -X POST https://api.sutramx.com/monitors/<monitor-id>/heartbeat/rotate \
  -H "Authorization: Bearer $SUTRAMX_API_KEY"
json
{ "heartbeat_url": "https://api.sutramx.com/heartbeat/<new-token>" }

The Heartbeat panel#

The monitor's page shows a Heartbeat panel with:

  • Heartbeat URL with a Copy button
  • Crontab example, a ready-to-paste line for your schedule
  • Last heartbeat, or "None received yet"
  • Next expected, with the grace period. It turns red ("Overdue") when the deadline has passed
  • Schedule (shown with its time zone) and Expected every
  • Rotate URL

Examples#

In each example, replace the URL with your own heartbeat URL. Keep it in an environment variable or secret rather than in source control.

crontab#

Ping only when the job succeeds (&&), with a 10-second timeout and 3 retries:

bash
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 https://YOUR-HEARTBEAT-URL > /dev/null

Report failures as well as successes:

bash
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL" > /dev/null || curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL?status=fail" > /dev/null

Bash script with duration and exit code#

bash
#!/usr/bin/env bash
set -u
HEARTBEAT_URL="https://YOUR-HEARTBEAT-URL"

start=$(date +%s%3N)
/usr/local/bin/backup.sh
exit_code=$?
duration=$(( $(date +%s%3N) - start ))

if [ "$exit_code" -eq 0 ]; then
  curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL?duration_ms=$duration" > /dev/null
else
  curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL?status=fail&duration_ms=$duration" > /dev/null
fi
exit "$exit_code"

date +%s%3N needs GNU date. On macOS, use $(($(date +%s) * 1000)) instead.

curl#

bash
# Success
curl -fsS -m 10 --retry 3 https://YOUR-HEARTBEAT-URL

# Failure
curl -fsS -m 10 --retry 3 "https://YOUR-HEARTBEAT-URL?status=fail"

# HEAD works too
curl -fsS -m 10 -I https://YOUR-HEARTBEAT-URL

Python#

Uses the requests package (pip install requests):

python
import os, time, requests

HEARTBEAT_URL = os.environ["HEARTBEAT_URL"]

def run_job():
    ...  # your work here

start = time.monotonic()
try:
    run_job()
    status = "ok"
except Exception:
    status = "fail"
    raise
finally:
    duration_ms = int((time.monotonic() - start) * 1000)
    try:
        requests.get(HEARTBEAT_URL, params={"status": status, "duration_ms": duration_ms}, timeout=10)
    except requests.RequestException:
        pass  # never let monitoring break the job

Node.js#

Node 18 and later include fetch:

ts
const HEARTBEAT_URL = process.env.HEARTBEAT_URL!;

async function ping(status: 'ok' | 'fail', durationMs: number) {
  const url = `${HEARTBEAT_URL}?status=${status}&duration_ms=${durationMs}`;
  try {
    await fetch(url, { signal: AbortSignal.timeout(10_000) });
  } catch {
    // never let monitoring break the job
  }
}

const start = Date.now();
try {
  await runJob();
  await ping('ok', Date.now() - start);
} catch (error) {
  await ping('fail', Date.now() - start);
  throw error;
}

GitHub Actions#

Ping at the end of a scheduled workflow. The if: failure() step reports a failed run. Store the URL as a repository secret named SUTRAMX_HEARTBEAT_URL.

yaml
name: Nightly export
on:
  schedule:
    - cron: '0 3 * * *'   # GitHub schedules are UTC too

jobs:
  export:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: ./scripts/export.sh

      - name: Heartbeat (success)
        if: success()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.SUTRAMX_HEARTBEAT_URL }}"

      - name: Heartbeat (failure)
        if: failure()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.SUTRAMX_HEARTBEAT_URL }}?status=fail"

Use the same cron expression in SutramX. GitHub can start scheduled workflows several minutes late under load, so set a grace period that allows for it, for example 30 minutes.

Kubernetes CronJob#

Store the URL in a Secret:

bash
kubectl create secret generic sutramx-heartbeat --from-literal=url='https://YOUR-HEARTBEAT-URL'

Then ping after the main command:

yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: nightly-report
spec:
  schedule: "0 2 * * *"
  timeZone: "Etc/UTC"
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      backoffLimit: 0
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: report
              image: your-registry/report:latest
              env:
                - name: HEARTBEAT_URL
                  valueFrom:
                    secretKeyRef:
                      name: sutramx-heartbeat
                      key: url
              command: ["/bin/sh", "-c"]
              args:
                - |
                  if /app/run-report; then
                    curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL"
                  else
                    curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL?status=fail"
                    exit 1
                  fi

The container image needs curl. Set timeZone (Kubernetes 1.27 and later) so the cluster schedule matches the UTC schedule in SutramX.

Create a heartbeat monitor with the API#

bash
curl -X POST https://api.sutramx.com/monitors \
  -H "Authorization: Bearer $SUTRAMX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Nightly database backup",
    "type": "cron",
    "config": {
      "cron_expression": "0 2 * * *",
      "grace_seconds": 1800
    },
    "tags": ["backups"]
  }'

Read the monitor back with GET https://api.sutramx.com/monitors/<id>. Heartbeat monitors include a heartbeat_url field. See the REST API for authentication and the full request format.

Pausing and maintenance#

  • While a heartbeat monitor is paused, pings still get 200 OK but are ignored. After you resume it, SutramX waits for the next scheduled run, so a run missed while paused does not trigger an alert.
  • A maintenance window holds back alerts for a missed run. If the job is still missing when the window ends, the alert goes out then.

Common questions#

How do I test a new heartbeat monitor? Run the job once by hand, or call the URL with curl. Last heartbeat in the Heartbeat panel updates within seconds.

My job runs at irregular times. Use the widest cron expression that covers when it can run, with a generous grace period. For example, a job that runs "sometime each night" can use 0 0 * * * with a grace period of 8 hours.

Can one URL be used by several jobs? Each run's ping counts for the schedule, so several jobs sharing a URL hide each other's failures. Create one monitor per job.

What happens if the job pings twice in one run? Nothing bad. Extra pings just update the last-heartbeat time. Keep below 60 pings a minute per URL.

How quickly am I alerted about a missed run? Shortly after the run's due time plus grace. SutramX re-evaluates heartbeat monitors every couple of minutes, then your normal alert routing applies.

Heartbeats are not arriving. See Troubleshooting.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.