Skip to content

ProductHeartbeats

Know when the job that should have run did not.

Cron jobs, backups and workers ping a URL when they run. Observer catches the run that failed, hung or never started, and alerts you like any other metric.

A week of a nightly backup

Period
1 day
Grace
15 min
Deadline
1 day 15 min

Healthy under 87,300 s

Unhealthy over 87,300 s

The value is the seconds since the last success. A failed run is unhealthy whatever its age.

Monday and Tuesday succeed in about four minutes. Wednesday exits with code 2 and the check stays failed until Thursday's run succeeds. On Friday cron never fires: at 02:19 the check goes late, and it recovers when the job is re-run by hand at 08:51. Saturday and Sunday succeed.

No run since Thursday. Late at 02:19, the deadline.

Console / Metrics / Nightly backup

Check

LateEvaluated every minute and on each ping.

Last ping

1d agoSuccess, ran 4 min 31 s

Next success due

overdue, 50m agoLast success plus period and grace

Schedule

every 1 dayGrace 15 min, max run 30 min
Sample pings, judged by Observer's heartbeat evaluator. The cells below the week are the console's, as they read at the moment you pick.

What it catches

Every way a job goes quiet, named.

Each state carries its own reason, so the alert says what happened: the run failed, it hung, or it never started.

Late

No success by the deadline.

The deadline is the last success plus the period and a grace you set. A job that stops running reads late instead of quiet.

Failed

The job said so.

Ping /fail or send a non-zero exit code. The check stays failed until the next successful run, and up to 10 KB of output posted with the ping is kept beside it.

Running too long

Started, never finished.

Ping /start when the run begins. A run open longer than its maximum turns the check unhealthy before the deadline arrives.

Waiting

No data until the first ping.

A new check reads No data, never unhealthy, until the job reports for the first time.

Setup

One line in the job.

Every check gets its own URL. The console writes the call for wherever the job runs, with the URL filled in, then waits for the first ping.

Console / Metrics / Nightly backup

Check

Waiting for first pingNothing is evaluated until it arrives.

Last ping

–No pings yet

Next success due

–Last success plus period and grace

Schedule

every 1 dayGrace 15 min, max run 30 min

Ping URL

Your job calls this when it finishes. Treat it like a password: anyone holding it can report runs for this check.

Waiting for the first ping

Run your job once, or call the URL with curl. This page checks for it every few seconds; the check stays No data, never unhealthy, until the first ping lands.

Ping URL
https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1

Append /start when a run begins, /fail or /<exit code> when it ends. POST a body (up to 10 KB) to attach output.

Send pings

Copy a pattern for where your job runs. Pings never block or fail the job.

Shell
curl -fsS -m 10 --retry 3 -o /dev/null https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1/start
if /usr/local/bin/backup.sh > /tmp/nightly-backup.log 2>&1; then
  curl -fsS -m 10 --retry 3 -o /dev/null https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1
else
  # Report the exit code and the last 10 KB of output.
  code=$?
  tail -c 10000 /tmp/nightly-backup.log | curl -fsS -m 10 --retry 3 -o /dev/null --data-binary @- https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1/$code
fi

One URL

The URL is the credential.

/api/heartbeat/<token>

GET, POST or HEAD all count. Nothing else to sign or configure, and the console rotates the URL when it leaks.

Outcomes

Start, success, failure.

Append /start when a run begins and /fail or the exit code when it ends. A bare ping is a success.

As code

Or in observer.yaml.

Declare checks next to the rest of your metrics with source_type heartbeat and apply them from CI.

Private networks

Jobs with no way out ping the agent.

A job that cannot reach the internet pings the Observer agent next to it instead. The agent forwards each ping over the connection it already has.

Console / Metrics / Ledger export

Check

UpEvaluated every minute and on each ping.

Last ping

29m agoSuccess, ran 41 s

Next success due

in 23hLast success plus period and grace

Schedule

every 1 dayGrace 15 min, max run 30 min

Recent pings

What your job reported, newest first.

KindDurationExitMessageTime
Successvia agent41 s––2026-10-04 23:30:44 GMT+0
Startvia agent–––2026-10-04 23:30:03 GMT+0
Successvia agent39 s––2026-10-03 23:30:39 GMT+0
Startvia agent–––2026-10-03 23:30:00 GMT+0

Exact lateness

The time it reached the agent.

Each ping waits in the agent's own queue on disk and is recorded at the moment the agent received it, so a delay of up to an hour on the way to Observer never makes a run look late.

Scoped

Only your organization's checks.

The agent forwards over its own key, and Observer only accepts pings for checks in the agent's organization.

If the agent stops

You still hear about it.

agent 1.6.0

Lateness is judged in Observer, not on the agent. If the agent goes offline its relayed checks go late, and the agent offline alert fires.

Then

The same path as every metric.

A heartbeat is a metric. Everything Observer does with a metric, it does with a missed job.

Everywhere a metric goes

On pages, in alerts, in SLOs.

A check is a metric whose value is the time since the last success. Put it on a status page, route its alerts, give it an SLO.

Recovery

Recovered on the ping.

A successful ping is evaluated as it arrives, so a late or failed check recovers then, not on the next tick.

Plans

Each check counts as one metric.

Every plan includes heartbeats, with nothing extra to buy. Free covers 10 metrics, Starter 50, Pro 500.

Let the telemetry write your status page.

PlanFree
CardNot needed
First public metricAbout 10 min
All three healthyAPICheckoutWebhooks