ProductHeartbeats
Know when the job that should have run did not.
Cron jobs, backups and workers ping a URL when they run. Observer catches the run that failed, hung or never started, and alerts you like any other metric.
A week of a nightly backup
- Period
- 1 day
- Grace
- 15 min
- Deadline
- 1 day 15 min
Healthy under 87,300 s
Unhealthy over 87,300 s
The value is the seconds since the last success. A failed run is unhealthy whatever its age.
No run since Thursday. Late at 02:19, the deadline.
Check
LateEvaluated every minute and on each ping.Last ping
1d agoSuccess, ran 4 min 31 sNext success due
overdue, 50m agoLast success plus period and graceSchedule
every 1 dayGrace 15 min, max run 30 minWhat it catches
Every way a job goes quiet, named.
Each state carries its own reason, so the alert says what happened: the run failed, it hung, or it never started.
Late
No success by the deadline.
The deadline is the last success plus the period and a grace you set. A job that stops running reads late instead of quiet.
Failed
The job said so.
Ping /fail or send a non-zero exit code. The check stays failed until the next successful run, and up to 10 KB of output posted with the ping is kept beside it.
Running too long
Started, never finished.
Ping /start when the run begins. A run open longer than its maximum turns the check unhealthy before the deadline arrives.
Waiting
No data until the first ping.
A new check reads No data, never unhealthy, until the job reports for the first time.
Setup
One line in the job.
Every check gets its own URL. The console writes the call for wherever the job runs, with the URL filled in, then waits for the first ping.
Check
Waiting for first pingNothing is evaluated until it arrives.Last ping
–No pings yetNext success due
–Last success plus period and graceSchedule
every 1 dayGrace 15 min, max run 30 minPing URL
Your job calls this when it finishes. Treat it like a password: anyone holding it can report runs for this check.
Run your job once, or call the URL with curl. This page checks for it every few seconds; the check stays No data, never unhealthy, until the first ping lands.
https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1Append /start when a run begins, /fail or /<exit code> when it ends. POST a body (up to 10 KB) to attach output.
Send pings
Copy a pattern for where your job runs. Pings never block or fail the job.
curl -fsS -m 10 --retry 3 -o /dev/null https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1/start
if /usr/local/bin/backup.sh > /tmp/nightly-backup.log 2>&1; then
curl -fsS -m 10 --retry 3 -o /dev/null https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1
else
# Report the exit code and the last 10 KB of output.
code=$?
tail -c 10000 /tmp/nightly-backup.log | curl -fsS -m 10 --retry 3 -o /dev/null --data-binary @- https://use.observer/api/heartbeat/hb7Qx2mK9pLw4RtZ8vNc3Fd1/$code
fiOne URL
The URL is the credential.
/api/heartbeat/<token>GET, POST or HEAD all count. Nothing else to sign or configure, and the console rotates the URL when it leaks.
Outcomes
Start, success, failure.
Append /start when a run begins and /fail or the exit code when it ends. A bare ping is a success.
As code
Or in observer.yaml.
Declare checks next to the rest of your metrics with source_type heartbeat and apply them from CI.
Private networks
Jobs with no way out ping the agent.
A job that cannot reach the internet pings the Observer agent next to it instead. The agent forwards each ping over the connection it already has.
Check
UpEvaluated every minute and on each ping.Last ping
29m agoSuccess, ran 41 sNext success due
in 23hLast success plus period and graceSchedule
every 1 dayGrace 15 min, max run 30 minRecent pings
What your job reported, newest first.
| Kind | Duration | Exit | Message | Time |
|---|---|---|---|---|
| Successvia agent | 41 s | – | – | 2026-10-04 23:30:44 GMT+0 |
| Startvia agent | – | – | – | 2026-10-04 23:30:03 GMT+0 |
| Successvia agent | 39 s | – | – | 2026-10-03 23:30:39 GMT+0 |
| Startvia agent | – | – | – | 2026-10-03 23:30:00 GMT+0 |
Exact lateness
The time it reached the agent.
Each ping waits in the agent's own queue on disk and is recorded at the moment the agent received it, so a delay of up to an hour on the way to Observer never makes a run look late.
Scoped
Only your organization's checks.
The agent forwards over its own key, and Observer only accepts pings for checks in the agent's organization.
If the agent stops
You still hear about it.
agent 1.6.0Lateness is judged in Observer, not on the agent. If the agent goes offline its relayed checks go late, and the agent offline alert fires.
Then
The same path as every metric.
A heartbeat is a metric. Everything Observer does with a metric, it does with a missed job.
Everywhere a metric goes
On pages, in alerts, in SLOs.
A check is a metric whose value is the time since the last success. Put it on a status page, route its alerts, give it an SLO.
Recovery
Recovered on the ping.
A successful ping is evaluated as it arrives, so a late or failed check recovers then, not on the next tick.
Plans
Each check counts as one metric.
Every plan includes heartbeats, with nothing extra to buy. Free covers 10 metrics, Starter 50, Pro 500.