Skip to content

Scheduled Jobs

Heartbeat monitoring

Heartbeat monitoring is expecting a periodic signal from a job and alerting when the signal is late or missing.

What is heartbeat monitoring?

Heartbeat monitoring describes expecting a periodic signal from a job and alerting when the signal is late or missing. In reliability work, the label is useful only when it maps to a measurable check, a clear owner, and a next action when expectations break. Without that operational meaning, the phrase becomes decoration in dashboards and status updates.

Why it matters

Heartbeat monitoring matters because teams need a precise shared meaning for expecting a periodic signal from a job and alerting when the signal is late or missing. Vague language turns incidents into arguments about words instead of fixes.

When everyone uses the same definition, alerts, status updates, and post-incident reviews stay aligned.

How it works

In practice, expecting a periodic signal from a job and alerting when the signal is late or missing shows up as a concrete signal you can measure or communicate. Operators define what good looks like, watch for deviations, and record what happened when expectations break.

The useful version of heartbeat monitoring is operational: it changes who gets notified, what customers see, or which metric a team reviews after an incident.

Practical example

Imagine a team operating around backup job pings a heartbeat URL each night. When observed behavior stops matching the definition of heartbeat monitoring, the team treats that change as a reliability event with a clear owner and next step.

Common misconception

Heartbeat monitoring is the same as website monitoring

That reading usually collapses distinct ideas into one slogan. Keep heartbeat monitoring tied to observable behavior so the definition stays useful under pressure.

How Fajita handles this

Fajita heartbeat monitors watch for expected pings. See heartbeat monitoring.

Frequently asked questions

What should call the heartbeat URL?

The job itself, after successful work completes, or at a stage you define clearly.

What if a job runs long?

Use a grace period that matches realistic runtime so late work does not page incorrectly.

Related documentation

Was this definition clear?