Skip to content

Monitoring

Service health

Service health is the overall condition of a service relative to expected behavior. Teams use the term to keep checks, alerts, and reviews precise.

What is service health?

Service health describes the overall condition of a service relative to expected behavior. In reliability work, the label is useful only when it maps to a measurable check, a clear owner, and a next action when expectations break. Without that operational meaning, the phrase becomes decoration in dashboards and status updates.

Why it matters

Service health matters because teams need a shared, precise meaning for the overall condition of a service relative to expected behavior. Vague language turns incidents into arguments about words instead of fixes.

When everyone uses the same definition, alerts, status updates, and post-incident reviews stay aligned.

How it works

In practice, the overall condition of a service relative to expected behavior shows up as a concrete signal you can measure or communicate. Operators define what good looks like, watch for deviations, and record what happened when expectations break.

The useful version of service health is operational: it changes who gets notified, what customers see, or which metric a team reviews after an incident.

Practical example

Imagine a team running checks against operational versus degraded for api.example.com. When the observed behavior stops matching the definition of service health, the team treats that change as a reliability event with a clear owner and next step.

Common misconception

Service health is a single binary up or down bit

That reading usually collapses distinct ideas into one slogan. Keep service health tied to observable behavior so the definition stays useful under pressure.

How Fajita handles this

Fajita expresses health through monitor state and status-page components, including degraded and down.

Related documentation

Was this definition clear?