Skip to content

Reliability Metrics · MTTD

Mean time to detect

Mean time to detect is the average time from failure start to detection.

What is mean time to detect?

Mean time to detect describes the average time from failure start to detection. In reliability work, the label is useful only when it maps to a measurable check, a clear owner, and a next action when expectations break. Without that operational meaning, the phrase becomes decoration in dashboards and status updates.

Why it matters

Mean time to detect matters because teams need a precise shared meaning for the average time from failure start to detection. Vague language turns incidents into arguments about words instead of fixes.

When everyone uses the same definition, alerts, status updates, and post-incident reviews stay aligned.

How it works

In practice, the average time from failure start to detection shows up as a concrete signal you can measure or communicate. Operators define what good looks like, watch for deviations, and record what happened when expectations break.

The useful version of mean time to detect is operational: it changes who gets notified, what customers see, or which metric a team reviews after an incident.

Practical example

Imagine a team operating around MTTD of eight minutes across incidents. When observed behavior stops matching the definition of mean time to detect, the team treats that change as a reliability event with a clear owner and next step.

Common misconception

MTTD starts when someone opens a laptop

That reading usually collapses distinct ideas into one slogan. Keep mean time to detect tied to observable behavior so the definition stays useful under pressure.

How Fajita handles this

Faster detection usually means better external monitoring coverage and alert delivery.

MTTD

MTTD = Total time from failure start to detection ÷ Number of incidents
  • Failure start can be hard to know precisely.
  • Use consistent clocks and incident markers.

Related documentation

Was this definition clear?