Skip to content

Category

Reliability Metrics

Reliability metrics quantify how often a service works as expected and how quickly teams detect and recover from failure.

Why this category matters

Shared numbers let teams compare periods, set goals, and explain impact. Without clear definitions, uptime and MTTR become marketing language instead of operational tools.

Recommended learning order

  1. UptimeUptime is the time a service was available during a period.
  2. Uptime percentageUptime percentage is availability expressed as a percentage of eligible monitored time.
  3. AvailabilityAvailability is the share of time a service was able to fulfill its intended function.
  4. Mean time to detectMean time to detect is the average time from failure start to detection.
  5. Mean time to recoveryMean time to recovery is the average time from failure to restored service.

Foundational terms

Advanced terms

All terms in Reliability Metrics