Skip to content

Incidents

Partial outage

Partial outage is a failure that affects some features, regions, or customers but not the whole service.

What is partial outage?

Partial outage describes a failure that affects some features, regions, or customers but not the whole service. In reliability work, the label is useful only when it maps to a measurable check, a clear owner, and a next action when expectations break. Without that operational meaning, the phrase becomes decoration in dashboards and status updates.

Why it matters

Partial outage matters because teams need a precise shared meaning for a failure that affects some features, regions, or customers but not the whole service. Vague language turns incidents into arguments about words instead of fixes.

When everyone uses the same definition, alerts, status updates, and post-incident reviews stay aligned.

How it works

In practice, a failure that affects some features, regions, or customers but not the whole service shows up as a concrete signal you can measure or communicate. Operators define what good looks like, watch for deviations, and record what happened when expectations break.

The useful version of partial outage is operational: it changes who gets notified, what customers see, or which metric a team reviews after an incident.

Practical example

Imagine a team operating around billing API down while marketing site stays up. When observed behavior stops matching the definition of partial outage, the team treats that change as a reliability event with a clear owner and next step.

Common misconception

Partial outage means only one server restarted

That reading usually collapses distinct ideas into one slogan. Keep partial outage tied to observable behavior so the definition stays useful under pressure.

How Fajita handles this

Describe partial outages clearly on status pages so customers know what still works.

Was this definition clear?