Category
Incidents
Incidents are verified periods of degraded or failed service that a team tracks from detection through recovery and review.
Why this category matters
A failed check is a signal. An incident is the shared record of what broke, who is working it, and when the service returned to a healthy state.
Recommended learning order
- IncidentIncident is a tracked period of degraded or failed service.
- Incident verificationIncident verification is confirming a failure is real before treating it as an incident.
- Incident severityIncident severity is a ranked label for how bad an incident is.
- Recovery confirmationRecovery confirmation is evidence that a service has returned to healthy behavior after a failure.
- Post-incident reviewPost-incident review is a blameless review of detection, response, and prevention after an incident.
Foundational terms
- IncidentIncident is a tracked period of degraded or failed service.
- Incident managementIncident management is the process of detecting, coordinating, resolving, and reviewing service failures.
- Incident verificationIncident verification is confirming a failure is real before treating it as an incident.
- False positiveFalse positive is an alert or incident that fired when the service was actually fine.
- Recovery confirmationRecovery confirmation is evidence that a service has returned to healthy behavior after a failure.
Advanced terms
- FlappingFlapping is rapid oscillation between healthy and failed states.
- Incident reopeningIncident reopening is returning a resolved incident to an active state when failure returns.
- Root-cause analysisRoot-cause analysis is structured investigation into why an incident happened.
- Post-incident reviewPost-incident review is a blameless review of detection, response, and prevention after an incident.
All terms in Incidents
- Degraded performance
- False negative
- False positive
- Flapping
- Incident
- Incident acknowledgment
- Incident assignment
- Incident detection
- Incident management
- Incident recap
- Incident reopening
- Incident resolution
- Incident response
- Incident severity
- Incident timeline
- Incident verification
- Major outage
- Partial outage
- Post-incident review
- Recovery confirmation
- Root-cause analysis
- Service disruption