The definitions
- MTTR — mean time to recoveryTotal downtime divided by the number of incidents: the average length of an outage from first user impact to full recovery.
- MTBF — mean time between failuresWorking time divided by the number of incidents: on average, how long the service runs before the next failure.
- MTTD — mean time to detectTotal time from first impact to first alert, divided by incidents. It is the part of MTTR that monitoring alone can shrink.
- AvailabilityOne minus downtime over the period, as a percentage. The number your SLA is written in.
Averages hide the shape
Two services can share an MTTR of 25 minutes while behaving completely differently: one has a single two-hour outage a quarter, the other a five-minute blip every week. The first needs faster rollback and better runbooks; the second needs root-cause work. Look at MTBF next to MTTR, and keep the incident list, not just the averages.
If MTTD is a third of MTTR, a third of every outage is spent waiting for someone to notice. A one-minute check interval with multi-region confirmation typically brings detection under two minutes; the check interval planner shows the trade-off against the load you add.
What to count as an incident
Count failures that users could see. A monitor flapping for thirty seconds on one region is not an incident; a checkout returning errors for ten minutes is. Be consistent across periods so the trend means something, and measure downtime from the user's side rather than from when the on-call engineer acknowledged the page.