Free tool

MTTR & MTBF calculator

Four numbers that describe how reliable a service really is: how often it fails, how long it stays down, how quickly you find out, and what that adds up to.

Count user-facing failures only. Flapping checks and false alarms are not incidents.
From first user impact to full recovery, summed across all incidents.
From first impact to the first alert, summed. Leave blank to skip MTTD.
MTTR — mean time to recovery
25 min

A typical incident took this long from first impact to recovery. About 5.0 min of that passed before anyone was alerted.

MTBF15 days
MTTD5.0 min
Availability99.884%
Incidents / month2.0
Which number to work on

A high MTTR with a low MTBF means rare, long outages: invest in runbooks, rollbacks and faster alerts. A low MTTR with a short MTBF means frequent small failures: fix the root causes rather than getting faster at restarting things.

The definitions

  • MTTR — mean time to recoveryTotal downtime divided by the number of incidents: the average length of an outage from first user impact to full recovery.
  • MTBF — mean time between failuresWorking time divided by the number of incidents: on average, how long the service runs before the next failure.
  • MTTD — mean time to detectTotal time from first impact to first alert, divided by incidents. It is the part of MTTR that monitoring alone can shrink.
  • AvailabilityOne minus downtime over the period, as a percentage. The number your SLA is written in.

Averages hide the shape

Two services can share an MTTR of 25 minutes while behaving completely differently: one has a single two-hour outage a quarter, the other a five-minute blip every week. The first needs faster rollback and better runbooks; the second needs root-cause work. Look at MTBF next to MTTR, and keep the incident list, not just the averages.

Detection is the cheapest part to fix

If MTTD is a third of MTTR, a third of every outage is spent waiting for someone to notice. A one-minute check interval with multi-region confirmation typically brings detection under two minutes; the check interval planner shows the trade-off against the load you add.

What to count as an incident

Count failures that users could see. A monitor flapping for thirty seconds on one region is not an incident; a checkout returning errors for ten minutes is. Be consistent across periods so the trend means something, and measure downtime from the user's side rather than from when the on-call engineer acknowledged the page.

Shrink the part of MTTR you control

SutramX confirms outages across regions and pages within a minute, so detection stops being the slow part.