Guide · 9 min read

MTTR, MTTD, MTTA and MTBF Explained, with Formulas and a Worked Example

Incident metrics are only useful if everyone means the same thing by them. MTTR alone has four common meanings.

Last updated

In short

MTTD (mean time to detect) is the average time from a failure starting to an alert. MTTA (mean time to acknowledge) runs from alert to a person responding. MTTR usually means mean time to recovery: failure to service restored. MTBF (mean time between failures) is the average working time between incidents. Each is a total divided by the incident count.

The incident timeline behind every metric

Every incident passes through the same moments, and each metric measures the gap between two of them.

  • Failure starts: users are first affected.
  • Detected: monitoring, or a customer, notices and an alert goes out.
  • Acknowledged: a person confirms they are on it.
  • Recovered: the service works for users again, usually through a rollback, restart, failover or fix.
  • Resolved: the underlying cause is fixed so it will not recur, which can be days later.

Every "mean time to" metric is the total of one of these gaps across a set of incidents, divided by the number of incidents.

MTTD: mean time to detect

MTTD = total time from failure start to detection ÷ number of incidents.

It measures how long problems exist before anyone knows about them. It is the part of an outage spent with nobody working on it, which makes it the easiest to justify reducing. It is also the hardest to measure honestly: the failure start time usually has to be reconstructed afterwards from logs, error rates or the first failed check.

MTTA: mean time to acknowledge

MTTA = total time from alert to acknowledgement ÷ number of incidents.

It measures how quickly a person picks up an alert. A high MTTA usually points to routing (alerts going to a channel instead of a person), noise (people have learned to ignore alerts), or coverage gaps (nobody on call at night or at weekends).

MTTR: the four meanings

The R in MTTR stands for different words in different teams and tools, and the numbers they produce are not comparable.

  • Mean time to repairThe time spent actually fixing the fault, from starting the repair to finishing it. It comes from equipment maintenance and leaves out detection and response.
  • Mean time to recovery (or restore)From failure start until the service works for users again. This is the most common meaning for online services and the one closest to what users experience.
  • Mean time to respondFrom the alert until the team is working on the problem, or in some definitions until it is mitigated. It overlaps with MTTA and leaves out detection.
  • Mean time to resolveFrom failure start until the incident is fully closed, including the fix that stops it recurring. It can be far longer than recovery, because a quick rollback ends the outage but not the work.
Write the definition down

Before you report MTTR, decide which start and end points you use and put them next to the number. A team that moves from "repair" to "recovery" will see its MTTR jump without anything getting worse.

MTBF: mean time between failures

MTBF = total working time ÷ number of failures, where working time is the period minus downtime.

It measures how often things break. Some tools instead average the gap between one incident starting and the next one starting; that version includes the downtime, so it is larger than the first by exactly the MTTR. Either works if you stay consistent. MTTF (mean time to failure) is the related term for things that are replaced rather than repaired.

MTBF and MTTR together give availability: availability = MTBF ÷ (MTBF + MTTR). Raising availability means failing less often, recovering faster, or both.

A worked example

Take a 30-day month (43,200 minutes) with three incidents. The times are in minutes.

  • Incident A: a bad deploy, mid-morningDetected 4 minutes after failure start, acknowledged 5 minutes after the alert, recovered by rollback 40 minutes after failure start.
  • Incident B: a database failover, overnightDetected after 2 minutes, but acknowledged 18 minutes after the alert because the page went to a chat channel. Recovered after 75 minutes.
  • Incident C: checkout failing behind a healthy home pageDetected after 12 minutes, by a customer email, because no monitor checked checkout. Acknowledged after 3 minutes, recovered after 35.
  • MTTD = (4 + 2 + 12) ÷ 3 = 6 minutes.
  • MTTA = (5 + 18 + 3) ÷ 3 ≈ 8.7 minutes.
  • MTTR (recovery) = (40 + 75 + 35) ÷ 3 = 50 minutes, for 150 minutes of downtime in total.
  • MTBF = (43,200 − 150) ÷ 3 = 14,350 minutes, just under 10 days.
  • Availability = 14,350 ÷ (14,350 + 50) ≈ 99.65%, the same as (43,200 − 150) ÷ 43,200.

Breaking the 50-minute MTTR down shows where the time went: on average 6 minutes undetected, about 8.7 minutes waiting for a person, and about 35.3 minutes diagnosing and fixing. Each part has a different remedy: incident C needs a monitor on checkout, incident B needs alerts that reach a person, and incident A needs a faster rollback. The free MTTR and MTBF calculator does this arithmetic from your own incident count, downtime and period.

How to measure them from your own incidents

You do not need a special tool to start. A spreadsheet with one row per incident and five timestamps is enough, as long as each timestamp always comes from the same source.

  • Failure startThe first failed check, or the moment the error rate or latency graph left its normal range. Reconstruct it during the postmortem rather than using the alert time.
  • DetectionThe first alert sent, or the time of the first customer report if monitoring never fired. Note which, because the second kind is a monitoring gap.
  • AcknowledgementThe time someone acknowledged the alert in your paging or monitoring tool, not when they first replied in chat.
  • RecoveryWhen the monitor recovered or users could complete the affected action again.
  • ResolutionWhen the follow-up fix shipped, if you track mean time to resolve as well.

Review the numbers monthly or quarterly alongside the incident list itself. The trend, and the incidents that drive it, tell you far more than any single month's figure.

How do you reduce each metric?

Reducing MTTD

  • Shorten the check interval on critical paths. The worst-case detection time equals the interval and the average is about half of it.
  • Monitor what users do, not only a /health route: login, checkout and the API calls your app makes, with content assertions.
  • Confirm failures from more than one region. It adds a few seconds but stops false alarms that teach people to ignore real ones.
  • Add heartbeats for scheduled jobs, which fail without producing any request to watch.

Reducing MTTA

  • Route urgent alerts to the person on call, through a channel that wakes them: push, SMS or a phone call rather than email at 3am.
  • Add an escalation policy, so an unacknowledged alert reaches a second person after a fixed delay.
  • Cut noise. Every alert that needs no action makes the next real one slower to acknowledge.
  • Let people acknowledge from where the alert arrives, such as a button in chat.

Reducing the repair part of MTTR

  • Make rollback fast and routine, and put risky changes behind feature flags you can switch off.
  • Keep runbooks for known failure modes, next to the alert that triggers them.
  • Record deploys and config changes where responders can see them; "what changed?" is usually the first question.
  • Post a status page update early, so support requests do not compete with the fix for attention.

Increasing MTBF

  • Run a postmortem on significant incidents and ship the action items that remove the cause.
  • Look for repeats: the same monitor or the same cause appearing month after month.
  • Use an error budget to decide when to slow feature work in favour of reliability.

Pitfalls when reporting incident metrics

  • Averages hide the shapeOne four-hour outage and a dozen two-minute blips can produce similar averages but need different fixes. Keep the incident list next to the means, and look at medians or the worst incident too.
  • Small numbers swingWith three incidents a month, one bad night doubles MTTR. Compare quarters, or rolling windows, before declaring a trend.
  • Undetected incidents are missingAn outage nobody noticed never enters the data, so poor monitoring makes MTTD look better, not worse.
  • What counts as an incident changes the resultDecide whether partial outages, single-region failures and degraded performance count, and apply the rule the same way every period.
  • Targets get gamedIf MTTR becomes a target, incidents start getting closed early. Use these numbers to find where time goes, not to rank people.

Where SutramX shows these numbers

A SutramX incident opens when an outage is confirmed, so the time from the first failed check to the incident is bounded by the check interval plus the confirmation re-checks. Each monitor's page shows MTBF as the average gap between incident starts, once it has at least two incidents. The Reliability page shows MTTR per monitor as the average incident duration, as part of its health score. The incident list exports to CSV with started_at, acknowledged_at, resolved_at and duration_seconds for each incident, which is enough to calculate MTTA and MTTR in a spreadsheet.

Next steps

Keep reading

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.