What an error budget is
If you promise 99.9% availability over 30 days, you are also promising no more than 43.2 minutes of downtime. That allowance is the error budget. It turns a vague goal into a number a team can spend, on risky deploys, on migrations, on the inevitable incident, and it tells you when to stop spending.
Burn rate, in one line
Burn rate is how fast you consume the budget compared with spending it evenly across the window. A burn rate of 1 means you will land exactly on target. A burn rate of 2 means the budget is gone halfway through the window. Google's SRE workbook alerts on a 14.4× burn over one hour (which would eat 2% of a 30-day budget in sixty minutes) and raises a lower-priority ticket at 1× over three days.
- Budget healthy, burn rate under 1Ship. This is the signal that it is safe to take risk: feature flags on, migrations scheduled, experiments running.
- Burn rate over 1, budget not yet goneSlow down. Prioritise reliability work and keep deploys small until the rate drops back.
- Budget exhaustedFreeze non-essential changes. Every further minute of downtime is a breached promise, which is exactly the moment to stop adding new ways to break things.
Internal health checks that return 200 while customers see errors do not count as uptime. Measure availability from outside your network, from more than one region, and with the same request a real client would make. Our SLA, SLI and SLO guide walks through choosing an indicator that matches what users experience.
Related
The SLA uptime calculator converts any percentage into allowed downtime per day, week, month and year. The MTTR and MTBF calculator tells you whether the budget is being spent on a few long incidents or many short ones.