Reliability insights
The Reliability page explained: SLOs and error budgets, dependency mapping, tail latency, anomalies, region fingerprints and health scores.
The Reliability page turns your check history into reliability numbers: SLOs with error budgets and burn rates, a dependency graph that suppresses noisy downstream alerts, tail latency per region, latency anomalies, incident region fingerprints and a health score for every monitor. This page explains what each widget shows and exactly how SutramX calculates it.
Open it from the sidebar: Checks & insights → Reliability IQ (https://app.sutramx.com/dashboard/reliability).
Plan availability#
Anyone in the workspace can open the Reliability page and see tail latency, anomalies, region fingerprints and health scores. Two parts depend on your plan:
| Feature | What it unlocks |
|---|---|
| SLO tracking | Creating, editing and deleting SLOs, plus burn rates and error budgets |
| Dependency mapping | Seeing, adding and removing dependencies between monitors |
Both are included on Growth and Pro. On plans without them, the forms are disabled and show "SLO tracking isn't included in your plan" or "Dependency mapping isn't included in your plan" with an Upgrade to enable it link. See Plans & limits for the current lineup.
Time window#
Use the window selector at the top right to choose 7d, 14d, 30d or 90d (default 30d). The window applies to tail latency, region fingerprints and health scores. SLOs always use their own window, and latency anomalies always use a fixed 14-day baseline.
Below the header, a Data sources line tells you where each number came from for the chosen window:
- For windows up to 13 days, uptime, flakiness and percentiles come straight from individual checks ("Raw checks").
- For longer windows, uptime combines individual checks for the most recent 13 days with daily rollups for older days.
- For windows longer than 14 days, p95/p99 come from daily aggregates. Each day's percentile is weighted by that day's check count. This is an approximation, because percentiles can't be combined exactly across days.
- Flakiness always uses only the most recent 13 days of individual checks, because counting status flips needs every check.
Summary tiles#
The four tiles at the top of the page:
| Tile | What it shows |
|---|---|
| Average health score | Mean of every monitor's health score in the window, out of 100 |
| Latency anomalies | Number of monitors currently flagged as anomalous |
| SLOs burning too fast | Number of enabled SLOs whose fast and slow burn rates are both above their thresholds |
| Dependencies | Number of dependency links in the workspace |
SLOs & error budgets#
An SLO (service level objective) is an availability target for one monitor over a rolling window, for example "99.9% of checks succeed over 30 days". Each monitor can have one SLO. Saving a second SLO for the same monitor replaces the first.
Create an SLO#
- Open Reliability IQ and find the SLOs & error budgets panel.
- Choose a Monitor.
- Set Target (%), Window (days), Fast burn threshold and Slow burn threshold.
- Leave Track this SLO checked so burn rates are calculated.
- Click Save SLO.
To change an SLO, click Edit on it, adjust the values and click Update SLO. You can't change the monitor of an existing SLO. Delete it and create a new one instead. Delete asks for confirmation. Only burn-rate tracking stops; check history isn't affected.
| Field | What it does | Default / limits |
|---|---|---|
| Monitor | The monitor whose checks the SLO measures | Required, one SLO per monitor |
| Target (%) | Share of checks that must succeed | 99.9; 90 to 99.999 |
| Window (days) | Rolling window the target applies to | 30; whole days, 1 to 30 |
| Fast burn threshold | Burn-rate multiple that trips the short window | 2; 0.1 to 100 |
| Slow burn threshold | Burn-rate multiple that trips the full window | 1; 0.1 to 100 |
| Track this SLO | When off, the SLO is kept but no burn rate is calculated (shown as "not tracked") | On |
How the error budget is calculated#
For an SLO with target T and window W:
- Error budget = W × (100% − T). For example, 99.9% over 30 days allows 30 × 24 × 60 × 0.001 = 43.2 minutes of failed checks.
- Consumed = W × observed error rate, where error rate = failed checks ÷ total checks in the window.
- Error budget left = budget − consumed, shown as minutes and as a percentage of the budget. It never goes below 0%.
- The budget is exhausted when nothing is left and the window contains at least one check.
The SLO counts any check that isn't up as an error, including degraded checks (reachable but slow or partially failing). Checks run during maintenance windows or while a monitor is paused are not excluded. That makes an SLO stricter than the uptime percentage on a monitor's own page (see Uptime math).
Example: 99.9% over 30 days with 0.05% of checks failing consumes 30 days × 0.0005 = 21.6 minutes. That leaves 21.6 of 43.2 minutes, or 50% of the budget.
Burn rates#
A burn rate says how fast you're spending the budget: burn rate = observed error rate ÷ allowed error rate. A burn rate of 1.0 uses up the budget exactly by the end of the window. A burn rate of 2.0 uses it up in half the window.
SutramX measures two windows for each SLO:
- Slow window: the full SLO window.
- Fast window: 1/12 of the SLO window, clamped to between 5 and 60 minutes. For the 1–30 day windows you can set in the dashboard, this is always the last 60 minutes.
An SLO shows Burning too fast when the fast burn rate is at or above the fast threshold and the slow burn rate is at or above the slow threshold. Otherwise it shows Within budget. Requiring both windows filters out short blips (fast only) and old damage that has already stopped (slow only).
Each SLO card shows Fast burn: 0.00× (threshold 2.0), Slow burn, Error budget left with remaining and total minutes, and a budget meter.
Dependency graph#
A dependency says that one monitor (the upstream, or parent) must be healthy for another (the downstream, or child) to work. For example, your database monitor is upstream of your API monitor.
When an upstream monitor has an open incident, a downstream monitor that goes down still gets an incident, but its down alert is suppressed. The incident records that it was suppressed because a dependency was down. One root cause then pages you once instead of once per affected service.
Add a dependency#
Only the workspace owner can add or remove dependencies, because they silence alerts for downstream monitors.
- In the Dependency graph panel, choose Upstream (parent) and Downstream (child).
- Click Add dependency.
Rules:
- A monitor can't depend on itself.
- Both monitors must be in the current workspace.
- A dependency that would create a cycle (A → B → A, directly or through other monitors) is rejected with "This dependency would create a cycle".
- Adding the same pair again is harmless. It updates the existing link.
To remove a link, click Remove next to it in the list and confirm. The downstream monitor then alerts on its own again when the upstream one is down.
Arrange the graph#
Each node is a monitor. Paused monitors are drawn greyed out, and arrows point from upstream to downstream. Drag nodes to arrange them, or focus a node with Tab and move it with the arrow keys. Positions are saved for the whole workspace, so everyone sees the same layout. Press Enter on a node (or click it) to open its details.
Tail latency (p95 / p99)#
The chart shows the 12 monitor-and-region pairs with the worst p99 response time in the window. Click Show as table to see every monitor and region with its p95, p99 and Samples count.
- p95: 95% of checks were faster than this.
- p99: 99% of checks were faster than this.
A big gap between p95 and p99 means occasional slow outliers, such as cold starts, garbage collection pauses or a slow network path in one region. Percentiles are calculated per probe region, so one slow region doesn't get hidden by fast ones.
Latency anomalies#
A monitor is flagged when its latest response time is above its 14-day mean + 2 standard deviations. The baseline needs at least 20 response-time samples from the last 14 days. Each anomaly shows Latest … vs. threshold ….
This check reflects the most recent response, so a flagged monitor clears as soon as a normal response comes in. It doesn't change with the window selector.
Incident region fingerprints#
For each incident in the window (up to the 100 most recent), SutramX looks at checks from every probe region from 10 minutes before the incident started until it resolved. For an ongoing incident, it uses the 10 minutes after the start. It then classifies the incident:
| Fingerprint | Rule |
|---|---|
| Global — most regions failed | 3 or more regions failed, and at least 60% of the regions that reported |
| Multi-region | 2 or more regions failed, but not enough to count as global |
| Single region | Exactly 1 region failed |
| Localized — no region reported failures | Regions reported, but none of their checks failed in that period |
| Unknown — raw data no longer retained | No per-region data, or the incident started more than 14 days ago |
The failing regions are listed after the date. A single-region fingerprint usually points to a network path or regional provider issue rather than your service. A global one usually means the service itself is down. For how regions confirm outages before an incident opens, see Regions & confirmation.
Monitor health score#
Each monitor gets a score from 0 to 100 for the window:
score = 100
− (100 − uptime%) × 0.6
− min(20, incidents × 2)
− min(15, MTTR minutes ÷ 10)
− min(15, flakiness × 100)- uptime%: checks with status up ÷ all checks in the window. A monitor with no checks counts as 100%.
- incidents: incidents that started in the window.
- MTTR: mean time to recovery, the average incident duration in minutes. Ongoing incidents count up to now.
- flakiness: status flips per check, counted separately per region. Degraded counts as up here, so slow-but-reachable checks don't count as flips.
The score is never below 0. The chart shows the 20 lowest-scoring monitors, and the list below shows every monitor with Score · Uptime · incidents · MTTR. A monitor that bounces up and down scores lower than one with a single outage of the same total length.
Worked example: 99.5% uptime, 3 incidents, 12-minute MTTR and a flakiness of 0.02 gives 100 − 0.3 − 6 − 1.2 − 2 = 90.5.
Monitor details#
Click a monitor in the health score list, or a node in the dependency graph, to open its details:
- Health score, Uptime, Fast / slow burn (or "No SLO") and Error budget left.
- Latency by region: p95/p99 per region.
- Latency anomalies for that monitor.
- Incident region fingerprints for that monitor's incidents.
Uptime math#
Uptime appears in several places, and the rules differ slightly on purpose. If two numbers disagree, this is usually why:
| Where | Degraded checks | Paused / maintenance time | Rounding |
|---|---|---|---|
| Monitor page and uptime history | Count as up | Excluded | Rounded down to 3 decimals, never up to 100% |
| Reliability health score | Count as down | Included | 3 decimals |
| SLO error budget | Count as errors | Included | 2 decimals |
| Weekly report email | Count as failed | Included | 2 decimals, shown as 100% at 99.995% and above |
A time range with no checks has no uptime value on the monitor page, rather than 100%.
To convert an uptime target into allowed downtime, multiply the period by (100% − target). For 30 days (43,200 minutes):
| Target | Allowed failure per 30 days |
|---|---|
| 99% | 432 minutes (7.2 hours) |
| 99.5% | 216 minutes (3.6 hours) |
| 99.9% | 43.2 minutes |
| 99.95% | 21.6 minutes |
| 99.99% | 4.32 minutes |
SutramX measures uptime as a share of checks, not of wall-clock time. A failed check stands in for the time until the next one, so a shorter check interval gives a more precise figure.
Usage cost tracking#
On plans with cost analytics (Growth and Pro), the workspace owner sees a Usage costs section on the Billing tab of Plans & billing. It shows an estimated infrastructure cost of your monitoring:
- Cost attribution for the last 7 days, 30 days or 90 days, with Estimated cost, Checks, Incidents and Notifications. Click Refresh cost data to recalculate.
- Budget thresholds (US$): enter a Threshold name and Amount (US$), then click Add threshold. Thresholds added here are monthly and cover the whole workspace. When the estimated cost for the current period reaches a threshold, SutramX records one budget alert for that period.
Use the API#
The page uses the same REST endpoints you can call with an API key (see REST API):
# Reliability overview for the last 30 days (days: 7, 14, 30 or 90; anything else uses 30)
curl -H "Authorization: Bearer $SUTRAMX_API_KEY" \
"https://api.sutramx.com/reliability/overview?days=30"
# Create or replace the SLO for a monitor
curl -X POST "https://api.sutramx.com/reliability/slo-targets" \
-H "Authorization: Bearer $SUTRAMX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"monitor_id": "8f0c7d2e-1b2a-4c3d-9e8f-0a1b2c3d4e5f",
"target_percentage": 99.9,
"window_minutes": 43200,
"burn_rate_fast_threshold": 2,
"burn_rate_slow_threshold": 1,
"is_enabled": true
}'
# List dependencies
curl -H "Authorization: Bearer $SUTRAMX_API_KEY" \
"https://api.sutramx.com/reliability/dependencies"Through the API, window_minutes accepts 60 to 43,200 (1 hour to 30 days). Other endpoints: GET /reliability/slo-targets, PUT and DELETE /reliability/slo-targets/{id} and GET /reliability/monitor/{monitorId}?days=30 for a single monitor (same days values). Adding (POST /reliability/dependencies) and removing (DELETE /reliability/dependencies/{id}) dependencies needs the workspace owner's signed-in session, so API keys can't change them.
Common questions#
Why does my SLO show 0% budget left when the monitor page says 99.95% uptime? The monitor page counts degraded checks as up and leaves out maintenance and paused time. The SLO counts both against you. See Uptime math.
Why is a fingerprint "Unknown"? Per-region check data for incidents that started more than 14 days ago isn't used for fingerprints. Switch to a shorter window to focus on recent incidents.
Why don't I see a burn rate for my SLO? Track this SLO is off (the SLO shows "not tracked"), or your plan doesn't include SLO tracking.
Does a dependency hide the downstream incident? No. The incident is still recorded and visible. Only the down alert is suppressed while the upstream monitor has an open incident.
Related
Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.