Reliability Insights

Reliability Insights: Cause, Impact and Error Budgets

Knowing something is down is step one. SutramX also helps you answer the next questions: which change caused it, what else it took down, and how much error budget is left.

In short

From the Growth plan ($29/mo), SutramX adds reliability insights: deploy events from GitHub, GitLab or any CI are matched to incidents that start within two hours, SLO targets track each monitor's error budget, monitor dependencies suppress downstream alerts when an upstream fails, and weekly email reports summarise uptime and incidents.

Who it's for

For production engineering teams that run against SLOs and want to see how much error budget is left. Each monitor can have an availability SLO over a rolling window, with its error budget and fast and slow burn rates. SLO tracking and dependency mapping are included on Growth and Pro. Tail latency, latency anomalies, region fingerprints and health scores on the Reliability page are visible to everyone in the workspace.

Which deploy caused this?

Most incidents follow a change. Send deploy events to SutramX from GitHub, GitLab or a generic signed webhook, and each incident lists the deploys that happened shortly before it started, ranked by how close in time they were and whether the service name matches.

The result is a heuristic, not a verdict — but it turns the first ten minutes of an incident from "what changed?" into "was it this?".

ProbefiresCheckfailsQuorumconfirmsAlertdispatches

SLOs and error budgets

An uptime percentage on its own is a rear-view mirror. An SLO target turns it into a budget you can spend: when the budget is healthy, ship; when it is burning fast, slow down and fix reliability first.

Dependencies and weekly reports

Map which monitors depend on which. When a shared database or upstream API fails, SutramX records the downstream failures as suppressed rather than paging you once per affected service.

Weekly reports summarise uptime, incidents and trends by email, so reliability gets reviewed on a schedule instead of only after something breaks.

Specifications

PlansGrowth, Pro and Enterprise
Deploy sourcesGitHub, GitLab, generic signed webhook, or the API
Correlation windowDeploys up to 2 hours before an incident starts
SLO targetsPer-monitor availability targets with error budget
DependenciesDownstream alert suppression when a parent monitor is down
Usage cost trackingEstimated monitoring usage cost with budget alerts

Learn more

Related capabilities

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.