Alerts & Incidents

Why this alert & flakiness score

Every SutramX incident explains itself — region votes, the quorum rule, the failure class, the alert decision and a fault verdict — and every monitor gets a 7- and 30-day flakiness score.

Every incident in SutramX has a Why this alert panel. It shows the evidence behind the incident and the decision SutramX made about alerting, so you don't have to work it out from raw check results. The same reasoning is summed up in a one-line verdict that is also added to your alerts.

The explanation is built with fixed rules from the checks and timeline events recorded for the incident. No AI is involved, so the same evidence always gives the same explanation. It is included on every plan, Free too.

Where to find it#

  • Incident page: open any incident from Incidents or from an alert link. The Why this alert panel is near the top.
  • Monitor page: the Why card near the top of the page, above the stats row, shows the latest verdict, the quorum rule, one chip per region (for example "Frankfurt: Up") and the monitor's flakiness. While an incident is open it links to the full panel (Why this alert →).
  • API, AI assistants and the terminal: GET /incidents/:id/explanation, GET /monitors/:id/explanation and GET /monitors/:id/flakiness with any API key (see API reference); the MCP tool sutramx_explain_incident (see MCP server); and sutramx why <monitor-or-incident> (see CLI).

What the panel shows#

Why this alert
Region results, the quorum rule, the failure and the alert decision add up to a one-line verdict and a fault badge. The same evidence always gives the same explanation.

The verdict#

One sentence that sums up the incident, for example:

  • "Down from 2 of 3 regions: HTTP 503 server error (Arizona still up)"
  • "Down from 3 of 3 regions: Timeout (all regions agree)"
  • "… likely not your fault: Stripe API degraded for several SutramX customers"
  • "Only Mumbai saw timeouts; quorum not met (2 needed), no alert sent" (on the monitor page, when no incident opened)

Fault verdict#

A badge next to the verdict says where the evidence points:

BadgeMeaning
Confirmed at your serviceEnough regions confirmed a failure at your endpoint itself: DNS, TLS, connection, timeout, an HTTP error status, a redirect or content check, or a slow response
Likely not your faultA vendor your endpoint depends on is failing for several SutramX customers (see Vendor health), or the problem is on an Indian ISP network (see Last-mile checks)
Our checker, not your siteThe evidence points at a SutramX checker, for example every region was inconclusive
(no badge)The evidence doesn't point clearly anywhere. Everything that was seen is still shown

Facts#

FactWhat it tells you
Quorum ruleHow many regions had to agree, for example "2 of 3 regions must agree", and whether a cross-region confirmation ran. See Regions & confirmation
FailureThe failure class in plain words and its scope: One region, Several regions or All regions
Alert decisionAlert sent, or why not (see below)

Failure classes:

ClassLabel
DNSDNS resolution failed
TLSTLS / certificate error (more specific when known, for example "TLS certificate expired")
ConnectionConnection failed
TimeoutTimeout
5xxHTTP 5xx server error (for example "HTTP 503 server error")
4xxHTTP 4xx client error
Other statusUnexpected HTTP status
RedirectRedirect error
ContentContent check failed
SlowResponse time over the limit
429Rate limited (HTTP 429)
Bot protectionBlocked by bot protection
CheckerInconclusive (checker-side problem)
OtherCheck failed

Region votes#

A table with one row per region: Region, Result (Up, Down, Blocked, Inconclusive or No result), What it saw (error class and HTTP status), Timings and Checked. Regions whose result counted towards opening the incident are marked (confirmed). Blocked and inconclusive results never vote.

Alert decision#

If no alert went out, or it was delayed, the panel says why. Reasons include:

ReasonWhat it means
MaintenanceA maintenance window was active
Quiet hours / deployment windowAlerts were paused for that window
DependencyAn upstream monitor this one depends on was already down
Error filterAn alert filter on the monitor excluded this failure
FlappingThe monitor was changing state too often; changes were collapsed into one incident
Checker holdSutramX briefly held the alert because many incidents opened across all its customers at once (possibly its own network)
Quorum not metNot enough regions agreed
No recipients / your preferencesNobody was set to receive it, or notification preferences muted it
RecoveredThe monitor recovered before the alert went out, for example when the last re-check before sending found it answering again (see How we verify an outage)
Opened manuallyThe incident was opened through the API, so no alert was sent
Workspace suspendedThe workspace was suspended
Channel failuresDelivery to a channel failed; the incident timeline has the details

Contributing findings#

Extra evidence that may explain the incident, labelled Third-party vendor, Internet provider (last mile) or SutramX checker. The findings SutramX thinks most likely are shown separately in a Likely cause box right under the verdict.

The verdict in your alerts#

Down alerts include the one-line verdict as Why: … on:

  • Email, Slack, Microsoft Teams, Discord, Google Chat, Mattermost and Telegram
  • PagerDuty (in custom_details.explanation) and Opsgenie (in the alert description)
  • Webhooks and Zapier, in the explanation field of the payload. It is null on recovery events or when no verdict is available. See Webhooks

SMS, WhatsApp and voice alerts stay short and don't include it.

Flakiness score#

Every monitor gets a flakiness score from 0 to 100, for the last 7 days and the last 30 days, shown on the monitor page's Why card. Higher means noisier.

ScoreBand
0–9Stable
10–29Some noise
30–59Flaky
60+Very flaky

The score adds up these signals, each with its own cap:

  • Unconfirmed failures: failed checks that other regions didn't confirm.
  • Short incidents: incidents that recovered on their own within 5 minutes.
  • Flapping: episodes where the monitor kept changing state.
  • Inconclusive checks: checks that couldn't reach a result.

The top reason is shown under the score; hover the score badge to see up to three. A window with fewer than 20 checks shows "not enough data" instead of a score. Scores are refreshed every few minutes.

A flaky monitor is worth tuning: raise the failure threshold, check from more regions, or loosen a response-time limit. See Monitor settings.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.