Why this alert & flakiness score
Every SutramX incident explains itself — region votes, the quorum rule, the failure class, the alert decision and a fault verdict — and every monitor gets a 7- and 30-day flakiness score.
Every incident in SutramX has a Why this alert panel. It shows the evidence behind the incident and the decision SutramX made about alerting, so you don't have to work it out from raw check results. The same reasoning is summed up in a one-line verdict that is also added to your alerts.
The explanation is built with fixed rules from the checks and timeline events recorded for the incident. No AI is involved, so the same evidence always gives the same explanation. It is included on every plan, Free too.
Where to find it#
- Incident page: open any incident from Incidents or from an alert link. The Why this alert panel is near the top.
- Monitor page: the Why card near the top of the page, above the stats row, shows the latest verdict, the quorum rule, one chip per region (for example "Frankfurt: Up") and the monitor's flakiness. While an incident is open it links to the full panel (Why this alert →).
- API, AI assistants and the terminal:
GET /incidents/:id/explanation,GET /monitors/:id/explanationandGET /monitors/:id/flakinesswith any API key (see API reference); the MCP toolsutramx_explain_incident(see MCP server); andsutramx why <monitor-or-incident>(see CLI).
What the panel shows#
The verdict#
One sentence that sums up the incident, for example:
- "Down from 2 of 3 regions: HTTP 503 server error (Arizona still up)"
- "Down from 3 of 3 regions: Timeout (all regions agree)"
- "… likely not your fault: Stripe API degraded for several SutramX customers"
- "Only Mumbai saw timeouts; quorum not met (2 needed), no alert sent" (on the monitor page, when no incident opened)
Fault verdict#
A badge next to the verdict says where the evidence points:
| Badge | Meaning |
|---|---|
| Confirmed at your service | Enough regions confirmed a failure at your endpoint itself: DNS, TLS, connection, timeout, an HTTP error status, a redirect or content check, or a slow response |
| Likely not your fault | A vendor your endpoint depends on is failing for several SutramX customers (see Vendor health), or the problem is on an Indian ISP network (see Last-mile checks) |
| Our checker, not your site | The evidence points at a SutramX checker, for example every region was inconclusive |
| (no badge) | The evidence doesn't point clearly anywhere. Everything that was seen is still shown |
Facts#
| Fact | What it tells you |
|---|---|
| Quorum rule | How many regions had to agree, for example "2 of 3 regions must agree", and whether a cross-region confirmation ran. See Regions & confirmation |
| Failure | The failure class in plain words and its scope: One region, Several regions or All regions |
| Alert decision | Alert sent, or why not (see below) |
Failure classes:
| Class | Label |
|---|---|
| DNS | DNS resolution failed |
| TLS | TLS / certificate error (more specific when known, for example "TLS certificate expired") |
| Connection | Connection failed |
| Timeout | Timeout |
| 5xx | HTTP 5xx server error (for example "HTTP 503 server error") |
| 4xx | HTTP 4xx client error |
| Other status | Unexpected HTTP status |
| Redirect | Redirect error |
| Content | Content check failed |
| Slow | Response time over the limit |
| 429 | Rate limited (HTTP 429) |
| Bot protection | Blocked by bot protection |
| Checker | Inconclusive (checker-side problem) |
| Other | Check failed |
Region votes#
A table with one row per region: Region, Result (Up, Down, Blocked, Inconclusive or No result), What it saw (error class and HTTP status), Timings and Checked. Regions whose result counted towards opening the incident are marked (confirmed). Blocked and inconclusive results never vote.
Alert decision#
If no alert went out, or it was delayed, the panel says why. Reasons include:
| Reason | What it means |
|---|---|
| Maintenance | A maintenance window was active |
| Quiet hours / deployment window | Alerts were paused for that window |
| Dependency | An upstream monitor this one depends on was already down |
| Error filter | An alert filter on the monitor excluded this failure |
| Flapping | The monitor was changing state too often; changes were collapsed into one incident |
| Checker hold | SutramX briefly held the alert because many incidents opened across all its customers at once (possibly its own network) |
| Quorum not met | Not enough regions agreed |
| No recipients / your preferences | Nobody was set to receive it, or notification preferences muted it |
| Recovered | The monitor recovered before the alert went out, for example when the last re-check before sending found it answering again (see How we verify an outage) |
| Opened manually | The incident was opened through the API, so no alert was sent |
| Workspace suspended | The workspace was suspended |
| Channel failures | Delivery to a channel failed; the incident timeline has the details |
Contributing findings#
Extra evidence that may explain the incident, labelled Third-party vendor, Internet provider (last mile) or SutramX checker. The findings SutramX thinks most likely are shown separately in a Likely cause box right under the verdict.
The verdict in your alerts#
Down alerts include the one-line verdict as Why: … on:
- Email, Slack, Microsoft Teams, Discord, Google Chat, Mattermost and Telegram
- PagerDuty (in
custom_details.explanation) and Opsgenie (in the alert description) - Webhooks and Zapier, in the
explanationfield of the payload. It isnullon recovery events or when no verdict is available. See Webhooks
SMS, WhatsApp and voice alerts stay short and don't include it.
Flakiness score#
Every monitor gets a flakiness score from 0 to 100, for the last 7 days and the last 30 days, shown on the monitor page's Why card. Higher means noisier.
| Score | Band |
|---|---|
| 0–9 | Stable |
| 10–29 | Some noise |
| 30–59 | Flaky |
| 60+ | Very flaky |
The score adds up these signals, each with its own cap:
- Unconfirmed failures: failed checks that other regions didn't confirm.
- Short incidents: incidents that recovered on their own within 5 minutes.
- Flapping: episodes where the monitor kept changing state.
- Inconclusive checks: checks that couldn't reach a result.
The top reason is shown under the score; hover the score badge to see up to three. A window with fewer than 20 checks shows "not enough data" instead of a score. Scores are refreshed every few minutes.
A flaky monitor is worth tuning: raise the failure threshold, check from more regions, or loosen a response-time limit. See Monitor settings.
Related
Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.