Incidents
How SutramX opens and closes incidents automatically, and how to acknowledge, snooze, resolve, annotate and write postmortems for them.
An incident is SutramX's record of one outage of one monitor. It opens automatically when a monitor is confirmed down, collects everything that happens while it is open (checks, alerts, escalation steps, notes), and closes automatically when the monitor recovers. You use the incident page to coordinate the response, post public updates to your status page and write the postmortem.
How incidents open and close#
SutramX opens and closes incidents for you. There is no New incident button in the dashboard: every incident belongs to a monitor.
Opening#
An incident opens when a monitor's failure is confirmed:
- A check fails. A single failed check doesn't open an incident.
- The monitor has to fail the number of consecutive checks set by its failure threshold, and enough regions have to agree it is down (the down quorum). See Regions & confirmation.
- Once both conditions are met, SutramX opens the incident, records which regions confirmed it, and sends the down alert to your alert channels.
Every incident explains itself: the Why this alert panel shows each region's result, the quorum rule, the failure class, the alert decision and a verdict that is also included in the alert. See Why this alert & flakiness score.
A monitor can have only one open incident at a time. A degraded check (slow but reachable) never opens an incident.
If alerts are suppressed when the incident opens (for example by a maintenance window, quiet hours, a deployment window, an upstream dependency that is already down, an alert filter or your notification preferences), the incident is still recorded, but it is marked Suppressed with the reason and no down alert goes out. If the suppression ends while the monitor is still down, the down alert is sent then.
Closing#
An incident resolves automatically when the monitor passes the number of consecutive checks set by its recovery threshold and enough regions agree it is back up. SutramX then:
- Records the total down time.
- Sends a recovery alert, but only if a down alert was actually delivered for this incident.
- Ends any escalation that is still running.
- Emails status page subscribers that the incident is resolved, if they were told it started.
Flapping#
If a monitor goes down again within a few minutes of an automatic recovery, SutramX reopens the same incident instead of creating a new one, and marks it Flapping. One flapping episode stays one incident, and its down time adds up only the time the monitor was actually down, not the up gaps in between.
While an incident is flapping, repeat down alerts are held back so you aren't paged on every bounce. If the monitor then stays down for a full cooldown period after the last bounce, and the last thing you were told was "recovered", the down alert is sent again. The Flapping chip clears once the monitor has stayed up for the cooldown period.
Incident states#
The list and the incident page show a main status plus extra state chips.
| State | What it means |
|---|---|
| Ongoing | The incident is open: the monitor is still confirmed down. |
| Resolved | The monitor recovered and SutramX closed the incident. |
| Resolved manually | Someone on your team resolved it with Resolve. |
| Suppressed — reason | Alerts were suppressed for this incident. The reason is shown, for example maintenance window, quiet hours, deployment window, upstream dependency down, alert filter, notification preferences, flapping cooldown or opened manually. |
| Acknowledged by name | Someone acknowledged the incident. Hover the chip to see when. |
| Snoozed until time | Escalation and repeat alerts are paused until this time. |
| Flapping | The monitor has been going up and down. |
The incidents list#
Open Incidents in the sidebar to see every incident across your monitors, newest first, 25 per page.
| Control | What it does |
|---|---|
| Status tabs | All, Ongoing, Resolved and Suppressed. |
| Search | Matches the monitor name or URL. |
| Monitor | Shows incidents for one monitor only. |
| Started from / Started to | Limits the list to incidents that started in a date range (in your profile's time zone). |
| Clear filters | Resets the filters. |
| Export | Export CSV or Export JSON downloads the incidents that match the current filters (up to 5,000). |
Each row shows the monitor, its state chips, when it started and how long it lasted (for ongoing incidents, how long so far). The list refreshes itself every 30 seconds while any incident on the page is ongoing, and every 45 seconds otherwise.
The CSV export has these columns: id, monitor_name, monitor_url, started_at, resolved_at, duration_seconds, status, alert_suppressed, suppression_reason, is_flapping, confirmed_from, acknowledged_at, acknowledged_by, resolved_manually and postmortem.
The incident page#
Click an incident to open its page. The header shows the monitor name (a link to the monitor), the status, the state chips and the regions that confirmed the outage, for example "Confirmed down from: FRA1 (Frankfurt, Germany), AZ (Arizona, USA) (2/2)". If the incident created a GitHub issue through the GitHub integration, a GitHub issue link appears too.
Below the header, four cards show Started, Resolved, Down for (or Down time once resolved) and Acknowledged. An ongoing incident page refreshes itself every 30 seconds.
| Section | What it shows |
|---|---|
| Timeline | Every event in order: incident started, alerts sent, skipped or suppressed (with the reason), escalation steps, acknowledgements, snoozes, notes, maintenance windows and the resolution. Use Load more for older events. |
| Notes | Internal notes and public status page updates. See Notes and public updates. |
| Postmortem | Your write-up, once the incident is resolved. See Postmortems. |
| Incident insights | Facts computed from your monitoring data, plus an optional AI summary. See Root-cause information. |
| Status page updates | AI-drafted public updates you review before publishing. See AI incident assist. |
| Escalation | The escalation policy in use, which steps were sent and when the next one fires. See Escalation & on-call. |
| Notifications | Each down and recovery alert per channel, with its status: Sent, Queued, Retrying, Failed or Skipped, and the error if there was one. |
| Runbook | The runbook attached to this incident, with a checkbox per step. |
| Happened around the same time | Deploys and other incidents around the same time, as suggestions based on timing only. |
| Check logs | The individual checks around the incident, with status and details. |
Acknowledge an incident#
Acknowledging tells your team (and your escalation policy) that someone is on it.
- Open the incident.
- Click Acknowledge.
When you acknowledge an incident:
- Any active escalation stops: no further escalation steps are sent.
- The acknowledgement is passed on to PagerDuty and Opsgenie if those integrations are connected.
- The incident shows Acknowledged by your name, and the timeline records it.
You can only acknowledge an ongoing incident, and only once. You can also acknowledge from the Acknowledge button on a Slack alert (see Slack, Teams & chat apps). If a flapping incident goes down again, it needs a fresh acknowledgement.
Snooze alerts#
Snoozing pauses escalation and repeat alerts for an ongoing incident without resolving it. It is useful when you know about the problem and are working on it.
- Open the incident.
- Under Snooze alerts for, pick 5 min, 15 min, 30 min, 1 h or 4 h.
- Click Snooze.
The incident shows Snoozed until the end time. Escalation steps that fall due during the snooze wait until it ends. To end it early, click Unsnooze. Resolving the incident clears the snooze.
Only the workspace owner can snooze an incident. Any member can unsnooze. Through the API, a snooze can be between 5 and 1,440 minutes.
Resolve an incident manually#
Normally you don't need to: SutramX resolves the incident when the monitor recovers. Resolve manually when you know the incident is over, for example after you paused or fixed a monitor that was failing for a reason you don't care about.
- Open the ongoing incident and click Resolve.
- Optionally fill in Resolution note (optional), for example what fixed it. Up to 5,000 characters.
- Click Resolve in the dialog.
The incident is marked Resolved manually, the down time is recorded, escalation ends, and a recovery alert is sent only if a down alert was delivered. Your resolution note is saved as an internal note.
Notes and public updates#
Use Notes on the incident page to keep a record of what your team found and did, and to post updates to your status page.
- Type in Add a note (up to 5,000 characters).
- Tick Public (show on status page) if this is an update for your customers. Leave it unticked for an internal note.
- Click Add note.
| Note type | Who sees it |
|---|---|
| Internal | Only your team, in the dashboard. |
| Public | Your team, and visitors of every status page that shows this monitor. It appears under the incident, without the author's name. |
Notes show the author and time. To remove a note, click the bin icon. Deleting a public note also removes it from your status pages.
Public notes are the way to post incident updates to a status page. You can write them yourself here, or let AI draft them and publish after review: see AI incident assist. If your status page offers extra languages, you can add translations of each public note: see Status pages overview.
How incidents appear on status pages#
An incident appears on every published status page that shows its monitor:
- While it is ongoing, at the top of the page under the status banner, as "Monitor name is unavailable", with its public notes.
- Once resolved, under Incident history as "Monitor name was unavailable", with how long it lasted and its public notes.
Internal notes, postmortems, alert details and who acknowledged the incident are never shown publicly. To put a monitor on a status page, see Status pages overview.
Root-cause information#
SutramX does not guess a root cause on its own, but the incident page collects the evidence:
- Incident insights shows Where it failed (each region's checks, failures, first failure and recovery), Error types (for example timeouts, DNS resolution failures, TLS errors, HTTP 5xx), the median response time before and during the incident, Deploys around the start and other monitors that Failed at the same time.
- Happened around the same time lists deploys and other incidents close in time, each with a heuristic score. These are matches by timing, not proof of cause.
- Check logs shows the raw checks around the incident.
- An AI summary can turn these facts into a short explanation with a likely cause and a confidence level, on plans that include it. See AI incident assist.
Deploy matches need deployment correlation, which depends on your plan. See Deployment correlation.
Runbooks#
A runbook is a checklist of steps for a type of incident. You create runbooks in the Runbook Automation section of Alerts → On-call & escalation. When an incident opens, SutramX suggests a runbook that matches the monitor's type or service tag; click Attach runbook on the incident page to use it. Tick each step as you complete it; the time it was completed is recorded.
Postmortems#
Once an incident is resolved, a Postmortem section appears on its page.
- Write your postmortem in Write-up. It supports simple Markdown:
#headings,-lists,**bold**and backticks for inline code. Up to 50,000 characters. - Click Save postmortem.
To change it later, click Edit. Saving an empty write-up clears the postmortem. The section shows when it was last updated.
You can only add a postmortem to a resolved incident. Postmortems are internal: they appear in the dashboard and in incident exports, but not on your status pages. On plans that include it, you can start from an AI-drafted postmortem: see AI incident assist.
Using the API#
Incidents are available through the REST API with an API key. For example, to list ongoing incidents:
curl "https://api.sutramx.com/incidents?status=ongoing&page_size=50" \
-H "Authorization: Bearer sk_your_api_key"The response is paginated:
{
"items": [
{
"id": "6f1c2a9e-3b7d-4c2a-9a51-0e8f4d2b7c10",
"monitor_id": "1a2b3c4d-5e6f-4a1b-8c9d-0e1f2a3b4c5d",
"monitor_name": "Checkout API",
"started_at": "2026-09-30T14:02:11.000Z",
"resolved_at": null,
"alert_suppressed": false,
"is_flapping": false,
"acknowledged_at": null
}
],
"total": 1,
"page": 1,
"page_size": 50
}The status filter accepts all, ongoing, resolved, suppressed and acknowledged. Other useful endpoints are POST /incidents/{id}/acknowledge, POST /incidents/{id}/resolve (optional body {"note": "..."}), POST /incidents/{id}/notes (body {"body": "...", "public": true}) and GET /incidents/export?format=csv. Snoozing needs the workspace owner and isn't available with an API key. See REST API and the API reference.
Common questions#
Why did a monitor fail but no incident open? One failed check isn't enough. The failure has to repeat for the monitor's failure threshold and be confirmed by enough regions. Until then the monitor may show as degraded. See Regions & confirmation.
Why is there an incident but I got no alert? Check the state chips and the timeline. A Suppressed incident was recorded without alerting, and the timeline shows why, for example "Down alert not sent — no verified alert contacts" or a maintenance window. The Notifications panel shows each channel's delivery status.
Why didn't I get a recovery alert? Recovery alerts are only sent when a down alert was delivered, so you never get a "recovered" message for an outage you weren't told about.
Can I delete an incident? No. Incidents are your outage history and feed your uptime numbers. You can resolve them and add notes.
Can I create an incident by hand? Not from the dashboard. Through the API, POST /incidents with a monitor_id opens an incident for a monitor that has no open incident. It is marked Suppressed — opened manually and sends no down alert unless SutramX's own checks later confirm the outage.
Related
Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.