Alerts & Incidents

How alerting works

When SutramX sends an alert, who receives it, how duplicates and flapping are handled, and how maintenance and quiet hours silence alerts.

SutramX alerts you when a monitor goes down and again when it recovers. This page explains the whole path, from a failed check to a message on your phone, so you can predict exactly who gets notified, on which channel, and when.

The alert lifecycle#

Every alert belongs to an incident. A monitor that fails its checks opens an incident, and the incident drives the notifications:

  1. A check fails. One failed check never alerts on its own. SutramX waits until the monitor's failure threshold and its multi-region confirmation agree that it is really down (see Regions & confirmation).
  2. An incident opens. SutramX records the incident and decides whether the down alert may go out now. It may be held back by a maintenance window, quiet hours, a dependency, an error-type filter or your notification preferences (see When alerts are held back).
  3. The down alert goes out. Right before sending, SutramX re-checks the monitor once more; if the site already answers again, the alert waits briefly and is dropped if it really recovered (see How we verify an outage). It goes to every email recipient of the monitor and every connected channel whose routing covers the monitor, all at the same time. Browser push is sent too.
  4. Escalation starts, if the monitor has an escalation policy. Steps fire on their delays until someone acknowledges or the incident resolves (see Escalation & on-call).
  5. The monitor recovers. Once the recovery threshold and confirmation agree, the incident resolves and a recovery alert goes to the same places, with the total downtime.

Alerts are delivered reliably: a restart on our side never drops one. Every delivery attempt is recorded on the incident page under Notifications with the status Queued, Sent, Retrying, Failed or Skipped.

One incident, every channel at once
The down alert goes to every email recipient and every connected channel whose routing covers the monitor, all at the same time.

Events that notify you#

EventEmailChat apps & webhookPagerDuty / OpsgenieSMS, WhatsApp, voicePush
Monitor down (incident opened)YesYesTrigger / createYesYes
Monitor recovered (incident resolved)Yes, if enabledYesResolve / closeYesYes (optional)
Incident acknowledgedNoNoAcknowledgeNoNo
Escalation stepEmail stepsIf the step targets the channelIf the step targets itIf the step targets itNo
SSL certificate / domain expiringIf enabledYesNoNoNo
Maintenance started / endedIf enabledYesNoNoNo
Website health issues, third-party status changesYesYesNoNoNo
ISP outage from last-mile checksYesYesNoNoNo

"Chat apps & webhook" means Slack, Microsoft Teams, Discord, Google Chat, Mattermost, Telegram and custom webhooks (Zapier receives incident alerts only). "Push" means browser push. Paging tools (PagerDuty, Opsgenie) and ticketing (GitHub issues) only ever receive real incidents, never notices.

Where alerts go#

There are two independent fan-outs for every alert.

Email recipients#

Email alerts for a monitor go to the union of:

  • The monitor's own recipients, set on Alerts → Recipients or from a monitor's menu with Alert recipients….
  • The monitor group's recipients, set with Edit group emails on Alerts → Recipients. They get alerts for every monitor in the group.
  • The workspace's verified alert contacts: the workspace owner's account email, added automatically at sign-up. They get email alerts for every monitor.

Each address receives one email per alert, even if it appears in several lists. If no address on the monitor is deliverable, the alert goes to the email of the account that owns the monitor.

Connected channels#

Channels are connected once per workspace on Alerts → Channels & API. Each connection has its own routing:

Routing optionWho it alerts for
All monitors (default)Every monitor in the workspace
Selected groupsMonitors in the chosen monitor groups
Selected monitorsOnly the chosen monitors

To change it, open the connection's menu, choose Routing, pick an option and click Save routing. You can connect up to 10 connections of each channel type (for example, three Slack channels routed to different teams).

Recovery alerts to PagerDuty, Opsgenie and GitHub ignore routing: if a page was opened, it is always resolved, even if you changed the routing in between.

Push notifications#

Browser push is personal. It goes to the browsers of workspace members who turned it on, for every monitor in the workspace. See Email & push.

Recipient confirmation#

SutramX only emails addresses that agreed to receive alerts. When you add an address that isn't a teammate's verified account email:

  1. SutramX sends it a one-time Confirm alert emails message.
  2. The recipient clicks Yes, send me alerts (or No, decline these emails).
  3. Until they confirm, the address is listed as Pending verification and receives no alerts.
RuleValue
Confirmation link valid for72 hours
Wait before resending to the same address10 minutes
Confirmation emails per workspace50 per rolling 24 hours
Confirmation emails per address (across all workspaces)5 per rolling 24 hours

Teammates with a verified account email are confirmed automatically. If a confirmation expires, click Resend next to the address on Alerts → Recipients.

Every alert email has an unsubscribe link (and a one-click unsubscribe header). An address that unsubscribes is shown as Unsubscribed and stops receiving alerts from that workspace. Only the recipient can resubscribe, from the same link. An address that declined is shown as Declined.

Deduplication#

SutramX makes sure one event produces one notification per destination:

  • Each incident sends at most one down alert and one recovery alert to each email address and each channel connection. A retried check, an overlapping worker or a restart can't page you twice.
  • Retries after a provider error reuse the same delivery, so a retried SMS or call isn't sent (or charged) twice.
  • A recovery alert is sent only if the down alert actually went out. You never get a "recovered" message for an outage you were never told about.

Flapping monitors#

A monitor that goes down again shortly after recovering is flapping. Instead of opening a new incident each time, SutramX reopens the incident it just resolved and suppresses the repeat down alert during a short cooldown (five minutes by default).

If the monitor then stays down for a full cooldown after the reopen, and the last thing you heard was "recovered", SutramX sends the down alert again so a real outage is never hidden. If you were last told "down", no extra page is sent.

Repeat alerts and reminders#

SutramX doesn't repeat the down alert on a timer while an incident stays open. To keep notifying people until someone responds, assign an escalation policy to the monitor: each step fires after its delay until the incident is acknowledged or resolved.

When alerts are held back#

A down alert can be held back for these reasons. The incident is still recorded either way.

ReasonWhat happens to the alert
Active maintenance window covering the monitorHeld. If the monitor is still down when the window ends, the down alert is sent then.
Quiet hours or a deployment windowHeld. Sent once the window ends if the monitor is still down.
The monitor's parent dependency is already downHeld while the parent is down. Sent afterwards if this monitor is still down.
The failure's HTTP error type isn't selected in Alert on these HTTP errorsNot sent for this incident.
Monitor Down Alerts is turned off in Notification PreferencesNot sent on any channel.
A flapping cooldownSee Flapping monitors.
The incident is snoozedNo re-alerts or escalation until the snooze ends. See Incidents.

Escalation steps that come due during a maintenance window, quiet hours or a deployment window are postponed and re-checked every 5 minutes, so on-call isn't paged during planned work.

Quiet hours and deployment windows#

Set these up in the Quiet Hours & Deployment Windows section of Alerts → On-call & escalation:

FieldWhat it doesDefault / limits
Window nameLabel shown in the listRequired, up to 255 characters
TypeQuiet hours or Deployment window (they behave the same for alerting)Quiet hours
TimezoneTimezone the start and end times are inUTC
Start time / End timeDaily window; an end before the start runs overnight (for example 22:00 to 06:00)Required
Days of weekDays the window starts on; leave all unselected to apply every dayEvery day
Applies toAll monitors, Selected monitors or Selected groupsAll monitors

Only the workspace owner can create, edit or delete windows.

Notification preferences#

The Notification Preferences section of Alerts → On-call & escalation has workspace-wide switches:

PreferenceControlsDefault
Monitor Down AlertsDown alerts on every channel. Off means no down alert and no escalation.On
Monitor Recovery AlertsRecovery emails only. Channels always get the recovery so pages get resolved.On
Weekly ReportsThe weekly summary emailOn
SSL Certificate ExpiryExpiry reminder emails only. Chat channels get reminders whenever the monitor's expiry notifications are on.Off
Maintenance AlertsMaintenance start/end emails onlyOn

Testing your setup#

  • A monitor's email recipients: on the monitor page, open the menu and choose Send test notification. It emails the monitor's first alert recipient.
  • A channel: on Alerts → Channels & API, open the connection's menu and choose Send test. Most channels are tested automatically right after you connect them.
  • Push: use Send test in the Browser push alerts card.

Channel availability by plan#

Email, Slack, Microsoft Teams, Discord, Google Chat, Mattermost, Telegram and push are included on every plan. Webhooks, Zapier, GitHub, PagerDuty, Opsgenie, SMS, WhatsApp, voice calls, escalation policies and on-call schedules depend on your plan. A locked channel shows the plan it needs on its card. See Plans & limits for the current table.

Channel access is checked again at send time. If your plan changes, alerts to channels it no longer includes are skipped, and the skip is recorded on the incident.

A channel card that says Not available yet means the channel can't be used on SutramX right now. Nothing needs to change on your account.

Troubleshooting#

  • No email arrived: check that the address shows Confirmed on Alerts → Recipients, isn't Unsubscribed, and that Monitor Down Alerts is on. The incident's Notifications list shows every attempt and why it was skipped.
  • No SMS, WhatsApp or call: the number must be verified with the 6-digit code sent to it when you connected it. See SMS, WhatsApp & voice.
  • A channel shows Failing: the last 3 deliveries failed, and the workspace owner was emailed about it. The connection shows Last success and the last error. Fix the cause (for example a deleted webhook), then use Send test.
  • The alert came late: check whether a maintenance window, quiet hours or a dependency held it back, or whether the site answered the last re-check before sending (see How we verify an outage). The incident timeline records the reason.
  • No recovery alert: a recovery is only sent if the down alert went out. Recovery emails also need Monitor Recovery Alerts turned on.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.