Alerts & Incidents

How we verify an outage before alerting you

The checks SutramX runs before a down alert goes out, including a fresh re-check right before sending, what happens when the site answers again, and how to read it on the incident timeline.

Coming soon. This page documents a feature that is not generally available yet. Details may change before launch. Questions? Email support@sutramx.com.

A down alert should mean your site is really down. SutramX checks a failure several times, in several ways, before it tells you, and it checks once more right before the alert is sent. This page explains each step in plain words and how to see what happened on the incident page.

The full chain#

A failure has to get through every step below before a down alert reaches you.

  1. A quick re-test inside the same check. When a check fails with a network error, a server error (5xx) or a slow response, the same location tries again a few seconds later before recording the failure. A blip that clears in a second never becomes a failed check.
  2. Your failure threshold. The location must record as many failed checks in a row as the monitor's Failure threshold (default 1).
  3. Several locations must agree. On a monitor checked from more than one location, enough of them must see it down at the same time before an incident opens. See Regions & confirmation.
  4. A second opinion for single-location monitors. When a monitor is checked from one location only, a failure seen there is not treated as an outage straight away. SutramX asks other locations to check it once. If they see it up, no incident opens.
  5. A fresh re-check right before the alert. Once the incident has opened and the down alert is ready, SutramX checks the monitor one more time, from one of its locations, just before sending.

Steps 1 to 4 decide whether an incident opens. Step 5 decides whether the alert should still go out at that moment.

The re-check before sending#

The last re-check catches outages that ended while the earlier steps were running, such as a server that restarted or a deploy that finished.

  • Still failing: the alert goes out straight away and says that it was re-verified (see the example below).
  • The site answers: the alert waits briefly and SutramX keeps checking. If the site has really recovered, the alert is dropped and the incident resolves without paging anyone. If it fails again, the alert goes out.
  • Can't tell (for example the re-check itself couldn't run): the alert goes out. When in doubt, SutramX sends.

A down alert that was re-verified carries a line like this:

text
Re-verified at 12:03 UTC from Frankfurt: still failing (HTTP 503).

The re-check typically adds a few seconds to the alert.

Escalation waits while a monitor recovers#

If the monitor has an escalation policy and a step comes due while the monitor is answering again, the step waits instead of paging the next person. If the monitor recovers, the step never fires. If the outage continues, the step goes out as usual.

Certificate and domain expiry warnings#

Before an SSL certificate or domain expiry reminder is sent, SutramX looks at the certificate or domain again. If you have already renewed it, you don't get a warning for the old date.

Heartbeat and browser checks#

  • Heartbeat (cron) monitors: SutramX can't make your job run, but right before alerting it looks again for a ping. If your job checked in late (after the deadline but before the alert went out), the alert waits and is dropped once the monitor is back up.
  • Browser checks: a browser check runs a whole scripted journey, so it isn't re-run from another location before alerting. Its own result decides.

Alerts that arrive late are dropped#

Alerts are queued for delivery and retried if Slack, email or another channel is briefly unavailable. Just before each delivery SutramX checks the incident again. If the outage is already over, or a newer alert for the same outage has already gone out, the late "down" alert is not delivered, and that channel doesn't get the matching "recovered" message either. You never get a DOWN message after the RECOVERED one.

Other notices are double-checked too#

NoticeWhat is checked before it is sent
Checks are being blocked (a firewall or bot protection is refusing our checks)The block has to repeat: two blocked checks in a row from one location, or blocked from two locations
New broken pages or links (website health)Each newly broken address is loaded once more; anything that works again is left out, and if nothing is left no notice is sent
Lighthouse score droppedThe drop has to show again on the next scheduled audit
A provider reports a problem (third-party status)The problem has to still be there on the next poll of the provider's status page (a few minutes). Short blips are dropped. "Resolved" updates are never delayed
Certificate or domain expiryThe certificate or domain is looked up again if our data is more than a few minutes old

A real outage never goes unreported#

Checking twice protects you from false alarms. The opposite mistake, a real outage that nobody hears about, is worse, so SutramX also watches for it:

  • Every channel failed. If Slack, a webhook or another channel can't deliver a down alert after its retries, the workspace owner gets it by email (and on the phone app if installed), marked Delayed alert and naming the channel that failed. This happens even if you turned down emails off because a channel carries your alerts. You also get a separate email that the integration is failing.
  • The alert never went out. An open outage with no delivered alert after about 15 minutes is sent again, or sent to the owner if the channels are broken.
  • Nobody is set up to receive alerts. The owner is emailed.
  • A maintenance window or quiet hours ended while the outage was still going. The held alert is sent even if no new check comes in.
  • A dependent check stays silent only while its parent monitor's own alert actually reached someone. If the parent's alert didn't go out, the child alerts.
  • Checks stopped running for a monitor. SutramX restarts them, and if that doesn't help you get one notice: "We couldn't check Shop for 12 minutes". This is never shown as a fake outage.
  • Some locations agree, others don't. If a site has been failing from some check locations for 15 minutes without enough locations agreeing to open an incident, you get one Partial outage notice, for example "Shop is down from Mumbai and Singapore for 15 minutes". It doesn't count against uptime, and it is never sent when the problem is our own checker.
  • A server that keeps answering "too many requests" (HTTP 429) without a bot-protection page counts as failing after 3 checks in a row from the same location. A real bot or firewall challenge is still shown as blocked, not down.

When the owner got an outage through this safety net, they also get the "recovered" message when it ends. You only hear "recovered" for alerts you actually received.

Reading the incident timeline#

Every step is recorded on the incident's timeline, so you can see why an alert came when it did. The Why this alert panel sums it up.

Timeline eventWhat it means
Outage re-checked — site answered, alert held back (down_alert_verifying)The fresh re-check right before sending found the site answering, so the alert is waiting while SutramX checks again. Shows the location and the attempt number
Outage re-verified — still failing, alert sent (down_alert_verified)The re-check confirmed the site is still failing, with what it saw (for example HTTP 503 or a timeout), and the alert went out. Also recorded when the re-check couldn't tell, because then the alert is sent anyway
Outage not settled by re-checks — alert sent anyway (down_alert_verify_timeout)The site kept going back and forth until the wait limit, so the alert was sent anyway. Shows how many re-checks ran and how long the alert waited
Escalation step held — the monitor is recovering (escalation_step_deferred)An escalation step came due while the monitor was recovering, so it is waiting rather than paging (at most twice, a minute each)
Delayed alert sent to the workspace owner (down_alert_missed_recovered)The safety net found that the down alert had reached nobody and alerted the owner directly (or re-sent it through your channels)
Recovery sent to the workspace owner (recovery_fallback_sent)The owner got this outage through the safety net, so they also got the "recovered" message
Late down alert not sent (down_alert_dropped_stale)A queued down alert was still waiting for delivery on one channel when the outage ended (or a newer alert replaced it), so it was dropped

The names in brackets are the event types, for anyone reading incident timelines through the REST API.

Common questions#

Will verification make my alerts slow? Usually it adds a few seconds. It can add more only while your site is answering again, and never more than about 5 minutes.

Can a real outage be hidden? No. If the site keeps failing, or SutramX can't tell, the alert is sent. An alert is only dropped when the site has really recovered.

I still got an alert for a short blip. If the site was down when the re-check ran, the alert is correct for that moment. To ignore short blips, raise the Failure threshold or check from more locations. See Monitor settings.

Do I need to turn this on? No. It happens automatically; there is nothing to set up.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.