Monitors

Regions & confirmation

How SutramX schedules checks, which probe regions run them, how a quorum of regions confirms an outage, timeouts, flapping, and how to allowlist the checker.

SutramX checks each monitor from one or more probe regions and, when a monitor uses more than one, only opens an incident when enough of them agree that it is down. This page explains how checks are scheduled, where they run from, how a monitor moves between Up, Degraded and Down, and how to let the checker through your firewall.

How checks are scheduled#

Every monitor has a Check interval. On each interval, each region assigned to the monitor sends its own request and records its own result. Regions do not share answers, so a network problem near one region never leaks into another region's result.

SettingValues
Check interval15 seconds, 30 seconds, 1, 2, 3, 5, 10 or 15 minutes
Fastest interval you can pickDepends on your plan. Faster options are not offered; the picker shows your plan's fastest interval
Default for new monitors5 minutes in the dashboard. Through the API, 2 minutes, or your plan's fastest interval if that is slower

Regions check on their own schedule, so their checks are not exactly in step. A monitor on a 1-minute interval checked from 3 regions gets about 3 checks a minute, spread across the minute.

Heartbeat monitors are the exception. SutramX does not probe them from regions; it evaluates them itself, from one source, against the schedule you entered.

Probe regions#

These are the probe regions SutramX runs right now:

RegionCodeCountryContinent
FRA1 (Frankfurt, Germany)fra1GermanyEurope
AZ (Arizona, USA)usa-az-probe——
IN (Mumbai)in-mumbaiIndiaAsia

Only locations with a live checker are offered. The New York location (nyc1) was retired on 2026-10-03, and monitors that used it were moved to Frankfurt (fra1) automatically. A request that still names nyc1 is rejected with a message telling you to use fra1 instead.

How many regions a monitor uses#

Your plan decides how many regions each monitor is checked from:

  • The table below shows how many of the live locations above each plan uses today.
  • If a location is retired, the next one in line takes its place.
PlanFreeStarterGrowthPro
Limits
Price (global, USD, incl. taxes)Free$12/mo$29/mo$59/mo
Price (India, INR, incl. GST)Free₹599/mo₹1,499/mo₹2,999/mo
Monitors50100200400
Minimum check interval3 min1 min30 s15 s
Probe regions per monitor1333
Regions that must agree before an incident opens1233
Status pages1310Unlimited
API keys13Unlimited
Team membersUnlimitedUnlimitedUnlimitedUnlimited
Multi-step API check steps3510
AI generations per month150500
Browser check runs per month5003,000
SMS credits per month25100300
WhatsApp credits per month1003001,000
Voice call credits per month25100
Alerting
Email alerts
Slack alerts
Discord alerts
Telegram alerts
Microsoft Teams alerts
SMS alerts
Voice call alerts
WhatsApp alerts
Webhook alerts
PagerDuty
Opsgenie
Escalation policies
On-call schedules
Browser push alerts
Status pages
Custom status page domains
White-label status pages
Private, password-protected status pages
Status pages behind SSO
Multi-language status pages
Reliability insights
SLO tracking
Monitor dependencies
Usage cost tracking
Weekly reports
Deployment correlation
Monitoring
DNS record monitors
MCP server monitors
Multi-step API checks
Browser checks
Website health checks
Third-party service status
Scheduled checks from Indian ISP networks, with ISP-only outage alerts
AI incident assist
AI incident summaries
AI postmortem drafts
AI-drafted status page updates
Account
SSO / SAML
Support tickets
Priority support

Live from the SutramX plan catalogue. Annual billing and current offers are on the pricing page.

See Plans & limits for what each plan includes.

Choose locations for a monitor#

You can choose which live locations check a monitor, for example only the regions where your users are. Every plan, Free included, can pick from all live locations; the plan sets how many each monitor uses.

  • When creating a monitor: use Check from in the Add new monitor dialog.
  • On an existing monitor: open it and click Locations, tick the locations you want, then click Save.

Rules:

  • You can pick any live location, up to your plan's number of locations per monitor, and you need at least one.
  • With no selection, the monitor uses your plan's default locations.
  • If none of your chosen locations is available any more, the monitor falls back to the plan's default locations rather than going unchecked.
  • On a plan with one location, clicking a different location swaps it in.

Through the API, send regions (a list of location codes) when you create a monitor, or use PUT /monitors/<id>/regions:

bash
curl -X PUT https://api.sutramx.com/monitors/<monitor-id>/regions \
  -H "Authorization: Bearer $SUTRAMX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "regions": ["fra1", "usa-az-probe"] }'

A location outside your plan is rejected with Monitoring location not available on your plan: … (available: …).

From a failed check to an incident#

A monitor goes Down (and an incident opens) only after all of these steps pass.

  1. Each region checks independently. A single check fails when, for example, the status code is unexpected, a keyword is missing, the response is too slow for your Max response time (ms), or the connection fails.
  2. Quick re-test. When a check fails with a network error, a server error (5xx) or a slow response, the same location re-tests once a few seconds later, inside the same interval, before recording the failure. Network errors are re-tested with a fresh DNS lookup. A connection reset that clears in a second never becomes a failed check. Failures that a re-test would not change, such as a TLS error, a 4xx status or a missing keyword, are not re-tested; other regions confirm them instead.
  3. Failure threshold. The region must have recorded the number of consecutive failed checks set in Failure threshold (1–10, default 1). Raise it for targets that are known to blip.
  4. Quorum. Enough live regions must report the monitor down at the same time. Each region's vote is its most recent result, if it is no older than two check intervals plus 30 seconds (at least 1 minute). A region with no result that recent has no vote.
  5. Alerting. The incident opens and your alert routing and escalation policy take over. Right before the down alert is sent, the monitor is re-checked once more (see How we verify an outage). Maintenance windows, quiet hours and monitor dependencies are applied at this point.

Quorum: how many regions must agree#

Each plan has a number of regions that must agree before an outage is confirmed. The live table above lists it per plan. If you pick fewer locations for a monitor than the quorum, the quorum is capped at the number you picked. A monitor checked from one location needs one failing region.

Regions vote, quorum decides
Example with three regions and a quorum of 2: a failed check is retried, and the incident opens only when enough regions report the monitor down at the same time.

Example: a monitor checked from 2 regions with a quorum of 2. If 1 region cannot reach your site but the other can, no incident opens. If both fail at the same time, it does. The failed checks from individual regions still show on the monitor's page.

On a monitor checked from more than one location, the header shows how many locations are reporting and how many must agree, for example 2/2 locations reporting · 2 must agree to confirm an outage.

When a probe region is offline#

Only regions with a live probe get a vote. If so many of a monitor's regions are offline that the normal quorum can no longer be reached, the monitor header shows Some locations not reporting so you can see that coverage is reduced, and a stricter rule applies so a real outage is still reported:

  • Every remaining live region must report the monitor down, on at least 2 consecutive checks each.
  • If the built-in checker has a fresh result and sees the monitor up, no incident opens.
  • When both hold, the incident opens and the alert and incident timeline say it was confirmed with reduced coverage, naming how many regions confirmed it and how many were missing.

A single failed check from the one remaining region is never enough on its own. Checker address changes, which can take a region offline briefly, are announced on sutramx.com/bot.

Recovery#

Recovery is deliberately easier than going down:

  1. The region must record the number of consecutive passing checks set in Recovery threshold (1–10, default 1).
  2. A majority of the monitor's live regions must see it up again, never more than the down quorum.

That way an incident does not stay stuck open just because one region is still offline.

Monitor states#

StateMeaning
UpThe latest check passed
DegradedThe monitor responded but slower than Degraded above (ms), or the latest check from a region failed without enough regions agreeing to open an incident
DownAn incident is open: the failure was confirmed by the quorum
MaintenanceA maintenance window covering the monitor is in progress
PausedYou (or a plan change) paused the monitor. No checks run
PendingNo check has finished yet
BlockedThe latest check was refused by the target's firewall or bot protection, so availability could not be verified. See How SutramX reports a block
InconclusiveThe latest check failed on SutramX's side, not the target's, and no region has a recent real result

Slow is not down. A response slower than Degraded above (ms) is recorded as degraded and counts as up for the quorum, so it never opens an incident. If slow should count as down, set Max response time (ms) instead; slower responses then fail the check. Both settings are under Show advanced options on HTTP and API monitors.

Timeouts#

Each check type has its own timeout. Whatever you set, a single attempt can never take longer than the check interval minus 3 seconds, so a check always finishes before the next one starts. The editor shows the effective value.

Monitor typeTimeout settingRangeDefault
HTTP and APITimeout (seconds)1–60 s30 s in the dashboard; 10 s if not set through the API
PingTimeout (seconds) per packet1–30 s10 s, with 4 packets per check (1–10)
TCP portConnect timeout (seconds)1–30 s10 s
UDP portResponse timeout (seconds)1–30 s10 s
DNSTimeout1–30 s5 s
Multi-step APITimeout for all steps together1–60 s30 s
MCP serverTimeout (seconds)1–60 s15 s
HeartbeatNone. Uses the grace period instead

A timed-out check is a failed check, and it goes through the same threshold and quorum steps as any other failure.

Flapping#

If a monitor goes down again within 5 minutes of recovering, SutramX reopens the same incident and marks it as flapping instead of opening a new one. The repeat down alert is held back so a flapping target does not flood your channels. If the monitor then stays down for the full 5 minutes, and the last thing you were told was "recovered", you are alerted again.

The flapping mark clears once the monitor has stayed up for 5 minutes. To reduce flapping alerts on a noisy target, raise the Failure threshold and Recovery threshold.

Allowlisting the checker#

If a firewall, WAF, CDN bot protection or rate limiter blocks SutramX, your monitor reports failures while the site is fine. Allowlist the checker in one of these ways.

By user agent#

Every check sends this user agent unless you set your own User-Agent header on the monitor:

http
User-Agent: Mozilla/5.0 (compatible; SutramX-Monitor/1.0; +https://sutramx.com/bot)

Match on the substring SutramX-Monitor. It survives version changes.

The website health crawler identifies itself separately as SutramX-SiteHealth.

By IP address#

The current checker IP addresses are published at:

  • sutramx.com/bot, with step-by-step instructions for Cloudflare, Vercel, AWS WAF, nginx, Apache and WordPress security plugins
  • sutramx.com/bot/ips.txt, a plain-text list with one address per line for scripts and firewall automation

The list changes when SutramX adds or retires a location. If you allowlist by IP, refresh it regularly, or new locations will be blocked when they go live.

By secret header#

The most precise option: add a custom request header with a secret value to your HTTP or API monitor (see HTTP & keyword monitors), then allow requests that carry that header in your firewall. It keeps working if checker addresses change.

How SutramX reports a block#

When a check is refused with HTTP 403, 429 or 503 and the response carries a known bot-protection signature (Cloudflare, Vercel, AWS WAF, Akamai, Imperva, Sucuri, DataDome, Kasada and others), SutramX records it as Blocked by bot protection (WAF) rather than as a plain 4xx or 5xx:

text
Blocked by Cloudflare bot protection (HTTP 403) — the target refused the check, so its availability could not be verified. Allowlist the checker in the target’s firewall: https://sutramx.com/bot#allowlisting

The check still counts as failed. If you cannot allowlist, you can stop it from alerting by unticking Blocked by bot protection (WAF) under Alert on these HTTP errors in the monitor editor. See Troubleshooting.

Common questions#

My site was down in one country but I got no alert. On a multi-region plan, a failure seen by fewer regions than the quorum does not open an incident. You can still see that region's failed checks on the monitor's page. If one market matters on its own, create a separate monitor checked only from that location.

Why did a region fail when the site was fine? Usually a regional network path problem, a CDN edge issue, or a firewall that blocks some checker addresses but not others. The quorum is there so these don't page you.

Do more regions mean more traffic to my site? Yes: one request per interval from each region, plus an occasional re-test after a failure. See the behaviour notes on sutramx.com/bot.

Can I make SutramX alert on the first failure from any region? Choose a single location for the monitor and keep Failure threshold at 1. You lose the protection against false alerts that the quorum gives.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.