How the planner calculates
- Worst-case detectionOne full interval: the failure starts just after a passing check.
- Average detectionHalf an interval, because a failure is equally likely to start at any point between checks.
- Requests per minute60 ÷ interval in seconds × monitors × regions. This is the steady load your checks put on your endpoints.
- Checks per monthThe same rate over a 30-day month: 2,592,000 seconds ÷ interval × monitors × regions.
The figures assume the outage persists and ignore confirmation. In practice, add the failure threshold (each extra required failure adds up to one more interval) and a few seconds for the regions to agree. When a check fails, SutramX also re-tests once from the same location about two seconds later before counting it, so a brief connection reset never becomes an alert.
A worked example
Twenty monitors checked every 60 seconds from 2 regions make 40 requests a minute, or about 1.73 million checks a month. An outage is noticed after 30 seconds on average and 60 at worst. Raise the failure threshold to 2 and the worst case becomes about two minutes. Halve the interval to 30 seconds with the same threshold and you are back to roughly one minute, with 80 requests a minute. Either is light load for a cheap health endpoint; the choice is about how many minutes of an outage you are willing to miss.
Detection time is half your outage
Total incident duration is detection plus response plus recovery. Teams pour enormous effort into the last two and then leave a five-minute check interval in place, which sets a floor on the first that no amount of on-call excellence can overcome.
If your checks run every five minutes, your average outage is already two and a half minutes old before anyone could possibly know. Moving to sixty seconds removes two minutes from every single incident you will ever have.
Not every endpoint deserves the same interval
- Revenue pathsCheckout, payment callbacks, and login. 15 to 30 seconds — the cost of downtime dwarfs the cost of checking.
- Core APIsAnything other services depend on. 60 seconds is a sensible default.
- Marketing pagesImportant but not transactional. 5 minutes is usually plenty.
- Batch and internal toolsLow urgency. 15 minutes — the longest interval SutramX offers — avoids noise for things nobody uses at 3am.
Multiply interval by monitor count by region count and the numbers grow quickly. Ten monitors at 30-second intervals across five regions is 100 requests a minute against your own infrastructure, continuously, forever. Make sure the endpoints you point at are cheap to serve.
Prefer cheap endpoints for frequent checks
A health endpoint that runs three database queries and renders a template is a bad monitoring target at high frequency — you are adding meaningful load precisely when the system is already struggling.
The better pattern is a lightweight endpoint that verifies critical dependencies with minimal work, returns a small response, and is excluded from caching. Save the expensive, realistic transaction checks for a slower interval.
Keep the timeout in step with the interval too. In SutramX a single attempt can never take longer than the interval minus 3 seconds, so a check every 15 seconds times out after at most 12 seconds. If your endpoint is sometimes slower than that, it is a poor fit for a very short interval.
Intervals on SutramX plans
The shortest interval you can set depends on your plan: 15 seconds on Pro, 30 seconds on Growth, 60 seconds on Starter and 3 minutes on Free. You can always choose a longer one, up to 15 minutes. Paid plans confirm a failure from more than one region before alerting, as described in regions and confirmation and on the multi-region monitoring page.
For a longer discussion, read choosing check intervals and reducing false alerts. Detection time feeds straight into MTTD, covered in MTTR and MTTD explained, and the downtime cost calculator puts a price on each minute.
Frequently asked questions
Why is average detection time half the interval?
A failure is equally likely to start at any moment between two checks. On average it starts halfway through, so it waits half an interval for the next check. In the worst case it starts just after a passing check and waits a full interval.
Does a failure threshold slow detection down?
Yes. A threshold of N needs N consecutive failed checks, which adds up to N − 1 extra intervals before an alert. A 30-second interval with a threshold of 2 detects about as fast as a 60-second interval with a threshold of 1, and it is much less likely to alert on a single blip.
Will frequent checks get blocked by my firewall or CDN?
They can, if a rate limiter or bot protection treats the checker as abusive. Allowlist the SutramX user agent or its published IP addresses so failures you see are real.
What about cron jobs and scheduled tasks?
They do not have an interval in this sense. A heartbeat monitor waits for your job to call it, and alerts when a run is due and no ping arrives within the grace period, which defaults to 10% of the time between runs.