Troubleshooting
Fix common SutramX problems: blocked checks, false or missing alerts, heartbeats, SSL, timeout and DNS errors, custom domains, API errors and payments.
Find your symptom below. Each entry lists what you see, why it happens and how to fix it. Error messages are quoted exactly as SutramX shows them, so you can search this page for the text you see.
When you open a failing monitor, the check history shows the error for each region. Start there: it usually tells you which section below applies.
Monitor blocked by a firewall, WAF or Cloudflare#
Symptoms
- Checks fail with Blocked by bot protection (WAF) and a message like
Blocked by Cloudflare bot protection (HTTP 403) — the target refused the check, so its availability could not be verified. - Or checks fail with
Received HTTP 403,Received HTTP 429orReceived HTTP 503while the site works in your browser. - Some regions pass and others fail.
Cause. Your firewall, CDN bot protection (Cloudflare, Vercel, AWS WAF, Akamai, Imperva, Sucuri and others), a WordPress security plugin or a rate limiter is refusing SutramX's requests. Datacenter traffic is often challenged by default.
Fix
- Allowlist the checker. The simplest rule matches the user agent substring
SutramX-Monitor. For a stricter rule, allowlist the IP addresses at sutramx.com/bot/ips.txt. - Follow the step-by-step instructions for Cloudflare, Vercel, AWS WAF, nginx, Apache and WordPress at sutramx.com/bot. On Cloudflare, see the notes below.
- Or add a custom request header with a secret value to the monitor and allow that header in your firewall. See HTTP & keyword monitors.
- Exempt the checker from rate limiting too. Every region a monitor uses checks on each interval, which can trip a strict limit.
On Cloudflare:
- Free plan with Bot Fight Mode on: Bot Fight Mode cannot be bypassed by WAF custom rules, a Skip action or Page Rules, and a user agent rule does not help. Add an IP Access Rule with the action Allow for each address in /bot/ips.txt (Security → WAF → Tools → IP Access Rules); Bot Fight Mode does not challenge a request an IP Access Rule has already allowed. Or turn Bot Fight Mode off.
- Pro and above (Super Bot Fight Mode): create a custom rule matching the SutramX addresses with the action Skip and select All Super Bot Fight Mode rules.
- A WAF custom rule expression on its own only skips managed rules, rate limiting and your custom rules, not Bot Fight Mode.
- SutramX has applied to Cloudflare's Verified Bots program and signs its requests with Web Bot Auth; until Cloudflare approves it, use the rules above.
If you can't change the firewall, untick Blocked by bot protection (WAF) under Alert on these HTTP errors in the monitor editor. Checks still fail, but no alerts are sent. See Regions & confirmation.
False or unexpected down alerts#
Symptoms. You got a down alert, but the site looked fine when you checked.
Causes and fixes
| Cause | Fix |
|---|---|
| Firewall or bot protection blocked the checker | See the section above |
| A short blip (deploy, restart, cold start) that cleared within a minute | Raise Failure threshold to 2 or 3 under Show advanced options. The monitor must then fail that many checks in a row before alerting |
| Timeout too short for a slow endpoint | Raise Timeout (seconds). See Timeouts |
| Max response time (ms) is set and the site was slow | Use Degraded above (ms) instead if slow should show amber but not alert |
| Keyword or status code rule no longer matches the page | Check Expected status codes, Required keyword and Forbidden keywords against what the page returns now |
| The site really was down in some places | Check which regions failed in the incident: "Confirmed down from: …" |
| Monitor on a single location | A single location has no other region to confirm a failure. Use more locations if your plan includes them |
On multi-region plans, an incident only opens when enough regions agree. See Regions & confirmation.
A monitor keeps going up and down#
If a monitor goes down again within 5 minutes of recovering, the incident is reopened and marked as flapping, and the repeat alert is held back. Raise Failure threshold and Recovery threshold to smooth out an unstable target, and look for the underlying cause (an overloaded server, a failing instance behind a load balancer).
A monitor shows Degraded#
Degraded means either the response was slower than Degraded above (ms), or a check failed in some region without enough regions agreeing to open an incident. Degraded never sends a down alert. Open the monitor to see which region failed and why.
Not receiving alerts#
Symptoms. An incident opened (it shows on Incidents), but you got no email, chat message or page.
Open the incident first. If its alert was held back, it says why:
| Reason shown | Meaning | Fix |
|---|---|---|
| maintenance window | A maintenance window covered the monitor | Expected. If it is still down when the window ends, the alert goes out then |
| quiet hours / deployment window | A quiet-hours or deployment window was active | Check Alerts → On-call & escalation → Quiet Hours & Deployment Windows |
| notification preferences | Monitor Down Alerts is switched off | Turn it on under Alerts → On-call & escalation → Notification Preferences |
| alert filter | The error class is unticked under Alert on these HTTP errors | Tick it in the monitor editor |
| upstream dependency down | A parent monitor was already down | Expected; you were alerted for the parent |
| flapping cooldown | The monitor recovered and failed again within 5 minutes | You are alerted again if it stays down for 5 minutes |
If nothing was held back, check delivery:
- Email recipients are verified. Alerts only go to verified addresses. New recipients get a one-time confirmation email that must be accepted. Check the alert recipients list.
- Spam folder. Add SutramX's sending address to your contacts or allowlist.
- Unsubscribed. If a recipient used the unsubscribe link in an alert email, they no longer get alert email for the workspace.
- Chat and webhook channels. Check the connection is still active, and that it is routed to this monitor or its group.
- SMS, WhatsApp and voice. These use alert credits. When credits run out, those channels stop sending. See Alert credits.
- Escalation policy. If the monitor uses an escalation policy, later steps only fire after their delay and stop once someone acknowledges.
Use Send test notification in the monitor page's menu to test delivery end to end.
No recovery alert#
A recovery alert is only sent if a down alert was sent for that incident. If the outage started and ended inside a maintenance window or quiet hours, neither is sent. With Monitor Recovery Alerts switched off, recovery emails stop but connected channels still get the recovery.
Heartbeat not received#
Symptoms. A heartbeat monitor reports No heartbeat received for the run expected at … although the job ran, or Last heartbeat says "None received yet".
Test the URL by hand from the machine that runs the job:
curl -v "https://YOUR-HEARTBEAT-URL"| Response | Cause | Fix |
|---|---|---|
200 OK but the monitor still alerts | The ping arrives later than the deadline, the schedule is wrong, or the monitor is paused (pings to a paused monitor are ignored) | Check the schedule and grace period below, and resume the monitor |
404 Unknown heartbeat URL | The URL was rotated, mistyped, or the monitor was deleted | Copy the current URL from the monitor's Heartbeat panel |
400 status must be "fail" or "ok" | Unsupported status value | Use ok or fail. See Heartbeat & cron monitors |
429 Too many heartbeats for this monitor; slow down | More than 60 pings a minute to one URL | Ping once per run |
No response, timeout or Could not resolve host | The job's host can't reach the internet, or an outbound proxy or firewall blocks it | Allow outbound HTTPS to the SutramX API host |
Other common causes:
- Time zone. The cron expression is always UTC. A job scheduled at 02:00 local time in a UTC+5:30 zone runs at 20:30 UTC the day before. Convert the schedule to UTC.
- Grace period too short. Ping at the end of the job and set Grace period (minutes) longer than the job's longest normal run.
- The ping never runs. With
job && curl …, curl only runs if the job exits with code 0. Check the job's own logs. - curl is missing in a minimal container image. Install it or use another HTTP client.
- The job reported a failure. A
?status=failping keeps the monitor down until a successful ping arrives. The check showsJob reported failure at ….
SSL and TLS errors#
Symptoms. HTTP checks fail with one of these errors:
| Error | Cause | Fix |
|---|---|---|
TLS certificate has expired | The certificate's end date has passed | Renew the certificate and check auto-renewal |
TLS certificate hostname does not match the requested host | The certificate doesn't cover the host in the monitor's URL | Use a certificate that includes the host, or monitor the host the certificate is issued for |
TLS certificate is self-signed | Self-signed certificate, or a private CA | Use a publicly trusted certificate. For internal test systems, untick Verify TLS certificates |
TLS certificate has been revoked | The CA revoked the certificate | Issue a new certificate |
A common cause of "it works in my browser" is a server that doesn't send its intermediate certificate. Browsers often fill the gap; other clients don't. Configure the server to send the full chain.
Unticking Verify TLS certificates makes the check pass despite certificate problems, so only do it for systems where that is acceptable. For warnings before a certificate expires, see SSL & domain expiry.
Timeouts#
Symptoms. Checks fail with:
Connection timed out before establishing a socket to the upstream host: no connection could be made. Usually a firewall silently dropping the checker's traffic, or the server is down.Read timed out while waiting for the upstream response body: the server connected but was too slow to send the response.Request timed out before a complete response was received: the whole request took longer than the timeout.
Fix
- If only some regions time out, it's often a firewall or geo-blocking rule. Allowlist the checker (see above).
- Raise Timeout (seconds) (1–60 s for HTTP and API).
- Remember the timeout is capped at the check interval minus 3 seconds. On a 15-second interval the effective timeout is at most 12 seconds; the editor shows the effective value. Use a longer interval for slow endpoints.
- Point the monitor at a lightweight health endpoint rather than a heavy page.
DNS failures#
Symptoms. Checks fail with DNS resolution failed, or a monitor can't be saved with We couldn't find … Check the address for typos, or that its DNS record exists.
Causes and fixes
- Typo in the host name. Check the URL or host on the monitor.
- The record doesn't exist publicly. SutramX resolves names with public DNS. Names that exist only on your internal DNS or in a hosts file can't be checked.
- Recently changed records. Wait for the record's TTL to pass.
- DNS provider outage. If all regions fail at once for several monitors on the same domain, check your DNS provider's status.
- Expired domain. See SSL & domain expiry.
Target resolves to a non-public address and cannot be monitored means the host resolves to a private, loopback or internal IP address. SutramX only checks public addresses. Expose a public health endpoint, or use a heartbeat monitor where your internal service pings SutramX instead.
Status page custom domain not working#
Symptoms. Verify fails, or the custom domain doesn't load your status page.
Custom domains need two DNS records, both shown in the status page's Custom domain panel:
| Type | Name | Value |
|---|---|---|
| TXT | _sutramx-verify.<your domain> | The verification token shown in the panel |
| CNAME | <your domain> | The target shown in the panel |
| Message | Fix |
|---|---|
We could not find the TXT record yet. | Add the TXT record. DNS changes can take a few minutes, sometimes up to an hour. Some DNS providers add your domain automatically, so enter only the part before it (for example _sutramx-verify.status) |
A TXT record exists but its value does not match. | Copy the value exactly, without quotes or spaces |
The TXT record is correct, but the CNAME does not point to the target yet. | Point the CNAME to the target shown. On Cloudflare, set the record to "DNS only" (grey cloud); a proxied record hides the CNAME |
The DNS lookup failed or timed out. | Temporary. Try again in a few minutes |
Custom domain is already in use | Another status page already uses the domain. Remove it there first |
Use a domain you own, such as status.yourcompany.com | SutramX's own domains can't be used |
After verification, the HTTPS certificate is issued automatically on the first HTTPS visit, which can take a moment. Use Check SSL in the panel to see its state. Custom domains are a plan feature (Growth and Pro); after a downgrade the domain stops serving, and the page stays available at its /status/ URL. See Branding & custom domains.
API errors#
Errors are returned as JSON with an error message and, for most errors, a code:
{ "error": "monitors limit reached", "code": "ENTITLEMENT_LIMIT_REACHED", "details": { "resource": "monitors", "plan": "free", "current": 50, "limit": 50, "upgradePath": "/subscription" } }400 Bad Request#
{"error": "Validation failed", "errors": [{"field": "interval_seconds", "message": "…"}]} lists each invalid field. Other 400 errors explain the problem in error, for example Invalid interval. Minimum is 60 seconds for starter plan.
401 Unauthorized#
error | Fix |
|---|---|
No token provided | Send the key as Authorization: Bearer sk_…. The Bearer prefix is required |
Invalid API key | The key is wrong or was deleted. Create a new one |
This API key is disabled because your plan's API key limit was reduced (API_KEY_DISABLED) | Your plan now allows fewer keys. Upgrade, or delete other keys and use one that is still active |
curl https://api.sutramx.com/monitors -H "Authorization: Bearer $SUTRAMX_API_KEY"403 Forbidden#
code | Meaning | Fix |
|---|---|---|
FEATURE_NOT_AVAILABLE | The feature isn't in your plan. details names the plan you need | Upgrade, see Plans & limits |
ENTITLEMENT_LIMIT_REACHED | You hit a plan limit, such as monitors or status pages | Delete something you don't need, or upgrade |
WORKSPACE_OWNER_REQUIRED | Only the workspace owner can do this. API keys never act as the owner | Ask the owner, or do it from the owner's dashboard session |
SESSION_AUTH_REQUIRED | The endpoint needs a signed-in session; API keys are not accepted | Use the dashboard |
WORKSPACE_ACCESS_DENIED | For example API key is not valid for the requested workspace: an API key only works for the workspace it was created in | Don't send a different X-Workspace-Id, or use that workspace's key |
This API key cannot manage alert channels or exports. Create an API key with automation access in Settings → API keys. means the key's access level is too low for that endpoint. See REST API.
409 Conflict#
WORKSPACE_PENDING_DELETION: This workspace is scheduled for deletion. Cancel the deletion to make changes.
429 Too Many Requests#
You sent too many requests in a time window. Responses include standard RateLimit-* headers; wait for the reset and retry with backoff. Limits apply per endpoint group, for example:
| Endpoint group | Limit |
|---|---|
| Creating, updating and deleting monitors | 120 requests per 15 minutes per IP address |
| Run check now and test checks | 30 per 5 minutes per user |
| Test notifications | 10 per 15 minutes per user |
| Heartbeat pings | 60 per minute per heartbeat URL, 600 per minute per IP address |
Payment problems#
Payment failed. The dashboard shows "Your last payment didn't go through". Your monitors keep running while the payment is retried. Go to Plans & billing → Billing and click Update payment method (on Razorpay, Complete payment). If the grace period ends without a successful payment, the account moves to the Free plan.
Monitors paused after a downgrade or failed payment. Monitors over your plan's limits are paused, not deleted. On Free, a banner says how many. Upgrade to resume them, or choose which ones run with Choose monitors to resume on Billing.
Upgrade still pending. "Your plan change is still waiting for payment confirmation." The payment provider can take a few minutes to confirm. You stay on your current plan until it succeeds. If your browser blocked the payment page, use the "Open payment page" link.
Card or UPI declined. "The payment didn't go through, so nothing changed." Try another card or payment method, or check with your bank. SutramX never sees your card number, CVV or UPI PIN, so we can't see why a bank declined.
For invoices, refunds and plan changes see Payments & invoices and Upgrade, downgrade & cancel.
Website health run failed#
| Message | Fix |
|---|---|
robots.txt disallows … for our crawler. | Allow SutramX-SiteHealth in robots.txt, or choose another start page |
robots.txt returned HTTP 5xx, so the site was not crawled | Fix robots.txt; an erroring robots.txt means "crawl nothing" |
The start page returned HTTP … / The start page could not be loaded: … | Check the Start URL and that the crawler isn't blocked |
A check of this site is already running | Wait for the current run to finish |
See Website health.
Still stuck?#
Email support@sutramx.com with the monitor or status page name, the time of the problem (with time zone) and the exact error message. See Getting help.
Related
- Regions & confirmation
- How alerting works
- Heartbeat & cron monitors
- REST API
- FAQ
- Guides: Why Your Monitor Reports 403 While Your Site Works, Why Is My Website Down? and Site Not Loading on Jio or Airtel
- Free tools: Is It Down?, DNS Checker and HTTP Status Code Reference
- More from SutramX: Checker IP addresses and user agent
Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.