Prove your uptime, do not just claim it
If you publish an SLA, at some point a customer will dispute a month. When that happens, availability data measured by your own infrastructure is not persuasive — it is the defendant testifying on its own behalf.
Third-party monitoring from outside your network gives you an independent record. It also protects you in the other direction: it lets you demonstrate that an incident a customer experienced was on their side, not yours.
Watch the auth path separately
Authentication tends to be the most heavily depended-upon service and the one most likely to fail in ways that are invisible to a health check. Token signing keys rotate. Session stores fill up. Identity provider integrations expire.
Monitor an actual login round trip against a synthetic account, not just whether the login page renders. The page rendering proves your frontend is alive; completing a login proves the system works.
Regional failures hit specific customers
A SaaS product with international customers will eventually have a regional problem: a CDN edge serving stale assets in one metro, a routing change that black-holes one country, a cloud zone degrading.
From a single monitoring location, all of these look perfectly healthy — while an entire customer segment cannot work. Checking from several continents is what surfaces them before the support tickets do.
Tie incidents to the deploy that caused them
SaaS teams ship often, and most incidents follow a change. On Growth and above, SutramX takes deploy events from GitHub, GitLab or any CI webhook and lists the likely culprit on the incident itself, next to SLO tracking that shows how much error budget the incident spent.
Because team seats are unlimited on every plan, support, product and engineering can all see the same incident timeline without anyone rationing logins.
A starting set of monitors
Login: a multi-step API check (on Starter, Growth and Pro) that posts the credentials of a synthetic account to your login endpoint, saves the returned token into a variable, then calls an authenticated endpoint such as /api/me and asserts that the response contains the account's ID. Keep the password in the check's encrypted Secrets, never in the step itself. On a plan without multi-step checks, an API monitor with a long-lived test token in a Bearer header covers most of the same ground.
Public API: an API monitor on a cheap, read-only endpoint. The JSON API guided preset is a good start: GET, expect 200, a content-type of application/json, forbidden keywords such as "error" and "stack trace", a 1,500 ms maximum response time and a failure threshold of 2. Add an assertion on one field your customers rely on, so a response that is technically valid but empty still fails.
Background jobs: billing runs, digest emails, data exports and queue workers. Give each a cron heartbeat monitor with its cron expression and time zone. The job calls its heartbeat URL when it succeeds, or adds ?status=fail when it fails, and SutramX alerts you if a run is due and no ping arrives within the grace period.
Vendors: follow the official status of your payment, email and cloud providers with third-party status (on Starter, Growth and Pro). On every plan, when a widely used vendor API starts failing for many SutramX customers at once, your incident is marked "Likely external", so you are not debugging your own code during someone else's outage.
Decide what pages a person
Not every failed check is worth waking someone. Page for login, the public API and anything that takes money. Send the marketing site, docs and internal tools to a chat channel. Use the Degraded above setting to mark slow responses without opening an incident, and a failure threshold of 2 on targets that occasionally blip, so one bad check never wakes anyone.
Escalation policies (on Growth and Pro) keep notifying until someone acknowledges: for example, post to Slack and email the on-call engineer immediately, then page PagerDuty if nobody has acknowledged after 10 minutes. On-call schedules on Pro decide who "the on-call engineer" is at any hour. Schedule maintenance windows for planned work, so the incident is still recorded but nobody is paged and your uptime is not charged for it.
Run a status page your customers trust
Add one component per thing a customer recognises (Login, API, Dashboard, Webhooks), not one per internal service. Group them into sections if you have several products. The page updates itself from your monitors, so an outage shows up without anyone remembering to post it.
During an incident, post short public updates: what is affected, what you are doing, and when you will update next. Visitors can subscribe by email with double opt-in, and they get a message when an incident starts and when it is resolved. Scheduled maintenance appears on the page in advance.
Every plan includes at least one status page (Free has 1 status page). Custom domains such as status.example.com are on Growth and Pro, and password-protected pages for a single enterprise customer are on Growth and Pro. The page is hosted by SutramX, away from your own infrastructure, so it stays reachable when your product is not.
Common questions
How quickly will we hear about an outage? Within roughly one check interval, plus the time it takes to confirm the failure. The shortest intervals are 15 s (Pro), 30 s (Growth), 60 s (Starter) and 3 min (Free). On paid plans an outage is confirmed by more than one region before anyone is alerted, which keeps one region's network problem from paging you.
Can SutramX data back an SLA report? Uptime is measured from outside your network, as the share of checks that passed. Daily uptime history is kept for 90 days on Free, 12 months on Starter and 24 months on Growth and Pro. SLO tracking with error budgets and burn rates is on Growth and Pro. Check history and incidents can be exported.
Should we monitor staging? Yes, but route it to its own channel and leave it out of escalation. A staging monitor that pages people at night teaches them to ignore pages.
Features, guides and tools for this setup
- Features: API Monitoring, Multi-Region Probes, Public Status Pages and Reliability Insights
- Guides: SLA, SLI & SLO Explained, Monitoring APIs Effectively, Status Page Best Practices, Incident Response Basics and On-Call Rotation
- Free tools: SLA Uptime Calculator, Error Budget Calculator and Incident Message Generator
- Integrations: Slack alerts setup, Microsoft Teams alerts setup and PagerDuty alerts setup
- Use cases: Monitoring for E-Commerce and Monitoring for Developers & Indie Hackers
- Comparisons: SutramX vs Better Stack and SutramX vs Opsgenie