SaaS Platforms

Monitoring for SaaS Platforms

When you sell software as a service, availability is not a technical metric. It is the product. Your customers built workflows on top of you, and an outage stops their business as well as yours.

What tends to go wrong

  • Your SLA is contractualEnterprise agreements carry financial penalties. You need an independent record of availability, not just your own server logs.
  • Login is the whole productIf authentication is down, everything is down — regardless of how healthy the rest of your fleet looks.
  • Customers integrate your APIA breaking change or slow endpoint cascades into your customers' systems, and they will notice before you do.
  • Multi-tenant blast radiusA problem affecting one tenant shard can be invisible in aggregate metrics while a subset of customers is completely down.

What to monitor

EndpointWhy it mattersSuggested interval
Login / authenticationTotal product outage if unavailable. Highest priority endpoint you own.15–30s
Public API rootCustomer integrations depend on it. Assert on response shape, not just status.30–60s
Application dashboardThe main authenticated surface your users load daily.60s
Webhook delivery endpointSilent failures here break customer automations without any visible error.60s
Marketing siteNot transactional, but a down homepage damages trust and blocks signups.5m

Prove your uptime, do not just claim it

If you publish an SLA, at some point a customer will dispute a month. When that happens, availability data measured by your own infrastructure is not persuasive — it is the defendant testifying on its own behalf.

Third-party monitoring from outside your network gives you an independent record. It also protects you in the other direction: it lets you demonstrate that an incident a customer experienced was on their side, not yours.

Watch the auth path separately

Authentication tends to be the most heavily depended-upon service and the one most likely to fail in ways that are invisible to a health check. Token signing keys rotate. Session stores fill up. Identity provider integrations expire.

Monitor an actual login round trip against a synthetic account, not just whether the login page renders. The page rendering proves your frontend is alive; completing a login proves the system works.

Regional failures hit specific customers

A SaaS product with international customers will eventually have a regional problem: a CDN edge serving stale assets in one metro, a routing change that black-holes one country, a cloud zone degrading.

From a single monitoring location, all of these look perfectly healthy — while an entire customer segment cannot work. Checking from several continents is what surfaces them before the support tickets do.

Tie incidents to the deploy that caused them

SaaS teams ship often, and most incidents follow a change. On Growth and above, SutramX takes deploy events from GitHub, GitLab or any CI webhook and lists the likely culprit on the incident itself, next to SLO tracking that shows how much error budget the incident spent.

Because team seats are unlimited on every plan, support, product and engineering can all see the same incident timeline without anyone rationing logins.

A starting set of monitors

Login: a multi-step API check (on Starter, Growth and Pro) that posts the credentials of a synthetic account to your login endpoint, saves the returned token into a variable, then calls an authenticated endpoint such as /api/me and asserts that the response contains the account's ID. Keep the password in the check's encrypted Secrets, never in the step itself. On a plan without multi-step checks, an API monitor with a long-lived test token in a Bearer header covers most of the same ground.

Public API: an API monitor on a cheap, read-only endpoint. The JSON API guided preset is a good start: GET, expect 200, a content-type of application/json, forbidden keywords such as "error" and "stack trace", a 1,500 ms maximum response time and a failure threshold of 2. Add an assertion on one field your customers rely on, so a response that is technically valid but empty still fails.

Background jobs: billing runs, digest emails, data exports and queue workers. Give each a cron heartbeat monitor with its cron expression and time zone. The job calls its heartbeat URL when it succeeds, or adds ?status=fail when it fails, and SutramX alerts you if a run is due and no ping arrives within the grace period.

Vendors: follow the official status of your payment, email and cloud providers with third-party status (on Starter, Growth and Pro). On every plan, when a widely used vendor API starts failing for many SutramX customers at once, your incident is marked "Likely external", so you are not debugging your own code during someone else's outage.

Decide what pages a person

Not every failed check is worth waking someone. Page for login, the public API and anything that takes money. Send the marketing site, docs and internal tools to a chat channel. Use the Degraded above setting to mark slow responses without opening an incident, and a failure threshold of 2 on targets that occasionally blip, so one bad check never wakes anyone.

Escalation policies (on Growth and Pro) keep notifying until someone acknowledges: for example, post to Slack and email the on-call engineer immediately, then page PagerDuty if nobody has acknowledged after 10 minutes. On-call schedules on Pro decide who "the on-call engineer" is at any hour. Schedule maintenance windows for planned work, so the incident is still recorded but nobody is paged and your uptime is not charged for it.

Run a status page your customers trust

Add one component per thing a customer recognises (Login, API, Dashboard, Webhooks), not one per internal service. Group them into sections if you have several products. The page updates itself from your monitors, so an outage shows up without anyone remembering to post it.

During an incident, post short public updates: what is affected, what you are doing, and when you will update next. Visitors can subscribe by email with double opt-in, and they get a message when an incident starts and when it is resolved. Scheduled maintenance appears on the page in advance.

Every plan includes at least one status page (Free has 1 status page). Custom domains such as status.example.com are on Growth and Pro, and password-protected pages for a single enterprise customer are on Growth and Pro. The page is hosted by SutramX, away from your own infrastructure, so it stays reachable when your product is not.

Common questions

How quickly will we hear about an outage? Within roughly one check interval, plus the time it takes to confirm the failure. The shortest intervals are 15 s (Pro), 30 s (Growth), 60 s (Starter) and 3 min (Free). On paid plans an outage is confirmed by more than one region before anyone is alerted, which keeps one region's network problem from paging you.

Can SutramX data back an SLA report? Uptime is measured from outside your network, as the share of checks that passed. Daily uptime history is kept for 90 days on Free, 12 months on Starter and 24 months on Growth and Pro. SLO tracking with error budgets and burn rates is on Growth and Pro. Check history and incidents can be exported.

Should we monitor staging? Yes, but route it to its own channel and leave it out of escalation. A staging monitor that pages people at night teaches them to ignore pages.

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.