E-Commerce

Monitoring for E-Commerce

In e-commerce, downtime has an exact price, and you can calculate it. Every minute the checkout is broken is revenue that does not arrive and, worse, customers who go somewhere else.

What tends to go wrong

  • Checkout failure is silentCustomers rarely report a broken checkout. They abandon the cart and you see a conversion dip days later.
  • Peak traffic is when it breaksSales events and seasonal peaks are exactly when infrastructure fails and exactly when downtime costs the most.
  • Payment callbacks are invisibleA failing webhook from your payment provider does not show up on any page. Orders just quietly stop being confirmed.
  • Third-party dependenciesPayment gateways, shipping APIs, and tax services can fail independently of your own stack.

What to monitor

EndpointWhy it mattersSuggested interval
Checkout pageDirectly revenue-bearing. Any failure here has an immediate, measurable cost.30s
Payment callback / webhookSilent failure mode — orders stop confirming with no visible error anywhere.60s
Product detail pageThe most-trafficked page type and the main entry point from search.60s
Cart / basket APIBreaks the path to checkout even when the checkout page itself is fine.60s
Search endpointFrequently returns 200 with zero results when the index is broken.2m

Monitor the funnel, not just the homepage

A homepage check tells you the site is serving traffic. It tells you nothing about whether anyone can actually buy something. The failure modes that cost money live deeper in the funnel, on pages that are harder to reach and therefore less frequently checked.

Every step from product page to cart to checkout to payment confirmation deserves its own monitor, because each one can fail independently while the others look fine.

Search returning nothing is an outage

Site search is the clearest example of a soft failure. When the search index is broken or empty, the endpoint returns HTTP 200 with an empty result set. Every conventional uptime monitor reports the site as perfectly healthy.

Meanwhile, every customer who searches sees no products and leaves. A body assertion requiring a known result is what turns this from an invisible revenue leak into an alert.

Get ready before the peak, not during it

The week before a major sale is when monitoring coverage should be reviewed and tightened, not the morning of. Increase check frequency on revenue paths, verify alert routing reaches someone who will be awake, and confirm your status page is reachable and up to date.

It is also worth checking that certificate expiry dates do not fall inside your peak window — a surprising number of seasonal outages are just a certificate that lapsed on the worst possible weekend.

Example monitor settings

Checkout page: an HTTP monitor using the Website HTML guided preset (status 200, 301 or 302, a content-type of text/html, forbidden keywords such as "error" and "exception", a 2,500 ms maximum response time and a failure threshold of 2). Add a Required keyword that only appears when the page works, such as the label on your pay button. A checkout that renders an error message with a 200 status then fails the check.

Product page: the same preset on a product you always stock, with the product name as the Required keyword. If that product sells out regularly, pick a different one, or the monitor will report a stock change as an outage.

Search: request a search URL for that same product and require its name in the response. This is the check that catches an empty index returning 200.

Cart: an API monitor that asserts on the cart endpoint's status and JSON. With multi-step API checks (on Starter, Growth and Pro) you can add a test product to a cart and read the cart back in one check. Use a test product, and never complete a real payment from a monitor.

Order confirmation: a payment callback cannot easily be simulated without creating orders. Instead, run a scheduled job every hour that counts orders confirmed in the last hour. If there were any, it calls a cron heartbeat URL; if there were none during trading hours, it calls the same URL with ?status=fail. A missing or failed ping opens an incident, which catches the silent case where payments succeed but orders never confirm.

Payment, shipping and tax providers

Your checkout depends on services you do not run. Third-party status (on Starter, Growth and Pro) follows the official status pages of more than 70 services, including Stripe, AWS and Cloudflare, and sends their incidents through your normal alert channels. You can watch only the components you use, such as a payment provider's API and webhooks.

Official status pages are often slow to admit a problem, so keep your own checks too. On every plan, when a widely used vendor API starts failing for many SutramX customers at once, your incident is marked "Likely external", which saves time during a sale when every minute of diagnosis counts.

Who gets told, and how

Send checkout, cart and order confirmation to a channel someone watches closely, and the rest to a quieter one. Each chat connection has its own routing, so this takes two connections, not two tools. SMS and WhatsApp alerts with monthly credits are on Starter, Growth and Pro, which helps when the person on call is away from a laptop.

For a sale, an escalation policy (on Growth and Pro) makes sure an outage is not left sitting in a channel: notify the team, then text or page the on-call person if nobody acknowledges within a few minutes. If you freeze deploys for the event, schedule maintenance windows only for the work you actually plan, because alerts for monitors inside a window are held back.

A status page for shoppers

Shoppers do not care about your API. Name components in their words: Storefront, Checkout, Order tracking, Customer accounts. When something breaks, a short public update ("Some card payments are failing. Orders placed with other methods are not affected.") prevents a pile of duplicate support emails.

Announce planned maintenance ahead of time, and let customers subscribe to incident emails. A custom domain such as status.example.com is on Growth and Pro; otherwise the page lives on a SutramX address that keeps working when your store does not.

Common questions

Will monitoring add load during a peak? Very little. Each check is one request from each region at each interval: ten monitors every 30 seconds from two regions is 40 requests a minute, which is negligible next to sale traffic. Point frequent checks at pages that are cheap to serve.

How often should checkout be checked? As often as your plan allows. The shortest intervals are 15 s (Pro), 30 s (Growth), 60 s (Starter) and 3 min (Free). Detection time is about one interval plus confirmation, and for checkout the cost of an undetected minute is far higher than the cost of the checks.

Does a 200 mean the store works? No. A 200 only means the server answered. Required and forbidden keywords, and assertions on API responses, are what tell you the page showed what a customer needs to see.

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.