Alerting & Escalation

Alerting and Escalation Without the Noise

The hard part of alerting is not delivery. It is making sure the message that arrives at 3am is worth waking up for, and that it reaches someone who can act.

In short

SutramX alerts only after a failure is confirmed. Email, Slack, Microsoft Teams, Discord, Telegram, Google Chat and Mattermost work on every plan, webhooks from Starter, and PagerDuty and Opsgenie from Growth; SMS and WhatsApp use monthly alert credits on paid plans. Each incident sends one alert thread with the failure reason and confirming regions, then a recovery notice.

Who it's for

For small DevOps and engineering teams that need escalation and on-call without paying per responder. Team seats are unlimited on every plan, including everyone on the rota. Escalation policies, which keep notifying the next person, channel or PagerDuty/Opsgenie until someone acknowledges, are included on Growth and Pro. On-call schedules, which decide who is on duty, are included on Pro.

Alert fatigue is a reliability problem

A team that has learned to ignore its alerts is less safe than a team with no alerts at all, because the second team knows it is flying blind. The first believes it is covered.

Every false alarm spends a little of the trust that makes the next real alert effective. Treating alert precision as a feature — rather than sending everything and letting humans filter — is the only way that trust survives contact with production.

IncidentconfirmedSlackMicrosoft TeamsGoogle ChatDiscordTelegramEmailWebhooksPagerDutyOpsgenie

Choosing a channel by urgency

Not every notification deserves the same intrusiveness. Matching channel to severity is what keeps the loud channels meaningful.

  • Chat (Slack, Discord, Microsoft Teams, Google Chat, Mattermost, Telegram)Team-visible, low friction. Good default for most incidents during working hours.
  • EmailBest for low-urgency and informational events such as an early certificate-expiry warning.
  • WebhookFor routing into your own systems — ticketing, automation, custom on-call logic.
  • Pager (PagerDuty, Opsgenie)Reserved for genuine wake-someone-up severity, with real escalation policies behind it. These tools can also phone or text your on-call engineer.

What a good alert contains

An alert should be actionable on its own. Which monitor, what specifically failed, how long it has been failing, which regions confirmed it, and a direct link to the detail view. If the first thing a responder has to do is go find out what the alert means, the alert is incomplete.

Specifications

ChannelsEmail, Slack, Discord, Telegram, Microsoft Teams, Google Chat, Mattermost (all plans); webhooks (Starter+); PagerDuty, Opsgenie (Growth+)
EscalationEscalation policies (Growth+); on-call schedules (Pro)
Trigger conditionCross-region confirmed state change
DeduplicationOne thread per incident
Recovery alertsSent on confirmed recovery
PayloadMonitor, failure reason, duration, confirming regions, deep link

Learn more

Related capabilities

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.