Alerts & Incidents

Escalation & on-call

Build escalation policies that keep notifying people until an incident is acknowledged, and on-call rotations that route alerts to whoever is on duty.

An escalation policy makes sure an outage gets a human response. After the first alert, it works through a list of steps (email the on-call engineer, then text the team lead, then page PagerDuty), each after a delay, until someone acknowledges the incident or it resolves. On-call schedules decide who "the on-call engineer" is at any moment.

Both are on Alerts → On-call & escalation. Escalation policies and on-call schedules depend on your plan, and on-call schedules can need a higher plan than escalation policies; see Plans & limits. Only the workspace owner can create, edit or delete them.

How escalation progresses#

  1. A monitor with an escalation policy goes down and its down alert is sent to all its normal recipients and channels.
  2. The policy starts. Step 1 fires after its delay, counted from the down alert.
  3. Each following step fires after its own delay, counted from the previous step.
  4. Escalation stops as soon as the incident is acknowledged or resolved.
  5. After the last step, escalation is complete. Steps don't repeat; the incident page shows "All steps have been sent but the incident is still ongoing."

Example policy:

StepDelayChannelTarget
1ImmediatelyEmailCurrent on-call
210 minutes after previous stepSMS+1XXXXXXXXXX
315 minutes after previous stepPagerDutyConnected PagerDuty integration

With this policy, an incident that nobody acknowledges emails the on-call engineer when it opens, texts the lead 10 minutes later, and pages PagerDuty 25 minutes after the first alert.

Escalation until someone answers
The example policy above: email the on-call engineer, text the lead after 10 minutes, page PagerDuty 15 minutes later. Acknowledging at minute 12 stops the remaining steps.

Escalation only starts when the down alert is actually sent. An incident whose alert is held back (for example by a maintenance window) starts escalating once the alert goes out.

Create an escalation policy#

  1. Go to Alerts → On-call & escalation and scroll to Escalation Policies.
  2. Enter a Policy name, for example "Production critical".
  3. Configure the first step, then click Add step for each further step.
  4. Leave Policy enabled ticked.
  5. Click Create Policy.
FieldWhat it doesDefault / limits
Policy nameName shown in monitor settingsRequired, up to 255 characters
Minutes after the alert (step 1) / Minutes after previous stepDelay before this step fires0 to 1,440 minutes; 0 = immediately. New policies start at 0, added steps at 5.
ChannelHow this step notifiesEmail, Slack, Discord, Microsoft Teams, Google Chat, Mattermost, Telegram, PagerDuty, Opsgenie, Webhook, WhatsApp, SMS, Voice call
Send to (Email only)Current on-call or Email addressCurrent on-call
Email addressAddress to email when Send to is Email addressMust be a valid email
Number to text / Number to call (SMS, Voice call)Phone number for this stepE.164 format, for example +1XXXXXXXXXX
Policy enabledDisabled policies stay assigned but never runEnabled

A policy can have up to 20 steps. Use the arrows to reorder steps and the bin icon to remove one.

What each step channel does#

ChannelWho is notified
Email, Current on-callWhoever the on-call schedules say is on duty when the step fires
Email, Email addressThat address, once it has confirmed alert emails
SMS / Voice callThe number on the step, once it is verified (see below). Uses 1 SMS or voice call credit.
Slack, Teams, Discord, Google Chat, Mattermost, Telegram, Webhook, WhatsAppEvery connection of that channel, ignoring routing
PagerDuty / OpsgenieEvery connection of that tool, on the same page as the original alert

Chat, webhook and WhatsApp steps need the channel to be connected on Alerts → Channels & API. The policy form warns you when a step's channel isn't connected or isn't on your plan, with a link to Open Integrations or Upgrade.

Verify SMS and voice numbers#

A new number on an SMS or voice call step must be verified before it is paged. When you click Create Policy or Save Changes, SutramX texts a 6-digit code to each new number. The step then shows Code sent or Unverified, and a panel under the policy says alerts to that number are paused. Enter the code and click Verify. On voice call steps you can choose Call me with the code instead.

The code expires after 10 minutes, and you can request a new one after 60 seconds. A number that is already verified in the workspace, for example as an SMS or voice connection, needs no new code. If a step comes due before its number is verified, the step is skipped and the workspace owner is emailed. See SMS, WhatsApp & voice for the full rules.

Escalation messages#

Escalation messages say the incident is still open:

  • Email: subject "Escalation step 2: Checkout API incident is still open", with the step number, how long it has been open, the current on-call person (on Current on-call steps), and an Acknowledge incident button.
  • SMS: [ESCALATION step 2] Checkout API is still DOWN. plus the incident link.
  • Voice: "Escalation. Checkout API is still down."
  • Chat and webhooks: a down alert whose detail reads "Escalation step 2: still down".

Assign a policy to a monitor#

Escalation is opt-in for each monitor:

  1. Open the monitor and click Edit.
  2. Choose the policy in Escalation policy.
  3. Click Save changes.

When you add a monitor, you can also pick the policy under Show advanced options → Who gets alerted.

Default policy in that list means no escalation policy: the monitor sends its normal alerts but doesn't escalate. Disabled policies are shown with "(disabled)" and don't run.

Acknowledging an incident#

Acknowledging tells SutramX someone is on it. It stops the escalation immediately. You can acknowledge from:

  • The incident page in the dashboard, with Acknowledge.
  • The Acknowledge button on Slack alerts from a one-click Slack connection.

The Acknowledge incident button in escalation emails opens the incident page.

The first acknowledgement is also sent to PagerDuty and Opsgenie, so their on-call stops being notified too. Acknowledging doesn't resolve the incident: the recovery alert still goes out when the monitor comes back up.

If a flapping monitor goes down again after recovering, the incident needs a fresh acknowledgement, and when its down alert is sent again the escalation restarts from step 1.

When steps are delayed or skipped#

SituationWhat happens
The incident is snoozedSteps wait until the snooze ends
A maintenance window, quiet hours or a deployment window covers the monitorThe step is postponed and re-checked every 5 minutes
The monitor is pausedEscalation doesn't run
Current on-call but nobody is on dutyYour first PagerDuty connection is paged instead. Without PagerDuty, the step waits 5 minutes and tries again, up to 6 times (about 30 minutes). If nobody is on duty by then, the step is skipped and the policy moves on.
The email recipient hasn't confirmed alert emailsThe email is skipped (PagerDuty is paged instead if connected) and the policy moves on
The email can't be deliveredPagerDuty is paged instead if connected; otherwise the step is retried every 5 minutes
The channel isn't connected, isn't on your plan, or its credits are used upThe step is skipped and the policy moves on to the next step
The SMS or voice number on the step isn't verifiedThe step is skipped, the workspace owner is emailed, and the policy moves on
A temporary delivery errorThe step is retried every 5 minutes
Your plan no longer includes escalation policiesEscalation stops

Every step, skip and delay is recorded on the incident timeline, and the incident page shows an Escalation card with the policy's progress (Active, Acknowledged, Resolved or All steps sent).

On-call schedules#

An on-call schedule rotates a list of people through on-call duty. Escalation steps that email Current on-call use it to pick the recipient.

Create a schedule#

  1. Go to Alerts → On-call & escalation and find On-Call Rotation Scheduling.
  2. Enter a Schedule name, for example "Primary SRE rotation".
  3. Choose the Timezone and the Rotation start date and time.
  4. Set the Rotation interval (hours), for example 168 for weekly shifts.
  5. Under On-Call Members, click Add Member and enter each person's Name and Email, in rotation order.
  6. Optionally add Shift Windows and Holiday Overrides.
  7. Leave Schedule enabled ticked and click Create Schedule.
FieldWhat it doesDefault / limits
Schedule nameLabel for the scheduleRequired, up to 255 characters
TimezoneTimezone for shift windows and holidaysUTC
Rotation startWhen the first member's turn begins (entered in your local time)Required
Rotation interval (hours)How long each member is on call before the next takes over1 to 720 hours
On-Call MembersPeople in the rotation, in order1 to 100 members; name and email required
Shift WindowsDay, Start, End and Handoff (min) of on-call hoursOptional; none means on call at all times
Holiday OverridesDates (with an optional name) when this schedule has nobody on callOptional
Schedule enabledDisabled schedules are ignoredEnabled

How the current on-call is chosen#

  • Rotation: members take turns in list order. Each turn lasts the rotation interval, counted from the rotation start. After the last member, it starts again with the first.
  • Shift windows: if a schedule has shift windows, someone is on call only during them. A window whose end is before its start runs overnight. Without shift windows, the rotation covers every hour.
  • Holidays: on a holiday date (in the schedule's timezone), the schedule has nobody on call.
  • Several schedules: SutramX checks enabled schedules starting with the most recently updated and uses the first one that has someone on duty.

The top of the section shows Current on-call: with the person's name, email and source ("rotation", "override" or "rotation (handoff window)"), or "No active on-call assignee".

Handoff (min) marks the minutes around the start and end of a shift as a handoff period, shown as "rotation (handoff window)". Escalation emails still go to the current member, not the outgoing or incoming one.

Edit or delete a schedule#

Click Edit on a schedule to change it, then Save Schedule. Delete removes the schedule with its members, shifts and holidays. Escalation steps that email Current on-call may then have nobody to notify.

Common questions#

Can the on-call person be texted or called instead of emailed? Current on-call is available for Email steps only. SMS and voice steps go to the fixed number on the step.

Does escalation repeat after the last step? No. Add more steps (up to 20) with longer delays if you need reminders over a longer period.

Do escalation steps follow channel routing? No. A step goes to every connection of its channel, so it reaches the right place even if routing excludes the monitor.

Will a resolved incident keep escalating? No. Resolving the incident, automatically or manually, stops its escalation.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.