An on-call rotation decides who responds to alerts at any given time. Most teams start with a weekly rotation, a primary and a secondary responder, an escalation policy that pages the next person when an alert is not acknowledged within a set time, and a written handoff at each shift change. Follow-the-sun suits teams spread across time zones.
What is an on-call rotation for?
Alerts are useless if nobody is responsible for them. An on-call rotation names one person, at every hour you choose to cover, who will acknowledge an alert, assess it and either fix the problem or pull in the people who can. Everyone else can stop watching their phone.
Being on call does not mean fixing everything alone. The on-call person's job is to respond quickly, stop the bleeding where they can, and escalate when they cannot.
Common rotation patterns
- Weekly rotationEach person covers a full week, then hands over. Simple to plan and remember, with few handoffs. The cost is a long, tiring stretch when the week is busy. Handing over mid-week, for example on Tuesday morning, keeps the handoff away from weekends and Monday releases.
- Shorter shiftsDaily shifts, or splitting the week into weekday and weekend blocks, spread the load more evenly but add handoffs. Useful when on-call is heavy or when people want weekends protected.
- Follow-the-sunTeams in different time zones each cover their own working hours, so nobody is paged at night. It needs at least two or three locations with people able to handle the same systems, and careful handoffs between them.
- Primary and secondaryThe primary takes every page; the secondary is paged only when the primary does not acknowledge in time, or is pulled in for big incidents. Often the secondary is next week's primary, so they already know what is going on.
- Business hours and out-of-hours tiersNot every alert needs a 3am response. Many teams page out of hours only for urgent, user-facing failures and let everything else wait for working hours.
How many people does a rotation need?
With a weekly rotation, each person is on call one week in every N, where N is the number of people. Two people means being on call every other week, which few people can keep up for long. Google's Site Reliability Engineering book suggests at least eight engineers for a single-site rotation covering 24 hours a day, or six per site for a two-site rotation, so that each person is on call rarely enough to recover.
Small teams rarely have that, and that is fine if the coverage promise matches the headcount. A team of three can run business-hours on-call well; it should be wary of promising round-the-clock response to every alert.
Handoffs
Most on-call mistakes happen at the seams between shifts: an incident that was "nearly fixed", a deploy that is still rolling out, a noisy monitor everyone was told to ignore. A short written handoff closes the gap.
- Open incidents and anything still being watched.
- Changes in flight: deploys, migrations, config changes and maintenance windows during the next shift.
- Alerts that fired and what was done about them, including any that were noise and need fixing.
- Known risks: a big customer launch, a dependency having a bad week, a certificate or domain due to expire.
- Anything the outgoing person was asked to follow up on.
Write it where the next person will find it, in the incident channel or a shared document, and keep the format the same every week so it takes five minutes.
Escalation policies
An escalation policy is the safety net under the rotation. It says who gets notified next, and after how long, if an alert is not acknowledged. Without one, a missed page waits until someone happens to notice.
- Step 1, immediatelyThe primary on-call, through a channel that reaches them: push or email in the day, SMS or a phone call at night.
- Step 2, after 10 to 15 minutesThe secondary on-call, or the team lead.
- Step 3, after a further 15 minutesA wider group or an engineering manager, as a last resort.
Pick delays that give a sleeping person time to wake up, read the alert and acknowledge it, but not so long that a missed page leaves a real outage unattended. Acknowledging should stop the escalation, so people are not paged for an incident someone is already handling. And make sure the last step reaches a person who can act, not a channel nobody reads.
Alert fatigue
The fastest way to ruin a rotation is to page people for things that do not need them. Every unnecessary page teaches the on-call engineer to respond a little more slowly to the next one.
- Page only for problems that are urgent and affect users. Send everything else to a ticket queue or a daytime channel.
- Confirm failures from more than one location before paging, so a single network blip does not wake anyone.
- Review every page at the end of each shift: was it real, was it actionable, did it go to the right person?
- Track pages per shift and night-time pages per person. A rising count is a reliability problem to fix, not a cost of doing business.
- Give people time off or a late start after a bad night, without them having to ask.
Compensation and fairness
Being on call restricts what people can do with their evenings and weekends, even when nothing breaks. Teams recognise that in different ways:
- A flat allowance per shiftSimple and predictable, and it pays for the restriction as well as the work.
- Pay for time worked out of hoursRewards the actual disruption, though it can feel like being paid for things breaking.
- Time off in lieuTime back for nights and weekends spent working incidents.
In some countries, employment law treats restrictive standby time as working time, or sets rules on rest periods after night work, so check local rules before designing a scheme. Fairness matters as much as money: share holidays evenly, make shift swaps easy, and include senior engineers and managers in the rotation rather than leaving it to the newest hires.
On-call for small teams
Two to five people cannot run a rotation like a large company, but they can still run one that works.
- Be explicit about coverage. "Business hours, plus best effort at night for full outages" is a legitimate policy if you write it down.
- Page out of hours only for your most critical tier: the website being down, checkout or login failing, the API returning errors.
- Use phone calls or SMS for night-time pages, and only for those.
- Make the escalation path short: the on-call person, then a founder or lead.
- Keep a status page, so customers can see you know about a problem even when the fix takes time.
- Invest in automatic recovery, such as restarts, health-checked deploys and rollback, so fewer problems need a human at all.
Escalation and on-call schedules in SutramX
SutramX includes escalation policies on Growth and Pro and on-call schedules on Pro. Every plan has unlimited team members, with no per-seat or per-responder pricing.
- Escalation policiesUp to 20 steps, each with a delay of 0 to 1,440 minutes after the previous one and a channel: email, Slack, Discord, Microsoft Teams, Google Chat, Mattermost, Telegram, PagerDuty, Opsgenie, webhook, WhatsApp, SMS or voice call. A policy is assigned per monitor and starts when the down alert is sent.
- Stopping escalationEscalation stops as soon as the incident is acknowledged or resolved. You can acknowledge on the incident page or with the Acknowledge button on Slack alerts, and the first acknowledgement is passed on to PagerDuty and Opsgenie. Steps do not repeat after the last one.
- On-call schedulesA schedule rotates members (up to 100) in order, each for a rotation interval of 1 to 720 hours, such as 168 for weekly shifts, in a time zone you choose. Optional shift windows limit on-call to certain hours, with a handoff period around each change, and holiday overrides mark days with nobody on call.
- Who gets pagedEmail steps can send to Current on-call, which the schedules resolve when the step fires. SMS and voice steps go to the fixed number on the step and use credits. If a Current on-call step finds nobody on duty, it pages your first PagerDuty connection if you have one; otherwise it retries for about 30 minutes, then skips to the next step.
- DelaysSnoozes, maintenance windows, quiet hours and deployment windows postpone steps, and every step, skip and delay is recorded on the incident timeline.
Holiday overrides and shift windows leave hours with nobody on call. Make sure those hours either have a later escalation step that always reaches someone, or a PagerDuty connection as a fallback. Full details are in the escalation and on-call documentation.