Guide · 10 min read

Blameless Postmortem Template, with a Filled-In Example

An incident costs you the outage either way. A good postmortem is how you get something back for it.

Last updated

In short

A blameless postmortem records what happened during an incident, why the system allowed it, and what will change, without assigning personal fault. A useful template covers a summary, impact, a timeline, root cause and contributing factors, what went well and badly, and action items that each have one owner and a due date.

What is a blameless postmortem?

A postmortem, also called an incident review or post-incident report, is a written account of an incident: what happened, what it affected, why it happened, and what will change so it does not happen again in the same way.

Blameless means the review looks for the conditions that let a mistake turn into an outage rather than for the person who made it. People in an incident act on the information they have at the time. If someone ran the wrong command, the useful questions are why that command was so easy to run against production, and why nothing caught its effect sooner. Once a review becomes about who is at fault, people leave out the details that would have prevented the next incident.

Blameless is not consequence-free

Blameless does not mean nobody is accountable. People are accountable for the action items they own, and for being honest in the review. They are not punished for a reasonable decision that turned out badly.

When should you write one?

Decide the triggers in advance, so writing a postmortem is routine rather than a signal that someone is in trouble. Common triggers:

  • Any outage visible to customers above a set length or severity.
  • Any data loss, or any security incident.
  • An incident that needed escalation beyond the on-call engineer.
  • A near miss that could easily have been serious.
  • Any incident where someone on the team asks for one.

Write the draft within a few days, while logs are still retained and people remember why they made each decision. Review it together in a short meeting, then publish it where the whole team can find it.

The postmortem template

Copy these sections in order. Keep each one short; a postmortem that takes an hour to read will not be read.

1. Header

  • Title: a plain description of what users experienced, such as "Checkout errors for 38 minutes".
  • Date of the incident, author, reviewers and status (draft, in review, final).
  • Severity, using your own written definitions.

2. Summary

Three to five sentences a non-engineer can follow: what broke, for how long, who was affected, what fixed it, and the most important change you are making.

3. Impact

  • Who was affected and how: which features, which users or regions, what they saw.
  • Duration from first user impact to recovery.
  • Measurable effects you actually have data for: failed requests, failed orders, support tickets, SLA or error budget consumed.
  • What was not affected, which matters as much to readers.

4. Timeline

Timestamped events in one time zone, usually UTC, from the triggering change to full recovery. Include when the failure started, when it was detected and how, when it was acknowledged, key decisions and the evidence behind them, customer communication, mitigation and recovery. Note the gaps too: twenty minutes with nothing happening is a finding.

5. Root cause and contributing factors

Describe the technical trigger, then the conditions that let it cause harm. Incidents rarely have one cause. Asking "why was that possible?" several times, sometimes called the five whys, is a useful prompt, but stop at the factors you can change and resist forcing a single chain. Group contributing factors by type: the change itself, detection, response, tooling and process.

6. What went well, what went badly, where we got lucky

Write down what worked so you keep doing it. Luck deserves its own heading: an incident that happened during working hours, or a backup that existed by accident, is a risk you have not fixed yet.

7. Action items

Each item is a specific change with one owner, a due date, a priority and a link to the ticket in your normal work tracker. Prefer changes to systems over reminders to people: "add a check that blocks the deploy" lasts longer than "be more careful". Two or three items that ship beat fifteen that do not.

8. Lessons learned and follow-up

One or two broader lessons for other teams, and a date to check the action items were completed.

Example: a filled-in postmortem (fictional)

The example below is invented to show the template in use. The company, people and numbers are not real.

Checkout errors for 38 minutes after a payment configuration change

Severity 2. Status: final. Author: on-call engineer. Reviewers: payments team, support lead.

Summary

On a Tuesday afternoon, a configuration change sent checkout's payment requests to the payment provider's sandbox endpoint instead of production. For 38 minutes, every card payment failed with a generic error, while the rest of the shop worked normally. A monitor on the checkout API detected the failure 3 minutes after the change. The team rolled the configuration back. We are adding a startup check that refuses to run production with sandbox credentials, and separating the two configurations.

Impact

  • All card payments failed from 14:02 to 14:40 UTC. Browsing, search and accounts were unaffected.
  • Customers saw "Payment could not be processed. Please try again."
  • Support received a burst of tickets about failed payments; none reported being charged.
  • The month's checkout error budget was exceeded.

Timeline (UTC)

  • 13:58A change to rotate the payment API key is merged. It also changes the endpoint setting, copied from the staging configuration.
  • 14:02The change deploys. Payment requests start failing.
  • 14:05The checkout API monitor confirms the failure from several regions and opens an incident. The on-call engineer is paged.
  • 14:08The on-call engineer acknowledges and opens the incident channel.
  • 14:12A status page update says payments are failing and other features work.
  • 14:15 to 14:30Investigation focuses on the payment provider, whose status page shows no incident. Provider logs are not accessible to on-call.
  • 14:31A payments engineer joins and spots the sandbox endpoint in the deployed configuration.
  • 14:36The configuration is rolled back.
  • 14:40Payments succeed again. The monitor recovers and the incident resolves. The status page is updated to monitoring, then resolved at 15:10.

Root cause and contributing factors

  • TriggerThe production payment endpoint was replaced with the sandbox endpoint during a key rotation.
  • The changeStaging and production settings live in one file with similar names, so a copied block is easy to miss in review.
  • PreventionNothing checked that production was using production credentials and endpoints. The service started normally with a sandbox endpoint.
  • DetectionDetection was fast, because the checkout API had its own monitor with a content assertion.
  • ResponseOn-call had no access to the provider's request logs and no runbook for payment failures, so 16 minutes went on ruling out the provider.

What went well, what went badly, where we got lucky

  • Went well: detection in 3 minutes; a status update within 10 minutes of the change; a clean rollback.
  • Went badly: the deploy diff did not make the endpoint change obvious; diagnosis depended on one person who happened to be available.
  • Lucky: the incident happened during working hours, when a payments engineer was online.

Action items

  • P1: startup guardThe service refuses to start in production with sandbox credentials or endpoints. Owner: payments team lead. Due: within one week.
  • P1: separate configurationSplit staging and production payment settings into separate files with distinct names. Owner: payments engineer. Due: within two weeks.
  • P2: runbookWrite a payment-failure runbook that starts with "what changed?" and links to the deploy history. Owner: on-call engineer. Due: within two weeks.
  • P2: accessGive on-call engineers read access to the payment provider's request logs. Owner: engineering manager. Due: within one month.

Lessons learned

A per-endpoint monitor made detection fast; the slow part was diagnosis. Configuration that differs between environments needs the same review safeguards as code.

Common postmortem mistakes

  • Naming individuals as the cause. Describe roles and decisions instead.
  • Stopping at "human error". It is where the investigation starts, not where it ends.
  • Writing it weeks later, from memory, when the logs have expired.
  • Action items without owners or dates, or kept in the document instead of the work tracker.
  • Never checking whether the action items shipped.
  • Publishing internal detail to customers. A public incident report is a separate, shorter document.

Postmortems in SutramX

Every resolved SutramX incident has a Postmortem section for your write-up, in simple Markdown, next to the incident's timeline of checks, alerts, escalation steps, acknowledgements and notes. Postmortems are internal and never shown on status pages.

On Pro, a Postmortem draft panel can generate a first draft from the incident's monitoring data. The draft has a title, summary, impact, a UTC timeline, root-cause hypotheses (each with its evidence and a confidence of high, medium or low), contributing factors, action items with a priority and an owner role, and lessons learned. It is blameless by design: it describes systems and decisions, and action items name a role such as "on-call engineer", never a person.

Nothing is saved as your postmortem until someone edits the draft and clicks Use as postmortem, which converts it to Markdown. The AI only knows what the monitoring data and your notes show, so treat its hypotheses as leads to check. Drafts are generated only when someone clicks the button, from a redacted fact sheet that leaves out credentials, contact details and request bodies, and each draft counts as one AI generation from the plan's monthly allowance (500 a month on Pro). AI incident summaries and status page update drafts are on Growth and Pro. Details are in the AI incident assist documentation.

Next steps

Keep reading

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.