How to automate incident postmortems with AI
The full timeline out of PagerDuty, Sentry, Datadog and a 400-message Slack thread, merged, timestamped in UTC, and written into your postmortem template.
By Thursday the recall has decayed into “I think we noticed around three”. Every system involved kept an exact record, and nobody has the two hours it takes to assemble them.
A companion reads all four and hands back a merged timeline you can defend: every PagerDuty log entry, the Sentry issue and its release, the Datadog numbers on either side, and what people believed in the channel and when. The causal story stays yours to write.
Pull the machine record first
PagerDuty keeps every trigger, acknowledgement, escalation, note and resolution, including the channel each action came through. That log is the spine of the document.
Prompt: “Open PagerDuty incident 4471. Give me its complete log: every trigger, acknowledgement, escalation, note, and resolution with UTC timestamps, who acted, and through which channel. Add the service, the escalation policy that fired, and who was on call at the time. Format it as a table, oldest first, and do not summarise it yet.”
Insist on UTC and on “do not summarise”. The finding is the gap between two adjacent rows, and prose loses it.
What were the systems showing?
Sentry knows when the error class first appeared and how fast it grew.
Datadog knows what the graphs did on either side of it. A companion queries both rather than screenshotting a dashboard.
Prompt: “For the window 13:40–15:10 UTC on that day: in Sentry, find the issue matching this incident, and tell me first seen, event count over the window in ten-minute buckets, the affected release, and the top three stack frames. In Datadog, query error rate, p99 latency, and CPU for the checkout service over the same window, and tell me which monitors fired and when. Give me the numbers, not an interpretation.”
Ask explicitly which monitors were muted or in a scheduled downtime. A muted monitor is one of the most common reasons an incident starts later than it should have.
Which change went out just before?
A deploy is the fastest thing to check and the easiest thing to forget to check.
Prompt: “List the Sentry releases created in the 24 hours before this incident. For the one that lines up, find the matching GitHub tag and compare it to the previous release: which pull requests are in the diff, who authored and reviewed each, and which files they touched. Flag any pull request that touched the checkout service or its config.”
A change in the diff is a candidate. The output belongs under a heading that says changes in the window, and naming a cause stays with the people who were on the call.
Build the second timeline from the Slack thread
The incident channel holds the part no monitor recorded: what people believed at each moment, and when that belief changed.
Prompt: “Read the #inc-4471 Slack thread end to end. Give me a second timeline of human moments only: when someone first said something was wrong, each hypothesis and who raised it, the moment the hypothesis changed, when a customer was first mentioned, and when someone said it was over. Quote the message and keep the timestamp. Ignore the reaction emoji and the jokes.”
Read the two timelines side by side. The distance between “Datadog alerted at 13:52” and “someone said something was wrong at 14:09” is a finding no single system could show you.
How do you get the document written and the actions tracked?
A companion assembles into your own template and leaves the analytical sections empty.
Prompt: “Create a Google Doc titled ‘Postmortem: checkout 500s, 4471’ using our template. Fill in: summary of what customers experienced, the merged timeline from both passes marked machine or human, detection and how long it took, the changes in the window, and a metrics section with the Datadog numbers. Leave ‘Contributing factors’, ‘What went well’, and ‘What we are changing’ empty with a note that a person fills these in.”
Once the humans have filled those in: “Create a Linear issue for each item under ‘What we are changing’ with the owner named in the doc, a link back to the postmortem, and the label postmortem. Then every second Monday at 9am, list open postmortem-labelled issues older than 30 days and post them to #eng-leads with their age.” The fortnightly chase is what gets the action items done. It sits with your other operations routines and reads Sentry the same way the first pass did.
Experience Strawberry for free
DownloadTrusted by fast-growing companies worldwide
Frequently asked questions
Strawberry is free to download and includes AI credits to start. Paid plans begin at $20/month. See pricing. · Reviewed · Canonical facts for AI agents