← Back to Work
Process file · Ops Ops

What’s inside: company memory · monitoring

The post-mortem drafted from the timeline, not from memory

The question this file answersWhy does the post-mortem cost an engineer another evening after the outage is already fixed?

Fits: engineering teams running 5–50+ incidents a month with alerting in PagerDuty, Datadog, Grafana or New Relic, deployments in GitHub or GitLab, and a post-mortem that someone is supposed to write afterwards.

Not for: teams with an incident a quarter, where the write-up is a rare and considered piece of work — nor for the root-cause analysis, which stays with the engineer.

Typical day

What the desk looks like today

Typical, from the software playbook — not a client's day. Five to fifty incidents a month is the playbook's range for a mid-size operation. An alert fires in PagerDuty or Datadog; the on-call engineer investigates, correlates logs with what was deployed, finds the cause, applies the fix — and then the second job begins: rebuilding the timeline from the incident channel, writing the post-mortem, updating the runbook. It is done last, late, by the tiredest person. In a bad week the write-ups slip and the next on-call meets the same failure without the notes.

What changes

What Monday looks like after

The morning after an incident. The on-call engineer opens a post-mortem that already has its skeleton: the timeline from first alert to resolution, the deployment that preceded it, the runbook entry that applied. What remains to write is the part only they know — why the system behaved that way and what should change — written in daylight, not at midnight. Stakeholders get a summary the same morning, sent by a person. You can see which incidents have a finished post-mortem and which runbooks were updated. Atlassian (2023) measured 2–3 hours saved per incident; the fix and the architectural call are still made by the same people.

Typical, not a measured client result. Every figure here comes from the playbook source named below.

2–3 h

per incident spent assembling the write-up — Atlassian (2023)

Before: someone rebuilds the incident timeline from the chat channel at midnight. After: Atlassian (2023) reports 2–3 hours saved per post-mortem draft and PagerDuty (2023) 25–40% from assisted correlation — root cause stays a person.

How this file is built

Atlassian 'IT Service Management' (2023) reports automated post-mortem drafting saving 2–3 hours per incident; PagerDuty 'State of Digital Operations' (2023) reports AI-assisted correlation reducing MTTR by 25–40%. Published figures, not measured by us; playbook range 30–50%. The fix stays engineering.

What we install

What we put in front of the systems you already run

Alerting keeps firing from PagerDuty, Datadog, Grafana or New Relic; the on-call rota does not change, and GitHub or GitLab and Jira stay put. Around them we add a correlation and drafting step through their APIs:

  1. when an incident opens, the alerts, deployments and incident-channel messages are gathered into one timeline
  2. runbook entries matching the alert pattern are surfaced to the on-call engineer mid-incident
  3. at close, a post-mortem draft is written from the timeline, with a suggested root-cause category and blanks only the engineer can fill
  4. a stakeholder summary and a status-page update are drafted for a person to send
  5. the draft is filed in Jira.

We start with your highest-volume alert sources.

What stays human — and what this will not do

Root cause on a failure nobody has seen before. The fix. Architecture. The stakeholder call on a major incident.

What can go wrong — and what we do about it

Correlation depends on tagging: services named inconsistently across tools produce timelines with gaps, and the first weeks go on naming, not drafting. A suggested root-cause category is a suggestion; treated as the answer it makes the post-mortem worse than none, so it is labelled a guess and the causal section stays blank for the engineer. Older monitoring tools without APIs mean partial timelines, stated on the draft, not filled in. The 30–50% range covers assembly and correlation; PagerDuty's MTTR figure is about resolution time, not write-ups; root cause on a novel failure is outside both.

How long it takes, and what we need from you

Audit, about two weeks (€1.5–3K): we read the last quarter's incidents and post-mortems, count how many got written, how late, and check the APIs of your alerting, source control and chat tools. Pilot, 4–6 weeks (€10–20K) — medium complexity in the software playbook, mostly correlating sources cleanly: one alert source, every draft reviewed by the on-call engineer, nothing sent out without a person. Production: all services, runbook updates proposed. From you: API access to PagerDuty, GitHub or GitLab, the incident channel, your template.

The path: free 60-second estimate → free 20-minute review → paid audit of this one process (€1.5–3K, typically two weeks) → pilot with your people in the loop (€10–20K, weeks, not quarters). No transformation programme. Prices are public, on the services page →

This is about you if…
What does this mean in euros?

That depends on your volumes and wage costs — this page will not invent the number. The free 60-second estimate runs that calculation from your answers, with every multiplier sourced.

Get your free savings estimate 60 seconds · no sales call Or write first → Map an ops process like this one — free, 60 seconds →

Not a named Aperanda client. Process file · Ops.

Deep-dive process file. Volumes, weeks and sources come from the industry playbook; nothing here is a named client.

All process files