← Back to Work
Process file · Ops
Ops
What’s inside: company memory · monitoring
The post-mortem drafted from the timeline, not from memory
The question this file answersWhy does the post-mortem cost an engineer another evening after the outage is already fixed?
Fits: engineering teams running 5–50+ incidents a month with alerting in PagerDuty, Datadog, Grafana or New Relic, deployments in GitHub or GitLab, and a post-mortem that someone is supposed to write afterwards.
Not for: teams with an incident a quarter, where the write-up is a rare and considered piece of work — nor for the root-cause analysis, which stays with the engineer.
Typical day
What the desk looks like today
Typical, from the software playbook — not a client's day. Five to fifty incidents a month is the playbook's range for a mid-size operation. An alert fires in PagerDuty or Datadog; the on-call engineer investigates, correlates logs with what was deployed, finds the cause, applies the fix — and then the second job begins: rebuilding the timeline from the incident channel, writing the post-mortem, updating the runbook. It is done last, late, by the tiredest person. In a bad week the write-ups slip and the next on-call meets the same failure without the notes.
What changes
What Monday looks like after
The morning after an incident. The on-call engineer opens a post-mortem that already has its skeleton: the timeline from first alert to resolution, the deployment that preceded it, the runbook entry that applied. What remains to write is the part only they know — why the system behaved that way and what should change — written in daylight, not at midnight. Stakeholders get a summary the same morning, sent by a person. You can see which incidents have a finished post-mortem and which runbooks were updated. Atlassian (2023) measured 2–3 hours saved per incident; the fix and the architectural call are still made by the same people.
Typical, not a measured client result. Every figure here comes from the playbook source named below.
2–3 h
per incident spent assembling the write-up — Atlassian (2023)
Before: someone rebuilds the incident timeline from the chat channel at midnight. After: Atlassian (2023) reports 2–3 hours saved per post-mortem draft and PagerDuty (2023) 25–40% from assisted correlation — root cause stays a person.
How this file is built
Atlassian 'IT Service Management' (2023) reports automated post-mortem drafting saving 2–3 hours per incident; PagerDuty 'State of Digital Operations' (2023) reports AI-assisted correlation reducing MTTR by 25–40%. Published figures, not measured by us; playbook range 30–50%. The fix stays engineering.
What we install
What we put in front of the systems you already run
Alerting keeps firing from PagerDuty, Datadog, Grafana or New Relic; the on-call rota does not change, and GitHub or GitLab and Jira stay put. Around them we add a correlation and drafting step through their APIs:
- when an incident opens, the alerts, deployments and incident-channel messages are gathered into one timeline
- runbook entries matching the alert pattern are surfaced to the on-call engineer mid-incident
- at close, a post-mortem draft is written from the timeline, with a suggested root-cause category and blanks only the engineer can fill
- a stakeholder summary and a status-page update are drafted for a person to send
- the draft is filed in Jira.
We start with your highest-volume alert sources.
What stays human — and what this will not do
Root cause on a failure nobody has seen before. The fix. Architecture. The stakeholder call on a major incident.
What can go wrong — and what we do about it
Correlation depends on tagging: services named inconsistently across tools produce timelines with gaps, and the first weeks go on naming, not drafting. A suggested root-cause category is a suggestion; treated as the answer it makes the post-mortem worse than none, so it is labelled a guess and the causal section stays blank for the engineer. Older monitoring tools without APIs mean partial timelines, stated on the draft, not filled in. The 30–50% range covers assembly and correlation; PagerDuty's MTTR figure is about resolution time, not write-ups; root cause on a novel failure is outside both.
How long it takes, and what we need from you
Audit, about two weeks (€1.5–3K): we read the last quarter's incidents and post-mortems, count how many got written, how late, and check the APIs of your alerting, source control and chat tools. Pilot, 4–6 weeks (€10–20K) — medium complexity in the software playbook, mostly correlating sources cleanly: one alert source, every draft reviewed by the on-call engineer, nothing sent out without a person. Production: all services, runbook updates proposed. From you: API access to PagerDuty, GitHub or GitLab, the incident channel, your template.
The path: free 60-second estimate → free 20-minute review → paid audit of this one process (€1.5–3K, typically two weeks) → pilot with your people in the loop (€10–20K, weeks, not quarters). No transformation programme. Prices are public, on the services page →
This is about you if…
- Do incidents fire alerts in PagerDuty, Datadog or Grafana, with a chat channel each?
- Does the on-call engineer rebuild the timeline and write the post-mortem afterwards?
- Are deployments and tickets in GitHub, GitLab, Jira or Linear a system could read?
What does this mean in euros?
That depends on your volumes and wage costs — this page will not invent the number. The free 60-second estimate runs that calculation from your answers, with every multiplier sourced.
Not a named Aperanda client. Process file · Ops.
Deep-dive process file. Volumes, weeks and sources come from the industry playbook; nothing here is a named client.
All process files