Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Incident Management
    Postmortem
    SRE

    Automatic Incident Postmortems from Alarms and Work Orders

    July 3, 2026
    10 min read

    Incident postmortems can be generated automatically when the incident-response system captures the evidence while the incident is happening. The source v3.2 design assembles the event timeline, links alarms and work orders, produces a four-part draft, assigns improvement actions to owners and due dates, and sends the reviewed postmortem into the knowledge base.

    The source example says the initial draft can be generated in about 40 seconds. That speed comes from having the material already structured during response, so the system never has to reconstruct the incident from memory after everyone has moved on.

    Why are postmortems difficult to write manually?

    The source explains the problem directly: postmortems often do not get written because gathering the logs and evidence is too time-consuming. The incident may be spread across:

    • Monitoring alarms
    • Work orders
    • Change records
    • Chat messages
    • Automation logs
    • Device events
    • Task state
    • Operator notes

    After the service is restored, the team has to reconstruct the sequence, and that can take longer than the actual remediation. Important details are forgotten, timestamps conflict, and the final document becomes a narrative based on memory instead of an evidence-based record.

    The source design solves this by collecting the material during the response process.

    What should be captured during the incident?

    Capture the timeline as the response progresses. The source incident flow records each operational step. A card-level example goes like this: a hardware alarm appears, multiple alarms consolidate into one incident, and the affected tasks and inference instances are identified. The node is isolated after authorization, tasks are rescheduled, the work order closes, and the review enters the knowledge base.

    That sequence already contains the skeleton of the postmortem. The automation only has to organize the recorded events into that story.

    How do alarms feed the postmortem?

    Alarms provide the technical evidence around detection and symptoms. The source AIOps model keeps raw alarms even when it consolidates them into a smaller number of incidents, which is important for postmortem generation. The draft can include:

    • First relevant alarm
    • Related alarm sequence
    • Consolidated incident
    • Downstream symptoms
    • Alarm timestamps
    • Affected objects

    The raw evidence remains available for drill-down. The postmortem should not copy hundreds of alarm lines into the main narrative. It should use the alarm history to reconstruct what the system observed and when.

    How does the incident timeline get generated?

    The incident platform records state changes and actions with timestamps, and the source postmortem page says the draft is built from the event timeline. That timeline can include:

    • Detection
    • Incident creation
    • Assignment
    • Escalation
    • Diagnosis
    • Approval
    • Remediation
    • Validation
    • Closure

    The exact timeline depends on the incident. The design principle is that operational actions create events automatically. If the team updates the work order, executes an approved automation, or changes a resource state, that action should already have a timestamp, and the postmortem generator can order those events into a coherent sequence.

    How do work orders contribute?

    Work orders provide responsibility, action, and resolution context. The source incident and workflow model links alarms with work orders and writes execution results back into the work order. That means a postmortem can retrieve:

    • Who owned the incident
    • Which team handled it
    • Which action was requested
    • Which action was approved
    • What was executed
    • What result was recorded
    • When the work order closed

    This beats asking the engineer to retype the response after the incident, because the work order itself becomes part of the evidence.

    How do remediation actions contribute?

    Remediation actions explain what the team or automation actually did to restore service. The source operations model divides actions into automatic, semi-automatic, and manual risk levels, and it also keeps operation audit. A postmortem can therefore show:

    • Recommended action
    • Authorization
    • Execution time
    • Execution identity
    • Result
    • Rollback if any
    • Validation

    That helps answer one of the most important postmortem questions, which is what fixed the incident. It also reveals whether the first action failed and another action was required.

    What four sections does the source-generated postmortem contain?

    The v3.2 source uses a four-part structure:

    • Timeline reconstruction
    • Root cause and contributing factors
    • What went well
    • Improvement items

    The structure works because it separates evidence from judgment. The timeline says what happened, the root-cause section explains why, the "what went well" section identifies controls that worked, and the improvement section turns the incident into follow-up work.

    The source also says root cause and contributing factors are listed separately, which helps avoid reducing a complex incident to one simplistic cause.

    How should root cause be handled automatically?

    The system can bring forward the root-cause evidence collected during incident analysis, but the final judgment should remain reviewable. The source AIOps design produces a likely root cause, a confidence level, and the evidence behind it, and the postmortem generator can use that as a starting point. It can also include:

    • Topology evidence
    • Time sequence
    • Related change
    • Historical incident
    • Hardware event

    The source postmortem page states the larger principle plainly: automation gathers the material, while people make the judgment. The draft should therefore distinguish recorded evidence from the final accepted conclusion.

    What are contributing factors?

    Contributing factors are conditions that made the incident more likely, harder to detect, or harder to recover even if they were not the primary root cause. The source postmortem structure explicitly separates root cause and contributing factors.

    Examples depend on the incident. Possible evidence can come from:

    • Missing monitoring
    • Old configuration
    • Capacity pressure
    • Slow escalation
    • Unclear ownership
    • Failed automation
    • Insufficient redundancy

    The source does not prescribe a fixed taxonomy, so the reviewer should choose the factors supported by the recorded evidence.

    Why include what went well?

    Incident review should preserve effective controls as well as record failures, and the source postmortem structure includes "what went well" for that reason. That section can capture that the alarm detected the issue quickly, topology identified the affected service, the backup path worked, checkpoint recovery protected training progress, the correct runbook was matched, or on-call escalation reached the right engineer.

    Reliable operations depend on keeping the controls that worked as much as on adding new ones after every incident. A balanced review creates a clearer improvement plan.

    How are improvement items tracked?

    The source assigns each improvement item to a person, due date, and status, and sends automatic reminders when an item becomes overdue. I think this is one of the strongest parts of the design, because a postmortem without tracked actions can become a document nobody uses.

    Each improvement should therefore answer what will change, who owns it, when it is due, and what its current status is. The source turns postmortem improvement into operational work instead of leaving it as prose.

    Why should postmortems enter the knowledge base?

    Incident experience should be reusable. The source automatically places the postmortem into the knowledge base, which lets future operators and the AI assistant retrieve similar incidents. A later engineer can ask whether we have seen this failure before, what the root cause was, what action worked, and which improvement was supposed to prevent recurrence. Postmortem work becomes organizational memory that way.

    For the assistant layer, how AI assistants use alarms, work orders, runbooks, and infrastructure documents to help IT operations teams explains how historical incidents can improve later troubleshooting.

    How does automatic postmortem generation improve data quality?

    It reduces the gap between the incident and the review. The source content collection makes a similar point: timeline, changes, alarms, handling actions, impact data, and conclusions should be accumulated during response rather than recreated afterward.

    When evidence is captured automatically, timestamps are more accurate, work-order actions are already linked, audit logs identify who executed what, the affected resource is known, and the operator spends less time copying data. None of that guarantees the analysis is correct, but it improves the evidence available for the analysis.

    Should chat messages be included?

    The broader source content says chat records can be part of the material that naturally accumulates during response, but the v3.2 postmortem page specifically mentions event timeline, alarms, and work orders. If chat integration exists, it can provide additional context.

    The source does not define a required chat-ingestion capability for the v3.2 postmortem feature, so the core postmortem should not depend on chat. It should work from structured operational records first.

    How should changes be included?

    Change history is useful because a recent change may be related to the incident. The source governance and data foundation retain:

    • Configuration changes
    • Before and after values
    • Operation audit
    • Time

    The postmortem generator can place relevant changes on the incident timeline. That does not prove the change caused the incident, but it gives reviewers evidence to weigh.

    For the audit process, how IT teams can create an auditable change management process for infrastructure operations explains why exact before and after values matter.

    How does an incident postmortem connect to SRE metrics?

    The postmortem can explain why reliability metrics changed. The source SRE layer tracks:

    • MTTD
    • MTTR
    • Automatic-remediation ratio
    • Error budget
    • Burn rate

    The incident timeline contains the events needed to explain MTTD and MTTR, the service-impact record explains error-budget consumption, and the remediation history shows whether automation helped. The improvement items then become reliability work, which closes the loop from measurement to incident to improvement.

    Can the system generate the final postmortem automatically?

    It can generate the first draft automatically, but the source does not remove human review. The source says the material is automatically assembled and people make the judgment, and I think that is the correct boundary.

    A system can reconstruct timestamps reliably, link alarms and work orders, and propose a root cause based on existing analysis. Questions such as whether the process was appropriate, what the organization should change, and whether the incident was preventable still require accountable human review.

    What should the final review workflow look like?

    A source-grounded workflow starts when the incident closes and the system assembles the timeline, alarms, work orders, and remediation into a generated draft. A reviewer confirms root cause and contributing factors, the team confirms what went well, and improvement actions receive owners and due dates. The postmortem is then approved and enters the knowledge base, while open improvements remain tracked until closed. The source provides all of these pieces.

    What should an automatic postmortem interface show?

    A practical view can show:

    • Generated timeline
    • Linked alarms
    • Linked work orders
    • Root cause
    • Contributing factors
    • What went well
    • Improvement items
    • Owner
    • Due date
    • Status
    • Knowledge-base publication

    The source v3.2 example says this draft can be produced in about 40 seconds once the incident material is available. A platform example that uses this automatic-assembly model is Sensaka.

    If I were automating postmortems, I would focus less on making the final prose sound polished and more on making the evidence impossible to lose. Capture every state change, alarm, work order, approval, remediation action, and validation during the incident. Once those events are structured, generating the first draft is easy. The valuable human work is deciding what the evidence means and what must change next.

    Frequently Asked Questions

    What does the source automatically assemble into a postmortem?

    The source builds the draft from the event timeline and automatically links alarms and work orders, then structures the result into timeline, root cause and contributing factors, what went well, and improvement items.

    How fast does the source example generate the first draft?

    The v3.2 source says the postmortem draft is generated in about 40 seconds once the incident material is already captured in the platform.

    Does automatic generation replace human judgment?

    No. The source explicitly frames automation as gathering and assembling the evidence. People still judge root cause, contributing factors, what went well, and which improvements should be accepted.