KITE 2025 New Product Award — Local IT | SACEEC

What Is a Post-Incident Review? Purpose, Agenda, and Lessons Learned

An incident that teaches you nothing is just damage. A good post-incident review turns it into a permanent improvement. Here is exactly how to run one.

Free PDF GuideDownload this guide as a PDF

When an incident is finally contained, the temptation is to exhale, close the ticket, and move on. That is exactly when the most valuable work gets skipped. A post-incident review (PIR) is the structured conversation that turns a painful event into a permanent improvement: clearer controls, faster detection, fewer surprises next time. Run it badly and it becomes a blame session that people learn to dread and avoid. Run it well and it is one of the highest-return rituals in your entire risk and resilience programme. This guide explains what a PIR is, how to keep it blameless, when to hold one, the agenda that actually works, and how to make sure the lessons survive past the meeting.

Watch: how to run a blameless post-incident review Watch: How to run a post-incident review (short walkthrough)
i

What You'll Learn

By the end of this guide you will understand what a post-incident review is and why it matters, how to run it as a blameless exercise, which incidents warrant one, a tested agenda covering timeline through action items, how to write up lessons learned so they stick, how to drive follow-through, and a complete PIR template you can adapt today.

What a Post-Incident Review Is

A post-incident review, also called a post-mortem, after-action review, or retrospective, is a structured session held shortly after an incident is resolved to understand what happened, why, how the response went, and what should change. It is not a status update, and it is not a disciplinary hearing. It is a learning exercise whose only product is a shared, accurate account of the event and a short list of concrete improvements.

A PIR sits at the end of the incident lifecycle. The earlier stages (detection, triage, containment, and recovery) are covered in what good incident management looks like, and every incident you review should already be captured in your incident register. The review is where the organisation extracts the maximum value from an event it has already paid for.

A good PIR answers four questions without flinching:

  • What actually happened? A factual, time-stamped account, free of assumptions.
  • Why did it happen? The contributing causes, not just the trigger.
  • How well did we respond? What helped, what got in the way, what we got lucky on.
  • What will we change? A small number of owned, dated actions.

Why It Has to Be Blameless

Whether your PIRs produce value comes down mostly to one thing: whether people feel safe being honest in them. The moment a review becomes a hunt for "who caused this", three things happen. People stop volunteering information, the real contributing causes stay hidden, and your detection of the next incident gets slower because nobody wants to be the one who raises their hand.

A blameless review assumes that everyone acted reasonably given the information and tools they had at the time. When someone "made a mistake", a blameless PIR treats that as a signal that the system allowed the mistake: a missing guardrail, a confusing interface, an unclear runbook, an alert that fired too late. You fix the system, not the person.

i

Pro Tip

Replace "who" questions with "what" and "how" questions. Instead of "Why did you deploy on a Friday?", ask "What in our process made a Friday deploy seem like the right call, and what would have made the risk visible?" Same facts, completely different psychological safety.

Blameless does not mean accountability-free. Actions still get owners and dates, and systemic patterns of negligence are still escalated through the right channels. But that escalation happens through management, not through the review. The PIR itself stays focused on learning.

Want the full framework with worked examples?

When to Hold One

You cannot run a full review for every minor hiccup, and you should not skip one just because the incident is now over. Tie the decision to severity and to whether there is something worth learning. A practical trigger matrix:

Trigger Review Depth Timing
Critical / major incident (significant outage, data loss, regulatory notification) Full PIR with formal write-up Within 3 to 5 business days of resolution
Moderate incident with customer or financial impact Standard PIR, lightweight write-up Within 1 to 2 weeks
Minor incident that revealed a surprising weakness Short review, action items only At the next team retro
Recurring low-severity incidents (same type, repeated) Thematic review across the cluster Quarterly or on trend trigger
Near-miss (no impact, but easily could have) Short review, treat as a free lesson Within 1 to 2 weeks

That last row matters more than people expect. Near-misses are incidents that gave you all the warning and none of the damage. Reviewing them is the cheapest learning you will ever buy. If you spot the same near-miss type repeatedly, that is a strong signal to dig in. See how to spot incident trends over time.

!

Important

Do not hold the review while the incident is still live. The job during an incident is to contain and recover; the job afterwards is to learn. Mixing the two leads to half-baked containment and a rushed, defensive review. Wait until the dust has settled, but not so long that memories fade. Three to five days is the sweet spot for major incidents.

The Agenda That Works

A PIR should be time-boxed. Sixty to ninety minutes for a major incident is plenty if it is run well. The facilitator should ideally be someone who was not directly responsible for the response, so the conversation stays neutral. Walk through these five blocks in order.

1. Build the Timeline

Start by reconstructing what happened, minute by minute, from the first signal to full recovery. Pull timestamps from your monitoring, chat logs, and the incident register so the timeline is anchored in evidence rather than memory. Capture detection time, escalation time, key decisions, and resolution time. The timeline is the factual spine everything else hangs from.

2. What Happened (and Why)

Once the timeline is agreed, work through the contributing causes. Use a technique like the "five whys" to move past the immediate trigger to the underlying conditions. Most serious incidents have several contributing causes, not one root cause, so name them all.

3. What Worked

Deliberately spend time on what went right. Did an alert fire as designed? Did the runbook hold up? Did the on-call rotation respond fast? Naming what worked tells you which controls to protect and reinforces good practice. Skipping this step is a common reason reviews feel punishing.

4. What Didn't Work

Now the friction points: where detection was slow, where the runbook was wrong, where ownership was unclear, where a control failed. Keep it about the system. This is where the richest improvement ideas come from.

5. Action Items

Convert the discussion into a short list of specific, owned, dated actions. Resist the urge to generate twenty actions; five well-chosen ones that actually get done beat twenty that languish. Each action needs a named owner (a person, not a team) and a due date.

Agenda Block Time Output
Timeline reconstruction 15 to 20 min Agreed, time-stamped sequence of events
What happened & why 20 min List of contributing causes
What worked 10 min Controls and behaviours to preserve
What didn't work 15 min Weaknesses and friction points
Action items 15 min 3 to 6 owned, dated actions

Documenting Lessons Learned

A review that lives only in people's heads decays within weeks. The write-up is what makes the lesson durable and searchable. Keep it concise, a one-to-two page document for most incidents, but make sure it captures the essentials: a short summary, the timeline, contributing causes, what worked and didn't, the action items, and any links to the incident register entry and affected risks.

The most under-used field is "lessons learned" written as transferable statements. "Our database failover took 40 minutes because the runbook referenced a decommissioned host" is a fact. "Runbooks must be tested against live infrastructure quarterly" is a lesson, because it transfers to other systems and other teams. Write the lesson, not just the fact.

Example

Lesson Learned, Done Two Ways

Weak: "The on-call engineer didn't know who to escalate to, so it took a while to get the right people."

Strong: "Escalation paths for payment-system incidents were undocumented, adding roughly 25 minutes to response. Lesson: every tier-1 service needs a documented, tested escalation path in the runbook, reviewed at each on-call handover."

Store the write-up somewhere searchable and linked from the incident record. When a similar incident happens in eighteen months, the responder should be able to find your review and stand on its shoulders.

Follow-Through

This is where most PIR programmes quietly fail. The meeting happens, the document gets written, the action items get listed, and then nobody tracks them to completion. Six months later the same incident recurs, and the review notes show the fix was "agreed" but never done.

Treat PIR action items like any other commitment. They go into a tracker with owners and due dates, they get reviewed in regular operational meetings, and overdue ones get escalated. A simple discipline is to start every PIR by reviewing the open actions from previous reviews. If old actions are still open, that is itself a finding.

!

Important

An action item without an owner and a date is a wish, not a commitment. If you cannot assign a real owner in the room, the action is too vague, so break it down until you can. And if your reviews keep producing the same action item across multiple incidents, escalate the underlying gap as a risk in its own right rather than re-discovering it every quarter.

Closing the loop is also what connects incidents back to your wider risk picture. The systemic weaknesses a PIR surfaces often belong in your risk register as control gaps or new risks, a connection explored in detail in how incidents connect to your risk register.

Example PIR Template

Here is a complete, copyable template for a standard post-incident review. Adapt the fields to your context, but keep the blameless framing and the action-item discipline.

Example Template

Post-Incident Review: INC-2026-0184, Payment Gateway Outage

Severity: Major (customer-facing outage, ~90 minutes)

Date of incident: 22 May 2026  |  Review held: 26 May 2026

Facilitator: J. Mthembu (Risk & Resilience), not part of the response team

Summary: A configuration change to the payment gateway introduced an invalid timeout value, causing transactions to fail intermittently for 90 minutes before rollback restored service.

Timeline: 09:12 change deployed · 09:31 error-rate alert fired · 09:48 incident declared · 10:05 cause identified · 10:42 rollback complete · 10:45 service confirmed healthy.

Contributing causes: (1) Change validation did not check timeout bounds; (2) Alert threshold was set high, delaying detection by ~19 minutes; (3) Rollback runbook referenced an outdated dashboard.

What worked: On-call responded within 3 minutes of escalation; rollback procedure ultimately succeeded; customer comms went out promptly.

What didn't: Detection was too slow; the runbook's dashboard link was dead; no pre-deploy validation of config bounds.

Action items:

• Add config-bound validation to the deploy pipeline (Owner: A. Naidoo, Due: 12 Jun 2026)
• Lower payment error-rate alert threshold and test (Owner: T. Dube, Due: 5 Jun 2026)
• Audit and fix all runbook dashboard links (Owner: S. Patel, Due: 19 Jun 2026)

Lessons learned: Configuration changes to tier-1 services need automated bound-checking before deploy; alert thresholds for revenue-critical paths should target sub-5-minute detection.

Linked risk register entry: R-2026-031 (Payment platform availability), likelihood reassessed upward.

Common Mistakes to Avoid

1. Turning It Into a Blame Session

The fastest way to kill a PIR programme is to use it to assign fault. Once people fear the review, they hide information, and you lose the very honesty the exercise depends on. Keep it systemic.

2. Skipping the "What Worked" Section

Reviews that only catalogue failures feel like punishment and miss half the lesson. Knowing which controls held is as valuable as knowing which failed, because it tells you what to protect.

3. Generating Too Many Actions

Twenty action items is a to-do list nobody finishes. Pick the three to six changes that matter most and actually drive them to done.

4. Never Closing the Loop

Action items that are listed but never tracked are the single most common failure mode. Track them like any other commitment, and review open ones at the start of the next PIR.

5. Holding It Too Late (or Too Early)

Wait until the incident is fully resolved, but not so long that memories fade. For major incidents, three to five days is right.

6. Not Connecting Findings to Risk

A PIR that ends at "we fixed the bug" misses the bigger picture. The weaknesses it surfaces often belong in your risk register, where they can be tracked and prevented across the whole organisation.

Key Takeaways

Summary

  • A post-incident review turns a resolved incident into permanent improvement by examining what happened, why, how the response went, and what to change.
  • The review must be blameless, fixing the system rather than the person, or people stop being honest and the value collapses.
  • Trigger a full PIR for major incidents, lighter reviews for moderate ones, and don't skip near-misses, because they are free lessons.
  • Run a tight agenda: timeline, what happened and why, what worked, what didn't, and a short list of owned, dated actions.
  • Write durable lessons learned, track action items to completion, and feed systemic findings back into your risk register.

Frequently Asked Questions

What is the difference between a post-incident review and a root cause analysis?

A root cause analysis is one technique used inside a review to identify contributing causes. A post-incident review is the broader session that includes the timeline, what worked, what didn't, action items, and follow-through. Root cause analysis is part of it, not a replacement for it.

Who should attend a post-incident review?

The people who responded to the incident, anyone who owns the affected systems or processes, and a neutral facilitator who was not part of the response. Keep it small enough for honest discussion. For major incidents you may also invite a risk or compliance representative to capture register and reporting implications.

How do I keep a review blameless when someone clearly made a mistake?

Treat the mistake as a system signal. Ask what allowed it: a missing guardrail, an unclear runbook, a confusing interface, a missing check. Fixing the condition prevents the next person from making the same error. Genuine misconduct is handled separately through management, not in the review itself.

Should I run a review for near-misses?

Yes, a short one. A near-miss gave you the warning without the damage, which makes it the cheapest learning available. Reviewing near-misses is one of the most effective ways to prevent the full-blown version, and recurring near-misses are a strong signal to investigate a deeper trend.

How long should a post-incident review take?

Sixty to ninety minutes is enough for most major incidents if the timeline is prepared in advance. Minor incidents can be covered in fifteen to thirty minutes at a regular team retro. The write-up usually takes another hour or two for a major incident.

What happens to the action items after the review?

They go into a tracker with named owners and due dates, get reviewed in regular operational meetings, and overdue ones get escalated. Start each new review by checking the status of prior actions, since outstanding ones are a finding in themselves. Systemic gaps that keep recurring should be promoted to your risk register.

Save this guide for later

Download the PDF version to read offline or share with your team.

Co-Founder & ERM Practitioner

An enterprise risk management practitioner with experience across healthcare, public sector, and regulated environments. Phumi focuses on translating ERM frameworks into practical, decision-relevant processes.

Co-Founder & ERM Practitioner

Specialises in enterprise risk management through risk assessments, data analysis, and mitigation planning. Contributes to compliance oversight, risk reporting, and monitoring of key risk indicators.