Debrief: how to run one that changes the next time

Last updated: · Reviewed quarterly

A debrief is a structured review of one event, run soon after it, by the people who were in it, comparing what was meant to happen with what did. It is not a status meeting, not a post-mortem written for somebody else, and not the sprint retrospective. Two meta-analyses have measured it and both found the same thing: teams that debrief perform measurably better than teams that do not, and the effect gets larger when the review is anchored to a record of the event rather than to what four people remember.

What a debrief is, and what it is not

The word does double duty. In ordinary usage it means being questioned after a mission, and a large share of the people who look it up want no more than that. The practice is narrower and more useful: a facilitated review of a discrete event, held close to it in time, in which the participants themselves are the audience.

Three features distinguish it from the reviews it gets confused with. The scope is one event rather than a period of time. The proximity is minutes or hours rather than weeks. And the people in the room were in the event, which means the review is for them and not a report for somebody upstairs. Take away any one of those and it becomes a different meeting with different economics.

ReviewScopeTimingAudience
DebriefOne event: a call, a launch, a session, an incidentImmediately, or the same dayThe people who were in it
RetrospectiveA period of work, usually a sprintOn a fixed cadenceThe team, about its own process
Post-mortemA finished project or a failureWeeks later, once it is overOften somebody other than the participants
Status meetingOngoing workRecurringWhoever needs to know

The retrospective boundary is the one worth holding. A retrospective asks how the team is working; a debrief asks what happened in that one thing and why it differed from the plan. They are complements, not substitutes, and the formats for the first are covered in standup and retrospective formats.

Where the practice comes from

The formal version is the US Army's After Action Review, and its lineage is unusually traceable. S.L.A. Marshall, a combat historian in the Second World War, began interviewing units as a group immediately after engagements rather than collecting individual accounts later. The Army Research Institute turned the method into doctrine over the following decades, and the Center for Army Lessons Learned published A Leader's Guide to After-Action Reviews as TC 25-20 in 1993, which is still the document everything else descends from.

Two things in that doctrine are usually lost in translation. The first is that an AAR is explicitly not a critique and not an evaluation: it is a professional discussion in which rank is set aside, and the leader who ran the event is a participant rather than the judge. The second is that most AARs are informal. The formal version with a facilitator, a prepared timeline and a written product is the exception; the standard case is a twenty-minute conversation held on the spot.

The commercial adaptation that matters is Marilyn Darling, Charles Parry and Joseph Moore's Learning in the Thick of It in the Harvard Business Review, which observed that organisations copying the AAR usually copy the meeting and drop the part that makes it work: the review is supposed to feed the planning of the next event, not to be filed.

The four questions, worked

The doctrine is four questions. They look trivial and each of them has a characteristic way of going wrong.

  1. What was supposed to happen? The intent, stated as it was stated at the time. This question fails when nobody can produce the original intent, so the room reconstructs a version of it that conveniently matches the outcome. If the intent was never said out loud, the debrief has nothing to compare against and quietly becomes a discussion of how everyone feels it went.
  2. What actually happened? The sequence of events, agreed before anyone interprets them. This question fails when it is skipped as obvious. It is not obvious: people in the same room routinely disagree about the order of events, and the disagreement is the most informative thing in the meeting.
  3. Why were there differences? The analysis. This is the question the whole exercise exists for, and it is the one that gets four minutes because the first two ate the timebox. It also fails in a subtler way covered further down: answered from memory, after the outcome is known, it produces a tidier and more inevitable story than the one the room was actually in.
  4. What can we learn? The output. This fails when it produces a list of observations instead of one change with a name and a date against it. A debrief that ends in observations has produced nothing, however good the discussion was.

What the evidence says

Unusually for a meeting practice, this one has been measured twice at scale, and the second measurement is recent.

Scott Tannenbaum and Christopher Cerasoli's meta-analysis in Human Factors in 2013 pooled the available studies of team and individual debriefs and found that debriefing improved subsequent performance by roughly 25 per cent over control groups. That figure is the one most often quoted, and it holds across medical teams, military units and student groups alike.

Nathanael Keiser and Winfred Arthur Jr's larger meta-analysis in the Journal of Applied Psychology in 2021 is the more useful one, because it went looking for the conditions under which the effect is larger or smaller. The headline effect was substantial, and two moderators matter for anyone designing the practice rather than defending it. First, reviews anchored to objective records of the event were consistently more effective than reviews conducted from recollection alone. Second, self-led debriefs worked, which is the finding that makes the practice affordable, but they were the case where the objective record mattered most: a team reviewing itself, without an outside facilitator, has nothing to correct a shared misremembering except the record.

One counterintuitive result is worth planning around. Longer is not better in the way you would expect. Debriefs running well beyond twenty minutes tended to help individuals more than they helped the team, which is an argument for a short, tightly scoped review after every event rather than a long one after every third.

Running one in twenty minutes

The informal AAR is the version worth adopting, because it is the version people will actually keep doing. It has five moving parts and none of them require preparation.

Who is in the room. The people who were in the event, and nobody else. The presence of anyone with authority over the participants changes what gets said, which is why the doctrine is explicit about setting rank aside and why a debrief with the sponsor in it is a different meeting.

When. As close to the event as you can manage. The value decays fast, and the decay is not gentle: the reconstruction problem described below starts the moment the outcome is known.

How long. Twenty minutes, held to. The four questions get roughly three, five, eight and four, which deliberately gives the analysis question the largest share and is the opposite of how these meetings usually run.

Who leads. Anyone. Rotating it is better than fixing it, because the person who led the event has the strongest incentive to lead the review toward a comfortable answer.

What comes out. One change, with one owner and one date. Not a list. If the room genuinely produced three, write three and accept that two of them will not happen; a single owned change that gets made is worth more than a page of observations that get filed. The discipline is the same one that makes a retrospective produce anything, and the argument for one owner rather than a team is in action items.

The debrief in the fields that engineered it

Two professions have spent decades refining this, and both assume something a company meeting almost never has.

Medicine. Simulation training in healthcare produced PEARLS, Walter Eppich and Adam Cheng's blended framework, which sequences a reactions phase, a description phase, an analysis phase and a summary, and lets the facilitator choose between learner self-assessment, focused facilitation and directive teaching depending on the gap being addressed. The lightest tool from the same tradition is plus-delta: what worked, what would you change. It is two columns and it is the version to reach for when twenty minutes is not available. Cheng and colleagues' paper on plus-delta is worth reading precisely because it takes a technique that looks too simple seriously.

Aviation. Line Oriented Flight Training debriefs were studied directly by NASA's flight cognition group. Key Dismukes, Kimberly Jobe and Lori McDonnell's training manual and their analysis of instructor technique found that the instructors who produced the most crew participation talked least, asked open questions, and let silence run. That is a transferable finding, and it is the single most common failure of a debrief led by the person who ran the event.

What both fields assume is the record. Nobody debriefs a simulation from memory when the session was filmed, and nobody debriefs a flight without the data. Meetings are the exception, and the reason is mechanical rather than philosophical.

Four variants you will meet

The four questions are the same every time. The intent statement is not, and getting it right is most of the work.

EventThe intent to compare againstThe trap
A meetingWhat the meeting was called to decide or produceReviewing whether it felt useful rather than whether it produced the decision
A launchThe dated plan, and the number it was supposed to moveWaiting for the metrics, by which time nobody recalls the launch
An incidentThe expected behaviour of the system, and the expected response timeSliding into blame, which the doctrine explicitly forbids
A research sessionThe questions the session was meant to answerSkipping it because the recording exists, so the debrief is postponed to analysis and never happens

The research-session debrief between moderator and note-taker, held in the ten minutes after the participant leaves, is one of the highest-value debriefs anyone runs and one of the first to be dropped when the schedule tightens. Its whole purpose is to capture the interpretation while it is still fresh, before six sessions blur into one.

For incidents specifically, the analysis question is where 5 Whys and similar techniques belong, and those are covered alongside the retrospective formats rather than repeated here.

Memory is the weak link

The debrief happens ten minutes after the event, which is precisely when nobody has a record yet. The notes are half written, the recording is still uploading, and if the conversation happened at a desk or in a huddle room there was never a recording to begin with. So the team debriefs the only thing available, which is four people's recollection of a thing they were participants in.

That is a worse source than it feels like. Laura Guilbault, Fred Bryant, Justin Brockway and Emil Posavac's meta-analysis of hindsight bias pooled 95 studies and 252 independent effect sizes and found a mean effect around .39. Hindsight bias is specifically the distortion of recalled prior judgement toward the known outcome, which is a precise description of question three going wrong. Asked why there were differences, after the outcome is known, a room reliably produces a cleaner and more inevitable account than the one it was actually in. The people who were most surprised at the time are the most confident afterwards that they saw it coming.

This is where the objective-record moderator in the 2021 meta-analysis stops being an abstraction. If the intent was stated at the top of the call, it is in the record verbatim, and question one has an answer instead of a reconstruction. If commitments were made, they are in the record too, and the characteristic failure of four people remembering four versions of what was agreed becomes a retrieval problem rather than an argument.

This is the problem Earkeep was built for. It records your day continuously on your own device, transcribes it there, and writes the result to plain files you own, so a debrief run at 15:10 about the 14:00 call has the transcript of the 14:00 call, without anyone having remembered to start a recording or invite anything to the meeting. You mark the stretch of the day that was the call, and the lines are there in order with timestamps. An agent working against that stretch can draft the first version of the debrief document, which is worth doing mainly because it spends the twenty minutes on question three instead of on transcription. Getting consent to record is yours to do, and there is a page on telling people you are recording.

The claim not to make is the one about attribution. The record cannot tell you who said any of it. A debrief that needs to know who committed to something is not a question the transcript answers, and this is survivable, because the doctrine is explicitly about the group's actions rather than individual accountability. The record supplies the what. The room supplies the who.

What debriefing also means, and why the distinction matters

There is a second practice with the same name and a very different evidence base, and conflating them has done real harm. Critical Incident Stress Debriefing is a single-session psychological intervention delivered after a traumatic event. The Cochrane review by Suzanna Rose, Jonathan Bisson, Rachel Churchill and Simon Wessely found no evidence that single-session psychological debriefing prevents post-traumatic stress disorder, and some evidence that it may increase risk in certain groups. Its conclusion was that compulsory debriefing of victims should cease.

None of that bears on the performance review this page describes, which is a discussion of events rather than an intervention on emotional processing. But if your organisation runs both, name them differently. A team that has heard that "debriefing does not work" has usually heard about the wrong one.

Where this fits

A debrief is worth running after events that were meant to achieve something specific, which is not every meeting on the calendar. Deciding which meetings should exist at all comes first, and that argument is in meeting overload. The cadence-based cousin of the debrief, the sprint retrospective, is in standup and retrospective formats, along with the ceremony's place in the wider framework covered by Agile, Scrum and Kanban. The industrial version of the same loop, the control phase that checks whether an improvement held, is in Lean Six Sigma, CPM and PERT. And what happens to the output of many debriefs over a year, once they accumulate into something you can search, is from what was said to what you know.

Sources