Ideas

Retrospective games, and the one problem they are all solving

The first answer sets the frame, and most retro formats leave you to work around it yourself.

· Updated

Most retrospectives are settled by whoever answers first. Somebody says the sprint felt rushed, and from that moment the room is editing that sentence instead of writing its own. Lorenz and colleagues ran repeated estimation tasks with 144 people and showed each of them what the others had guessed. The estimates converged, the group's collective error did not fall, and confidence went up anyway.

That is a retrospective in miniature. The team agrees faster, feels better about agreeing, and gets no closer to what actually happened.

The fix is not a better question and not a livelier format. What breaks a retro is anchoring. The first answer sets the frame, the second adjusts to the first, and by the fifth person you are collecting agreement with whoever spoke quickest. Some of the formats below have a way around that built in, and the rest need one added. The difference is worth knowing before you pick one.

Why retrospectives turn into agreement

Anchoring is the bias where a number you have just heard drags your own estimate towards it. Onuki, Honda and Ueda define it as a case where exposure to some piece of information affects the numerical estimate you make next. The experiment everyone quotes is Tversky and Kahneman's, from 1974, and this is Grau and Bohner's account of it. People first asked whether African states were more or less than 10% of the United Nations went on to estimate 25%. People given 65% as the comparison estimated 45%. A wheel spun in front of them had produced the number.

Grau and Bohner, whose two studies with 184 participants tested where the effect comes from, describe it as stable over time and independent of the participants' motivation or expertise. Caring about the answer does not protect you from it, and neither does knowing the subject.

A single small ink-black square sits alone at the far left of a wide empty page. Ten more squares run away to the right: the nearest is filled warm ochre and barely bigger, and the nine behind it are outlines that grow steadily

A retro is not a quiz about the United Nations, so the honest question is whether the same thing happens to a team discussing its own month. The group research points the same way, though it runs on invented cases rather than on a team's own work. De Wilde, Ten Velden and De Dreu write that under conformity pressures, group members are biased towards exchanging commonly known information and away from exchanging unique information. Their groups were three people solving a fictional murder. Moser and colleagues ran the same kind of task with 174 students picking a candidate from a fictional shortlist. Shared facts were more likely to be introduced into the discussion than the ones a single person held.

The thing a retro exists to find is the fact only one person has. Anchoring is the mechanism that keeps it in their head.

What going round the table costs you

Going round the table is the fallback, either as the whole format or as the readout after the writing. It carries two measured costs before anybody has said anything difficult.

The first is production blocking. Paulus, Baruah and Kenworthy put it in one line: in face-to-face group settings, only one person can effectively share ideas at one time. Waiting is not free. You spend it holding your own thought rather than having the next one, and groups of four end up generating about half as many ideas as the same four people working alone.

Eight outlined figures stand side by side on a single ground line, and only the first has an open mouth and empty hands, with a small ochre card at its feet. Each of the other seven holds a filled card flat against the chest, slate blue for the third and sixth figures and warm ochre for the rest.

The second is evaluation apprehension, which is the polite name for knowing you are being judged. Zhou and colleagues ran three experiments on it. Exposure to other people's ideas and evaluation apprehension can produce deficits in the number and the categories of ideas, without changing how novel those ideas are. Participants, the authors suggest, may have paid so much attention to others' ideas that it limited their views to a small number of categories.

A retro raises the second of those past anything a lab can stage. Your manager is often in the room, the answer has your name on it, and the thing worth saying is usually the thing that makes somebody look bad.

Do retrospectives actually work

The ceremony is mandated, which is not the same as being useful. The Scrum Guide gives the Sprint Retrospective the purpose of planning ways to increase quality and effectiveness, and timeboxes it to three hours for a one-month sprint. The twelfth principle behind the Agile Manifesto says the team reflects at regular intervals on how to become more effective, then tunes and adjusts its behaviour accordingly.

The evidence that structured reflection works is strong, and it comes from the debrief literature rather than from agile. Keiser and Arthur pooled 83 studies covering 955 teams and 4,684 individuals and put the overall effect of an after-action review at d = 0.92. As effect sizes go, that is a large one.

There is a condition on that number. A debrief reviews one specific performance, and a sprint retrospective usually reviews a fortnight from memory. Matthies, Dobrigkeit and Hesse point out that the retrospective activities proposed so far often do not rely on project data, depending solely on the perceptions of team members. Memory is the thing anchoring bends.

So bring a record. Pull the tickets, the incidents, the dates the scope moved. Put them up before anyone speaks. The games below are for the part a record cannot give you, which is what people made of it.

The change that fixes almost every retro format

Write first, alone, in silence. Reveal everything at once. Then talk.

That one change removes two of the three costs and cuts the third. Nobody waits for a turn, so nothing is blocked. Nobody reads an answer before writing their own, so no card is anchored on another. Nothing carries a name while it is being written, which is most of the nerves gone. Paulus and colleagues describe two of those three for methods where everyone contributes at once, plus one the list above leaves out. There is less production blocking, and less evaluation apprehension because the method minimises awareness of the sources of the ideas. Motivation goes up too, because everybody stays active. Anchoring is not one of their variables, so that part of the case rests on the studies further up.

The third cost is cut rather than cured. On a team of six, people can recognise each other's phrasing, and the anonymity of a card is thinner than it looks.

Twelve blank cards filled solid slate blue lie in three tidy rows of four on the left. The same twelve lie on the right as outlines, in the same three rows of four, each one tilted a few degrees

The mechanic already has a name in agile. It is the silent writing phase, and plenty of teams run it on stickies or a shared board and then undo it by reading the notes out one at a time. That readout is the round-robin again, with an extra step.

Reveal is this shape as an activity. Everyone answers on their own phone, no other participant sees a card until you turn it over, and the cards are anonymous unless you deliberately choose otherwise. You see the text before the room does, so that you can drop a card that should not go on a projector. You can flip them all at once or step through them one at a time, and stepping through puts a beat between the cards. Its Silent retro template carries the prompt "What is the one thing we should stop doing?" and lets each person send two answers.

Two answers rather than one is deliberate. Paulus and colleagues describe the ordinary search for ideas as tapping the most obvious or common ideas first, with the rarer ones surfacing later. The second card is the one you asked for.

When a word cloud helps, and when it quietly lies

A live word cloud is the wrong instrument for the hard question. Words appear as they arrive and grow as they repeat. Anyone still typing can read what is already up, and from that point the room is agreeing rather than answering. That is the anchoring problem with better typography. Text that is already written has none of it, so paste three retros' worth of cards into a word cloud generator to see what keeps coming back.

One of the word cloud templates asks "What should we stop doing?". Ask that one through a reveal instead. The word cloud earns its place a few minutes earlier, on the question where convergence is the point rather than the failure.

That question is the temperature check. Sum up the month in one word and you get a picture of the shared feeling in about forty seconds, each word sized by how many people typed it. Repeats grow, which is the whole signal. Six people typing "rushed" is a finding, and it is a finding you want the room to see before it decides what to talk about.

The stage below is that mechanic: one word each, arriving live, and the repeats growing as they land. Run it, read it out loud, then go silent for the real question.

On the big screen

Word cloud

What's one word that describes your ideal meeting?

Waiting for responses...

Retrospective games, and what each one is for

Most of what follows is a container for the same three questions: what happened, what it means, what we do next. The rest are smaller. A check-in tells you what kind of room you are in, and a voting round only does the third question. Pick by the problem you have, not by the format you have not tried.

If the problem is Run
The team says everything is fine The silent stop-doing round
Everything is "it went badly" and nothing is specific Four Ls
Nobody mentions risk Sailboat
Mood is off and nobody says why Mad, Sad, Glad
Twelve actions, none finished Dot voting, then one owner
The room looks bored Lean coffee retro
You have no idea who is in the room Explorer, shopper, vacationer, prisoner
Two people disagree and will not say so Tier list of the improvements

Every row here wants a silent writing phase in front except the check-in, where the private pick is already the whole activity. The top row is nothing but that phase, and where it changes a game the entry says so.

Two vertical columns of small marks, where the left column clusters tightly in the middle and the right column splits into two clumps at top and bottom

Sailboat

The boat is the team. Wind is what pushed you forward, anchors are what held you back, rocks are the risks you can see ahead, and the island is where you are trying to get to. Everyone adds to all four.

What it is for: getting risk and destination into the same picture as the complaints. The rocks are why it survives. Most prompts that look forward ask what to do next, and the rocks ask what could go wrong.

How it fails: the anchors fill up and the rocks stay empty, because complaints are easier than forecasts. Collect all four in one silent phase, then open the discussion on rocks rather than on anchors.

Mad, Sad, Glad

Three buckets. What made you angry, what disappointed you, what you were happy about. Written privately, then revealed.

What it is for: the emotional signal a metrics review cannot see. It is the right pick after a bruising quarter, a launch that slipped, or a reorganisation. Put the dates on screen first, so that a mood attaches to an event.

How it fails: it collects mood and stops. Every item in the mad column needs one follow-up question, which is what happened and on what date. Feelings are the index. The event is the entry.

Start, Stop, Continue

Three columns of behaviour, phrased as actions rather than observations.

What it is for: producing something a person can do on Monday. Every column is already phrased as an action, which is the shortest route from a complaint to a change.

How it fails: the stop column holds the real information and it is the one people will not sign. Run stop anonymously and the other two openly. If the anonymous column is the one that fills up, the problem was never the format.

Four Ls: liked, learned, lacked, longed for

Four prompts, one round, everyone writes to all four.

What it is for: feedback that never gets specific. Lacked and longed for pull apart two things a single "went badly" column welds together: a thing that was missing, and a thing somebody wants.

How it fails: liked and learned fill up and the other two stay thin. Ask for at least one item in each of the four before you reveal.

Starfish: keep, more of, less of, start, stop

Five columns where Start, Stop, Continue has three. The two extra ones are degrees rather than switches, which is closer to how a process actually changes.

What it is for: tuning a system that basically works. Use it when the team is fine and you want it better, not when something is on fire.

How it fails: five columns times eight people is a wall of items and no time to sort it. Cap it at two items per column per person, and spend the back half on the two fullest columns.

Timeline

Put the dates on the wall first: releases, incidents, the day the scope moved, the week two people were out. Then ask people to mark where their weeks went well or badly.

What it is for: the memory problem. It puts the record on the board rather than on the slide you opened with, which is the answer to the criticism that these sessions run purely on perception. It also stops the hour collapsing onto the most recent week.

How it fails: building it live is its own workshop and it eats the hour. Bring the dates from the tickets and the incident log, and open with them on screen.

Temperature check

One word each, at the same moment, into a shared picture. Forty seconds.

What it is for: knowing whether this is a repair session or a tuning session. The answer changes which format you run and which question you open with.

How it fails: people treat it as the retro. It picks the question and it does not answer one.

Explorer, shopper, vacationer, prisoner

Everyone privately picks the word that describes how they arrived. Explorers want to dig, shoppers will take one good idea, vacationers are glad to be out of their inbox, and prisoners would rather be elsewhere. Facilitators call it ESVP.

What it is for: knowing what you are working with. A room of four prisoners needs a different session from a room of four explorers, and otherwise you find that out too late to do anything about it.

How it fails: with names attached, picking prisoner in front of your manager costs something, so a room of explorers tells you nothing. Anonymous or skip.

Hopes and fears

Two prompts, written privately, revealed together. What do you want out of the next quarter, and what are you quietly worried about.

What it is for: a new team, a new project, or the first session after a leadership change. Fear is the half people do not volunteer out loud, so it only works written down and revealed together.

How it fails: it produces a beautiful list and no owner. Convert at least one fear into a named check next month.

Fist of five

Everyone shows one to five fingers at the same instant, on a question like how confident you are in the plan. Mountain Goat Software's scale runs from five, a great idea, down to one, a project-threatening decision, and everybody votes on a count of three.

What it is for: a number you can track across sessions, and a fast disagreement detector. An even split between fives and twos means the plan has two readings in the room.

How it fails: hands drift up one at a time and the slower ones match the room. Count out loud so that every hand appears together, and never take it after the manager has shown theirs. A Reveal round set to flip all the cards at once does that, and you see every number rather than an average. A live poll is quicker to set up, but its bars fill as the votes land, so whoever votes last has already read the room.

Lean coffee retro

Everyone writes topics, the room votes, the top topic gets a short timebox, and at the end of it the room votes again on whether to keep going. Take the keep-going vote on a count of three, or the first hand decides it.

What it is for: teams who have run every format twice and are bored. It has almost no fixed content, so there is less of it to wear out, and the agenda comes from the room.

How it fails: with no facilitator watching the clock, the first topic eats the hour. The timer is the format.

Dot voting

Each person gets a small fixed number of votes to spend on the items collected earlier. Three is a good number. They can put all three on one item.

What it is for: choosing what to fix. It turns a long list into one decision without a debate about the list.

How it fails: run straight after an open discussion of the same items, it re-counts the loudest voice. Vote on what to discuss before the discussion. When you have to vote on an outcome afterwards, collect it silently and show it all at once. The same rule holds for every voting round: silent and simultaneous, or it is theatre.

Tier list of improvements

Take the improvements the team proposed and have everybody rank them at once into rows, privately, then show the board.

What it is for: finding disagreement. A ranking exercise where everyone agrees is a slow way of learning nothing. The interesting item is the one half the room put at the top and half put at the bottom.

How it fails: it becomes a popularity contest for the loudest suggestion. Tier List is built for the other outcome. It computes the spread of placements for each item as well as the average, so a wide spread surfaces the argument the room was avoiding. Ranking something other than your own process is a warm-up rather than a decision, and there are ready-made tier lists for that.

The one-improvement rule

Not a game. A rule you apply at the end of any of them. Pick exactly one improvement, name the person, name the date it gets checked.

What it is for: retros that generate twelve actions and complete none. One finished change beats twelve recorded ones.

How it fails: nobody re-reads it. Put it at the top of the next session, before the new list.

The silent stop-doing round

One prompt, anonymous, two answers each, revealed at once: what is the one thing we should stop doing.

What it is for: the team that says everything is fine. It is the one to run when you have time for a single prompt, and it works only if nothing carries a name.

How it fails: somebody asks who wrote the third card. Do not answer, and do not let anyone else.

Retro of the retro

Every fifth session, make the retro itself the subject. What here is worth the hour, and what is habit.

What it is for: stopping a format from outliving its usefulness. It is also the session where shortening the retro is a proposal rather than a complaint.

How it fails: it turns into a complaint about meetings in general. Keep it to this meeting.

Which of these are ceremony

Some of it is ceremony, and pretending otherwise is how retros lose the room.

A format with no record behind it is a memory test. Mad, Sad, Glad six weeks after the events, with no dates in front of anybody, collects the last fortnight and calls it the quarter. Attach it to a timeline.

A format with no decision after it is a support group. Four Ls, Hopes and fears, Sailboat and Starfish all produce a wall of items and no owner unless somebody makes the last ten minutes about choosing. If you have twenty minutes rather than sixty, spend them on one prompt and one decision, and drop the multi-column board entirely.

A format that collects opinions in public is measuring the loudest person. That includes any round-robin, any show of hands that goes up one at a time, and any vote taken once the manager has said which way they lean.

And the honest one: the boat, the starfish, the drawing. What the boat contributes is two prompts that look forward, the rocks and the island, and you would get those from two questions on a slide. The work is done by the silent phase, the simultaneous reveal, and the record you brought. The picture is packaging, and packaging is not nothing, because a team that enjoys the session is likelier to turn up to the next one. Just do not confuse it with the mechanism.

Blameless retrospectives and the prime directive

Norm Kerth's prime directive gets read aloud at the start of a lot of retros. It says that everyone did the best job they could, given what they knew at the time, their skills and abilities, the resources available, and the situation at hand. It sounds like a nicety. It is a correction for a specific bias.

Lea and colleagues gave 212 people the same three incident scenarios with the same investigation findings, and changed only how badly the patient ended up. As the outcome got worse, participants judged the staff involved as more responsible for causing it. Punitive recommendations rose from 5% at no or low harm to 8% where the outcome involved death. The actions were identical. The verdict moved with the ending.

Two rows of six identical figures sit one above the other, each standing behind a blank panel with its hands on the top edge. A small dark circle sits at the end of the upper row and a much larger ochre circle at the end of the lower one. A heavy line runs back from the large circle along the whole length of the lower row

Your retro is run knowing the ending. That is the point of it, and it is also why the room will be harder on the person whose decision preceded the bad week than the evidence supports.

Google's rule for postmortems is the operational version. For a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team. Their reason is practical rather than kind: where finger pointing and shaming prevail, people will not bring issues to light for fear of punishment. DORA states the chain in one sentence, that by removing blame you remove fear, and by removing fear you enable teams to surface problems.

That is the same variable Amy Edmondson has spent a career measuring. Bahadurzada, Kerrissey and Edmondson describe it as a state of low interpersonal risk that helps people ask questions, request help, and admit mistakes. Across 14,943 healthcare workers, they found it positively associated with safety improvement. Google's own team research put psychological safety first in order of importance among the five dynamics it found across 180 teams.

Read the directive out loud if you like it. Then act on what it implies, which is to make the first pass anonymous, so that nobody has to be brave in order to be honest. None of the studies above tested anonymity itself, so that step follows from the mechanism rather than from a measurement.

How to run one, and what to do when it stalls

Sixty minutes is the normal budget for a two-week sprint.

The shape. Five to ten minutes on the record and the temperature check. Five minutes of silent writing. Ten minutes reading the reveal together. Twenty minutes on the two or three things the reveal surfaced. Ten minutes choosing one change and naming its owner. Esther Derby and Diana Larsen's Agile Retrospectives is where the staged version of this comes from.

Fourteen small rectangles are scattered across the upper half of the page. Thirteen are blank outlines and the fourteenth is filled warm ochre and carries a small figure. A single line runs from that filled one down to a lone dark circle in the lower right corner

The order matters more than the format. Record, then feeling, then silent writing, then discussion, then decision. Move discussion earlier and every later step measures the discussion instead of the team.

Timebox the whole thing. The Scrum Guide caps the Sprint Retrospective at three hours for a one-month sprint and says shorter sprints usually get shorter events. An hour is enough for a fortnight. A retro that regularly overruns is a retro people start declining.

Everyone writes, including you. A facilitator who only facilitates is a person whose observations never enter the record. Write your own answers before you read anyone else's, the same as everybody.

Remote and hybrid. Silent-then-reveal survives remote better than most formats, because the tool holds every card back until you turn it over, which a shared wall cannot do. In a hybrid room, put every person on their own phone, including the ones sitting together. The remote half then gets the same session rather than a camera pointed at a whiteboard.

When the reveal is thin. Two possibilities, and they need different answers. If the cards are sparse, the prompt was usually too broad, so ask a narrower one straight away: name one thing that took longer than it should have. If the cards are polite, the anonymity is not believed, and encouragement will not fix that in the moment. Move to a numeric vote, and rebuild the trust across the next few sessions instead.

When one item dominates. Somebody names the big thing and the room spends the hour there. That is often correct. Say so out loud, park the other items in writing, and promise them the front of the next session.

When the same item comes back. Third month running, same complaint, no progress. Stop collecting and change the question to why the last two attempts failed. A repeated item is rarely a retro problem. It is usually a decision nobody in the room has the authority to make, and naming that is the useful output.

When to skip it. During a live incident, run the incident. Immediately after redundancies, an activity in front of the room reads as management deciding how people should feel. And a two-week sprint where nothing changed does not need an hour. Ask the one prompt, read the cards, go.

When it lands flat. Do not apologise for the session. Make the first specific observation yourself, about a decision you made and got wrong. A facilitator who is visibly wrong has shown the room that being wrong here costs nothing.

FAQ

Common questions

What are the best retrospective games for a team meeting?

Start, Stop, Continue for producing actions, Sailboat when you need risk on the table as well as complaints, and Mad, Sad, Glad after a hard quarter. Run any of them with a silent writing phase first and a simultaneous reveal. The format matters far less than whether the first answer anchors the rest.

How long should a sprint retrospective take?

About an hour for a two-week sprint. The Scrum Guide caps it at three hours for a one-month sprint and says shorter sprints usually get shorter events. Spend ten minutes on the record, five writing in silence, most of the middle on discussion, and the last ten choosing one change and naming who owns it.

Why do retrospectives feel like a waste of time?

Usually because they run on memory and produce no decision. Matthies and colleagues note that the retrospective activities proposed so far often rely on the perceptions of team members rather than on project data. Bring the tickets, the incidents and the dates, then finish by choosing one improvement with an owner and a check date.

What is a silent retrospective?

Everyone writes their answers privately, nothing is shown until all of them are in, and then every answer appears at once. That removes the wait for a turn and the anchor from whoever went first, and it takes your name off the card while you are writing.

How do you run a retrospective with a remote or hybrid team?

Put every person on their own device, including the ones sitting in the same room. Write silently, reveal together, discuss afterwards. Simultaneous private writing survives a remote session better than most formats, and it stops the conference room from becoming the meeting while the remote half watches.

Should retrospectives be anonymous?

The first pass should be. Anonymity is what makes it cheaper to write the thing somebody would not say with their manager present. Google's postmortem guidance is blunt that where blame prevails, people will not raise issues for fear of punishment. Once the items are on the board, the discussion can be perfectly open.

What is the retrospective prime directive?

Norm Kerth's line, read at the start of a session. Regardless of what we discover, everyone did the best job they could given what they knew at the time, their skills and abilities, the resources available and the situation at hand. It is a correction for outcome bias: Lea and colleagues showed people judge identical actions more harshly when the ending was worse.

Run this with your own team

Start free, no credit card. Your audience joins from their phones with a code — nothing to install.