The Retrospective Problem
Retrospectives decay into complaint or silence for structural reasons, not cultural ones. How to retro on the system rather than the team, what to do about actions nobody owns, and when to stop running them.
Every retrospective eventually arrives at one of two terminal states. In the first, the same three complaints are raised every fortnight — the environments, the dependency on the platform team, the requirements arriving late — and everyone nods, and nothing happens, and the meeting becomes a scheduled venting slot that leaves people flatter than when they walked in. In the second, nobody says anything much. The facilitator works through a format, the team offers a few procedural observations, two mild actions are recorded, and everyone returns to work relieved that it was short.
Both states are usually diagnosed as cultural problems. The team is negative; the team is disengaged; the team needs better facilitation. This diagnosis is wrong often enough to be worth treating as the exception rather than the rule. Both states are rational adaptations to the same underlying fact: the retrospective has no reliable route to changing anything that matters.
If you raise a real problem six times and nothing changes, the sensible response is to stop raising it. If the only problems that can be fixed are small ones, the sensible response is to discuss only small problems. What looks like cynicism or apathy is usually accurate learning about how much power the meeting has, and no amount of format innovation will correct a power problem. The fix lives in the structure around the retrospective, not inside it.
Why the decay is rational
The retrospective is the adaptation step of an empirical loop. Adaptation requires the authority to change something, and the impediments that most constrain a delivery team are overwhelmingly outside its boundary: shared environment contention, a queue at an architecture review, a release window controlled elsewhere, unclear priorities originating two levels up, a hiring freeze, a dependency on a team with its own backlog.
The team can see these clearly. It can do nothing about any of them. So the retrospective faces a choice each fortnight between discussing things it cannot change, which is demoralising, and discussing things it can change, which are by construction the least significant things available. Complaint is the first path; procedural triviality is the second. Silence is what you get when the team has tried both.
This is the same failure identified in empiricism as the engine: inspection without the authority to adapt is surveillance. The retrospective is where that gap becomes personal, because the people doing the inspecting are also the people living with the unresolved consequence.
Psychological safety is a precondition, not an output
Amy Edmondson's research established the term, and it has since been diluted into a synonym for being nice to each other. It is not that. Psychological safety is the shared belief that raising a problem, admitting an error or disagreeing with someone senior will not be held against you. It is a belief about consequences, and beliefs about consequences are formed by observing consequences.
Which means you cannot install it with a workshop. You install it by what happens the first time somebody says something inconvenient in front of a manager. If that person is contradicted, corrected, or later told they were "not being constructive", the team will have all the data it needs, and the next four retrospectives will be pleasant and useless.
Three structural factors dominate:
Who is in the room. A line manager who conducts performance reviews for the attendees changes what gets said, regardless of how approachable they are. This is not a comment on the individual; it is a fact about the incentive structure. If system-level problems need a manager present, consider having them attend for a specific agenda item and leave, rather than sitting through the whole session.
What happens to what is said. If retrospective content leaks into performance conversations even once, the meeting is finished as a source of truth. Make the confidentiality rule explicit, and then actually keep it, including when the content is embarrassing to someone senior.
Whether disagreement has ever changed an outcome. Safety is built by visible instances of someone pushing back and being right, and the organisation adjusting. One of those is worth a year of workshops.
Retro on the system, not on the team
Most retrospectives are conducted as a discussion about the team's behaviour during the last two weeks. This is the wrong unit of analysis, because delivery performance is dominated by the system the team works in rather than by the effort it applies. A team that is waiting nine days for a review is not going to improve its outcome through better collaboration.
Shift the object of inspection from people to work. Bring the actual data and look at it together.
Ageing work in progress. Which items sat longest, and in which state? This is the single most productive artefact to bring, because it converts "things felt slow" into "these four items spent eleven days waiting for test". The argument moves from perception to evidence.
The cycle time distribution and its outliers. Not the average — the long tail. Pick the three slowest items of the last month and reconstruct what happened to each. The pattern is usually identical across all three and is usually structural.
Blocked and waiting states. Count the transitions into "blocked" and what caused each. If eight of eleven blocks trace to one external dependency, you have located the problem and you no longer need anyone's opinion about it.
Unplanned work. How much arrived, from where, and who absorbed it. Teams routinely underestimate this and then blame themselves for not delivering the plan.
This is where flow metrics earn their place. They are not primarily a reporting instrument; they are the material that makes a retrospective about something other than feelings.
The action items nobody owns
The most common structural defect in retrospective practice is generating six to ten actions, recording them in a document that lives outside the working backlog, assigning most of them to "the team", and reviewing them once at the start of the next retrospective, where everyone agrees they are still relevant.
Four rules fix the great majority of this.
One or two actions, never more. A team that generates eight actions will complete none. A team that generates one will usually complete it. The discipline of choosing forces a judgement about what actually matters, which is the valuable part of the exercise.
A named individual, not the team. Not because that person does all the work, but because diffuse ownership is no ownership. Someone has to be answerable for whether it happened.
A date, and a place in the real backlog. If the action is worth doing it is worth competing for capacity alongside everything else. Actions kept in a separate document are actions nobody has decided to fund.
Reviewed first, not last. Open each retrospective by checking the previous action. If it did not happen, that is the most important discussion available to the team and it should displace the rest of the agenda. Either the action was not really important, or something prevented it — and the something is itself the impediment.
Route the impediments you cannot fix
The part that is almost always missing is a route for problems the team cannot solve. Without one, those problems accumulate in the retrospective until the team stops raising them. With one, they become someone else's backlog item, which is where they belong.
| Impediment type | Who can actually resolve it | Where it should go |
|---|---|---|
| Team practice, tooling within the team, working agreements | The team | Sprint backlog, this sprint |
| Dependency on one other team, environment contention | A named manager or the other team's lead | Direct request with a date, tracked |
| Shared platform capacity, review bottlenecks, release process | Engineering leadership | Standing impediment forum, reviewed fortnightly |
| Funding model, org structure, incentive design | Executive sponsor | Transformation backlog, quarterly |
The second and third rows are where the Scrum Master role either proves its value or reveals itself as administration. Logging an impediment is not resolving it. The role's actual measure is how many structural constraints were removed in the last quarter, and by whom.
Where the fourth row dominates, the retrospective is telling you something about the organisation rather than about the team, and that signal is worth reading carefully rather than filtering out — see resistance is information.
When to stop running them
This will be unpopular with people whose job title includes the word agile, but it follows directly from the framework's own logic: a practice that is not producing adaptation is not serving empiricism, and continuing it out of compliance is exactly the behaviour the approach was meant to eliminate.
If six consecutive retrospectives have produced no change that anyone outside the meeting could detect, stop running them on a fixed cadence. Say why, publicly, and replace them with something that does produce change:
Event-driven reviews. Run a proper review after each significant event — a release that went badly, an incident, a feature reaching customers, a dependency that blocked you for a week. These have concrete material, genuine stakes, and a natural audience, and they tend to produce real actions because the cost of the event is fresh.
A monthly flow review with authority present. Half an hour with the ageing chart, the cycle time distribution and one decision-maker who can commit to a change. Fewer, better conversations beat fortnightly ritual.
A deep quarterly session. Two hours, prepared properly, with data gathered in advance and a sponsor who has agreed to act on one output. This is where structural change actually gets decided.
Continuous improvement in the moment. The best-performing teams adjust constantly — someone notices the review queue is backing up and the team changes its working agreement that afternoon. If your team does this, its retrospective may genuinely be redundant, and saying so is a sign of maturity rather than backsliding.
The distinction that matters is between stopping because the practice is uncomfortable and stopping because the practice is inert. The first is avoidance. The second is judgement. The test is whether you replace it with something that produces change, and whether you are prepared to resume when circumstances shift — after a reorganisation, a new joiner cohort, a painful release.
What to do on Monday
Go back through the last six retrospectives and list every action recorded. Mark each one: done, quietly dropped, or still open. Take the percentage that were done and put it on the wall.
If the number is low, do not attempt to improve compliance. Instead, take every item from those six sessions and sort it into the four rows of the table above. The distribution tells you what kind of problem you have. A team whose impediments cluster in row one has a practice problem it can solve this week. A team whose impediments cluster in rows three and four has an organisational problem, and the correct next action is to get thirty minutes with the person who owns those constraints — with the list in hand — rather than to run another retrospective.
Then cap actions at one per session and review it first thing next time. Nothing else changes. Four sessions later, compare the completion rate. In most teams it goes from near zero to near one, which is a larger improvement in adaptation than any format change will ever deliver.