Sprint Planning That Is Not Theatre
How to run planning so it produces a goal the team can act on rather than a number somebody will be held to. Covers sprint goals, capacity versus forecast, right-sizing instead of points, and the usual failure modes.
Sit in on sprint planning at a typical organisation and you will watch four hours spent producing a number that nobody will use, a list nobody will remember by Thursday, and a sprint goal written in the last three minutes to describe a selection that was made on other grounds. The meeting is long, poorly attended by the people whose decisions matter, and universally disliked. Teams tolerate it because it is in the framework.
The tolerance is misplaced, because the meeting as run is usually solving the wrong problem. Most sprint planning is organised around the question "how much can we fit", which is an allocation problem. The question the framework actually asks first is "why is this sprint valuable", which is a strategy problem. Get the second one right and the first becomes straightforward. Get the second one wrong and no amount of estimation rigour will save the sprint, because you will have a well-calibrated plan to do a collection of unrelated things.
The good news is that planning is one of the cheapest things to fix, because the change is largely to the order of the conversation rather than to anything structural. A team that agrees its goal before selecting its work will run a materially different sprint from one that does it the other way round, and the difference shows up within a fortnight.
What the session is actually for
Three outputs, in order, and the order is load-bearing.
Why this sprint is valuable. A single sprint goal, stated in a sentence, that someone outside the team could understand and that will still be true on day eight. It is a statement about the outcome, not a summary of the tickets.
What can be done. A forecast of the backlog items the team believes will serve that goal. A forecast, not a commitment — the distinction is the difference between a planning tool and a performance instrument, and it is covered properly in Scrum as written versus Scrum as practised.
How the work will get done. Enough of a technical plan that the team is not starting from zero on Monday afternoon, and enough shared understanding that the first item can begin immediately. Not a full task breakdown of every item — the later items will be re-planned anyway as the earlier ones teach you something.
Everything that does not serve one of those three outputs is either refinement that should have happened earlier or reporting that should not be happening at all.
The sprint goal is the unit of commitment
A good sprint goal gives the team a basis for making decisions without convening a meeting. That is its whole function. On day six, when someone discovers that an item is larger than expected, the goal tells them whether to push through, cut scope within the item, or drop it entirely. Without a goal, every such decision escalates, and the team's autonomy is nominal.
Good goals are outcome-shaped and singular: a new customer can complete registration without contacting support; the checkout path survives peak load without manual intervention; we know whether users will accept the shorter onboarding flow. Note that the last one commits to obtaining knowledge rather than shipping a feature, which is legitimate and underused.
Bad goals are a list wearing a sentence's clothing: complete the items in the sprint backlog; progress the payments epic and fix outstanding defects. If your goal contains the word "and" joining two unrelated outcomes, you have two goals, which means you have none — because the point of a goal is to resolve trade-offs, and a goal that permits every trade-off resolves nothing.
The hard part is that a genuine goal requires the Product Owner to have decided something. A team whose backlog is a merged stream of requests from eleven stakeholders cannot write a coherent goal because there is no coherent thesis to express. In that situation, the missing goal is a symptom; the problem is Product Owner authority.
Capacity is an input; a forecast is an output
These get conflated constantly, and the conflation is why planning meetings are long.
Capacity is arithmetic. How many people, how many days, minus leave, minus the standing commitments — support rota, interviews, that architecture forum somebody has to attend. It takes four minutes to calculate and it should be known before the meeting starts. Capacity tells you the size of the container.
A forecast is a probabilistic statement about what will fit in that container, derived from evidence about what has fitted in similar containers before. It is not a sum of estimates. It is a range with a confidence attached, and the honest version sounds like "we are reasonably confident of these five, we expect to get through these three as well, and these two are unlikely".
Teams that confuse the two end up doing arithmetic on invented quantities: converting story points to hours, hours to person-days, person-days to capacity, and arriving at a precise answer built entirely out of guesses. The precision is manufactured at every step, and the resulting number is treated as more reliable than the historical throughput data sitting unused in the same tool.
The alternative is to forecast from throughput. If the team has finished between six and eleven items in each of the last ten sprints, then the answer to "how much can we take" is bounded by that range regardless of how the items are sized in points, and the interesting planning conversation is about which items rather than how many. This is the practical application of flow metrics to a sprint cadence.
Why the estimation ritual decays
The sequence is consistent across organisations and it takes about a year.
Estimation is introduced as a shared-understanding exercise — the value is in the discussion when two people give wildly different numbers, not in the number itself. That is a genuinely good practice. Then someone notices that the numbers aggregate into velocity. Velocity gets reported upward. Reported numbers get compared between teams. Comparison creates pressure. Pressure creates inflation, and within a few quarters the same work is worth more points than it was, the discussion has stopped happening because everyone knows the "right" answer, and planning poker has become a ceremony for generating a figure whose only consumer is a slide. The full mechanism is set out in velocity is not a performance metric.
The deeper problem is that effort estimation has a poor return even when it is honest. Estimates of complex work are dominated by the variance, not the mean, and the variance is precisely what estimation cannot capture. An item estimated at three that turns out to take nine days was not mis-estimated through carelessness; it contained a discovery, and discoveries are not estimable by definition.
Right-sizing beats sizing
The more robust practice is to stop estimating how big things are and start making them a consistent size.
Right-sizing means splitting every item until it is small enough that the team is confident it can be finished — to the Definition of Done — within a few days. Items that cannot be split to that size are not ready to be planned; they are ready to be investigated, and the investigation is itself an item with a time box.
Once items are roughly uniform, counting works. Throughput becomes a meaningful measure because the units are comparable. Forecasting becomes a matter of looking at the distribution of the last ten sprints rather than summing invented quantities. And the splitting conversation, which is where the real design thinking happens, replaces the estimating conversation, which is where the arguing happens.
| Points-based planning | Right-sized counting | |
|---|---|---|
| Planning question | How big is this? | Is this small enough yet? |
| Time spent | On calibration and debate | On splitting and sequencing |
| Forecast basis | Sum of estimates against velocity | Historical throughput distribution |
| Cross-team comparison | Tempting and meaningless | Obviously meaningless |
| Effect of an outsized item | Absorbed into an average | Visible immediately as an ageing item |
| Main failure mode | Point inflation under pressure | Splitting into slices with no user value |
That last row is a real risk and worth guarding against. Splitting for size alone produces technical slices — "build the API", "build the front end" — that cannot be finished independently and reintroduce the internal waterfall. Each slice should still change something a user or an operator could observe, even if only behind a flag.
How to run the session
For a two-week sprint, ninety minutes is usually sufficient when refinement has been happening continuously. If planning routinely takes four hours, the problem is upstream: items are arriving unrefined and the team is doing analysis in a large-group setting, which is the most expensive possible venue for it.
Before. The Product Owner arrives with a proposed goal and an ordered backlog whose top items have been discussed at least once. Capacity is already calculated. Any item whose feasibility is genuinely unknown has been raised beforehand, not sprung in the room.
First twenty minutes: the goal. The Product Owner proposes; the team interrogates. Is it achievable, is it singular, what would make it fail. Expect to rewrite it. If the goal does not survive this conversation, the rest of the meeting is pointless and you should stop and fix the goal rather than proceeding to selection.
Next forty minutes: selection. Pull items from the top of the ordered backlog, testing each against the goal. Items that do not serve the goal need an explicit justification to be included — there will always be some, and that is fine, but the default should be exclusion. Stop when the forecast reaches the lower end of the historical throughput range, not the upper end.
Final thirty minutes: the plan for the first item or two. Not all of them. Enough to start, and enough to surface any technical disagreement while everyone is present. The team leaves knowing who is doing what on Monday.
End with a falsification question. What would have to be true for this sprint to fail? Write the answers down. They are your risk register for the fortnight and they take two minutes to produce.
Common failure modes
Planning as allocation. Each person is assigned their own items to fill their own capacity. The result is a sprint containing five parallel streams that finish simultaneously at the end, with all integration and review compressed into the last two days. Plan the work, not the people.
Selecting to fill capacity. If the team consistently takes on the maximum it could theoretically do, it has no slack for the unplanned work that arrives in every sprint, and it will fail its goal most fortnights. Plan to the level you can achieve reliably and use the remaining capacity for improvement work, which will otherwise never happen.
The invisible standing load. Support, incidents, code review for other teams, security queries. In most organisations this is a substantial share of real capacity and it is absent from planning, which is why planned work reliably overruns. Measure it for three sprints and subtract it honestly.
Refinement smuggled into planning. If the meeting includes the phrase "so what does this actually mean", the item was not ready. Refinement is a continuous activity, not a slot, and it should involve two or three people rather than nine.
Nobody with authority present. If the questions that arise in planning have to be answered by someone not in the room, the sprint starts with open dependencies and the goal is a hope. Either the Product Owner is available or planning is provisional.
What to do on Monday
Change one thing in the next planning session: agree the goal before opening the backlog. Do not allow any item to be discussed until a single sentence is written on the board and the team has argued about it. This inverts the usual order and it will feel awkward once.
Then cut the meeting in half. Whatever the current length, halve it and enforce it. The constraint will force the analysis upstream into refinement where it belongs, and will surface immediately which items were arriving unready.
Finally, pull the last ten sprints of completed-item counts out of your tool and write the range on the wall. Next planning, forecast against that range rather than against a sum of estimates, and compare the result after four sprints. In most teams the throughput-based forecast is no worse — frequently better — and it costs approximately nothing to produce.