Outcomes Over Output In Practice
How to move from a feature roadmap to outcome commitments without losing the trust of stakeholders who want dates. Writing measurable outcomes, choosing leading indicators, and the ways OKRs fail.
There is a particular kind of quiet that follows a successful delivery which changed nothing. The team shipped everything on the roadmap. The dates held. The quality was fine. And six months later, the numbers the investment was justified by look exactly as they did before. Nobody is at fault in any way that can be pointed at, which is precisely what makes the situation so difficult to address. Every individual decision was defensible and the aggregate was waste.
This happens because the organisation committed to output — a list of things to build — rather than to outcomes, a change in what people do or what the business achieves. Output commitments are attractive because they are easy to write, easy to track and easy to prove you met. Outcome commitments are harder in every one of those dimensions, which is why most organisations claim to have made the shift while their planning documents remain lists of features with dates beside them.
"Outcomes over output" has been repeated enough to have lost its edge, and the phrase is not the useful part. The mechanics are: how to write an outcome a team can steer by, what to measure when the real effect takes a year to appear, how to keep outcome commitments honest, and what to say to a stakeholder who reasonably wants to know what they are getting and when.
Why "we delivered everything and nothing changed" happens
Three mechanisms, each independently sufficient.
The feature was a hypothesis and nobody treated it as one. Someone believed that if the product did X, customers would do Y. That belief was reasonable and it was never tested. By the time X was built, the belief had been converted through a series of planning artefacts into a requirement, at which point it was no longer something that could be wrong — only something that could be late.
The list was assembled from requests, not from a problem. Roadmaps that originate as an aggregation of stakeholder asks have no theory behind them. They are internally consistent — each item has a sponsor who wanted it — and collectively incoherent, because no single customer problem is fully solved by any of them.
Success was defined as completion. If the reward structure asks "did you deliver what you said", then the rational strategy is to promise deliverable things and deliver them. Nobody in that system is incentivised to stop halfway through a feature that has stopped looking worthwhile, and stopping is the highest-return decision available in product work.
The third one is the deepest, because it is structural rather than intellectual. Teams will optimise for what is inspected. If the quarterly review inspects a list, you will get a list, no matter how many times the word outcome appears in the strategy deck.
What makes an outcome measurable
Most things written on a whiteboard under the heading "outcomes" are not outcomes. They are either outputs in disguise, or aspirations with no attached observation.
A usable outcome has three properties. It describes a change in behaviour or in a business result, not a thing that exists. It names a direction and a magnitude. And it is observable within a timeframe short enough to steer by — which usually means within the quarter, not at the end of the year.
| Written as | Why it fails | Rewritten |
|---|---|---|
| Launch the new onboarding flow | An output; completion is guaranteed and proves nothing | Increase the proportion of new accounts reaching first successful transaction within a week |
| Improve the customer experience | No direction, no magnitude, no observation | Reduce repeat contacts for the same issue within thirty days |
| Increase engagement | Undefined behaviour; invites metric gaming | Increase the share of accounts using the reporting feature at least weekly |
| Become the market leader in onboarding | Correct ambition, wrong altitude for a team's quarter | Cut median time from signup to first value, with a target range |
| Reduce technical debt | An activity, not a result | Reduce median lead time for changes in the billing subsystem |
Note what the rewrites have in common: each names a specific population, a specific behaviour, and a direction. A team can generate ten candidate solutions against any of them and choose between them on evidence. Against the left column, a team can only execute.
Magnitude is where most teams get stuck, because they do not know what a realistic improvement looks like and fear committing to a number they will miss. Two things help. State a range rather than a point, with the low end a result worth having and the high end an ambitious one. And be explicit that the first quarter of any new measure is calibration — you are establishing what movement is even possible. A target set with no baseline is a guess, and treating a guess as a commitment teaches everyone to set soft targets forever.
Leading and lagging indicators
The uncomfortable arithmetic of outcome work is that the things the business cares about — revenue, retention, market share, cost to serve — move slowly and are influenced by many factors outside the team's control. A team that commits only to those has no steering signal during the quarter and no defensible attribution at the end of it.
The resolution is a chain. The lagging indicator is what the organisation actually wants. The leading indicator is something observable within weeks that you have reason to believe precedes it. The team commits to the leading indicator and the organisation holds the lagging one.
Take retention. It is lagging by definition — you only know a customer stayed after the period in which they could have left. But the behaviours that precede it are observable much sooner: whether a new account reached genuine value in the first fortnight, whether usage broadened beyond one person, whether the customer built the product into a recurring process. Those are steerable this month.
The discipline that makes this honest is stating the causal belief out loud. "We believe accounts that complete a second successful workflow in the first month retain materially better, and so we are working to increase that proportion." Written that way, the indicator is falsifiable: if it moves and retention does not, your model of the business was wrong and you have learned something valuable. If nobody writes the belief down, the leading indicator quietly becomes a target in its own right, and the team optimises a proxy that has come loose from the thing it was proxying for.
How OKRs fail
Objectives and key results are the most widely adopted vehicle for outcome commitments, and most implementations produce more heat than movement. The failures are consistent enough to be listed.
Output dressed as key results. "Key result: ship the mobile app." This is the dominant failure and the hardest to police, because teams under pressure to commit to something achievable will naturally drift toward things they control. A key result that is fully within the team's control is almost certainly an output.
Cascading until nothing means anything. The executive objective is decomposed into departmental objectives, then team objectives, then individual objectives, with each layer adding interpretation. By the fourth level, teams are writing key results that satisfy a template and have no traceable relationship to anything a customer would notice. Cascade the direction; let teams write their own key results against it.
Too many. An organisation with five objectives and four key results each has twenty commitments, which is not a set of priorities but a comprehensive description of everything currently underway. The value of the instrument is entirely in what it excludes.
Tied to compensation. The moment an OKR affects a bonus, the negotiation about the target becomes more important than the work, and every target will be set at a level that is comfortably achievable. Ambition and appraisal cannot occupy the same instrument.
Quarterly ritual, no mid-quarter response. OKRs are set in week one, ignored until week eleven, and graded in week twelve. The grading conversation is where the least value is. The value is in week five, when the indicator is not moving and the team changes approach in response. If there is no mechanism for that, the framework is an elaborate reporting overhead.
No baseline. Committing to "increase X by twenty percent" without knowing the current value or its natural variance means you cannot tell success from noise. Measure first, commit second.
Negotiating with stakeholders who want a date and a scope
The honest starting point is that the stakeholder's request is usually legitimate. A campaign has a booking deadline; a partner integration has a contractual date; a regulatory change has a commencement date. Treating every request for certainty as an agile maturity problem is both wrong and a fast way to lose the argument. What works is to separate the types of commitment rather than pretend only one exists.
Where a date is genuinely fixed, commit to the date and negotiate the scope. This is not a failure of outcome thinking; it is a constraint, and constraints are normal. What you refuse is fixing date, scope and quality simultaneously, because that combination is arithmetically unavailable and the missing variable gets paid out of quality, where nobody has to sign for it.
Where the date is soft, offer predictability rather than precision. Forecast from your own flow data as a range with a confidence level, not as a single date. A stakeholder told "we are eighty-five percent confident of landing between the second and fourth week of March" has more usable information than one told "the fourteenth of March", and, importantly, will not be surprised in March.
Where the request is really about anxiety, sell visibility. Many demands for detailed plans are requests for reassurance that something is happening. Reassurance delivered continuously — working software every two weeks, a live view of progress toward the outcome — costs far less than reassurance delivered as documentation, and it is more convincing.
Commit to the problem, timebox the investment. The most robust form of outcome commitment for genuinely uncertain work is: we will spend this team for this quarter attempting to move this indicator, we will show you evidence every fortnight, and if it is not moving by the midpoint we will change approach or stop. That is a real commitment with real accountability, and it is frequently more attractive to a senior stakeholder than a feature list, because it caps their downside.
What an outcome roadmap looks like
A roadmap expressed as outcomes is not a list of features with the dates removed. It has a different structure, and the structure is what carries the meaning.
It is organised in horizons rather than months. The near horizon — this quarter — carries committed outcomes with named indicators, baselines and target ranges, plus the solutions currently being tried, explicitly marked as bets rather than commitments. The middle horizon carries outcomes only, with no solutions attached, because deciding now how you will achieve something nine months away discards the information you will have by then. The far horizon carries the problems you believe matter, without dates.
For each committed outcome it states the customer or business problem, the indicator, the baseline, the target range, the causal belief connecting indicator to business result, and the review cadence. That is roughly a page per outcome, and four to six of them describe strategy more fully than a hundred-row feature spreadsheet.
What disappears is the implied promise that everything listed will be built. What appears is an explicit statement of what the organisation is trying to change and how it will know. Stakeholders lose the comfort of a specific feature in a specific month and gain the ability to see mid-quarter whether the investment is working — information the feature roadmap never provided, since a feature roadmap can be perfectly green right up to the moment it is revealed to have accomplished nothing.
One transitional format is worth knowing. Keep the feature roadmap for a cycle, but add two columns: what outcome each item is meant to move, and what evidence exists that it will. Fill them in honestly. The rows that cannot be filled in are the conversation, and it is more productive than any attempt to argue the format away.
What to do on Monday
Take your current roadmap and, beside each item, write the change in customer or business behaviour it is meant to produce, and how you would know. Do it without softening. Some proportion of the list will have no answer, and that is the finding — not a reason for embarrassment, just the current state made visible.
Then pick the single largest item with no answer and go and ask its sponsor what they expect to be different once it exists. Do not challenge, just ask, and write down what they say. If they have a clear answer, you have found your first outcome and a willing partner. If they do not, you have just saved a quarter of engineering capacity, and you found out at the cheapest possible moment.