Decision Latency Is A Delivery Metric
The elapsed time between a question being raised and an answer being given is a queue, and usually the largest one nobody measures. How to make it visible, and why a leader's calendar constrains the whole organisation's throughput.
Most delivery organisations have spent a decade getting good at measuring the work. Cycle time, throughput, deployment frequency, lead time for change, flow efficiency — the instrumentation is mature, the dashboards exist, and the teams have learned to read them. What almost none of them measure is the thing that sits in the middle of every one of those flows, consuming more elapsed calendar than any build step, and generating no telemetry at all.
That thing is a decision. Somebody raised a question, and somebody else had to answer it, and between those two events the work stopped.
Decision latency is the elapsed time between a decision being needed and an answer being given and communicated. It is a queue in exactly the sense that why busy teams are slow teams describes: work arrives, waits for a scarce server, and departs. The scarce server in this case is a person, usually a senior one, usually running at very high utilisation, with a weekly arrival pattern dictated by a calendar designed for something other than throughput. Every property of queueing theory applies. High utilisation produces non-linear wait. Batching arrivals into a weekly forum produces predictable multi-day delay. Variability in service time produces a long right tail.
The reason this is uncomfortable is that the server is you. Team-level flow problems can be delegated downward and framed as a delivery maturity issue. Decision latency cannot. It is a direct, measurable property of how senior people arrange their attention, and it is the one queue in the system that no team can fix on its own behalf.
Why nobody sees it
The mechanism of invisibility is almost banal: nobody timestamps the moment a question was raised.
Work items have creation dates. Tickets have state transitions. Pull requests have opened-at fields. But a decision is typically raised in a conversation, a Slack thread, a comment on a document, or an agenda item proposed for a meeting three weeks out. There is no record of when it became blocking. By the time it is formally captured — as a risk in a steering pack, or an agenda line, or an escalation — days or weeks of waiting have already elapsed and have been silently absorbed into whatever ticket was blocked.
So the delay shows up in the data, but not as a decision. It shows up as a long cycle time on a story, an item sitting in an ambiguous "blocked" column, a sprint that failed to complete, or a quarter that came in under plan. The organisation then performs a root cause analysis that concludes the estimates were optimistic, the requirements were unclear, or the team lacked focus. Every one of those conclusions is available, defensible, and points away from the calendar of the person running the review.
There is a second reason it stays hidden: the cost is borne entirely downstream. The person who takes four weeks to decide experiences four weeks of ordinary busy work. The teams waiting experience context-switching, speculative work that may be discarded, and quiet erosion of belief that anything they raise will be answered. Cause and pain are separated by enough organisational distance that the connection is never made in the room where it could be fixed.
What a decision costs while it waits
A decision that has not been made is not neutral. It is actively generating cost, in four distinct ways, and the costs compound rather than add.
The blocked work idles or goes speculative. The team either stops — in which case you are paying for capacity that produces nothing — or proceeds on an assumption. Proceeding on an assumption is usually the responsible-looking choice and is frequently the more expensive one, because if the decision lands the other way, everything built on the assumption is rework. The longer the wait, the more expensive the rework, which means the cost of a late decision rises with the square of something rather than linearly with time.
Work in progress rises. A team that cannot finish item A does not stand idle; it starts item B. This is the mechanism by which decision latency silently degrades every other flow metric in the system. The blocked item stays in flight, the new item joins it, and Little's Law does the rest: work in progress goes up, throughput does not, so cycle time rises for everything, including the items that had nothing to do with the original question.
Context decays. The people who framed the question lose the detail. When the answer arrives six weeks later, re-acquiring the context costs real hours, and the answer frequently no longer fits the situation that produced it, which generates a second question.
The option value of the decision decays. Cost of delay is not evenly distributed across time. Many decisions are cheap if made early and expensive if made late, because the set of reversible choices shrinks as commitments accumulate around the gap. A decision deferred is not a decision preserved; it is a decision made by default, in whatever direction the existing momentum was already travelling.
The fourth point deserves emphasis because it inverts the intuition that deferring is prudent. Deferring is prudent when the deferral buys information. It is expensive when it buys only calendar. Most executive-level deferrals buy calendar, and the distinction is easy to test: ask what specific new information will exist at the later date that does not exist now. If the answer is vague, you are not deferring a decision, you are queueing it.
How to measure it this week
You do not need a tool, a workshop or a maturity model. You need two timestamps and the discipline to record the first one.
Start a single list — a spreadsheet is entirely adequate — with five columns: the question, the date it was first raised, who needs to answer it, the date the answer was given and communicated, and what was blocked in the meantime. Populate it going forward only; a historical reconstruction costs weeks of argument about dates and teaches you nothing.
The rule that makes this work is the definition of "raised": the date a person with the problem first asked someone who could answer it. Not the date it hit an agenda. Not the date it was escalated. The forum date is the outcome you are trying to measure, so it cannot also be the start.
After four weeks you will have enough data to produce the only chart that matters here: a distribution of decision latency, not an average. Report the eighty-fifth percentile alongside the median, for exactly the reason that flow metrics argues for percentiles over means. The median tells you what a typical decision costs. The tail tells you what your organisation actually plans against, because teams do not plan against the typical case — they plan against the case that burned them last time, and they add buffer accordingly. That buffer is real money and it is invisible in every budget you have.
| What to record | Why it matters |
|---|---|
| Date raised | The only number nobody currently captures; without it there is no metric |
| Date answered and communicated | Answered privately and not communicated is not answered |
| Decision owner | Reveals whether one person is the server for most of the queue |
| Work blocked | Converts latency into cost of delay you can put a number against |
| Route taken | Shows which decisions needed a forum and which merely found one |
The last row is where the surprises are. Categorise each decision after the fact as one of three types: reversible and low-stakes, irreversible and high-stakes, or genuinely uncertain in the Cynefin sense of requiring a probe before an answer is possible. Then look at how long each type waited. In most organisations the reversible, low-stakes decisions wait roughly as long as the irreversible ones, because they travel the same governance route. That single finding usually justifies the whole exercise.
The calendar as a constraint
Here is the version of this argument that is hardest to hear and most reliably true.
In a great many organisations, the binding constraint on delivery throughput is not engineering capacity, testing capacity, or the platform. It is the availability of about eight people's attention. Every significant decision routes through one of them. Their calendars are booked at something close to full utilisation, three weeks out. And because a fully utilised server produces asymptotic wait times, the practical effect is that every question entering that queue waits between one and four weeks regardless of its urgency, its size, or its value.
Theory of Constraints is unambiguous about what to do with a bottleneck: you protect it, you do not load it with work that could be done elsewhere, and you subordinate the rest of the system to it. Applied honestly to an executive calendar, that means the majority of what currently occupies it should not be there. Status updates consume the constraint and produce no decision. Information-sharing meetings consume the constraint to transmit what a document would transmit better. Approval of work that someone closer to the information could approve consumes the constraint to add a rubber stamp and three days of latency.
The arithmetic of a weekly forum is worth stating plainly, because it is usually treated as a scheduling detail rather than a design decision. If decisions can only be made at a weekly meeting, and questions arrive uniformly through the week, then the average wait before the question is even considered is half a week — around three and a half days — before you add the time to get onto the agenda, the probability of being deferred to the following week, and the time for the outcome to be written down and communicated. A single weekly gate is rarely under four days of average latency and frequently closer to eight. Two such gates in sequence, which is common, and you have built a fortnight of pure wait into every decision the organisation makes. Nobody chose that. It is an emergent property of two recurring calendar entries. The full treatment of this is in the operating cadence.
Reducing latency without reducing quality
The objection to all of this is that fast decisions are bad decisions. Sometimes true, and the distinction between the cases is the whole craft.
Separate the reversible from the irreversible. Most decisions are reversible at modest cost. Those should be made by whoever has the information, immediately, with a default of yes and an obligation to write down what was decided. Reserve deliberation for the small number of decisions that genuinely cannot be undone.
Set a decision service level and publish it. Something like: any question raised to this leadership group receives an answer, a named owner, or an explicit date by which it will be answered, within five working days. Note the three acceptable outputs. "Not yet, and here is when" is a legitimate answer and resolves the latency problem almost as well as a decision, because it lets the team plan.
Default to a decision at the point of information. A decision made by the team with full context in an hour is usually better than the same decision made in a forum with a summary slide in four weeks. Pushing authority down is not a cultural gesture; it is moving the server closer to the arrivals.
Make the absence of a decision visible. Put an ageing counter on open decisions and show it in the same review where delivery progress is shown. An item that has been waiting thirty-one days for an answer should be as conspicuous as a red delivery status, because it is the same thing.
Answer in writing, once. A decision communicated verbally to three people is not communicated. Write the decision, the reasoning and the date somewhere durable — architecture decision records are the engineering version of this habit, and it generalises upward unchanged.
What to do on Monday
Open a spreadsheet and write down every question currently waiting on you. Not projects — questions. For each one, find the date it was first raised, by asking the person who raised it rather than by looking in your inbox. You will find at least one that has been open longer than you would have guessed, and probably one that the raiser has quietly given up on and worked around.
Then pick the oldest, answer it today however imperfectly, and say out loud in your next leadership meeting how long it had been waiting. That single act establishes decision latency as a real metric better than any dashboard, because it demonstrates the two things the organisation needs to believe: that the clock starts when the question is asked, and that the person at the top of the escalation path is inside the measurement rather than outside it.