# Reading the numbers
*Decide in advance what you'd act on*
---
The monthly review is the same meeting in every business I've run: revenue against last month, cost against budget, this line ahead and that one behind, and a story attached to every movement. Most of the stories are noise. I don't mean the data is bad, or that the people explaining it are careless. I mean the movement itself usually carries no information about the machine underneath, and we explain it anyway, because we're meaning-making machines and the meeting was booked.
Most of what work consists of is people, some data, and a period of time, put together in a room, and it would be nice to think each of the three was chosen deliberately. Usually the meeting is whatever length the calendar offered, the attendee list just grew, and the data is whatever the pack happens to contain. There's a whole literature that says decline the meeting, ask for the agenda, and maybe that works somewhere. In practice, in a business where relationships matter and things need to get done, when someone calls you into a meeting about something that matters to them, you go. Being that dogmatic about it just doesn't really work. The meeting isn't the problem. The problem is what we do with the numbers once we're in the room, because having assembled the people and the data, somebody has to say what it all means.
## Ask what it's measured against
Every variance is a comparison, and the thing being compared against is usually chosen by default. Month-on-month assumes all months are equal, which they aren't: they differ by working days, seasonality, period-end effects. Year-on-year assumes last year was representative, and the irony is that last year was no more normal than this one. It had its own one-offs and distortions, which everybody has since forgotten, and so did last month. Comparing against plan is at least a comparison against something somebody chose on purpose.
So a variance often tells you more about the choice of baseline than about the month. And the aggregate underneath can be hiding the real story entirely: I've seen a flat top line that segmentation showed was really two stories, a growing core and one region being squeezed by its funder, cancelling each other out in the total.
Manufacturing worked this out a long time ago. Any complex system produces variation in its outputs, and rather than explain each one, the discipline built a language for it: tolerances, thresholds, the difference between a reading inside normal behaviour and one outside it. Finance and knowledge work face much the same dynamics, ours arising from customers and processes rather than machines, and we've never really adopted the language. There's no agreed answer to when a movement deserves escalation, when it deserves a conversation, and when it's simply fine. So the movements get discussed every month, whichever way they went, because by definition there's always a variance against something.
## What reacting to noise costs
Getting this wrong costs more than the meeting.
The first cost is the industry that grows up around it. There's always an explanation available, so a finance team asked for one every month will supply one every month, and after a while a serious amount of capable people's time goes into reconciling movements that nobody reads, nobody tests, and that change nothing about what anyone does next. Some of the reconciliation is owed, to be fair - your bank and your auditors need theirs - but that's a different document from the one management steers by.
The second cost is worse, and it belongs to being a particular size. If your calibration is off and you act on a wiggle, you set machinery going: initiatives get briefed, owners named, teams re-pointed. In a small business that's recoverable, because you can stop it quickly and aim somewhere else next month. Once you're mid-sized, the briefing takes a month or two to reach the people who actually have to do it, and by then you've held another monthly review and started another set of things. The business is still absorbing last month's instruction while receiving this month's, the two aren't obviously related, nobody further down the chain can tell you which of them still matters, and the whole place works flat out without moving anything forward.
## A chart instead of a table
The fix costs almost nothing. A twelve-month trend with a band around the average tells you more in five seconds than any snapshot variance table. A point inside the band gets noted. A point outside it, or three months trending one way, gets investigated. The question changes from "why is this month down on last?" to "is anything unusual happening?", and most months the answer is no, and the meeting moves on.
I personally like these charts, and I'd concede that some of the formal machinery feels too mathematical for a monthly pack. The intuition is the powerful part: look at the trend rather than the point, and agree in advance what would be worth acting on.
> [!note]- For those so inclined
> The formal version is statistical process control, and the chart is a process behaviour chart: limits calculated from the mean and the moving range, rules for runs and shifts. Worth knowing it exists; probably more than most packs need.
All of it depends on agreeing the threshold before the number lands. One you reach for afterwards just backs up whatever you already wanted to do, in either direction.
## Which number are you asking for?
Estimates have the same problem more quietly, and it costs you more.
Your team says the project will take two months, and they're being honest: two months is their middle estimate, the point where they'd bet even odds. Could it land in one month? Almost impossible, because there's a floor on how fast work can go. Could it take five? Easily - a dependency slips, a requirement changes, a key person gets pulled onto something urgent. There's no ceiling on delay. That asymmetry means the average outcome sits above the honest middle, and the gap is bigger than people expect: the early finishes save you weeks, the late ones cost you months. As a rule of thumb, plan against the team's number times 1.2 for routine work, and times 1.5 where the work is genuinely new.
Business cases and delivery plans land on my desk the same way: they always show things delivered on time and strong returns. So my baseline is that things get delayed and plans under-deliver, and I nudge from there on the credibility of the paper. The better thought out it is, the smaller the adjustment.
> [!note]- For those so inclined
> Task durations tend to follow a lognormal distribution, bounded on the left, unbounded on the right, so the mean sits above the median, and the gap widens with uncertainty. That's where the 1.2 and 1.5 come from.
The practical move is to hold two numbers at once, doing two different jobs. The team plans against the honest middle, because that's the number everyone coordinates around, and if you gave them the padded one they'd plan to it and take it. The board gets the mean, because across a year of projects that's the number that makes the budget land roughly on plan. The project that taught me this wasn't a failure: the board heard two months, it took three and a half, nobody had done anything wrong, and everybody was annoyed. Two months was the right estimate for the team and three was the closer number for the board. The problem was one number doing both jobs.
## Ten interviews aren't a percentage
Small samples are the third version of the same problem. A product team runs ten user interviews, seven prefer option A, and the readout says "overwhelming favourite". It feels immediately actionable and it isn't, because the sample you need grows fast as the effect you're hunting shrinks: halve the difference and you need four times the people. Small differences, the kind most product changes actually produce, need serious samples to confirm. Only dramatic ones show up in small groups.
Show a redesign to thirty users and the old version to thirty others, and twenty of the thirty prefer the new one: 67%, and the slide says strong preference. Thirty per group sounds like a small version of a proper test. It can only detect a gap of thirty-odd points, and your redesign almost certainly isn't that much better than what it replaced. A 67% reading on thirty people sits inside the range you'd expect by chance if the two designs were identical. Which is not the same as saying the redesign isn't better. It may well be. The reading just can't tell you.
> [!note]- For those so inclined
> A workable rule of thumb: the sample you need per group is about 16 × p(1−p) / d², where p is the baseline rate, d the absolute difference you want to detect, and the constant bakes in the usual confidence and power. The d² is what ruins your plans. At a 4% conversion baseline, detecting a lift to 4.8% needs nearly ten thousand visitors per variant; run it with twelve hundred and you have roughly a 15% chance of seeing the effect even when it's real. You'd conclude the change didn't work, kill it, and move on.
What small samples are good for is words, not percentages. Ten interviews might surface the phrase people actually use, or the same friction point mentioned seven times unprompted, or an objection you hadn't thought of, and those are as informative at five people as at five hundred. Many decisions can't wait for statistical significance, and some never could: you'll never get three thousand enterprise buyers into a pricing study. Do the qualitative work well, make the decision, and be honest about what you don't know.
## What a business case is really saying
A business case is where all of this lands at once. Six months of development at £500k, fifty sales at £30k, a 3x payback. Every figure in it is a single point standing in for a range, and the ranges are lopsided in the same direction as the estimates above: costs have a floor and no ceiling, sales ramp late and start discounted. Rerun that case with nine months of development and thirty early deals at £25k and the 3x has evaporated to breakeven. The second set of assumptions is no more true than the first. The point is how much room the headline had to fall.
So the useful version of the slide says "between 1x and 3x, depending mainly on how long development actually takes", and the conversation changes from approve-or-reject to a constructive debate about what needs to be true. The front page of the last three-year plan I signed off carried a short list of exactly that - no unplanned exits, delivery on time, no adverse change in policy - and the sign-off conversation was about the list, not the spreadsheet.
Decide in advance what would make you act, whether it's a band on a chart, a multiplier on an estimate, or a list on the front of a plan. Then let most months be what most months are: no root cause to find, nothing to action, nobody to blame. That's okay.
---