The dashboard is unambiguous, and it is good news. You shipped the new onboarding on Tuesday, and by Friday the conversion rate is up nine percent. Somebody posts the chart in the channel with a rocket under it. The line steps up right at the deploy — exactly the shape you hope for — and for once the story writes itself: we changed the thing, the number went up, the change worked.
Then someone who has been burned before does the boring thing. They split the rate by cohort — new users on one line, returning users on the other — and the good news comes apart in their hands. New users converted worse after the deploy. Returning users converted worse too. There is no third group. Every cohort on the chart went down, and the total went up.
Nobody typed a wrong number. Both readings are true at the same time. That is the unsettling part, and it has a name.
The average is a mixture, not a measurement
A conversion rate feels like a property of the product — turn a dial, the rate responds. It isn’t. An overall rate is a weighted average of the rates inside it, and a weighted average has two inputs, not one: the rates, and the weights. The rates are how well each kind of user converts. The weights are how much of your traffic each kind of user makes up. Move either one and the headline number moves, and from the outside the two causes are indistinguishable.
Here’s the release, in the only four numbers that matter. Before Tuesday, new users converted at
4.0% and returning users at 9.0%, and new users were the bulk of the traffic — about 70% of it. The
blended rate was 0.70 × 4.0 + 0.30 × 9.0 = 5.5%. After the deploy, new users slipped to 3.4% and
returning users to 8.2% — both down. But the release also included a “welcome back” prompt that
pulled far more returning users into the flow, and the mix tilted to roughly 45% new, 55% returning.
Blend those: 0.45 × 3.4 + 0.55 × 8.2 = 6.0%.
Down in every part, up in total. The 5.5 → 6.0 climb — the nine percent everyone celebrated — is real arithmetic. It just isn’t the arithmetic anyone thought they were reading. What improved wasn’t how the product converts. It’s that a larger slice of the audience was the slice that always converted well.
What actually moved was the mix
Once you see it, the mechanism is almost rude in its simplicity. Returning users have always converted at roughly twice the rate of new ones. If you do nothing to the product at all and merely change the doorway so that more returning users walk through it, the blended rate rises. You can make the top-line number go up by making the product worse for everybody, as long as you shift enough weight toward the group that was already doing well.
That is not a freak event. It is the default behavior of any release that touches acquisition, targeting, an email, a banner, a change in what a link says — anything that alters who arrives rather than what happens once they do. Most releases touch both at once. So the top-line rate is forever being tugged by two hands, and the aggregate faithfully reports their sum while telling you nothing about which hand pulled.
A blended metric can’t distinguish a better product from a richer mix. It was never designed to. Averaging is exactly the step that throws that information away.
Why the aggregate can’t warn you
The cruelty of the paradox is that the overall number gives no sign it’s misleading you. It doesn’t flicker or star itself. A rate that rose because the product got better and a rate that rose because the mix shifted look identical at the top level — same line, same step, same rocket. The warning lives entirely in the decomposition, and the decomposition is precisely the thing a single headline metric has already collapsed by the time you read it.
Which means “conversion is up nine percent” is not a finding. It’s a question wearing the costume of an answer. Up for whom? Up because they converted better, or up because more of the good converters showed up? The blended number cannot be interrogated, because the two stories that produce it are already blended beyond separation. You can only get the answer back by refusing to average in the first place — by keeping the cohorts apart and looking at the weights on purpose.
Reading the number the way that can’t lie
This is the kind of failure code-results is built to refuse. Its whole job is tying a deploy to the business numbers that moved after it, and a paradox like this is how that job goes wrong when you’re not careful: you credit the release with a win no cohort actually had, ship more like it, and quietly erode the product while the chart keeps climbing.
So code-results doesn’t hand you the blended rate and call it attribution. It reads the metric the only way that survives a mix shift. It splits the number by cohort before it reports it, so “up nine percent” can never hide “down in every segment.” It watches the weights themselves as a metric in their own right — a change in the mix is an event, not background — so a release that moved who showed up is flagged as exactly that. And when it does blend, it can hold the mix fixed at last week’s weights and re-blend, which strips the composition effect out and leaves the honest question: with the same audience, did each cohort get better or worse? On this release, that answer is worse, plainly, and the celebration was premature.
None of that is exotic statistics. It’s the discipline of never letting an average be the last word, because an average is where the evidence goes to hide.
The total is not the team
It’s tempting to treat the top-line number as the scoreboard and everything under it as detail. The paradox is the standing reminder that it’s the other way around. The cohorts are where the product actually lives — real people, having a real experience that got better or worse. The total is just a summary of them, weighted by an audience mix you may have changed by accident and read as a triumph.
Up in total, down in every part is not a rare glitch to guard against. It’s what a blended metric does whenever the mix moves, which is often. The fix isn’t a cleverer aggregate. It’s the humility to split the number, look at the weights, and let the parts outvote the whole — because the whole, left to speak alone, will tell you that you won a race that everyone running actually lost.