Do six user interviews and something quietly dangerous happens on the way to the readout. You sit with the transcripts, the six people blur into a composite, and out the other side walks a single figure — the user — who wants roughly what the middle of your six wanted. It feels like synthesis. It is actually an average, and an average of six is a very small number of people pretending to be everyone.

The trouble isn’t that the composite is imprecise. It’s that the composite can point at a region of the map where nobody actually stands. Average a group that’s really two groups and you don’t get a blurry version of either — you get a third person, invented by the arithmetic, who belongs to neither. Build the product for that invented person and you ship something a little wrong for absolutely everyone, which is a worse place to be than a little wrong for half and right for the other half.

The mean is a mask

Picture a spectrum: on one end, people who reach for a single feature every day; on the other, people who touch a bit of everything, rarely. Interview six and they scatter — a couple here, a couple there, a stray in the middle. With only six dots the eye can’t tell a gap from noise, so you do the natural thing and take the center. The center looks like a moderate, all-rounder user. Reasonable. Plannable. A persona is born.

Now run the same study a hundred deep. The scatter resolves. Two dense clusters rise on either side and the middle empties out — and the average you trusted is sitting in the valley between them, on ground where almost no one is. It was never describing a real user. It was describing the midpoint of two real users who have nothing to do with each other.

A mean is a summary that assumes there’s one thing to summarize. When there are two, it stops being a description and becomes a decoy — a confident little marker planted exactly where you should not build.

This is the failure that hides behind small-sample research, and it’s different from the one we usually warn about. It’s not that six people can’t tell you anything — they can. It’s that six people can’t tell you what kind of distribution you’re standing in, and the shape is the part that decides what to do next. Is this one audience or two? A long tail or a clean split? You cannot see the shape of a crowd through a keyhole, and six interviews is a keyhole.

Distribution is the deliverable

The instinct to average comes from scarcity. When each interview costs a scheduled call, a no-show buffer, and an afternoon of notes, you can afford so few that a single summary is all the data will bear — there simply aren’t enough points to trust a second cluster over sampling noise. So the format of the finding gets quietly dictated by the cost of the finding. Expensive research produces personas because personas are what you can defend on ten interviews. The number of conversations you can afford decides whether you’re allowed to see a distribution at all.

Change the cost and you change what you’re permitted to conclude. This is the whole reason AI User Interviews exists — not to make interviews cheaper as a line item, but to move research off the calendar so the cap comes off how many you can run. At a hundred-plus conversations a week, the second cluster stops being a hunch you can’t act on and becomes a hill you can see and count. The output is no longer one averaged stranger; it’s there are two groups, here’s the split, here’s what each actually wanted. You get to hand a builder the shape instead of the midpoint.

And the shape is what a builder can use. “Our average user wants a simpler default and more power” is a contradiction you’ll spend a roadmap trying to satisfy. “Forty percent want one feature to be flawless and sixty percent want to poke at everything” is two coherent bets you can actually make — and, just as usefully, it tells you the tidy compromise in the middle is the one thing to not build, because it’s aimed at the empty valley.

Enough of the right thing, not more of the wrong one

None of this is an argument that bigger is automatically better, and it’s worth saying so plainly, the same way we try to stay honest about when a signal is real versus noise. A hundred sloppy interviews with a leading script give you a hundredfold-confident wrong answer, which is the most expensive kind. Volume doesn’t launder bad questions; it multiplies them. The point of scale here is specifically to resolve structure — to tell one hill from two — and past the point where the clusters have stabilized, another fifty conversations mostly confirm what you already have. That’s the same discipline as knowing when you’ve heard enough: you run until the shape holds still, not forever.

So the goal isn’t “as many interviews as possible.” It’s enough to see the distribution instead of a decoy — which is almost always more than a scarce, calendar-bound process will ever let you run, and almost always fewer than a nervous team thinks it needs once the tap is open. The right number is the one where a second look doesn’t move the hills. Below it, you’re averaging in the dark. Above it, you’re paying for reassurance.

Why a studio this small cares

We didn’t arrive at this from theory; we kept getting fooled by our own composites. We build and run our own products — that’s the one rule the studio runs on — and the version of this mistake we made most was averaging a handful of user notes into a single “they” and then arguing about what “they” wanted, when “they” were plainly two camps we’d mashed together to keep the number of interviews small enough to afford. The fights were really about a person we’d invented.

The average will always be computable; that was never the problem. The problem is that a mean quietly promises there’s one crowd to describe, and it keeps that promise even when it’s wrong — handing you a confident marker planted in the emptiest part of the room. Talk to enough people that the room’s real shape shows through, and you stop building for the phantom in the middle and start building for the people who are actually standing there. There are usually more of them than you thought, and they were never all the same.