Two users sit for the same study a week apart. The first lights up about a feature; the second shrugs at it. That contrast goes in the deck as a finding — reactions to the feature are mixed — and the team plans around it. But before you build anything on that sentence, it’s worth asking a question nobody asks: were they answered the same question, in the same way, by the same interviewer having the same day? Because if they weren’t, the difference in what they said might be a difference in how they were asked, and you have no way to tell the two apart.

Research treats the respondent as the variable and everything else as fixed. The sample varies, the answers vary, and you read the spread as signal about your users. What that quietly assumes is that the instrument — the interview itself — held still between sessions. In a human interview, it never does.

The interviewer is an instrument, and it drifts

Think about what actually happens across a day of user interviews. The script says “walk me through how you’d get started,” but by the third session it’s become “so you’d probably start here, right?” — the same question, worn into a leading one without anyone noticing. Tone shifts: a warm interviewer pulls warmer answers, a rushed one pulls shorter ones. Follow-ups fork — one participant gets a curious “say more about that,” the next gets a nod and a move to the next line. And the tenth interview of the day is run by a tired person who has already half-decided what they’re going to hear.

None of that is incompetence. It’s what it means for a person to conduct a conversation. But it means the measuring device changed shape between every reading, and in any other kind of measurement that would end the discussion. You would not trust a scale that weighed a little differently each time you stepped on it, then average the numbers and call the average a fact.

A user interview is a measurement, and the interviewer is the instrument. For decades we varied the one thing an instrument is supposed to hold still, then read the wobble as news about our users.

Hold the instrument still, and the noise separates from the signal

This is the part of AI User Interviews that’s easy to miss under the headline about speed. Yes, an AI interviewer means you can run a hundred conversations at once instead of ten in a week. But the quieter, more important thing is that it asks the same question the same way — the hundredth time as neutrally as the first, with the same patience, the same follow-up logic, no drift, no fatigue, no leaning in. The instrument stops moving.

And once the instrument is constant, a real thing happens to your data: the interviewer’s contribution to the variance goes to zero, and what’s left over is the respondent. When two answers differ now, they differ because the people differ — not because one of them got the warm interviewer on a good day and the other got the tired one reading a question that had quietly turned leading. You have, maybe for the first time, an actual controlled comparison. The difference in the answers finally means what everyone always pretended it meant.

A constant instrument can be constantly wrong — and that’s the discipline

Here’s where a tool like this gets oversold, so let me undersell it deliberately, the way we try to keep code-results honest about correlation. Holding the interviewer still removes the random error — the session-to-session wobble. It does nothing about systematic error. Write a leading question and an AI will ask it leadingly a hundred times in a row, with beautiful consistency. Standardization doesn’t launder a bad instrument; it just stops the instrument from being a new source of noise on every read.

That’s not a caveat that weakens the case — it’s the whole reason the case matters. A constant instrument is the thing that finally makes your questions the variable you’re testing, instead of hiding your bad questions inside the noise of who happened to ask them. When the asking is controlled, a leading question shows up as a suspiciously uniform answer, and you can go fix the script. When the asking is random, that same bias is smeared across the session-to-session variance and you never see it. Control isn’t a guarantee of truth. It’s what makes your own errors legible enough to correct — which is the only kind of research that improves.

Why a two-person studio cares who’s asking

We didn’t come to this from a methods textbook. We build and run our own products — the one rule the studio runs on — which means we do our own user research, badly, the way small teams do: a handful of calls, squeezed between shipping, each one run a little differently because we were a little different people on each of those afternoons. And then we’d argue about what a difference in two answers meant, with no way to rule out that the difference was just us.

The fix wasn’t to interview harder. It was to notice that the interviewer had been an uncontrolled variable the entire time, sitting right in the middle of the experiment, and to build the thing that holds it still. That’s the same move as running research as an instrument you leave on rather than an event you schedule: once the asking costs almost nothing and never varies, the interview stops being a performance and starts being a measurement. The answers were always supposed to be about your users. Now they can be — because the interviewer finally stopped being the variable.