Write out a research script and read it back and you’ll notice something a little deflating: the questions are the easy part. “What were you trying to do?” “Where did you get stuck?” “What did you do next?” Anyone can write those, and once they’re written they never change — you can print them, hand them to a hundred people, and get a hundred tidy answers to exactly the questions you thought to ask. That’s a survey. It is a fine thing to have. It is almost never where the finding is.

The finding is in the sentence after the answer. Someone says “I gave up and just exported it to a spreadsheet,” and the whole interview is what you do with that. A survey files it and moves to the next item. An interviewer leans in — why the spreadsheet? what did the tool make you feel you couldn’t do? what would have kept you inside it? — and two or three questions later, none of which existed a minute ago, you arrive at the actual reason, which was never the spreadsheet at all. The script got you to the door. Everything worth knowing was on the other side of it, reached by questions you could not have written in advance because they depend entirely on the answer you just heard.

A survey treats every answer as a destination. An interview treats it as a fork. The scripted question is a place you decided to stand; the follow-up is where you go once you’ve heard where they actually are.

The expensive part is the part that never scaled

Here is the quiet economics of user research, and the reason it stayed a boutique craft for so long. The cheap part — the script — is infinitely reproducible. The expensive part — the follow-up — is the one thing you cannot mass-produce, because each one is a fresh, improvised reaction to a specific human being’s specific sentence. A good interviewer is not someone with better questions on the page. They’re someone who hears “just exported it to a spreadsheet” and knows, in the moment, that this is a door and not a wall, and knows which way to push on it.

So teams did the rational thing and stopped paying for it. At scale you send the survey, because the survey scales and the interview doesn’t. You get a thousand answers to the questions you already knew to ask, and zero to the ones the answers should have raised. The follow-up — the whole reason the method works — got quietly designed out of the method the moment you needed more than a dozen of them. Everyone knew this was the trade. Nobody thought it was optional, because the improvising part lived in a person, and a person can only sit through so many conversations before “why?” goes flat.

That’s the assumption AI User Interviews exists to break. Not “the AI reads the script” — reading the script was never the hard part or the valuable part. The AI does the follow-up. It hears the spreadsheet sentence and decides, right there, that it’s worth a why; it asks a real one, grounded in what was actually said; it listens to that answer and decides whether to dig again or whether the reason underneath has finally stopped moving. The scripted questions are the trunk everyone already has. What we had to build was the branching.

Knowing when to stop digging

The failure mode isn’t a machine that can’t follow up. It’s one that follows up forever — a chatbot that responds to “the spreadsheet was easier” with an eager “interesting, tell me more!” for the ninth time, mining a seam that ran out three questions ago. A real interviewer has a sense the script can’t give them: the feel for when a why has bottomed out, when the person has told you the actual reason and the next probe would just be you refusing to hear it. Descend too little and you’ve run a survey with extra steps. Descend too much and you’ve worn out a respondent to hear the same thing louder.

So the mechanism has two moves, not one. Push: this answer is a surface, there’s a cause under it, ask. And stop: this answer is bedrock, the reason has stopped moving, thank them and go. Getting the first without the second gives you a tireless interrogator, which is worse than a survey, not better. The craft we were actually trying to reproduce lives in the judgment between them — and because the model never gets bored, never gets defensive at the ninth why, and asks the hundredth follow-up with exactly the attention it gave the first, the ceiling that made teams give up on interviews at scale simply isn’t there.

What you get back

Run it and the transcript doesn’t read like a survey with the answers pasted in. It reads like an interview — questions that quote the sentence before them, threads that go three deep on the moments that mattered and one deep on the ones that didn’t, the same improvised descent a good researcher would have made, a hundred times over, each one specific to the person on the other end. The script is still in there, holding the trunk steady so a hundred conversations stay comparable. But the part you came for was never the script. It was the fork after every answer, and the nerve to walk through it — and then the sense to know when you’d reached the bottom and could stop.