---
name: interview-guide-build
description: Produces the question set for user discovery interviews and expert interviews, structured so that answers can be wrong rather than merely encouraging. Builds questions from seven forms that anchor to specific past events, costs already paid, or artefacts the person can open, and removes the six forms that produce flattery. Names eleven ways a question leads a witness and rewrites each by removing the premise. Orders the guide into phases with a contamination line, before which the product may not be mentioned in the questions or in the calendar invite, and includes length limits, sampling counts per segment, a stopping rule, and a separate variant for expert interviews where the incentive inverts. This skill should be used before booking user or expert interviews, or after a round of conversations that produced encouragement and no facts.
---

# Interview guide build

## The claim this skill is built on

An interview guide is a measurement instrument, and a question that cannot be answered wrongly measures nothing.

This is the whole difference between a set of conversations that changes what a team builds and a set that confirms what they had already decided. It is not about asking more questions, or better rapport, or being sceptical of what you hear. It is about question form, decided in advance, in writing, before anyone is in the room.

The reason it has to be decided in advance is that you cannot repair it afterwards. A constructed answer and a remembered one arrive in identical fluent sentences, and no amount of care during analysis separates them.

## Why the obvious approach fails

The default guide is a list of topics phrased as questions: how do you currently handle this, what would you like to see, would you use a tool that does the thing. Every one of those is a compliment generator, for three separate reasons that stack.

People are poor predictors of their own future behaviour, so any hypothetical answer is a guess offered with unearned confidence. People are polite, and they are more polite to the person who built the thing, who is usually the one asking. And a question about general habits is answered from a reconstructed norm rather than from memory, so it produces a tidy average that no actual day matched.

The output is a set of transcripts full of encouragement, from which any conclusion can be drawn, which means none can.

## Two families of question form

### Forms that produce falsifiable answers

1. **Last instance.** "Walk me through the last time you did this. When was it?" A date is checkable and can be wrong. Also note the answer "I cannot remember the last time", which is data, not a failed question.
2. **Cost already paid.** "What did you do about it, how long did that take, and what did it cost?" Costs leave traces in calendars, invoices, tickets and headcount. Either the trace exists or it does not.
3. **Prior attempts.** "What have you already tried, and what happened to each attempt?" Someone who has tried nothing has a problem they can live with, and that is the most useful thing they will tell you.
4. **Artefact request.** "Can you show me the spreadsheet you use? Can we open it now?" What people show differs from what they describe, in one predictable direction: descriptions are tidier and shorter than the real thing. This comes from contextual inquiry, published by Beyer and Holtzblatt in 1998, whose model is that the interviewer is an apprentice watching the work rather than a questioner.
5. **Timeline reconstruction.** The switch interview from the jobs-to-be-done tradition: reconstruct the purchase or adoption as a sequence of events. First thought, passive looking, the triggering event, active looking, the decision, first use, and who else was involved. The same tradition names four forces acting on any switch: the push of the current situation, the pull of the new option, the anxiety about the new, and the habit holding the present in place. Anchoring to events makes gaps visible, and gaps are where the real trigger hides.
6. **Negative case.** "When did you last decide this was not worth bothering with?" This finds the boundary of the problem, which nobody volunteers because it is not a complaint. Related to the critical incident technique, published by Flanagan in 1954, which collects specific incidents with a defined outcome rather than opinions about performance.
7. **Commitment and referral.** "Who else has this that I should speak to?" and then an ask that costs something. There is a currency ladder: time, reputation, money. An introduction costs reputation, which is why it is worth more than agreement, which costs nothing.

A note on going deeper: laddering, published by Reynolds and Gutman in 1988, takes an answer about an attribute and asks why it matters, repeatedly, until it reaches a consequence and then a value. Three or four rungs is the practical limit before it becomes uncomfortable, and the useful rung is usually the second.

### Forms that produce flattery

1. **Hypothetical future.** "Would you use it?" Predicts nothing, and the yes is free.
2. **General habit.** "How do you usually do this?" A reconstruction, not a memory.
3. **Feature preference.** "Which of these would you want?" Everything sounds good when nothing has a price or a trade attached.
4. **Self-quoted price.** "What would you pay for this?" A number they have no way to know and no obligation to honour. Willingness to pay is measured by asking for money, not by asking about money.
5. **Approval seeking.** "Does that make sense?" The only comfortable answer is yes.
6. **Compound.** Two questions sharing one verb. They answer the easier half and it gets filed against both.

**The cut test.** For each question, write down the answer that would make you abandon your current hypothesis. If no possible answer to that question could do it, the question is not research. Cut it, or move it to the warm-up where it can do no harm.

## Eleven ways a question leads a witness

1. **Presupposition.** The answer is inside the question. "How frustrating is the export step?" has already established that it is frustrating.
2. **Loaded premise.** Assumes an unestablished fact. "What do you do when the sync fails?" assumes it fails.
3. **Acquiescence.** Agree and disagree formats pull agreement, and pull it harder from people with less power in the room or a reason to want the meeting to go well.
4. **Social desirability.** Anything with a respectable answer gets the respectable answer. Questions about how often people write tests, take backups or read the documentation are the classic cases.
5. **Anchoring.** Any number you say first becomes the reference point for every number they say afterwards, including prices and time estimates.
6. **Demand characteristics.** The person works out what you are hoping to hear and supplies it. Strongest when you are visibly the maker, which is most of the time in early discovery.
7. **Framing.** Gain-framed and loss-framed versions of the same question produce different answers from the same person.
8. **Order priming.** An earlier question makes a concept salient and it colours everything after. This is the mechanism behind the contamination line below.
9. **Closed option set.** Offering three options removes the fourth, and the fourth was the one you needed.
10. **Absolute terms.** "Do you always" invites a denial and loses the actual frequency, which was the number you wanted.
11. **Interviewer echo.** Repeating back a stronger version of what they said and having it confirmed. "So it is a real blocker for you?" converts a mild complaint into a quotable one, and the quote is yours rather than theirs.

**The read-aloud test.** Say each question out loud and ask what a person who disagrees with your premise would say. If there is no comfortable way to disagree, the question leads. Fix it by removing the premise rather than by softening it: "How frustrating is the export step?" becomes "Tell me about the last time you exported something."

## The ordering rule

The guide runs in phases, with a line in it.

- **Phase 0. Screen on behaviour, not on title.** The screener asks what the person did in the last month, not what their role is called. Role titles do not predict behaviour across organisations, and a wrong-sample interview is the most expensive mistake available and the cheapest to prevent.
- **Phase 1. Context.** Three to five minutes. Role, team, tools, scope. Nothing evaluative.
- **Phase 2. The problem, in specific past instances.** Last instance, timeline, what happened next. This is where the evidence lives.
- **Phase 3. The workaround and what it costs.** What they built, bought, hired, or endured, and the negative case.
- **The contamination line.**
- **Phase 4. The product, the concept, the demo.** Only here, and only if it is needed at all. Everything from this point is data about their reaction to your thing, and none of it is data about the problem.
- **Phase 5. Commitment and referral.** Ask for something that costs them: a follow-up with the person who actually does the work, a copy of the document, a pilot slot, a date, a deposit.

**Why asking earlier contaminates everything after it.** Three things happen at once the moment the person learns what you are building, and none can be undone. They begin filtering their answers for relevance to your thing, so the parts of their life that do not fit stop being mentioned. Politeness engages, hardest against the person who built it. And their memory of the problem is now retrieved through the frame of your solution, so even a sincere later answer is a different answer to the one they would have given.

The reason this is an ordering rule rather than a preference: you cannot re-run the person. Each participant is a single-use instrument, and in a sample of six or eight that loss is a large fraction of the study.

**This includes the calendar invite.** An invite that says "feedback on our new scheduling feature" has contaminated the interview before it starts. Invite wording is part of the instrument, so write it with the guide: name the topic area and the length, not the product.

## Numbers that shape a guide

- **Eight to twelve primary questions for a forty-five to sixty minute session.** Each good question spawns three to five follow-ups, and the follow-ups carry the evidence. A twenty-five question guide is a survey read aloud, and it guarantees there is no room for any of them.
- **Talk ratio.** The other person should produce roughly three quarters to four fifths of the words. Past a quarter of the words, you are running a demo.
- **Silence.** Count three seconds after they stop. The second half of an answer usually arrives after the pause, and it is usually the useful half.
- **Five to eight interviews per distinct segment**, before patterns repeat. The count is per segment and never in total, so three segments is fifteen to twenty-four conversations, not eight.
- **Stopping rule.** Stop a segment when two consecutive interviews produce no new codes. If the tenth still produces new codes, what you called one segment is more than one.
- **A correction worth carrying.** The widely repeated claim that five users are enough comes from work by Nielsen and Landauer published in 1993 on how many participants are needed to find most usability problems in a single interface. It is not a claim about discovery interviews, market segments, or anything statistical, and using it to justify five conversations spread across four segments misapplies a real result.
- **Notes.** Record verbatim quotes with a timestamp, and mark every note as observed, quoted, or inferred. A paraphrase written during the session is already an interpretation and the original wording is gone.

## The expert interview variant

The incentive inverts. A user wants to be helpful. An expert wants to be authoritative, and an authoritative answer is available for every question including the ones they cannot know.

Six moves that change:

- Ask what they have personally observed, with numbers they have seen themselves, rather than industry figures they have read.
- Ask for the counterexample: when does this fail, and who does it not apply to.
- Ask what changed their mind, and roughly when. An expert who has never updated is reporting a position rather than a practice.
- Ask them to rank rather than rate. Ranking forces a trade and rating does not, so five items rated highly tells you nothing and five items ordered tells you their model.
- Ask what the consensus in their field gets wrong. This is frequently the question they have been waiting to be asked and it is where the specific detail lives.
- Send the topic in advance, never the questions, for anything you want unrehearsed.

And one prohibition: never ask an expert to predict. Ask what they have seen. Prediction confidence rises with expertise more reliably than prediction accuracy does, so a forecast from an expert arrives with a credibility it has not earned.

## The decision rule for each drafted question

Classify every question before the guide is finished:

- **Falsifiable.** Names a specific past event, a cost that was paid, or an artefact you can ask to see. Keep it.
- **Leading.** Matches one of the eleven patterns. Rewrite by removing the premise, not by softening it, then reclassify.
- **Unfalsifiable.** No possible answer changes what you believe. Cut it, or demote it to Phase 1.
- **You cannot tell.** Apply both tests. Write down the answer that would make you abandon your hypothesis; if you cannot write one, it is unfalsifiable. Then read it aloud and ask whether a person who disagrees with your premise has a comfortable way to say so. If it survives both, keep it and mark it for review after the first two interviews, because the form of a question is easiest to judge from the answers it actually produced.

## Worked example, compressed

**Situation.** A team is adding recurring scheduling to a project management tool and has drafted six questions.

**The draft, classified.**

1. "How do you currently handle recurring work?" General habit. Unfalsifiable.
2. "How frustrating is it when a recurring task gets missed?" Presupposition and a loaded premise.
3. "Would you use a feature that creates recurring tasks automatically?" Hypothetical future, plus demand characteristics, plus it names the product in question three.
4. "Which of these three options would you prefer?" Closed option set.
5. "How much would you pay for that?" Self-quoted price, anchored by whatever they were shown in question four.
6. "Does that make sense?" Approval seeking.

Verdict on the draft: six questions, none falsifiable, five leading patterns present, and the product introduced at question three, which contaminates the remaining half of the session regardless of how those questions were worded.

**The rebuild.** Nine questions, seven of them above the line.

1. Context: what is your role, and how many people's work do you track?
2. Tell me about the last piece of work that had to happen again on a schedule. When was that?
3. Walk me through the last time one of those was missed. What happened next, and who noticed?
4. What did you do about it afterwards?
5. Where do those live now? Can we open it while we talk?
6. What have you already tried for this, and what happened to each thing you tried?
7. When did you last decide a repeating item was not worth tracking at all?
8. Contamination line, then the concept: what is your reaction, and what would stop you using it?
9. Can I come back in two weeks with a working version, and would you bring the person who maintains that sheet?

**Verdict.** Nine questions, seven above the line, five that produce a dated event or an opened artefact, one negative case, and two commitment asks with different currencies. What the rebuild gives up is the ability to ask early whether they would use it, which was the question the team most wanted answered and the one no interview can answer. That trade should be stated to the team before the round starts, or someone will add question three back.

## Failure modes

**The pitch in disguise.** A demo with pauses in it. The tell is the talk ratio and the fact that every session ends warmly.

**The feature poll.** Options presented, preferences counted, and the result treated as demand. Preference between free things is not demand.

**The compliment banked as validation.** "That sounds really useful" recorded as a yes. It is a politeness token and it has no information in it.

**Premise leakage.** A question that assumes the thing you are trying to establish, so the transcript contains your assumption in their voice.

**The general habit filed as a memory.** "I usually check it every morning" written into the notes as behaviour, when nobody was asked about a specific morning.

**Talking over the pause.** The interviewer fills three seconds of silence and loses the second half of the answer, permanently, and never knows it happened.

**The spoken survey.** A guide so long there is no room for follow-ups, so every answer stays at the first level and nothing is ever pursued.

**Sampling the enthusiastic.** Recruiting from your own followers, your existing customers or your network and calling it the market. The tell is that the recruiting channel never appears in the write-up.

**The invite that contaminates.** The interview is spent before it starts because the calendar invite named the feature.

**Paraphrased notes.** No verbatim quotes, so six weeks later nobody can tell what the person actually said, and the strongest available evidence is somebody's memory of their own summary.

**No commitment ask.** Every conversation ends in agreement, nobody was asked for anything that costs, and a round of interviews produces no information about whether anyone would act.

## What this skill does not do

- It does not conduct the interview. Talk ratio, silence, and choosing the right follow-up happen live and account for most of the difference between a good session and a wasted one.
- It does not recruit, and it cannot see whether your sample is the right one. Recruitment error is larger than question error and it is invisible in the transcripts.
- It does not analyse. Coding, theming, counting how many people said a thing and keeping the evidence trail are separate work with their own methods and tools.
- It is not a usability test protocol. Those are task-based, mostly silent, and scored on what people do rather than what they say.
- It does not cover consent, recording law, data protection or institutional ethics review, all of which are binding in some settings and none of which are optional there.
- It cannot tell you whether the answer you got is representative. Five people who all said the same thing may be five people from the same place.
