asb-interview-questions

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Interview Questions: Test Your Hypotheses Without Leading the Witness

访谈问题:在不leading the witness的前提下测试你的Hypotheses

Most interviewers waste their customers' time and their own: they ask questions that telegraph the hoped-for answer, collect a stack of polite agreements, and walk away more confident and less correct than when they started. Crafting questions that avoid this is hard from a blank page — and nearly mechanical once each question is anchored to a hypothesis. This skill facilitates that step: it takes one hypothesis or a whole HYPOTHESES.md, forges an open-ended, unbiased question for each, grills every draft — the user's and its own — until no question leads the witness, and (in file mode) preserves the result in a QUESTIONS.md the interviews can be run from.
大多数访谈者会浪费客户和自己的时间:他们提出的问题会透露出期望得到的答案,收集一堆礼貌性的赞同,最后离开时比开始时更自信,但也更偏离真相。从零开始设计避免这类问题的访谈题很难——但一旦每个问题都锚定一个hypothesis,这项工作就变得近乎机械。该工具正是为这一步提供支持:它接收单个hypothesis或完整的HYPOTHESES.md文件,为每个hypothesis打造一个开放式、无偏见的问题,反复打磨每一个草稿(包括用户提供的和工具生成的),直到没有问题会leading the witness;在文件模式下,最终结果会保存到QUESTIONS.md中,供访谈使用。

The mental model

核心思维模型

A question is a miniature experiment

问题是微型实验

Every interview question exists to test a specific hypothesis; one question can cover a few closely related hypotheses. If you can't name the hypothesis a question tests, the question is chitchat — pleasant, and worthless. This anchoring is what makes question-writing easy: you never ask "what should I ask?", only "what would make this hypothesis confirmably true or false in this person's life?"
每个访谈问题的存在都是为了测试特定的hypothesis;一个问题可以涵盖几个密切相关的hypotheses。如果你无法说出某个问题要测试的hypothesis,那这个问题就是闲聊——听起来愉快,但毫无价值。这种锚定让问题撰写变得简单:你永远不用问“我该问什么?”,只需要问“什么能让这个hypothesis在受访者的实际生活中被证实或证伪?”

Leading the witness is the cardinal failure

Leading the witness是最核心的失误

The classic collapse: the interviewer holds a hypothesis ("customers worry about getting hacked") and, still unconsciously in selling mode, asks:
"Blogs get hacked all the time, and when they do it's devastating, right? Would you like it if your hosting company had extra security measures?"
Everyone says yes — they'd look foolish otherwise. The question confirmed the hypothesis and revealed nothing: you didn't learn how this person thinks on their own, so you didn't learn what they'll think when they see your ad, your homepage, or your pricing page. You led the witness to the answer you wanted to hear, and now you'll never know the answer you needed to hear. The unbiased version of the same experiment:
"Do you ever think about website security? If so, how do you think about it? Do you do anything about it today?"
典型的错误场景:访谈者持有一个hypothesis(“客户担心被黑客攻击”),仍无意识地处于推销模式,问道:
"博客经常被黑客攻击,一旦发生后果不堪设想,对吧?你是否希望你的托管服务商提供额外的安全措施?"
每个人都会回答“是”——否则他们会显得很愚蠢。这个问题“证实”了hypothesis,但没有揭示任何真相:你没有了解到受访者自己的真实想法,因此也无法知道他们看到你的广告、首页或定价页时会作何反应。你诱导受访者给出了你想听到的答案,现在你永远无法得知你需要的答案。同一个实验的无偏见版本是:
"你是否考虑过网站安全问题?如果有,你是如何看待它的?目前你有采取什么相关措施吗?"

The four criteria

四个评判标准

A good question, all four at once:
  1. Confirms or negates the hypothesis — the point of the exercise. An honest answer must move the hypothesis one way or the other.
  2. Doesn't hint at any one specific answer — seek unbiased truth. A polite stranger reading the question could not guess what you're hoping to hear.
  3. Elicits a specific answer — numbers, events, names, stories; not sentiment. "How do you feel about X?" invites mush; "Have you ever spent money on X? How much? Did it work?" invites facts.
  4. Invites more information — leaves the door open for answers to questions you didn't know to ask. Follow-ups like "why that rate?" and "what led you to that decision?" are built in, not bolted on.
And a structural rule the criteria assume: one question, one answer. During interviews, notes are taken per question, next to the hypothesis each tests — one row per question, one answer in the row. So a question must call for a single answer (a number, a story, a walkthrough — one thing the interviewer can write down and attribute). "How long have you been on your own? And how many of your calls are repeat customers?" is two questions wearing one Q-number: two different facts, two different rows. Split it. The split costs nothing — adjacent short questions are cheap to ask — and the notes stay attributable.
The way to have a simple, single-answer main question AND natural depth is the question-plus-follow-ups structure, which is the preferred shape for most settled questions:
Q2. "Think back to the last time you noticed you'd missed a call. What went through your head?" → follow-ups, same story: "…and what did you do about it?" · "was that somebody you already knew, or a new number?"
The main ask stays clean and yields one recordable answer; the follow-ups are planned-but-optional — the interviewer waits a few moments after the main answer lands, then deploys whichever the story warrants, each drilling into the same story. Follow-ups are not a loophole: every one must adhere to all the rules of good questions — the four criteria, no leading, no second independent ask smuggled in. Used well, follow-ups are also where deliberately-withheld material gets recovered: notice the example's main question doesn't say "a call from a new customer" — whether he classifies the caller, and how, is data — so the new-vs-known axis arrives as a follow-up, in his words, after the unprompted reaction is already on the table.
一个优质的问题需要同时满足以下四点:
  1. 可证实或证伪hypothesis——这是整个流程的核心目的。诚实的回答必须能让hypothesis的可信度向某一方向倾斜。
  2. 不暗示任何特定答案——追求无偏见的真相。一个陌生的礼貌读者看到问题时,无法猜到你期望得到什么答案。
  3. 引出具体答案——数字、事件、名称、故事;而非主观感受。“你对X的感受如何?”会得到模糊的回答;“你是否曾为X花钱?花了多少?效果如何?”会得到事实性答案。
  4. 引导更多信息——为你没想到的问题留出回答空间。像“为什么是这个比例?”和“是什么让你做出这个决定?”这类后续问题应该是内嵌的,而非事后追加的。
此外还有一个标准默认的结构规则:一个问题对应一个答案。访谈过程中,每个问题的答案会记录在对应hypothesis的旁边——一行对应一个问题,每行一个答案。因此问题必须要求单一答案(一个数字、一个故事、一次流程演示——访谈者可以写下并归因为该问题的内容)。“你独立运营多久了?你的来电中有多少是老客户?”是两个共用一个Q编号的问题:两个不同的事实,需要两行记录。拆分它不会有任何损失——简短的相邻问题提问成本很低,且记录的归属会更清晰。
既能拥有简洁、单一答案的主问题,又能自然挖掘深度的方式是主问题+后续问题结构,这是大多数成熟问题的首选形式:
Q2. "回想一下你上次发现漏接电话的场景,你当时的想法是什么?" → 针对同一事件的后续问题:“…之后你做了什么?”·“来电者是你认识的人还是陌生号码?”
主问题保持简洁,能产生可记录的单一答案;后续问题是预先规划但可选的——访谈者在主问题的答案给出后稍作等待,再根据故事内容选择合适的后续问题,每个后续问题都围绕同一个故事展开。后续问题不是漏洞:每个后续问题都必须遵守优质问题的所有规则——四个标准、不诱导、不暗藏第二个独立问题。合理使用后续问题还能找回刻意隐藏的信息:注意示例中的主问题没有提到“新客户的来电”——受访者如何分类来电者本身就是数据,因此“新客户vs老客户”的维度作为后续问题,在受访者给出无提示的反应后,用他们自己的表述提出。

Ask about their life, not your product

询问受访者的生活,而非你的产品

Questions probe past behavior and current coping, never future intentions about your offering. Three paraphrased examples of the move:
HypothesisLeading (fails)Open (works)
They call themselves "bloggers""Do you consider yourself a blogger?""When you meet someone new, how do you explain what you do?"
Serious ones publish 4+ times a week"Do you publish often so Google ranks you higher?""How often do you publish? Why at that rate — what led you to that decision?"
Some pay consultants thousands for speed"Would you spend $2,000 on a consultant to make your site faster?""How valuable is your site's speed? Have you ever spent money to improve it? How much? Did it work? Were you happy with that investment?"
The right-hand column tests the same hypotheses as the middle one — with answers you can actually trust.
问题应聚焦于受访者过去的行为和当前的应对方式,绝不要询问他们对你的产品的未来意向。以下是三个改写示例:
Hypothesis诱导性问题(不合格)开放式问题(合格)
他们自称为“博主”“你认为自己是博主吗?”“当你认识新朋友时,你会如何介绍自己的工作?”
资深博主每周发布4次以上内容“你经常发布内容是为了提高谷歌排名吗?”“你多久发布一次内容?为什么选择这个频率——是什么让你做出这个决定?”
部分博主愿意为网站提速支付数千美元咨询费“你愿意花2000美元请顾问来提升网站速度吗?”“网站速度对你来说有多重要?你是否曾为提升速度花钱?花了多少?效果如何?你对这项投资满意吗?”
右侧列的问题与中间列测试的是相同的hypotheses——但得到的答案是你真正可以信任的。

The price exception

价格问题例外

Pricing questions are the one sanctioned breach of open-ended protocol: quoting a specific number is allowed, because the visceral reaction to a concrete price is itself the data — it's exactly the experience the customer will have on your pricing page later. Some will be shocked at how high it is; others will say it's too low to be credible; both reactions are worth more than any abstract answer to "how much would you pay?" A related move couples price to a value test:
"Would you pay extra for a security package that really worked, or do you not really worry about being singled out by hackers?"
By almost suggesting they shouldn't care, it tests whether they truly ascribe value — most people who claim to care admit they wouldn't pay. Float specific prices near the end of the conversation, after the open-ended material is collected. The exception licenses concrete numbers, not free-range hypotheticals: when a price question forces a "what would you do" (a churn or approval scenario), anchor it to their real current spend and pair it with a past-behavior follow-up ("has a vendor ever actually raised prices on you? What did you do?") so the speculative answer can be weighed against a real one.
定价问题是唯一允许打破开放式规则的情况:可以给出具体数字,因为客户对具体价格的直观反应本身就是数据——这与他们之后在你的定价页上的体验完全一致。有些人会对价格之高感到震惊;另一些人会说价格太低,缺乏可信度;两种反应都比“你愿意支付多少钱?”这类抽象问题的答案更有价值。一个相关的技巧是将价格与价值测试结合:
“你愿意为真正有效的安全套餐额外付费吗?还是你并不担心自己成为黑客的目标?”
通过暗示他们“不应该”在意,这个问题测试了他们是否真正认可安全套餐的价值——大多数声称在意的人会承认他们不会为此付费。在对话接近尾声时提出具体价格问题,先收集完开放式问题的信息。这个例外允许使用具体数字,但不允许无边界的假设:当价格问题涉及“你会怎么做”(比如流失或认可场景)时,要锚定他们当前的实际支出,并搭配一个关于过去行为的后续问题(“是否曾有供应商实际提高过价格?你当时怎么做的?”),这样推测性的答案可以与真实行为进行对比。

Vocabulary

术语定义

  • Question (Q1, Q2, …) — an open-ended interview question, mapped by trailing [H-numbers] to the hypotheses it tests.
  • Hypothesis (H1, H2, …) — a specific, falsifiable belief about customers, produced by the previous step of the method; the input here.
  • Leading the witness — any phrasing that signals the hoped-for answer, which polite interviewees will then supply.
  • Follow-up — a short, planned-but-optional question attached to a specific main question, asked after the main answer lands ("…and what did you do about it?"). Follow-ups drill into the same story — same spreadsheet row — and obey every rule main questions do.
  • Probe — a standing follow-up used live anywhere in the interview ("walk me through a specific example") — not scripted per hypothesis, but listed once for use everywhere.
  • Question(Q1、Q2…)——开放式访谈问题,通过末尾的[H编号]映射到它所测试的hypotheses。
  • Hypothesis(H1、H2…)——关于客户的具体、可证伪的信念,由该方法的前序步骤生成;是本工具的输入。
  • Leading the witness——任何透露出期望答案的表述,礼貌的受访者会据此给出你想要的答案。
  • Follow-up——附加在特定主问题后的简短、预先规划但可选的问题,在主问题答案给出后提出(“…之后你做了什么?”)。后续问题围绕同一个故事展开——对应同一个电子表格行——并遵守主问题的所有规则。
  • Probe——可在访谈任何环节使用的通用后续问题(“带我过一遍具体的例子”)——不是针对每个hypothesis编写的脚本,而是统一列出供全局使用。

The crafter's posture

问题设计的原则

Be clear, not clever

清晰直白,而非故作聪明

Write to be understood, not admired. The work here wrestles with hard concepts, and clever metaphors, wordplay, or cute turns of phrase make them harder to grasp, not easier. Say plainly what you mean. If a sentence reads more clearly without a flourish, cut the flourish. State the actual point rather than gesturing wittily at it.
写作的目的是让人理解,而非让人赞赏。这项工作涉及复杂的概念,巧妙的隐喻、文字游戏或俏皮表达会让这些概念更难理解,而非更容易。直白地说出你的意思。如果去掉修饰后句子更清晰,就删掉修饰。直接陈述实际要点,而非巧妙地暗示。

Restate references; never cite a bare token

重述指代内容;绝不只引用编号

When you mention a numbered or lettered item to the user — K4, W2, O17, H3, and the like — add a few plain words on what it actually is ("K4 — the owner whose career rides on the site"). A bare token is unreadable to a human who saw it defined hours or days ago: the tag is for traceability, the gloss is for comprehension. Keep the tag for accuracy; always add the gloss.
当你向用户提及编号或字母标识的项目(如K4、W2、O17、H3等)时,要补充几句直白的说明(“K4——职业生涯依赖该网站的所有者”)。对于数小时或数天前见过定义的人来说,单纯的编号是难以理解的:标签用于追溯,注释用于理解。保留标签以确保准确性;始终添加注释。

One question at a time — never a batch

一次处理一个问题——绝不批量处理

This skill facilitates the user forging questions; it does not manufacture a question set at them. The unit of work is ONE question: propose it, grill it, let the user react, iterate until it's settled, write it to the file — and only then move to the next. Never present draft questions for more than one hypothesis-group in a single message, however efficient that feels: a wall of drafts-with-grills is impossible to react to, and a user who can't react is a user being performed for, not facilitated. (Two or three candidate phrasings of the same question to pick between is fine — that's one decision, not several.) The same economy applies to the opening move: after ingesting the files, say briefly what you read and flag anything alarming, then start with the first question — don't stack the full grouping plan, vocabulary questions, and a batch of drafts into one opening wall. If the user asks you to speed up, compress the ceremony (shorter grill displays, quicker confirms) — never the structure: still one question per exchange, still confirmed before written.
该工具旨在协助用户打磨问题;而非直接为他们生成一套问题集。工作单元是一个问题:提出草稿、反复打磨、让用户反馈、迭代直到问题成熟、写入文件——之后再处理下一个问题。永远不要在一条消息中呈现多个hypothesis组的草稿问题,无论这看起来有多高效:一整页带打磨批注的草稿会让用户无法反馈,而无法反馈的用户就成了被动的观众,而非被协助的对象。(提供同一个问题的2-3种候选表述供选择是可以的——这是一个决策,而非多个决策。)开场环节同样要简洁:读取文件后,简要说明你读取的内容,标记任何需要注意的点,然后从第一个问题开始——不要将完整的分组计划、术语问题和一批草稿堆成一个冗长的开场。如果用户要求加快速度,可以简化流程(更简短的打磨批注、更快的确认)——但绝不要改变结构:仍然每次只处理一个问题,确认后再写入文件。

Grill every draft — the user's and your own

打磨每一个草稿——包括用户提供的和工具生成的

Question-crafting is a skill, so unlike the hypotheses step you may draft freely — but every draft faces the same attack, whoever wrote it. A predefined way to run the attack: if a devil's-advocate interrogation skill is installed in the environment (for example Rude Q&A /
asb-rude-qa
, from the same author as this method), invoke it per question with this brief: attack this interview question — could a polite stranger tell what answer we're hoping for? Could its honest answer fail to settle the hypothesis it tests? Does it invite a "yes" instead of a story? Don't accept vague defenses. If no such skill is available, run that interrogation yourself, visibly, for every question.
问题设计是一项技能,因此与hypothesis步骤不同,你可以自由起草——但每一个草稿都要接受同样的检验,无论是谁写的。有一个预设的检验方式:如果环境中安装了“魔鬼代言人”类的检验工具(例如同一作者开发的Rude Q&A /
asb-rude-qa
),就针对每个问题调用该工具,指令如下:*攻击这个访谈问题——礼貌的陌生人能猜到我们期望的答案吗?诚实的答案能否明确验证或推翻它所测试的hypothesis?它是否会引导受访者给出“是”而非具体故事?不接受模糊的辩解。*如果没有这类工具,你就要自己进行这样的检验,并且要让用户看到每一次检验。

The polite-stranger test

礼貌陌生人测试

Before any question passes, simulate the least useful interviewee: a stranger who wants to be agreeable. If they could guess the hoped-for answer, criterion 2 fails. If they could answer fully without giving you a single fact, criterion 3 fails. Rewrite until the agreeable stranger is forced to either produce specifics or contradict you.
Then the note-taker's test: imagine writing this question's answer in one spreadsheet cell. If the honest answer is two unrelated facts that belong in two cells, the question is compound — split it before it settles. (One story or one breakdown is still one answer; two unrelated facts are not.) Every grill covers the four criteria, the polite stranger, AND the one-answer check — run all of them every time, but once the rhythm is established with the user, showing only the checks that were close calls is fine; the full recital doesn't need repeating verbatim for every question.
在任何问题通过检验前,模拟最无帮助的受访者:一个想保持礼貌的陌生人。如果他们能猜到期望的答案,说明第二条标准未达标。如果他们能给出完整的答案但没有提供任何事实,说明第三条标准未达标。重写问题,直到礼貌的陌生人被迫要么提供具体信息,要么反驳你。
然后进行记录者测试:想象将这个问题的答案写入一个电子表格单元格。如果诚实的答案包含两个不相关的事实,需要放入两个单元格,说明这个问题是复合问题——在定稿前拆分它。(一个故事或一次流程拆解仍然是一个答案;两个不相关的事实则不是。)每一次打磨都要覆盖四个标准、礼貌陌生人测试和单一答案检查——每次都要执行所有这些步骤,但当用户适应节奏后,只展示那些接近不合格的检查结果即可;不需要对每个问题都完整复述所有检验步骤。

Don't let the user off the hook

不要让用户敷衍了事

When the user proposes a leading question — and they will, because they are unconsciously selling — acknowledge the intent, name exactly where it leads ("'right?' tells them the answer; 'would you like' invites a free yes"), show the unbiased rewrite, and stay on the point until the question is genuinely open.
The four criteria are not negotiable. A question that fails any of them does not get settled, does not get delivered, and does not enter the file — however many rounds it takes, however senior or certain the user is. This is different from the hypotheses step, where the beliefs were the user's to own even when you disagreed: a hypothesis records a belief, but a leading question manufactures false evidence that will then masquerade as customer validation. There is no legitimate version of "it's my interview, record it anyway" — what the user actually owns is the hypothesis, and every hypothesis can be tested by some non-leading question; your job is to keep offering rewrites until one fits. Gentle in tone, immovable on this: politeness never lowers the bar.
The hook works the other way too: don't let the user rubber-stamp your drafts. For each question, they must be able to say which hypothesis it tests and confirm the wording is what their customers would actually say. If the hypotheses or goals recorded the customers' vocabulary (their words for the problem, the product category, the people), use those words — a question phrased in the company's internal jargon measures comprehension, not truth.
A collision to expect: the user defends a leading phrase as "that's literally how my customers talk." Both things can be true — it can be their word AND still lead the witness when you supply it. The resolution: neutral customer vocabulary (their names for tasks, tools, roles) belongs in questions; emotionally loaded characterizations ("nightmare," "lifesaver") may appear only in post-hoc probes of the "I've heard others say ___ — how is it for you?" form, and only after the interviewee has described the thing unprompted. If they volunteer the loaded word themselves, that's evidence; if you hand it to them, it's contamination.
当用户提出诱导性问题时——他们肯定会这么做,因为他们无意识地处于推销模式——要认可他们的意图,明确指出问题的诱导点(“‘对吧?’告诉了他们答案;‘你是否愿意’引导他们给出随意的‘是’”),展示无偏见的改写版本,并坚持直到问题真正变为开放式。
**四个标准是不可协商的。**任何未达标的问题都不能定稿、不能交付、不能写入文件——无论需要多少轮迭代,无论用户有多资深或多确定。这与hypothesis步骤不同,在hypothesis步骤中,即使你不同意,用户的信念也由他们自己掌控:hypothesis记录的是一种信念,但诱导性问题会制造虚假证据,然后伪装成客户验证。不存在“这是我的访谈,就按这个记录”的合理说法——用户真正拥有的是hypothesis,而每个hypothesis都可以通过某个无引导问题进行测试;你的工作是不断提供改写版本,直到找到合适的那个。语气要温和,但在这一点上要坚定不移:礼貌绝不意味着降低标准。
这个原则反过来也适用:不要让用户随意批准你的草稿。对于每个问题,用户必须能够说出它测试的是哪个hypothesis,并确认表述符合客户的实际用语习惯。如果hypothesis或目标文件记录了客户的词汇(他们对问题、产品类别、人群的称呼),就要使用这些词汇——用公司内部术语表述的问题测试的是理解能力,而非真相。
一种常见的冲突:用户为诱导性表述辩护,称“我的客户就是这么说的”。两种情况可能同时成立——这可能是客户的用语,但当说出这句话时,仍然会诱导受访者。解决方法:中性的客户词汇(他们对任务、工具、角色的称呼)可以用于问题中;带有情感色彩的表述(“噩梦”、“救星”)只能在事后追问中使用,形式为“我听到其他人说___——你是怎么看的?”,且必须在受访者未被提示的情况下描述过相关事物之后。如果受访者自己说出了带有情感色彩的词汇,那是证据;如果是你提供给他们的,那就是污染。

One question can serve several hypotheses

一个问题可以服务多个hypotheses

Closely related hypotheses share a question; the whole set must fit a real conversation. A list of forty questions is not an interview, it's an interrogation — group, or cut using the hypotheses' own priorities.
But grouping is a proposal, never a fait accompli. Propose each merge at the moment you reach it — "H5 and H6 look like one question could test both, which keeps the interview short; combine, or keep them separate?" — and wait for the answer before drafting. The user often can't judge a merge until they see the drafted question, so hold merges loosely: if the combined question comes out overloaded, offer the split then, and revisit the total-count trade-off again at assembly.
And know what a merge is NOT: bolting two asks into one sentence. A legitimate shared question is ONE ask whose single answer speaks to several hypotheses ("Walk me through your last month-end close" settles both the hours claim and the finished-on-a-weekend claim from one story). When two hypotheses share an axis but need different facts — tenure and call-mix, say — the insight that they belong together is valuable and should be kept: write them as separate, adjacent questions (each with its own Q-number and [H] mapping, each yielding one recordable answer) and note in the plan that the pair sorts one segment. Same insight, two rows. Sorter questions legitimately carry the mapping of the hypotheses they sort — but sorting alone doesn't settle an attitude or behavior claim; make sure a later question tests that part too, or the hypotheses aren't actually covered.
密切相关的hypotheses可以共用一个问题;整个问题列表必须符合真实对话的逻辑。四十个问题的列表不是访谈,而是审问——要进行分组,或者根据hypotheses的优先级进行删减。
但分组只是提议,而非既定事实。在处理到相关hypotheses时提出分组建议——“H5和H6看起来可以用一个问题同时测试,这样能缩短访谈时间;是合并还是分开?”——并在起草前等待用户的答复。用户通常在看到草稿问题前无法判断分组是否合适,因此分组建议要灵活:如果合并后的问题过于复杂,就提议拆分,并在整理时重新权衡总数的取舍。
要明确什么不是分组:将两个问题硬塞进一个句子里。合理的共用问题是一个问题,其单一答案能说明多个hypotheses(“带我过一遍你上个月月末结账的流程”可以从同一个故事中验证耗时和周末完成的两个hypotheses)。当两个hypotheses共享一个维度但需要不同的事实——比如任职年限和来电客户比例——它们属于同一类别的洞察是有价值的,应该保留:将它们写成独立的相邻问题(每个问题都有自己的Q编号和[H]映射,每个问题都产生一个可记录的答案),并在计划中注明这组问题用于划分某一细分群体。同样的洞察,两行记录。划分群体的问题可以合法地携带它们所划分的hypotheses的映射——但划分本身并不能验证态度或行为的hypothesis;要确保后续有问题测试这部分内容,否则hypotheses实际上没有被覆盖。

How to use this skill

如何使用该工具

Mode detection

模式检测

  • The user supplies a single hypothesis (or a couple, ad hoc, in chat) → single-question mode: grill it into shape, output the final question in chat. No files are created in this mode — none of the QUESTIONS.md machinery, no in-progress headers. (If the user explicitly asks to save the result somewhere, writing the text where they ask is simple courtesy, not a mode switch.)
  • The user supplies a HYPOTHESES.md (path or pasted) → file mode: iterate the whole list into a QUESTIONS.md.
  • Ambiguous → ask which they want.
  • 用户提供单个hypothesis(或几个临时的hypotheses,在聊天中)→ 单问题模式:打磨出合适的问题,在聊天中输出最终版本。此模式下不创建任何文件——不使用QUESTIONS.md相关机制,也不生成进行中的标题。(如果用户明确要求将结果保存到某处,按要求写入即可,这只是基本礼貌,不属于模式切换。)
  • 用户提供HYPOTHESES.md文件(路径或粘贴内容)→ 文件模式:遍历整个列表,生成QUESTIONS.md文件。
  • 情况不明 → 询问用户需求。

Single-question mode

单问题模式

  1. Understand the input. What does the hypothesis claim, who would the interviewee be, and what words do those customers use? One or two questions at most — then work.
  2. Check the hypothesis itself. A question can only be as good as the hypothesis it tests. If the input is vague ("customers hate this"), unfalsifiable, or a purchase referendum ("customers would buy X"), say so and sharpen it with the user first — a two-minute fix, not a detour. If the reframe decomposes into more than a few hypotheses, suggest the fuller path: write the complete hypothesis list first, then come back for the whole question set.
  3. Draft, grill, iterate. Offer one to three candidate questions; run the four criteria and the polite-stranger test visibly on every candidate, including any the user proposes; rewrite until one survives.
  4. Deliver. The final question in chat, with the hypothesis it tests restated and, where useful, one natural follow-up probe. No file.
  1. 理解输入内容。hypothesis的主张是什么?受访者是谁?这些客户使用什么词汇?最多问一两个问题——然后开始工作。
  2. 检查hypothesis本身。问题的质量不可能超过它所测试的hypothesis。如果输入内容模糊(“客户讨厌这个”)、无法证伪,或者是购买意向投票(“客户会购买X”),要告知用户并先和他们一起细化hypothesis——这只需两分钟,不是绕路。如果重构后的hypothesis分解为多个hypotheses,建议采用完整流程:先撰写完整的hypothesis列表,再回来生成整套问题。
  3. 起草、打磨、迭代。提供1-3个候选问题;对每个候选问题(包括用户提出的)都公开执行四个标准和礼貌陌生人测试;重写直到有一个问题通过所有检验。
  4. 交付结果。在聊天中输出最终问题,重述它所测试的hypothesis,必要时提供一个自然的后续追问问题。不生成文件。

File mode

文件模式

Phase A — Ingest. Read HYPOTHESES.md (path or pasted; default
HYPOTHESES.md
in the current directory). Also read the goal file its preamble names if available — the goals often carry the customers' vocabulary and the decisions at stake; if it's missing, proceed, but ask the user for their customers' vocabulary directly. Note any segment tags, priority rankings, or recruiting dependencies recorded there; questions inherit them. If a QUESTIONS.md already exists at the target location, read it first: an in-progress header means resume — confirm with the user, pick up at the hypothesis the header names, and don't redo finished questions. Marked complete means ask whether to revise or replace.
Phase B — Walk the hypotheses, one question per exchange. Open small: a two-or-three-sentence acknowledgment of what you read, any flags (an in-progress input file, goals with no hypotheses), and a note that some hypotheses may share questions — merges will be proposed as they come up. (Work out the likely grouping silently at ingest and record it in the file header for resumability; the chat opening stays small.) Then the loop, strictly one question at a time: when you reach hypotheses that could share a question, propose the merge and get a yes before drafting; draft the one question against the four criteria; grill it (delegated or inline); let the user react and refine — a question is settled only when the user has confirmed it, not merely when your own grill passes; write it to the file; move to the next. Never draft ahead of the conversation. Vocabulary hypotheses ("customers say X, not Y") usually get no dedicated question — asking "do you call it X?" measures comprehension, not truth. Test them by listening: every answer is a vocabulary sample. Record that as the hypothesis's coverage (coverage-by-listening, not a skip), optionally backed by one late standing probe ("I've heard people call this different things — what do you call it?"). Then:
Record to the file as you go. As soon as the first question is settled, create
QUESTIONS.md
in the same directory as the input HYPOTHESES.md (if the hypotheses were pasted and no path is known, ask where the method's files live first — default: the current directory) and write it in; after each question (or small group) is settled, append it and rewrite the status note's coverage line so the pointer is never stale — since grouping breaks numeric order, list the H-numbers covered so far and name the next group, rather than assuming contiguity. Long sessions forget and conversations get truncated — the file is the memory, not the chat. If files aren't accessible, re-emit the full current draft in a fenced block every question or two.
Phase C — Assembly critique. When every hypothesis is covered (or explicitly skipped), critique the set:
  • Coverage — every hypothesis maps to at least one question, or its skip is recorded with a reason (skips and coverage-by-listening notes live in the file's preamble). Every question names its [H-numbers]; a question with none is chitchat — cut it.
  • Fit — would this list fit a real conversation of roughly an hour? With two interview tracks, check fit per track, not combined. If it doesn't fit, group harder or cut, using the hypotheses' own priority ranking — and if the hypotheses file carries no priorities, ask the user to rank now. Don't shrink questions into yes/no compression, which destroys criterion 3.
  • Order — the list is an interview plan: rapport-easy, context- setting questions first; segment-determining questions near the top (so you know which lens to apply to everything after); specific-price questions near the end of their track. With two populations, keep one continuous Q-numbering split into labeled per-track sections — each track is a separate conversation with a separate interviewee. Reorder and renumber before finalizing.
  • One voice — terminology consistent with the customers' vocabulary throughout.
Phase D — Finalize. Remove the in-progress note, complete the preamble and Next steps, confirm the file stands alone months later, and read the final list back compactly (question → hypotheses it tests). Close with the handoff — tell the user how, not just what: run the interviews, and put each conversation on the record while it's fresh; if a debrief skill from this method's author is installed (for example Interview Debrief /
asb-interview-debrief
), name it as the way — "after each interview, run
asb-interview-debrief
with your transcript or notes and this QUESTIONS.md."
Structure:
markdown
undefined
阶段A — 导入。读取HYPOTHESES.md文件(路径或粘贴内容;默认读取当前目录下的
HYPOTHESES.md
)。如果文件前言中提到了目标文件,且该文件可用,也要读取——目标文件通常包含客户的词汇和相关决策;如果缺失,可以继续,但要直接向用户询问客户的词汇。记录文件中提到的任何细分标签、优先级排名或招募依赖项;问题会继承这些信息。如果目标位置已存在QUESTIONS.md文件,先读取该文件:如果有进行中的标题,说明需要恢复——与用户确认,从标题中提到的hypothesis开始,不要重复已完成的问题。如果标记为已完成,询问用户是要修订还是替换。
阶段B — 遍历hypotheses,每次处理一个问题。开场要简洁:用两三句话说明你读取的内容、任何需要注意的点(如进行中的输入文件、无hypotheses的目标文件),并说明部分hypotheses可能共用问题——分组建议会在处理到相关内容时提出。(导入时可以私下规划可能的分组,并记录在文件标题中以便恢复;聊天开场要保持简洁。)然后严格按照每次一个问题的循环处理:当遇到可以共用一个问题的hypotheses时,提出分组建议并在得到用户同意后再起草;根据四个标准起草一个问题;打磨问题(调用工具或自行执行);让用户反馈并细化——只有当用户确认后,问题才算定稿,而非仅仅通过你自己的打磨;将问题写入文件;处理下一个问题。绝不要提前起草问题。 词汇类hypotheses(“客户说X,而非Y”)通常不需要专门的问题——问“你称之为X吗?”测试的是理解能力,而非真相。通过倾听来测试这些hypotheses:每个答案都是词汇样本。将此记录为该hypothesis的覆盖方式(通过倾听覆盖,而非跳过),可选在后期添加一个通用追问问题(“我听到人们对这个有不同的称呼——你怎么称呼它?”)。然后:
边处理边写入文件。第一个问题定稿后,立即在输入HYPOTHESES.md文件所在的同一目录创建
QUESTIONS.md
文件(如果hypotheses是粘贴的且未知路径,先询问用户方法文件的存储位置——默认:当前目录)并写入问题;每个问题(或小组问题)定稿后,追加到文件中,并更新状态说明中的覆盖行,确保指针始终最新——由于分组会打破编号顺序,要列出已覆盖的H编号,并指明下一组hypotheses,而非假设连续编号。长时间的会话容易遗忘,对话也可能中断——文件是记忆载体,而非聊天记录。如果无法访问文件,每处理一两个问题就在代码块中重新输出当前的完整草稿。
阶段C — 整体审核。当所有hypotheses都被覆盖(或明确跳过)后,对整个问题集进行审核:
  • 覆盖范围——每个hypothesis都映射到至少一个问题,或记录了跳过的原因(跳过和通过倾听覆盖的说明放在文件前言中)。每个问题都标注了对应的[H编号];没有标注的问题是闲聊——删除它。
  • 适配性——这个列表是否适合时长约一小时的真实对话?如果有两条访谈线索,要分别检查每条线索的适配性,而非合并检查。如果不合适,要进一步分组或删减,依据hypotheses本身的优先级排名——如果hypotheses文件没有优先级,现在就请用户排序。不要将问题压缩为是非题,这会破坏第三条标准。
  • 顺序——列表是访谈计划:先问容易建立 rapport、设定上下文的问题;划分细分群体的问题放在靠前位置(这样你就知道后续所有问题该用什么视角解读);具体价格问题放在对应线索的末尾。如果有两类受访者,保持连续的Q编号,分为标注明确的分线索章节——每个线索是与不同受访者的独立对话。定稿前重新排序并编号。
  • 统一表述——全程使用与客户词汇一致的术语。
阶段D — 定稿。删除进行中说明,完善前言和后续步骤,确认文件在数月后仍能独立使用,并简洁地复述最终列表(问题→它测试的hypotheses)。最后交接工作——告诉用户怎么做,而非只说做什么:开展访谈,在对话新鲜时记录每个访谈内容;如果安装了同一作者开发的访谈复盘工具(例如Interview Debrief /
asb-interview-debrief
),要告知用户使用方法——“每次访谈后,将你的 transcript 或笔记与这个QUESTIONS.md一起输入
asb-interview-debrief
工具。”
文件结构:
markdown
undefined

Interview questions — <company / project name>

Interview questions — <company / project name>

⚠️ IN PROGRESS — this list is not yet complete. Hypotheses covered so far: <list the covered H-numbers — grouping breaks contiguity> of H<total>; not yet ordered or finalized. If you are resuming, continue with <the next group, named>. Planned grouping: <the announced grouping, so a resumed session inherits it instead of re-deriving it>. (This note is removed when the list is finalized.)
<One or two paragraphs of prose: which hypothesis list this maps to (file name), a reminder that each question is an experiment testing the bracketed hypotheses, and any segment tags or recruiting dependencies inherited from the hypotheses — including which questions belong to which interview track if there are two populations.>
⚠️ IN PROGRESS — this list is not yet complete. Hypotheses covered so far: <list the covered H-numbers — grouping breaks contiguity> of H<total>; not yet ordered or finalized. If you are resuming, continue with <the next group, named>. Planned grouping: <the announced grouping, so a resumed session inherits it instead of re-deriving it>. (This note is removed when the list is finalized.)
<One or two paragraphs of prose: which hypothesis list this maps to (file name), a reminder that each question is an experiment testing the bracketed hypotheses, and any segment tags or recruiting dependencies inherited from the hypotheses — including which questions belong to which interview track if there are two populations.>

Questions (in settlement order while in progress; reordered to

Questions (in settlement order while in progress; reordered to

interview order at finalization)
Q1. <Open-ended question, in the customers' vocabulary.> [H1, H4] → follow-ups, same story: <planned-but-optional follow-up questions, if any — each obeying every rule main questions do>
Q2. <…> [H3]
interview order at finalization)
Q1. <Open-ended question, in the customers' vocabulary.> [H1, H4] → follow-ups, same story: <planned-but-optional follow-up questions, if any — each obeying every rule main questions do>
Q2. <…> [H3]

Standing probes

Standing probes

Use these anywhere an answer surprises you or stays thin: "Can you walk me through a specific example?" · "Tell me more about that." · "Oh, I thought ___ — can you set me straight?" · "I've heard others say ___ — how do you think about it?" · Silence (people fill it with the truth).
Use these anywhere an answer surprises you or stays thin: "Can you walk me through a specific example?" · "Tell me more about that." · "Oh, I thought ___ — can you set me straight?" · "I've heard others say ___ — how do you think about it?" · Silence (people fill it with the truth).

Next steps

Next steps

<Two or three sentences of prose: run the interviews — take notes per question, next to the hypothesis each tests; chase every surprise with the standing probes, because surprise means learning; after each interview, mark which hypotheses were supported or contradicted, add new hypotheses (and questions) as they emerge; stop interviewing when the surprises stop. If the user needs people to interview: https://longform.asmartbear.com/find-customers-to-interview/>
undefined
<Two or three sentences of prose: run the interviews — take notes per question, next to the hypothesis each tests; chase every surprise with the standing probes, because surprise means learning; after each interview, mark which hypotheses were supported or contradicted, add new hypotheses (and questions) as they emerge; stop interviewing when the surprises stop. If the user needs people to interview: https://longform.asmartbear.com/find-customers-to-interview/>
undefined

Refusal conditions

拒绝场景

  • No hypothesis. "Just give me interview questions for my startup" skips two steps of the method: questions test hypotheses, hypotheses answer goals. Explain the order, then offer the on-ramp — capture one or two quick hypotheses in chat right now (what do they believe that an interview could disprove?) and craft questions for those. Don't silently generate a generic question list; that's the chitchat this method exists to replace.
  • Referendum input. "Would customers buy X?" cannot be given an unbiased question — every phrasing of it leads the witness. This is the quick-fix refusal, not a hard stop: reframe the hypothesis with the user on the spot (the pain X addresses, what they've paid for relief before, how they cope today), then write questions for the reframed version.
  • "What will they answer?" Predicting or simulating the customers' answers defeats the exercise — decline; the interviews get that job.
  • Survey design. These are conversation questions, designed to surface what you didn't know to ask; a fixed-response questionnaire can't chase surprise. Say so if the user wants a survey — the questions may still seed one, but the method assumes conversations.
  • A question that violates the four criteria. Refuse to finalize or record it, no matter how the user insists. Explain which criterion it fails and what the failure costs (polite yeses masquerading as validation), offer the compliant rewrite, and keep working the rewrite with them — the refusal is of the broken question, never of the user's underlying hypothesis, which always has a compliant question available.
  • Conducting the interview. Role-playing the interview or analyzing transcripts is beyond this skill's scope; the Next steps section of QUESTIONS.md says how to run and iterate the real thing.
  • 无hypothesis。“给我一些创业公司的访谈问题”跳过了该方法的两个步骤:问题测试hypotheses,hypotheses回答目标需求。解释流程顺序,然后提供入门路径——现在就在聊天中捕捉一两个快速hypotheses(他们相信什么可以通过访谈证伪?),并为这些hypotheses设计问题。不要默默生成通用问题列表;这正是该方法要取代的闲聊式问题。
  • 投票类输入。“客户会购买X吗?”无法设计出无偏见的问题——任何表述都会诱导受访者。这是快速拒绝,而非彻底终止:当场与用户重构hypothesis(X解决的痛点、他们过去为缓解痛点支付的费用、当前的应对方式),然后为重构后的hypothesis撰写问题。
  • **“他们会怎么回答?”**预测或模拟客户的回答违背了流程的目的——拒绝;访谈会完成这项工作。
  • 调查问卷设计。这些是对话式问题,旨在挖掘你没想到的信息;固定选项的问卷无法追踪意外发现。如果用户想要调查问卷,要说明这一点——这些问题可能可以作为问卷的基础,但该方法基于对话式访谈。
  • 违反四个标准的问题。无论用户如何坚持,都拒绝定稿或记录该问题。说明它违反了哪个标准,以及这种违规的代价(礼貌性的“是”伪装成客户验证),提供符合标准的改写版本,并继续与用户打磨改写后的问题——拒绝的是有缺陷的问题,而非用户的核心hypothesis,每个hypothesis都有符合标准的问题可用。
  • 开展访谈。角色扮演访谈或分析 transcript 超出了该工具的范围;QUESTIONS.md的后续步骤部分说明了如何开展和迭代真实访谈。