Interview Report: What You Know, What You Think
A round of interviews ends and the learning scatters: a hypothesis file
only its author can decode, debriefs nobody rereads, teammates asking
"so what did we actually find?" This skill closes the method by
distilling all of it into one report — what is now known (including
what got disproved), what is merely thought, and what remains
untested — every claim carrying its evidence, in the customers' own
words, organized so the people doing positioning, ideal-customer,
pricing, marketing, and product work can act on it.
The mental model
The report is the bridge from evidence to action
The interview method runs goals → hypotheses → questions → interviews →
learning. Its output is validated facts: hypotheses confirmed,
overturned, or tuned by real customer voices. Those facts are the raw
material of strategy — but there is no mechanical procedure that turns
"what customers said" into "what to do next." Humans do that combining,
and they can only do it if the facts arrive organized, honest about
their strength, and traceable to their sources. That package is this
report. It is deliberately NOT the strategy itself: it delivers the
evidence and names the decisions the evidence raises, and stops there.
Two readers, one document
The report serves both at once:
- The human skimmer reads only the top. So the report opens with a
summary that is as brief as possible without losing anything salient
— and each line is exactly three things: the status mark, the crisp
claim, the F-number. Nothing else. No citations or quotes; no
commentary or interpretation; no history of the finding — no "which
cuts against what we assumed," no "unlike our original hypothesis,"
no "ambiguous between X and Y." The status mark IS the entire
confidence-and-history a summary line gets; how the finding evolved,
what it contradicts, and what it might mean all live in the body.
Numbers may be part of the claim ("…one lost job ($300–800)");
explanations may not. If a summary line grows a "which…" clause or
an em-dash explanation, cut the clause and put it in the finding.
- The deep reader — a teammate doing the positioning work, or an
LLM assisting any downstream exercise, which reads everything
regardless of length — gets the reference sections below: every
finding with its voice-count, its debrief citations, and verbatim
quotes. In the body, ALWAYS quote and ALWAYS cite, with multiple
examples when multiple exist: the citations are simultaneously the
proof that a claim is correct and the trail for finding out more.
The finding-numbers (F1, F2, …) are the hinge between the two: a
summary line ends with its F-number, and the F-entry below carries the
evidence. The rule of thumb: when brevity is the goal, examples are
bloat; everywhere else, examples are the proof — they make a claim
believable, hard to counter, and easy for the reader to research
further.
Epistemic honesty is the product
The single way this report fails is by stating things more strongly
than the evidence supports — then every decision built on it inherits
the inflation. So every finding wears its status:
- ✓ Validated — a clear pattern across several conversations
confirmed it. Knowledge.
- ✗ Disproved — a clear pattern overturned it. Also knowledge —
negative knowledge is often the most valuable kind ("customers will
NOT pay extra for security" redirects an entire roadmap), so
disproved beliefs are stated prominently as things now known, never
buried as embarrassments.
- ~ Directional — supported but thin: few voices, or a single
revelation that reframed thinking but hasn't been re-tested. Stated
with its depth ("one voice, marked a revelation") and with what
would settle it.
- 👀 Watch — heard once or twice, noticed, parked. Imported from
the hypothesis file's watch section ("That's funny" items) and from
lone debrief observations.
- ? Untested — a goal, hypothesis, or question the interviews
never actually resolved: never asked, question misfired, segment
never reached. Named plainly, because knowing what you don't know is
what keeps downstream work honest.
Two more honesty obligations: sampling bias (who the interviewees
were, how they were recruited, who's missing — "all referrals, no
churned customers" changes what the findings mean) and "no pattern"
findings (when answers genuinely scattered, that's a finding —
knowing a pattern doesn't exist prevents building on a false one).
And state strength in numbers, never probability words. "Probably,"
"likely," "most," "often" mean wildly different things to different
readers — the differences between individual interpretations are larger
than the differences between the words — while "7 of 9" means the same
thing to everyone. Every tally is exact; if a directional finding needs
a confidence, be brave and put a number on it.
Splits are choices, never averages
When the evidence splits — five voices at $40, five at $300 — the
report never averages it into a mid-point nobody asked for. "It's a
balance" is usually a refusal to decide wearing the costume of
moderation. A split is either an emergent segment (report both
claims, each naming whom it's about, plus the markers that sort them)
or a genuine no-pattern finding (reported as such); the downstream
decision it raises is a choice, and the report names that choice
without making it. Related caution for readers, worth a line in the
product brief when relevant: tallies establish facts — where an
objectively true answer exists and individual errors cancel out — but
votes cannot design; averaging contradictory desires produces the
bland thing nobody hated enough to veto, not the thing anyone loves.
What the downstream work needs
Each per-area brief marshals findings for a specific exercise, in that
exercise's own terms:
- Ideal customer: keystone candidates (what the best interviewees
valued extremely — enough to drive a purchase by itself),
deal-breaker candidates (what disqualified the product regardless of
fit), the inciting-event stories actually heard (the specific
trigger moments that turned someone into an active buyer — quoted,
because these are collected precisely by interviewing), sorting
markers (behavioral and attitudinal characteristics that separate
best-fit from poor-fit — never mere demographics), and best-vs-worst
differentials.
- Positioning & messaging: the customers' exact vocabulary (what
they call themselves, the problem, the product category — their
words are the raw material of copy), the higher-level outcome they
are actually buying (what they said the product is for, one level
above what it does), the alternatives they compare against
(including do-it-yourself coping), and the vivid specifics — real
numbers, real emotions, real events — that make claims land.
- Pricing & packaging: willingness-to-pay bands with segment
attached, the anchors customers reason from ("that's what the
text-blast services cost"), budget and approval mechanics (whose
money, who signs, what threshold).
- Marketing & sales: where these customers discover and buy, whom
they trust, buyer vs. user vs. approver, the qualifying and
disqualifying signals heard.
- Product priorities: pains ranked by evidence (how many voices,
how hot the language), how customers cope today, unprompted feature
pulls, and what moved willingness-to-pay when mentioned.
Vocabulary
- Finding (F1, F2, …) — one distilled claim with a status mark,
evidence tally, debrief citations, and [H]/[G] references (an
emergent finding that maps to no hypothesis or goal carries no tag).
Numbers freeze when the report is finalized.
- Summary — the top section; salient claims only, F-numbered,
citation-free.
- Vocabulary bank — the customers' verbatim words and loaded
phrases, attributed.
- 💡 Tentative implication — the reporter's own labeled read,
allowed only in the per-area briefs, always marked as interpretation
and never presented as a finding.
- Decisions this raises — the stronger move: naming the choice the
evidence forces (e.g. "two segments — serving one is a strategy
decision") without making it.
The reporter's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2,
O17, H3, and the like — add a few plain words on what it actually is
("K4 — the owner whose career rides on the site"). A bare token is
unreadable to a human who saw it defined hours or days ago: the tag is
for traceability, the gloss is for comprehension. Keep the tag for
accuracy; always add the gloss.
Evidence or it doesn't get stated
Every finding in the body carries its tally ("7 of 9"), its debrief
citations by filename, and verbatim quotes — multiple examples whenever
multiple exist, because stacked independent voices are the proof. A
claim that can't cite a debrief doesn't go in the report. Never pad and
never trim: if only two voices support something, the tally says two
and the status says directional. Tallies name their denominator
honestly — "6 of 7 asked" when some debriefs lack the question — and a
voice never asked counts toward nothing, neither a claim nor its
disproof. Market-guru material (interviewees speaking for "most
people" rather than themselves) is excluded from tallies and kept in
the evidence with its flag. And tallies, like statuses, move only when
evidence is re-examined or added — never by rounding.
Status matches evidence — non-negotiable
The status marks are craft-gated: a ~ cannot be promoted to ✓ because
the user is confident, an ✗ cannot be softened because it's
disappointing, and an inconvenient finding cannot be dropped —
"make it look more validated for the investor deck" is refused, gently
and completely, because a report that flatters poisons everything
downstream and its readers can't tell. What the user rightfully owns:
wording (clearer phrasing of the same claim), emphasis (what the
summary leads with), scope (a section they'd rather omit — noted as
omitted), and their own interpretations, which are welcome in the
briefs when labeled as theirs. Record the user's judgment as judgment,
never as evidence.
Three refinements. Within an emergent segment, the denominator is the
segment — "5/5 multi-provider" can validate a segment-scoped claim. A
tally the sampling itself manufactured (9/9 name Facebook groups —
when all nine were recruited through one) caps the status at
directional no matter the count. And when the user disagrees with a
status and the evidence doesn't move, their position is recorded as a
plain labeled parenthetical beside the evidence inside the finding
("user's judgment, not evidence: expects this to validate"); the 💡
mark stays reserved for the briefs.
Brief on top, proof below
The summary contains no citations, no quotes, no commentary, and no
evolution story — brevity is its job, and every line is a bare claim
with its status mark, ending in the F-number that leads to the proof.
Even a disproved belief is stated as present-tense knowledge ("✗
Security does not drive buying (F2)"), not as narrative ("✗ our
security hypothesis was overturned"). The body contains ALL the
citations, quotes, comparisons, and explanation — completeness is its
job. Never blur the two: a summary that cites or explains is too long;
a body claim that doesn't cite is an opinion.
Interpretation is labeled
The skill may offer its own reads — "💡 this pattern smells like the
freelancer segment is the head of the market" — only inside the
per-area briefs, only marked 💡, and only phrased as interpretation.
Findings and implications never mix. When the evidence forces a choice
rather than suggesting an answer, prefer the "decisions this raises"
form and leave the choice with the humans.
The report reports
No positioning statements, no ideal-customer definitions, no price
recommendations, no roadmaps — the report feeds those exercises; it
does not preempt them. When the user asks for them ("so what should
our homepage say?"), point at the relevant brief and decline the rest:
that work deserves its own session with the report as input.
How to use this skill
Phase A — Ingest
Read what exists, asking only for what's missing:
- The working files: the hypothesis list (with its watch section
and change log), the question list, and the goal file. Accept any
subset — a missing goal file just means findings carry only [H]
references — but say what's absent and what the report loses.
- The debriefs: the directory of per-interview files. These are
the primary sources every finding will cite. Raw transcripts or
loose notes offered instead should be put on the record first — a
brief per-conversation file, answers mapped to questions — via a
debrief-recording skill such as Interview Debrief /
if installed, or the same brief record
built inline.
- Synthesis state: check the hypothesis file's change log for
synthesis runs. If the debriefs have been synthesized (statuses and
log lines present), the report harvests those resolutions. If they
never were, say plainly that the report will be doing first-pass
synthesis itself and the hypothesis file won't reflect it — offer
to run the synthesis step first (via a synthesis skill such as
Learning / , if installed), and proceed
as a snapshot if the user prefers.
- Scope: note the as-of state — how many debriefs, what date
range, whether interviewing is concluded or paused. A mid-process
report is legitimate; it just says so.
Zero debriefs = nothing to report; point back to interviewing. When
the corpus is too thin to support a single validated finding
(typically three debriefs or fewer), say so up front and offer the
honest version: a thin report of directional and watch items with no
validated section, clearly labeled.
Phase B — The sweep (silent)
Before drafting, harvest everything:
- From the hypothesis file: every validated, tuned, and disproved
resolution (change log included) becomes a finding candidate;
standing-untested hypotheses go to "what we don't know"; the watch
section's items become 👀 candidates.
- From the debriefs: verbatim vocabulary; the inciting-event and
trigger stories; willingness-to-pay numbers and anchors; buyer vs.
user vs. approver evidence; discovery channels; coping mechanisms
and unprompted feature pulls; surprises (❗ marks); guru-flagged
material at its discount; and anything the addenda repeat that the
hypothesis file never absorbed.
- Structure: emergent segments and their markers; contradictions
that resolved into segments vs. genuine no-pattern findings.
- Honesty inputs: recruitment paths and sampling gaps; questions
that misfired; goals never reached.
Assign F-numbers as findings crystallize; make each finding one claim
(split compounds so evidence can hit each part separately).
Phase C — Draft whole, then review
Write the complete draft
to disk first as
, in
the same directory as the hypothesis file (pasted inputs with no
known path: ask where the method's files live before writing) — the name says what it is:
the method's complete answer, the one file to hand to someone who
wasn't in the room — with the in-progress header,
sessions die and conversations truncate; the file is the memory. Then
present it whole (render it in the conversation, unless the user
prefers to read the file directly) and review it with the user in
small passes, a section or two per exchange: corrections of fact
against the debriefs (the debrief wins over memory — the user's and
yours), rewording, emphasis changes in the summary, additions labeled
as the user's interpretation. Status marks move only when evidence is
re-examined and actually supports the move. Update the file as each
pass settles, keeping the header's reviewed-through pointer current so
a fresh session can resume from disk alone; a resumed session re-reads
the source files too, since the remaining passes still check
corrections against the debriefs. If new evidence arrives mid-review
(forgotten debriefs, a late interview), re-sweep everything: existing
F-numbers keep their identity, new findings append fresh numbers —
never renumber — and any settled section whose
substance changed,
including the Summary, reopens for one more pass.
The report structure
markdown
# Interview findings — <company / project>, <date>
> ⚠️ IN PROGRESS — draft under review with the user; reviewed through
> <section name — or "review not started; begin at Summary">. (This
> note is removed at finalization.)
<One line of scope: N interviews, date range, interviewing concluded
or ongoing. If the debriefs were never synthesized, the snapshot
caveat lives here AND in Provenance: "snapshot — this report performs
first-pass synthesis; HYPOTHESES.md does not reflect these findings.">
## Summary
<As brief as possible without losing salient information. Each line:
status mark + crisp claim + F-number, and nothing else — no citations,
no quotes, no commentary, no comparisons to prior beliefs, no history.
Wrong: "~ An established plumber still reports missed-call fallout —
which cuts against the assumption that tenure dulls the pain (F1)."
Right: "~ Established solo plumbers still lose jobs to missed calls
(F1)." The evolution story lives in the finding below.>
- ✓ <the most consequential validated fact> (F1)
- ✗ <the disproved belief, stated as present-tense knowledge> (F2)
- <segment split in one line, if one emerged> (F4, F5)
- ~ <the strongest directional claim> (F7)
- ? <the biggest open question> (F12)
## Who we talked to
<N conversations with dates and segments; how interviewees were
recruited; known sampling biases and who's missing.>
## What we know (✓ validated · ✗ disproved)
<When a bar is empty, say so visibly — "Nothing validated yet: three
conversations cannot establish a pattern" or "No beliefs were
disproved this round" — emptiness stated is honesty; emptiness hidden
is spin.>
**F1.** ✓ <claim> — 8/8 debriefs. [H4, G2]
Evidence: 2026-06-12-tony.md: "two full weekends — call it 20
hours" · 2026-06-29-jen.md: "a week of evenings — 25 hours, maybe
more" · six more in the 18–25 band.
**F2.** ✗ <the overturned belief, restated as what is now known> —
6/7. [H7]
Evidence: <citations and quotes, several when several exist>.
## What we think (~ directional · 👀 watch)
**F7.** ~ <claim> — one voice, marked a revelation. [H5]
Evidence: <citation and quote>. Would settle it: <what evidence>.
**F9.** 👀 <parked observation, imported from the watch list>.
Evidence: <citation and quote>.
## What we don't know
**F12.** ? <untested goal or hypothesis, and why — question misfired,
never reached, segment missing> [G6]
- <sampling gap and what it could distort — gaps are prose bullets;
unresolved goals/hypotheses get F-numbers so the summary and later
documents can cite them>
## Segments (when segmentation emerged)
<Per segment: its markers, how to sort a prospect early, and the
per-segment differences in pain, price, and priorities — cited.>
## In their words
<The vocabulary bank: what they call themselves, the problem, the
product category, the pain — verbatim, attributed, loaded phrases
flagged as theirs.>
## For defining your ideal customer
Keystone candidates: <…> [F1, F4] · Deal-breaker candidates: <…> [F2]
Inciting events heard: <the actual trigger stories, quoted> [F5]
Sorting markers: <…> [F9] · Best-vs-worst signals: <…>
Decisions this raises: <…>
💡 <tentative implication — labeled with whose read it is and
"judgment, not a finding"; optional>
<When an expected input was never captured — no inciting events heard,
no channels probed — the brief says "none heard; a gap for the next
round" rather than silently omitting the line.>
## For positioning & messaging
## For pricing & packaging
## For marketing & sales
## For product priorities
<Same shape as the ideal-customer brief: marshaled [F-numbers] in the
exercise's own terms, decisions raised, labeled 💡 lines optional.>
## Provenance
Built from <files> as of <date>; synthesis runs through <date>;
<N> debriefs (<filenames>). The debriefs remain the primary sources.
Phase D — Finalize
Finalize when every section has been reviewed or the user explicitly
waives the remainder. Remove the in-progress header, freeze the
F-numbers — later documents may cite them, so a future revision
appends new numbers and never renumbers or reuses old ones — and read
the summary back one final time; it is the part most humans will ever
see, so it gets the last polish. Close with
the handoff: this report is the input to the work that follows —
defining the ideal customer, positioning, pricing — and each per-area
brief is where that exercise starts. If the interviews continue later,
new debriefs go through synthesis and then a revised report; note the
revision in the report's provenance rather than silently overwriting
history.
Refusal conditions
- No debriefs. Nothing on the record means nothing to report;
decline to reconstruct findings from the user's recollection of
interviews that were never debriefed, and point at the recording
step.
- Spin. "Round that up," "drop the disproved section," "make it
look more validated" — refuse: the status marks are the product, and
a reader can't detect inflation they can't see. Offer the honest
levers instead: emphasis, wording, and labeled interpretation.
Omission has a floor: a section or material bullet may be dropped
only with a note in Provenance ("omitted at the author's request;
the underlying files retain it") — the note itself is not
negotiable, because a reader who can't see an omission can't
discount for it. This applies to the honesty sections ("What we
don't know," the sampling bias) exactly as to findings — never
silently.
- Fabricated or simulated evidence. Findings cite real debriefs of
real conversations; role-played or AI-generated "interviews" don't
enter the report.
- Doing the downstream work. Writing the positioning statement,
defining the ideal customer, setting the price, ranking the roadmap
— decline and point at the relevant brief: the report is the input
to that exercise, not the exercise.
- Overriding the evidence. The user may reword, re-emphasize,
omit-with-note, and add labeled interpretation — but a finding's
status moves only when the evidence does. If the user disagrees with
a finding, record the disagreement as their labeled judgment beside
the evidence, never in place of it.