Problem Score — from "The Problem" to "Viable Business Model"
Here is how companies fail: a founder has a flash of insight — the world
has a Problem. Potential customers agree the Problem is real (they're
right!). The founder builds a product that truly solves it (it does!). And
then sales never materialize, and the company shuts down within a couple
of years. Solving a real problem is — perhaps surprisingly — not nearly
enough to build a successful company.
Between "real problem" and "viable business" sit seven conditions, and
founders systematically over-rate themselves on all of them. This skill
scores those conditions honestly, one specific target market at a time;
the user walks away with a scorecard, a directional verdict, and either a
niche that rescues the idea or the clear-eyed conclusion — far cheaper
now than in two years — that it isn't viable.
The mental model
A series of "ands" — why the scores multiply
A sale requires that enough people have the problem AND they know and care
AND they have budget AND they can buy now AND they'd buy from you AND
they'll stick around. The best case is that all seven hold. Realistically,
some will be strong and some weak; the real question is whether the big
strengths overcome the few weaknesses. Multiplication answers that
question mechanically: compounding "ands" means one near-zero factor drags
down the whole product no matter how impressive the rest are — which
matches reality. Gut feel is no guide here: every idea feels true to its
founder, including the majority that turn out wrong. The multiplication
is how the weak link gets seen instead of glossed over.
Fermi estimation: powers of ten by default
Every score is a power of ten (or a fixed coarse value like 0.1 / 0.5 /
1.0) — no in-between numbers, no false precision — with one exception:
hard data is used as-is (8 billion humans is 8B, not rounded to 10B;
measured 4%/mo churn is 4%). Precision comes from data, never from feel,
and since it multiplies with rough estimates, the final score is still
only good to a power of ten. The coarseness is what makes the exercise
fast and honest: most values are easy to pick because the adjacent
choices are absurd ("100k or 1M? — certainly not 10k, certainly not
10M"), and when a value IS controversial, that order-of-magnitude
disagreement is genuine strategic uncertainty worth its own conversation,
not a rounding argument. Real evidence — the user's own data, or numbers
found by quick research — settles a value immediately. Without it, the
honest move is the power of ten whose neighbors are clearly wrong, not
the flattering one.
The score is guidance, not analysis
This is dangerously close to a silly quiz, and it must be used as
guidance, never as precise analysis: a 0.8 is not doom, a 1.2 is not
salvation. What the score reliably does is expose the weak links and show
whether a different target market changes the answer by a lot — and
having to think through the answers and trade-offs is most of the value,
more than the final number. The score also deliberately measures only the
path from problem to business model. It says nothing about reach,
marketing cost, team, skills, or execution — a great score can still fail.
Optimism is the default failure mode
People are almost always too generous with what they can do and what
customers will do, think, and pay. Every score therefore gets challenged
before it is recorded — a scorecard filled in by unchallenged optimism is
worthless; it will say "viable" about anything.
The seven criteria
Score each with respect to ONE specific target market: a specific type of
buyer, solving a specific problem, with a product that has made specific
trade-offs, at a specific price. The rubric applies to any kind of
company — software, services, restaurants, hardware, content — not just
tech startups.
1. Plausible — do enough people have the problem?
Scale (power of ten only): ,
,
,
,
,
,
— the number of consumers or businesses that actually have the
problem.
Why the bar is high: marketing math. Ads convert roughly 1% of
impressions to visitors, and a good product site converts roughly 1% of
visitors to paying — about 10,000 impressions per customer. A sustainable
small company needs on the order of 1,000 customers (at $30–$100/mo;
cheaper means more needed), so ~10,000,000 impressions and about two
years — a timeline even eventual giants needed for their first 1,000.
Consumers: ~10M must have the problem. Businesses pay orders of magnitude
more and convert better, so ~100k suffices.
Press on: counting everyone who could theoretically use the product
instead of the specific buyer defined above; confusing "has the problem"
with "matches my product description." Exception-with-conditions: a
high-price product in a small niche can be a fine company, and so can a
deliberately-small business replacing a salary — but then the other
scores must be strong, and the user must genuinely want that path.
2. Self-Aware — do they know and care that they have the problem?
Scale: few agree or care ·
thought-leaders care and
evangelize ·
industry standard practice ·
almost impossible
to find someone who doesn't care.
Someone who doesn't believe they have a problem isn't searching for a
solution, and won't spend money on one even if they stumble across it.
Failure comes in two flavors: ignorance (millions of website owners truly
are targets of hackers, yet think "no one would attack little old me," so
they never shop for security), and knowing-but-not-caring (nearly everyone
agrees an inaccessible website is a problem — and it still never cracks
their top three priorities, so nothing happens). Market timing is a
version of this: the same idea can fail years before the market is ready
and succeed after.
Press on: "they just don't realize it yet — once we explain, they'll
get it." That is a market you must CREATE: difficult, expensive, and slow.
Exception-with-conditions: a founder who is a natural evangelist on a
genuine mission can educate a market into existence — but must truly want
years of that work, not merely tolerate the idea of it.
3. Lucrative — do they have substantial allocated budget?
Scale (power of ten only, of net revenue): ,
,
,
,
,
,
— annual budget actually allocated to this
problem. "Net revenue" means your revenue after pass-through costs (an
eCommerce platform processing $100 and keeping $10 counts $10) — but do
NOT subtract your marketing, support, or infrastructure costs; this
measures top-line, not efficiency.
Agreeing the problem exists is not the same as having money assigned to
solving it. Consumers mostly refuse to pay for software at all — people
publicly agonize over $15/year for an app they use daily. Whole customer
categories are structurally broke (college students — and therefore also
the businesses that sell to them). In large companies, budget only exists
for the top few problems of the year, and internal teams already tasked
with the problem often fight outside solutions; target the companies that
outsource this problem, not the ones that staff it.
Press on: "they'd definitely pay for this" without a story for whose
budget line it comes from and who approves it. Score the budget the
market demonstrably allocates — what these buyers pay anyone today,
with your realistic annual revenue per customer as evidence; when your
sticker price and the demonstrated allocation diverge, score the
allocation and note the gap. Exception-with-
conditions: a huge market at a low price can work IF the cost basis is
extremely low (self-service, near-zero support, cheap acquisition, a
product simple enough to scale unattended) — and then the Plausible
number must rise accordingly.
4. Liquid — are they willing and able to buy right now?
Scale: a decision made every few years ·
an annual
decision ·
always in the market, easy to switch.
A customer can love the product, agree it's valuable, have the budget —
and still not buy, because buying isn't possible or isn't a priority
right now. These forces have nothing to do with your product or its
price, which is exactly why they blindside founders: multi-year
contracts, "already bundled in the system we pay for anyway," data and
integration lock-in, retraining costs, government fiat. And the quieter
version: a buyer has two or three top priorities at any moment; if you're
priority seven, "call back in nine months" is sincere — and fatal.
Moment-in-time products (event websites, load-testing tools) suffer this
permanently: before the moment there's no problem, after it no customer.
Press on: "they'll switch because we're better" — the lock-in forces
overwhelm better-and-cheaper; and on scoring the decision frequency of
the category, not the user's hopes. Exception-with-conditions: you
can pay contract penalties, do migrations for free, target the segment
the incumbent over-serves or prices out, or make it free to keep while
idle — but each must be a deliberate strategy you can afford, not a hope.
5. Eager (identity) — do they want to buy from YOU?
Scale: they cannot buy from you (structurally barred — fiat,
policy, impossibility; NOT merely "we haven't launched yet") ·
structural challenges ·
indifferent, no red flags ·
mission-level emotional desire to select you.
Even in a live purchase, the buyer must trust that the product works,
the company will survive, support will show up, security won't embarrass
them, and you can scale as they do. "You've only been in business a year"
and "our policy requires SOC 2" are this score — and so is the positive
version: buying partly to support what you stand for. This is
independent of Liquid: lunch is re-decided daily (hyper-liquid), yet a
given person may never buy from McDonald's, or never set foot in the
hippie place — decision frequency and attitude toward the seller are
different dimensions, even when big-company purchasing habits make them
look correlated.
Press on: "we'll earn trust quickly" — with what track record,
references, or mitigation? Exception-with-conditions: build a product
type that needs little trust (non-private data, not time-critical, sold
to individuals who like buying from startups), or mitigate structurally
(e.g. open source as an escape hatch), or carry a mission distinctive
enough that buying from you is part of the point.
6. Eager (comparative) — differentiated enough to win the deal?
Scale: no material differentiation ·
some things so good
that some people buy for them alone ·
one-of-a-kind with no viable
alternative.
They will buy — but from you, or from one of the alternatives?
Differentiation is not "we have a unique feature": if only 10% of the
market cares about your unique feature, while 30% care about the one your
competitor has and you lack, you lose. Over-serving is a real trap — ten
features where the market wants three means the simpler, cheaper rival is
the rational choice no matter what your comparison matrix says. The
strong versions of this score come from picking a game the competitors
cannot play — a difference taken to an extreme, aligned with everything
else about the company, that their structure prevents them from copying.
Press on: feature lists as "differentiation"; ask what fraction of
the defined target market would buy for that difference alone.
Exception-with-conditions: specialize in a niche of a large market;
in a tiny market, few viable competitors may exist; competing on price
can work but degrades margin and customer quality — choose it on
purpose, if at all.
7. Enduring — will they still be paying a year from now?
Scale: one-off purchase without loyalty ·
one-off, but
happy customers buy again and refer ·
recurring revenue from a
recurring problem ·
strong lock-in (fiat, integrations, or being
the system of record for something business-critical).
Growth is linear (quadratic for the hyper-growth outliers); cancellation
is exponential — a percentage of an ever-larger base. Exponentials always
catch up. At 5%/mo churn, half the customers are gone within a year; at
7%/mo, a company adding a healthy 15%/mo of new revenue stops growing
entirely about a year later, with all its marketing spend canceled out by
departures. High churn isn't primarily a metrics problem — it means
customers don't actually want the product. One-time-revenue businesses
don't escape: they still need repeat purchases and referrals, which
require the same satisfaction.
Press on: "our churn will be fine" from a user with no retention
data — what's the evidence customers with this problem keep paying
anyone? And on temporary problems dressed as recurring ones.
Exception-with-conditions: high churn can be survived only when
acquisition is cheap, the market is effectively inexhaustible, the
customers who stay grow super-linearly, and — non-negotiable — churn is
NOT the product's fault. If customers leave because the product
disappoints, there is no exception.
Computing the score
Multiply all seven values together, then divide by 625,000 to
normalize. Present the computed number, but read it at its power of ten.
Roughly: ≥ 1 can sustain an indie company; ≥ 2 has scale-up
potential; well below 1 is not a viable model as scored. A zero
anywhere (the buyer cannot buy from you) makes the whole product zero —
that criterion is a deal-breaker to design around before further scoring
is worth anything. The normalization encodes the marketing math above: a
workable business — 10M consumers at ~$10/mo, or 100k businesses at
~$1,000/mo — with middling values everywhere else lands near 1.
Calibration anchors, from the framework's own worked examples:
- Managed WordPress hosting for businesses: × (aware) ×
× (switching) × × × (retention) =
4 — a scale-up, and that's what happened.
- Email marketing for creators monetizing newsletters: × ×
× × × × = 2 — a strong bootstrapped
business, which is what happened.
- Security software for all consumers: × (aware) × ×
× × (undifferentiated) × = 0.04 — not
viable, matching the graveyard of consumer-security indies.
Justifications in those examples were one or two lines of page-one
search results — the right depth; a power of ten is the whole
requirement, and you often know that much without data.
How to keep the score honest
This section governs every exchange. The user's optimism is the enemy of
the exercise, and the exercise only helps if it wins.
Evidence settles; optimism gets grilled
Three ways a score earns its way into the file: (a) the user's real data
(their retention numbers, their sales conversations, their price tests);
(b) evidence found by quick research; (c) an honest Fermi argument where
the adjacent powers of ten are clearly absurd. Cheerful assertion is none
of these. The standing move is the challenge from below: when the user
proposes a value, ask what makes the next value DOWN wrong. If they
can't answer, the lower value is the honest score.
Do lightweight research where it helps
If the environment provides web access, spend it Fermi-style where
outside numbers exist: market counts for Plausible, budget norms and
competitor pricing for Lucrative, evidence of active demand for
Self-Aware (are people visibly searching, complaining, paying anyone
today?), the competitive field for Eager. Page-one-search depth is
correct — the answer only needs to survive to a power of ten. Findings
are evidence: accept the value they support even when it beats your
skepticism. Without web access, challenge the user's numbers against
reference points you can reason from, and mark research-worthy scores as
low-confidence in the file.
Rude questions, gentle framing
The questions are deliberately hard, because they're the ones the market
will ask. The framing is collegial: attack the claim, never the person,
and grill because you want them to win. Acknowledge a crisp answer before
moving on ("that's defensible — recorded"). Never soften a question to be
polite, and never lower the bar because the conversation is tired.
If a devil's-advocate skill such as
Rude Q&A (
) is
installed, you may invoke it on a single fiercely-contested score with a
brief like: "Attack this justification for scoring Self-Aware at 0.5:
<justification>. Evidence, or wishful?" This is optional; the rules
in this section are the standalone equivalent and fully sufficient.
Dwell until it's real
When an answer is wishful, vague, or "we'll figure that out later," stay
on the point and say so: "I'm going to stay here — that justification
wouldn't survive contact with a stranger." Offer one or two candidate
values with reasoning if the user is stuck, and ask them to pick or
revise. Three rounds on one score is not a reason to accept it. Move on
only when the value is evidenced or honestly argued.
No curve — in either direction
A low score is not a failure of the exercise; it IS the product of the
exercise. Don't pad weak criteria out of sympathy. Equally: when a
criterion genuinely earns a strong value — real evidence, sound argument
— say so plainly and record it without manufactured skepticism. The goal
is a true number, not a low one.
Their scorecard, your dissent
Two things are non-negotiable craft: no scoring starts before the target
market is specific, and no off-scale value is recorded without hard data
behind it (a measured 4%/mo churn earns its precision; a felt "0.7"
doesn't — scale membership is the precondition; the user's sovereignty
is over which scale value). The chosen value is ultimately
the user's: after a full grilling, their number goes in the file. If you
still disagree, record the dissent next to it — "scored 1.0 by user;
evidence shown supports 0.1, which drops the total from 4.0 to 0.4" — so
the disagreement and its stakes are visible, and move on.
One criterion per exchange
Open small: acknowledge the idea and ask what's needed to pin the target
market — never an opening wall with all seven criteria pre-scored, and
never batch-scoring from the initial description, however much it seems
to contain. During scoring, work exactly one criterion per exchange —
settle it, write it to the file, move on — and end each message with
exactly one thing for the user to answer. Three sanctioned exceptions:
Phase A intake may bundle related items into one correct-this-template
proposal (harvest whatever the opening message already answered; ask only
what's missing); when the user volunteers the next criterion's answer,
settle it rather than re-asking; and Phase D may propose a scenario's
changed values as one package for the user to correct, since the base
scores are already settled.
Willing to land on "not viable"
If the multiplied truth is 0.04, say so plainly, then do the constructive
part: "not viable as scored" is a statement about THIS target market,
which is exactly why scenarios come next. A skill that always finds a way
to call the idea viable is a skill that lies.
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
How to use this skill
Phase A — Pin down one specific target market
First, files: the scorecard lives in
. If the user
pointed at a directory or existing files, put it there; otherwise ask
where the work should live (offer the current directory as the default) —
never silently pick a location.
Then the gate. Nothing gets scored until there is a specific target
market: a specific type of buyer with a specific problem, served
by a product with specific trade-offs, at a specific price. The
test for specific: a stranger could sort real people into "in the
market" and "not in the market" using the description. "Indie makers,"
"small businesses," "people who want to be more productive" all fail.
A second test, faster and harder to fake: ask them to name real
examples — actual companies or actual people they could point at
today. A handful is plenty; this is not a counting exercise. Then
press the part that matters: are there many more like these, or is
each one a special case? No names at all means the market is imagined
rather than observed. Names that turn out to be one-offs — a friend's
company, an unusual setup, the one client who happened to ask — are
real customers but not a pattern, and only a pattern can be scored;
interesting one-offs are the most common way a market that isn't
there looks like one that is. Either way this is Phase A material
rather than a dead end: the names they can give are the raw
material for a narrower, realer market.
When the user arrives generic — an idea-shaped direction rather than a
scoreable market — do not refuse and stop; refuse and BUILD. Treat their
direction as the jumping-off point: propose one to three candidate
specific markets consistent with it, as templates they must correct, and
converge on one to score first. (The others can become scenarios later.)
If they have no price yet, make them pick one to score at — price
determines the business model, and an unpriced idea can't be scored.
Also establish in this phase: what the user's ambition is
(salary-replacing indie vs. scale-up — it changes how the threshold
reads), and what real evidence they already have (customers, data,
interviews, waitlists).
Create
as soon as the target market is settled — it is
the first settled item, and the file is the memory of the exercise, not
the chat.
Phase B — Score the seven criteria, one per exchange
For each criterion in order: state it in one line with its scale, ask for
the user's value and justification, research where useful, grill per the
posture rules, settle the value, and append it to the file with its
justification (one or two lines, like the calibration examples), its
evidence class —
,
,
; tag mixed evidence
with the strongest class that materially supports the value, and append
"low-conf" when research was warranted but unavailable — and any dissent
line. Write concessions and self-corrections into the justification
("user came down from 1M") — a resumed session must be able to tell a
grilled score from a rubber-stamped one.
Phase C — Compute and deliver the verdict
Multiply, divide by 625,000, and deliver the verdict before any remedies:
the number, what it means against the user's stated ambition, and — most
importantly — the one or two weak links that dominate the result. Remind
the user what the score does NOT measure (reach, execution, costs, team).
Verdict first, whole; negotiating findings one at a time as they land is
how optimism creeps back in.
Phase D — Scenarios: narrow, re-score, or face the truth
A bad or marginal score is the beginning of the useful part, not the end.
Narrowing the target market is very often the right move. It feels
like shrinking ambition; it usually isn't — focus concentrates every
other score. Dropping from "all consumers" to a sharply-defined niche
typically costs one or two powers of ten on Plausible while raising
Self-Aware, Lucrative, and both Eager scores by more than that combined.
And targeting the bullseye doesn't forfeit the rest of the market: the
buyers adjacent to a sharply-drawn ideal customer respond to the same
clear positioning, so the effective market is many times larger than the
niche itself. Propose one or two candidate niches, re-walk ONLY the
scores that change, and record each scenario side by side in the file —
noting when a scenario is really a different product (a narrower market
at a higher price with different trade-offs is a different business).
The boundary: the same buyer narrowed or re-priced is a scenario in this
file; a different buyer type (a marketplace's other side, a different
persona) gets its own scorecard — a second top-level section or file.
For a multi-sided business, every side must clear the bar; the weaker
side is the business's constraint, and note where one side's scores
quietly assume the other side already exists.
Exception paths are the other lever: each criterion has one, and each
comes with conditions. Offering one means asking whether the user
genuinely wants that path — the evangelist's decade of educating a
market, the price-fighter's margins — not whether they'll nod at it.
Or face the truth. If no scenario reaches viability, say so: this
idea, in every market the user is willing to serve, is not a viable
business as scored — and it's better to know now, with time and money
left to find a better idea. That is a successful outcome of this
exercise.
Close by pointing forward: the scorecard's
-class scores are
guesses that customer conversations can convert to evidence — open-ended
interviews with the defined buyer, testing whether they know they have
the problem, what budget it comes from, and when they'd buy. Re-score as
evidence arrives; the file supports it.
Resuming
If
exists when the skill loads, read it and continue
from the
pointer in its status note — nothing settled gets
re-asked.
The scorecard file
markdown
# Problem Score — <idea, in a few words>
⚠️ IN PROGRESS — next: <criterion or phase>; <any plan a resumed session
must inherit, e.g. "scenario B re-scores only Plausible, Lucrative">
<!-- remove this note when final -->
## Target market (Scenario A)
Buyer: <specific> · Problem: <specific> · Trade-offs: <the deliberate
ones> · Price: <$X> · Ambition: <indie|scale-up> · Evidence: <what's real>
## Scores (Scenario A)
| Criterion | Value | Justification | Class |
| :-- | :-- | :-- | :-- |
| Plausible | 1M | <one or two lines> | [research] |
| Self-Aware | 0.1 | <…> — dissent: user says 0.5; evidence supports 0.1 (total 1.2 → 0.24) | [fermi] |
**Total: <product> ÷ 625,000 = <score>** — <one-line verdict>
## Scenario B — <the niche>
<same shape; only changed scores re-justified, others carried over>
## Verdict & next steps
<verdict across scenarios; the weak links; what the score does not
measure; which [fermi] scores to verify with interviews; the decision>
Refusal conditions
- No specific target market. A direction ("something for indie
makers") is a jumping-off point for Phase A construction, never a
scoring target. Refuse the score, not the user.
- Multiple buyer types in one scorecard. A marketplace's two sides,
or "SMBs and enterprises," are different markets with different
scores. Score each separately; refuse the blended average.
- Validation theater. If the user signals they want a good score —
or the launch is already committed and no answer would change it —
name that, and offer to proceed only on honest terms.
- Precision demands. Feel-based in-between values ("0.7-ish") are
refused — precision must come from hard data. And in either direction,
refuse to treat 0.9 vs. 1.1 as meaningfully different verdicts; the
tool is directional.
- Execution questions. Reaching customers, ads, hiring, fundraising —
outside what this score measures; say so rather than stretch the rubric.