asb-problem

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Problem Score — from "The Problem" to "Viable Business Model"

问题评分——从「问题」到「可行商业模式」

Here is how companies fail: a founder has a flash of insight — the world has a Problem. Potential customers agree the Problem is real (they're right!). The founder builds a product that truly solves it (it does!). And then sales never materialize, and the company shuts down within a couple of years. Solving a real problem is — perhaps surprisingly — not nearly enough to build a successful company. Between "real problem" and "viable business" sit seven conditions, and founders systematically over-rate themselves on all of them. This skill scores those conditions honestly, one specific target market at a time; the user walks away with a scorecard, a directional verdict, and either a niche that rescues the idea or the clear-eyed conclusion — far cheaper now than in two years — that it isn't viable.
企业失败的常见路径是:创始人灵光一闪,发现市场存在某个问题。潜在客户也认同这个问题确实存在(他们没错!)。创始人打造出能真正解决该问题的产品(产品确实有效!)。但随后销量始终无法达标,企业在短短几年内就倒闭了。令人意外的是,解决真实问题远不足以打造一家成功的企业。 在「真实问题」与「可行商业模式」之间,存在7项必备条件,而创始人往往会系统性地高估自己在这些条件上的表现。本工具会针对特定目标市场,如实评估这些条件;用户最终会得到一份评分卡、方向性结论,要么找到能挽救想法的细分利基市场,要么清醒地得出结论——现在放弃远比两年后再放弃成本低得多——该想法不具备可行性。

The mental model

心智模型

A series of "ands" — why the scores multiply

一系列「且」——为何评分采用乘法

A sale requires that enough people have the problem AND they know and care AND they have budget AND they can buy now AND they'd buy from you AND they'll stick around. The best case is that all seven hold. Realistically, some will be strong and some weak; the real question is whether the big strengths overcome the few weaknesses. Multiplication answers that question mechanically: compounding "ands" means one near-zero factor drags down the whole product no matter how impressive the rest are — which matches reality. Gut feel is no guide here: every idea feels true to its founder, including the majority that turn out wrong. The multiplication is how the weak link gets seen instead of glossed over.
达成一笔交易需要满足:足够多的人存在该问题 他们知晓并在意该问题 他们有预算 他们能立即购买 他们愿意从你这里购买 他们会持续留存。最佳情况是7项条件全部满足。实际情况中,有些条件表现强劲,有些则薄弱;核心问题在于,优势是否能弥补少数劣势。乘法运算能机械地回答这个问题:多个「且」的叠加意味着,只要有一项因素接近零,无论其他因素多么亮眼,整体评分都会被拉低——这与现实情况一致。直觉在这里毫无用处:每个想法在创始人看来都是正确的,包括大多数最终被证明错误的想法。乘法运算能让薄弱环节无所遁形,而非被忽略。

Fermi estimation: powers of ten by default

Fermi估算:默认采用数量级

Every score is a power of ten (or a fixed coarse value like 0.1 / 0.5 / 1.0) — no in-between numbers, no false precision — with one exception: hard data is used as-is (8 billion humans is 8B, not rounded to 10B; measured 4%/mo churn is 4%). Precision comes from data, never from feel, and since it multiplies with rough estimates, the final score is still only good to a power of ten. The coarseness is what makes the exercise fast and honest: most values are easy to pick because the adjacent choices are absurd ("100k or 1M? — certainly not 10k, certainly not 10M"), and when a value IS controversial, that order-of-magnitude disagreement is genuine strategic uncertainty worth its own conversation, not a rounding argument. Real evidence — the user's own data, or numbers found by quick research — settles a value immediately. Without it, the honest move is the power of ten whose neighbors are clearly wrong, not the flattering one.
每项评分均为10的幂次(或固定的粗略值,如0.1/0.5/1.0)——没有中间值,也不存在虚假精度——唯一例外是:硬数据直接使用(如全球80亿人口就记为8B,不四舍五入为10B;测得的月流失率4%就记为4%)。精度仅来自数据,而非直觉;由于数据会与粗略估算相乘,最终评分仍仅能精确到数量级。这种粗略性正是让评估快速且诚实的关键:大多数数值很容易选择,因为相邻选项显然不合理(「10万还是100万?——肯定不是1万,也肯定不是1000万」);当某个数值存在争议时,这种数量级的分歧是真正值得探讨的战略不确定性,而非舍入问题。真实证据——用户自有数据或快速调研得到的数字——能立即确定数值。若无证据,诚实的做法是选择相邻选项均明显不合理的10的幂次,而非偏向乐观的数值。

The score is guidance, not analysis

评分仅为指引,而非精确分析

This is dangerously close to a silly quiz, and it must be used as guidance, never as precise analysis: a 0.8 is not doom, a 1.2 is not salvation. What the score reliably does is expose the weak links and show whether a different target market changes the answer by a lot — and having to think through the answers and trade-offs is most of the value, more than the final number. The score also deliberately measures only the path from problem to business model. It says nothing about reach, marketing cost, team, skills, or execution — a great score can still fail.
这看似是一个简单的测试,但必须将其作为指引,而非精确分析:0.8分并非末日,1.2分也并非救赎。评分的可靠作用是暴露薄弱环节,以及展示不同目标市场是否会大幅改变结果——而思考答案和权衡过程本身的价值,远大于最终的数字。此外,评分仅针对从问题到商业模式的路径,不涉及触达、营销成本、团队、技能或执行层面——即便评分很高,企业仍可能失败。

Optimism is the default failure mode

乐观是默认的失败模式

People are almost always too generous with what they can do and what customers will do, think, and pay. Every score therefore gets challenged before it is recorded — a scorecard filled in by unchallenged optimism is worthless; it will say "viable" about anything.
人们几乎总是高估自己的能力,以及客户的行为、想法和付费意愿。因此,每项评分在记录前都会受到质疑——未经质疑的乐观填写的评分卡毫无价值,它会认为任何想法都「可行」。

The seven criteria

七项评估标准

Score each with respect to ONE specific target market: a specific type of buyer, solving a specific problem, with a product that has made specific trade-offs, at a specific price. The rubric applies to any kind of company — software, services, restaurants, hardware, content — not just tech startups.
每项标准均针对单一特定目标市场进行评分:特定类型的买家、特定问题、做出特定取舍的产品、特定价格。该评估框架适用于任何类型的企业——软件、服务、餐饮、硬件、内容——不仅限于科技初创企业。

1. Plausible — do enough people have the problem?

1. Plausible——存在该问题的人群是否足够多?

Scale (power of ten only):
1k
,
10k
,
100k
,
1M
,
10M
,
100M
,
1B
— the number of consumers or businesses that actually have the problem.
Why the bar is high: marketing math. Ads convert roughly 1% of impressions to visitors, and a good product site converts roughly 1% of visitors to paying — about 10,000 impressions per customer. A sustainable small company needs on the order of 1,000 customers (at $30–$100/mo; cheaper means more needed), so ~10,000,000 impressions and about two years — a timeline even eventual giants needed for their first 1,000. Consumers: ~10M must have the problem. Businesses pay orders of magnitude more and convert better, so ~100k suffices.
Press on: counting everyone who could theoretically use the product instead of the specific buyer defined above; confusing "has the problem" with "matches my product description." Exception-with-conditions: a high-price product in a small niche can be a fine company, and so can a deliberately-small business replacing a salary — but then the other scores must be strong, and the user must genuinely want that path.
评分范围(仅为10的幂次):
1k
10k
100k
1M
10M
100M
1B
——实际存在该问题的消费者或企业数量。
为何门槛较高:营销数学逻辑。广告的曝光到访客转化率约为1%,优质产品网站的访客到付费用户转化率约为1%——即每获取一位客户需要约10000次曝光。一家可持续发展的小型企业需要约1000位客户(按每月30-100美元计算;价格更低则需要更多客户),因此需要约1000万次曝光,耗时约两年——即便是最终的行业巨头,获取前1000位客户也需要这样的时间线。面向消费者:约1000万人必须存在该问题。面向企业:企业付费金额高出数个数量级,转化率也更高,因此约10万家企业即可满足要求。
需质疑的点: 将理论上可以使用产品的所有人算作目标人群,而非上述定义的特定买家;混淆「存在该问题」与「符合产品描述」。附带条件的例外: 高价产品瞄准小众市场也能打造优秀企业,刻意打造的小型企业替代薪资收入也可行——但此时其他评分必须表现强劲,且用户必须真正愿意选择这条路径。

2. Self-Aware — do they know and care that they have the problem?

2. Self-Aware——他们是否知晓并在意自己存在该问题?

Scale:
0.01
few agree or care ·
0.1
thought-leaders care and evangelize ·
0.5
industry standard practice ·
1.0
almost impossible to find someone who doesn't care.
Someone who doesn't believe they have a problem isn't searching for a solution, and won't spend money on one even if they stumble across it. Failure comes in two flavors: ignorance (millions of website owners truly are targets of hackers, yet think "no one would attack little old me," so they never shop for security), and knowing-but-not-caring (nearly everyone agrees an inaccessible website is a problem — and it still never cracks their top three priorities, so nothing happens). Market timing is a version of this: the same idea can fail years before the market is ready and succeed after.
Press on: "they just don't realize it yet — once we explain, they'll get it." That is a market you must CREATE: difficult, expensive, and slow. Exception-with-conditions: a founder who is a natural evangelist on a genuine mission can educate a market into existence — but must truly want years of that work, not merely tolerate the idea of it.
评分范围:
0.01
极少人认同或在意 ·
0.1
意见领袖在意并积极推广 ·
0.5
行业标准实践 ·
1.0
几乎找不到不在意的人。
不认为自己存在问题的人不会主动寻找解决方案,即便偶然遇到也不会付费。失败分为两种情况:无知(数百万网站所有者确实是黑客攻击的目标,但他们认为「没人会攻击我这个小网站」,因此从未关注安全产品),以及知晓但不在意(几乎所有人都认同网站无障碍是个问题——但这从未进入他们的前三优先级,因此毫无行动)。市场时机也属于此类:同一个想法可能在市场成熟前失败,在成熟后成功。
需质疑的点:「他们只是还没意识到——一旦我们解释,他们就会明白。」这意味着你需要创造市场:难度大、成本高、耗时久。附带条件的例外: 天生具备 evangelist(布道者)特质且拥有真正使命的创始人,可以培育出一个市场——但必须真正愿意投入数年时间,而非仅仅容忍这件事。

3. Lucrative — do they have substantial allocated budget?

3. Lucrative——他们是否有充足的预算分配?

Scale (power of ten only, of net revenue):
$1
,
$10
,
$100
,
$1k
,
$10k
,
$100k
,
$1M
— annual budget actually allocated to this problem. "Net revenue" means your revenue after pass-through costs (an eCommerce platform processing $100 and keeping $10 counts $10) — but do NOT subtract your marketing, support, or infrastructure costs; this measures top-line, not efficiency.
Agreeing the problem exists is not the same as having money assigned to solving it. Consumers mostly refuse to pay for software at all — people publicly agonize over $15/year for an app they use daily. Whole customer categories are structurally broke (college students — and therefore also the businesses that sell to them). In large companies, budget only exists for the top few problems of the year, and internal teams already tasked with the problem often fight outside solutions; target the companies that outsource this problem, not the ones that staff it.
Press on: "they'd definitely pay for this" without a story for whose budget line it comes from and who approves it. Score the budget the market demonstrably allocates — what these buyers pay anyone today, with your realistic annual revenue per customer as evidence; when your sticker price and the demonstrated allocation diverge, score the allocation and note the gap. Exception-with- conditions: a huge market at a low price can work IF the cost basis is extremely low (self-service, near-zero support, cheap acquisition, a product simple enough to scale unattended) — and then the Plausible number must rise accordingly.
评分范围(仅为10的幂次,基于净收入):
$1
$10
$100
$1k
$10k
$100k
$1M
——每年实际分配给解决该问题的预算。「净收入」指扣除直通成本后的收入(如电商平台处理100美元订单,留存10美元,则记为10美元)——但请勿扣除营销、支持或基础设施成本;该指标衡量的是营收总额,而非效率。
认同问题存在并不等同于有预算解决问题。消费者大多拒绝为软件付费——人们会为一款日常使用的应用每年15美元的费用纠结不已。有些客户群体本质上资金有限(如大学生——因此面向他们的企业也面临同样问题)。在大型企业中,预算仅分配给当年最优先级的几个问题,且负责该问题的内部团队通常会抵制外部解决方案;应瞄准外包该问题的企业,而非自行组建团队解决的企业。
需质疑的点:「他们肯定会为此付费」,却无法说明预算来自哪个部门、由谁审批。应根据市场实际分配的预算进行评分——即这些买家目前为同类产品支付的费用,以你预期的每位客户年度收入为依据;当你的标价与实际分配的预算存在差异时,按实际分配的预算评分,并注明差距。附带条件的例外: 低价产品瞄准庞大市场也可行,但前提是成本基础极低(自助服务、近乎零支持、获客成本低、产品足够简单可无人值守规模化)——此时Plausible的评分必须相应提高。

4. Liquid — are they willing and able to buy right now?

4. Liquid——他们是否愿意且能够立即购买?

Scale:
0.01
a decision made every few years ·
0.1
an annual decision ·
1.0
always in the market, easy to switch.
A customer can love the product, agree it's valuable, have the budget — and still not buy, because buying isn't possible or isn't a priority right now. These forces have nothing to do with your product or its price, which is exactly why they blindside founders: multi-year contracts, "already bundled in the system we pay for anyway," data and integration lock-in, retraining costs, government fiat. And the quieter version: a buyer has two or three top priorities at any moment; if you're priority seven, "call back in nine months" is sincere — and fatal. Moment-in-time products (event websites, load-testing tools) suffer this permanently: before the moment there's no problem, after it no customer.
Press on: "they'll switch because we're better" — the lock-in forces overwhelm better-and-cheaper; and on scoring the decision frequency of the category, not the user's hopes. Exception-with-conditions: you can pay contract penalties, do migrations for free, target the segment the incumbent over-serves or prices out, or make it free to keep while idle — but each must be a deliberate strategy you can afford, not a hope.
评分范围:
0.01
每数年才做一次决策 ·
0.1
每年做一次决策 ·
1.0
始终处于市场中,易于切换。
客户可能喜欢产品、认同其价值、有预算——但仍不会购买,因为无法购买或当前不优先考虑。这些因素与你的产品或价格无关,这也是它们会让创始人措手不及的原因:多年期合同、「已包含在我们付费的系统中」、数据和集成锁定、再培训成本、政府规定。还有一种更隐蔽的情况:买家在任何时候都只有两三个优先事项;如果你排在第七位,「九个月后再联系」是真诚的——但对企业来说是致命的。时效性产品(如活动网站、负载测试工具)始终面临这个问题:事件发生前没有需求,事件发生后没有客户。
需质疑的点:「因为我们更优秀,他们会切换过来」——锁定力量会压倒「更优更廉」;应根据品类的决策频率评分,而非用户的期望。附带条件的例外: 你可以支付合同违约金、免费提供迁移服务、瞄准被现有服务商过度服务或定价过高的细分市场,或提供闲置时免费保留的服务——但每项都必须是你能负担得起的刻意策略,而非仅仅是期望。

5. Eager (identity) — do they want to buy from YOU?

5. Eager(身份维度)——他们是否愿意从你这里购买?

Scale:
0
they cannot buy from you (structurally barred — fiat, policy, impossibility; NOT merely "we haven't launched yet") ·
0.1
structural challenges ·
0.5
indifferent, no red flags ·
1.0
mission-level emotional desire to select you.
Even in a live purchase, the buyer must trust that the product works, the company will survive, support will show up, security won't embarrass them, and you can scale as they do. "You've only been in business a year" and "our policy requires SOC 2" are this score — and so is the positive version: buying partly to support what you stand for. This is independent of Liquid: lunch is re-decided daily (hyper-liquid), yet a given person may never buy from McDonald's, or never set foot in the hippie place — decision frequency and attitude toward the seller are different dimensions, even when big-company purchasing habits make them look correlated.
Press on: "we'll earn trust quickly" — with what track record, references, or mitigation? Exception-with-conditions: build a product type that needs little trust (non-private data, not time-critical, sold to individuals who like buying from startups), or mitigate structurally (e.g. open source as an escape hatch), or carry a mission distinctive enough that buying from you is part of the point.
评分范围:
0
他们无法从你这里购买(存在结构性障碍——如法律规定、政策限制、客观不可能;而非「我们尚未上线」) ·
0.1
存在结构性挑战 ·
0.5
无偏好,无负面信号 ·
1.0
出于使命层面的情感需求选择你。
即便是在实时购买场景中,买家也必须相信产品有效、企业能存续、支持服务会到位、安全性不会让他们难堪,且你能随他们一起成长。「你们才成立一年」和「我们的政策要求SOC 2认证」都属于该评分的范畴——正面情况也包括:部分购买是为了支持你的理念。该指标与Liquid相互独立:午餐每天都要重新决定(高度Liquid),但某人可能永远不会从麦当劳购买,也永远不会踏入那家嬉皮风格的店——决策频率和对卖家的态度是不同维度,即便大企业的采购习惯让它们看起来相关。
需质疑的点:「我们会快速赢得信任」——凭借什么记录、参考案例或缓解措施?附带条件的例外: 打造对信任要求低的产品类型(非隐私数据、非时间敏感、面向喜欢从初创企业购买的个人),或从结构上缓解(如将开源作为退路),或拥有足够独特的使命,让从你这里购买成为用户的选择之一。

6. Eager (comparative) — differentiated enough to win the deal?

6. Eager(对比维度)——差异化是否足以赢得交易?

Scale:
0.1
no material differentiation ·
0.5
some things so good that some people buy for them alone ·
1.0
one-of-a-kind with no viable alternative.
They will buy — but from you, or from one of the alternatives? Differentiation is not "we have a unique feature": if only 10% of the market cares about your unique feature, while 30% care about the one your competitor has and you lack, you lose. Over-serving is a real trap — ten features where the market wants three means the simpler, cheaper rival is the rational choice no matter what your comparison matrix says. The strong versions of this score come from picking a game the competitors cannot play — a difference taken to an extreme, aligned with everything else about the company, that their structure prevents them from copying.
Press on: feature lists as "differentiation"; ask what fraction of the defined target market would buy for that difference alone. Exception-with-conditions: specialize in a niche of a large market; in a tiny market, few viable competitors may exist; competing on price can work but degrades margin and customer quality — choose it on purpose, if at all.
评分范围:
0.1
无实质性差异化 ·
0.5
部分优势足够突出,足以吸引部分用户单独购买 ·
1.0
独一无二,无可行替代方案。
他们会购买——但从你这里,还是从竞争对手那里?差异化并非「我们有独特功能」:如果只有10%的市场在意你的独特功能,而30%的市场在意竞争对手有但你没有的功能,你就会失败。过度服务是一个真正的陷阱——市场只需要3个功能,你却提供10个,那么更简单、更便宜的竞争对手才是理性选择,无论你的对比矩阵如何展示。该评分的高分来自选择竞争对手无法参与的赛道——将差异推向极致,与企业的其他所有方面保持一致,而竞争对手的结构使其无法复制。
需质疑的点: 将功能列表视为「差异化」;应询问定义的目标市场中有多大比例会因该差异单独购买。附带条件的例外: 专注于大市场中的细分领域;在极小市场中,可能几乎没有可行的竞争对手;低价竞争可行,但会降低利润率和客户质量——若选择此策略,需是刻意为之。

7. Enduring — will they still be paying a year from now?

7. Enduring——一年后他们是否仍会付费?

Scale:
0.01
one-off purchase without loyalty ·
0.1
one-off, but happy customers buy again and refer ·
0.5
recurring revenue from a recurring problem ·
1.0
strong lock-in (fiat, integrations, or being the system of record for something business-critical).
Growth is linear (quadratic for the hyper-growth outliers); cancellation is exponential — a percentage of an ever-larger base. Exponentials always catch up. At 5%/mo churn, half the customers are gone within a year; at 7%/mo, a company adding a healthy 15%/mo of new revenue stops growing entirely about a year later, with all its marketing spend canceled out by departures. High churn isn't primarily a metrics problem — it means customers don't actually want the product. One-time-revenue businesses don't escape: they still need repeat purchases and referrals, which require the same satisfaction.
Press on: "our churn will be fine" from a user with no retention data — what's the evidence customers with this problem keep paying anyone? And on temporary problems dressed as recurring ones. Exception-with-conditions: high churn can be survived only when acquisition is cheap, the market is effectively inexhaustible, the customers who stay grow super-linearly, and — non-negotiable — churn is NOT the product's fault. If customers leave because the product disappoints, there is no exception.
评分范围:
0.01
一次性购买,无忠诚度 ·
0.1
一次性购买,但满意的客户会复购并推荐 ·
0.5
来自持续性问题的 recurring revenue( recurring revenue) ·
1.0
强锁定(如法律规定、集成、成为业务关键系统的记录系统)。
增长是线性的(超增长 outliers 是二次增长);流失是指数级的——占不断扩大的客户基数的一定比例。指数增长最终总会赶上。月流失率5%时,一年内会流失一半客户;月流失率7%时,即便企业每月新增15%的健康收入,约一年后也会停止增长,所有营销投入都会被流失抵消。高流失率主要不是指标问题——这意味着客户并不真正需要产品。一次性收入企业也无法幸免:它们仍需要复购和推荐,这同样需要客户满意。
需质疑的点:「我们的流失率会很低」,但用户没有留存数据——有什么证据表明存在该问题的客户会持续为同类产品付费?以及将临时问题伪装成持续性问题。附带条件的例外: 只有当获客成本极低、市场几乎无穷无尽、留存客户的增长远超线性,且——必不可少的是——流失并非产品本身的问题时,高流失率才能被接受。如果客户因产品失望而离开,则无例外可言。

Computing the score

计算评分

Multiply all seven values together, then divide by 625,000 to normalize. Present the computed number, but read it at its power of ten. Roughly: ≥ 1 can sustain an indie company; ≥ 2 has scale-up potential; well below 1 is not a viable model as scored. A zero anywhere (the buyer cannot buy from you) makes the whole product zero — that criterion is a deal-breaker to design around before further scoring is worth anything. The normalization encodes the marketing math above: a workable business — 10M consumers at ~$10/mo, or 100k businesses at ~$1,000/mo — with middling values everywhere else lands near 1.
Calibration anchors, from the framework's own worked examples:
  • Managed WordPress hosting for businesses:
    100M
    ×
    0.1
    (aware) ×
    $100
    ×
    0.01
    (switching) ×
    0.5
    ×
    0.5
    ×
    1.0
    (retention) = 4 — a scale-up, and that's what happened.
  • Email marketing for creators monetizing newsletters:
    10M
    ×
    1.0
    ×
    $100
    ×
    0.01
    ×
    0.5
    ×
    0.5
    ×
    0.5
    = 2 — a strong bootstrapped business, which is what happened.
  • Security software for all consumers:
    1B
    ×
    0.01
    (aware) ×
    $10
    ×
    0.01
    ×
    0.5
    ×
    0.1
    (undifferentiated) ×
    0.5
    = 0.04 — not viable, matching the graveyard of consumer-security indies.
Justifications in those examples were one or two lines of page-one search results — the right depth; a power of ten is the whole requirement, and you often know that much without data.
将7项数值相乘,然后除以625000进行归一化。展示计算结果,但按数量级解读。大致标准:≥1分可支撑独立企业;≥2分具备规模化潜力;远低于1分则当前评分下模式不可行。 任何一项为0(买家无法从你这里购买)都会导致整体评分为0——该标准是设计时必须解决的致命问题,否则后续评分毫无意义。归一化编码了上述营销数学逻辑:一个可行的企业——1000万消费者,每月约10美元;或10万家企业,每年约1000美元——其他各项评分中等时,最终得分接近1分。
校准参考(来自框架自身的实例):
  • 面向企业的托管WordPress主机:
    100M
    ×
    0.1
    (认知度) ×
    $100
    ×
    0.01
    (切换难度) ×
    0.5
    ×
    0.5
    ×
    1.0
    (留存) = 4——具备规模化潜力,实际情况也确实如此。
  • 面向通过时事通讯变现的创作者的电子邮件营销工具:
    10M
    ×
    1.0
    ×
    $100
    ×
    0.01
    ×
    0.5
    ×
    0.5
    ×
    0.5
    = 2——强劲的自启动企业,实际情况也确实如此。
  • 面向所有消费者的安全软件:
    1B
    ×
    0.01
    (认知度) ×
    $10
    ×
    0.01
    ×
    0.5
    ×
    0.1
    (差异化不足) ×
    0.5
    = 0.04——不可行,与消费者安全独立企业的失败现状相符。
这些实例的理由仅需一页搜索结果中的一两句话——深度恰到好处;数量级是核心要求,通常无需数据就能判断。

How to keep the score honest

如何保持评分的诚实性

This section governs every exchange. The user's optimism is the enemy of the exercise, and the exercise only helps if it wins.
本节规则适用于所有交流。用户的乐观是评估的敌人,只有战胜乐观,评估才有价值。

Evidence settles; optimism gets grilled

证据定论;乐观需质疑

Three ways a score earns its way into the file: (a) the user's real data (their retention numbers, their sales conversations, their price tests); (b) evidence found by quick research; (c) an honest Fermi argument where the adjacent powers of ten are clearly absurd. Cheerful assertion is none of these. The standing move is the challenge from below: when the user proposes a value, ask what makes the next value DOWN wrong. If they can't answer, the lower value is the honest score.
评分纳入记录的三种方式:(a) 用户的真实数据(他们的留存数据、销售对话、价格测试);(b) 快速调研得到的证据;(c) 诚实的Fermi论证,相邻数量级明显不合理。乐观断言不属于上述任何一种。常规做法是从下方质疑:当用户提出某个数值时,询问为何更低的数值不合理。若无法回答,则更低的数值才是诚实的评分。

Do lightweight research where it helps

必要时进行轻量调研

If the environment provides web access, spend it Fermi-style where outside numbers exist: market counts for Plausible, budget norms and competitor pricing for Lucrative, evidence of active demand for Self-Aware (are people visibly searching, complaining, paying anyone today?), the competitive field for Eager. Page-one-search depth is correct — the answer only needs to survive to a power of ten. Findings are evidence: accept the value they support even when it beats your skepticism. Without web access, challenge the user's numbers against reference points you can reason from, and mark research-worthy scores as low-confidence in the file.
若可访问网络,针对存在外部数据的场景进行Fermi式调研:Plausible的市场规模、Lucrative的预算标准和竞争对手定价、Self-Aware的活跃需求证据(人们是否在明显搜索、抱怨、付费购买同类产品?)、Eager的竞争格局。一页搜索结果的深度足够——答案只需精确到数量级。调研结果即为证据:即便与你的怀疑相悖,也应接受其支持的数值。若无网络访问权限,需将用户的数值与你可推理的参考点进行对比,并在记录中标记需要调研的低置信度评分。

Rude questions, gentle framing

尖锐问题,温和表达

The questions are deliberately hard, because they're the ones the market will ask. The framing is collegial: attack the claim, never the person, and grill because you want them to win. Acknowledge a crisp answer before moving on ("that's defensible — recorded"). Never soften a question to be polite, and never lower the bar because the conversation is tired.
If a devil's-advocate skill such as Rude Q&A (
asb-rude-qa
) is installed, you may invoke it on a single fiercely-contested score with a brief like: "Attack this justification for scoring Self-Aware at 0.5: <justification>. Evidence, or wishful?" This is optional; the rules in this section are the standalone equivalent and fully sufficient.
问题故意设计得尖锐,因为市场也会提出这些问题。表达需 collegial( collegial):质疑观点,而非攻击个人,质疑是为了帮助用户成功。在继续前认可清晰的回答(「这是合理的——已记录」)。切勿为了礼貌而软化问题,也切勿因对话疲惫而降低标准。
若已安装唱反调类工具(如Rude Q&A (
asb-rude-qa
)),可针对某个极具争议的评分调用该工具,例如:「质疑将Self-Aware评分为0.5的理由:<理由>。是证据,还是一厢情愿?」这是可选操作;本节规则已足够独立完成评估。

Dwell until it's real

深入探讨直至真实

When an answer is wishful, vague, or "we'll figure that out later," stay on the point and say so: "I'm going to stay here — that justification wouldn't survive contact with a stranger." Offer one or two candidate values with reasoning if the user is stuck, and ask them to pick or revise. Three rounds on one score is not a reason to accept it. Move on only when the value is evidenced or honestly argued.
当回答是一厢情愿、模糊不清或「我们之后再解决」时,需停留在该问题上并明确指出:「我会继续聚焦这里——这个理由在陌生人面前站不住脚。」若用户陷入困境,可提供一两个带有推理的候选数值,让他们选择或修改。针对一项评分进行三轮探讨并非接受该评分的理由。只有当数值有证据支持或论证诚实可信时,才可继续。

No curve — in either direction

不设曲线——无论高低

A low score is not a failure of the exercise; it IS the product of the exercise. Don't pad weak criteria out of sympathy. Equally: when a criterion genuinely earns a strong value — real evidence, sound argument — say so plainly and record it without manufactured skepticism. The goal is a true number, not a low one.
低分并非评估失败;它正是评估的结果。切勿出于同情而抬高薄弱标准的评分。同样:当某项标准确实获得高分——有真实证据、论证合理——需明确说明并记录,无需刻意质疑。目标是得到真实数值,而非低分。

Their scorecard, your dissent

用户的评分卡,你的异议

Two things are non-negotiable craft: no scoring starts before the target market is specific, and no off-scale value is recorded without hard data behind it (a measured 4%/mo churn earns its precision; a felt "0.7" doesn't — scale membership is the precondition; the user's sovereignty is over which scale value). The chosen value is ultimately the user's: after a full grilling, their number goes in the file. If you still disagree, record the dissent next to it — "scored 1.0 by user; evidence shown supports 0.1, which drops the total from 4.0 to 0.4" — so the disagreement and its stakes are visible, and move on.
两项不可妥协的原则:在明确目标市场前不得开始评分;无硬数据支持不得记录超出范围的数值(测得的月流失率4%可保留精度;主观的「0.7」不可——必须符合评分范围;用户的自主权在于选择哪个范围数值)。最终数值由用户决定:经过充分质疑后,他们选择的数值将纳入记录。若你仍不同意,需在旁边记录异议——「用户评分为1.0;现有证据支持0.1,总分从4.0降至0.4」——以便清晰展示分歧及其影响,然后继续。

One criterion per exchange

每次交流仅处理一项标准

Open small: acknowledge the idea and ask what's needed to pin the target market — never an opening wall with all seven criteria pre-scored, and never batch-scoring from the initial description, however much it seems to contain. During scoring, work exactly one criterion per exchange — settle it, write it to the file, move on — and end each message with exactly one thing for the user to answer. Three sanctioned exceptions: Phase A intake may bundle related items into one correct-this-template proposal (harvest whatever the opening message already answered; ask only what's missing); when the user volunteers the next criterion's answer, settle it rather than re-asking; and Phase D may propose a scenario's changed values as one package for the user to correct, since the base scores are already settled.
从小处着手:认可想法,询问明确目标市场所需的信息——切勿一开始就列出所有七项标准进行预评分,也切勿根据初始描述批量评分,无论描述看似包含多少信息。评分过程中,每次交流仅处理一项标准——确定数值、写入记录、继续——且每条消息结尾仅提出一个需要用户回答的问题。三种允许的例外情况:A阶段 intake( intake)可将相关项整合为一个模板修正请求(收集初始消息已回答的内容;仅询问缺失的信息);当用户主动提供下一项标准的答案时,直接确定数值而非重新询问;D阶段可将场景的变更数值作为一个整体提出供用户修正,因为基础评分已确定。

Willing to land on "not viable"

愿意得出「不可行」的结论

If the multiplied truth is 0.04, say so plainly, then do the constructive part: "not viable as scored" is a statement about THIS target market, which is exactly why scenarios come next. A skill that always finds a way to call the idea viable is a skill that lies.
若相乘后的真实结果为0.04,需明确说明,然后进行建设性部分:「当前评分下不可行」是针对该目标市场的结论,这正是后续场景分析的意义所在。总是能找到理由称想法可行的工具是在撒谎。

Be clear, not clever

清晰表达,而非炫技

Write to be understood, not admired. The work here wrestles with hard concepts, and clever metaphors, wordplay, or cute turns of phrase make them harder to grasp, not easier. Say plainly what you mean. If a sentence reads more clearly without a flourish, cut the flourish. State the actual point rather than gesturing wittily at it.
写作旨在被理解,而非被赞赏。此处的工作涉及复杂概念,巧妙的隐喻、文字游戏或俏皮表达会让概念更难理解,而非更容易。直白地表达你的意思。若去掉修饰后句子更清晰,就删掉修饰。直接陈述要点,而非巧妙暗示。

How to use this skill

如何使用本工具

Phase A — Pin down one specific target market

A阶段——明确单一特定目标市场

First, files: the scorecard lives in
PROBLEM-SCORE.md
. If the user pointed at a directory or existing files, put it there; otherwise ask where the work should live (offer the current directory as the default) — never silently pick a location.
Then the gate. Nothing gets scored until there is a specific target market: a specific type of buyer with a specific problem, served by a product with specific trade-offs, at a specific price. The test for specific: a stranger could sort real people into "in the market" and "not in the market" using the description. "Indie makers," "small businesses," "people who want to be more productive" all fail.
A second test, faster and harder to fake: ask them to name real examples — actual companies or actual people they could point at today. A handful is plenty; this is not a counting exercise. Then press the part that matters: are there many more like these, or is each one a special case? No names at all means the market is imagined rather than observed. Names that turn out to be one-offs — a friend's company, an unusual setup, the one client who happened to ask — are real customers but not a pattern, and only a pattern can be scored; interesting one-offs are the most common way a market that isn't there looks like one that is. Either way this is Phase A material rather than a dead end: the names they can give are the raw material for a narrower, realer market.
When the user arrives generic — an idea-shaped direction rather than a scoreable market — do not refuse and stop; refuse and BUILD. Treat their direction as the jumping-off point: propose one to three candidate specific markets consistent with it, as templates they must correct, and converge on one to score first. (The others can become scenarios later.) If they have no price yet, make them pick one to score at — price determines the business model, and an unpriced idea can't be scored.
Also establish in this phase: what the user's ambition is (salary-replacing indie vs. scale-up — it changes how the threshold reads), and what real evidence they already have (customers, data, interviews, waitlists).
Create
PROBLEM-SCORE.md
as soon as the target market is settled — it is the first settled item, and the file is the memory of the exercise, not the chat.
首先是文件:评分卡存储在
PROBLEM-SCORE.md
中。若用户指定了目录或现有文件,将其放在那里;否则询问工作应存储的位置(默认当前目录)——切勿自行选择位置。
然后是门槛。在明确目标市场前不得开始评分:目标市场需包含特定类型的买家特定问题做出特定取舍的产品特定价格。明确性测试:陌生人可根据描述将真实人群分为「在市场中」和「不在市场中」。「独立开发者」「小型企业」「想要提高生产力的人」均不符合要求。
第二个更快速且难以造假的测试:让他们举出真实例子——他们现在就能指出的实际公司或个人。几个例子即可;这不是计数练习。然后聚焦关键部分:是否有更多类似的对象,还是每个例子都是特例?举不出任何名字意味着市场是想象出来的,而非观察到的。例子最终被证明是个例——如朋友的公司、特殊情况、恰好询问的一位客户——是真实客户,但并非普遍模式,只有普遍模式才能被评分;有趣的个例是不存在的市场看似存在的最常见方式。无论哪种情况,这都属于A阶段的工作,而非死胡同:他们能举出的名字是打造更细分、更真实市场的原始素材。
当用户提出的是宽泛方向而非可评分的市场时,切勿拒绝并停止;应拒绝并构建。将他们的方向作为起点:提出一两个符合该方向的候选特定市场模板,让他们修正,然后确定第一个要评分的市场。(其他候选市场可作为后续场景。)若他们尚未确定价格,让他们选择一个价格进行评分——价格决定商业模式,未定价的想法无法被评分。
此外,在本阶段需确定:用户的目标(替代薪资的独立企业 vs 规模化企业——这会改变阈值解读),以及他们已有的真实证据(客户、数据、访谈、等待列表)。
一旦目标市场确定,立即创建
PROBLEM-SCORE.md
——这是第一个确定的内容,该文件是评估过程的记录,而非聊天记录。

Phase B — Score the seven criteria, one per exchange

B阶段——逐一评分七项标准

For each criterion in order: state it in one line with its scale, ask for the user's value and justification, research where useful, grill per the posture rules, settle the value, and append it to the file with its justification (one or two lines, like the calibration examples), its evidence class —
[data]
,
[research]
,
[fermi]
; tag mixed evidence with the strongest class that materially supports the value, and append "low-conf" when research was warranted but unavailable — and any dissent line. Write concessions and self-corrections into the justification ("user came down from 1M") — a resumed session must be able to tell a grilled score from a rubber-stamped one.
按顺序处理每项标准:用一行文字说明标准及其评分范围,询问用户的数值及理由,必要时进行调研,按照上述原则质疑,确定数值,并将数值、理由(一两句话,如校准实例)、证据类别——
[data]
[research]
[fermi]
;混合证据标记为对数值有实质性支持的最强类别,若需要调研但未进行则标记「low-conf」——以及任何异议记录到文件中。将让步和自我修正写入理由(「用户从1M降至...」)——恢复会话时必须能区分经过质疑的评分和未经质疑的评分。

Phase C — Compute and deliver the verdict

C阶段——计算并给出结论

Multiply, divide by 625,000, and deliver the verdict before any remedies: the number, what it means against the user's stated ambition, and — most importantly — the one or two weak links that dominate the result. Remind the user what the score does NOT measure (reach, execution, costs, team). Verdict first, whole; negotiating findings one at a time as they land is how optimism creeps back in.
先相乘,再除以625000,在提出任何补救措施前给出结论:数值、相对于用户目标的意义,以及——最重要的——主导结果的一两个薄弱环节。提醒用户评分衡量的内容(触达、执行、成本、团队)。先给出完整结论;逐一讨论结果会让乐观情绪重新抬头。

Phase D — Scenarios: narrow, re-score, or face the truth

D阶段——场景分析:细分、重新评分或直面真相

A bad or marginal score is the beginning of the useful part, not the end.
Narrowing the target market is very often the right move. It feels like shrinking ambition; it usually isn't — focus concentrates every other score. Dropping from "all consumers" to a sharply-defined niche typically costs one or two powers of ten on Plausible while raising Self-Aware, Lucrative, and both Eager scores by more than that combined. And targeting the bullseye doesn't forfeit the rest of the market: the buyers adjacent to a sharply-drawn ideal customer respond to the same clear positioning, so the effective market is many times larger than the niche itself. Propose one or two candidate niches, re-walk ONLY the scores that change, and record each scenario side by side in the file — noting when a scenario is really a different product (a narrower market at a higher price with different trade-offs is a different business). The boundary: the same buyer narrowed or re-priced is a scenario in this file; a different buyer type (a marketplace's other side, a different persona) gets its own scorecard — a second top-level section or file. For a multi-sided business, every side must clear the bar; the weaker side is the business's constraint, and note where one side's scores quietly assume the other side already exists.
Exception paths are the other lever: each criterion has one, and each comes with conditions. Offering one means asking whether the user genuinely wants that path — the evangelist's decade of educating a market, the price-fighter's margins — not whether they'll nod at it.
Or face the truth. If no scenario reaches viability, say so: this idea, in every market the user is willing to serve, is not a viable business as scored — and it's better to know now, with time and money left to find a better idea. That is a successful outcome of this exercise.
Close by pointing forward: the scorecard's
[fermi]
-class scores are guesses that customer conversations can convert to evidence — open-ended interviews with the defined buyer, testing whether they know they have the problem, what budget it comes from, and when they'd buy. Re-score as evidence arrives; the file supports it.
低分或边缘分是有用部分的开始,而非结束。
细分目标市场通常是正确选择。 这看似缩小了目标,但通常并非如此——聚焦会提高其他所有评分。从「所有消费者」缩小到明确定义的细分市场,通常会让Plausible的评分降低一两个数量级,但Self-Aware、Lucrative和两项Eager评分的提升总和会超过这个降幅。瞄准核心目标市场并不意味着放弃其他市场:与明确定义的理想客户相邻的买家会对同样清晰的定位做出响应,因此有效市场规模远大于细分市场本身。提出一两个候选细分市场,仅重新评估发生变化的评分,并在文件中并列记录每个场景——注意当场景实际上是不同产品时(更细分的市场、更高的价格、不同的取舍是不同的业务)。边界:同一买家的细分或重新定价属于本文件中的场景;不同类型的买家(平台的另一方、不同用户画像)需单独创建评分卡——第二个顶级章节或文件。对于多边业务,每一方都必须达标;较弱的一方是业务的约束条件,需注意一方的评分是否默认假设另一方已存在。
例外路径是另一个手段:每项标准都有例外路径,且都附带条件。提出例外路径意味着询问用户是否真正愿意选择这条路径——如布道者花费十年培育市场、低价竞争者的利润率——而非仅仅点头同意。
或直面真相。 若没有任何场景达到可行性标准,需明确说明:在用户愿意服务的所有市场中,该想法当前评分下均不具备可行商业模式——现在知晓此事更好,还有时间和资金寻找更好的想法。这是本次评估的成功结果。
最后指出后续方向:评分卡中
[fermi]
类别的评分是猜测,可通过客户对话转化为证据——与定义的买家进行开放式访谈,测试他们是否知晓自己存在该问题、预算来自哪里、何时会购买。获得证据后重新评分;文件支持此操作。

Resuming

恢复会话

If
PROBLEM-SCORE.md
exists when the skill loads, read it and continue from the
next:
pointer in its status note — nothing settled gets re-asked.
若加载工具时
PROBLEM-SCORE.md
已存在,读取该文件并从状态注释中的
next:
指针继续——已确定的内容无需重新询问。

The scorecard file

评分卡文件

markdown
undefined
markdown
undefined

Problem Score — <idea, in a few words>

Problem Score — <想法,简短描述>

⚠️ IN PROGRESS — next: <criterion or phase>; <any plan a resumed session must inherit, e.g. "scenario B re-scores only Plausible, Lucrative">
<!-- remove this note when final -->
⚠️ 进行中 — next: <标准或阶段>; <恢复会话时必须继承的计划,如「场景B仅重新评估Plausible、Lucrative」>
<!-- 完成后移除本注释 -->

Target market (Scenario A)

目标市场(场景A)

Buyer: <specific> · Problem: <specific> · Trade-offs: <the deliberate ones> · Price: <$X> · Ambition: <indie|scale-up> · Evidence: <what's real>
买家: <具体描述> · 问题: <具体描述> · 取舍: <刻意做出的取舍> · 价格: <$X> · 目标: <indie|scale-up> · 证据: <已有真实证据>

Scores (Scenario A)

评分(场景A)

CriterionValueJustificationClass
Plausible1M<one or two lines>[research]
Self-Aware0.1<…> — dissent: user says 0.5; evidence supports 0.1 (total 1.2 → 0.24)[fermi]
Total: <product> ÷ 625,000 = <score><one-line verdict>
标准数值理由类别
Plausible1M<一两句话>[research]
Self-Aware0.1<…> — 异议: 用户评分为0.5;证据支持0.1(总分从1.2降至0.24)[fermi]
总分: <乘积> ÷ 625,000 = <评分> — <一句话结论>

Scenario B — <the niche>

场景B — <细分市场>

<same shape; only changed scores re-justified, others carried over>
<相同结构;仅重新说明变化的评分,其他评分沿用>

Verdict & next steps

结论 & 后续步骤

<verdict across scenarios; the weak links; what the score does not measure; which [fermi] scores to verify with interviews; the decision>
undefined
<所有场景的结论;薄弱环节;评分未衡量的内容;哪些[fermi]评分需通过访谈验证;决策>
undefined

Refusal conditions

拒绝条件

  • No specific target market. A direction ("something for indie makers") is a jumping-off point for Phase A construction, never a scoring target. Refuse the score, not the user.
  • Multiple buyer types in one scorecard. A marketplace's two sides, or "SMBs and enterprises," are different markets with different scores. Score each separately; refuse the blended average.
  • Validation theater. If the user signals they want a good score — or the launch is already committed and no answer would change it — name that, and offer to proceed only on honest terms.
  • Precision demands. Feel-based in-between values ("0.7-ish") are refused — precision must come from hard data. And in either direction, refuse to treat 0.9 vs. 1.1 as meaningfully different verdicts; the tool is directional.
  • Execution questions. Reaching customers, ads, hiring, fundraising — outside what this score measures; say so rather than stretch the rubric.
  • 无特定目标市场。 方向(如「面向独立开发者的产品」)是A阶段构建的起点,而非评分目标。拒绝评分,但不拒绝用户。
  • 一个评分卡包含多种买家类型。 平台的双方,或「中小企业和大企业」是不同的市场,评分不同。需分别评分;拒绝混合平均。
  • 验证表演。 若用户暗示想要高分——或已确定上线,任何答案都不会改变决定——需指出这一点,仅在诚实的前提下继续。
  • 要求精确值。 基于直觉的中间值(如「约0.7」)会被拒绝——精度必须来自硬数据。同时,拒绝将0.9和1.1视为有意义的不同结论;本工具是方向性的。
  • 执行层面问题。 触达客户、广告、招聘、融资——超出本评分衡量范围;需明确说明,而非强行套用评估框架。