investigate-anything
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseInvestigate anything
任意对象调查入门
The front door. Everything downstream is faster than the thinking that should
precede it, which is why most bad investigations are not collection failures —
they are framing failures. You can run forty tools against a name and produce a
confident dossier on the wrong person. The work here is deciding what question
you are answering, what would count as an answer, and what would prove you
wrong.
这是调查工作的入门环节。后续所有操作的速度都快于前期应做的思考,这也是大多数失败的调查并非收集环节出错,而是框架设定失误的原因。你可能用四十种工具查询一个姓名,却自信地整理出一份关于错误对象的档案。此环节的核心是确定你要解答的问题、什么可算作答案,以及什么能证明你的结论错误。
Core vocabulary
核心术语
Used across every skill in this repo, defined only here.
- Selector — one identifiable data point: name, handle, email, phone, domain, IP, wallet, hash, plate, IMO, company number.
- Pivot — turning one selector into new selectors (email → breach record → reused handle → forum profile → real name). An investigation is a chain of pivots. Every pivot is also a chance to jump onto a different person.
本仓库中所有技能都会用到以下术语,仅在此处定义:
- Selector(识别符):一个可识别的数据点:姓名、账号、邮箱、电话、域名、IP、钱包地址、哈希值、车牌、IMO编号、公司注册号。
- Pivot(关联跳转):将一个识别符转化为新的识别符(例如:邮箱→泄露记录→复用账号→论坛资料→真实姓名)。调查的本质就是一系列关联跳转的链条。每一次跳转都有可能转向另一个无关对象。
Step 1 — Authorized scope
步骤1:确定授权调查范围
Read ../../ETHICS.md, then write down five things:
- Subject — the specific entity, distinguished from anyone with a similar name. Write the discriminators you will use ("the J. Okonkwo who is a director of company 09xxxxxx", not "J. Okonkwo").
- Objective — see Step 2.
- In bounds — selector types, sources, and whether interaction is allowed.
- Out of bounds — the named things you will not do: logging into anything belonging to the subject, contacting them, family members, medical or religious data, and any selector unrelated to the objective.
- Jurisdiction — whose law governs you, the subject, and the data. In the EU/UK, aggregating scattered public facts about a living person is processing personal data and needs a lawful basis and minimisation.
If you cannot state who authorized this and on what basis, stop. Not "proceed
carefully" — stop. An unauthorized investigation cannot be fixed later by a
good report.
Done when all five are written down and you can name the specific action
that would put you out of bounds.
阅读../../ETHICS.md,然后写下以下五项内容:
- 调查对象:明确的实体,需与同名对象区分开。写下你将使用的区分标识(例如:“担任公司注册号09xxxxxx董事的J. Okonkwo”,而非“J. Okonkwo”)。
- 调查目标:详见步骤2。
- 允许范围:允许使用的识别符类型、信息来源,以及是否允许与调查对象产生交互。
- 禁止范围:明确列出你不会执行的操作:登录调查对象的任何账号、联系调查对象或其家属、获取医疗或宗教相关数据,以及任何与调查目标无关的识别符。
- 管辖权限:约束你、调查对象及数据的所属法律。在欧盟/英国,汇总关于自然人的分散公开信息属于处理个人数据,需具备合法依据并遵循最小化原则。
如果你无法说明谁授权了此次调查以及授权依据,请立即停止。不是“谨慎推进”,而是直接停止。未经授权的调查,即便后续报告质量再高也无法弥补其合规性缺陷。
完成标志:五项内容全部书面记录,且你能明确说出哪些具体操作会超出调查范围。
Step 2 — Frame an answerable question
步骤2:构建可解答的问题
"Find out about X" is not an objective; it has no stopping condition, so it
terminates when you get bored or when you find something that feels like a
result. Rewrite the request until it has a subject, a decision it feeds, and a
condition that would settle it.
| Vague | Answerable |
|---|---|
| Investigate this company | Does this supplier have an undisclosed owner subject to sanctions, and does it operate from the address on the invoice? |
| Who is this account | Is the operator of |
| Look into this domain | Is |
Then write the negative: what finding would mean no. If nothing could, the
question is unfalsifiable and you will confirm it whatever you see.
Done when the objective is one sentence, and you have written what a "no"
answer would look like.
“调查X的相关情况”并非明确的目标;它没有终止条件,只会在你感到厌烦或找到看似有价值的结果时停止。重新梳理需求,直到它包含明确的调查对象、服务的决策事项,以及可判定任务完成的条件。
| 模糊表述 | 可解答问题 |
|---|---|
| 调查这家公司 | 该供应商是否存在未披露的受制裁所有者?其实际经营地址是否与发票上的地址一致? |
| 这个账号属于谁 | |
| 调查这个域名 | |
然后写下否定情况:什么样的发现会得出“否”的结论。如果不存在这样的情况,说明该问题无法证伪,无论你发现什么都会去印证预设结论。
完成标志:调查目标凝练成一句话,且你已明确记录“否”的结论对应的具体情况。
Step 3 — Write the collection plan before collecting
步骤3:收集前制定收集计划
Aimless pivoting feels productive because every pivot yields something. A
collection plan is the list of questions, each mapped to the source most
likely to settle it, ranked by cost and intrusiveness — so you notice when
you're three hours into an interesting branch that answers nothing.
For each question: the indicator that would answer it, the source or skill that
produces it, whether it is passive, and what you do if it comes back empty.
Passive-first, always: exhaust archives, registries, and logs before anything
that touches the subject. Plans are revised as you learn — the point is that
deviations become visible.
Done when each objective question has a named source or skill and a
first/fallback order.
漫无目的的关联跳转看似高效,因为每次跳转都会产生新信息。收集计划是一份问题清单,每个问题都对应最有可能解答它的信息来源,并按成本和侵入性排序——这样你就能及时发现自己在一个有趣但无关的分支上浪费了三个小时。
针对每个问题:记录能解答它的指标、生成该指标的来源或技能、是否为被动收集,以及如果没有结果该如何处理。始终优先采用被动收集:在接触调查对象的任何信息前,先穷尽档案、注册信息和日志。计划可随调查进展调整——关键是让偏离计划的行为变得可见。
完成标志:每个目标问题都对应明确的来源或技能,且已确定优先/备选顺序。
Step 4 — Route by starting selector
步骤4:根据初始识别符选择工作流
Type the workflow skill's name.
| You have | Run |
|---|---|
| A person's name or real identity | |
| A company, brand, or invoice entity | |
| A domain, website, or IP | |
| A username or handle | |
| An email address | |
| A phone number | |
| A photo or video to locate or verify | |
| A social profile you already attribute | |
No clear starting point? Start with the selector that is both unique and
indexed — email and domain beat name and handle, because names collide and
handles are claimed by strangers.
Technique skills load themselves when you describe what you are doing. Before
touching anything the subject controls, run .
Keep the case in a graph from the first pivot: .
investigate-without-getting-madegraph-the-networkDone when the routed workflow has been run and its findings are recorded
with sources.
输入对应工作流技能的名称。
| 你拥有的信息 | 运行的技能 |
|---|---|
| 个人姓名或真实身份 | |
| 公司、品牌或发票主体 | |
| 域名、网站或IP | |
| 用户名或账号 | |
| 邮箱地址 | |
| 电话号码 | |
| 需定位或验证的照片/视频 | |
| 已完成归因的社交账号 | |
没有明确的初始识别符?选择同时具备唯一性和可索引性的识别符——邮箱和域名优于姓名和账号,因为姓名容易重复,账号可能被陌生人注册。
当你描述正在执行的操作时,技术技能会自动加载。在接触调查对象控制的任何内容前,先运行技能。从第一次关联跳转开始,用图谱记录案件:。
investigate-without-getting-madegraph-the-network完成标志:已运行选定的工作流,且所有发现都已记录来源。
Step 5 — Grade sources as you collect, not afterwards
步骤5:收集时同步评级来源,而非事后评级
Use the Admiralty (NATO-style) scheme: a letter for the source and a
number for the information, graded independently, on every item.
- Letter A–F: the source's track record and access. A = reliable history, no doubt of authenticity; F = cannot be judged.
- Number 1–6: whether the content is confirmed by other independent sources, and whether it is logical in itself. 1 = confirmed elsewhere; 6 = cannot be judged.
The independence matters: a corporate registry filing is B2, a well-run
newspaper report of that filing is B2 at best, and an anonymous forum post
repeating the newspaper is D3 — not new corroboration. Grade the source you
actually touched, not the source it claims to have. Full grid, worked
examples, and the common misgradings:
reference/source-grading.md.
Done when every retained finding carries a two-character grade.
采用海军部(北约式)评级体系:对每条信息分别标注来源评级字母和信息评级数字。
- 字母A–F:代表来源的可信度和获取渠道。A=可信度高,真实性无疑问;F=无法判定。
- 数字1–6:代表内容是否得到其他独立来源的佐证,以及自身是否符合逻辑。1=已被其他来源证实;6=无法判定。
独立性至关重要:企业注册文件的评级为B2,报道该文件的知名报纸文章评级最高为B2,而重复该报道的匿名论坛帖子评级为D3——这并非新的佐证。请对你实际接触的来源进行评级,而非它声称的来源。完整评级表、实操示例及常见错误评级请参考:reference/source-grading.md。
完成标志:所有保留的发现都带有两位字符的评级。
Step 6 — Test hypotheses against each other
步骤6:用竞争假设相互验证
Analysis of Competing Hypotheses (ACH) exists because the natural mode of
investigation — pick the likeliest story, look for support — always succeeds.
Support is easy to find for any plausible story.
List every hypothesis including the boring ones ("it is a different person with
the same name", "the account was sold", "the shared IP is shared hosting").
Build a matrix of evidence against hypotheses, and for each cell ask only
whether the evidence is consistent with that hypothesis. Then work by column:
the hypothesis with the fewest inconsistencies wins, not the one with the most
support. Evidence consistent with every hypothesis has no diagnostic value
— the subject having a LinkedIn does not distinguish anything. A handful of
diagnostic items beats a hundred consistent ones. Worksheet and a filled
example: reference/ach-worksheet.md.
Done when you have listed at least one hypothesis you did not want and
recorded what evidence would refute your favoured one.
竞争假设分析法(ACH)的存在是因为调查的自然模式——选择最可能的结论,寻找支持证据——总会“成功”。任何看似合理的结论都很容易找到支持证据。
列出所有假设,包括看似无关的假设(例如:“这是同名的另一个人”“账号已被出售”“共享IP来自共享主机”)。构建证据与假设的矩阵,针对每个单元格仅判断证据是否与该假设一致。然后按列分析:不一致情况最少的假设获胜,而非支持证据最多的假设。与所有假设都一致的证据无诊断价值——调查对象拥有LinkedIn账号无法区分任何情况。少量具有诊断价值的证据胜过大量一致的证据。工作表及填写示例请参考:reference/ach-worksheet.md。
完成标志:你已列出至少一个你不希望成立的假设,并记录了能推翻你偏好假设的证据。
Where this goes wrong
常见误区
Confirmation bias, OSINT edition. You are given a name and told the person
works in logistics. You find a logistics profile and stop asking whether it is a
different person with the same name. The tell is that your discriminators
disappear once you find a candidate — you selected on them to find, then
stopped applying them to test. Fix: before you search, write the attributes
the true subject must have and the ones they cannot have; check every candidate
against both. Rejections are findings and belong in the report.
Circular reporting. Three sources agree, so you grade it confirmed. All
three copied one blog post, or all three pull from the same aggregator or the
same leaked dataset. This is the single most common cause of confident wrong
attribution, and it is invisible unless you look for it. For each corroborating
source, find its origin: check publication dates in order, look for identical
phrasing or a copied typo, and check whether the "independent" people-search
sites resell the same broker feed — see . Independent
means different collection, not different websites.
dig-through-data-brokersStale data presented as current. Registries, WHOIS, and broker records
carry the date they were captured, not today's truth. Record the observation
date next to every fact; archive the page via .
read-deleted-pagesSelector drift. Each pivot carries the risk that you have changed people. A
chain of five pivots each 90% likely is a coin flip. Re-anchor: after every
pivot, state which confirmed selector ties the new one to the subject.
Tool output as evidence. An enumerator's hit list, a breach aggregator's
match, a face-search score — these are leads. The tool did not verify identity;
it matched a string or a vector.
OSINT版确认偏差:有人给你一个姓名,告诉你此人从事物流行业。你找到一个物流从业者的资料后,就不再核实是否是同名的另一个人。这种情况的特征是,一旦找到候选对象,你之前设定的区分标识就被抛之脑后——你用这些标识去寻找对象,却不用它们去验证对象。解决方法:搜索前,写下目标对象必须具备的属性和绝对不能有的属性;每找到一个候选对象都要对照这两组属性检查。排除候选对象的过程也是调查结果的一部分,需写入报告。
循环报道:三个来源的信息一致,所以你判定为已确认。但实际上这三个来源都复制了同一篇博客文章,或者都来自同一个数据聚合器或泄露数据集。这是导致自信地得出错误归因的最常见原因,除非刻意追查否则难以发现。对于每个佐证来源,找到其信息源头:按发布日期排序检查,寻找相同措辞或复制的拼写错误,核实那些“独立”的人物搜索网站是否转售同一经纪人的数据源——可参考技能。“独立”指的是不同的收集渠道,而非不同的网站。
dig-through-data-brokers将过时数据呈现为当前数据:注册信息、WHOIS记录和经纪人数据标注的是采集日期,而非当前的真实情况。为每个事实记录观察日期;通过技能存档页面。
read-deleted-pages识别符偏移:每一次关联跳转都有可能转向无关对象。五次跳转每次准确率90%的话,最终结果的准确率就像抛硬币一样低。解决方法:每次跳转后,明确说明哪个已确认的识别符将新识别符与目标对象关联起来。
将工具输出当作证据:枚举工具的命中列表、泄露数据聚合器的匹配结果、人脸搜索得分——这些都只是线索。工具并未验证身份,它只是匹配了字符串或特征向量。
Confidence grading
置信度评级
Applies repo-wide unless a technique skill says otherwise.
- Confirmed — two or more genuinely independent sources (different collection, not different sites), or one authoritative primary record such as a signed registry filing, plus nothing contradicting.
- Probable — one strong source, or several weak ones that survived a circular-reporting check, with the alternative hypotheses tested and weaker.
- Unconfirmed — a single uncorroborated lead. Say so in the report; do not quietly promote it because later text depends on it.
- Rejected — contradicted. Record it and why.
除非某一技术技能另有说明,否则本评级适用于整个仓库的所有内容。
- 已确认:两个或多个真正独立的来源(不同的收集渠道,而非不同网站),或一份权威的原始记录(如签署的注册文件),且无矛盾信息。
- 大概率:一个可靠来源,或多个通过循环报道检查的弱来源,且已测试过其他假设并证明其可能性更低。
- 未确认:单一未佐证的线索。需在报告中明确说明;不要因为后续内容依赖它就悄悄提升其置信度。
- 已排除:存在矛盾信息。记录该情况及排除原因。
When to stop
何时停止调查
Stop when the objective question is answered to the confidence the decision
requires, or when you can document that available open sources cannot answer it
and name what would (a records request, a subpoena, interviews). "We could not
establish X, having checked A, B and C" is a deliverable, and often the honest
one. Stop also when you cross a scope boundary — that is a re-authorization
event, not a judgement call to make mid-flow.
Done when the objective is answered or documented as unanswerable, and
has been run.
write-the-intel-brief当调查目标问题已得到决策所需置信度的答案,或者你能证明现有开源资源无法解答该问题并明确所需的其他手段(如调取记录、传票、访谈)时,停止调查。“我们已检查A、B、C,但无法确定X”也是一项交付成果,而且往往是最诚实的结论。当你超出调查范围边界时也需停止——这需要重新授权,而非在调查过程中自行判断。
完成标志:调查目标已得到解答或被记录为无法解答,且已运行技能。
write-the-intel-briefWorked example
实操示例
Objective: is the supplier on this invoice controlled by the
ex-director of a barred entity? Discriminator: a director DOB month/year.
Nordvale Tradingx-ray-a-companyCompeting hypotheses: (1) same person, (2) different A. Kestrel, (3) name used
as a nominee. Diagnostic item: the barred entity's filings and Nordvale's list
the same unusual accountancy firm, and the domain in the invoice footer shares a
registrant email with the barred entity's old site ().
That is inconsistent with (2), consistent with (1) and (3). Reported as
probable, with the nominee hypothesis flagged as untested, and the gap named:
beneficial ownership is not disclosed in that jurisdiction.
who-owns-this-domain调查目标:发票上的供应商是否由某受限实体的前董事控制?区分标识:董事的出生年月。
Nordvale Trading运行技能返回的注册记录显示,董事为“A. Kestrel”,出生年月匹配,地址为邮件转发服务地址。两个人物搜索网站和一个商业名录网站都显示A. Kestrel的家庭地址相同——看似相互佐证,但追查后发现这三个来源都包含同一经纪人数据源中的街道名称拼写错误。属于循环报道,评级为D3并被排除。
x-ray-a-company竞争假设:(1) 同一人;(2) 另一位A. Kestrel;(3) 姓名被用作代持。关键诊断信息:受限实体的文件和Nordvale的文件列出了同一家不常见的会计师事务所,且发票页脚的域名与受限实体旧网站的注册邮箱相同(通过技能核实)。这与假设(2)矛盾,与假设(1)和(3)一致。最终报告结论为“大概率”,同时标注代持假设未得到验证,并明确指出信息缺口:该司法辖区未披露实际所有权信息。
who-owns-this-domainPivots
关联跳转
Every workflow feeds while it runs and
at the end. Selector-to-skill routing is Step 4; the
technique skills each list their own pivots.
graph-the-networkwrite-the-intel-brief每个工作流运行时都会同步到技能,结束时会运行技能。步骤4是根据识别符选择对应技能;各技术技能会列出自身支持的关联跳转。
graph-the-networkwrite-the-intel-briefLegal notes
法律注意事项
Public availability is not permission. Data-protection law applies to
aggregation of public personal data, and the aggregate is more sensitive than
any part. Terms of service govern automated collection even where the data is
public, and scraping disputes turn on authorization and contract, not on
whether the page was visible. Never authenticate to, probe, or send traffic at
systems belonging to the subject without written authorization.
公开可获取并不代表有权使用。数据保护法适用于汇总公开个人数据的行为,且汇总后的信息比单个部分更敏感。即使数据是公开的,自动化收集也需遵守服务条款,抓取争议的判定依据是授权和合同,而非页面是否可见。未经书面授权,切勿对调查对象所属系统进行身份验证、探测或发送流量。