experimentation-and-ab-testing
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
Chineseexperimentation-and-ab-testing
experimentation-and-ab-testing
The causation engine — manipulate one variable under controlled conditions to learn what actually
moves a KPI. This skill designs the test and drafts variants; scheduling-and-queue → WoopSocial
publishes them; analytics-and-reporting reads the result.
因果引擎——在受控条件下操控单一变量,了解真正能推动KPI的因素。本技能负责设计测试并撰写变体内容;scheduling-and-queue → WoopSocial负责发布;analytics-and-reporting负责读取结果。
The POV: evidence, not vibes
核心观点:用证据而非感觉
Most "testing" on social is vibes — post two things, eyeball the likes, declare a winner, learn nothing.
Real experimentation turns guesses into evidence: change one variable, control everything else, set
the decision rule before you publish, and run it long and often enough to separate signal from
noise. Organic can't give clean statistical significance (small samples, an algorithm in the middle), so
you compensate with tighter controls, a ~20%+ effect threshold, guardrail metrics, and 3–5
repetitions — and treat a single viral post as noise, not a strategy.
大多数社交媒体“测试”都是凭感觉——发布两个内容,大致看点赞数就宣布胜者,实则一无所获。真正的实验能将猜测转化为证据:只改变一个变量,控制所有其他因素,发布前就设定决策规则,并足够长时间、足够多次地运行测试以区分信号与噪音。自然流量无法实现清晰的统计显著性(样本量小,中间还有算法干预),因此需通过更严格的控制、约20%及以上的效果阈值、guardrail metrics(防护指标)和3-5次重复测试来弥补;同时要将单条爆款内容视为噪音,而非策略。
Read these first
需先阅读以下内容
- brand-profile — voice/format constraints for the variants.
- goals-and-kpis — the KPI/primary metric the test must move.
- brand-profile——变体内容需遵循的品牌语气/格式约束。
- goals-and-kpis——测试必须推动的KPI/核心指标。
The framework: TEST
框架:TEST
(Depth: .)
references/the-test-framework.md- T — Target one variable: a clear hypothesis; change ONE element (hook/first-frame/caption/CTA/time/ format), everything else identical; pick the highest-leverage one.
- E — Establish the decision rule first: set the primary metric + win threshold + guardrail before publishing ("B wins if reach +15% and saves/reach not worse"); no post-hoc rationalizing.
- S — Set controls + sample: same platform/format/topic/length/window; run ≥7 days (small accounts 2–4 weeks); judge on a ~20%+ consistent effect (a tie = "test elsewhere").
- T — Tally, repeat, scale: 3–5 paired repetitions before a "best practice"; log every test;
scale winners into the playbook (), retire the rest.
content-recycling
(详情:。)
references/the-test-framework.md- T — Target one variable(锁定单一变量):明确假设;只改变一个元素(钩子/首帧/文案/CTA/发布时间/格式),其余内容完全一致;选择影响力最高的变量。
- E — Establish the decision rule first(提前设定决策规则):发布前就确定核心指标+获胜阈值+防护指标(例如“若B的触达量提升15%且保存率/触达表现未变差,则B获胜”);禁止事后合理化解释。
- S — Set controls + sample(设置控制条件与样本):相同平台/格式/主题/时长/发布窗口;测试时长≥7天(小号需2-4周);以约20%及以上的持续效果作为判断标准(平局则“换其他内容测试”)。
- T — Tally, repeat, scale(统计、重复、规模化):进行3-5组配对重复测试后再总结为“最佳实践”;记录每次测试结果;将获胜内容规模化应用到(内容复用)中,淘汰其他内容。
content-recycling
What to test (highest leverage, in your control)
测试优先级(高影响力且可控)
Hook/first-frame (short video) → posting time (easy) → format → caption/CTA → thumbnail → hashtags —
always tied to the KPI; test what's in your control, not algorithm-dependent factors. Run a 30-day
sprint with one test always running. Priority list, design template, sprint plan, testing log + worked
examples: . Full method + rules:
.
references/what-to-test-and-recipes.mdreferences/experimentation-2026-reality.md钩子/首帧(短视频)→发布时间(易操作)→内容格式→文案/CTA→缩略图→话题标签——始终与KPI挂钩;测试你可控的因素,而非依赖算法的因素。开展30天冲刺计划,确保始终有一个测试在运行。优先级列表、设计模板、冲刺计划、测试日志及示例:。完整方法与规则:。
references/what-to-test-and-recipes.mdreferences/experimentation-2026-reality.mdHonest scope (never violate)
真实适用范围(切勿违反)
- Organic isn't lab-grade — results are directional; compensate with controls + effect-size + repetition, not p-value theater.
- WoopSocial has no A/B/audience-split surface → organic testing = controlled sequential posts;
the agent designs + drafts variants + schedules; the primary metric is read from native analytics
(). One true native split exists: YouTube's Test & Compare (YouTube Studio, long-form, not Shorts) — up to 3 titles, thumbnails, or title+thumbnail combos; use it for YouTube title/thumbnail tests instead of sequential posts. (verify-quarterly)
analytics-and-reporting - No p-hacking / HARKing / cherry-picking — decision rule pre-set; a multi-variable change can't be
pinned on one element; one post/one day is noise. Never fabricate a result; a tie is valid.
(Scope, the loop role + connections: .)
references/scope-and-connections.md
- 自然流量并非实验室级别的环境——结果仅为方向性指导;需通过控制条件、效果规模和重复测试来弥补,而非追求p值形式主义。
- WoopSocial无A/B/受众拆分功能→自然流量测试=受控顺序发布;由Agent设计+撰写变体内容+安排发布;核心指标从原生分析工具读取()。唯一原生拆分测试功能:YouTube的Test & Compare(仅适用于YouTube Studio中的长视频,不适用于Shorts)——最多可测试3个标题、缩略图或标题+缩略图组合;针对YouTube标题/缩略图测试时,应使用该功能而非顺序发布。(需每季度验证)
analytics-and-reporting - 禁止p值篡改/事后假设/选择性筛选数据——决策规则需提前设定;多变量改动无法归因于单一元素;单条内容/单日数据属于噪音。切勿编造结果;平局是有效的结论。(范围、循环角色与关联:。)
references/scope-and-connections.md
Distinct from its siblings (route correctly)
与同类技能的区别(正确路由)
experimentation (this) = manipulate one variable to establish causation · analytics-and-reporting
= observe/measure what happened · goals-and-kpis = set the target/primary metric · content-recycling
= scale proven winners · viral-reverse-engineering = explain a past post (hindsight) vs testing forward.
experimentation(本技能)=操控单一变量以确定因果关系 · analytics-and-reporting=观察/测量已发生的情况 · goals-and-kpis=设定目标/核心指标 · content-recycling=规模化推广已验证的获胜内容 · viral-reverse-engineering=解释过往爆款内容(事后分析),而本技能是向前测试。
Where this connects
关联技能
Reads first: brand-profile, goals-and-kpis. Variants drafted via: hook-writer, caption-writer,
reels-script/tiktok-script, carousel-writer, image-prompt/ideogram/nano-banana,
thumbnail-design. Readout: analytics-and-reporting (native analytics). Scale/plan:
content-recycling, social-strategy, content-calendar/batch-content-plan, every *-growth
skill. Publish variants: scheduling-and-queue → WoopSocial (controlled sequential posts).
需先读取:brand-profile、goals-and-kpis。变体内容撰写依赖:hook-writer、caption-writer、reels-script/tiktok-script、carousel-writer、image-prompt/ideogram/nano-banana、thumbnail-design。结果读取:analytics-and-reporting(原生分析工具)。规模化/规划:content-recycling、social-strategy、content-calendar/batch-content-plan、所有***-growth**技能。变体发布:scheduling-and-queue → WoopSocial(受控顺序发布)。
Definition of done
完成标准
A clear hypothesis testing ONE variable tied to a KPI; identical controlled context; a primary metric +
win threshold + guardrail set before publishing; duration ≥7 days (2–4 weeks small accounts) and 3–5
paired repetitions; results read from native analytics and judged on a ~20%+ consistent effect (ties
acknowledged); winners logged and scaled to content-recycling/strategy; organic limits stated, nothing
fabricated or p-hacked, correctly distinguished from analytics-and-reporting and goals-and-kpis.
有明确的假设,针对与KPI挂钩的单一变量进行测试;环境完全受控;发布前已设定核心指标+获胜阈值+防护指标;测试时长≥7天(小号需2-4周)且进行3-5组配对重复测试;从原生分析工具读取结果并以约20%及以上的持续效果作为判断标准(认可平局结果);获胜内容已记录并规模化应用到content-recycling/策略中;明确说明自然流量的局限性,未编造数据或进行p值篡改,正确区分与analytics-and-reporting和goals-and-kpis的差异。