scripting-and-storyboarding
The pre-production system — start from the spine, write both columns, envision every beat, number the
shots, and edit on paper first. The agent writes the plan; the human shoots and judges; WoopSocial
publishes the finished video. (Craft skill — no tool file.)
The POV: videos are won in pre-production — the cheapest edit is on paper
Chaotic shoots and endless edits are a pre-production deficit, not a talent deficit: the plan is "where you
catch problems before they become expensive fixes on set or in post." Four truths carry this skill. (1) A
words-only script plans half the video — the visual half gets improvised, and improvisation defaults to a
static talking face; the two-column AV script (audio | visual) writes what's seen, beat by beat, so the
change-every-few-seconds rhythm is planned, not hoped for. (2) The storyboard is a decision document, not
art — its only question is "what's on screen at this beat?", answered by a shot table, stick figures, or AI
previz frames (2026 pipelines make character-consistent boards in minutes — the tested benchmark: 12 frames
from ~3 hours to 40–60 minutes, with explicit camera intent and a character bible as the craft rules).
(3) The shot list regrouped BY SETUP is the batching trick — two weeks of content in one afternoon comes
from sorting shots by location/framing/outfit, never by video. (4) Runtime math doesn't negotiate —
~130–150 wpm means a 300-word script is a 2-minute video; the stopwatch pass and the paper cut happen before
the shoot. And the 2026 twist: AI video sequences need more pre-production, not less — every panel becomes
a generation's shot brief, and previz frames cost cents where generations cost credits.
Read these first
- idea-generation-and-ideation — the idea being productionized.
- short-form-video-script (or the long-form skill) — the retention spine this turns into production.
- brand-profile + design-and-templates — visual language, end cards.
The framework: SCENE
(Depth:
references/the-scene-framework.md
.)
- S — Start from the spine: one job, audience, format target, beat outline — before any words.
- C — Columns: audio + visual: the two-column AV script with [VO]/[SFX]/[Action] tags; if the visual isn't
written, it will be improvised; no "me talking" twice in a row.
- E — Envision every beat: the cheapest board that answers "what's on screen" — shot table → stick figures
→ AI previz frames (explicit camera intent, character bible + ref image, micro-beats, directional not
final); skip boards honestly where a shot list suffices; for AI video, each panel = the shot brief.
- N — Number the shots: shot list with setups → regroup by setup for batching (sequence setups, buffer
takes on hooks, label footage ); runtime math + feasibility pass.
- E — Edit on paper first: table read with a stopwatch, cut the middle not the spine, the honesty pass
(no staged-as-candid, no lifted scripts) — then hand off to the shoot, the edit, and WoopSocial.
The reality (verify-quarterly)
Stable craft: the two-column AV script (broadcast standard), ~130–150 wpm runtime math, the
storyboard-as-decision-document, the by-setup batch regroup. The 2026 AI layer (attribute): script→board
pipelines matured (Boords — free tier, sign-off layer; Storyboarder.ai — 250K+ creators, animatics; LTX,
mStudio, Studiovity; free Wonderunit); character consistency is now baseline; the practitioner benchmark cut
12-frame boards from ~3 hrs to 40–60 min using LLM-beats → visual directives → seeded frames — with tags
([VO]/[SFX]/[Action]), explicit camera intent (
"the AI will guess — don't make it"), locked palette, and a
character bible; boards stay directional; traditional illustration runs ~$50–300/frame (the economics behind
the shift). AI-video sequences: panel = shot brief (luma's DREAM); board first, generate second.
Attribute
all; verify-quarterly. Full detail:
references/scripting-and-storyboarding-2026-reality.md
; the templates
and two worked examples:
references/templates-and-examples.md
.
Honest scope (never violate)
- The agent writes every planning artifact (beats, AV script, board/briefs, shot list, batch plan, timing
pass, footage map); the human shoots/generates, judges every take, and approves — the agent cannot see
footage or operate a camera and never fabricates "that take works." AI previz follows the image rules
(no unpermitted likeness; original/consented characters; disclosure if frames publish).
- Production integrity: no staged-as-candid content (actors posing as unaffiliated strangers = deceptive
endorsement; skits are fine disclosed); no lifted scripts (structure study yes, verbatim no); honest
runtime/feasibility. WoopSocial publishes the finished video; it does not script, storyboard, shoot, or
edit. (Full scope:
references/scope-and-connections.md
.)
Distinct from its siblings (route correctly)
scripting-and-storyboarding (this) = the pre-production system · short-form-video-script = the
retention-words craft (WATCH writes the spine; SCENE productionizes it) · talking-head-and-piece-to-camera
= the on-camera delivery · capcut / descript = the edit executing the AV script's visual plan · luma /
ai-video = the generations whose shot briefs the panels become · flux / image-prompt = previz frames +
reference stills · storytelling-and-narrative = the narrative angle/WHAT this schedules into shots ·
idea-generation-and-ideation = supplies the idea.
Where this connects
Reads first: idea-generation-and-ideation + short-form-video-script/long-form + brand-profile +
design-and-templates. Feeds: the shoot day (the human), talking-head-and-piece-to-camera, luma
(shot briefs), capcut/descript (footage map + AV script), content-calendar (the batch). Publishes via:
the finished video → scheduling-and-queue → WoopSocial. Measure with: shoot efficiency + edit speed +
published retention via analytics-and-reporting — never fabricated.
Definition of done
A shootable, editable plan: the spine confirmed (one job, beat outline from the script skill), the two-column
AV script written with every visual beat specified ([VO]/[SFX]/[Action] tags; no "me talking" twice in a row),
the storyboard produced at the cheapest sufficient fidelity (shot table, stick figures, or AI previz frames
with explicit camera intent + a character bible — directional, never final art; skipped honestly where a shot
list suffices), a numbered shot list regrouped by setup with the batch plan, buffer takes, and a labeled
footage map, the runtime verified by stopwatch math (~130–150 wpm) and cut on paper before the shoot, and the
honesty pass held (no staged-as-candid, no lifted scripts, feasibility stated straight); AI-video panels
written as generation shot briefs with previz-before-credits economics; the human shooting and judging,
the edit receiving a clean handoff, and the finished video publishing via WoopSocial; nothing staged as
real, nothing plagiarized, no fabricated production claims; and correctly distinguished from
short-form-video-script, talking-head-and-piece-to-camera, capcut/descript, and luma.