seedance-ai-avatar
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseAI Avatar — Digital Persona Video Prompts
AI Avatar——数字角色视频提示词
A complete system for crafting Seedance 2.0 prompts on Higgsfield that feature AI avatars, digital personas, virtual presenters, and synthetic characters as the primary subject. This skill covers everything from photorealistic digital humans to stylized 3D characters, including hook patterns, environment design, camera work, lighting, and audio direction.
一套完整的系统,用于在Higgsfield平台上撰写以AI Avatar、数字角色、虚拟主持人和合成角色为核心主题的Seedance 2.0提示词。该技能涵盖从photorealistic human(超写实人类)到stylized 3D(风格化3D)角色的所有内容,包括钩子模式、环境设计、运镜、灯光和音频指导。
1. Input Specs
1. 输入规范
Before generating prompts, gather the following from the user:
Character definition
- Avatar style: photorealistic human / stylized 3D / 2D animated / abstract-geometric
- Gender presentation, approximate age, and visual identity descriptors (if applicable)
- Skin tone, hair, wardrobe notes, or reference aesthetics
- Personality register: corporate, playful, authoritative, futuristic, warm
Content purpose
- Use case: product demo, brand spokesperson, educational content, virtual influencer post, sales video, onboarding clip
- Platform destination: Instagram Reel, YouTube, LinkedIn, TikTok, website hero
- Desired emotional response: trust, excitement, curiosity, aspiration
Technical requirements
- Clip duration: 2–4 sec / 6–10 sec / 15–30 sec
- Aspect ratio: 9:16 vertical / 16:9 landscape / 1:1 square
- Motion intensity: static with subtle animation / active gestures / high-energy movement
Brand constraints
- Color palette or forbidden colors
- Brand keywords or phrases the avatar should embody visually
- Any style references (films, games, other AI avatars)
生成提示词前,请从用户处收集以下信息:
角色定义
- Avatar风格:photorealistic human / stylized 3D / 2D animated / abstract-geometric
- 性别呈现、大致年龄和视觉特征描述(如有)
- 肤色、发型、着装细节或参考美学风格
- 性格定位:商务风、活泼风、权威风、未来风、温暖风
内容用途
- 使用场景:产品演示、品牌代言人、教育内容、虚拟网红帖、销售视频、入职培训片段
- 投放平台:Instagram Reel、YouTube、LinkedIn、TikTok、网站首屏
- 期望情感反馈:信任、兴奋、好奇、向往
技术要求
- 片段时长:2–4秒 / 6–10秒 / 15–30秒
- 宽高比:9:16竖屏 / 16:9横屏 / 1:1正方形
- 动作强度:静态带微动画 / 主动手势 / 高能量动作
品牌约束
- 调色板或禁用颜色
- Avatar视觉上需体现的品牌关键词或短语
- 任何风格参考(电影、游戏、其他AI Avatar)
2. Philosophy — Why AI Avatars Are the Future of Scalable Content
2. 理念——为何AI Avatar是可扩展内容的未来
The uncanny valley problem — and how to avoid it
The uncanny valley is the psychological discomfort triggered when a synthetic human looks almost real but falls short in ways the brain cannot consciously articulate. Eyes that do not quite track. Skin that lacks micro-texture. Movement that is smooth but wrong. The brain pattern-matches to illness, deception, or death.
There are two reliable exits from the uncanny valley:
Exit 1: Go fully photorealistic. Push render quality, skin subsurface scattering, eye moisture, micro-expression, and hair fidelity past the threshold where the brain accepts the character as real. This requires explicit prompt language: "hyper-photorealistic skin texture," "subsurface light scattering on face," "wet corneal reflection," "micro-expressions," "natural breath movement."
Exit 2: Commit to stylization. A clearly stylized character — 3D cartoon, cel-shaded, geometric — never triggers uncanny valley because the brain never expected biological realism. Own the artifice. Make it beautiful and internally consistent.
The danger zone is the middle ground: a character that is clearly rendered but aspires to realism without achieving it. Avoid half-measures. Either push all the way to photorealism or lean fully into a defined non-realistic aesthetic.
Additional uncanny valley avoidance rules:
- Never describe eye movement without describing blink timing
- Always specify natural weight and momentum in gestures
- Avoid "perfect symmetry" — real faces are asymmetric
- Include micro-imperfections: a strand of hair out of place, slight asymmetry in the smile
- Ground the avatar with realistic environmental interaction: light falloff, shadow contact, hair affected by subtle air movement
恐怖谷问题——以及如何规避
恐怖谷指的是当合成人类看起来近乎真实,但在大脑无法有意识地表达的方面存在不足时,引发的心理不适。比如眼神追踪不自然、皮肤缺乏微纹理、动作流畅但违和。大脑会将其与疾病、欺骗或死亡模式匹配。
有两种可靠的方法可以摆脱恐怖谷:
方法一:完全走超写实路线。将渲染质量、皮肤次表面散射、眼部湿润度、微表情和毛发逼真度提升到大脑认可角色为真实的阈值以上。这需要明确的提示词语言:“hyper-photorealistic skin texture(超写实皮肤纹理)”、“subsurface light scattering on face(面部次表面光散射)”、“wet corneal reflection(角膜湿润反光)”、“micro-expressions(微表情)”、“natural breath movement(自然呼吸动作)”。
方法二:坚持风格化。明确风格化的角色——3D卡通、赛璐珞着色、几何风格——永远不会触发恐怖谷,因为大脑从未期望其具备生物真实性。接受其人工属性,使其美观且内部一致。
危险区域是中间地带:一个明显经过渲染但追求真实感却未达标的角色。避免折中方案。要么完全推向超写实,要么完全采用明确的非写实美学风格。
额外的恐怖谷规避规则:
- 描述眼球运动时,务必同时描述眨眼时机
- 始终指定手势的自然重量和动量
- 避免“完美对称”——真实的面部都是不对称的
- 加入微瑕疵:一缕凌乱的头发、微笑时轻微的不对称
- 通过真实的环境互动让Avatar更真实:光线衰减、接触阴影、头发受微风影响
3. Avatar Style Spectrum
3. Avatar风格谱系
Photorealistic Human
Photorealistic Human
The most persuasive style for commercial and B2B content. Optimized for spokesperson videos, testimonials, and brand ambassadors.
Prompt markers: "photorealistic digital human," "hyper-realistic skin with visible pores," "subsurface scattering skin shader," "natural asymmetry in facial features," "wet corneal highlight," "professional studio grooming," "fine hair strands rendered individually"
Best for: SaaS product demos, financial services, healthcare, e-commerce spokesperson
Risks: Highest uncanny valley exposure — demand maximum fidelity or shift styles
最适合商业和B2B内容的风格。优化用于代言人视频、推荐视频和品牌大使内容。
提示词标记:“photorealistic digital human(超写实数字人类)”、“hyper-realistic skin with visible pores(可见毛孔的超写实皮肤)”、“subsurface scattering skin shader(次表面散射皮肤着色器)”、“natural asymmetry in facial features(面部特征自然不对称)”、“wet corneal highlight(角膜湿润高光)”、“professional studio grooming(专业工作室妆容)”、“fine hair strands rendered individually(单独渲染的细发丝)”
最佳适用场景:SaaS产品演示、金融服务、医疗保健、电商代言人
风险:恐怖谷暴露风险最高——要求最高逼真度或切换风格
Stylized 3D Character
Stylized 3D Character
Distinct personality with artistic intentionality. Strong brand identity potential. Avoids uncanny valley entirely.
Prompt markers: "stylized 3D character with smooth simplified geometry," "large expressive eyes," "exaggerated proportions with cartoon appeal," "cel-shaded render," "clean edge lighting," "Pixar-adjacent character design," "bold color palette skin tones"
Best for: Consumer brands, gaming, entertainment, apps, youth-oriented products
Risks: May feel low-trust for serious B2B content — calibrate style to audience
具有艺术目的性的独特个性。强大的品牌塑造潜力。完全规避恐怖谷。
提示词标记:“stylized 3D character with smooth simplified geometry(几何平滑简化的风格化3D角色)”、“large expressive eyes(大而富有表现力的眼睛)”、“exaggerated proportions with cartoon appeal(具有卡通吸引力的夸张比例)”、“cel-shaded render(赛璐珞着色渲染)”、“clean edge lighting(清晰边缘光)”、“Pixar-adjacent character design(皮克斯风格相近的角色设计)”、“bold color palette skin tones(大胆调色板肤色)”
最佳适用场景:消费品牌、游戏、娱乐、应用程序、面向年轻人的产品
风险:对于严肃的B2B内容可能缺乏可信度——根据受众调整风格
2D Animated
2D Animated
Flat or semi-flat illustrated character with motion. Highly distinctive, excellent for social media and explainer content.
Prompt markers: "2D animated character in motion graphics style," "flat illustration aesthetic," "bold outline strokes," "limited animation with expressive poses," "vector art quality," "anime-influenced design," "looping idle animation"
Best for: EdTech, productivity tools, startup explainers, social media content
Risks: Hard to convey facial nuance — lean on body language and motion design
带有动作的扁平或半扁平插画角色。极具辨识度,非常适合社交媒体和讲解类内容。
提示词标记:“2D animated character in motion graphics style(动态图形风格的2D动画角色)”、“flat illustration aesthetic(扁平插画美学)”、“bold outline strokes(粗轮廓线条)”、“limited animation with expressive poses(具有表现力姿势的有限动画)”、“vector art quality(矢量艺术质感)”、“anime-influenced design(受动漫影响的设计)”、“looping idle animation(循环待机动画)”
最佳适用场景:教育科技、生产力工具、初创企业讲解视频、社交媒体内容
风险:难以传达面部细微表情——依赖肢体语言和动态设计
Abstract / Geometric
Abstract / Geometric
The avatar is not a humanoid but a geometric entity — particle systems, fluid simulations, crystalline structures — that communicates through movement and light. Most experimental.
Prompt markers: "abstract digital entity formed from particles of light," "geometric crystalline structure with inner luminescence," "fluid simulation avatar morphing between states," "data visualization made sentient," "void-space entity with no fixed form"
Best for: Tech companies, AI products, concept content, high-concept brand films
Risks: Low relatability — pair with clear voice-over or text overlay
Avatar不是人形,而是几何实体——粒子系统、流体模拟、晶体结构——通过运动和光线进行交流。最具实验性。
提示词标记:“abstract digital entity formed from particles of light(由光粒子构成的抽象数字实体)”、“geometric crystalline structure with inner luminescence(具有内部发光的几何晶体结构)”、“fluid simulation avatar morphing between states(在不同状态间变形的流体模拟Avatar)”、“data visualization made sentient(具有感知能力的数据可视化)”、“void-space entity with no fixed form(无固定形态的虚空实体)”
最佳适用场景:科技公司、AI产品、概念内容、高概念品牌影片
风险:亲和力低——搭配清晰的旁白或文字叠加
4. 2-Second Hook Patterns
4. 2秒钩子模式
The first two seconds determine watch-through rate. For avatar content, there are four proven hook architectures.
前两秒决定完播率。对于Avatar内容,有四种经过验证的钩子架构。
The Glitch-to-Perfect
glitch-to-perfect(从故障到完美)
Opens with visual corruption — static, scan lines, fragmented pixels, digital artifacts — then resolves into a clean, confident avatar. Communicates: synthetic origin, premium output.
Prompt language: "opens with heavy RGB chromatic aberration and scan line distortion, digital glitch artifacts fragmenting the face, resolves at 1.5 seconds into pristine clarity, sharp focus snapping to perfect render quality"
Use when: Launching an AI product, announcing new technology, establishing digital-native brand identity
以视觉损坏开场——静态噪点、扫描线、像素碎片、数字伪影——然后过渡到清晰、自信的Avatar。传达:合成起源、优质输出。
提示词语言:“opens with heavy RGB chromatic aberration and scan line distortion, digital glitch artifacts fragmenting the face, resolves at 1.5 seconds into pristine clarity, sharp focus snapping to perfect render quality(开场带有严重的RGB色差和扫描线失真,数字故障伪影使面部碎片化,1.5秒时恢复至极致清晰,焦点快速切换到完美渲染质量)”
适用场景:AI产品发布、新技术公告、建立数字原生品牌形象
The Digital Birth
Digital Birth(数字诞生)
The avatar materializes from nothing — particles assembling, wireframe filling, data streams converging into form. Communicates: creation, emergence, future.
Prompt language: "avatar assembles from scattered luminous particles converging to a central point, wireframe mesh fills with texture and color, form resolves from abstract to fully rendered in under two seconds, volumetric glow during formation"
Use when: Product launches, brand debut videos, origin story content, intro sequences
Avatar从无到有——粒子聚集、线框填充、数据流汇聚成形态。传达:创造、涌现、未来。
提示词语言:“avatar assembles from scattered luminous particles converging to a central point, wireframe mesh fills with texture and color, form resolves from abstract to fully rendered in under two seconds, volumetric glow during formation(Avatar由分散的发光粒子汇聚到中心点组装而成,线框网格填充纹理和颜色,形态在两秒内从抽象变为完全渲染,形成过程中带有体积光)”
适用场景:产品发布、品牌亮相视频、起源故事内容、开场序列
The Style Morph
Style Morph(风格转换)
The avatar rapidly shifts visual style — sketch to 3D to photorealistic, or multiple artistic styles in sequence — before landing on a final defined look. Communicates: versatility, creativity, range.
Prompt language: "avatar transitions through four visual styles in rapid succession: pencil sketch, flat 2D illustration, stylized 3D, photorealistic render, each held for 0.4 seconds, final state held as primary look, transitions use dissolve with chromatic fringing"
Use when: Creative agencies, design platforms, AI tools, content creation products
Avatar快速切换视觉风格——从草图到3D再到超写实,或依次切换多种艺术风格——最终定格在确定的外观。传达:多功能性、创造力、多样性。
提示词语言:“avatar transitions through four visual styles in rapid succession: pencil sketch, flat 2D illustration, stylized 3D, photorealistic render, each held for 0.4 seconds, final state held as primary look, transitions use dissolve with chromatic fringing(Avatar快速依次切换四种视觉风格:铅笔草图、扁平2D插画、风格化3D、超写实渲染,每种风格持续0.4秒,最终状态作为主视觉,切换时使用带有色差的溶解效果)”
适用场景:创意机构、设计平台、AI工具、内容创作产品
The Fourth Wall Break
Fourth Wall Break(打破第四面墙)
Avatar makes direct, deliberate eye contact with the camera, creating the immediate impression of personal address. Can begin mid-gesture or mid-sentence — in media res. Communicates: directness, confidence, presence.
Prompt language: "avatar turns to face camera directly with deliberate intention, locks eye contact with subtle confident smile, slight lean forward, no introduction or preamble, as if resuming a conversation"
Use when: Sales content, coaching products, high-conversion landing page videos, personal brand content
Avatar直接、刻意地与镜头对视,立即营造出直接对话的感觉。可以从手势中途或句子中途开始——即切入正在进行的场景。传达:直接、自信、存在感。
提示词语言:“avatar turns to face camera directly with deliberate intention, locks eye contact with subtle confident smile, slight lean forward, no introduction or preamble, as if resuming a conversation(Avatar刻意转身直面镜头,带着微妙自信的微笑锁定目光,微微前倾,无介绍或开场白,仿佛继续一场对话)”
适用场景:销售内容、辅导产品、高转化率落地页视频、个人品牌内容
5. Environment Design
5. 环境设计
The environment shapes perception of the avatar's authority, energy, and brand world.
环境塑造了Avatar的权威性、活力和品牌世界的感知。
Void / Abstract Space
Void / Abstract Space(虚空/抽象空间)
The avatar exists in pure black, pure white, or a gradient void with subtle environmental effects (fog, particles, light beams). Maximum focus on the character. Minimum distraction.
Prompt language: "infinite dark void background, subtle volumetric fog at floor level, no horizon line, directional key light from upper left, ambient fill light from right, deep shadows behind avatar"
Best for: Premium product launches, minimalist brands, close-up spokesperson content
Avatar存在于纯黑、纯白或带有微妙环境效果(雾、粒子、光束)的渐变虚空之中。最大程度聚焦角色,最小化干扰。
提示词语言:“infinite dark void background, subtle volumetric fog at floor level, no horizon line, directional key light from upper left, ambient fill light from right, deep shadows behind avatar(无限黑暗虚空背景,地面带有微妙体积雾,无地平线,左上方向主光,右侧环境补光,Avatar身后深阴影)”
最佳适用场景:高端产品发布、极简品牌、特写代言人内容
Virtual Studio
Virtual Studio(虚拟工作室)
A clearly designed, branded environment suggesting a broadcast or content studio — but digital. Can include branded color washes, abstract architecture, floating graphic elements.
Prompt language: "virtual broadcast studio environment with clean geometric architecture, brand color accent lighting on background panels, subtle depth-of-field blur on background, professional lighting rig casting clean key light, polished floor with soft reflection"
Best for: B2B, SaaS, media companies, professional services
经过明确设计的品牌化环境,暗示广播或内容工作室——但为数字环境。可包含品牌色洗墙、抽象建筑、浮动图形元素。
提示词语言:“virtual broadcast studio environment with clean geometric architecture, brand color accent lighting on background panels, subtle depth-of-field blur on background, professional lighting rig casting clean key light, polished floor with soft reflection(具有简洁几何建筑的虚拟广播工作室环境,背景面板带有品牌色重点照明,背景带有微妙景深模糊,专业灯光设备投射清晰主光,抛光地板带有柔和反射)”
最佳适用场景:B2B、SaaS、媒体公司、专业服务
Futuristic Office / Corporate Space
Futuristic Office / Corporate Space(未来办公室/企业空间)
A recognizable workspace — desk, window, shelving — but elevated with futuristic design language. Grounds the avatar in a professional context while signaling technological advancement.
Prompt language: "near-future office environment with clean architectural lines, large floor-to-ceiling window with city view, ambient blue-white environmental light, holographic display elements visible in background, depth-of-field blur keeping avatar in sharp focus"
Best for: Enterprise, finance, leadership content, thought leadership
可识别的工作空间——办公桌、窗户、置物架——但采用未来主义设计语言升级。将Avatar置于专业背景中,同时彰显技术先进性。
提示词语言:“near-future office environment with clean architectural lines, large floor-to-ceiling window with city view, ambient blue-white environmental light, holographic display elements visible in background, depth-of-field blur keeping avatar in sharp focus(具有简洁建筑线条的近未来办公室环境,大型落地窗可俯瞰城市,环境蓝白光,背景可见全息显示元素,景深模糊使Avatar保持清晰焦点)”
最佳适用场景:企业、金融、领导力内容、思想领导力
Neon Cyber Environment
Neon Cyber Environment(霓虹赛博环境)
High-contrast urban digital environments — neon-lit corridors, data centers, glowing grid floors, cyberpunk architecture. High energy, high visual impact.
Prompt language: "neon-lit cyberpunk environment with deep shadows and saturated color pools of magenta and cyan, glowing geometric floor grid extending to horizon, electric particle effects drifting through air, haze atmosphere with volumetric light shafts"
Best for: Gaming, crypto/web3, youth brands, entertainment, tech startups
高对比度的城市数字环境——霓虹走廊、数据中心、发光网格地板、赛博朋克建筑。高能量、高视觉冲击力。
提示词语言:“neon-lit cyberpunk environment with deep shadows and saturated color pools of magenta and cyan, glowing geometric floor grid extending to horizon, electric particle effects drifting through air, haze atmosphere with volumetric light shafts(霓虹照亮的赛博朋克环境,带有深阴影和品红、青色饱和色块,发光几何地板网格延伸至地平线,空气中漂浮着电粒子效果,雾霭氛围中带有体积光束)”
最佳适用场景:游戏、加密货币/Web3、青年品牌、娱乐、科技初创企业
6. Camera Techniques for Avatar Content
6. Avatar内容的运镜技巧
Camera work communicates status, intimacy, and energy for digital characters as much as for human subjects.
Framing fundamentals
- Close-up (face-fill frame): maximum trust and intimacy — best for direct address and emotional moments
- Medium (chest-up): most versatile — shows gesture, expression, and upper body simultaneously
- Wide (full body): establishes character scale and presence — best for action or environment reveals
- Rule of thirds: place avatar eyes on upper third horizontal line — avoid dead center unless intentional
Camera movement vocabulary
Slow push-in: builds anticipation and intimacy. "Camera slowly pushes in from medium to close-up over 4 seconds, no shake, smooth dolly motion"
Ken Burns on static frame: creates life from a still render. "Subtle slow zoom in with slight rightward drift, mimicking handheld documentary intimacy"
Orbit / arc move: reveals the avatar's three-dimensionality. "Camera arcs 30 degrees around avatar from left to right at shoulder height, smooth constant speed, avatar rotates slightly to maintain eye contact"
Reveal push: avatar is revealed as camera moves. "Camera starts on environment detail, pushes forward to reveal avatar in foreground, avatar turns to address camera at full reveal"
Match cut / style cut: avatar holds a pose that cuts perfectly to a different style or angle. Requires specifying the exact pose to hold at cut point.
Depth of field guidance
- Photorealistic avatars: always specify shallow depth of field — it reads as cinematic and elevates perceived quality
- Stylized avatars: can run sharper focus — their visual world tends toward graphic clarity
- Prompt language: "f/1.8 equivalent shallow depth of field, background in soft bokeh blur, avatar in sharp focus from eyes to chin"
Stabilization language
- For authority and trust: "locked-off tripod shot, zero camera shake"
- For energy and immediacy: "subtle handheld movement, barely perceptible breathing motion on camera"
- For cinematic premium: "smooth gimbal-stabilized movement, fluid and controlled"
运镜对数字角色的地位、亲密感和活力的传达,与对人类主体的作用一样重要。
构图基础
- 特写(面部填满画面):最大程度的信任和亲密感——最适合直接对话和情感时刻
- 中景(胸部以上):最通用——同时展示手势、表情和上半身
- 全景(全身):确立角色比例和存在感——最适合动作或环境展示
- 三分法:将Avatar的眼睛置于上三分线——除非刻意,否则避免居中
运镜词汇
缓慢推镜:营造期待感和亲密感。“Camera slowly pushes in from medium to close-up over 4 seconds, no shake, smooth dolly motion(镜头在4秒内从中景缓慢推至特写,无抖动,平滑轨道运动)”
静态画面的Ken Burns效果:为静态渲染赋予生命力。“Subtle slow zoom in with slight rightward drift, mimicking handheld documentary intimacy(微妙缓慢放大并轻微向右偏移,模仿手持纪录片的亲密感)”
环绕/弧形移动:展示Avatar的三维立体感。“Camera arcs 30 degrees around avatar from left to right at shoulder height, smooth constant speed, avatar rotates slightly to maintain eye contact(镜头在肩部高度从左向右环绕Avatar旋转30度,匀速平滑,Avatar轻微旋转以保持目光接触)”
推镜展示:镜头移动时露出Avatar。“Camera starts on environment detail, pushes forward to reveal avatar in foreground, avatar turns to address camera at full reveal(镜头从环境细节开始,向前推镜以展示前景中的Avatar,完全露出时Avatar转身面向镜头)”
匹配剪辑/风格剪辑:Avatar保持一个姿势,完美切入不同风格或角度。需要指定剪辑点的精确姿势。
景深指导
- 超写实Avatar:始终指定浅景深——具有电影感,提升感知质量
- 风格化Avatar:可使用更清晰的焦点——其视觉世界倾向于图形清晰度
- 提示词语言:“f/1.8 equivalent shallow depth of field, background in soft bokeh blur, avatar in sharp focus from eyes to chin(等效f/1.8浅景深,背景柔和散景模糊,Avatar从眼睛到下巴保持清晰焦点)”
稳定语言
- 体现权威和信任:“locked-off tripod shot, zero camera shake(固定三脚架拍摄,无镜头抖动)”
- 体现活力和即时性:“subtle handheld movement, barely perceptible breathing motion on camera(微妙手持移动,镜头带有几乎难以察觉的呼吸式运动)”
- 体现电影级高端感:“smooth gimbal-stabilized movement, fluid and controlled(平滑云台稳定运动,流畅可控)”
7. Lighting for Avatar Content
7. Avatar内容的灯光设计
Holographic Glow
Holographic Glow(全息发光)
The avatar emits or is surrounded by luminous energy — neon, bioluminescent, data-light. Strongly signals AI/digital nature.
Prompt language: "avatar lit from within by soft blue-white holographic glow, light emanating from chest area and face, ambient environmental light fills from all sides evenly, subtle rim light of electric cyan outlines the silhouette, no hard shadows"
Avatar自身发光或被发光能量包围——霓虹、生物发光、数据光。强烈暗示AI/数字属性。
提示词语言:“avatar lit from within by soft blue-white holographic glow, light emanating from chest area and face, ambient environmental light fills from all sides evenly, subtle rim light of electric cyan outlines the silhouette, no hard shadows(Avatar由内部柔和蓝白全息光照明,光线从胸部和面部散发,环境光均匀填充四周,电青色微妙轮廓光勾勒剪影,无硬阴影)”
Studio Clean
Studio Clean(专业工作室灯光)
Professional three-point lighting setup. Key light, fill light, rim light. The standard for commercial spokesperson content.
Prompt language: "professional three-point lighting: strong key light from upper left at 45 degrees, soft fill light from right reducing shadow contrast, tight rim light from behind right creating clean separation from background, color temperature 5600K daylight-balanced"
专业三点布光设置:主光、补光、轮廓光。商业代言人内容的标准配置。
提示词语言:“professional three-point lighting: strong key light from upper left at 45 degrees, soft fill light from right reducing shadow contrast, tight rim light from behind right creating clean separation from background, color temperature 5600K daylight-balanced(专业三点布光:左上45度强主光,右侧软补光降低阴影对比度,右后方窄轮廓光使Avatar与背景清晰分离,色温5600K日光平衡)”
Neon Rim
Neon Rim(霓虹轮廓光)
Hard, colored rim light from behind creates dramatic silhouette and strong graphic identity. Background colors bleed onto subject.
Prompt language: "neon rim lighting with magenta light from behind left and cyan from behind right, strong color separation, dark front face with colored halo around silhouette, moody underexposure on face revealing only key features"
后方强烈彩色轮廓光营造戏剧性剪影和鲜明图形标识。背景颜色渗透到主体上。
提示词语言:“neon rim lighting with magenta light from behind left and cyan from behind right, strong color separation, dark front face with colored halo around silhouette, moody underexposure on face revealing only key features(霓虹轮廓光,左后方品红光,右后方青光,强烈色彩分离,面部前方较暗,剪影周围有色环,面部低调光仅显示关键特征)”
Volumetric Atmosphere
Volumetric Atmosphere(体积氛围光)
Light behaves as physical substance — fog, haze, dust particles catching beams. Creates depth, atmosphere, and cinematic quality.
Prompt language: "volumetric light shafts cutting through atmospheric haze at 45-degree angle, dust and particle matter visible in light beams, soft falloff from bright to dark, environmental light wrapping avatar with soft fill from bounced volumetric scatter"
Lighting rules for photorealistic avatars specifically:
- Always include a catch light in the eyes — a small specular highlight
- Specify skin tone interaction with colored light: "warm amber light warming skin to golden tones"
- Avoid flat frontal lighting — it eliminates the three-dimensionality that makes a render feel real
- Include shadow direction consistency between avatar and environment
光线表现为物理物质——雾、霾、尘埃粒子捕捉光束。营造深度、氛围和电影质感。
提示词语言:“volumetric light shafts cutting through atmospheric haze at 45-degree angle, dust and particle matter visible in light beams, soft falloff from bright to dark, environmental light wrapping avatar with soft fill from bounced volumetric scatter(体积光束以45度角穿透大气霾,光束中可见尘埃和粒子,从亮到暗柔和衰减,环境光通过体积散射反弹为Avatar提供软补光)”
超写实Avatar专属灯光规则:
- 眼睛中始终包含眼神光——小镜面高光
- 指定肤色与彩色光的互动:“warm amber light warming skin to golden tones(暖琥珀光将皮肤暖化为金色调)”
- 避免平面前方照明——会消除使渲染真实的立体感
- 确保Avatar与环境的阴影方向一致
8. Sound Design for Avatar Content
8. Avatar内容的音效设计
Sound design for AI avatar videos operates in three layers.
AI Avatar视频的音效设计分为三个层次。
Synthetic Voice Integration
Synthetic Voice Integration(合成语音整合)
The avatar's voice is its most direct communication channel. Even in a video prompt (which describes the visual), the voice direction informs character and performance.
Voice direction markers to include in prompts or production notes:
- "Voice has measured, authoritative cadence with subtle warmth"
- "Speech pace is deliberate — one beat pause before key statements"
- "Slight synthetic processing on voice: clean reverb tail, light high-frequency shimmer"
- "Voice texture is clear and close-mic'd, no room ambience"
- "AI-native voice quality: intelligible, precise articulation, slight harmonic brightness"
Avatar的声音是最直接的沟通渠道。即使在视频提示词(描述视觉内容)中,语音指导也会影响角色和表现。
提示词或制作说明中需包含的语音指导标记:
- “Voice has measured, authoritative cadence with subtle warmth(语音节奏沉稳、权威,带有微妙温暖感)”
- “Speech pace is deliberate — one beat pause before key statements(语速从容——关键陈述前停顿一拍)”
- “Slight synthetic processing on voice: clean reverb tail, light high-frequency shimmer(语音带有轻微合成处理:清晰混响尾音,轻微高频光泽)”
- “Voice texture is clear and close-mic'd, no room ambience(语音质感清晰,近距离收音,无房间环境音)”
- “AI-native voice quality: intelligible, precise articulation, slight harmonic brightness(原生AI语音质量:清晰易懂,发音精准,略带谐波明亮感)”
Digital Ambient
Digital Ambient(数字环境音)
Environmental sound that signals the digital world the avatar inhabits.
Textures to specify:
- "Sub-bass hum of active server infrastructure, very low level"
- "High-frequency tonal shimmer suggesting energy fields"
- "Subtle digital breathing sound: periodic soft clicks and chirps"
- "Near-silence with occasional data-ping transient"
- "Low-level static white noise suggesting signal presence"
暗示Avatar所处数字世界的环境音。
需指定的音效质感:
- “Sub-bass hum of active server infrastructure, very low level(活跃服务器基础设施的次低音嗡鸣,音量极低)”
- “High-frequency tonal shimmer suggesting energy fields(暗示能量场的高频 tonal shimmer)”
- “Subtle digital breathing sound: periodic soft clicks and chirps(微妙数字呼吸声:周期性轻咔哒声和啁啾声)”
- “Near-silence with occasional data-ping transient(近乎静音,偶尔出现数据 ping 瞬态音)”
- “Low-level static white noise suggesting signal presence(低电平静态白噪音,暗示信号存在)”
Futuristic Music
Futuristic Music(未来感音乐)
Background music that contextualizes the avatar's world and emotional register.
Music direction markers:
- "Minimal electronic score with synthetic textures, no live instruments"
- "Tempo-matched to avatar's speaking rhythm, approximately 90 BPM"
- "Cinematic trailer energy: percussive hits synchronized to visual cuts"
- "Ambient generative music: evolving pads, no melodic line, pure atmosphere"
- "Clean tech-house pulse at low level, provides forward momentum without distraction"
Sound design principle for AI avatars: sound should feel designed, not organic. Real-world ambience (birds, traffic, HVAC) undermines the digital world effect. Every sonic element should feel intentional and synthetic.
为Avatar的世界和情感基调提供背景的音乐。
音乐指导标记:
- “Minimal electronic score with synthetic textures, no live instruments(极简电子配乐,合成质感,无真实乐器)”
- “Tempo-matched to avatar's speaking rhythm, approximately 90 BPM(节奏匹配Avatar的说话节奏,约90 BPM)”
- “Cinematic trailer energy: percussive hits synchronized to visual cuts(电影预告片能量感:打击乐与视觉剪辑同步)”
- “Ambient generative music: evolving pads, no melodic line, pure atmosphere(氛围生成音乐:渐变铺垫,无旋律线,纯粹氛围)”
- “Clean tech-house pulse at low level, provides forward momentum without distraction(低音量清晰科技浩室节拍,提供前进动力且不分散注意力)”
AI Avatar音效设计原则:音效应经过设计,而非自然真实。现实世界环境音(鸟鸣、交通、空调)会破坏数字世界效果。每个音效元素都应具有目的性和合成感。
Platform Optimization
平台优化
- TikTok (9:16): High energy, fast cuts acceptable. Stylized 3D or 2D avatars perform best. Keep under 15s.
- Instagram Reels (9:16): Higher visual polish expected. Photorealistic avatars work well here. Cover frame must show avatar clearly.
- YouTube (16:9): Longer format, more room for establishing shots. Enterprise/educational use cases.
- LinkedIn (16:9 or 1:1): Professional, corporate-appropriate. Photorealistic only. Subtle animation, minimal visual effects.
- TikTok(9:16): 高能量,允许快速剪辑。Stylized 3D或2D Avatar表现最佳。时长控制在15秒以内。
- Instagram Reels(9:16): 对视觉质感要求更高。Photorealistic Avatar在此平台表现良好。封面帧必须清晰展示Avatar。
- YouTube(16:9): 长格式,有更多空间用于开场镜头。企业/教育场景适用。
- LinkedIn(16:9或1:1): 专业、符合企业调性。仅使用Photorealistic风格。微动画,尽量减少视觉特效。
9. Complete Example Prompts
9. 完整示例提示词
Example 1 — SaaS Product Spokesperson (Photorealistic)
示例1——SaaS产品代言人(Photorealistic)
Photorealistic AI avatar of a professional woman in her early 30s, South Asian appearance,
shoulder-length dark hair, wearing a sharp minimal blazer in charcoal, positioned center-frame
in medium shot from chest up.
Avatar is placed in a virtual studio environment: clean architectural background with frosted
glass panels catching diffuse blue-white ambient light, polished light floor with subtle
reflection of subject, background held in soft bokeh blur.
Lighting: professional three-point setup — strong key light from upper left at 45 degrees,
warm fill from the right reducing shadow contrast, thin rim light from behind creating clean
silhouette separation from background. Color temperature 5600K.
Camera: static tripod shot at eye level, no movement, locked and authoritative. Shallow
depth of field with background in smooth bokeh.
Avatar's expression is open and confident — slight genuine smile, direct eye contact with
camera, a catch light visible in both eyes. Subtle asymmetry in the smile. A single strand
of hair rests slightly displaced to the right, adding naturalism.
At frame open she has just finished a thought and takes a brief natural breath before
beginning to speak. Head has a slight, natural tilt of attentiveness. No glitch effect —
clean professional entry.
Skin rendered with visible pore texture, subsurface light scattering at cheek and brow
areas, natural lip moisture highlight, eye corneal wetness with point source reflection.
Micromotion: almost imperceptible chest movement from breathing, barely visible blink at
second 2, slight weight shift that creates minimal shoulder movement. Everything reads as
alive, not frozen.
Color grade: clean and cool, slight desaturation, professional broadcast color treatment.Photorealistic AI avatar of a professional woman in her early 30s, South Asian appearance,
shoulder-length dark hair, wearing a sharp minimal blazer in charcoal, positioned center-frame
in medium shot from chest up.
Avatar is placed in a virtual studio environment: clean architectural background with frosted
glass panels catching diffuse blue-white ambient light, polished light floor with subtle
reflection of subject, background held in soft bokeh blur.
Lighting: professional three-point setup — strong key light from upper left at 45 degrees,
warm fill from the right reducing shadow contrast, thin rim light from behind creating clean
silhouette separation from background. Color temperature 5600K.
Camera: static tripod shot at eye level, no movement, locked and authoritative. Shallow
depth of field with background in smooth bokeh.
Avatar's expression is open and confident — slight genuine smile, direct eye contact with
camera, a catch light visible in both eyes. Subtle asymmetry in the smile. A single strand
of hair rests slightly displaced to the right, adding naturalism.
At frame open she has just finished a thought and takes a brief natural breath before
beginning to speak. Head has a slight, natural tilt of attentiveness. No glitch effect —
clean professional entry.
Skin rendered with visible pore texture, subsurface light scattering at cheek and brow
areas, natural lip moisture highlight, eye corneal wetness with point source reflection.
Micromotion: almost imperceptible chest movement from breathing, barely visible blink at
second 2, slight weight shift that creates minimal shoulder movement. Everything reads as
alive, not frozen.
Color grade: clean and cool, slight desaturation, professional broadcast color treatment.Example 2 — Virtual Influencer Hook (Stylized 3D, Digital Birth Hook)
示例2——虚拟网红钩子(Stylized 3D,Digital Birth Hook)
Stylized 3D avatar character: gender-neutral presentation, large luminous eyes with
teal irises, smooth simplified skin geometry with clean subsurface warmth, short silver
hair rendered in individual strands, wearing a graphic oversized jacket in deep purple
with gold accent trim.
Opens with The Digital Birth hook: scattered luminous gold and teal particles
in void space, drifting inward as if being called together, density increasing,
wireframe mesh of the character's face emerging at center, particles filling in
surface texture, color and detail resolving over 1.8 seconds into the fully rendered
stylized avatar.
At full resolution the avatar opens its eyes — both irises catch light simultaneously —
and looks directly into the camera with a slow, deliberate smile. This moment is held
for 0.5 seconds in silence before any other motion.
Environment: pure void background, deep black with barely visible particle drift,
no horizon, rim lighting of teal from behind left and gold from behind right creating
dual-color silhouette framing that matches character palette.
Camera: medium close-up, centered, no movement during formation sequence, then very
slow push-in beginning at full resolution moment, arriving at close-up by second 5.
Lighting: rim glow from environment colors, soft fill from below and in front,
no hard shadows on face — clean, graphic, beautiful.
Music/sound: during particle formation, high-frequency shimmer building in intensity,
resolving to a single clean synth tone at full avatar reveal. Then minimal electronic
ambient pulse begins beneath subsequent content.
Cel-adjacent render with clean edge highlights, smooth gradient skin tones, bold
stylistic intentionality throughout. Never attempts photorealism.Stylized 3D avatar character: gender-neutral presentation, large luminous eyes with
teal irises, smooth simplified skin geometry with clean subsurface warmth, short silver
hair rendered in individual strands, wearing a graphic oversized jacket in deep purple
with gold accent trim.
Opens with The Digital Birth hook: scattered luminous gold and teal particles
in void space, drifting inward as if being called together, density increasing,
wireframe mesh of the character's face emerging at center, particles filling in
surface texture, color and detail resolving over 1.8 seconds into the fully rendered
stylized avatar.
At full resolution the avatar opens its eyes — both irises catch light simultaneously —
and looks directly into the camera with a slow, deliberate smile. This moment is held
for 0.5 seconds in silence before any other motion.
Environment: pure void background, deep black with barely visible particle drift,
no horizon, rim lighting of teal from behind left and gold from behind right creating
dual-color silhouette framing that matches character palette.
Camera: medium close-up, centered, no movement during formation sequence, then very
slow push-in beginning at full resolution moment, arriving at close-up by second 5.
Lighting: rim glow from environment colors, soft fill from below and in front,
no hard shadows on face — clean, graphic, beautiful.
Music/sound: during particle formation, high-frequency shimmer building in intensity,
resolving to a single clean synth tone at full avatar reveal. Then minimal electronic
ambient pulse begins beneath subsequent content.
Cel-adjacent render with clean edge highlights, smooth gradient skin tones, bold
stylistic intentionality throughout. Never attempts photorealism.Example 3 — AI Presenter, Enterprise Tech Launch (Futuristic, Fourth Wall Break Hook)
示例3——AI主持人,企业科技发布(未来风,Fourth Wall Break Hook)
Photorealistic digital human avatar, male, mid-40s, Northwestern European appearance,
close-cropped dark hair with salt-and-pepper at temples, wearing a clean white shirt
under a structured dark navy jacket, no tie. Commands immediate authority.
Opens in media res with The Fourth Wall Break: avatar is already mid-turn toward camera,
as if he became aware of our presence. He completes the turn in the first 0.8 seconds,
makes direct, unhurried eye contact, and leans forward almost imperceptibly — the body
language of a person with something important to say, not a presentation to deliver.
Environment: near-future corporate interior, large window occupying the entire background,
out-of-focus cityscape in deep-focus bokeh, the room's architecture is clean and minimal —
single floating desk edge visible at frame right, ambient indirect ceiling light from above,
blue-hour exterior light bleeding through window glass and catching the side of his face.
Lighting: ambient daylight from window provides cool atmospheric fill, key light from
left at 45 degrees slightly warmer to balance the cool ambient, crisp but narrow rim
light from the window side. High dynamic range — bright window, controlled shadows on
face, catch light clearly visible in both eyes.
Camera: low eye-level camera — slightly below eyeline, reinforcing the subject's authority.
After initial static hold (0–3 seconds), a nearly imperceptible slow push-in begins,
arriving 8% closer by second 10. This movement is felt subconsciously, not noticed.
Skin: maximum photorealistic fidelity — visible pore texture, subsurface warmth at
cheeks and forehead, five o'clock shadow rendered at individual hair level, natural
lip texture, oil/moisture on brow area.
Micromotion: regular natural blinks every 3–4 seconds, subtle jaw movement as if
processing a thought before speaking, minimal but present chest breath movement,
a very slight asymmetric quality to his resting expression — the face of someone
thinking, not posing.
Sound: room ambience is complete silence except for a very low server-hum that suggests
he exists at the intersection of physical and digital. No music until he begins speaking.
Color grade: slightly desaturated with lifted blacks, clean digital cinema aesthetic,
neutral whites, no stylization that would undermine the photorealistic character intent.Photorealistic digital human avatar, male, mid-40s, Northwestern European appearance,
close-cropped dark hair with salt-and-pepper at temples, wearing a clean white shirt
under a structured dark navy jacket, no tie. Commands immediate authority.
Opens in media res with The Fourth Wall Break: avatar is already mid-turn toward camera,
as if he became aware of our presence. He completes the turn in the first 0.8 seconds,
makes direct, unhurried eye contact, and leans forward almost imperceptibly — the body
language of a person with something important to say, not a presentation to deliver.
Environment: near-future corporate interior, large window occupying the entire background,
out-of-focus cityscape in deep-focus bokeh, the room's architecture is clean and minimal —
single floating desk edge visible at frame right, ambient indirect ceiling light from above,
blue-hour exterior light bleeding through window glass and catching the side of his face.
Lighting: ambient daylight from window provides cool atmospheric fill, key light from
left at 45 degrees slightly warmer to balance the cool ambient, crisp but narrow rim
light from the window side. High dynamic range — bright window, controlled shadows on
face, catch light clearly visible in both eyes.
Camera: low eye-level camera — slightly below eyeline, reinforcing the subject's authority.
After initial static hold (0–3 seconds), a nearly imperceptible slow push-in begins,
arriving 8% closer by second 10. This movement is felt subconsciously, not noticed.
Skin: maximum photorealistic fidelity — visible pore texture, subsurface warmth at
cheeks and forehead, five o'clock shadow rendered at individual hair level, natural
lip texture, oil/moisture on brow area.
Micromotion: regular natural blinks every 3–4 seconds, subtle jaw movement as if
processing a thought before speaking, minimal but present chest breath movement,
a very slight asymmetric quality to his resting expression — the face of someone
thinking, not posing.
Sound: room ambience is complete silence except for a very low server-hum that suggests
he exists at the intersection of physical and digital. No music until he begins speaking.
Color grade: slightly desaturated with lifted blacks, clean digital cinema aesthetic,
neutral whites, no stylization that would undermine the photorealistic character intent.10. Prompt Rules for Avatar Content
10. Avatar内容提示词规则
Rule 1: Commit to a style lane and stay in it.
Mixed signals between photorealism and stylization produce uncanny valley. Every detail in the prompt must serve the same aesthetic target. If you write "photorealistic skin" and "cartoon eyes" in the same prompt, the output will be incoherent.
Rule 2: Describe aliveness, not perfection.
Perfect symmetry, frozen expressions, and stillness read as dead or rendered. Include asymmetry, micro-motion, imperfection, and breath. "A strand of hair slightly displaced" does more for believability than five sentences about render quality.
Rule 3: Eyes anchor everything.
The eyes are the first thing viewers evaluate for believability and connection. Always specify: catch light presence, iris color and texture detail, corneal wetness, natural blink timing, and where the gaze is directed. Underspecified eyes produce the single most common uncanny valley failure.
Rule 4: Ground the avatar to its environment.
The avatar should interact with its environment through light and shadow — light from a window should fall on the skin, a neon environment should tint the avatar's face. A character that floats in front of its background without environmental interaction reads as a composite, not a presence.
Rule 5: Camera language is character language.
How the camera frames and moves around an avatar communicates personality and status before the character does anything. A locked-off shot says authority. A slow push-in says intimacy. A low angle says power. A slight handheld wobble says authenticity. Choose intentionally.
Rule 6: Use sound as world-building, not decoration.
Synthetic ambience and designed sound textures establish that this is a digital world, which paradoxically makes the avatar more believable — you are setting context for synthetic existence rather than trying to pass it off as real.
Rule 7: Specify the hook in the first sentence of the prompt.
Seedance and most AI video models weight the opening of the prompt heavily for the opening of the video. Put the hook description first — what happens in seconds 0–2 — before any other detail.
Rule 8: Duration shapes detail density.
For 2–4 second clips: focus entirely on one moment, one expression, one action. Every detail must earn its place.
For 6–10 second clips: allow for a single transition or beat — hook, then hold.
For 15–30 second clips: you can describe an arc — opening state, middle development, closing beat.
Rule 9: Avoid "AI-looking" descriptions.
Paradoxically, describing an avatar as "AI-generated" or "digital" in the prompt often produces visual clichés — blue grids, floating data, lens flares on eyes. If the goal is photorealism, describe a real human. If the goal is stylized, describe a designed character. Let the aesthetic communicate the AI identity, not the text.
Rule 10: Test photorealism with the close-up rule.
If you are writing a photorealistic avatar prompt, ask: does this description hold up in extreme close-up? If the answer is no — if the pores, eye detail, hair texture, and micro-expression are underspecified — push harder on those elements before finalizing the prompt.
规则1:坚持一种风格并贯彻始终。
超写实与风格化之间的混合信号会引发恐怖谷。提示词中的每个细节都必须服务于相同的美学目标。如果在同一个提示词中写“超写实皮肤”和“卡通眼睛”,输出结果会不协调。
规则2:描述生命力,而非完美。
完美对称、冻结的表情和静止状态会显得僵硬或无生命。加入不对称、微动作、瑕疵和呼吸动作。“一缕凌乱的头发”比五句关于渲染质量的描述更能提升真实感。
规则3:眼睛是核心。
眼睛是观众评估真实感和连接感的第一要素。始终指定:眼神光存在、虹膜颜色和纹理细节、角膜湿润度、自然眨眼时机,以及注视方向。眼睛描述不明确是恐怖谷失败最常见的原因。
规则4:让Avatar与环境融为一体。
Avatar应通过光线和阴影与环境互动——窗户的光线应落在皮肤上,霓虹环境应使Avatar的面部带有色彩。一个悬浮在背景前而无环境互动的角色会显得是合成的,而非真实存在。
规则5:运镜语言即角色语言。
镜头如何构图和围绕Avatar移动,在角色有所动作之前就传达了其个性和地位。固定镜头代表权威。缓慢推镜代表亲密。低角度代表力量。轻微手持晃动代表真实感。要有目的地选择。
规则6:将音效用作世界构建,而非装饰。
合成环境音和经过设计的音效纹理确立了这是一个数字世界,反而使Avatar更可信——你是在为合成存在设定背景,而非试图伪装成真实。
规则7:在提示词的第一句指定钩子。
Seedance和大多数AI视频模型会重点关注提示词的开头部分,以生成视频的开场。将钩子描述放在最前面——即0-2秒发生的事情——然后再添加其他细节。
规则8:时长决定细节密度。
对于2-4秒的片段:完全聚焦于一个瞬间、一种表情、一个动作。每个细节都必须有存在的意义。
对于6-10秒的片段:允许单个过渡或节拍——钩子,然后保持。
对于15-30秒的片段:可以描述一个完整的弧线——开场状态、中间发展、结尾节拍。
规则9:避免“AI感”描述。
矛盾的是,在提示词中将Avatar描述为“AI生成”或“数字”往往会产生视觉陈词滥调——蓝色网格、浮动数据、眼睛上的镜头光晕。如果目标是超写实,描述一个真实人类。如果目标是风格化,描述一个经过设计的角色。让美学传达AI属性,而非文字。
规则10:用特写规则测试超写实度。
如果你正在撰写超写实Avatar提示词,请自问:这个描述在极端特写镜头下是否成立?如果答案是否定的——如果毛孔、眼睛细节、头发纹理和微表情描述不足——在最终确定提示词前,进一步细化这些元素。