# How We Cracked Seedance 2.0 Prompting and NEVER Have to Regenerate AI UGC Videos...
**作者**: Adrian Solarz
**日期**: 2026-05-07T21:30:55.000Z
**来源**: [https://x.com/adriansolarzz/status/2052501455169568875](https://x.com/adriansolarzz/status/2052501455169568875)
---

many operators using Seedance 2.0 are burning through credits regenerating the same clip 3 to 4 times before they get something usable.
but we figured out the prompting structure that gets us usable output on the first or second generation consistently.
so in this article, i'll be breaking down exactly how the Seedance 2.0 prompting structure works in production:
- the 5-beat formula
- the dialogue embedding rule
- the brand name pronunciation defense, the selfie POV framing
- and the real-brand prop strategy that's the biggest AI-detection defense.
if you've been running Seedance 2.0 with prompts that produce inconsistent output, this is what to do:
## The 15-Second Hard Limit Most Operators Misunderstand

before getting into the prompting structure, there's a hard technical limit that every operator needs to internalize because most people work against it without realizing.
Seedance 2.0 caps every single generation at 15 seconds.
there's no workaround for this at the model level. you cannot prompt your way into a 30-second clip. every generation produces a 15-second output maximum and that's the ceiling.
operators trying to generate longer clips by prompting for "30-second clip showing X" produce 15-second outputs that try to compress 30 seconds of action into the available window, which produces visual chaos and wastes the generation entirely. the prompt fights the model's hard limit and loses every time.
How to Build Longer Content
longer content gets built by stitching multiple 15-second clips together in post.
generate each shot as its own clip up to 15 seconds. use a frame from the first clip as a reference image for the second clip (this locks character and location consistency across the stitched sequence). stitch the clips together in CapCut, Premiere, or Final Cut. match the color grade and grain across all clips so the assembled output reads as a single coherent piece.
a 30-second AI UGC reel is 2 stitched 15-second clips. a 60-second longform piece is 4 stitched clips. the workflow scales linearly with the desired duration, but every individual generation stays under the 15-second ceiling.
How to Use the Full 15 Seconds Effectively
within a single 15-second generation, the prompting structure also matters for how the duration gets used.
front-load all character and environment details early in the prompt because the model weights earlier tokens more heavily when constructing the visual baseline of the generation.
use 1 continuous camera move to carry the story arc (push-in, arc, pull-back) rather than describing "cut to" transitions. the model gets confused by transition descriptions and produces output that doesn't match either side of the cut cleanly.
change lighting in-world rather than changing locations. a 15-second generation that depicts dawn light bleeding through blinds gradually brightening the scene works cleanly. a 15-second generation that tries to move from a bedroom to a kitchen mid-clip produces visual chaos.
once you internalize the 15-second cap, your prompting structure changes. you stop trying to fit complex multi-beat narratives into single generations and start designing each generation as 1 specific scene that unfolds within the 15-second window, with longer narratives built through stitched sequences.
## The 5-Beat Prompt Formula Every Generation Follows

every Seedance 2.0 prompt that consistently produces usable output on the first generation follows the same 5-beat formula.
Beat 1: Subject With Specific Visual Details
every prompt opens with the subject and specific visual details: age, hair color and style, clothing, accessories, ethnicity.
vague subject descriptions produce variable subjects across regenerations. specific subject descriptions produce consistent subjects that match what the operator was actually trying to create.
"a woman" produces wildly different outputs across generations. "a 28-year-old woman with shoulder-length brown hair, wearing an oversized white t-shirt and gold hoop earrings" produces output that consistently lands in the visual ballpark of what was specified.
Beat 2: Setting and Environment
one concrete, evocative location rather than abstract style references.
abstract style references produce wildly inconsistent renderings because the model has to interpret what the abstract reference means visually. concrete location descriptions give the model a specific visual target.
"a luxury aesthetic" is abstract. "a sunlit kitchen with marble countertops, a wooden cutting board with cherry tomatoes, and morning light streaming through the window above the sink" is concrete.
the concrete description produces the luxury aesthetic the abstract description was trying to evoke, because specific visual elements compose into the aesthetic rather than the model having to guess what the aesthetic means.
Beat 3: Action Within the Clip Duration
1 to 2 clear primary motions plus a secondary beat or two, calibrated to the clip duration.
operators stuffing prompts with 6 to 8 different actions produce visual chaos because the model can't fit all the actions into the duration cleanly even with the 15-second ceiling.
the formula is 1 primary action that anchors the clip and 1 to 2 secondary beats that add dimension. for a longer clip approaching the 15-second cap, you can extend to 2 primary actions with smooth transitions between them. "she takes a sip of coffee while looking out the window, then turns to camera and starts talking" is 2 primary actions with a clean transition between them, which reads as natural rather than rushed.
for shorter clips (5 to 8 seconds), 1 primary action plus 1 secondary beat is the right density. the prompt structure stays the same regardless of clip duration. only the action density inside the prompt scales with duration.
Beat 4: Camera Direction
explicit camera moves using named cinematography terms, not vague adjectives.
models trained on film data respond to film vocabulary. "tracking shot," "dolly push-in," "orbit," "arc around," "crane up," "whip pan," "handheld with natural shake," "locked-off," "first-person POV" all produce predictable, controllable camera behavior.
vague terms like "interesting camera movement" produce unpredictable output because the model has to guess what "interesting" means in this context.
Beat 5: Lighting Plus Style Plus Mood Tags
stacked at the end of the prompt rather than scattered throughout.
lighting carries more mood information than almost any other variable. giving lighting its own dedicated section at the end of the prompt produces stronger output than burying it in a list of other scene descriptors.
named lighting terms work best: "golden hour," "blue hour," "magic hour," "rim light," "back light," "volumetric haze," "warm tungsten," "cool fluorescent," "neon accents," "soft window light from camera left."
## Dialogue Must Be Embedded Inline

this is the most common failure mode for new Seedance 2.0 operators.
they write the visual scene description, then write a separate "voiceover script" section below. the model treats these as 2 separate inputs and produces output where the dialogue and visual don't sync correctly.
dialogue must be embedded inline using the "says [exact line]" attribution pattern.
the wrong approach:
"a 28-year-old woman in a sunlit kitchen takes a sip of coffee.
Voiceover script: 'I genuinely cannot believe how this product changed my routine.'"
the right approach:
"a 28-year-old woman in a sunlit kitchen takes a sip of coffee, sets the cup down, looks directly at the camera with excited wide eyes, and says, 'I genuinely cannot believe how this product changed my routine.'"
The Rules for Dialogue Embedding
use "says" attribution every time. she says, he asks, the voice finishes. tie each spoken line to its matching visual moment in the same sentence.
put brand names on visible-speaker beats where the camera sees the speaker's mouth. lip-sync forces cleaner pronunciation than voiceover-over-cutaway beats produce.
avoid voiceover-over-cutaway beats for brand names. the model mumbles brand names when the speaker isn't visible because there's no visual lip-sync constraint forcing clean pronunciation.
end punchy. one clean CTA line, not 3 redundant ones. cramming multiple CTAs into the clip duration produces audio that runs over the time limit and gets cut off, regardless of whether you're generating at 8 seconds or 15 seconds.
## The Brand Name Pronunciation Defense
this is the single biggest pronunciation problem in Seedance 2.0 production, and it has a structured solution.
TTS in Seedance 2.0 mispronounces brand names constantly. UGC becomes "UGG" (like the boots). new brand names get slurred into surrounding dialogue. short acronyms get parsed as words before letters. multi-syllable brand names get compressed into single mumbled words.
the fix is the 3-layer pronunciation defense applied to every brand name mention.
Layer 1: Phonetic Spelling
spell spoken acronyms phonetically: UGC becomes "You Gee See" or "U G Cee." UGCMaxxing becomes "U G Cee Maxxing" or "You Gee See Maxxing." SaaS becomes "Sass" (as a word) or "S. A. A. S." (as letters). AI becomes "A. I." with dots. CEO becomes "C. E. O."
multi-syllable brand names get split: Cantina becomes "Can Teena" (two words, not one). Klaviyo becomes "Klay Vee O." each syllable that the model tends to mumble gets explicit phonetic guidance.
Layer 2: Enunciation Direction
add enunciation direction before the spoken line:
"she slows down slightly to clearly enunciate the brand name and says..."
"the voice is clear and loud over the shot, each word distinct..."
"she says while clearly emphasizing each syllable..."
these directions tell the model to slow the audio specifically at the brand name position, which produces cleaner pronunciation than letting the audio run at default pace.
Layer 3: Repetition With Dash
say the brand name twice in the middle of the clip with a dash separating the repetitions:
"It's called Can Teena, Can Teena, and I'm literally obsessed."
this serves 3 functions simultaneously: it sounds natural in dialogue (real people repeat brand names when introducing them), it gives the TTS 2 chances to nail the pronunciation, and it teaches the viewer the correct pronunciation in case the model still mumbles slightly.
The Brand Name Placement Rule
brand names go on visible-speaker beats only.
the camera should see the speaker's mouth when the brand name is being said. lip-sync forces the model to produce cleaner audio because the visual constraint requires the lips to match the words.
voiceover-over-cutaway beats (where the speaker isn't visible) produce mumbled brand names because there's no visual constraint forcing clean pronunciation. never put a brand name in a voiceover beat. structure the script so brand names always land when the speaker is on camera.
## Selfie POV Framing That Actually Works

this is where most early Seedance 2.0 prompts fail in obvious ways.
operators write "phone held at arm's length" or "selfie shot of a woman" and get back a clip showing a woman literally holding a phone in front of her face. the model rendered the description literally.
the camera doesn't show the selfie POV. the camera shows a wide shot of someone holding a phone.
The Fix
phrase selfie POV explicitly:
"the camera IS her phone's front-facing selfie camera, we see her face filling the frame from the phone's perspective, no phone visible in the shot, just her looking directly into the lens as if talking to her followers."
every key phrase in that description matters:
"the camera IS her phone's front-facing selfie camera" establishes that the viewing perspective is the phone's perspective, not an external camera filming someone with a phone.
"face filling the frame from the phone's perspective" establishes the framing the selfie POV produces.
"no phone visible in the shot" explicitly removes the phone from the rendered output, preventing the literal interpretation that produces shots of someone holding a phone.
"looking directly into the lens" establishes eye contact with the camera, which is what selfie POV creates.
Stability Direction
for cloud phones in particular, real selfie videos are usually filmed with the phone propped up rather than handheld, which creates a specific stability characteristic.
for propped-phone stability: "subtle natural shake mimicking a propped-up phone."
for fully locked-off stability: "steady locked-off front-facing camera view."
for handheld instability (rare in real UGC, more common in vlog-style content): "subtle handheld shake, organic camera movement."
specifying the stability characteristic prevents the model from defaulting to either too-still (which reads as tripod-mounted production) or too-shaky (which reads as poorly-filmed amateur content).
## Real-Brand Props as the AI-Detection Defense

this is the single biggest realism multiplier in Seedance 2.0 prompting, and almost no operators use it correctly.
AI-generated video defaults to generic blurry backgrounds because logos are "safer" for the model to render. that exact safety is what tips viewers off as artificial. real environments contain visible brand logos everywhere: water bottles, coffee cups, product packaging, store signs, clothing tags, electronic devices.
always name 2 to 4 specific real brands visible in the scene.
Real-Brand Prop Library by Environment
bedroom or vanity scenes: Stanley tumbler (sage green), Target bag, Red Bull, Sephora bag, AirPods case, Glossier products, Celsius can.
kitchen scenes: Starbucks cup, Trader Joe's salad container, MacBook with stickers, Whole Foods bag, Brita pitcher.
car scenes: Liquid Death can in cup holder, AirPods case on dashboard, iPhone charger cable, Whole Foods bag on passenger seat.
street scenes: Starbucks cup, visible storefronts (Whole Foods, Trader Joe's), Sephora tote, New Balance sneakers, yellow cab.
podcast or studio scenes: Celsius can, Stanley tumbler, stickered MacBook, AirPods Max, LED strip lighting.
Why This Works Even When Logos Render Slightly Off
the shape and color carry most of the realism signal even if the logos render slightly off. a green Stanley tumbler shape on the desk reads as a Stanley tumbler even if the logo on the side is slightly distorted or simplified. the color and silhouette do the heavy lifting.
for high-stakes content where logo accuracy matters more, use a reference image of the actual product to lock the rendering. for most AI UGC content, the shape-and-color approach produces enough realism signal to pass the organic content test.
## On-Screen Text (Don't)
Seedance 2.0 cannot spell reliably. any request for readable text, logos, captions, or UI in frame renders as garbled gibberish.
never ask the model to render:
brand logos. add in post.
captions or subtitles. use CapCut auto-captions or equivalent.
app UIs. use a reference image and accept it'll be approximate.
website URLs in text. voiceover says it, the logo shows it in post.
platform logos. composite later.
The Workaround
generate the clip text-free. add logos, captions, and brand names in CapCut, Premiere, or After Effects. for karaoke-style word-at-a-time captions, CapCut auto-caption with the "bouncy" style works well.
for app reveal shots ("look at this app on my phone"), the production workflow is:
1. generate the selfie talking-head footage as one clip
2. film or screen-record the actual app separately
3. composite the real app screen into the insert beat in post
this is how 99% of real UGC creators handle app reveals anyway, so the assembled output reads as authentic.
## Trademark and Censorship Workarounds
named trademarked games and brands trigger censorship filters that either block the generation entirely or produce significantly degraded output.
the fix: describe the visual style, not the trademark name.
Common Trademark Workarounds
GTA 5 becomes: "stylized realism 3D animated cinematic in the visual language of an open-world crime drama video game cutscene."
Los Santos becomes: "sun-bleached West Coast city with palm trees."
Fortnite becomes: "battle royale aesthetic."
Marvel becomes: "superhero blockbuster look."
Nike becomes: "iconic striped sneakers" or use a generic real brand: "white New Balance sneakers."
The Underlying Principle
describe the visual DNA (proportions, shading, lighting, color grade) instead of naming the game or brand directly.
other phrases that produce the GTA aesthetic without triggering filters: "stylized realism 3D animation." "open-world video game cutscene aesthetic." "cel-shaded edges with realistic proportions." "early 2010s AAA action game cinematic." "neo-noir West Coast crime thriller rendered as stylized 3D game cinematic."
each of these communicates the visual style to the model without naming a specific trademarked property, which means the generation runs cleanly without censorship filter interference.
## Common Mistakes That Cause Regeneration Cycles

operators who can't get Seedance 2.0 to produce usable output on the first generation are typically making 1 or more of these specific mistakes:
stuffing too many micro-actions into a single clip. keep it to 1 primary action plus 1 to 2 secondary beats, calibrated to the duration. anything more produces visual chaos regardless of whether the clip is 5 seconds or 15 seconds.
writing scripts separately from the prompt. always embed dialogue inline using the "says [exact line]" pattern.
abstract style references. "Tesla Gigafactory meets heist movie" gets rendered as something completely different from what the operator imagined. describe concrete objects.
over-stacking style tags. pick 1 clear genre or aesthetic tag rather than combining 3 to 4 that conflict with each other.
polished lighting language in UGC prompts. just say "window light." cinematic lighting language produces cinematic results, which is the opposite of what AI UGC needs.
literal phone interpretation. write "the camera IS the phone's front-facing camera" rather than "phone held at arm's length."
expecting on-screen text to spell correctly. it won't. composite in post.
hyphenated acronyms. "U-G-C" gets read as "UGG." use "You Gee See" or "U G Cee."
voiceover brand names. put brand names on visible-speaker beats where the lip-sync constraint forces cleaner pronunciation.
generic backgrounds. always name 2 to 4 specific real visible brands in the scene.
## What This Means for Production Economics
operators running Seedance 2.0 with the prompting structure above generate usable output on the first or second generation consistently.
operators using the structure are producing 3 to 5x more finished clips per generation budget compared to operators producing through trial-and-error prompting that requires 4 to 6 regenerations per usable clip.
this changes the unit economics of AI UGC production significantly. at high production volumes, the difference between operators with mature Seedance 2.0 prompting disciplines and operators still figuring out the prompting structure compounds into thousands of dollars per month of differential generation costs.
building this prompting structure into the standard production workflow is what makes 30 to 50 reels per day per portfolio operationally and economically feasible. without the prompting discipline, the same production volume would require 3 to 5x the generation budget, which compresses margin to the point where the operation isn't sustainable at scale.
P.S. - if you just want us to implement this entire ai ugc structure for your campaigns instead...
DM me "SEE" on X (@adriansolarzz) and I'll show you exactly how we'd apply these principles to your specific offer.
- adrian
## 相关链接
- [Adrian Solarz](https://x.com/adriansolarzz)
- [@adriansolarzz](https://x.com/adriansolarzz)
- [5.6K](https://x.com/adriansolarzz/status/2052501455169568875/analytics)
- [@adriansolarzz](https://x.com/@adriansolarzz)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [5:30 AM · May 8, 2026](https://x.com/adriansolarzz/status/2052501455169568875)
- [5,618 Views](https://x.com/adriansolarzz/status/2052501455169568875/analytics)
---
*导出时间: 2026/5/8 13:34:11*
---
## 中文翻译
# 我们如何破解 Seedance 2.0 提示词,从此再也不用重新生成 AI UGC 视频……
**作者**: Adrian Solarz
**日期**: 2026-05-07T21:30:55.000Z
**来源**: [https://x.com/adriansolarzz/status/2052501455169568875](https://x.com/adriansolarzz/status/2052501455169568875)
---

许多使用 Seedance 2.0 的操作者在获得可用素材之前,往往要浪费积分对同一个片段重新生成 3 到 4 次。
但我们找到了一种提示词结构,能让我们在第一次或第二次生成时稳定地获得可用输出。
所以在这篇文章中,我将详细拆解 Seedance 2.0 的提示词结构在实际生产中是如何运作的:
- 5 节奏公式
- 对话嵌入规则
- 品牌名发音防御机制
- 自拍 POV 构图
- 以及作为最大的 AI 检测防御手段的“真实品牌道具策略”。
如果你在使用 Seedance 2.0 时,提示词产生的输出总是不稳定,请按照以下方法操作:
## 大多数操作者误解的 15 秒硬性上限

在深入讲解提示词结构之前,有一个硬性的技术限制是每个操作者必须内化于心的,因为大多数人都在不知不觉中与其对抗。
Seedance 2.0 将每一次生成都限制在 15 秒以内。
在模型层面没有变通的方法。你无法通过提示词获得一个 30 秒的片段。每次生成的输出最长只有 15 秒,这就是上限。
试图通过提示词要求“展示 X 的 30 秒片段”来生成更长视频的操作者,最终得到的只是 15 秒的输出,这会导致模型试图将 30 秒的动作压缩到有限的时间窗口内,从而产生视觉上的混乱,完全浪费了这次生成。提示词与模型的硬性上限对抗,结果总是以失败告终。
如何构建更长的内容
更长的内容是通过后期将多个 15 秒的片段拼接在一起而成的。
将每个镜头生成为独立的片段,最长 15 秒。使用第一个片段中的一帧作为第二个片段的参考图像(这可以锁定拼接序列中人物和地点的一致性)。在 CapCut、Premiere 或 Final Cut 中将这些片段拼接起来。匹配所有片段的色调和颗粒感,使拼接后的输出读起来像是一个连贯的整体。
一条 30 秒的 AI UGC Reels 是由 2 个拼接的 15 秒片段组成的。一个 60 秒的长视频是由 4 个拼接的片段组成的。工作流程会随着所需时长线性扩展,但每一次单独的生成都保持在 15 秒的上限之下。
如何有效利用完整的 15 秒
在单次 15 秒的生成中,提示词结构对于时长的利用也至关重要。
在提示词的前期就加载所有人物和环境细节,因为模型在构建生成的视觉基线时,对较早的 token 赋予更高的权重。
使用 1 个连续的镜头运动来承载故事弧线(推镜头、弧形运动、拉镜头),而不是描述“切至”这样的转场。模型会被转场描述搞混,导致生成的输出无法干净地匹配转场的任意一侧。
在世界观内改变光照,而不是改变地点。一个描绘晨光透过百叶窗逐渐照亮场景的 15 秒生成片段会非常干净。而一个试图在片段中途从卧室移动到厨房的 15 秒生成则会产生视觉混乱。
一旦你内化了 15 秒的上限,你的提示词结构就会改变。你不再试图将复杂的多节奏叙事塞进单次生成中,而是开始将每一次生成都设计为在 15 秒窗口内展开的 1 个特定场景,更长的叙事则通过拼接序列来实现。
## 每次生成都遵循的 5 节奏提示词公式

每一个都能在首次生成中稳定产生可用输出的 Seedance 2.0 提示词,都遵循同样的 5 节奏公式。
节奏 1:带有具体视觉细节的主体
每个提示词都以主体和具体的视觉细节开头:年龄、发色和发型、服装、配饰、种族。
模糊的主体描述会导致重新生成时主体出现变化。具体的主体描述会产生一致的主体,符合操作者实际想要创建的效果。
“一个女人”在多次生成中会产生截然不同的输出。“一位 28 岁的女性,留着齐肩的棕色头发,穿着一件超大的白色 T 恤,戴着金色圆环耳环”产生的输出始终能落在指定的视觉范围内。
节奏 2:场景与环境
一个具体的、能唤起画面的地点,而不是抽象的风格参考。
抽象的风格参考会产生极度不一致的渲染结果,因为模型必须从视觉上解读抽象参考的含义。具体的地点描述则为模型提供了具体的视觉目标。
“奢华美学”是抽象的。“一个阳光充足的厨房,有大理石台面,一个放着樱桃番茄的木质砧板,以及透过水槽上方的窗户射入的晨光”是具体的。
具体的描述能够呈现出抽象描述所试图唤起的奢华美学,因为具体的视觉元素组合成了这种美学,而不是让模型去猜测这种美学意味着什么。
节奏 3:片段时长内的动作
1 到 2 个清晰的主要动作,再加上一两个次要节奏,要根据片段时长进行校准。
在提示词中塞入 6 到 8 个不同动作的操作者会产生视觉混乱,因为即使有 15 秒的上限,模型也无法在时长内干净地容纳所有动作。
公式是:1 个作为片段锚点的主要动作,以及 1 到 2 个增加维度的次要节奏。对于接近 15 秒上限的较长片段,你可以扩展到 2 个主要动作,并在它们之间进行平滑过渡。“她一边看着窗外喝了一口咖啡,然后转向镜头开始说话”是 2 个主要动作,它们之间有着干净的过渡,读起来自然而不仓促。
对于较短的片段(5 到 8 秒),1 个主要动作加上 1 个次要节奏是合适的密度。无论片段时长如何,提示词结构保持不变。只有提示词内的动作密度随时长缩放。
节奏 4:镜头指导
使用命名的电影术语来描述明确的镜头运动,而不是模糊的形容词。
在电影数据上训练过的模型对电影词汇有反应。“追踪镜头”、“推轨镜头推近”、“轨道环绕”、“弧形环绕”、“起重机上升”、“快速摇摄”、“带有自然抖动的手持拍摄”、“固定镜头”、“第一人称 POV”都能产生可预测、可控的镜头行为。
像“有趣的镜头运动”这样模糊的术语会产生不可预测的输出,因为模型必须猜测“有趣”在这个语境下意味着什么。
节奏 5:光照 + 风格 + 情绪标签
堆叠在提示词的末尾,而不是分散在各处。
光照承载的情绪信息几乎比任何其他变量都多。在提示词末尾给光照一个专门的部分,比将其埋没在其他场景描述列表中能产生更强的输出效果。
命名的光照术语效果最好:“黄金时刻”、“蓝色时刻”、“魔幻时刻”、“轮廓光”、“逆光”、“体积雾”、“暖钨丝灯”、“冷荧光灯”、“霓虹点缀”、“来自摄像机左侧的柔和窗光”。
## 对话必须内联嵌入

这是新的 Seedance 2.0 操作者最常见的失败模式。
他们先写视觉场景描述,然后在下面写一个单独的“旁白脚本”部分。模型将这些视为两个独立的输入,产生的输出中对话和视觉画面无法正确同步。
对话必须使用“says [exact line](说[具体台词])”的归属模式内联嵌入。
错误的方法:
“一个阳光充足的厨房里,一位 28 岁的女性喝了一口咖啡。
旁白脚本:‘我真的无法相信这款产品改变了我的生活。’”
正确的方法:
“在一个阳光充足的厨房里,一位 28 岁的女性喝了一口咖啡,放下杯子,用兴奋的大眼睛直视镜头,并说,‘我真的无法相信这款产品改变了我的生活。’”
对话嵌入的规则
每次都使用“says(说)”归属。她说,他问,声音结束。将每一句口语与其匹配的视觉时刻 tied 在同一个句子中。
将品牌名称放在可见说话者的节奏上,即摄像机能看到说话者嘴巴的地方。口型同步能强制产生比旁白覆盖空镜更清晰的发音。
避免在品牌名称上使用旁白覆盖空镜。当说话者不可见时,模型会含糊地念品牌名,因为没有视觉口型同步的约束来强制清晰的发音。
结尾要有力。一句干净的 CTA(行动号召)台词,而不是 3 句冗余的。将多个 CTA 塞进片段时长会导致音频超时并被切断,无论你是生成 8 秒还是 15 秒。
## 品牌名称发音防御机制
这是 Seedance 2.0 生产中最大的发音问题,它有一个结构化的解决方案。
Seedance 2.0 中的 TTS(文本转语音)经常误读品牌名。UGC 会变成“UGG”(像那个靴子品牌)。新品牌名会被连读到周围的对话中。短缩写词会在字母之前被解析为单词。多音节品牌名会被压缩成含糊不清的单个词。
解决方法是对每一个品牌名称提及应用 3 层发音防御。
第 1 层:语音拼写
用语音方式拼写口语缩写:UGC 变成“You Gee See”或“U G Cee”。UGCMaxxing 变成“U G Cee Maxxing”或“You Gee See Maxxing”。SaaS 变成“Sass”(作为一个单词)或“S. A. A. S.”(作为字母)。AI 变成带点的“A. I.”。CEO 变成“C. E. O.”。
多音节品牌名被拆分:Cantina 变成“Can Teena”(两个词,而不是一个)。Klaviyo 变成“Klay Vee O”。模型容易含糊带过的每个音节都需要明确的语音指导。
第 2 层:发音指导
在口语台词前添加发音指导:
“她稍微放慢速度,清晰地念出品牌名称,并说……”
“声音在镜头上清晰响亮,每个词都清晰可辨……”
“她一边说,一边清晰地强调每个音节……”
这些指导告诉模型专门在品牌名称位置放慢音频速度,这比让音频以默认速度运行能产生更清晰的发音。
第 3 层:带破折号的重复
在片段中间说两遍品牌名称,中间用破折号分隔重复:
“它叫 Can Teena,Can Teena,我简直迷死它了。”
这同时起到了 3 个作用:它在对话中听起来很自然(真实的人在介绍品牌时会重复品牌名);它给了 TTS 两次机会来正确发音;即使模型仍然略有含糊,它也能教会观众正确的发音。
品牌名称放置规则
品牌名称仅用于可见说话者的节奏。
当说到品牌名称时,摄像机应该能看到说话者的嘴巴。口型同步迫使模型产生更清晰的音频,因为视觉约束要求嘴唇必须与单词匹配。
旁白覆盖空镜的节奏(说话者不可见)会导致品牌名称含糊不清,因为没有视觉约束强制清晰发音。永远不要把品牌名称放在旁白节奏中。构建脚本,使品牌名称总是落在说话者出现在镜头上的时候。
## 真正有效的自拍 POV 构图

这是大多数早期的 Seedance 2.0 提示词明显失败的地方。
操作者写“手臂伸直拿着手机”或“女人的自拍镜头”,得到的片段显示一个女人真的在脸前拿着手机。模型按字面意思渲染了描述。
摄像机没有显示自拍 POV。摄像机显示的是某人拿着手机的广角镜头。
修复方法
明确表述自拍 POV:
“摄像机就是她手机的前置自拍摄像头,我们从手机的视角看到她的脸填满画面,镜头中看不到手机,只有她直视镜头,就像在与粉丝交谈一样。”
该描述中的每个关键短语都很重要:
“摄像机就是她手机的前置自拍摄像头”确立了观看视角是手机的视角,而不是外部摄像机拍摄某人拿着手机。
“从手机的视角填满画面的脸”确立了自拍 POV 产生的构图。
“镜头中看不到手机”明确地从渲染输出中移除了手机,防止了产生某人拿着手机的镜头的字面解释。
“直视镜头”确立了与摄像机的眼神接触,这正是自拍 POV 所创造的。
稳定性指导
特别是对于云手机,真实的自拍视频通常是用手机支撑架拍摄的,而不是手持,这产生了一种特定的稳定性特征。
对于支架手机的稳定性:“模仿支撑架放置手机的细微自然抖动。”
对于完全固定的稳定性:“稳定的固定前置摄像头视角。”
对于手持不稳定性(在真实 UGC 中很少见,在 vlog 风格内容中较常见):“细微的手持抖动,有机的摄像机运动。”
指定稳定性特征可以防止模型默认为过于静止(读起来像三脚架拍摄的商业作品)或过于晃动(读起来像拍摄拙劣的业余内容)。
## 作为 AI 检测防御的真实品牌道具

这是 Seedance 2.0 提示词中最大的真实感倍增器,几乎没有人正确使用它。
AI 生成的视频默认为通用的模糊背景,因为渲染 logo 对模型来说更“安全”。而这种“安全”恰恰是向观众透露它是人造内容的线索。真实的环境到处都可见品牌 logo:水瓶、咖啡杯、产品包装、商店招牌、衣服标签、电子设备。
始终在场景中命名 2 到 4 个可见的具体真实品牌。
按环境分类的真实品牌道具库
卧室或梳妆台场景:Stanley 随行杯(鼠尾草绿)、Target 袋子、红牛、Sephora 袋子、AirPods 盒子、Glossier 产品、Celsius 罐。
厨房场景:星巴克杯子、Trader Joe's 沙拉容器、贴纸贴满的 MacBook、Whole Foods 袋子、Brita 水壶。
车内场景:杯架里的 Liquid Death 罐、仪表盘上的 AirPods 盒子、iPhone 充电线、副驾驶座上的 Whole Foods 袋子。
街头场景:星巴克杯子、可见的店面(Whole Foods、Trader Joe's)、Sephora 手提袋、New Balance 运动鞋、黄色出租车。
播客或工作室场景:Celsius 罐、Stanley 随行杯、贴纸贴满的 MacBook、AirPods Max、LED 灯带。
为什么即使 logo 渲染得有点不准,这招依然有效
即使 logo 渲染得稍微有点偏差,形状和颜色也承载了大部分的真实感信号。桌上的绿色 Stanley 随行杯形状就会被读作 Stanley 随行杯,即使侧面的 logo 稍微扭曲或简化了。颜色和轮廓起到了主要作用。
对于 logo 准确性更重要的关键内容,请使用实际产品的参考图像来锁定渲染。对于大多数 AI UGC 内容,形状加颜色的方法足以产生通过有机内容测试的真实感信号。
## 屏幕文字(不要做)
Seedance 2.0 无法可靠地拼写。任何对可读文字、logo、字幕或 UI 的请求都会渲染成乱码。
永远不要要求模型渲染:
品牌 logo。后期添加。
字幕或副标题。使用 CapCut 自动字幕或同等功能。
App UI。使用参考图像,并接受它会是近似的。
文本中的网站 URL。旁白说出它,logo 在后期展示。
平台 logo。稍后合成。
变通方法
生成不带文字的片段。在 CapCut、Premiere 或 After Effects 中添加 logo、字幕和品牌名称。