# The 5 AI UGC Formats We Run and When to Use Each
**作者**: Adrian Solarz
**日期**: 2026-07-23T20:50:15.000Z
**来源**: [https://x.com/adriansolarzz/status/2080395086379061587](https://x.com/adriansolarzz/status/2080395086379061587)
---

picking the wrong format is one of the best ways to waste good content.
a strong angle in the wrong shape underperforms a mediocre angle in the right one, because the format decides how the message actually reaches someone.
these are the 5 formats we run, what each one does, and exactly when to reach for it.
let's get into it.
## Format Follows the Content
before the 5, the rule that governs all of them.
the format is a decision, not a default, and it gets made from the content and the audience rather than from what you feel like making.
a transformation story wants a different shape than a product explanation, and forcing either into the wrong one costs you reach you already paid to produce.
so the question is always what this specific message needs, and the 5 formats below are the answers.
# Format 1: The Talking Head
a creator speaking directly to camera, no props, no demo, just a person making a point.
## The Mechanic
it runs on trust and relatability.
a person talking to you like a friend bypasses ad-skepticism in a way almost nothing else does, because that's exactly what real UGC looks like.
the whole format is built on the viewer believing there's a real person on the other side.
## When to Reach for It
for opinions, takes, and warnings, since a claim lands harder from a person than from a montage.
for cold audiences, because trust has to come before information, and a face builds it faster than anything.
and for building a recognizable creator, since repetition of the same person across many talking heads is what creates parasocial familiarity.
## How to Build It
this is the format where realism matters most, because the viewer is looking directly at a face for the whole clip.
the full realism stack rides on every generation, the texture clause, the anti-polish language, the named lighting, the expression clause with life in the eyes.
keep the action simple, small natural head movements and hand gestures, since busy motion is where generation still struggles.
and keep the sync clean, because drift between the lips and the audio is the tell that ends the clip.
## Where It Fails
a talking head with nothing to say is the weakest content there is.
the format has no visual interest carrying it, so the point has to be genuinely worth hearing, or there's nothing to hold the viewer at all.
# Format 2: The Demo
the creator showing the product working, in hand, in use, in the real world.
## The Mechanic
it runs on proof.
seeing a thing work is more convincing than being told it works, and the demo collapses the distance between claim and evidence.
## When to Reach for It
for products with a visible result, since anything with a before and after moment inside a single clip demos beautifully.
for solution-aware audiences, the people who know they have the problem and are comparing options, because the demo answers "does it actually work."
and for anything that's hard to explain in words, where 3 seconds of showing beats 30 seconds of describing.
## How to Build It
the difficulty here is hands and objects, which are still the weakest area in video generation.
keep the interaction simple, a single clear motion rather than an intricate sequence, and generate extra variants on any clip with detailed hand work.
for the tricky shots, graduate the clip to the premium tier where the fidelity is better, since a demo with distorted hands destroys the proof it exists to provide.
reference consistency matters too, because the product has to look like the same product in every shot.
## Where It Fails
products with no visible action don't demo, and forcing them into the format produces a clip of someone holding a box.
when there's nothing to show, pick a different format instead of showing nothing convincingly.
# Format 3: The Before and After
the transformation format, opening on a starting state and revealing the outcome.
## The Mechanic
it runs on the gap.
the distance between the 2 states is the story, and the viewer stays to close it, which makes this format naturally strong on completion.
relief and aspiration are the drivers, and both are heavy enough to carry a piece to the end.
## When to Reach for It
for anything with a visible outcome over time, skincare, fitness, home, finance, organization.
for problem-aware audiences, the people living in the before state who want to see the after is possible.
and when you have real proof, since the format depends entirely on the contrast being convincing.
## How to Build It
generate the 2 states as separate pieces with identical references and identical lighting, because the same recognizable person in both states is what sells the journey.
the only thing that changes between them is the emotional register, tired and frustrated in the before, relieved and quietly confident in the after.
any drift in the face, the room, or the light breaks the illusion that it's the same person across time, which is the entire premise.
## Where It Fails
an exaggerated transformation reads as fake instantly, and the audience punishes it hard.
the gap has to be believable, so a modest honest change outperforms a dramatic unbelievable one every time.
# Format 4: The Testimonial
a person recounting their own experience with the product or the outcome.
## The Mechanic
it runs on social proof.
people look to other people to decide what's safe and what's worth buying, and a testimonial borrows that confidence directly.
## When to Reach for It
for product-aware audiences, the people already considering it who need a reason to commit.
for purchases people think hard about, where the risk feels real and someone else having taken it first matters.
and for niches that run on trust, health, finance, anything where the viewer is asking who to believe.
## How to Build It
the tone is everything here, warm and earnest rather than energetic.
the expression clause matters more than usual, since a testimonial delivered with dead eyes is worse than no testimonial at all.
keep the language natural and specific, the details a real person would mention, because polished marketing phrasing kills the format's only advantage.
## Where It Fails
a testimonial that sounds scripted is worse than nothing, because it undermines trust instead of building it.
specificity is the fix, since real people mention odd particular details and ad copy never does.
# Format 5: The Slideshow
a sequence of still images with text overlays, swiped through instead of watched.
## The Mechanic
it runs on the swipe.
every swipe is an engagement signal, and the format is built from open loops that pull the viewer forward slide by slide.
it's also the cheapest format to produce at volume, since stills cost less than motion and skip the motion failure modes entirely.
## When to Reach for It
for list and reference content, since numbered and structured information is what people save.
for volume, because slideshows produce faster and cheaper than video, which makes wide testing affordable.
and when you want saves, since the format is the most save-friendly thing you can make.
## How to Build It
pick 1 of the 5 proven structures, hear me out, non-negotiables, green flags, ranked list, or before and after.
every slide loads the same 4 references, character, aesthetic baseline, scene, and slide template, because visual drift across slides is what gives automated slideshows away.
slide 1 is the hook and slide 2 is a second hook, since the algorithm can serve the piece starting from slide 2.
1 idea per slide, phrasing that pulls forward, and the final slide carries the keyword CTA.
## Where It Fails
slideshows are weak for anything needing motion or demonstration.
if the message depends on seeing something happen, stills can't carry it, and the format works against you.
## What Each Format Costs to Make
the formats aren't equal on production cost, and that shapes how you use them.
slideshows are the cheapest, since stills generate faster and cheaper than motion, which is why they carry the bulk of wide testing.
talking heads sit in the middle, a single subject with simple motion, so they generate reliably at the cheap tier.
demos cost the most, because hands and object interaction need extra variants and often a premium render to look right.
before and afters cost double by definition, since you're generating 2 matched states instead of 1 clip.
and testimonials cost like talking heads, with the extra care going into the expression rather than the render.
so the cheap formats carry the volume and the expensive ones get pointed at proven angles, which is the same routing logic applied at the format level.
## Matching Format to Awareness Stage
the cleanest way to choose is by where the viewer already is.
cold and unaware: the talking head or the slideshow, since both build interest before the product ever enters.
problem-aware: the before and after, because the viewer is living in the before and wants to see the after is real.
solution-aware: the demo, since they're comparing options and want proof it works.
product-aware: the testimonial, because they're deciding, and someone else's experience is what tips it.
run the mix across a batch and you're covering the whole funnel instead of hammering 1 stage.
## Mixing Formats Across a Batch
no single format carries an operation, so the daily slate runs several.
the mix gets set by the brief, weighted toward whatever's performing that week, and rotated when a format shows fatigue.
a batch that's all talking heads exhausts the audience, and a batch that's all slideshows leaves the trust-building formats unused.
variety across the slate also gives you comparative data, so you learn which formats your specific niche responds to instead of guessing.
## Testing Formats Against Each Other
when you're unsure which format a message wants, run it in 2.
hold the angle, the hook, and the creator steady, and change only the format across the pair.
whichever earns the better completion and save rate is the answer, and now you know for that type of message in your niche.
format testing is slower than hook testing and it's worth doing a few times, because the answer applies to every future piece of that type.
## 相关链接
- [Adrian Solarz](https://x.com/adriansolarzz)
- [@adriansolarzz](https://x.com/adriansolarzz)
- [716](https://x.com/adriansolarzz/status/2080395086379061587/analytics)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [4:50 AM · Jul 24, 2026](https://x.com/adriansolarzz/status/2080395086379061587)
- [716 Views](https://x.com/adriansolarzz/status/2080395086379061587/analytics)
---
*导出时间: 2026/7/24 11:08:32*
---
## 中文翻译
# 我们运行的 5 种 AI UGC 形式以及何时使用每种形式
**作者**: Adrian Solarz
**日期**: 2026-07-23T20:50:15.000Z
**来源**: [https://x.com/adriansolarzz/status/2080395086379061587](https://x.com/adriansolarzz/status/2080395086379061587)
---

选择错误的形式是浪费好内容的最佳方式之一。
错误形式下的强力切入角度,其表现往往不如正确形式下的平庸切入角度,因为形式决定了信息实际上如何触达受众。
以下是我们运行的 5 种形式,每种形式的作用,以及确切的使用时机。
让我们深入了解。
## 形式追随内容
在介绍这 5 种形式之前,先要了解支配它们所有形式的规则。
形式是一种决策,而非默认选项,它是根据内容和受众来决定的,而不是根据你的制作意愿来决定的。
一个转变故事需要一个不同于产品解释的形式,强行将任何一种内容塞入错误的形式,会让你损失已经花钱制作出来的触达机会。
所以问题永远是:这条特定的信息需要什么?下面的 5 种形式就是答案。
# 形式 1:口播
创作者直接对着镜头说话,没有道具,没有演示,只有一个人在阐述观点。
## 机制
它依靠信任和亲和力运行。
一个人像朋友一样和你交谈,几乎能避开所有的广告怀疑心理,因为这正是真实 UGC 的样子。
整个形式建立在观众相信屏幕另一端有一个真人的基础上。
## 何时使用
用于表达观点、看法和警告,因为相比蒙太奇剪辑,由人亲自提出的论点更有分量。
用于冷受众,因为在获取信息之前必须先建立信任,而一张面孔建立信任的速度比任何东西都快。
也用于打造可识别的创作者,因为在多个口播视频中反复出现同一张面孔,正是建立准社会关系熟悉感的关键。
## 如何构建
在这个形式中,真实感最重要,因为观众在整个片段中都会直视这张脸。
每一次生成都要加载完整的真实感堆栈:纹理条款、反修饰语言、指定灯光、眼神有神采的表情条款。
保持动作简单,通过自然的小幅头部移动和手势来表现,因为复杂的动作仍然是生成技术的弱项。
保持口型同步干净,因为嘴唇和音频之间的漂移是暴露片段真相的致命破绽。
## 它在哪里失效
一个无话可说的口播是所有内容中最弱的。
这种形式没有视觉趣味来支撑,所以观点必须真正值得聆听,否则没有任何东西能留住观众。
# 形式 2:演示
创作者展示产品在工作、在手边、在使用中、在现实世界中的样子。
## 机制
它依靠证据运行。
眼见为实,看到东西运作比听别人说它有效更有说服力,演示消除了主张和证据之间的距离。
## 何时使用
用于有可见结果的产品,因为任何在单个片段内包含“使用前和使用后”时刻的产品都非常适合演示。
用于有解决方案意识的受众,即那些知道自己有问题并正在比较选项的人,因为演示回答了“它真的有用吗?”这个问题。
也用于任何难以用语言解释的事物,因为 3 秒钟的展示胜过 30 秒的描述。
## 如何构建
这里的难点在于手和物体,这仍然是视频生成中最薄弱的环节。
保持交互简单,单一清晰的动作而不是复杂的序列,并对任何涉及精细手部动作的片段生成额外的变体。
对于棘手的镜头,将片段升级到保真度更高的高级层,因为一双变形的手会毁掉演示试图提供的证明。
参考一致性也很重要,因为产品在每个镜头中看起来必须是同一个产品。
## 它在哪里失效
没有可见动作的产品无法演示,强行将其套入这种形式只会产生一个有人拿着盒子的片段。
当没有什么可展示的时候,选择不同的形式,而不是令人信服地展示虚无。
# 形式 3:使用前后
转变形式,以初始状态开始并揭示结果。
## 机制
它依靠差距运行。
两种状态之间的距离就是故事,观众留下来是为了填补这个差距,这使得这种形式天然具有很强的完播率。
解脱感和渴望是驱动力,两者都足够强大,能将观众留住直到最后。
## 何时使用
用于任何随时间推移有可见结果的事物,如护肤、健身、家居、理财、整理。
用于有问题意识的受众,即生活在“使用前”状态并希望看到“使用后”是可能的人们。
当你拥有真实证据时,因为这种形式完全依赖于对比的可信度。
## 如何构建
使用相同的参考和相同的灯光将 2 种状态生成为独立的片段,因为在两种状态下出现同一个可识别的人是推销这一旅程的关键。
它们之间唯一改变的是情绪基调:使用前是疲惫和挫败,使用后是如释重负和沉静的自信。
面部、房间或光线上的任何漂移都会打破这是同一个人跨越时间的错觉,而这正是整个形式的前提。
## 它在哪里失效
夸张的转变会立刻被识破为虚假,观众会对此进行严厉的惩罚。
差距必须可信,因此一个适度的诚实的改变,每次都能胜过戏剧性的不可信的改变。
# 形式 4:证言
一个人讲述他们自己对产品或结果的经历。
## 机制
它依靠社会认同运行。
人们通过观察他人来决定什么是安全的、什么是值得购买的,证言直接借用了这种信心。
## 何时使用
用于有产品意识的受众,即已经在考虑但需要一个理由来承诺购买的人。
用于人们会深思熟虑的购买行为,在这些情况下风险感觉很真实,而有人先行尝试就显得很重要。
也适用于依靠信任运行的细分领域,如健康、金融,任何观众在问“该相信谁”的领域。
## 如何构建
语气是这里的一切,应该是温暖和真诚的,而不是充满活力的。
表情条款比往常更重要,因为用死气沉沉的眼睛传达的证言比没有证言更糟糕。
保持语言自然和具体,包含一个真实的人会提到的细节,因为经过修饰的市场营销措辞会扼杀这种形式唯一的优势。
## 它在哪里失效
听起来像照本宣科的证言比没有更糟,因为它破坏了信任而不是建立信任。
具体性是解决办法,因为真实的人会提到奇怪的细节细节,而广告文案绝不会。
# 形式 5:幻灯片
一系列带有文字叠加的静态图像,通过滑动浏览而不是观看。
## 机制
它依靠滑动运行。
每一次滑动都是一个参与信号,这种形式由未完结的悬念构建而成,拉动观众逐张向前浏览。
这也是批量生产成本最低的形式,因为静态图片比动态视频生成得更快更便宜,并且完全跳过了动态生成的失败模式。
## 何时使用
用于列表和参考内容,因为编号和结构化的信息正是人们会保存的内容。
用于批量生产,因为幻灯片制作得比视频更快更便宜,这使得广泛的测试变得负担得起。
当你想要收藏量时,因为这种形式是你能制作的最利于收藏的内容。
## 如何构建
从 5 种经过验证的结构中选择一种:“听我说”、“不可协商的要点”、“积极信号”、“排名列表”或“使用前后”。
每张幻灯片加载相同的 4 个参考:角色、美学基线、场景和幻灯片模板,因为幻灯片之间的视觉漂移是暴露自动生成的幻灯片的破绽。
第 1 张幻灯片是钩子,第 2 张幻灯片是第二个钩子,因为算法可能会从第 2 张幻灯片开始推送这个内容。
每张幻灯片 1 个观点,使用拉动向前的措辞,最后一张幻灯片承载关键词 CTA(行动号召)。
## 它在哪里失效
对于任何需要运动或演示的事物来说,幻灯片都很弱。
如果信息依赖于看到某事发生,静态图片无法承载它,这种形式会对你不利。
## 制作每种形式的成本
各种形式在生产成本上并不相等,这决定了你如何使用它们。
幻灯片最便宜,因为静态图片比动态视频生成得更快更便宜,这就是为什么它们承担了大部分广泛测试的任务。
口播处于中间位置,单一主体和简单动作,所以可以在廉价层可靠生成。
演示成本最高,因为手和物体交互需要额外的变体,通常还需要高级渲染才能看起来正确。
使用前后成本翻倍是注定的,因为你生成的是 2 个匹配的状态而不是 1 个片段。
证言的成本与口播相似,只是额外的心思花在表情上而不是渲染上。
因此,廉价的形式承担批量任务,昂贵的形式则用于经过验证的切入角度,这是在形式层面应用的同样的路由逻辑。
## 将形式与意识阶段相匹配
选择的最清晰方法是看观众已经处于哪个阶段。
冷受众和无意识受众:口播或幻灯片,因为两者都能在产品介入之前建立兴趣。
有问题意识受众:使用前后,因为观众生活在“使用前”的状态,想要看到“使用后”是真实的。
有解决方案意识受众:演示,因为他们正在比较选项,想要证明它有效。
有产品意识受众:证言,因为他们正在做决定,而他人的经历正是促成决定的关键。
在一批内容中混合运行这些形式,你就能覆盖整个漏斗,而不是只盯住一个阶段。
## 在一批内容中混合形式
没有单一的形式能支撑整个运营,所以每日的排片表会运行几种形式。
混合比例由简报设定,倾向于当周表现最好的形式,当某种形式显示出疲劳时进行轮换。
一批全是口播的内容会让受众感到厌倦,一批全是幻灯片的内容则会让建立信任的形式闲置。
排片表上的多样性也能给你提供对比数据,这样你就能了解你的特定细分领域对哪些形式有反应,而不是靠猜测。
## 测试不同形式
当你不确定一条信息需要哪种形式时,用两种形式运行它。
保持切入角度、钩子和创作者不变,只在这一对中改变形式。
whichever 获得更好的完播率和收藏率,就是答案,现在你就知道了针对你细分领域中的这类信息该用什么。
形式测试比钩子测试慢,但值得做几次,因为答案适用于未来每一个同类型的内容。
## 相关链接
- [Adrian Solarz](https://x.com/adriansolarzz)
- [@adriansolarzz](https://x.com/adriansolarzz)
- [716](https://x.com/adriansolarzz/status/2080395086379061587/analytics)
- [升级到 Premium](https://x.com/i/premium_sign_up)
- [7月 24, 2026 上午 4:50](https://x.com/adriansolarzz/status/2080395086379061587)
- [716 次观看](https://x.com/adriansolarzz/status/2080395086379061587/analytics)
---
*导出时间: 2026/7/24 11:08:32*