# How We're Going VIRAL by Automating TikTok Slideshows With Claude Co-work + GPT Image 2
**作者**: Adrian Solarz
**日期**: 2026-04-30T19:28:47.000Z
**来源**: [https://x.com/adriansolarzz/status/2049934006633021536](https://x.com/adriansolarzz/status/2049934006633021536)
---

most people in 2026 are pouring everything into TikTok videos and missing the format that the platform is actively pushing harder than almost anything else right now.
TikTok slideshows.
the algorithm is treating the format like a priority distribution channel, the audience is engaging with the swipe-through dynamic in ways that produce stronger save and share metrics than most video content, and the production costs are dramatically better than anything video-based at equivalent quality.
so we automate the entire system...
1. scripts and analytical intelligence through Claude Co-work.
2. static slide visuals through GPT Image 2.
3. distribution across the account portfolio.
the production volume across all of it sits at 30 to 50 slideshows per day per portfolio with a workflow that
## Why TikTok Slideshows Are Goated

TikTok's algorithm in 2026 weights signals differently across content formats, and the slideshow format is currently in what we call an algorithmic acceleration window.
the platform is actively trying to grow the format because slideshows produce engagement behaviors that are valuable to the broader platform ecosystem: longer per-impression dwell time as users swipe through slides, higher save rates than equivalent video content, and substantive comment engagement on specific slides that the algorithm weighs more heavily than emoji-only video reactions.
the swipe mechanic is the structural reason for these stronger metrics. a viewer who swipes through 6 slides is actively in your content for the duration of the swipe-through, which is meaningfully different from a viewer who passively watches a video and scrolls past at the first weak moment. the swipe forces a moment of decision at every slide, which produces the kind of active engagement the algorithm reads as evidence of genuine value.
beyond the engagement quality difference, the algorithm distributes slideshows from accounts with smaller followings more aggressively than equivalent video content right now.
this is the acceleration window.
operators who capture it now will benefit from the disproportionate distribution for several months before the format saturates and the algorithmic push normalizes.
the production economics make this acceleration window even more capturable. a slideshow takes a fraction of the production time of a video reel at equivalent quality, which means operators can produce at the volume the algorithmic push window rewards without the production capacity constraints that limit video output.
## The Slideshow Formats That Are Actually Working
before any production system gets built, the format selection has to be right. a well-executed production pipeline running the wrong format produces high-quality content that doesn't perform, because the format itself determines whether the content structure is aligned with what the algorithm is rewarding.
we've tested dozens of slideshow structures across client campaigns and there's a specific set that consistently produces the saves, sends, and substantive comment engagement that drives sustained distribution.
the personal story format opens with a hook slide that promises a specific personal outcome ("I lost 23 pounds in 11 weeks doing this"), walks through the discovery and the mechanism across 4 to 5 middle slides, and closes with a CTA slide that drives action. this format works exceptionally well because the curiosity gap created by the hook compels the viewer to swipe through the entire sequence to resolve the question of how the result was achieved.
the ranked list format presents 5 to 7 items in reverse order, with the strongest item revealed on the final slide. the curiosity gap of "what's number 1" pulls viewers through the entire slideshow, which produces the swipe-through velocity the algorithm reads as strong engagement. this format is particularly effective for content categories where the audience has existing expectations about what the top item should be.
the realization format opens with a specific moment of insight ("I realized why my skin wasn't getting better") and uses each subsequent slide to deepen the realization with specific details. this format generates strong save rates because viewers want to revisit the realization later when they're applying it to their own situation.
the controversial opinion format opens with a statement that contradicts conventional thinking in the niche and uses subsequent slides to defend the position with specific evidence. this format generates the highest comment engagement of any format we run because viewers either strongly agree (and want to add their support) or strongly disagree (and want to argue the position).
the routine breakdown format presents a specific routine or process across 5 to 6 slides, with each slide focusing on 1 step. these earn high save rates because the format is inherently reference-worthy. viewers save it as a guide they want to come back to when they're trying to implement the routine themselves.
these formats already exist in our slideshow library, which means the system isn't generating new strategy for each piece. it's executing proven structures with new content. that's the operational efficiency that makes 30 to 50 slideshows per day feasible without quality dropping.
## How Claude Co-work Generates the Scripts

the scripts and visual pairing notes for every slideshow get generated through Claude Co-work with the creative intelligence brief loaded into context.
the brief includes the persona document, the competitive landscape research, the unique mechanism articulation, and the transformation library. for the slideshow system specifically, the brief gets supplemented with format-specific guidance: the structural template for each slideshow format, the slide count parameters, and the visual pairing logic that connects each slide's text to a specific GPT Image 2 generation prompt.
the master prompt for personal story slideshows:
"using the persona document and creative brief loaded above, write a personal story slideshow script for TikTok. the format is 6 to 7 slides total. slide 1 is the opening hook line that promises a specific personal outcome with a timeline. slides 2 through 5 walk through the discovery and the mechanism, each one a single short text overlay that pairs with a specific visual. the final slide is a low-friction CTA. for each slide, include a GPT Image 2 generation prompt in brackets describing the exact static image to generate, including lighting direction, environmental context, composition, and mood. the visuals should maintain consistent aesthetic across all slides: warm interior lighting, authentic phone selfie feel, casual lifestyle environments, not studio quality."
the output is a complete slideshow script with embedded production prompts in approximately 60 seconds. the same prompt template runs across all 5 formats with the structural parameters adjusted for each.
The Variation Generation That Multiplies Output
once a slideshow has performed strongly in the portfolio, Co-work generates 5 to 6 variations from the single winning piece in a single session rather than producing 5 to 6 entirely new pieces from scratch.
the variation prompt:
"the following slideshow performed well based on the performance data in the performance data folder [paste script]. write 5 variations of this slideshow. each variation should maintain the format structure, the emotional register, and the overall narrative arc, but change the specific personal story details and the visual pairing notes. each variation should target a slightly different segment within the broader persona. preserve the closer line structure and the slide count."
the algorithm sees enough variation between accounts that nothing reads as duplicate content, and the underlying performance pattern that earned the original distribution carries through to the variations because the structural and emotional elements that drove engagement remain intact.
## How GPT Image 2 Generates the Slide Visuals

slideshows are static images, which means the visual generation tool needs to produce high-quality still images rather than video clips. GPT Image 2 is the right tool for this specific workflow because it handles the quality problems that make AI-generated slideshow visuals fail the organic content test in the first second.
Why GPT Image 2 Specifically Over Other Image Models
most AI image generation models in 2026 still struggle with 2 specific failure modes that are catastrophic for slideshow production at scale.
the first failure is product text rendering. for product-adjacent slideshows where the product needs to appear visually, packaging text needs to be legible and accurate. most image models produce garbled or incorrect text on product packaging, which immediately registers as AI to viewers who notice the inconsistency.
GPT Image 2 renders product text accurately, which means product reference shots in slideshows show the correct brand name and label content rather than AI-generated approximations.
the second failure is consistency across a multi-slide sequence. a slideshow needs aesthetic, lighting, and visual style coherence across 6 to 7 independently generated images. most image models produce strong individual images but fail to maintain the visual coherence across a full slide set that makes a slideshow feel like a single intentional piece. GPT Image 2's instruction-following capability handles this through the consistent aesthetic baseline in every slide prompt.
## The Technical Output Specifications
every slide gets generated at the correct specifications for TikTok slideshows in 2026.
aspect ratio is 9:16 portrait, specifically 1080x1920 pixels. this matches TikTok's native vertical format and ensures the slideshow displays correctly in the feed without cropping that would lose critical visual elements.
text overlay positioning follows a specific principle: the upper third or center of the frame is where slideshow text typically lands because viewers' attention enters the frame from the top when they swipe to a new slide. text overlaid in the bottom third risks being cut off on devices with different display configurations, and viewers' eyes don't naturally land there first on slideshow content the way they do on carousel content.
The Prompting Disciplines That Make GPT Image 2 Output Look Organic
scene direction over static description is the foundational discipline that separates GPT Image 2 output that passes the organic content test from output that registers as AI-generated.
"close-up mirror selfie of a woman in her late 20s holding her phone, soft warm bedroom lighting, casual oversized t-shirt, slightly tousled morning hair, authentic phone photo feel, not studio quality, 9:16 aspect ratio" produces output that looks like a real selfie someone took in their bedroom. "a woman taking a selfie" produces a generic stock-style image that fails the authenticity test immediately.
the consistent aesthetic instruction carries visual coherence across all slides in the slideshow. every GPT Image 2 prompt for every slide in the same slideshow should include the same aesthetic baseline: "warm interior lighting, authentic phone selfie or candid lifestyle feel, soft natural framing, organic not studio quality, casual everyday environments." this instruction is what makes the assembled slideshow feel like a single coherent visual narrative rather than a sequence of independently generated images.
skin texture specification for any slide that includes a human face. "realistic skin texture, visible pores around nose and cheeks, natural slight unevenness, no filter quality" in every prompt without exception. the smooth poreless skin that AI image models default to is the fastest route to viewers clocking the slideshow as synthetic.
anti-polish language replaces the produced aesthetic with the organic quality that registers as authentic. "casual phone framing, photographed in a real environment, not a professional set, organic not studio quality" pulls the output away from the polished aesthetic that triggers ad skepticism in cold audience viewers.
lighting as a dedicated clause rather than buried in a list of other scene elements. "soft warm bedroom light from window at left, casting natural shadows, golden hour quality, no harsh highlights, skin properly illuminated without overexposure" as a standalone element in every prompt that includes a human subject. lighting controls mood more than any other variable in image generation.
Generating 6 to 8 Variants Before Selecting
for slides that include human subjects, generate 6 to 8 variants from the same prompt before selecting the final image rather than accepting the first output and moving on.
the range of quality across variants from a single GPT Image 2 prompt is often significant. the weakest variant might have skin texture issues, expression problems, or compositional flaws that immediately register as AI-generated. the strongest variant passes the realism test in the first 2 seconds and is genuinely indistinguishable from a real phone selfie.
selecting from a pool of 6 to 8 variants takes slightly more time per slide but produces a finished slideshow quality level that the first-output-accepted approach cannot match.
## The Distribution Architecture Across the Account Portfolio

the distribution layer is where the production volume actually translates into reach. without the multi-account portfolio, 30 to 50 slideshows a day from a single account would face the structural distribution ceiling that limits any single TikTok account regardless of content quality.
we run client portfolios of 50 accounts to start, scaling based on CPM budget and niche characteristics. each account posts 1 to 3 slideshows daily, which sits in the algorithmic sweet spot for posting frequency on TikTok.
Trending Audio Matching as a Distribution Multiplier
TikTok slideshows can leverage trending audio as a distribution multiplier in ways Instagram carousels can't. a slideshow paired with audio currently in the algorithmic push window gets distributed faster and more broadly than the same slideshow with non-trending audio.
the audio research workflow runs at the start of every production batch. we identify which audio tracks are currently trending in the target niche, evaluate which ones fit the emotional register of the slideshows in the production batch, and assign appropriate audio to each slideshow before scheduling.
this 5 to 10 minute additional research per batch produces meaningful distribution lift because the algorithm uses audio fit as a ranking signal. a slideshow with trending audio relevant to its content category outperforms the same slideshow with random audio consistently enough to make the audio matching workflow non-negotiable.
The Account Warmup Architecture
every new account entering the portfolio goes through a structured warmup protocol before being deployed at full production volume. accounts that skip warmup and immediately post at high frequency are identifiable to TikTok's detection systems and tend to receive suppressed distribution from the start.
the protocol follows a behavioral ramp across the first 4 weeks: profile completion and engagement-only behavior in the first 3 days, simpler content posted every other day from day 4 to day 7, scaling to 1 post daily in week 2, 2 posts daily in week 3 with monitoring of non-follower reach, and full deployment in week 4 once the data confirms algorithmic trust is building.
Posting Time Staggering
posting time staggering across the portfolio prevents the simultaneous posting pattern that flags coordinated networks. the same slideshow format posting across 25 accounts at the same minute is identifiable to TikTok's detection systems as coordinated behavior.
the same format posting across 25 accounts spread across an 8-hour window with each account hitting its own optimal posting time looks like 25 independent accounts independently posting good content. each account's posting window gets calibrated based on the activity patterns of its specific audience to maximize first-hour engagement signals, which the algorithm uses to make initial distribution decisions.
## The Daily Production Rhythm
the production volume of 30 to 50 slideshows per day across 50 accounts sounds like it requires a team of 4 people working full days. with the right system it requires 1 operator working approximately 2 hours.
the analytical session (30 minutes): Co-work runs the analytical prompt against the previous day's performance data and produces the production brief for the current day's batch, identifying which formats and hook approaches the data shows are working strongest.
the production session (60 to 75 minutes): Co-work runs the format-specific script prompts to generate the day's slideshow scripts with embedded GPT Image 2 prompts. the GPT Image 2 generation produces all slide visuals in batches, with the variant selection discipline applied to character-included slides.
the scheduling session (15 to 20 minutes): the finished slideshows get distributed across the account portfolio with each piece assigned to its target account, paired with appropriate trending audio, and scheduled at that account's specific peak posting window.
2 hours per day. 30 to 50 slideshows. 50 accounts. the leverage comes from the Co-work pipeline handling the work that previously required multiple human operators, with the single operator's role shifted from production execution to quality review and strategic direction.
## What This System Produces at Scale
at 30 to 50 slideshows per day across 50 accounts, the operation is producing 1,000+ slideshows per month. each one is a save opportunity. each one is a potential DM share. each one is a chance at the algorithmic push that turns a single slideshow into millions of views.
the saves compound because TikTok uses save rate as a primary distribution signal, and a slideshow that earns saves continues distributing for days and weeks beyond the initial posting window. the DM shares compound by extending each slideshow's reach into networks the original distribution never touched. the comment engagement compounds through the algorithm's increased weighting of substantive comment activity.
at the volume the system runs, the compounding effect across 1,000+ monthly slideshows is the difference between an operation generating modest views and an operation generating the kind of consistent organic reach that produces 7-figure annual revenue without paid ads.
the funnel architecture receives all of this. the link in bio routes to a quiz funnel rather than a product page. the comment automation captures high-intent viewers into DM sequences. the email and SMS nurture flows convert the warm leads who didn't purchase immediately.
each layer of the funnel converts a different segment of the traffic the slideshows generate, which means the reach numbers translate into revenue rather than disappearing through a product page that converts 2% and loses 98%.
## What This Means for People Still on TikTok Video Only

TikTok video is still working. obviously.
the point isn't that slideshows replace video.
it's that operators running only video content are missing the format the platform is actively pushing harder right now and that produces stronger save and share metrics for the same production investment.
the operations winning across TikTok in 2026 are running both. video for cold discovery and rapid follower growth on certain content categories. slideshows for the engagement, save, and DM share metrics that drive sustained distribution and warm audience activation.
building the slideshow layer on top of an existing TikTok video operation isn't a major lift. the same Claude Co-work creative intelligence brief drives both. the same multi-account portfolio distributes both. the same funnel architecture converts both.
and that's literally it.
P.S. - if you just want us to implement this entire ai ugc structure for your campaigns instead...
DM me "TIKTOK" on X (@adriansolarzz) and I'll show you exactly how we'd apply these principles to your specific offer.
- adrian
## 相关链接
- [Adrian Solarz](https://x.com/adriansolarzz)
- [@adriansolarzz](https://x.com/adriansolarzz)
- [14K](https://x.com/adriansolarzz/status/2049934006633021536/analytics)
- [@adriansolarzz](https://x.com/@adriansolarzz)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [3:28 AM · May 1, 2026](https://x.com/adriansolarzz/status/2049934006633021536)
- [14.5K Views](https://x.com/adriansolarzz/status/2049934006633021536/analytics)
- [View quotes](https://x.com/adriansolarzz/status/2049934006633021536/quotes)
---
*导出时间: 2026/5/1 21:41:23*
---
## 中文翻译
# 我们如何通过 Claude Co-work + GPT Image 2 自动化 TikTok 幻灯片实现流量爆发
**作者**: Adrian Solarz
**日期**: 2026-04-30T19:28:47.000Z
**来源**: [https://x.com/adriansolarzz/status/2049934006633021536](https://x.com/adriansolarzz/status/2049934006633021536)
---

2026 年,大多数人将所有精力都投入到了 TikTok 视频中,却忽略了该平台目前正大力推广的一种形式。
TikTok 幻灯片。
算法正将这种形式视为优先分发渠道,受众对滑动浏览式的互动方式产生了比大多数视频内容更强的收藏和分享指标,而且在同等质量下,其制作成本远低于视频制作。
因此,我们要自动化整个系统……
1. 通过 Claude Co-work 生成脚本和分析情报。
2. 通过 GPT Image 2 生成静态幻灯片视觉图。
3. 跨账号矩阵进行分发。
整个系统的制作量达到每个矩阵每天 30 到 50 个幻灯片,其工作流程
## 为什么 TikTok 幻灯片是“神级”玩法

2026 年,TikTok 的算法对不同内容形式的信号权重分配不同,而幻灯片形式目前正处于我们所说的“算法加速窗口”。
平台正积极试图推动这种形式的发展,因为幻灯片能产生对更广泛的平台生态系统有价值的互动行为:随着用户滑动浏览幻灯片,单次展示的停留时间更长;收藏率高于同等视频内容;以及在特定幻灯片上产生的实质性评论互动——算法对这种评论的权重要高于仅包含表情符号的视频反应。
滑动机制是这些更强指标的结构性原因。一个滑过了 6 张幻灯片的观众,其注意力在滑动过程中是主动留在你的内容上的,这与被动观看视频并在出现第一个疲软时刻时划走的观众有着本质区别。滑动迫使观众在每一张幻灯片处做出一个决定,这产生的主动互动正是算法将其解读为具有真正价值的证据。
除了互动质量的差异外,目前算法对粉丝量较少的账号发布的幻灯片的分发力度,要远高于同等视频内容。
这就是加速窗口。
现在抓住这一机会的运营者,将在未来几个月内受益于这种不成比例的分发优势,直到这种形式饱和且算法推力恢复正常。
制作经济学使得这个加速窗口更容易被捕获。在同等质量下,制作一个幻灯片所需的时间只是视频短片的一小部分,这意味着运营者可以以算法推力窗口所奖励的体量进行生产,而不会受到限制视频输出的制作能力约束。
## 真正有效的幻灯片形式
在构建任何生产系统之前,形式选择必须正确。一个执行良好的生产流水线,如果选错了形式,只会产生表现不佳的高质量内容,因为形式本身决定了内容结构是否与算法奖励的东西相一致。
我们在客户活动中测试了数十种幻灯片结构,其中有一组特定的形式能够持续产生收藏、分享和实质性评论互动,从而推动持续的分发。
**个人故事形式**:以承诺特定个人结果的钩子幻灯片开头(例如“我在 11 周内减掉了 23 磅,方法就是这个”),通过 4 到 5 张中间幻灯片讲述发现的过程和机制,最后以促使行动的 CTA(行动号召)幻灯片结束。这种形式效果极佳,因为钩子制造的好奇心差迫使观众滑动整个序列,以解决“结果是如何实现的”这一疑问。
**排行榜形式**:以倒序呈现 5 到 7 个项目,最强的一项在最后一张幻灯片揭晓。“第一名是什么”的好奇心差拉动观众看完整个幻灯片,产生的滑动速度正是算法解读为强互动的信号。这种形式对于受众对榜首项目已有预期的内容类别特别有效。
**顿悟形式**:以一个具体的洞察时刻开头(例如“我意识到为什么我的皮肤一直没好转”),并利用随后的每一张幻灯片通过具体细节加深这种顿悟。这种形式能产生极高的收藏率,因为观众希望稍后在将其应用于自身情况时重温这种顿悟。
**争议观点形式**:以一个与该细分领域常规思维相矛盾的陈述开头,并利用随后的幻灯片用具体证据来捍卫该立场。这种形式产生的评论互动是我们所有形式中最高的,因为观众要么强烈同意(并想添加支持),要么强烈反对(并想要争论该立场)。
**流程拆解形式**:在 5 到 6 张幻灯片中展示一个特定的流程或步骤,每张幻灯片专注于 1 个步骤。这些内容能获得高收藏率,因为该形式本质上具有参考价值。观众将其作为指南保存下来,以便在尝试自己实施该流程时回看。
这些形式已经存在于我们的幻灯片库中,这意味着系统不会为每个作品生成新策略。它是用新内容执行经过验证的结构。正是这种运营效率,使得每天 30 到 50 个幻灯片的产量成为可能,且质量不会下降。
## Claude Co-work 如何生成脚本

每个幻灯片的脚本和视觉配对说明都是通过加载了创意简报的 Claude Co-work 生成的。
该简报包括人物设定文档、竞争格局研究、独特机制阐述和转化库。具体针对幻灯片系统,简报还补充了特定形式的指导:每种幻灯片形式的结构模板、幻灯片数量参数,以及将每张幻灯片的文本连接到特定 GPT Image 2 生成提示词的视觉配对逻辑。
**个人故事幻灯片的主提示词:**
“利用上方加载的人物设定文档和创意简报,为 TikTok 写一个个人故事幻灯片脚本。总共 6 到 7 张幻灯片。第 1 张是开场钩子句,承诺带有时间线的特定个人结果。第 2 张到第 5 张讲述发现和机制,每张都是一段简短的文字叠加,并配有一个特定的视觉描述。最后一张是低摩擦的 CTA(行动号召)。对于每张幻灯片,请在括号中包含一个 GPT Image 2 生成提示词,描述要生成的确切静态图像,包括光线方向、环境背景、构图和情绪。视觉效果应在所有幻灯片中保持一致的美学风格:温暖的室内灯光、真实的手机自拍感、休闲的生活环境、非影棚级质量。”
输出是一份包含嵌入式制作提示词的完整幻灯片脚本,大约耗时 60 秒。相同的提示词模板适用于所有 5 种形式,只需调整各自的结构参数。
**成倍增加产出的变体生成**
一旦某个幻灯片在矩阵中表现强劲,Co-work 会在单次会话中基于这单一爆款生成 5 到 6 个变体,而不是从头开始制作 5 到 6 个全新的作品。
**变体提示词:**
“根据表现数据文件夹中的表现数据 [粘贴脚本],以下幻灯片表现良好。为此编写 5 个变体。每个变体应保持形式结构、情感基调和整体叙事弧线,但改变具体的个人故事细节和视觉配对说明。每个变体应针对更广泛人物设定中略微不同的细分群体。保留结束语结构和幻灯片数量。”
算法看到的跨账号差异足以使其不被视为重复内容,而且使原作品获得分发的潜在表现模式会延续到变体中,因为驱动互动的结构和情感要素保持完整。
## GPT Image 2 如何生成幻灯片视觉图

幻灯片是静态图像,这意味着视觉生成工具需要生成高质量的静态图像而不是视频片段。GPT Image 2 是该特定工作流程的正确工具,因为它解决了那些导致 AI 生成的幻灯片视觉图在第一秒就无法通过自然内容测试的质量问题。
**为什么专门选择 GPT Image 2 而非其他图像模型**
2026 年,大多数 AI 图像生成模型仍在努力解决两个特定的失败模式,这对规模化制作幻灯片来说是灾难性的。
第一个失败是产品文字渲染。对于需要视觉展示产品的幻灯片,包装上的文字必须清晰易读且准确。大多数图像模型在产品包装上生成的文字乱码或不正确,这在注意到不一致的观众眼中会立即被识别为 AI。
GPT Image 2 能准确渲染产品文字,这意味着幻灯片中的产品参考镜头显示的是正确的品牌名称和标签内容,而不是 AI 生成的近似替代品。
第二个失败是跨多张幻灯片序列的一致性。一个幻灯片需要在 6 到 7 张独立生成的图像之间保持美学、光线和视觉风格的连贯性。大多数图像模型能生成强大的单张图像,但无法在整个幻灯片集中保持视觉连贯性,而正是这种连贯性让幻灯片感觉像一个单一的有意图的作品。GPT Image 2 的指令遵循能力通过在每个幻灯片提示词中包含一致的美学基线来解决这个问题。
## 技术输出规格
每张幻灯片都按照 2026 年 TikTok 幻灯片的正确规格生成。
纵横比为 9:16 竖屏,具体为 1080x1920 像素。这与 TikTok 的原生竖屏格式相匹配,确保幻灯片在信息流中正确显示,不会因裁剪而丢失关键的视觉元素。
文字叠加位置遵循一个特定原则:画面的上三分之一或中心是幻灯片文字通常放置的位置,因为当观众滑动到新幻灯片时,注意力是从顶部进入画面的。叠加在下三分之一的文字有在不同显示配置的设备上被切断的风险,而且观众的眼睛在看幻灯片内容时不像看轮播图那样自然地首先落在那里。
**让 GPT Image 2 输出看起来自然的提示词原则**
场景描述优于静态描述,这是区分能通过自然内容测试的 GPT Image 2 输出和被识别为 AI 生成的基础原则。
“20 多岁女性拿着手机的镜子自拍特写,柔和温暖的卧室灯光,宽松的 T 恤,略显凌乱的晨发,真实的手机照片感,非影棚级质量,9:16 纵横比”产生的输出看起来就像某人在卧室里拍的真实自拍。“一个正在自拍的女士”会产生通成的图库风格图像,立即无法通过真实性测试。
一致的美学指令在幻灯片的所有幻灯片中保持了视觉连贯性。同一个幻灯片中每张幻灯片的每个 GPT Image 2 提示词都应包含相同的美学基线:“温暖的室内灯光、真实的手机自拍或抓拍生活感、柔和的自然构图、有机非影棚级质量、日常休闲环境。”正是这条指令让组装好的幻灯片感觉像是一个单一连贯的视觉叙事,而不是一系列独立生成的图像。
皮肤纹理规范,对于任何包含人脸的幻灯片。“真实的皮肤纹理、鼻梁和脸颊周围可见的毛孔、自然的轻微不均匀、无滤镜质感”,每条提示词无一例外都要包含这一点。AI 图像模型默认生成的光滑无孔皮肤是让观众迅速识别出幻灯片是合成内容的最快途径。
反精致语言用有机质量替换生成的美学,从而注册为真实可信。“休闲的手机构图、在真实环境中拍摄、非专业布景、有机非影棚级质量”将输出从那种会让冷受众观众产生广告怀疑的精致美学中拉开距离。
光线作为一个专门的从句,而不是埋没在其他场景元素的列表中。“来自左侧窗户的柔和卧室光线,投射自然阴影,黄金时刻质感,无 harsh highlights,皮肤光线适当不过曝”,作为每个包含人物主体的提示词中的独立元素。在图像生成中,光线控制情绪的程度超过任何其他变量。
**在选择之前生成 6 到 8 个变体**
对于包含人物主体的幻灯片,在选择最终图像之前,应从同一提示词生成 6 到 8 个变体,而不是接受第一个输出并继续。
从单个 GPT Image 2 提示词生成的变体之间的质量差异通常很大。最弱的变体可能存在皮肤纹理问题、表情问题或构图缺陷,会立即被识别为 AI 生成。最强的变体在头 2 秒内就能通过真实感测试,真正与真实的手机自拍无法区分。
从 6 到 8 个变体的池中进行选择,每张幻灯片虽然多花了一点时间,但产生的成品幻灯片质量水平是接受第一个输出的方法无法比拟的。
## 跨账号矩阵的分发架构

分发层是制作量实际转化为触达的地方。如果没有多账号矩阵,单账号每天 30 到 50 个幻灯片将面临结构性分发上限,无论内容质量如何,任何单个 TikTok 账号都会受到此限制。
我们运行由 50 个账号组成的客户矩阵作为起点,并根据 CPM 预算和细分领域特征进行扩展。每个账号每天发布 1 到 3 个幻灯片,这处于 TikTok 算法最佳发布频率区间。
**热门音频匹配作为分发倍增器**
TikTok 幻灯片可以利用热门音频作为分发倍增器,这是 Instagram 轮播图做不到的。配对了当前处于算法推力窗口中的音频的幻灯片,比配对非热门音频的同一幻灯片分发得更快、更广。
音频研究工作流程在每批生产开始时运行。我们确定目标细分领域当前哪些音频正在热门,评估哪些适合该批次生产中幻灯片的情感基调,并在排期前为每个幻灯片分配合适的音频。
每批次额外花费这 5 到 10 分钟的研究时间会产生显著的分发提升,因为算法使用音频契合度作为排名信号。配对与其内容类别相关的热门音频的幻灯片,始终优于配对随机音频的同一幻灯片,这使得音频匹配工作流程成为不可或缺的环节。
**账号预热架构**
每个进入矩阵的新账号在以全量投产部署之前,都要经过结构化的预热协议。跳过预热并立即高频发布的账号会被 TikTok 的检测系统识别,并且往往从一开始就会受到分发压制。
该协议在前 4 周遵循行为 ramps:资料完善和互动