# I Studied Chinese AI Video Creators. They Are Not Prompting - They Are Building Factories.
**作者**: 0xTria
**日期**: 2026-06-18T07:12:57.000Z
**来源**: [https://x.com/0xTria/status/2067891751575019639](https://x.com/0xTria/status/2067891751575019639)
---

Holy shit. Chinese AI creators look like they are already living in 2028.
While most AI-video tutorials celebrate one decent six-second clip, creators on Douyin are connecting models into production lines that turn one product photo, one human performance, and one reusable workflow into dozens of ads. They replace products, clothes, actors, voices, languages, and locations without reshooting everything.
You cannot steal proprietary models or someone else’s identity. But you can copy the operating system: references, control sheets, storyboards, motion templates, constraints, automation, and batch production. That system can already be sold as e-commerce creatives, AI UGC, fashion videos, localized spokesperson ads, product demos, or virtual-influencer content.

## The pattern I kept seeing
The workflow is not:
```
Write a clever prompt → pray → regenerate 14 times.
```
It is:
```
Reference
→ Control artifact
→ Generation
→ Editing
→ Batch scaling
```
A reference supplies the real information: product, character, clothing, pose, or movement.
A control artifact converts it into something the next model can follow: a 3×3 storyboard, product feature sheet, character turnaround, clothing sheet, or motion reference.
The model receives less freedom and more direction. It knows what may change and what must stay fixed. Chinese creators are not asking AI to invent everything at once. They divide the job into stages and give each stage one responsibility.

## The Chinese AI factory stack
Tongyi Wanxiang
> Digital humans, character replacement, motion imitation, image and video generation
Kling AI
> Video generation and editing; product, outfit, and character replacement
GPT Image 2
> Storyboards, product sheets, character sheets, visual planning, and editing
Seedance 2.0
> Multimodal video generation from text, image, audio, and video references
Wan2.2
> Open video generation, image-to-video, and character animation workflows
ComfyUI
> Node-based assembly line for reusable workflows and batches |
MiniMax
> Multimodal coding, agents, video, audio, and code-based animation |
Claude Code + Skills
> Tool building, workflow logic, files, and reusable instructions |
Supabase
> Database, authentication, storage, and state for creator tools |
The point is not that one model wins. Each occupies a station. Kling edits footage. Wanxiang and Wan2.2 transfer characters or performance. GPT Image 2 creates planning sheets. Seedance turns references into a sequence. ComfyUI makes the process repeatable.
## Factory #1: one motion template becomes 100 product ads
Film a simple UGC template. A person holds a neutral object, points toward it, turns it toward the camera, or demonstrates a routine. Keep the background clean, lighting stable, and hands visible.
Upload a clean product reference and ask Kling O1 or a similar editing model to replace only the object:
```
Replace the object in the person’s hand with the product from the reference image.
Preserve the original hand position, finger grip, body motion, facial expression,
camera angle, lighting, shadows, and background.
Preserve the product’s exact shape, proportions, color, label, logo, and material.
Match the scale and perspective so it looks naturally held.
Maintain temporal consistency across the full clip.
```
The phrase **“replace only”** matters. So does the preservation list. Without it, the model may redesign the package, move the fingers, change the face, or alter the room.
Once the template works:
```
1 motion template
× 10 products
× 5 opening hooks
× 3 visual styles
= 150 creatives
```
Not all 150 will be good. The value is testing a wide creative surface without organizing 150 shoots.
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2067505830312686032)
>
> The $500 product shoot is turning into one prompt.
> A girl holds a dirty rag.
> Kling turns it into an iPhone, a cleanser tube, toothpaste, a cup, whatever product image you upload.
> Then it swaps her outfit and places her in a different scene while keeping the same motion, pose,
>
> 
The same logic works for clothing. Provide a clean garment reference and preserve the person, motion, camera, and environment. Add close references for fabric, collar, cuffs, print, and fit.
## Factory #2: storyboard first, video second
Most failed AI commercials begin with “make a premium cinematic ad.”
The Chinese workflow often begins with a visual contract: a 3×3 storyboard containing nine planned shots. It defines camera movement, product placement, lighting, mood, transitions, sound cues, and the final hero frame.
```
You are a world-class commercial director and visual planner.
Using the uploaded product reference, design a coherent 15-second advertisement
as a 3×3 storyboard with nine sequential panels.
For every panel specify:
- shot type and composition
- camera movement
- product position
- lighting and atmosphere
- key action
- transition
- sound-design note
Keep the product shape, logo, color, material, scale, and identity consistent.
Use one color palette and camera language.
End with a clean premium hero shot.
```
Generate the storyboard with GPT Image 2 or another image model, select the best version, and pass it into Seedance 2.0 or another reference-capable video model.
Now the video model does not have to invent the campaign, shot list, art direction, and product behavior simultaneously.
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2066398992284422265)
>
> ONE GUY CAN NOW CHARGE $3,000-$5,000 FOR PRODUCT ADS MADE IN ONE AI TAB
> The workflow is stupid simple:
> 1. Drop in a clean product photo
> 2. Ask AI for a 15-sec ad script
> 3. Generate a 3x3 storyboard with 9 shots
> 4. Pick the best frame
> 5. Paste the script into the video block
> 6.
>
> 
## Factory #3: products and characters get visual memory
A storyboard is not enough when the item occupies a small part of each frame. Logos blur. Faces drift. Clothes mutate. Packaging changes shape.
The fix is a **feature sheet**.
For a product, create front, side, back, macro texture, logo, material, and package views. For a person, create front, side, back, full-body, half-body, facial close-up, and neutral poses. For clothing, include fit, fabric, seams, cuffs, collar, and pattern.
```
Create a production-ready product reference sheet from the uploaded image.
Show front, side, back, three-quarter view, logo close-up, texture close-up,
material details, packaging geometry, and important functional features.
Use a neutral studio background.
Do not redesign, beautify, simplify, or invent details.
Preserve proportions, colors, typography, logo placement, and materials.
```
The prompt explains the task; the sheet protects identity.
This matters for AI influencers too. A useful virtual influencer is not one portrait. It is a small operating system:
```
persona.md
voice.md
visual_identity.md
content_rules.md
brand_rules.md
memory.md
```
The face, voice, topics, sponsor boundaries, wardrobe, and history should remain stable.
## Factory #4: motion is separated from identity
For fashion and creator content, motion is a reusable asset.
Start with a short reference clip: a simple dance, turn, product demonstration, reaction, or talking-head performance. Combine it with a character sheet and clothing sheet using Wan2.2 Animate, Tongyi Wanxiang, Kling, or a ComfyUI workflow.
```
Generate a video in which the character follows the movement and timing
of the reference video.
Preserve the face, hairstyle, body proportions, outfit, and identity.
Preserve the rhythm, pose transitions, gestures, and camera timing.
Keep the clothing design, fabric, color, and pattern unchanged.
Avoid face occlusion, limb distortion, texture flicker, and background drift.
```
Simple movement beats complex choreography. Five clean seconds are more reusable than twenty chaotic seconds with crossed arms, spinning hair, and constant occlusion.
One motion source can produce multiple models, outfits, languages, and campaign versions—provided you have permission to use the performance and are not fabricating endorsements.
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2066445283102249293)
>
> The Chinese guy who made $8,240 with Google Gemini is still making money
> Now he posts videos where AI turns him into a beautiful girl in revealing clothes
> At first, people thought it was just another Douyin seller trying to go viral
> But nobody realized the girl was not real
>
> 
> 
## ComfyUI is where content becomes a factory
A generation made manually in five browser tabs is still a trick. A saved ComfyUI graph is a production system.
The graph can ingest a product image, create a feature sheet, generate storyboard candidates, pass references into a video model, upscale the result, export structured files, and repeat across a product folder.
Expose variables such as product, hook, duration, aspect ratio, style, language, and CTA. Then batch them as a matrix:
```
/products × /hooks × /styles × /languages
```
This is the boring—and valuable—part. Production businesses run on naming, versioning, QA, and repeatability, not only impressive generations.
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2067105631266361428)
>
> A Chinese creator found a way to sell $3,000-$5,000 brand videos with no team, no camera crew, and no agency.
> no film crew.
> no editor.
> no storyboard artist.
> just one ComfyUI canvas turning product photos into finished AI ads.
> > GPT-image2 builds the director board: shots,
>
> 
## The coding layer nobody talks about
Not every animation should be rendered by a video model.
MiniMax and Claude Code can help turn references into SVG, CSS, Canvas, browser extensions, simple games, and interactive product experiences. Claude Skills can store repeatable animation or marketing rules. Supabase can hold users, results, storage, and state.
```
Use anticipation, squash and stretch, a clean motion arc,
ease-in/ease-out timing, follow-through, and subtle secondary motion.
Implement it in lightweight CSS and JavaScript.
```
Code-based motion is editable and reproducible—often a better fit for interfaces, explainers, or product pages.
## How to make money without selling AI slop
Do not sell “AI videos.” Sell a defined outcome:
- **Product Variation Pack:** 15–30 ads from one approved template.
- **Localized Spokesperson Pack:** one campaign adapted to several languages.
- **Fashion Motion Pack:** turnarounds, fabric close-ups, and short motion clips.
- **Storyboard-to-Ad Service:** product sheet, nine-shot plan, finished video, and exports.
- **Monthly Testing Engine:** new hooks and variations delivered every week.
The pitch is not “I use a secret Chinese model.”
It is:
> “Give me one approved product, one brand guide, and one licensed performance. I will deliver a controlled batch of testable creatives.”
Start with one niche and one template. Validate three outputs before generating thirty. Human-review logos, labels, anatomy, claims, synchronization, and brand compliance.
mybf.io ai boyfreind app
Features, prices, and regional availability change. Some tools require specific regions, accounts, or paid plans. Follow platform rules, advertising law, likeness rights, and consent requirements. Never use character replacement to impersonate someone or fake an endorsement.
## The seven-step playbook
```
1. Choose one repeatable commercial task.
2. Capture or license a clean motion template.
3. Build product, character, or clothing feature sheets.
4. Create a storyboard before multi-shot video.
5. Write strict preservation and negative constraints.
6. Validate a small batch with human QA.
7. Save the workflow, parameterize it, and scale.
```
That is the technology worth copying.
The advantage is no longer knowing the “best prompt.” Prompts are easy to copy and quickly become obsolete.
The durable advantage is owning a repeatable production system:
```
reference → control layer → generation → QA → batch scaling
```
Chinese AI creators are not waiting for one perfect model. They are combining imperfect models into factories—and those factories are already good enough to sell outcomes today.
## 相关链接
- [0xTria](https://x.com/0xTria)
- [@0xTria](https://x.com/0xTria)
- [12K](https://x.com/0xTria/status/2067891751575019639/analytics)
- [Tongyi Wanxiang](https://tongyi.aliyun.com/wanxiang/)
- [Kling AI](https://app.klingai.com/global/)
- [GPT Image 2](https://developers.openai.com/api/docs/models/gpt-image-2)
- [Seedance 2.0](https://seed.bytedance.com/en/seedance2_0)
- [Wan2.2](https://github.com/Wan-Video/Wan2.2)
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
- [MiniMax](https://www.minimax.io/)
- [Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview)
- [Skills](https://docs.anthropic.com/en/docs/claude-code/skills)
- [Supabase](https://supabase.com/docs)
- [Jun 18](https://x.com/0xTria/status/2067505830312686032)
- [6.4K](https://x.com/0xTria/status/2067505830312686032/analytics)
- [Jun 15](https://x.com/0xTria/status/2066398992284422265)
- [9.1K](https://x.com/0xTria/status/2066398992284422265/analytics)
- [Jun 15](https://x.com/0xTria/status/2066445283102249293)
- [52K](https://x.com/0xTria/status/2066445283102249293/analytics)
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
- [Jun 17](https://x.com/0xTria/status/2067105631266361428)
- [15K](https://x.com/0xTria/status/2067105631266361428/analytics)
- [mybf.io](https://mybf.io/)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [4:46 PM · Jun 19, 2026](https://x.com/0xTria/status/2067891751575019639)
- [12.4K Views](https://x.com/0xTria/status/2067891751575019639/analytics)
---
*导出时间: 2026/6/20 11:35:56*
---
## 中文翻译
# 我研究了中国的AI视频创作者。他们不是在写提示词——而是在建工厂。
**作者**: 0xTria
**日期**: 2026-06-18T07:12:57.000Z
**来源**: [https://x.com/0xTria/status/2067891751575019639](https://x.com/0xTria/status/2067891751575019639)
---

天哪。中国的AI创作者看起来仿佛已经生活在2028年。
当大多数AI视频教程还在为一个不错的6秒片段而欢呼时,抖音上的创作者们正在将模型连接成生产线:一张产品图、一段真人表演、一个可复用的工作流,就能转化出几十条广告。他们无需全部重拍,就能替换产品、服装、演员、声音、语言和场景。
你无法窃取专有模型或他人的身份。但你可以复制那个操作系统:参考图、控制表、分镜板、动作模板、约束条件、自动化和批量生产。这套系统本身已经可以作为电商创意素材、AI UGC(用户生成内容)、时尚视频、本地化代言人广告、产品演示或虚拟网红内容出售。

## 我反复看到的模式
他们的工作流不是这样的:
```
写一句机智的提示词 → 祈祷 → 重新生成14次。
```
而是这样的:
```
参考图
→ 控制组件
→ 生成
→ 剪辑
→ 批量规模化
```
参考图提供真实信息:产品、角色、服装、姿势或动作。
控制组件将其转化为下一个模型可以遵循的形式:一张3×3的分镜图、产品特征表、角色转稿图、服装表或动作参考图。
模型获得的自由度更少,得到的指导更多。它知道什么可以改变,什么必须保持不变。中国的创作者并不是要求AI一次性发明一切。他们将任务拆分为多个阶段,并赋予每个阶段单一职责。

## 中国AI工厂的技术栈
通义万相
> 数字人、角色替换、动作模仿、图像与视频生成
可灵 AI
> 视频生成与编辑;产品、服装和角色替换
GPT Image 2
> 分镜图、产品表、角色表、视觉规划和剪辑
Seedance 2.0
> 基于文本、图像、音频和视频参考的多模态视频生成
Wan2.2
> 开源视频生成、图生视频和角色动画工作流
ComfyUI
> 用于可复用工作流和批处理的基于节点的流水线
MiniMax
> 多模态编码、智能体、视频、音频和基于代码的动画
Claude Code + Skills
> 工具构建、工作流逻辑、文件和可复用指令
Supabase
> 面向创作者工具的数据库、认证、存储和状态管理
关键不在于哪个模型胜出。每个模型都占据一个工位。可灵负责剪辑素材。通义和Wan2.2负责传递角色或表演。GPT Image 2制作规划表。Seedance将参考图转化为序列。ComfyUI让整个流程可重复。
## 工厂 #1:一个动作模板变成100条产品广告
拍摄一个简单的UGC(用户生成内容)模板。一个人手持一个中性物品,指向它,将其转向镜头,或演示一套动作。保持背景干净、光线稳定、双手可见。
上传一张干净的产品参考图,并指示Kling O1或类似的编辑模型仅替换该物品:
```
将人手中的物品替换为参考图中的产品。
保留原始的手部位置、手指抓握方式、肢体动作、面部表情、
摄像机角度、光线、阴影和背景。
保留产品的精确形状、比例、颜色、标签、标志和材质。
匹配缩放和透视,使其看起来握持自然。
在整个片段中保持时间一致性。
```
**“仅替换”**这个词很重要。保留清单同样重要。没有这些,模型可能会重新设计包装、移动手指、改变面部表情或修改房间环境。
一旦模板跑通:
```
1个动作模板
× 10个产品
× 5个开场钩子
× 3种视觉风格
= 150条创意素材
```
并非这150条都会很完美。其价值在于无需组织150次拍摄,就能测试广阔的创意面。
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2067505830312686032)
>
> 价值500美元的产品拍摄正在变成一个提示词。
> 一个女孩拿着一块脏抹布。
> Kling把它变成了iPhone、洁面管、牙膏、杯子,或者你上传的任何产品图。
> 然后它替换她的服装,将她置于不同的场景中,同时保持相同的动作、姿势,
>
> 
同样的逻辑也适用于服装。提供一张干净的服装参考图,并保留人物、动作、摄像机和环境。为面料、领口、袖口、印花和版型添加特写参考。
## 工厂 #2:先分镜,后视频
大多数失败的AI商业广告始于“制作一个高端电影感广告”。
中国的流程通常始于一份视觉合同:一张包含9个预定镜头的3×3分镜图。它定义了摄像机运动、产品摆放、光线、氛围、转场、声音提示和最终的英雄镜头(hero frame)。
```
你是一位世界级的商业广告导演和视觉规划师。
利用上传的产品参考图,设计一段连贯的15秒广告,
以3×3分镜图的形式呈现,包含九个连续的画面。
对每个画面指定:
- 镜头类型和构图
- 摄像机运动
- 产品位置
- 光线和氛围
- 关键动作
- 转场
- 声音设计提示
保持产品的形状、标志、颜色、材质、比例和身份一致。
使用统一的色调和摄影语言。
以一个干净的高端英雄镜头结束。
```
使用GPT Image 2或其他图像模型生成分镜图,选择最佳版本,并将其输入到Seedance 2.0或其他支持参考的视频模型中。
此时,视频模型无需同时发明广告战役、镜头清单、艺术指导和产品表现。
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2066398992284422265)
>
> 一个人现在就可以凭借在一个AI标签页制作的产品广告收费3000-5000美元
> 工作流程简直蠢得简单:
> 1. 放入一张干净的产品照片
> 2. 向AI索要一个15秒的广告脚本
> 3. 生成一个包含9个镜头的3x3分镜图
> 4. 选出最好的画面
> 5. 将脚本粘贴到视频块中
> 6.
>
> 
## 工厂 #3:产品和角色拥有视觉记忆
当物品在每一帧中只占据很小一部分时,仅靠分镜图是不够的。标志会模糊。脸部会走样。衣服会变异。包装会变形。
解决方案是**特征表**。
对于产品,制作正面、侧面、背面、微距纹理、标志、材质和包装视图。对于人物,制作正面、侧面、背面、全身、半身、面部特写和中性姿势视图。对于服装,包括版型、面料、接缝、袖口、领口和图案。
```
从上传的图像创建一张生产级的产品参考表。
展示正面、侧面、背面、四分之三视图、标志特写、纹理特写、
材质细节、包装几何形状和重要功能特征。
使用中性的工作室背景。
不要重新设计、美化、简化或虚构细节。
保留比例、颜色、排版、标志位置和材质。
```
提示词解释任务,而参考表负责保护身份一致性。
这对AI网红也同样重要。一个有用的虚拟网红不仅仅是一张肖像。它是一个小型的操作系统:
```
persona.md (人设)
voice.md (声音)
visual_identity.md (视觉识别)
content_rules.md (内容规则)
brand_rules.md (品牌规则)
memory.md (记忆)
```
脸、声音、话题、赞助边界、衣橱和历史记录都应保持稳定。
## 工厂 #4:动作与身份分离
对于时尚和创作者内容,动作是一种可复用的资产。
从一段简短的参考片段开始:一段简单的舞蹈、转身、产品演示、反应或对着镜头说话的表演。使用Wan2.2 Animate、通义万相、Kling或ComfyUI工作流,将其与角色表和服装表结合。
```
生成一段视频,让角色遵循参考视频的动作和时机。
保留面部、发型、身体比例、服装和身份。
保留节奏、姿势过渡、手势和摄像机时机。
保持服装设计、面料、颜色和图案不变。
避免面部遮挡、肢体扭曲、纹理闪烁和背景漂移。
```
简单的动作胜过复杂的编舞。五秒干净的片段比二十秒充满双臂交叉、头发乱飞和频繁遮挡的混乱片段更具复用价值。
一个动作源可以产出多个模特、服装、语言和战役版本——前提是你拥有使用该表演的许可,并且没有伪造代言。
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2066445283102249293)
>
> 那个靠Google Gemini赚了8240美元的中国人还在赚钱
> 现在他发布视频,AI把他变成穿着暴露衣服的漂亮女孩
> 起初,人们以为这只是另一个试图走红的抖音卖家
> 但没人意识到那个女孩不是真的
>
> 
> 
## ComfyUI 是内容变为工厂的地方
在五个浏览器标签页里手动生成的生成仍然只是个戏法。一个保存好的ComfyUI图表则是一个生产系统。
该图表可以摄入产品图像,创建特征表,生成分镜候选项,将参考图传递给视频模型,放大结果,导出结构化文件,并针对产品文件夹重复上述过程。
暴露变量,如产品、钩子、时长、宽高比、风格、语言和号召性用语(CTA)。然后将它们作为矩阵进行批处理:
```
/产品 × /钩子 × /风格 × /语言
```
这是枯燥但有价值的一环。生产业务依赖于命名、版本管理、质量检查(QA)和可重复性,而不仅仅依靠令人印象深刻的生成效果。
> **0xTria@0xTria**: [原文链接](https://x.com/0xTria/status/2067105631266361428)
>
> 一位中国创作者找到了一种出售3000-5000美元品牌视频的方法,没有团队,没有摄制组,也没有代理商。
> 没有摄制组。
> 没有剪辑师。
> 没有分镜画师。
> 只有一个ComfyUI画布,将产品照片转化为成品AI广告。
> > GPT-image2构建导演板:镜头,
>
> 
## 没人谈论的代码层
并非所有的动画都应该由视频模型渲染。
MiniMax和Claude Code可以帮助将参考图转化为SVG、CSS、Canvas、浏览器扩展、简单游戏和交互式产品体验。Claude Skills可以存储可复用的动画或营销规则。Supabase可以保存用户、结果、存储和状态。
```
使用预判、挤压与拉伸、清晰的运动弧线、
缓入/缓出时机、跟随动作和微妙的次要运动。
使用轻量级的CSS和JavaScript来实现它。
```
基于代码的动作是可编辑且可复制的——对于界面、解释器或产品页面来说,这通常是一个更好的选择。
## 如何在不兜售AI垃圾的情况下赚钱
不要卖“AI视频”。卖一个定义明确的结果:
- **产品变体包:** 基于一个获批模板生成15-30条广告。
- **本地化代言人包:** 将一个战役适配为多种语言。
- **时尚动作包:** 转身图、面料特写和简短动作片段。
- **分镜转广告服务:** 产品表、九镜头方案、成品视频和导出文件。
- **月度测试引擎:** 每周交付新的钩子和变体。
推销辞令不是“我使用了一个秘密的中国模型”。
而是:
> “给我一个获批的产品、一本品牌指南和一个授权表演。我将交付一批可控的、可测试的创意素材。”
从一个细分领域和一个模板开始。在生成三十个之前,先验证三个输出。人工审核标志、标签、解剖结构、宣称、同步性和品牌合规性。
mybf.io ai boyfreind app
功能、价格和地区可用性可能会有变化。某些工具需要特定地区、账户或付费计划。请遵守平台规则、广告法、肖像权和同意要求。切勿使用角色替换功能来冒充他人或伪造代言。
## 七步操作手册
```
1. 选择一项可重复的商业任务。
2. 拍摄或授权一个干净的动作模板。
3. 构建产品、角色或服装特征表。
4. 在制作多镜头视频之前先创建分镜图。
5. 编写严格的保留指令和负面约束。
6. 通过人工质量检查(QA)验证小批量样本。
7. 保存工作流,将其参数化,并进行规模化扩展。
```
这才是值得复制的技术。
优势不再在于知道“最好的提示词”。提示词很容易复制,而且很快就会过时。
持久的优势在于拥有一个可重复的生产系统:
```
参考图 → 控制层 → 生成 → 质检 → 批量扩展
```
中国的AI创作者并没有等待一个完美的模型。他们正在将不完美的模型组合成工厂——而这些工厂现在已经足够好,可以出售成果了。