# Using Text-to-Video Models in Pre-Production
**作者**: LTX.io
**日期**: 2026-07-06T16:31:36.000Z
**来源**: [https://x.com/ltx_io/status/2074169403764084834](https://x.com/ltx_io/status/2074169403764084834)
---

Traditional storyboarding is slow. A single sequence of 20-30 boards can take a storyboard artist days to complete, and every revision cycle resets the clock. For independent filmmakers, advertising teams, and studios working under tight pre-production timelines, this bottleneck delays the entire production pipeline.
Text-to-video models change this. Instead of hand-drawing boards or waiting for an artist, you can generate rough video clips from script descriptions, assemble them into an animatic, and share moving pre-visualization with your team in hours. This guide covers the practical workflow for using text-to-video AI in pre-production: storyboarding, animatic assembly, and reference generation, using LTX-2.3 as the reference model.
## Why Text-to-Video Models Belong in Pre-Production
Pre-production is about communication and iteration, not final quality. The goal is to get everyone — director, DP, production designer, VFX supervisor — looking at the same reference before expensive shoot days begin. Text-to-video models produce rough but directionally accurate motion clips from text descriptions, which is exactly what pre-production needs.
The key difference from image-based storyboarding: generated clips show motion, camera behavior, and temporal flow. A storyboard frame shows composition. A generated clip shows how the camera moves through that composition and what the subject does during the shot. For directors who think in motion, this is a qualitative improvement in the communication tool.
## Building a Storyboard Generation Workflow
Translating Script to Prompts
Script lines don't map directly to generation prompts. A line like "INT. KITCHEN - MORNING - Maria pours coffee" needs to become a prompt that specifies visual composition, motion, and camera behavior. The LTX-2.3 prompting guide recommends keeping prompts under 200 words and structuring them chronologically: scene description first, then subject action, then camera behavior.
For storyboard generation, add camera direction to every prompt. Don't leave camera behavior implicit. If the shot is a medium close-up with a slow push-in, say so: "Medium close-up on Maria's hands as she pours coffee. Camera pushes in slowly. Morning light from the window falls across the counter."
Shot Types and Camera Language
LTX-2.3 responds to standard cinematographic camera descriptors: dolly in, pan left, tracking shot, static camera, pan right, aerial, low angle, close-up, wide shot. For pre-production boards, include the shot type and camera movement in every prompt. This produces clips that communicate visual intention to a DP or director more clearly than static images.
Useful prompt patterns for storyboarding:
• Establishing shot: "Wide aerial shot of [location]. Camera descends slowly toward [subject]. [Description of environment and time of day]."
• Over-the-shoulder dialogue: "Over-the-shoulder shot from [character A]'s perspective looking at [character B]. [Character B] speaks. Static camera. [Lighting and setting description]."
• Action beat: "[Subject] [performs action]. Camera tracks alongside at [distance/angle]. Motion is [speed]. [Key visual element to emphasize]."
Hardware and Setup for Storyboard Generation
For storyboarding with LTX-2.3, use the TI2VidTwoStagesPipeline for the highest quality output when you have GPU resources, or the distilled pipeline for rapid iteration. The distilled pipeline's 8-step inference is well-suited to storyboard generation, where you're generating many clips quickly and evaluating composition and motion rather than final image quality.
Note: Running LTX-2.3 locally requires a Linux system with CUDA 13+ and an Nvidia GPU with 80GB+ VRAM for the default configuration (32GB with FP8 quantization). For teams without this hardware, the hosted API provides the same generation capability without local infrastructure.
## Shot Iteration and Selection
Generate multiple takes per shot by varying the seed parameter. The same prompt with different seeds produces different camera angles, motion timing, and compositional variations. For pre-production, generate 3-5 takes per shot and select the one that best communicates the intended visual idea.
Organize generated clips by scene and shot number. A practical file naming convention: [scene_number]_[shot_number]_[seed]_v[take].mp4. This keeps takes traceable and makes assembling the animatic straightforward.
For shots that need longer duration, use LTX-2.3's Retake pipeline to regenerate a specific time range within a clip. If a 4-second clip has a good first 2 seconds but the motion goes wrong in the second half, Retake can regenerate just that section rather than regenerating the full clip. If a shot needs to be longer, use LTX-2.3's Retake pipeline to regenerate a specific time segment while preserving what works.
## Building the Animatic
An animatic is a sequence of storyboard clips cut together with rough timing that approximates the rhythm of the final edit. With generated video clips, you're building a moving animatic rather than a still-frame version, which communicates pacing and transition timing more clearly.
Assembly Workflow
Any video editing tool handles animatic assembly: Premiere Pro, DaVinci Resolve, Final Cut Pro, or even a simple timeline tool. Import your selected takes and cut them together to script timing. At this stage, the goal is rhythm and flow, not color or quality.
LTX-2.3 video frame counts must satisfy (F-1) % 8 == 0, giving valid clip lengths of 9, 17, 25, 33, 41, 49, 57, 65, 73, 81, 89, and 97 frames. At 25 fps, 97 frames is approximately 3.88 seconds. Plan shot durations around these constraints during storyboard generation rather than trimming in the edit, which wastes compute and may cut the motion before it completes.
Adding Reference Audio
Rough dialogue or scratch VO can go in at the animatic stage. LTX-2.3's audio-to-video pipeline can generate clips that respond to a reference audio track, useful when you want the generated motion to reflect the rhythm of dialogue or music. For the animatic, rough timing with scratch audio is sufficient. The goal is communicating intent, not final picture.
## Reference Clip Generation for VFX and Production Design
Beyond storyboarding, text-to-video models produce reference clips useful throughout pre-production:
• VFX pre-visualization: Generate rough clips showing intended VFX sequences — camera movements through environments, dynamic weather, vehicle behavior. These aren't final; they communicate scale, motion, and timing to the VFX team. LTX-2.3's image-to-video pipeline (TI2VidTwoStagesPipeline) generates motion from a concept art frame, giving VFX supervisors moving reference that approximates the final output.
• Location scouting reference: Generate clips showing how a described location should look and feel in the edit. These help production designers align on the visual target before physical location scouts begin.
• Camera test reference: Generate clips showing the intended camera movement for a shot. A dolly move, a handheld walk-and-talk, a crane rise — generating these before the shoot day lets the DP and director discuss execution with moving reference rather than verbal description alone.
## Cinematic Prompts for Pre-Production Quality
For pre-production reference, match your prompt language to the shot type and visual intent:
• Use film-standard camera descriptors: dolly, pan, tilt, track, push-in, pull-out, crane, aerial, handheld, tic, static, pan right, pan left, tracking shot. LTX-2.3 responds to these cinematographic directions through prompt language and can produce reference clips that approximate the intended camera movement.
• Describe lighting conditions specifically: "golden hour backlight", "overcast flat light", "harsh midday sun", "interior practical light from window". These influence the generated clip's visual quality and communicability.
• Reference a visual style if it's relevant to communicating the intended aesthetic: "shallow depth of field", "high contrast", "desaturated", "film grain".
LTX-2.3's multiple pipelines let you move between text-to-video and image-to-video as needed during pre-production. Generate an initial text-to-video clip to establish the scene, then use image-to-video to generate motion from a selected frame of concept art for shots where a specific visual is already designed. Assemble them into an animatic for review.
## Conclusion
Text-to-video generation compresses the pre-production storyboarding cycle from days to hours. The output isn't final-quality — it's communicative quality, which is what pre-production requires. For the animatic, the clips show timing, camera intention, and rough motion. For VFX and production design reference, they show scale and movement. For director-DP communication, they replace verbal description with something both parties can look at.
The workflow is: script to prompts, generate multiple takes per shot, select and assemble into animatic, iterate on timing and motion with the team. LTX-2.3's open-source pipeline and hosted API both support this workflow — the pipeline for teams with GPU access who want full control, the API for teams that prefer not to manage local inference infrastructure.
## 相关链接
- [LTX.io](https://x.com/ltx_io)
- [@ltx_io](https://x.com/ltx_io)
- [5.7K](https://x.com/ltx_io/status/2074169403764084834/analytics)
- [VFX](https://ltx.io/model/capabilities/vfx)
- [LTX-2.3 prompting guide](https://ltx.io/model/model-blog/ltx-2-3-prompt-guide)
- [TI2VidTwoStagesPipeline](https://ltx.io/model/model-blog/ltx-2-image-to-video-text-to-video-workflow)
- [hosted API](https://docs.ltx.video/)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [12:31 AM · Jul 7, 2026](https://x.com/ltx_io/status/2074169403764084834)
- [5,741 Views](https://x.com/ltx_io/status/2074169403764084834/analytics)
---
*导出时间: 2026/7/7 09:11:50*
---
## 中文翻译
# 在前期制作中使用文生视频模型
**作者**: LTX.io
**日期**: 2026-07-06T16:31:36.000Z
**来源**: [https://x.com/ltx_io/status/2074169403764084834](https://x.com/ltx_io/status/2074169403764084834)
---

传统的分镜制作非常缓慢。仅仅一组20-30个镜头的分镜,分镜师可能需要数天才能完成,而且每次修改都会重置时间表。对于独立电影制作人、广告团队以及在紧张的前期制作时间表下工作的工作室来说,这一瓶颈延误了整个制作流程。
文生视频模型改变了这一点。你无需手绘分镜或等待画师,而是可以直接从剧本描述生成粗略的视频片段,将它们组装成动态分镜(animatic),并在几小时内与团队分享动态的前期预演。本指南介绍了在前期制作中使用文生视频AI的实用工作流程:分镜制作、动态分镜组装以及参考素材生成,并以LTX-2.3作为参考模型。
## 为什么文生视频模型属于前期制作
前期制作的核心在于沟通与迭代,而非最终质量。目标是在昂贵的拍摄日开始前,让所有人——导演、摄影指导(DP)、美术指导、视效总监——都基于同一个参考达成共识。文生视频模型可以根据文本描述生成粗糙但方向准确的动态片段,这正是前期制作所需要的关键能力。
与基于图像的分镜相比,主要区别在于:生成的片段展示了动作、镜头行为和时间流。分镜画面展示的是构图;而生成的片段展示的是摄像机如何在该构图中移动,以及主体在镜头期间做了什么。对于那些通过动态思考的导演来说,这是一种沟通工具质的飞跃。
## 构建分镜生成工作流程
将剧本转化为提示词
剧本台词无法直接映射为生成提示词。像“INT. KITCHEN - MORNING - Maria pours coffee”(内景. 厨房 - 早晨 - Maria倒咖啡)这样的台词,需要变成一个指定了视觉构图、动作和镜头行为的提示词。LTX-2.3的提示指南建议将提示词控制在200字以内,并按时间顺序构建:先是场景描述,然后是主体动作,最后是镜头行为。
对于分镜生成,请在每个提示词中加入镜头指示。不要让镜头行为含糊不清。如果这是一个中近景并伴有缓慢的推镜头,请明确说明:“Maria倒咖啡时双手的中近景。摄像机缓慢推入。晨光从窗户洒在台面上。”
镜头类型与镜头语言
LTX-2.3可以识别标准的电影摄影镜头描述词:推镜头、左摇、跟踪镜头、固定机位、右摇、航拍、低角度、特写、广角镜头。对于前期制作分镜,请在每个提示词中包含镜头类型和运镜方式。这生成的片段能比静态图像更清晰地向摄影指导或导演传达视觉意图。
有用的分镜提示模式:
• 建立镜头:“[地点]的广角航拍镜头。摄像机缓慢向[主体]下降。[环境和时段描述]。”
• 过肩对话:“从[角色A]视角看向[角色B]的过肩镜头。[角色B]正在说话。固定机位。[光影和布景描述]。”
• 动作节拍:“[主体] [执行动作]。摄像机以[距离/角度]跟随移动。动作速度为[速度]。[需强调的关键视觉元素]。”
分镜生成的硬件与设置
使用LTX-2.3制作分镜时,如果你有GPU资源,请使用TI2VidTwoStagesPipeline以获得最高质量的输出;或者使用蒸馏管道进行快速迭代。蒸馏管道的8步推理非常适合分镜生成,因为你需要快速生成大量片段并评估构图与动态,而非追求最终图像质量。
注意:在本地运行LTX-2.3需要Linux系统、CUDA 13+以及Nvidia GPU,默认配置需80GB+显存(使用FP8量化时需32GB)。对于没有此类硬件的团队,托管API提供相同的生成能力,无需本地基础设施。
## 镜头迭代与选择
通过改变随机种子参数为每个镜头生成多个版本。相同的提示词搭配不同的种子会产生不同的镜头角度、运动时机和构图变化。对于前期制作,每个镜头生成3-5个版本,并选出最能传达预期视觉构思的那一个。
按场景和镜头编号组织生成的片段。一种实用的文件命名约定:[场景号]_[镜头号]_[种子]_v[版本].mp4。这能保持版本的可追溯性,并使动态分镜的组装变得简单。
对于需要更长时长的镜头,请使用LTX-2.3的Retake管道重新生成片段内的特定时间范围。如果一个4秒片段的前2秒很好,但后半部分的动态出了问题,Retake可以只重新生成该部分,而无需重新生成整个片段。如果镜头需要更长,请使用LTX-2.3的Retake管道重新生成特定的时间片段,同时保留有效的部分。
## 组装动态分镜
动态分镜是将分镜片段按粗略时间剪辑在一起的序列,近似于最终剪辑的节奏。使用生成的视频片段,你构建的是一个动态的动态分镜,而非静态帧版本,这能更清晰地传达节奏和转场时机。
组装工作流程
任何视频编辑工具都可以处理动态分镜的组装:Premiere Pro、DaVinci Resolve、Final Cut Pro,甚至是简单的时间轴工具。导入你选中的版本,按剧本时间剪辑在一起。在此阶段,目标是节奏和流畅度,而非色彩或质量。
LTX-2.3的视频帧数必须满足 (F-1) % 8 == 0,有效的片段长度为9、17、25、33、41、49、57、65、73、81、89和97帧。在25 fps下,97帧约为3.88秒。在分镜生成阶段就根据这些限制规划镜头时长,而不要在剪辑中进行修剪,因为修剪不仅浪费算力,还可能在动作完成前将其切断。
添加参考音频
粗略的对白或临时配音(VO)可以在动态分镜阶段加入。LTX-2.3的音生视频管道可以响应参考音轨生成片段,当你希望生成的动态反映对白或音乐的节奏时,这非常有用。对于动态分镜,使用临时音频配合粗略时机就足够了。目标是传达意图,而非最终画面。
## 为VFX和美术设计生成参考片段
除了分镜制作,文生视频模型还能生成在整个前期制作中有用的参考片段:
• VFX预演:生成展示预期VFX序列的粗略片段——环境中的镜头运动、动态天气、车辆行为。这些不是最终成片,但它们能向VFX团队传达规模、动态和时机。LTX-2.3的图生视频管道(TI2VidTwoStagesPipeline)可以从概念艺术帧生成动态,为视效总监提供近似最终输出的动态参考。
• 堪景参考:生成展示描述地点在剪辑中应有的视觉和感觉的片段。这些有助于美术设计师在实地堪景开始前统一视觉目标。
• 拍摄测试参考:生成展示镜头预期运镜方式的片段。推车移动、手持边走边谈、摇臂上升——在拍摄日前生成这些内容,让摄影指导和导演能够通过动态参考讨论执行方案,而不仅仅是口头描述。
## 适合前期制作质量的电影感提示词
对于前期制作参考,请将你的提示语言与镜头类型和视觉意图相匹配:
• 使用电影标准的镜头描述词:dolly(推拉)、pan(摇)、tilt(俯仰)、track(跟踪)、push-in(推入)、pull-out(拉出)、crane(摇臂)、aerial(航拍)、handheld(手持)、tic(微震)、static(固定)、pan right(右摇)、pan left(左摇)、tracking shot(跟踪镜头)。LTX-2.3能够通过提示语言响应这些电影摄影指示,并生成近似预期运镜的参考片段。
• 具体描述光照条件:“golden hour backlight”(黄金时刻逆光)、“overcast flat light”(阴天平光)、“harsh midday sun”(正午强光)、“interior practical light from window”(来自窗户的内景实景光)。这些会影响生成片段的视觉质量和传达效果。
• 如果与传达预期美学相关,请参考某种视觉风格:“shallow depth of field”(浅景深)、“high contrast”(高对比度)、“desaturated”(低饱和度)、“film grain”(胶片颗粒)。
LTX-2.3的多种管道让你在前期制作期间能够根据需要在文生视频和图生视频之间灵活切换。先生成一个文生视频片段确立场景,然后对于已经设计了特定视觉的镜头,使用图生视频从选定的概念艺术帧生成动态。最后将它们组装成动态分镜进行审核。
## 结语
文生视频生成将前期制作分镜周期从数天压缩到了数小时。其输出并非最终质量——而是沟通质量,这正是前期制作的要求。对于动态分镜,片段展示了时机、镜头意图和粗略动态。对于VFX和美术设计参考,它们展示了规模和运动。对于导演与摄影指导的沟通,它们用双方都能直观观看的内容取代了口头描述。
工作流程是:剧本转提示词,为每个镜头生成多个版本,筛选并组装成动态分镜,与团队一起迭代时机和动态。LTX-2.3的开源管道和托管API都支持此工作流程——管道适合拥有GPU资源且希望完全掌控的团队,API适合不愿管理本地推理基础设施的团队。
## 相关链接
- [LTX.io](https://x.com/ltx_io)
- [@ltx_io](https://x.com/ltx_io)
- [5.7K](https://x.com/ltx_io/status/2074169403764084834/analytics)
- [VFX](https://ltx.io/model/capabilities/vfx)
- [LTX-2.3 提示指南](https://ltx.io/model/model-blog/ltx-2-3-prompt-guide)
- [TI2VidTwoStagesPipeline](https://ltx.io/model/model-blog/ltx-2-image-to-video-text-to-video-workflow)
- [托管 API](https://docs.ltx.video/)
- [升级到 Premium](https://x.com/i/premium_sign_up)
- [12:31 AM · Jul 7, 2026](https://x.com/ltx_io/status/2074169403764084834)
- [5,741 观看](https://x.com/ltx_io/status/2074169403764084834/analytics)
---
*导出时间: 2026/7/7 09:11:50*