Text to video
Generating video from a written prompt alone — what it is good for, the coherence limits every model shares, and how to write a shot description.
Text-to-video produces a clip from a description. Every frame is invented, which makes it the most creatively open mode and the least controllable one.
What it is genuinely good for
- Exploring an idea before committing to a specific look.
- Generic or abstract footage — weather, textures, crowds, landscapes, b-roll where no particular thing must appear.
- Concept work where the point is the feeling rather than the specifics.
It is a poor fit whenever a particular product, person or layout must appear. Description cannot summon a specific thing — for that, fix a frame and use image to video.
Write it as a shot, not a story
A few seconds holds one action. The most common failure is a prompt describing a sequence of events, which the model compresses into incoherence rather than choosing one. A workable structure:
- Shot size and subject — "Close on a woman's hands kneading dough".
- The single action — what changes during the clip.
- Camera — static, slow push in, gentle orbit. One move only.
- Lighting — this carries most of the perceived production value.
- Setting — enough to place it, no more.
Limits shared by every model
| Limit | What you see | Work around it by |
|---|---|---|
| Duration coherence | Quality decays over the clip | Generating short, cutting several together |
| Hands and manipulation | Fingers merge; grip fails | Framing to avoid close hand work |
| Readable text | Signage warps frame to frame | Adding text in post |
| Object permanence | Things drift, merge or vanish | Simpler scenes, fewer objects |
| Specific likeness | The person is not your person | Supplying a reference frame |
Iterating without wasting credits
Video generation costs far more per run than images. Before spending on video, settle the look as a still — one image generation costs a fraction of one video generation, and the frame you approve becomes the input for a controlled animation. Treating text-to-video as the last step rather than the first is the single biggest saving available.
See the video generator overview and the model reference for which model suits which shot.
Common questions
How long can a text-to-video clip be?
Seconds rather than minutes, with the cap set per model. Longer output degrades, so most work is several short clips edited together.
Why does my clip ignore half the prompt?
Too many competing instructions for the duration. A few seconds holds one action — cut the prompt to one subject and one movement.
Can I get a specific person or product in the shot?
Not reliably from text. Generate or supply the frame containing it, then animate that with image-to-video.