Image to video
Animating a still image — the controllable way to generate video, how to write a motion prompt, and which starting frames animate well.
Image-to-video animates a picture you supply. Because the first frame is already decided, you control the look completely and ask the model only for motion. For most production work this is the mode that actually delivers.
Why it beats text-to-video for real work
Text-to-video asks the model to solve two problems at once: what the scene looks like, and how it moves. Image-to-video removes the first. That means:
- Your product, person or layout appears, because you put it there.
- Iteration is cheap. Fix the frame with image generations, which cost a fraction of video runs.
- Brand consistency is achievable, since every clip starts from an approved still.
- Fewer wasted generations, because you are only judging motion.
Which stills animate well
| Animates well | Animates badly |
|---|---|
| Implied motion already present — mid-stride, pouring, wind in fabric | Perfectly static studio shots with nothing to extend |
| Clear depth and separation between subject and background | Flat, busy compositions with no depth cue |
| Subject with margin around it | Subject cropped tight to the frame edge |
| Simple, readable scenes | Many small objects that can drift or merge |
| Soft even light | Hard complex shadow patterns |
Writing the motion prompt
Describe change over time, not the contents of the frame — the model can already see the frame, and re-describing it invites reinterpretation of something already correct.
- Subject motion — one action, with a speed.
- Camera — static, slow push, gentle orbit. One only.
- Secondary motion — hair, steam, fabric, foliage. This is what makes a still feel alive rather than warped.
A repeatable workflow
- Generate or shoot the frame and get it exactly right.
- Write a motion prompt describing one action plus one camera move.
- Generate short. Judge stability first, then whether the motion suits.
- Re-roll motion on the same approved frame rather than changing the frame.
- Cut several short clips together for anything longer than a few seconds.
Kling is the model most associated with strong image-to-video fidelity.
Common questions
Is image-to-video better than text-to-video?
More controllable, which for production work usually means better. Text-to-video is for exploring; image-to-video is for delivering something specific.
Why does my image barely move?
The starting frame has no implied motion. Stills with movement already suggested — mid-action, wind, flowing liquid — give the model something to extend.
Can I control where the camera goes?
To a degree. Name one camera move and keep it simple; stacking moves in a short clip produces incoherent results.