Text-to-video models reward prompts that read like a shot list. A few habits that consistently help:
- Lead with the subject and action: "a red fox trotting through fresh snow" beats "winter scene".
- Name the camera move: slow dolly-in, handheld follow, aerial pull-back, locked-off wide.
- Set the light and time of day: golden hour backlight, overcast soft light, neon at night.
- Describe the sound if the model generates audio: wind, footsteps crunching, distant traffic.
- Keep one idea per clip and stitch longer stories from several 5–10 second shots.
I have been testing these on Kling 4.0, a free online AI video generator that turns text into cinematic videos with native 4K and synced audio: https://kling4.org
A useful exercise is to write the same scene three ways (wide, medium, close-up) and compare how the model keeps the subject consistent between them.