Video prompts and image prompts are totally different beasts—you gotta learn how to “speak the language of motion.” I skimmed through that official guide on Civitai and jotted down what’s useful for me.
The general rule is 60 to 100 words—don’t make it too short or too wordy. Just lay out what’s happening in the scene, how the camera moves, what the lighting’s like, and the overall vibe.
For Wan2.1, this is the most practical part. It handles pull-back shots really well. The trick is to describe the scene first, then the camera movement, and finally what gets revealed after the camera pulls back. For example, start with a close-up of an Arctic explorer’s face, then pull back to show them alone in a blizzard, then pull back again to just endless ice. For tracking shots, make sure to use “camera follows” and be clear about what the subject is doing. Fast whip pans? It doesn’t really support those.
For Hunyuan, Tencent gave a formula: the simplest is subject + scene + action. If you want to get fancier, add camera language, atmosphere, and style on top. For multiple shots, use “Camera Switch” to cut between them, and connect two actions with “then.”
For lighting, just remember these: hard light gives high contrast, backlight creates silhouettes, and volumetric light is those beams in fog.