Yo, that Civitai guide on video generation prompts? Let me cut through the fluff and give you the real stuff.

Video prompts and image prompts are totally different beasts—you gotta learn how to “speak the language of motion.” I skimmed through that official guide on Civitai and jotted down what’s useful for me.

The general rule is 60 to 100 words—don’t make it too short or too wordy. Just lay out what’s happening in the scene, how the camera moves, what the lighting’s like, and the overall vibe.

For Wan2.1, this is the most practical part. It handles pull-back shots really well. The trick is to describe the scene first, then the camera movement, and finally what gets revealed after the camera pulls back. For example, start with a close-up of an Arctic explorer’s face, then pull back to show them alone in a blizzard, then pull back again to just endless ice. For tracking shots, make sure to use “camera follows” and be clear about what the subject is doing. Fast whip pans? It doesn’t really support those.

For Hunyuan, Tencent gave a formula: the simplest is subject + scene + action. If you want to get fancier, add camera language, atmosphere, and style on top. For multiple shots, use “Camera Switch” to cut between them, and connect two actions with “then.”

For lighting, just remember these: hard light gives high contrast, backlight creates silhouettes, and volumetric light is those beams in fog.

Honestly, the whole “pull-back” thing — starting with the scene and then adding motion — actually works pretty well. I’ve tried it a few times and the results are way more natural.

Yeah, quick panning shots aren’t supported. Tried messing around with it for ages, totally wasted my time.

Hunyuan’s formula is pretty easy to memorize.

60 to 100 words? No wonder my gens come out blurry—I always write way too short prompts.

As soon as you throw in that volumetric lighting description, the image quality instantly jumps up a notch.

Then I linked two actions together, pretty productive actually.