Been messing around with Kling 3.0 for a few days now, and honestly, you gotta treat it like a director to get the vibe right.

Back in the day, writing video prompts was just keyword stuffing: a person, a kitchen, nighttime, cool lighting. But after getting my hands on Kling 3.0, I completely changed my approach—it actually understands scene direction, not just a list of objects.

The biggest thing is it natively supports multi-shot output, up to six storyboard frames. Now I take the time to label each shot clearly: specify the shot type, subject, and movement for every single one, instead of cramming everything into one block. Terms like profile shot, macro close-up, tracking, POV, and shot-reverse shot—it actually gets them, automatically switching camera angles while keeping everything coherent.

Also, you gotta anchor your characters right from the start: unique names, consistent descriptions, no pronouns. For dialogue scenes, I write the action first, then the line—like “the black-suited agent slams his hand on the table,” then his yell—so the model knows who’s speaking. Don’t be vague with camera moves either: lock in follow, stop, pan, push—that’s how you get stable long takes. For the native audio, spell out who’s speaking, when, and what tone, so lip-sync and emotion match up. When you’re writing, think more about what the audience needs to see and feel—that beats just listing visual attributes any day.

The fact that it actually understands over-the-shoulder shots is insane, that’s a huge game-changer. Before, models were totally reliant on stitching things together.

Wait, “action first, then dialogue” — I never paid attention to that detail. No wonder my dialogue scenes always end up with the wrong faces saying the wrong lines. Gonna go fix that when I get back.

Yo anyone got a solid multi-shot prompt example? Been trying to piece one together but keep messing up the framing. Drop one if you got it, thanks!

Six mirrors at once is definitely convenient, but I’ve noticed that when you have too many lenses, the subject consistency starts to drift—especially noticeable on longer generations.

Hey, I’m just getting into this whole AI art thing. I’ve been using Midjourney for a bit, but I’m thinking about switching over to SD. I’ve got a 3060 with 12G VRAM, should be enough to run SDXL, right? I’ve heard ComfyUI is pretty good, but I’m not sure if I should start with that or just go with the regular SD WebUI. Also, I’ve seen people talking about Flux, but I have no idea what that is. Any tips for a noob? Thanks!

+1 to the above. If I have more than 4 shots, I usually just generate them separately and stitch them together. AI’s own timing can be really weird sometimes.

That “vibe tag” thing is seriously useful. Just slap “voice cracking” in there and the mood hits instantly.