You can go straight from a reference image to an animated video—this ComfyUI workflow saves so much hassle.

Back in the day, doing storyboard-to-video meant having to manually keyframe and animate every single frame—super time-consuming, and easy to mess up with visible glitches. Lately I saw someone using LTX 2.3’s image-condition LoRA to build a new workflow, and the idea is pretty clever: first, use an image generator to create a reference sheet for a character or concept—basically drawing multiple perspective panels on one image—then feed that whole reference sheet into the LTX setup.

It directly uses that image as a condition to generate a coherent animation video, no need to tweak frame by frame. For us content planners, the biggest value is compressing a messy multi-step process into just inputting a single reference image—communication and iteration become way cleaner.

Plus, I heard it can run on 6G VRAM, so even people with low-end cards can get started. I haven’t tested it myself yet, but I’m wondering—how consistent is the character in the generated video? Does it start to mess up the face after a few seconds?

Consistency is key, the quality of the reference sheet directly determines how the final output turns out.

6G can run this, so nice for low-spec users.

Frame-by-frame masking is an absolute nightmare, this setup saves so much hassle.

Multi-panel reference images are a technique worth learning.

The face issues depend on your prompt and the quality of the reference image. Try it out yourself and see.