MAGREF lets you control who shows up in the video using your own photos, and the multi-character consistency is actually pretty solid.

MAGREF’s selling point is: you throw in a few reference images, and the people/objects in the video stay exactly like that—no random faces popping up. Full name is Masked Guidance for Any-Reference, it locks identity using region masks plus channel concatenation.

It has three modes: single ID keeps one character throughout, multi-ID puts several different reference subjects in the same frame while keeping their own traits, and ID+object+background lets you build complex scenes. I tried multi-ID, and multiple characters on screen didn’t mix up faces at all.

A few pitfalls to note: don’t use blurry or pixelated input images; try to keep brightness and saturation consistent across all images, or the output colors will look messy; resolution only supports 480p and 720p; keep frame count under 81, or it’ll crash. Prompts don’t need to be too detailed—just nouns plus simple action words are enough. Sampling steps 4 to 6, CFG at 1 to 2, and if you’re not happy with the result, just tweak the prompt and rerun.

The whole “multiple IDs but no face blending” thing is a game changer. Before, if you tried to generate more than one face, it’d just fall apart.

Kinda annoying that the frame count is capped at 81.

You gotta keep the brightness and saturation consistent — I learned that the hard way, ended up with a mess of a output.