MAGREF’s selling point is: you throw in a few reference images, and the people/objects in the video stay exactly like that—no random faces popping up. Full name is Masked Guidance for Any-Reference, it locks identity using region masks plus channel concatenation.
It has three modes: single ID keeps one character throughout, multi-ID puts several different reference subjects in the same frame while keeping their own traits, and ID+object+background lets you build complex scenes. I tried multi-ID, and multiple characters on screen didn’t mix up faces at all.
A few pitfalls to note: don’t use blurry or pixelated input images; try to keep brightness and saturation consistent across all images, or the output colors will look messy; resolution only supports 480p and 720p; keep frame count under 81, or it’ll crash. Prompts don’t need to be too detailed—just nouns plus simple action words are enough. Sampling steps 4 to 6, CFG at 1 to 2, and if you’re not happy with the result, just tweak the prompt and rerun.