Spent two whole nights on this, but I finally got the InfiniteTalk Video2Video workflow running in ComfyUI. Time to share some real-world tips.
This thing is an audio-driven lip-sync model, and the biggest selling point is no video length limit—feed it an audio clip, and the lip movements track super accurately, plus the head and body motions are pretty natural. I used MultiTalk before, and the hands would constantly twitch or the face would mess up; InfiniteTalk is way more restrained in that department, way fewer distortions.
You need six core models to download. The two fp16 diffusion ones are the biggest VRAM hogs—Wan2_1-InfiniTetalk-Single and wan2.1_i2V_480p_14B both have to go in their respective folders. There’s also a Chinese wav2vec2 that auto-downloads, so no manual fuss.
VRAM-wise, I gotta warn you: the official minimum says 24G. My 4090 barely handles 480p, and the moment I try 720p it’s straight OOM. If you want higher res, you’re looking at 48G+. If you’re hitting VRAM limits, turn on low mem load and switch to the fp8 version—that’ll save your ass.
Prompts can be dead simple, but the input audio quality needs to be top-notch. First time I used a compressed mp3, and the lip sync came out blurry as hell.