The whole audio-video generation thing has had a bigger impact on talking-head videos than I initially thought. Before, you’d have to render the visuals first and then go back and sync the lip movements, and that middle step was both slow and prone to messing up. Now the model spits out both audio and visuals at once, so the pipeline gets way shorter. Minimax’s H3 has a version that can run locally, and the short-drama folks are already jumping on it—looks like they’re gearing up for mass production.
Talking-head content now basically has four paths: no face shown; lip-sync models, split between local and cloud; dedicated digital human models; and then straight-up video generation models. The first three are pretty mature, but the fourth has always been stuck on character consistency. Now that audio and video come out together, I’m planning to fill in that last gap and see if the character can hold up across multiple segments.