Just saw 36Kr’s news—Shengshu Tech officially launched the Vidu S1 real-time interactive model. What caught my eye is that it’s all about real-time interaction: you can have a live video call and use voice to steer the video direction. This is a whole different ball game from the old “type a prompt and wait for it to render slowly” approach.
Specs-wise, they claim it supports 540P (960x540), 25FPS, with a max of 42FPS. The initial avatar can be a real person, anime, or even a cute pet, and you can customize the voice—basically, you can quickly whip up a personalized interactive character.
I make short videos, and my first thought is that if this thing can really sync voice in real-time, it’d save a ton of effort for virtual streamers and interactive characters. But I’m wondering about the actual latency and whether the lip-sync is on point. Anyone here already tried it out?