Shengshu Tech just dropped the Vidu S1 real-time interaction model, it can do live video calls + voice control to change directions.

Just saw 36Kr’s news—Shengshu Tech officially launched the Vidu S1 real-time interactive model. What caught my eye is that it’s all about real-time interaction: you can have a live video call and use voice to steer the video direction. This is a whole different ball game from the old “type a prompt and wait for it to render slowly” approach.

Specs-wise, they claim it supports 540P (960x540), 25FPS, with a max of 42FPS. The initial avatar can be a real person, anime, or even a cute pet, and you can customize the voice—basically, you can quickly whip up a personalized interactive character.

I make short videos, and my first thought is that if this thing can really sync voice in real-time, it’d save a ton of effort for virtual streamers and interactive characters. But I’m wondering about the actual latency and whether the lip-sync is on point. Anyone here already tried it out?

The whole voice-controlled video direction thing is actually pretty fresh, haven’t seen anyone pull it off in real-time before.

540P at 25FPS isn’t high by today’s standards, but the real win is that it’s real-time. I get the trade-off.

Even cute pet avatars can be used as interactive characters—first thing that came to my mind was a pet account.

It all comes down to latency — if there’s even half a second of lag in a real-time call, the experience is completely ruined.

Shengshu’s been moving fast lately, the video space is getting super competitive.

If vtubers could do real-time voice-driven generation, the cost would drop a ton.

The 42 FPS thing is probably only under specific conditions, realistically it’s gonna be around 25 in daily use.