Text-to-video is getting a ton of capital moves lately—Vidu just closed another funding round, and on the flip side, that Huanlema thing is confirmed to be Alibaba’s product. Money and big tech are all piling in, which means everyone’s betting this line won’t hit a ceiling anytime soon.
What I’m more curious about is whether the model iteration pace will speed up noticeably after the funding hits, 'cause in this field, it’s like a major version every few months—fall behind and you’re out. Being tied to a big tech name has its perks too, like unlimited compute and distribution, so they can sustain the long burn.
The real turning point probably isn’t about who launches first, but who can nail generation stability and cost down to a scalable level.