From Sora to Kling, it feels like video AI still hasn’t had its “GPT moment.” I work in concept design, so I mainly use video models to bring static setups to life and check the vibe. Honestly, right now it’s more about “wow, that’s impressive” than “yeah, I can actually control this.”
By “GPT moment,” I mean that game-changing leap where anyone can casually use it and get consistent, usable results without endless gacha-style re-rolling. Current video models still have that heavy gacha feel—consistency, physics logic, and long shots are all still lacking. The direction’s definitely right, but we haven’t hit that tipping point yet. What do you guys think is the missing piece for that GPT moment—controllability or duration?