From Sora to Kling, it feels like video AI still hasn't had its GPT moment yet.

From Sora to Kling, it feels like video AI still hasn’t had its “GPT moment.” I work in concept design, so I mainly use video models to bring static setups to life and check the vibe. Honestly, right now it’s more about “wow, that’s impressive” than “yeah, I can actually control this.”

By “GPT moment,” I mean that game-changing leap where anyone can casually use it and get consistent, usable results without endless gacha-style re-rolling. Current video models still have that heavy gacha feel—consistency, physics logic, and long shots are all still lacking. The direction’s definitely right, but we haven’t hit that tipping point yet. What do you guys think is the missing piece for that GPT moment—controllability or duration?

“Gacha-style re-rolling” — that’s the perfect way to put it, so damn accurate.

I’m betting on controllability. Consistency is just too hard to nail down.

Long shot consistency is still a total pain in the ass.

Stunning results, but barely any control—one sentence sums up the current state.

Physics and logic are constantly messing up, characters clipping through stuff.

Not quite there yet, but heading in the right direction—same feeling here.

The real issue isn’t getting a god-tier result every now and then — it’s being able to consistently reproduce it.