Just saw someone say the best open-source text-to-image model has a new king overnight — the original SD team made it.

Looks like the best open-source text-to-image model just got dethroned. This new one’s made by the original Stable Diffusion team, and word is they’re dropping a SOTA video generation model next.

I’d heard rumors they went solo a while back. The open-source text-to-image scene’s always been a battlefield anyway—“best model” changes hands every few weeks. What I’m really curious about is the video side. If they actually deliver something solid, it could be a game-changer for people making motion assets. But for now it’s just talk—no concrete specs or release dates yet. Gonna keep an eye on it. Anyone got more intel? Drop it here.

The original team still has that pull, worth waiting to see what they cook up.

The title of “best open-source model” changes like five times a year lol.

Video generation is the real tough nut to crack. Image generation has been a dead end for a while now.

Let’s see the actual results first. Anyone can make a PPT.

Waiting for an update on this. As someone who makes motion assets, I really need this.