Looks like the best open-source text-to-image model just got dethroned. This new one’s made by the original Stable Diffusion team, and word is they’re dropping a SOTA video generation model next.
I’d heard rumors they went solo a while back. The open-source text-to-image scene’s always been a battlefield anyway—“best model” changes hands every few weeks. What I’m really curious about is the video side. If they actually deliver something solid, it could be a game-changer for people making motion assets. But for now it’s just talk—no concrete specs or release dates yet. Gonna keep an eye on it. Anyone got more intel? Drop it here.