Just saw this post from HUST and Alibaba Tongyi — they’re working on making image gen models allocate compute more intelligently. Haven’t dug into the details yet, but the idea alone gets me talking. The most annoying thing about text-to-image right now is that no matter how simple or complex the scene is, it always runs the same number of steps.
A plain solid-color background image and a super detailed scene eat up pretty much the same compute — total waste. If they can really dynamically allocate computation by region or by difficulty, doing less for simple parts and more for complex ones, that could cut inference costs significantly and speed up generation too.
More and more domestic universities and big tech companies are teaming up on low-level optimizations, and honestly, this kind of work is way more meaningful than just stacking parameters. I’ll come back and update once the paper or technical details drop.