In real projects, testing image generation tools makes leaderboards pretty much useless, because the leaderboard has no idea what kind of output you need to deliver. My low-tech method is to prep my own set of “benchmark images”: a typical e-commerce product shot for the company, a poster that must include Chinese text, a series requiring character consistency, and a complex scene with hands and text.
Every time a new tool drops, I run these fixed prompts through it—same prompts, same requirements. Whichever tool messes up the least on the stuff I actually do most, and gets fixed faster, that’s your tool. No one else’s benchmarks can replace the stress test of your own workflow.
I also made a version of that medical exam set. Fixed four images and ran them through a new tool, kept the ones that didn’t mess up as much.
You gotta use the Chinese one to really see if it works — so many models that claim they can do text fall apart the second you give them a long sentence.
The character consistency series is the best way to filter people out. The difference between those who can lock seeds and those who can’t is night and day.
Hands + text in complex scenes? Still barely any models can handle that.
Other people’s benchmarks don’t mean shit until you stress-test it with your own workload. That should be pinned.
I’ll also throw in a style reference image the client specifically asked for—generic leaderboards can’t even touch that level of detail.
Totally agree, running the same prompt and same requirements through each one is the only fair way to compare.
The quality of revisions matters more than making the first image look amazing—you really feel this when you’re rushing to meet a deadline.
One more thing: generation speed. When you’re doing a big batch, even one extra second per image is money.