The whole "benchmark" thing for AI image gen that everyone's hyped about? Honestly, it's probably just "polished misdirection."

Honestly, the whole “benchmark” thing in AI image generation right now feels like a beautifully crafted scam. I’ve been wanting to rant about this for a while.

Every time a new model drops, it’s all “I’m #1 on some leaderboard” or “my win rate crushes the competition by X%.” But when you actually go and generate images yourself, the top-ranked ones aren’t always the most user-friendly. Things like aesthetic quality, how well it follows your prompt, and whether the fingers look like a horror show—these gut-feel aspects are totally invisible in the rankings. Some evaluations are even done by models judging each other or a tiny sample of human raters, which is super easy to game with a little tweaking.

So, do you guys pick models based on the leaderboards, or do you just run your own side-by-side comparisons with the same prompt? Personally, I trust the rankings less and less. I’d rather build my own test set. What do you all think about this whole evaluation system mess?

Honestly, these ranking lists are a joke — the people making them and the ones topping them are usually the same circlejerk.

I only trust my own benchmark of 20 fixed prompts.

There’s no way to quantify aesthetics, but it’s literally the thing that decides whether your generated images are usable or not.

Manual scoring with small samples is way too sketchy—swap out the annotators and the results completely flip.

Picking models based on leaderboards is just asking for trouble. Better to do your own head-to-head with a fixed set of 20 prompts—at least then you’ll know exactly where it falls apart.