Honestly, the whole “benchmark” thing in AI image generation right now feels like a beautifully crafted scam. I’ve been wanting to rant about this for a while.
Every time a new model drops, it’s all “I’m #1 on some leaderboard” or “my win rate crushes the competition by X%.” But when you actually go and generate images yourself, the top-ranked ones aren’t always the most user-friendly. Things like aesthetic quality, how well it follows your prompt, and whether the fingers look like a horror show—these gut-feel aspects are totally invisible in the rankings. Some evaluations are even done by models judging each other or a tiny sample of human raters, which is super easy to game with a little tweaking.
So, do you guys pick models based on the leaderboards, or do you just run your own side-by-side comparisons with the same prompt? Personally, I trust the rankings less and less. I’d rather build my own test set. What do you all think about this whole evaluation system mess?