GPT Image 2 just hit #1 globally for text-to-image, beating out the old champ Nano Banana 2. I couldn’t find which benchmark or scoring dimensions they used, so I’m not about to guess the numbers.
As someone who works on algorithms, I take these rankings with a grain of salt. Text-to-image benchmarks are super sensitive to the prompt set and evaluation criteria—swap the test method and the same model can drop several spots. Being #1 on some leaderboard doesn’t mean it crushes every scenario, especially since Chinese and English prompts often give totally different results.
What really matters is blind testing with your own stuff: run the same batch of prompts across a few models and compare. If you’ve got the hardware, just do a round yourself—way more useful than staring at rankings. So between these two, which one’s giving you better outputs?