I’ve seen too many comparisons using academic datasets—honestly, they’re not that useful for real client work. This platform’s approach makes more sense to me: instead of synthetic tests, they score based on actual jobs run on the platform. This season, they tested 8 models across over 12,000 inference tasks, evaluating from three angles: quality, cost per image, and p95 latency. All the tests are on real scenarios clients actually pay for—corporate avatars, e-commerce product shots, creative portraits, that kind of stuff.
Quality scoring is automated: FID measures distribution similarity to a high-quality reference set, CLIP checks how well the output matches the prompt and reference images, plus 500 pairs of human blind A/B tests. The composite score is 40% FID, 30% CLIP, and 30% human evaluation. For anyone picking a model for actual jobs, this kind of data—grounded in real workloads—is way more reliable than leaderboard chasing.