Ran a head-to-head comparison on real client portrait orders, pitting Flux against SDXL and some closed-source models.

I’ve seen too many comparisons using academic datasets—honestly, they’re not that useful for real client work. This platform’s approach makes more sense to me: instead of synthetic tests, they score based on actual jobs run on the platform. This season, they tested 8 models across over 12,000 inference tasks, evaluating from three angles: quality, cost per image, and p95 latency. All the tests are on real scenarios clients actually pay for—corporate avatars, e-commerce product shots, creative portraits, that kind of stuff.

Quality scoring is automated: FID measures distribution similarity to a high-quality reference set, CLIP checks how well the output matches the prompt and reference images, plus 500 pairs of human blind A/B tests. The composite score is 40% FID, 30% CLIP, and 30% human evaluation. For anyone picking a model for actual jobs, this kind of data—grounded in real workloads—is way more reliable than leaderboard chasing.

Using real client orders to test this direction is the right call. Benchmark-chasing data is meaningless.

The client cares most about p95 latency.

A blind test with 500 pairs of samples should be enough to get a good idea.

Can you share the actual cost per image?

For corporate profile pics, closed-source is still the safer bet.

IMO setting the FID weight to 40% is a bit too high.

Yeah, putting 40% weight on FID is way too much. Clients don’t even look at that metric when judging whether a portrait looks good or not.