I’ve been testing Gemini 2.5 for e-commerce product images for the past 3-4 days. It’s really good at understanding complex prompts—you can dump the product selling points, composition, and atmosphere all at once, and it actually gets it, which is way smoother than a lot of pure text-to-image models.
But when it comes to the nitty-gritty details, like making sure the product logo isn’t distorted or the material reflections look realistic, it starts to mess up. Out of ten generations, maybe one is usable, so the ratio is still pretty low.
It’s fine for concept art or brainstorming ideas, but it’s too early to use it for final deliverables. Has anyone tried it for large-scale real-shot replacements? I wanna know how stable it actually is.
I’ve tried it too, the instruction following is really strong, but the final output still feels off.
Logo distortion is pretty much unavoidable with any text-to-image model.
When I had a bigger batch, I tried it for a week. Stability was so-so, maybe 2 or 3 out of 10 turned out decent.
Can’t get the material reflections right, basically useless for making digital product shots.
It’s unbeatable for brainstorming ideas, but when it comes to delivery, you gotta fall back on SD to cover your ass.
I’ve used it as a replacement for flat-lay clothing shots, but I still wouldn’t dare to use it directly for wrinkle details.
Honestly, when you factor in the cost of mass production with a 10-to-1 yield rate, it’s not really saving you anything.
OP’s right, the real cost is the manpower for picking images and redoing fixes.
I’ll give it that—it can handle complex compositions. Just not very consistent.