There’s this image that looks super simple — a person lying horizontally, going through a tire swing over a lake, body parallel to the water, forming a mirror image with their reflection. No matter how I tweaked it, I couldn’t get it right, until I used Ideogram 4 and set the position with a bbox.
Later I tried the same JSON on Krea 2, and it kept failing at first. After messing around for ages, I realized I was screwing myself over — I was using it on Tensor Art, and its prompt enhancer was quietly rewriting my prompt. Switched to running the exact same JSON locally, and it worked immediately.
So sometimes it’s not the model that’s the problem, the platform is messing with things behind the scenes. When composition requires precise spatial positioning, explicit coordinate control like bbox is way more reliable than pure text descriptions.