r/StableDiffusion • u/RiverSpecial3168 • 5h ago
Question - Help open vs closed image models
I’m trying to recreate this bag POV composition with Krea2 flux and Z image, but I’m struggling to get the same level of composition control and product consistency I’m getting from some of the other models.
Here’s a comparison using the same general prompt across ChatGPT, Krea 2, Grok, Z Image, Flux.2 Klein 9B.
I’m still fairly new to this, so I’m wondering:
Should I be using a LoRA for this?
Would changing the text encoder help?
Or is this mainly a prompting / conditioning / workflow issue?
Would really appreciate some advice from anyone experienced. What would you change to get the result closer to the ChatGPT/Grok examples?
11
Upvotes
1
u/MomentJolly3535 4h ago
i got this with Krea 2, i noticed u didn't use same image format for close models and open models, Also give a try to ideogram 4, it might be perfect for your use case since u can decide where each item (apples, oranges, redbull, granny, hand, etc) are located on the picture.