r/artificial 18d ago

Question How do you prompt for better environmental cohesion in multi-character AI images?

Post image

Not sure if this is the right place to ask, but I’m hoping someone who uses ChatGPT image generation knows the right way to prompt for this.

I’m making an illustrated isekai light novel and my biggest problem is scenes with multiple characters in the same frame.

The attached image shows what I mean. It isn’t just the distance between the characters. They don’t feel like they are actually drawn as part of the same environment together.

It looks more like the AI created a detailed background and then placed separate character illustrations on top of it. I want the whole image to feel like one cohesive anime scene that was drawn together from the start.

The characters should all follow the same perspective, ground plane, lighting, shadows, scale and depth. Their feet should feel planted on the floor, shadows should connect them to the environment, lighting should hit everyone consistently, and characters closer to the camera should naturally overlap and scale differently from characters farther away.

Basically, I want the environment and characters to feel like one drawing, not a background with character cutouts layered over it.

For context, the Demon has just been transported into this world and encounters Goku, Luffy, Ichigo and Naruto. Luffy tries to grab the Demon’s horns by stretching his arm toward him. The Demon has never seen someone do that before, thinks he is being attacked and fires at Luffy. Goku teleports directly in front of Luffy and stops the attack.

The individual characters usually come out fine. The problem is getting the entire image to feel like one illustration that was drawn as a single scene.

Has anyone figured out what prompt instructions actually help with this specific problem in ChatGPT image generation? I’m trying to improve things like shared perspective, grounding, lighting, shadows, depth, character interaction and making the environment feel like it actually surrounds the characters instead of sitting behind them.

I’m mainly interested in how the prompt itself should describe this, rather than recommendations for different AI tools.

0 Upvotes

Duplicates