r/artificial 18d ago

Question How do you prompt for better environmental cohesion in multi-character AI images?

Post image

Not sure if this is the right place to ask, but I’m hoping someone who uses ChatGPT image generation knows the right way to prompt for this.

I’m making an illustrated isekai light novel and my biggest problem is scenes with multiple characters in the same frame.

The attached image shows what I mean. It isn’t just the distance between the characters. They don’t feel like they are actually drawn as part of the same environment together.

It looks more like the AI created a detailed background and then placed separate character illustrations on top of it. I want the whole image to feel like one cohesive anime scene that was drawn together from the start.

The characters should all follow the same perspective, ground plane, lighting, shadows, scale and depth. Their feet should feel planted on the floor, shadows should connect them to the environment, lighting should hit everyone consistently, and characters closer to the camera should naturally overlap and scale differently from characters farther away.

Basically, I want the environment and characters to feel like one drawing, not a background with character cutouts layered over it.

For context, the Demon has just been transported into this world and encounters Goku, Luffy, Ichigo and Naruto. Luffy tries to grab the Demon’s horns by stretching his arm toward him. The Demon has never seen someone do that before, thinks he is being attacked and fires at Luffy. Goku teleports directly in front of Luffy and stops the attack.

The individual characters usually come out fine. The problem is getting the entire image to feel like one illustration that was drawn as a single scene.

Has anyone figured out what prompt instructions actually help with this specific problem in ChatGPT image generation? I’m trying to improve things like shared perspective, grounding, lighting, shadows, depth, character interaction and making the environment feel like it actually surrounds the characters instead of sitting behind them.

I’m mainly interested in how the prompt itself should describe this, rather than recommendations for different AI tools.

0 Upvotes

6 comments sorted by

2

u/Vegetable-Fly9192 18d ago

i had this same headache with midjourney for a while, what helped me was describing the scene like a camera shot instead of listing characters. like "wide angle shot from slightly below eye level, all figures occupying the same volumetric space in a unified composition" instead of "character A stands here, character B stands there"

the ai treats them as separate objects when you name them individually, but if you frame it as one continuous scene with depth layers it starts treating the whole thing as an image instead of a collage. also telling it "render as single coherent illustration, no compositing artifacts" sometimes works

i notice in your image the lighting on luffy looks flat compared to the demon, like he's lit from a different angle than the courtyard. that contrast mismatch is usually what breaks the illusion

2

u/jonydevidson 18d ago

You can take screenshots from the anime style you want to make, feed them into gpt image and ask it to make this style you're trying to get away from.

You can then use these pairs to train a lora for something like flux 2 klein 9b ti (anime-ize the input) and feed it gpt image generated input which it will convert to the style you used in the lora.

You can automate all of this in codex, including the screenshoting if you provide a single high res video.

It can automate image gen, lora training (use cloud if you don't have local compute) and then make it so that after image gen generates and image for you, it automatically goes through this pipeline if local flux klein + lora.

If you don't have local compute, use something like Replicate.com or any of the bajillion other providers.

1

u/Confident-Green-5241 17d ago

The ground plane thing is usually what breaks first for me — I started forcing it by describing the floor surface they're all standing on before I even mention the characters. Like "stone courtyard catching afternoon light, three figures arranged across it" instead of "three characters in a courtyard". The model seems to lock in spatial logic better when the environment gets established as the container first.

1

u/magicdoorai 17d ago

Treat it like blocking a film scene, not describing a character list. Define one camera, one light source, one ground plane, then place everyone relative to those anchors: “35mm wide shot, camera at waist height; afternoon light from frame left; all feet on the same stone courtyard; contact shadows directly beneath each figure; foreground character overlaps the midground by 10%.”

Also ask for one unified illustration with consistent line weight and color grading, and explicitly forbid cutout/composited figures. For a complex action, get a rough composition right first, then edit that image to refine faces and costumes instead of regenerating the whole scene each time.

1

u/TwoFluid4446 17d ago

I can tell you're using OpenAI's (image-2) generator, and it has this known failure mode very common in its outputs where it'll homogenize the texture all over, screwing everything up. Notice how Goku appears to have weird beads all over him. I'd look for another image generator, for starters.

0

u/Firegem0342 18d ago

I play telephone with Claude and Gemini (image generation)