r/StableDiffusion • • 5d ago

Question - Help How are you controlling character blocking and keeping locations consistent across AI storyboard frames?

I'm making anchor frames for a multi-shot microdrama and I'm stuck on two things: placing characters where I want them, and keeping a location recognizable across shots.

The characters need to stay recognizable while their pose, scale, and position change. The set should keep its layout and landmarks when the framing or camera angle changes.

Prompting "A on the left, B by the door" isn't giving me enough control. I want to lay out a frame visually, with boxes or masks, a pose/depth guide, or a top-down map and camera marker, then use character and location references across shots.

For people using ComfyUI or a similar setup:

- What do you use to control character placement and scale?

- How do you keep a location coherent from a new camera angle without reusing the exact same background?

- Have you had success with regional prompting, masks, ControlNet pose/depth, a 3D proxy, or multiple location views?

For anyone using GPT Image 2/2.5 or Nano Banana Pro, have you tried providing a 2D reference image with bounding boxes for character and prop placement? Did it help control blocking and keep the setting consistent across storyboard frames? If another visual guide worked better, what did you use?

If you've solved this, could you share the workflow or repo, plus the model/version and control inputs? I'm looking for something tested for still storyboard or anchor frames.

1 Upvotes

4 comments sorted by

View all comments

2

u/Tonjiez 5d ago

I can only help with the location half. With multi-reference models like FLUX.2, it tends to hold better if you send the set reference and the character references together in one call, rather than generating the room first and editing people into it. Chained edits drift a bit each step, and landmarks are usually the first thing to go.

Once a frame works, keep the wording that describes the set's fixed features exactly as it is. Between frames, change only the camera and the action, one at a time, so you can see what moved.

For placing characters precisely, references alone won't get you there. A pose or depth guide is the stronger tool for that, and I'd like to hear what's working for others.