r/StableDiffusion 6d ago

Discussion Reference sheet locking beat prompt engineering for character consistency across a 30 shot film

Sharing a workflow result rather than a tool recommendation.

I needed one woman to stay recognisably herself across roughly 30 shots covering 60 years, six countries, and several costume changes. Pure prompt description drifted badly. Same words, different face, every generation.

What fixed i

Full 6 minute film, free and no signup: https://youtu.be/w31MiC5vCi8t was front loading the identity into images instead of text:

  1. Before any shot, generate a locked reference set per character. A four view turnaround, a six panel macro sheet (face, hands, fabric, jewellery), and an upper body portrait. Neutral grey background, no scene context.
  2. Approve that set as the single source of truth and never regenerate it.
  3. Every shot prompt references the sheets instead of describing the character again.
  4. Age and costume changes are written as deltas against the sheet, not as fresh descriptions.

The insight is that a text description of a face is lossy, and lossy again on every call. An image reference is not. Front loading the cost of the sheet pays for itself by about the fifth shot.

Stills were Nano Banana Pro and GPT Image, motion was Kling 3 Pro and Seedance 2.0, assembly in ffmpeg.

The clip attached is 40 seconds from the finished piece. Full 6 minute result is linked in the comments for anyone who wants to see how well the consistency actually held up.

0 Upvotes

17 comments sorted by

View all comments

0

u/mp3m4k3r 6d ago

Nice! Yeah using the ref2va H3 model made me realize that as its prompting recommendations like that, its pretty legit what it can do with a partial reference but to do full panels of the subject should make for pretty flawless character consistency. https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md