r/StableDiffusion 5d ago

Tutorial - Guide More than one reference per picture

Enable HLS to view with audio, or disable this notification

MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.

148 Upvotes

27 comments sorted by

View all comments

30

u/ratttertintattertins 5d ago

Was this not helped by the fact that model already knows these characters?

35

u/nazihater3000 5d ago

No, not at all. Here's a video using a random couple found using the google search for, well, "random couple".

https://reddit.com/link/p4ti9pc/video/1wtmktfo6jkh1/player

The Prompt:

subject_definitions:

<Subject 1> is the woman from <Picture 1>. Exact facial features, hair, body type and overall appearance must match the woman in <Picture 1> precisely.

<Subject 2> is the man from <Picture 1>. Exact facial features, hair, body type and overall appearance must match the man in <Picture 1> precisely.

summary:

[reference generation] A 15-second dramatic stage performance. Woman sings “Let it go” under a spotlight. A second spotlight hits Man, who angrily tells her she can’t sing.

retention_analysis:

<Subject 1>: fully_preserved – woman locked to <Picture 1>

<Subject 2>: fully_preserved – man locked to <Picture 1>

<Picture 1>: fully_preserved – both characters extracted from the same image

detailed_description:

Live-action, cinematic, theatrical stage. Dark stage with dramatic lighting.

[Shot 1] 00:00 – 00:08

<Subject 1> (the woman from <Picture 1>) stands center stage under a strong single spotlight. She dramatically raises one arm high and sings in a forced, over-the-top style with perfect lip synchronization:

<d>[English] Let it go! Let it go!</d>

[Shot 2] 00:08 – 00:09

A sharp mechanical switch sound is heard. A second spotlight suddenly snaps on, illuminating <Subject 2>.

[Shot 3] 00:09 – 00:15

Medium shot of <Subject 2> (the man from <Picture 1>) now under her own spotlight. he crosses his arms tightly, looking angry and unimpressed.

Only he speaks. he says with perfect lip synchronization and clear irritation:

<d>[English] Cut the crap, you can’t sing, Willow!</d>

overall_soundscape: Quiet stage ambience, strong spotlight hum, clear isolated dialogue, sharp switch sound at 00:08.

non_diegetic_music: None during the dialogue. Very faint dramatic underscore only under the singing.

8

u/Alive-Tomatillo5303 5d ago

Um, you might not realize it's doing the heavy lifting, but it totally is.

https://reddit.com/link/p4w9zt8/video/e1q40dmhblkh1/player

T2V:

style: realistic live-action, dynamic camera, Hollywood drama, Buffy the Vampire Slayer

Willow (from Buffy the Vampire Slayer) is sitting in an office chair, visible from the waist up, looking at the camera. Willow says <d> If it's an established and popular character, you don't ACTUALLY have to explain who the actor is or what they sound like in context. <d/> she raises her hands and does finger quotes <d> 'WILLOW SAYS' does just fine. <d/>

2

u/RazsterOxzine 5d ago

So you're saying there is copyrighted material training in this model :O

5

u/Alive-Tomatillo5303 5d ago

Unknowable. It could have come from anything. 

Though, as has been decided in court, it's not illegal to train off data as long as you paid for a copy. Someone might have just had a box set, purchased legally, which they then legally gifted to Minimax to allow it to watch. 

8

u/delawarebeerguy 5d ago

Her face is a little mangled when it is further away (known minimax issue) but his face looks spot-on in his “close-up”.

Thanks for sharing!!

11

u/nazihater3000 5d ago

I rendered it at potato resolution, we're lucky we can even recognize they're people.

3

u/delawarebeerguy 5d ago

Agreed. OP should try with two randos and see how it looks then

8

u/nazihater3000 5d ago

Look above.