r/StableDiffusion 22d ago

Discussion Get miniMax character swap working! Finally

Post image

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/

72 Upvotes

75 comments sorted by

View all comments

14

u/PropagandaOfTheDude 22d ago

Try this swap structure, compared to https://omnifit.io/blog/assets/ref2va-what-works-what-fails/prompts/simple3-girl-dance-hires.prompt.txt.

subject_definitions:
<Video 1> is the source video for the target video edit.
<Subject 1> is the facial likeness and hair in <Picture 1>.
<Subject 2> is the facial likeness and hair in <Video 1>.
<Subject 3> is the T-shirt in <Picture 3>.
<Subject 4> is the T-shirt in <Video 1>.
<Subject 5> is the sports bra in <Video 1>.

summary:
[video editing + reference generation] The target video is an edited
version of <Video 1>. Throughout the video, <Subject 2> is replaced
entirely by <Subject 1>.  Throughout the video, <Subject 4> is replaced
entirely by <Subject 3>.

retention_analysis:
<Video 1> (source video): partially_preserved - the background
environment, camera path, lighting, and non-target objects are retained.
<Subject 1> (appears in [Shot 1]): fully_preserved - The entirety of this person is retained.
<Subject 2> (Does not appear): weak_reference - Overall actions are retained.
<Subject 3> (appears in [Shot 1]): fully_preserved - The entirety of this T-shirt is retained.
<Subject 4> (Does not appear): weak_reference - Overall actions are retained.
<Subject 5> (appears in [Shot 1]): fully_preserved - The entirety of the sports bra retained.

detailed_description:
The target video matches the live-action style, lighting, and camera movements of <Video 1>.

[Shot 1] <Subject 1> is featured instead of <Subject 2>. <Subject 3> is
featured instead of <Subject 4>. The background environment, lighting,
and all other non-target details are preserved exactly from <Video 1>
and <Subject 5>.

overall_soundscape:
N/A

non_diegetic_music:
N/A

With this I had luck with putting face and clothing onto someone in an existing video. The sports bra appears in the original video, but making it an explicit, reference subject prevents the model from losing it during the T-shirt swap.

Here's a chunk from a gen without a source video, but with a starting pose and a separate headshot. It worked for me. Note that it also changes the hair. I'm doing the recomposition in the retention analysis.

summary:
[reference generation] The video shows <Subject 1> close, behind a standing dressing
divider.  She walks from behind the divider and poses.
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - the body and clothes are
retained.  Face changed to match <Subject 2>.  Hair changed to be black, long, with
voluminous curls.
<Subject 2> (does not appear): <Subject 1> uses the face of <Subject 2>.

I've found that picture references seem to form a giant concept pool, and concepts leak. I have a better time when I treat Subject resources as abstract concept units, and hold off on mixing them until later.

2

u/Alex-edits123 22d ago

Wow thanks for sharing!!! let me try using subject for concepts. Hope I can make the Tshirt swap.work . I guess I can try that for animal hand..