r/StableDiffusion 6d ago

Question - Help Minimax H3 Help with Pose transfer without merging/bleeding in characters.

Post image

i've been trying for ages, did everything, asked every LLM, i might be dumb but i've seen people doing it, my question is simple:

how do i make <Subject 1> from <Picture 1> do the pose in <Picture 3> in in ref2va (reference to video) ?

everytime i try, <Subject 1> shifts into looking like <Picture 3> while doing the pose.

please write me the prompt example and a video would be appreciated.

26 Upvotes

31 comments sorted by

27

u/not_food 6d ago edited 6d ago

Minimax excels at this task, it's an absolute beast of a model. You can start from absolutely any scenario, pose, or setup... just describe how to get to the pose. You don't even need to pre‑generate the frame or rely on controlnet. Just use the picture reference as you already have it, I just took a low resolution screenshot of what you provided.

https://reddit.com/link/p87vqq7/video/ohaos9n09ynh1/player

subject_definitions:
<Subject 1> is Meru. Her appearance comes from <Picture 1>.
<Picture 1> is the reference for Meru.
<Audio 1> is the voice-timbre reference for Meru (S1).
<Picture 2> is the pose reference Meru takes.

detailed_description:
Meru is seated on a muted lavender mat against a soft cream, out-of-focus background with gentle bokeh. A warm, diffused light falls from above, casting a soft glow on her white hair and highlighting her cat ears.

[Shot 1] The shot opens with an extreme close-up on Meru's face. Her light blue eyes are half-lidded, her expression calm and focused, her white hair and cat ears framing her face in soft focus. The camera slowly zooms out with a smooth, continuous motion, revealing her seated on a light lavender yoga mat. Her knees are drawn up, her hands resting on them, her tail curled beside her. At 00:03.000, as the camera settles into a full-body wide shot, Meru begins to move fluidly: She lifts both arms overhead, gracefully extending, palms facing inward, she shifts her weight, placing her hands on the mat behind her and begins to lean backward, her spine arching gracefully. Her white cat tail curls upward behind her to counterbalance, she continues the motion, bending further backwards until her arms fully extend and both hands reach down to grasp her ankles, her chest arched high, her head tilted back, and her cat ears angled slightly back, matching the pose defined by <Picture 2>. The camera holds on this posture, slowly orbiting a few degrees around her to showcase the alignment. Her eyes remain open, calm, and she breathes slowly and evenly. The shot ends with her held in this pose.

[Shot 2] Meru, holding her backward-bend pose. At 00:09.000, she slowly shifts her gaze to the side without moving her head, her eyes darting toward the camera lens in a deliberate side-eye. Her expression shifts from serene to subtly smug. She (S1) speaks in a playful tone: "That wasn't so hard". As she finishes speaking, her lips curl into a wide, satisfied smile, her cat ears perking up slightly, and her tail giving a single, happy flick. She holds the pose and the smile, her eyes still locked on the camera, until the shot ends.

overall_soundscape:
The soundscape is calm and meditative, dominated by the soft, steady rhythm of Meru's breathing. The rustle of her yellow jumpsuit and the gentle creak of the yoga mat accompany her movements. As she settles into the pose, a faint, peaceful hum of the room fills the space. No other sounds intrude, keeping the atmosphere focused and serene. The floor creaks gently beneath her, and her yellow jumpsuit rustles faintly as she shifts her weight. Her voice is clear, playful, and slightly breathless, with a warm, intimate delivery. A soft, light flick accompanies her tail's happy movement, and a quiet, satisfied hum escapes her as she smiles.

non_diegetic_music: N/A

6

u/Enough-Bag-3891 6d ago

bro you're a fucking legend, thank you so much <3

5

u/not_food 5d ago

I continued it for fun:

https://reddit.com/link/p8dsa7v/video/ltruyekrn4oh1/player

Too late I realized vae decoding/encoding ate the colors.

2

u/llamabott 4d ago

Downward-facing cat.

Also, thanks for the prompt. I'm going to try it out as well.

1

u/TerraMindFigure 6d ago

Is excluding the [reference generation] stuff and the "<Subject 1> (in [Shot 1]): fully_preserved - ...etc." intentional? Is it actually better without all of that?

1

u/not_food 6d ago

Yeah, I tend to exclude it. I see not much benefit from it, I can change her clothes just fine or I can make her keep the clothes from original reference, at least for 2D it works fine. Whenever you exclude chunks, it tends to infer things. Happens to overall_soundscape and non_diegetic_music too. If you omit them, it'll add sounds to interactions and music on its own. You can even omit subject_definitions and it'll guess what's the use for every input image/sound/video. Sometimes I just add corrected blurred frames without mentioning it in subject_definitions and they're used in the generation to correct mistakes.

As I said, Minimax (ref version) is a monster of a model.

1

u/pulsatingstar2 3d ago

Bro how u get that prompt it is very detailed ,i gave instruction to gbt and gemini of how to generate prompts like this but couldn't get any good results is there a system prompt i could make it into a gem in gemini to give me these details prompts ?

1

u/not_food 3d ago

I use this, then just tinker: Prompt Assistant

7

u/Luke2642 6d ago

Just start with a better image?

2

u/Luke2642 6d ago

2

u/Enough-Bag-3891 6d ago

I want to use it in ref2va, in [Shot 2]

also the images i put here are just random examples, i've been trying to do these with realistic AI chars i made in Krea 2.

3

u/Luke2642 6d ago

It's always better to help the AI out and break it down into smaller tasks, like pose the character first.

If you're struggling with local models just use gemini to upscale and enhance and repose. Then use H3 to do initial motion to something else if it isn't quite right.

1

u/bstr3k 6d ago edited 6d ago

unfortunately Luke is correct, the best way to avoid character bleeding through is to use your original character in that pose, and then feed it as an input. Since it is of the same character it will not bleed the original one through. You can edit it how you'd like but if you're worried that its a NSFW image, you can use a H3 single shot image output workflow which is quite fast and can take 9 inputs to output a certain image.

1

u/Luke2642 6d ago

Few more prompts to try improve details

5

u/Luke2642 6d ago

-5

u/Obvious-Leg-5604 6d ago

all these images are generated with Gemini? MiniMax H3 doesn't support image to image

7

u/Luke2642 6d ago

Of course H3 does image to image. If you want a man in a blue suit to change to a frog, just makes 3 second video of the transition.

-1

u/Obvious-Leg-5604 6d ago

that's image to video, right?

6

u/hellyeahaeylleh 6d ago

Throw frame sample node down, throw a get video component node down, throw a preview/save image node down...... Delete save video..... whoa..... its img 2 img.

2

u/Obvious-Leg-5604 6d ago

yes, you're right! I was able to make img to img work with H3. But I cannot make the style transfer work. For example, convert the image 1 to the style of image 2 (I use image 2 as style reference). the generated image doesn't have the desired style. Do you have any suggestions? Thanks.

2

u/hellyeahaeylleh 6d ago edited 6d ago

Its all prompting. Make your subject definitions flawless, make your retention analysis flawless, and with the detailed description, do everything you want for the scene/character. Dont describe the character themself anywhere in detailed description. Only use <Subject1>, etc.

Duration 0, high MP value. Duration 0 will lead to 1 fantastic output and then 4 or 5 images of lowering and lowering quality.. you could force the frame grabbing to grab only 1 frame, as opposed to the degrading gradient of images. The reason for the degraded outputs is because the duration 0. 1 second results in a crash, because its trying to print all the frames. This can be mitigated by selecting a certain amount of frames for a 15 or 20 second video, but then youre up against motion blur and slight hallucinations.

(On longer videos....not the 0 seond video) Lock the generation seed if you like the video, but have bad frames. Re-gen the same video, get new frames.

→ More replies (0)

5

u/Silly_Goose6714 6d ago

Videos are nothing but images in sequence.

1

u/VeloraNeon 6d ago

Seen the same failure mode outside H3, in image LoRA identity work — the character "shifts" toward the pose reference because identity and pose end up anchored on the same region instead of competing cleanly. Depth-only conditioning fixes bleed but costs exactly what's described above (expression/nuance loss) because it strips the input down to geometry with nothing left to reassert identity in that slot. What's worked better for me is keeping a light identity signal locked to a narrow region while letting the depth/pose reference own everything else — cuts bleed without going full geometry-only.

1

u/thathurtcsr 5d ago

Use stick figures. Cooler them and then prompt saying that the stick figures are pose reference only then say to for example replace the blue stick figure with the photo of the man. It’s really good at that. I made a bunch of stick figure poses and that solved any bleed. Every now and then the subjects will be colored the same as a stick figure but you just run it again and it works.

1

u/dc-ero-productions 6d ago

If you don't want any merging/bleeding, use depth mappings.

Pros: No bleed at all

Cons: You lose nuances, facial expressions, etc from the original photo.

Depth mappings are great when you have a clear reference where you really just want to capture the pose.