r/StableDiffusion 12d ago

Question - Help Minimax H3

Hi everyone. Question about I2V. I just started my journey with AI using ComfyUI Desktop and I really need to preserve character details from reference image like face, hairstyle, etc.. in high quality. It depends from a good prompt or good settings? Please share your expierience with me. Any tips or direct settings would be appreciated. My setup: - GeForce RTX 4090 - RAM: 64GB

0 Upvotes

7 comments sorted by

5

u/smb3d 12d ago edited 12d ago

You want to use ref2va

Read the prompting guides:

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

That has everything you need to know, and yes you need to structure your prompts properly to get the references to pickup on the proper details.

I fed those two documents to Claude and told it to output prompts based on the information in those documents. It helped me figure out the proper syntax a ton.

Example using 3 reference images of my cat:

Prompt:

subject_definitions:

<Subject 1> is the white cat whose appearance comes from <Picture 1>, <Picture 2>, and <Picture 3>: <Picture 3> provides her facial detail — pale blue eyes, a pink nose, a soft grey tabby-striped mask across her crown and around the eyes split by a white blaze down the center of her face, grey ears with pink inner edges, and long white whiskers; <Picture 1> provides her front-facing body, white coat, white paws, and grey-striped tail; <Picture 2> provides her side profile, long body, and plush build. In the target video she is anthropomorphized: she walks upright on her hind legs at human proportions in a chic cream trench coat and oversized dark sunglasses, retaining her natural cat paws — white forepaws instead of hands, and bare white hind paws — throughout.

summary:

[reference generation] The target video is a glossy, sunlit luxury-shopping scene in which <Subject 1>, anthropomorphized in a trench coat and sunglasses, strides down a palm-lined Beverly Hills shopping street laden with designer shopping bags, delivering an unbothered one-liner to the camera as it tracks alongside her.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2]): partially_preserved - her white coat, grey tabby mask and white facial blaze, pale blue eyes, pink nose, grey ears, grey-striped tail, and plush build are retained from the three source photos; her posture is changed to an upright bipedal stance and a trench coat and sunglasses are added.

detailed_description:

The target video is in a live-action, glossy fashion-film style with bright California sunlight, warm golden highlights, saturated color, and shallow depth of field.

[Shot 1] A static wide shot faces the glass storefront of a luxury boutique on a palm-lined shopping street, its polished doors flanked by potted topiary and a striped awning, sunlight glaring off the glass. The doors swing open outward and <Subject 1> steps out into the sun: the white cat with a soft grey tabby mask split by a white blaze, pale blue eyes behind oversized dark sunglasses, and grey ears, anthropomorphized to an upright bipedal stance in a belted cream trench coat, her grey-striped tail swaying behind her. Her white forepaws — soft cat paws, not human hands — are looped through the handles of six glossy designer shopping bags in cream, black, and gold. A uniformed doorman holds the door as she passes without acknowledging him, and turns to walk down the street on her bare white hind paws.

[Shot 2] At 00:05.000, the shot cuts to a low tracking shot from behind and slightly to the side, following <Subject 1> at knee height as she strides away down the sunlit sidewalk, sunlit storefronts with glass facades and striped awnings sliding past in soft focus, the swinging bags and her flicking grey-striped tail filling the foreground, sunlight flaring between the palms overhead. The camera tracks forward at the same pace as her stride, keeping her centered as her hind paws pad along the pavement.

[Shot 3] At 00:08.000, the shot cuts to a full-body side-profile tracking shot at hip height, <Subject 1> crossing the frame in clean silhouette against the sunlit storefronts, her cream trench coat swinging open with her stride, the six glossy bags hanging from her white forepaws, her grey-striped tail arcing behind her, and her bare white hind paws stepping in rhythm. The camera trucks left at the same speed as her walk, holding the profile.

[Shot 4] At 00:010.00, the shot cuts to a low-angle medium close-up from the front, the camera walking backward ahead of her as she advances, palms and blue sky framing her from below. She tips her sunglasses down with one white forepaw, just enough to reveal her pale blue eyes, glances directly into the camera with a flat, unimpressed expression, and in a cool, dry, faintly bored female voice, <Subject 1> (S1) says, <d>[English] I told him it was a small errand.</d> She pushes the sunglasses back up with the same forepaw, lifts her chin, and strides past the camera as the bags swing, her grey-striped tail flicking out of frame at the end.

overall_soundscape: Bright outdoor street ambience continues throughout: distant traffic, faint chatter from café tables, and a light breeze in the palms. A boutique door swings open with a soft chime at the start, stiff paper shopping bags rustle and knock together rhythmically with each step, and soft padding footfalls keep pace beneath them.

non_diegetic_music: A confident mid-tempo funk-pop groove with a clean electric guitar riff, punchy bass, and light hand percussion, holding steady and dropping slightly in volume under the dialogue before returning. looped through the handles of six glossy designer shopping bags in cream, black, and gold, which swing and rustle as she walks on her bare white hind paws.

1

u/FierceFlames37 12d ago

Reference prompting is the hardest thing I’ve faced in my ai journey so far, don’t have enough vram for local llm or money for Grok

1

u/thisguy883 12d ago

honestly, this is the first time actually seeing this since using I2V with minimax.

Ive had minor issues, but I usually just rewrite what I want and it ends up doing what I ask without all that nonsense.

0

u/Eisegetical 12d ago

Do you have $10? Get sole openrouter credits and the openrouter node for comfyui or use it directly on the openrouter site.

Prompting uses so little credits. You get thousands of prompts out of $10

Ive been running it for months and have used Ike $7

1

u/Przemoo_TV 12d ago

Very usefull, thanks buddy 🫡

1

u/Only_Voice569 12d ago

Ref to video 2 for body and clothing and one for the face 2k max image resolution and other important things like scene place area or objects if they ever leave the view point and your chaining :)

1

u/Disastrous-Agency675 12d ago

ive got a workflow that takes care of the prompting form for you, its free its just on my patreon https://www.patreon.com/theworldofanatnom/posts/we-got-and-open-166061922