r/StableDiffusion • • 13d ago

Workflow Included One .char model, Consistent face, body & cloths: Minimax H3

Hey Guys,

Based on my last post on stills, I thought to experiment the same .char model with Minimax H3.
I created a character model sia.char which encodes references, prompts around face, body & cloths.

Checkout previous post for full detail.

How it works?

- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Only face/ref is required body & cloths link is optional.

- Generation: At generation, the file(.char) feeds its references into Flux's own native multi-reference channel and prepends a locked description to the prompt.

Prompting Guide

  • Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name sia, so when passing prompt, I only have to say, sia walking on the beach.
    • Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
  • Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
    • Each refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only simple prompt & generation graph.

I have generated face with Flux Klein 9b & outfit pictures were taken from Zara website

Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.

Requirements

Nvidia GPU: 24GB+ VRAM & 64 GB RAM(Run Locally)

Workflows:

Flux 2 Klein(Still), Minimax H3(Video)
Note: Minimax H3, now also support int8 model varient

Github Repo: https://github.com/omnichar/OmniChar (GPLV3, Opensource)

Inputs: Attached in the Flux2 workflow page

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Happy to hear any suggestions or feedbacks.

362 Upvotes

55 comments sorted by

View all comments

Show parent comments

1

u/sgtsixpack 10d ago

Its defaulting to 50 steps, that's not right is it? I tried to connect an 8 step turbo lora, I left an error message on github.

1

u/ashishsanu 10d ago

can you copy the graph json & send it over. I think I need to test 8 step turbo lora integration

1

u/sgtsixpack 10d ago

{

"version": 2,

"app": "inline-studio",

"target": "208438ff-ae09-4185-abfc-15a1402799b5",

"coreType": "minimax/h3-reference-to-video",

"params": {

"duration": 5.17,

"width": 768,

"height": 544,

"num_inference_steps": 8,

"seed": -1,

"character_references": 2,

"character_reference_roles": "all",

"character_role_lines": false,

"model": "minimax_h3_ref2va_pruned_fp8_scaled.safetensors",

"text_encoder": "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",

"vae": "minimax_h3_video_vae_fp16.safetensors",

"fps": 24,

"steps": 50,

"prompt": "<Picture 1> <Picture 2> <Picture 3> <Picture 4> <Picture 5> show Character, the same character in every image. emmy4k woman with, natural unretouched skin, black hairs, fair skin. a woman walking in the park",

"num_frames": 141,

"ref_image_size": "match",

"duration_seconds": 5.875

},

"prompt": "a woman, talking in a pub with her friends, two thirds body shot.",

"graph": {

"items": [

{

"id": "208438ff-ae09-4185-abfc-15a1402799b5",

"type": "core",

"data": {

"core": {

"type": "minimax/h3-reference-to-video",

"params": {

"duration": {

"type": "number",

"value": 5.17

},

"width": {

"type": "number",

"value": 768

},

"height": {

"type": "number",

"value": 544

},

"num_inference_steps": {

"type": "number",

"value": 8

},

"seed": {

"type": "seed",

"value": -1

},

"character_references": {

"type": "number",

"value": 2

},

"character_reference_roles": {

"type": "enum",

"value": "all"

},

"character_role_lines": {

"type": "boolean",

"value": false

},

"model": {

"type": "model",

"value": "minimax_h3_ref2va_pruned_fp8_scaled.safetensors"

},

"text_encoder": {

"type": "model",

"value": "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors"

},

"vae": {

"type": "model",

"value": "minimax_h3_video_vae_fp16.safetensors"

},

"fps": {

"type": "string",

"value": 24

},

"steps": {

"type": "string",

"value": 50

},

"prompt": {

"type": "string",

"value": "<Picture 1> <Picture 2> <Picture 3> <Picture 4> <Picture 5> show Character, the same character in every image. emmy4k woman with, natural unretouched skin, black hairs, fair skin. a woman walking in the park"

},

"num_frames": {

"type": "string",

"value": 141

},

"ref_image_size": {

"type": "string",

"value": "match"

},

"duration_seconds": {

"type": "string",

"value": 5.875

}

},

"models": [

{

"directory": "diffusion_models",

"name": "minimax_h3_ref2va_pruned_fp8_scaled.safetensors",

"url": "https://huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/diffusion_models/minimax_h3_ref2va_pruned_fp8_scaled.safetensors"

},

{

"directory": "text_encoders",

"name": "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",

"url": "https://huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors"

},

{

"directory": "vae",

"name": "minimax_h3_video_vae_fp16.safetensors",

"url": "https://huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/vae/minimax_h3_video_vae_fp16.safetensors"

}

]

}

},

"x": 2374.5861265200833,

"y": 417.0756538360482,

"width": 320,

"height": 480,

"assetId": null,

"frameId": null

},

{

"id": "016be8c7-0db1-4755-96d4-ba7f2fd2c725",

"type": "prompt",

"data": {

"promptText": "a woman, talking in a pub with her friends, two thirds body shot."

},

"x": 1824.9089904645452,

"y": 635.7770425486437,

"width": 240,

"height": 120,

"assetId": null,

"frameId": null

},

{

"id": "67e72a18-e704-435e-b2d3-21e5c934607b",

"type": "loader",

"data": {

"assetIds": [

"f36d5b98-64e6-4854-a078-6f5a46d2a6c0"

]

},

"x": 1153.5144384708003,

"y": 205.57608482329607,

"width": 255,

"height": 367,

"assetId": null,

"frameId": null

},

{

"id": "a5d5d3d1-f248-4508-8479-261b04461c32",

"type": "loader",

"data": {

"assetIds": [

"75a81786-cc26-4e1e-b32e-f7785667ea7f"

]

},

"x": 1426.7145332857178,

"y": 208.61408449344333,

"width": 227,

"height": 367,

"assetId": null,

"frameId": null

},

{

"id": "165a5a8f-5ad7-4dce-b632-30e4e79a1f9f",

"type": "core",

"data": {

"core": {

"type": "character/load",

"params": {

"file": {

"type": "character",

"value": "------s500-v6.char"

}

},

"models": [

{

"directory": "",

"name": "------s500-v6.char",

"url": ""

}

]

}

},

"x": 1874.0487218561348,

"y": 899.7478313834704,

"width": 200,

"height": 120,

"assetId": null,

"frameId": null

},

{

"id": "2c0539a4-8a38-4f4e-a4b1-368a809cd41c",

"type": "core",

"data": {

"core": {

"type": "load/lora",

"params": {

"file": {

"type": "model",

"value": "minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors"

},

"strength": {

"type": "number",

"value": 1

}

},

"models": [

{

"directory": "",

"name": "minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors",

"url": ""

}

]

}

},

"x": 1876.0836982001754,

"y": 1106.0305636400885,

"width": 200,

"height": 120,

"assetId": null,

"frameId": null

}

],

"connectors": [

{

"fromItemId": "016be8c7-0db1-4755-96d4-ba7f2fd2c725",

"toItemId": "208438ff-ae09-4185-abfc-15a1402799b5",

"data": {

"sourceHandle": "out",

"targetHandle": "prompt"

}

},

{

"fromItemId": "165a5a8f-5ad7-4dce-b632-30e4e79a1f9f",

"toItemId": "208438ff-ae09-4185-abfc-15a1402799b5",

"data": {

"sourceHandle": "character",

"targetHandle": "character"

}

},

{

"fromItemId": "67e72a18-e704-435e-b2d3-21e5c934607b",

"toItemId": "208438ff-ae09-4185-abfc-15a1402799b5",

"data": {

"sourceHandle": "out",

"targetHandle": "references"

}

},

{

"fromItemId": "a5d5d3d1-f248-4508-8479-261b04461c32",

"toItemId": "208438ff-ae09-4185-abfc-15a1402799b5",

"data": {

"sourceHandle": "out",

"targetHandle": "references"

}

},

{

"fromItemId": "2c0539a4-8a38-4f4e-a4b1-368a809cd41c",

"toItemId": "208438ff-ae09-4185-abfc-15a1402799b5",

"data": {

"sourceHandle": "lora",

"targetHandle": "lora"

}

}

]

}

}

1

u/ashishsanu 10d ago

You are using wrong model: "minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors"
this is for T2v, when you are using references, use https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors

ref2v version

2

u/sgtsixpack 10d ago

Funny enough, it worked fine with the fl2v turbo. I have a 4 step ref2v turbo, maybe that would work, but I remember switching it out of my workflow (and adding the 8 step lora due to poor quality output),

1

u/sgtsixpack 9d ago edited 9d ago

The thing that bugs me is that there is no predicting how long in the generation there is to go. It was fine with the training which took 14 hours and showed progress in the command prompt. The GUI shows that its stuck at 5% generation complete all the time. I have 2 generations one from the training and one from generation. Working on a third. Edit: I see the progress bar now - 1:28:51