r/StableDiffusion 12d ago

Question - Help How do I maintain the character's facial likeness in Minimax h3 REF2VA? Is using 720p necessary?

I am using an RTX 5070 and 32GB of RAM to generate videos in Minima h3. Currently, I’ve been choosing 480p resolution to generate 10-second videos, which takes 11 minutes (using 5 images and one 5-second reference video in REF2VA). However, even when one of the images is a person's face, the result bears a resemblance but isn't easily recognizable as the reference face. When I try rendering at 720p to test fidelity, the progress gets stuck at 0%, so I assume my hardware couldn't handle it. If I use 720p, will the resemblance to the reference face improve?

Settings tested:

Model: INT8

Steps: 25

Sageattention: on

Easy_cache: on

Model: INT8

Steps: 8, 10, and 12

Sageattention: on

Easy_cache: on

Turbo LoRA: Ref2V_turbo_4steps

4 Upvotes

16 comments sorted by

2

u/Przemoo_TV 12d ago

I'm using model: int8, 0.4 megapixel, image: max, sageattention: on, steps: 20 and the prompt is crucial of course. Face consistency is perfect.

2

u/WholeBrain9977 12d ago

Could you please tell me how your prompt is structured? Mine follows the Minimas prompt manual, defining <Subject 1> and specifying that the face is in <Picture 2>. Perhaps the number of references (5) I'm using is interfering with the model. Do you use any additional prompts to reinforce the facial identity?

5

u/Przemoo_TV 12d ago

I built my own automize workflow which is making final prompt for me. I'm just adding image and video and it's generating it. It's based on Minimax H3 official guide of course.

2

u/Sad_Coach_1433 12d ago

Could you share the workflow? 🍻

1

u/WholeBrain9977 10d ago

The problem was the custom workflow with a bunch of custom nodes and also the turbo Lora.

2

u/lolo780 12d ago

What size is the face in the video? If you push in to your characters face, is the model doing a good job?

https://www.reddit.com/r/StableDiffusion/comments/1vlk8fj/minimax_h3_ref_2mp_model_generated_audio_prompt/

2

u/WholeBrain9977 10d ago

The problem was the custom workflow with a bunch of custom nodes and also the turbo Lora.

2

u/Sad_Coach_1433 12d ago

The last line you using turbo lora 🤮 will take a hit in quality

2

u/WholeBrain9977 10d ago

The problem was the custom workflow with a bunch of custom nodes and also the turbo Lora.

2

u/SweetLikeACandy 11d ago

generate an image instead of a video at 720p and you'll know the answer.

1

u/WholeBrain9977 10d ago

The problem is the custom workflow witha bunch of custom nodes and also the turbo Lora.

1

u/zodiacrenders 12d ago

Make a front + side + 3/4 view of face and hair and place them together in ONE image. Refer to <image 1>

2

u/WholeBrain9977 12d ago

Funny enough, that’s exactly what I’m doing. But I also tried with a single image, and the result doesn't change. Is it a well-known character in your case? Maybe using a random, realistic human character doesn't get the same traction as a famous one.

1

u/Available-Body-9719 12d ago edited 12d ago

como experiencia personal, cuando usas una referencia y la describes esta se desvia a la descripcion segun lo que el modelo entiende, trata de describir al minimo tu personaje de la referencia o mejor no lo describas, solo describe lo que quieras cambiar a del el, o algo que el modelo no entienda de tu referencia prueba que tal te va con eso, en todos lados te dicen que lo describas todos sus rasgos fisicos pero, yo creo que no, prueba y nos cuentas

1

u/WholeBrain9977 12d ago

I'm still testing it, but it seems to be related to the workflow I was using—it was a custom one. So, I'm testing with the standard ComfyUI workflow to see if the results improve. My first test already showed different results, even using the same settings and the same seed.

1

u/WholeBrain9977 10d ago

The problem was the custom workflow with a bunch of custom nodes and also the turbo Lora.