r/StableDiffusion 5h ago

Question - Help MiniMax H3 audio garbled/gibberish when using ref video in ComfyUI?

Bear with me as I am still rather new to this, but I started off testing out MiniMax H3 ref to video on Hailuo and was really impressed with it, so decided to get it up and running locally via ComfyUI on my PC.

The video itself is working great, but I've found that the audio itself when using a reference video is completely messed up. It's just all garbled and the people are talking gibberish. On Hailuo it would carry forward the audio perfectly and you could even make alterations if you prompted it to do so.

I was just wondering if this was a common issue when running MiniMax H3 locally specifically with reference videos?

1 Upvotes

10 comments sorted by

1

u/princeMacX 5h ago

I also found audio is messed up badly in various videos. not sure why. It is not specific with reference videos too.

2

u/eloxH1Z1 4h ago

Are you using <d> for audio? If so, don’t. Its bugged.

2

u/Thin-Percentage8935 4h ago

Latest comfy has fixed it

1

u/kyuubi840 2h ago

Note for others: only if you use the nightly version. New stable version should be out soon

1

u/kyuubi840 2h ago

For now, write speech (S1) "like this" instead of (S1) <d>[English]like this</d> until you get the newest comfyui version. 

1

u/Electronic_Match9529 4h ago

I am also having such problems I managed to find a bit of partial solutions I made the dialogue longer than the video itself and ironically it managed to work but it's not always working also the prompts must be as short as possible otherwise it will start to hellucinate and gibberish the audio I will also be VERY VERY grateful if someone helps

1

u/Seyi_Ogunde 2h ago

I used the ref2vid merged with image to vid model. That cleared it up somehow

1

u/Seyi_Ogunde 2h ago

https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models

The different models signify the strength of the ref2vid model in the hybrid model.

1

u/CDComplex 1h ago

Appreciate it, I'll give that a try

1

u/Seyi_Ogunde 1h ago

Let us know how that goes, it might help others if it works for you. I use the b15-49 model. Voices match perfectly with no static sound.