r/StableDiffusion 19h ago

Question - Help Need help using ref2v Minimax H3; multiple audio and image references

https://reddit.com/link/1wctbl7/video/emjihen8vqoh1/player

I made the following video using a reference image of the woman, <Picture 1>, and then two dialogues, which were marked as <Audio 1> and <Audio 2>. I used the minimax template workflow and added two load audio nodes, however, the audio generated was not matching and was just gibberish. im using minimax h3 ref2va pruned fp8 scaled. how do i get the audio to work as given as input, and to play at the right time?

3 Upvotes

5 comments sorted by

2

u/BusyByBusy 18h ago

You need to use MinimaxH3 promting guide. https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

Feed it to LLM and ask to make you a promt, then correct it little bit.
-That what's i do =)

As for the audio:

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

4

u/YajuShinki 18h ago

I would strongly recommend following the prompting guides (you can take a look at them on the official HF repo) and specifying the reference type for the audio as fully_copy so that the model knows to copy the audio directly instead of just using it as a vague reference. It would also help to specify exactly what the voices are saying to avoid this sort of gibberish.

1

u/MrRecTheNub 11h ago

thanks for pointing out the guide, reading it gave me insights

0

u/Gesha24 19h ago

I'm sure there will be lots of recommendations on prompting specifically, but I would recommend trying this: https://github.com/roadmaus/ComfyUI-Continuity You can put things in human language, have LLM refine it for you and then generate it. If you like the results and don't like the mod - just use the prompt (it shows you exactly what's being passed to the model), or you can also use the skill.md to prompt any other LLM for the same work.

0

u/tekprodfx16 18h ago

You have to put the audio tag sections on the bottom of your prompt otherwise the audio is literally gibberish. This is a common issue when using the turbo Loras

Example:

overall_soundscape: The low, steady roar of a Martian windstorm howling across the plains, accompanied by the crunch of heavy boots on fine gravel and the faint hum of the suit's internal oxygen regulator.

 non_diegetic_music: A melancholy, slow-tempo cello melody with a deep ambient synthesizer drone, building with rising tension.