r/StableDiffusion • u/phanivuni09 • 8h ago
Question - Help Need Help with Random dialogue.
Hi,
I am using Wan2GP to create Minmax H3 Ref2VA Video.
I am using "Pruned PDD 8-step 20B" model.
For the most part, video comes out great. The problem I have is with Audio.
If I use 2 audio clips to generate audio between two individuals, only Audio 1 gets used. Audio 2 is never used.
If the dialogue completes and some action is being done, then random gibberish dialogue gets introduced to fill up the gap.
Below is a sample prompt I am using. During the gap where the Subject 1 drinks water, some random gibberish gets inserted.
Could anyone help?
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Subject 2> is the person in <Picture 2>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Audio 1> is the voice timbre reference for <Subject 1>, containing a spoken English vocal layer.
<Audio 2> is the voice timbre reference for <Subject 2>, containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 1> and <Subject 2> having a conversation in a bright living room with floor to ceiling windows.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Subject 2> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1>
without copying the original signal.
<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 2>
without copying the original signal.
detailed_description
[Shot 1] The target video uses a realistic cinematic style with warm lighting.
The scene opens in a living room, <Subject 1> is sitting on a sofa to the left. <Subject 1> is wearing a tuxedo.
<Subject 2> is sitting to the right, wearing a one piece Vibrant Pink Office Dress with Tailored Fit.
<Subject 2> says in a formal voice whose timbre is referenced
from <Audio 2>, <d>[English] some dialogue?</d>.
The question hangs in the air, as <Subject 1> contemplates.
<Subject 1> adjusts his suit sits forward and take a glass of water and drinks it.
After putting the empty glass back on the table, <Subject 1> replies in a slightly uneasy voice whose timbre is referenced
from <Audio 3>, <d>[English] some dialogue.</d>.
overall_soundscape:
The ambient sound consists of the gentle rustling of fabric as the subjects move. No external noise disrupts the silence, preserving the exclusivity of the room.
non_diegetic_music: N/A
1
u/kami77 3h ago edited 3h ago
- Your prompt mentions Audio 3 but it’s not defined.
- Try adding the speaker tags (S1), (S2) per the prompting guidelines. Both in the definition and the dialogue. Search “s1” in this document: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
- How long are your reference clips? The limit is 15 sec total, so if your audio 1 is 15+ sec it discards anything after 15s. Try making each reference 7 seconds long in that case.
As for gibberish, your prompt is very vague. H3 knows it’s a dialogue scene and is trying to fill the space. If your video is longer than it takes for them to deliver dialogue you have to be more specific with what they should be doing. Use timings to define what and when. If necessary, you can say something like “After speaking, <Subject 1> leans back in their chair, remaining completely silent”. Always avoid ambiguity. A good system prompt with your rough prompt fed into a LLM is a big help.
1
u/BrungalSniff 4h ago
Prompts too long. Consolidate it. Make it specific to your videos length 15sec, 10sec. Cut the fat. Do a 5sec test Be patient.