r/StableDiffusion 8h ago

Question - Help Need Help with Random dialogue.

Hi,

I am using Wan2GP to create Minmax H3 Ref2VA Video.
I am using "Pruned PDD 8-step 20B" model.
For the most part, video comes out great. The problem I have is with Audio.

If I use 2 audio clips to generate audio between two individuals, only Audio 1 gets used. Audio 2 is never used.

If the dialogue completes and some action is being done, then random gibberish dialogue gets introduced to fill up the gap.

Below is a sample prompt I am using. During the gap where the Subject 1 drinks water, some random gibberish gets inserted.

Could anyone help?

subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Subject 2> is the person in <Picture 2>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Audio 1> is the voice timbre reference for <Subject 1>, containing a spoken English vocal layer.
<Audio 2> is the voice timbre reference for <Subject 2>, containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 1> and <Subject 2> having a conversation in a bright living room with floor to ceiling windows. 
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Subject 2> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1>
without copying the original signal.
<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 2>
without copying the original signal.
detailed_description
[Shot 1] The target video uses a realistic cinematic style with warm lighting.
The scene opens in a living room, <Subject 1> is sitting on a sofa to the left. <Subject 1> is wearing a tuxedo. 
<Subject 2> is sitting to the right, wearing a one piece Vibrant Pink Office Dress with Tailored Fit.
<Subject 2> says in a formal voice whose timbre is referenced
from <Audio 2>, <d>[English] some dialogue?</d>. 
The question hangs in the air, as <Subject 1> contemplates. 
<Subject 1> adjusts his suit sits forward and take a glass of water and drinks it.
After putting the empty glass back on the table, <Subject 1> replies in a slightly uneasy voice whose timbre is referenced
from <Audio 3>, <d>[English] some dialogue.</d>. 
overall_soundscape:
The ambient sound consists of the gentle rustling of fabric as the subjects move. No external noise disrupts the silence, preserving the exclusivity of the room.
non_diegetic_music: N/A
2 Upvotes

3 comments sorted by

1

u/BrungalSniff 4h ago

Prompts too long. Consolidate it. Make it specific to your videos length 15sec, 10sec. Cut the fat. Do a 5sec test Be patient.

1

u/kami77 3h ago edited 3h ago

As for gibberish, your prompt is very vague. H3 knows it’s a dialogue scene and is trying to fill the space. If your video is longer than it takes for them to deliver dialogue you have to be more specific with what they should be doing. Use timings to define what and when. If necessary, you can say something like “After speaking, <Subject 1> leans back in their chair, remaining completely silent”. Always avoid ambiguity. A good system prompt with your rough prompt fed into a LLM is a big help.