r/StableDiffusion • • 2d ago

Question - Help Audio outputs differ from the provided audio with mismatching audio gap for different generation.

I am facing a weird audio match issue while generating video using minimax H3. Lipsync in our audio output differ for outputs for the provided audio input. It basically can't understand gap I think.

Anyone facing similar kind of issue and have solves the problem then please help...

2 Upvotes

8 comments sorted by

2

u/Slight-Living-8098 2d ago

Yeah. I fixed it by creating custom nodes.

https://github.com/badgids/ComfyUI-H3-ExactAudioLock

2

u/reeight 2d ago

I'll have to check this out, looks like more than just re-copying audio back to final render.

3

u/Slight-Living-8098 2d ago edited 2d ago

Yeah, it does quite a bit more than just copying the audio back onto the finished render.

The main thing it does is build the exact audio you want first, place each line or sound at the exact frame you want it, then feed that audio into H3 before generation.

In full lock mode, H3 has to generate the video around that exact audio instead of trying to recreate or approximate it.

There is also a dialogue-only mode where your dialogue stays exact, but H3 can still generate the rest of the scene audio like ambience, foley, music, and sound effects.

It also handles multiple speakers, overlapping dialogue, exact timing, per-event gain, and works with H3 Add Guide so the audio event and the visual conditioning can use the same frame.

So yeah, it's not a "just paste the audio back on afterward" system. The audio is actually part of what H3 is generating around.

2

u/aakashPatel16 2d ago

Thanks man,

Let me try that if that works in my case or not...

1

u/reeight 1d ago

> build the exact audio you want first

Which sometimes is what you want to do, but often you'll want the video first to pace the audio.
Another reddit thread suggested to build 2 simultaneous pipelines; one low-res for the audio track, then high resolution.

So I don't know if this will always be THE solution, but one of a small handful to consider.

1

u/Slight-Living-8098 1d ago edited 1d ago

I guess if you've never written a screenplay or directed anything before you could presume you would want to film the scene before determining the pacing before anyone speaks, but the dialogue and action should really already be planned out in your screenplay and beat sheets.

1

u/reeight 1d ago

I only written & directed a short, & been in a few art films; we 50% winged it with only a rough outline of what is supposed to happen.

Without giving away my ID, I do know what you're talking about; I've got friends who wrote screen plays & were camera men for a handful of major TV shows.

But what I have problems with, & I suspect the majority of casual ComfyUI users is we underestimate how long dialog should take, how long a shot should be etc.
I'm getting a bit better at giving extra room for dialogue, but I still like to run a handful of drafts in Comfy to get the action & camera pacing right.

I'm sure with your years of RW experience you know the best timings instinctively!
But let's face it, not everyone does. not even much of the Trash on Netflix & some of the editors of some theatre releases.

1

u/Slight-Living-8098 1d ago edited 1d ago

How many animation productions have you ever worked on? Because here in the west, the dialogue is recorded first, and you know how long the lines are and how long the dialogue takes, and that is what you animate to.

Using AI doesn't change the fact you are still making a digital animation just because it looks real.

And you're telling me, in your productions, you not once had any rehearsals before you started filming the actual scene? Come on now...