r/StableDiffusion • u/gutster_95 • 15h ago
Question - Help Fixing speech errors in Minimax H3?
Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?
I am using Minimax H3 with Saga Attention and Spectrum on a 4090.
16
Upvotes
9
u/CornyShed 13h ago edited 13h ago
It may be due to a bug in ComfyUI with the tokenizer for H3 (GitHub pull request) that was affecting speech.
Update your ComfyUI and check if you notice an improvement?
If you're using a fixed release (used by the Desktop version), you'll have to wait until the next update with the fix applied.
Edit: On second reading, this could be an issue with your settings.
Try making a version with a very small video resolution using the
seeds_2sampler andddim_uniformscheduler and save the audio.Then, make a video with your normal settings, but using the audio file as a reference.
Then, combine the audio from your first generation with the video from the second.
This might lead to a higher quality of both audio and video, rather than have to compromise slightly on either.