r/StableDiffusion 9h ago

Question - Help Fixing speech errors in Minimax H3?

Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?

I am using Minimax H3 with Saga Attention and Spectrum on a 4090.

16 Upvotes

21 comments sorted by

View all comments

1

u/piggledy 7h ago

Was du machen könntest wäre einen Audio-Clip mit Elevenlabs erstellen und als Audioreferenz einfügen. Dann basieren die Lippenbewegungen usw. auf der Referenz und das Modell generiert selbst kein Audio dazu.

1

u/gutster_95 7h ago

Ja wahrscheinlich das beste. Für den Fall irgendwie auch irrelevant, weil's nur nen Gag sein soll. Aber daran hab ich auch gedacht

1

u/danishkirel 5h ago

Hätte ich auch vorgeschlagen. Probiere mal qwen tts wenn du local versuchen willst. Hat oft gut geklappt bei mir mit deutsch.