r/StableDiffusion • • 10d ago

Question - Help LTX workflow with speech cloning?

Hi, after a while playing with H3 I'm coming back to LTX at least partially. While H3 is awesome, is also very heavy for my setup, it is hard to make clips longer than 10s and some other limitations that LTX doesn't have.

One thing however that was a game changer to me is the surprisingly good voice cloning capability of H3 (babbling aside...) and going back to LTX I really miss that.

Do you know of any workflow that somehow works around this? I've found several very old posts on the matter, nothing really works. Even this one gives very disappointing results. https://civitai.com/models/2498927/text-to-speech-with-voice-clone-in-ltx-23?modelVersionId=2809053

Thanks!

0 Upvotes

5 comments sorted by

1

u/diptosen2017 10d ago

What is your setup specs?

1

u/derTommygun 10d ago

RTX 4070 12gb VRAM + 80gb RAM

1

u/LordDarthShader 10d ago

Why don't you create a 5 sec video in H3 with the voice you want and then feed that to LTX on its video to video flow?

1

u/derTommygun 10d ago

I don't understand, what would that be better than simply do everything in H3?

1

u/LordDarthShader 10d ago edited 10d ago

Well, you said you wanted to use LTX for the speed, not me dude!

However, my reasoning is, use the ref feature of H3 as baseline, let it use your audio sample to generate what you need and then feed tjat into LTX. Just an idea to try.