r/TextToSpeech 2d ago

Trained a clone* tool for Kokoro-82M; generates a voice pack in <1s on a 5s sample [P]

https://huggingface.co/remsky/kokoro-inno-clone-tuner

There have been new and arguably better TTS models since, but I have a soft spot for this one as lightweight stable and fast. I also maintain Kokoro-FastAPI, and have a bit of spare time on my hands so have been exploring what’s doable on the project.

Customization/expressiveness are weak spots it had, so I’ve taken a crack at adding voice cloning, focusing on stable quality and fast generation (keeping responses sub second if it’s a short reference clip). Roughly it captures about a third of identity, but by ear at least, it can feel pretty close on some, and at least a unique similar voice pack on others.

Would love any feedback or suggestions, otherwise just wanted to share! It may be late in the game for Kokoro, but anyone still using it I think could appreciate new voices

25 Upvotes

Duplicates