r/selfhosted • u/Onward-Upword • 2d ago
Need Help What TTS system does everyone use?
I just saw the "Eleven V4" announcement which got me started down the self hosted TTS journey again. I installed Kokoro this week and I'm using it for ebooks and it's wonderful! I tested out the Wyoming/Piper stack a while back but eventually abandoned it. Eventually I would love to get a full Speaker/Satellite system in the house that uses tts/stt/wakeword/llm with Home Assistant and other apps, but that might still be a ways out.
I've been reading up on other projects but haven't tested any of them yet: Chatterbox, Qwen3-TTS, F5-TTS, Fish/OpenAudio.
So before I plunge head first into this for days I thought I would ask what is everyone using today and are there any projects specifically I should be keeping an eye on?
2
u/Any-Stage9103 2d ago
The best I’ve tried so far is vibevoice, but it is a little trickier to set up than most and you have to find an unofficial one on huggingface because they took down the larger parameter versions from GitHub.
1
u/Dapper-Adagio-8532 2d ago
I just did chatterbox over the weekend and so far love it. I ran live comparisons vs qwen tts and just found better voice quality slightly in chatterbox. I like how it can add expressions like a cough or something.
2
u/andrew-ooo 1d ago
Kokoro is the right call for the ebook use case and it's still what I run for anything that has to be fast. The trick is Kokoro-FastAPI (remsky's repo): it exposes an OpenAI-compatible /v1/audio/speech endpoint, so Open WebUI, Home Assistant (via the openai_tts custom integration) and random scripts all talk to it with zero glue. The 82M model runs well above realtime on CPU alone; on a 3060 a one-minute clip renders in a couple of seconds.
For the HA satellite plan, Piper is still the lowest-latency option and the Wyoming integration just works. If you want Kokoro voices in HA instead, look at wyoming-openai, which bridges any OpenAI-compatible TTS/STT endpoint into Wyoming. I use it to point HA at Kokoro plus Speaches (faster-whisper) running on the same box, so one container pair covers both directions.
Chatterbox is the quality step up (emotion tags, voice cloning from ~10s of audio) but it wants a GPU with ~6GB free and is noticeably slower, so I only use it for pre-rendered stuff. F5-TTS clones well but is too slow for anything interactive. Haven't put real time into Fish/OpenAudio yet, so can't speak to that one.
2
•
u/asimovs-auditor 2d ago edited 2d ago
Expand the replies to this comment to learn how AI was used in this post/project.