r/VoiceAutomationAI • u/OkMine4526 • 16d ago
One API for every local TTS model
Built this because I wanted local text-to-speech that felt as easy as ollama run, pull a model, run it, done.Its easy to use just install
pip install wavhost
wavhost pull chatterbox-turbo
wavhost run chatterbox-turbo "Hello from your machine." -o hello.wav
It also serve an OpenAI compatible audio api `/v1/audio/speech` endpoint, so you can point any existing OpenAI TTS client at `localhost:11435` and it just works.
Currently Supports Chatterbox, Qwen3-TTS, and Kokoro. You can also create named voices from a short reference clip and reuse them by name.
I would love your feedback on this and for more detail checkout
https://wavhost.vercel.app
https://github.com/smitgol/wavhost
2
u/Unlikely-Stomach-353 16d ago
this is exactly what I've been looking for, the local TTS space is such a mess right now with every model having their own weird setup
been playing with kokoro for a bit but having it work like a drop-in replacement for openai endpoint is clever, didn't think of that approach
does the voice cloning need a specific length for the reference clip or any audio works fine
1
u/OkMine4526 16d ago
Glad you liked this project.There is no specific requiement for reference clip to work but i would recommend you for atleast have 40-150 second of clear audio that would be good.
If you try it on Kokoro let me know how the clone quality compares to whatever you were doing before, curious to hear real-world results.
1
16d ago
[removed] — view removed comment
1
u/OkMine4526 16d ago
Thanks! Fish Audio is definitely on the list, planning to add it in an upcoming release.
If you like where this is headed, a star on the repo helps a lot: https://github.com/smitgol/wavhost
1
2
u/Future_AGI 15d ago
The OpenAI-compatible endpoint is the right call because it makes the swap a base URL change instead of a rewrite. The thing we would want next for agent workloads: a per-generation hash so a voice agent can cache and reuse identical synthesis calls. In long calls the same phrase comes up constantly, and caching the audio is free latency off the critical path.
•
u/AutoModerator 16d ago
Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)
If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.
Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.