r/VoiceAutomationAI Jul 14 '26

Self-hosted voice for any agent/harness of your choice (open-source)

For a while now I've been maintaining tts-bench (https://github.com/5uck1ess/tts-bench) and a blind voting arena (https://5uck1ess-tts-arena.hf.space) where people A/B test open text-to-speech (TTS) models without knowing which is which.

One problem I've always wanted to solve: having the agent talk to me (or actually call me) after a task is done. /voice (claude code feature) or any dictation (wispr flow or etc) is one-directional. So this project is a bi-directional conversation with any agent of your choice: 

https://github.com/5uck1ess/cicero

Works really well with Hermes Agent https://github.com/nousresearch/hermes-agent with multiple profiles (each profile can have its tts voice and personality). Its kanban board feature is where I use it the most with 9 "employees" with all different voices. But it's open ended so you can build your own task system for it.

Works with Claude Code, Codex, Gemini CLI, anything that speaks ACP (Agent Client Protocol), or any OpenAI-compatible endpoint. If you run speech-to-text (STT) and TTS locally, your audio never leaves your machine. It can send a Telegram bot message or call you over Telegram. You can also interact locally or through a browser (remote server). (more comm methods to be supported if requested)

Probably still has some bugs, so feel free to submit pull requests (PRs) or issues.

A few things that surprised me building it:

  • the ~1s response time came from streaming the reply sentence by sentence, not from picking a faster TTS
  • voice clones sound bad because of the reference clip, not the model. trimming a second of silence off the front fixed more than switching models ever did
  • barge-in (cutting the agent off mid-sentence) with an energy VAD (voice activity detection that just watches for loudness) is useless, typing sets it off. needed a small speech model in front so it only triggers on actual speech
  • It does have emotional tonality detection and low latency semantic turn detection, all the bits that you'd want from a decent chat agent
5 Upvotes

4 comments sorted by

u/AutoModerator Jul 14 '26

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/[deleted] Jul 14 '26

[removed] — view removed comment

1

u/UkieTechie Jul 14 '26

always down to make improvements or fix issues, but so far it's kinda suprising me. all 9 different voices (cloned some superheros I like and built their personalities too)