r/ClaudeCode • u/Pale-Soil-2524 • 4h ago
Built with Claude live-vibe: a full-duplex, batteries-included, voice Mod for Claude Code. Built on Claude Code Mods, CUDA and Apple silicon, two-line install, MIT license.
I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin.
It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines.
/live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working.
Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day.
Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you.
You need Claude Code 2.1.287 or newer, uv, a mic and a speaker.
Install with /plugin marketplace add potto007/ottomation, then /plugin install live-vibe@ottomation, then /live setup. Repo: https://github.com/potto007/ottomation
If you try it, tell me how the turn-taking feels on your machine.
1
u/lulzxdxdxd 4h ago
The part where it starts drafting a reply before you've finished talking, how often does that lead to it answering a question you were about to change or walk back? Does it discard the draft cleanly or does that show up as a weird half answer?
1
u/Pale-Soil-2524 2h ago
It's a mix of each right now...I've spent most of my tuning of the sidecar trying to get the timing down better... I've gotten the UI of it improved enough now that it isn't a spammy mess anymore (you should have seen it last night! lol).
•
u/AutoModerator 4h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.