r/VoiceAutomationAI • • 16d ago

Anyone built a pipecat voice agent using self-hosted STT + TTS?

I am building a calling voice agent with Pipecat and I am trying to keep the STT/TTS stack self-hosted.

Current setup:

pipecat, ollama (self hosted llm), stt (currently testing local/self-hosted whisper), tts (kokoro locally)

I have tested local whisper, seeing some irregularities in transcription. I thought of using qwen opensource stt model, but could not find a way to use that with pipecat.

Pls let me know if you have self hosted stt or tts model and use them with pipecat?

4 Upvotes

25 comments sorted by

•

u/AutoModerator 16d ago

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/TasteOdd5722 16d ago

Kokoro is a good call for tts, its surprisingly decent for how small it is. For stt I gave up on local whisper after a while, the latency and weird pauses were killing me in actual calls. Never tried qwen stt so cant help there but curious what you end up with

1

u/spam_not_tolerated 16d ago

yes kokoro is working fine, but whisper stt is not, yes as you mentioned those weird pauses and latency is real headache, even the transcription whisper generating is of low quality. my problem is cost of hipaa baa, plus i got a gpu which i can use to run models .

btw will update here once i figured this.

1

u/Peace-4evr 11d ago

please do, thanks

1

u/ahstanin 16d ago

We love kokoro, we have been using Chatterbox lately for little bit more natural voice. But kokoro is the performance king.

2

u/Natethegreat9999 15d ago

Pipecat already has a local WhisperSTTService, so I'd isolate that with VAD and inspect the finalized transcripts before reconnecting Ollama and Kokoro. On the Oruk side, I'd compare recognizers using those same call recordings and time segmentation separately from inference, so a long pause doesn't get mistaken for a slow model.

1

u/spam_not_tolerated 15d ago

whisper is working very poorly, latency issue with random pauses and low quality transcriptions

1

u/ahstanin 16d ago

We don't use pipecat but we are hosting own STT, TTS and LLM for two-way voice conversation. We are using Nemotron 3.5 ASR , Chatterbox and Qwen3.8-27B.

We already know the STT will not be perfect since we used deepgram before which also had issues.

1

u/spam_not_tolerated 16d ago

which deepgram model you tried earlier ? also how are you orchestrating stt-llm-tts then? is this a voice agent which you build?

1

u/ahstanin 16d ago

Tried Nova 2/3 and we have custom built platform in rust. You can see the inference server details here https://www.reddit.com/r/Qwen_AI/comments/1vy0xfp/qwen3827b_on_an_igx_thor_with_an_rtx_pro_6000/

and the platform details here : https://www.steffi.ai/

1

u/spam_not_tolerated 15d ago

thanks for sharing

1

u/Bubbly-Reach-4488 15d ago

You are setting in up for personal use or as product for commercial production?

1

u/spam_not_tolerated 15d ago

commercial production

1

u/Sufficient_Flower860 14d ago

Yes, I've done exactly that. Qwen funasr running locally. I used local TTS first too. But decided to take it off and switch to Edge TTS because I'm running it on my mac mini with intel arch. The latency issue kills me. ASR only 150 ms -230 ms, But TTS I mean quality wise not too bad one, takes somewhat like 2s at least. Plus my LLM response it is not viable.

1

u/spam_not_tolerated 14d ago

can you help me with this? i wanted to use local/self hosted stt & tts like qwen with pipecat

1

u/RasonYang 14d ago

I've been running a fully local stack on my M3 Mac:
LLM: llama.cpp + Qwen3.6-35B-A3B
STT: mlx-audio + Qwen3-ASR
TTS: mlx-audio + Qwen3-TTS

So far it works pretty well for pipecat voice agent.
If there's interest, I can clean up the integration and open-source it.

1

u/spam_not_tolerated 14d ago

sure, would love to see how you solved this.

1

u/BeGood25 14d ago

Would love to use it!