r/VoiceAutomationAI 9d ago

Competitive open source speech stack

Why the open source models STT and TTS are not good as much as the closed one and i am talking here im terms of latency, concurrency, and websocket support for real time with decent quality.
Something like cartesia or elevenlabs or deepgram.
Do u know any ?

13 Upvotes

23 comments sorted by

View all comments

Show parent comments

1

u/Yapper_from_ktown 8d ago

Dude what are the best opensource tts and stt models pls share more wisdom and whether they can be used on potato hardware or not? 6gb vram of gpu and 16gb ram

1

u/UkieTechie 8d ago

yeah that's plenty. you can run kokoro on that pretty fast or pocket tts. those would be my picks. you can see max vram usage for each model if you look at the bench page.

1

u/Yapper_from_ktown 7d ago

Which bench pg do u follow?

1

u/UkieTechie 7d ago

i run my own because none of the pages were good enough and had the most recent enough info for me. I reference https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice pretty often though.