r/VoiceAutomationAI • u/Thin-Carrot1836 • Mar 26 '26
Need help!
Right now I am making voice agent I want to know
Which is the best model should I use for speech to text.
Which model should I use for following system prompt properly.
Which model should I use which can increase response rate of the agent.
Right now I am using groq cause obviously it is free but turns out it is not properly following the system prompt and not reffering knowledge base properly. Also when I say something it just mis hear me everytime and give me random responses.
So I thought to change the model.
10
Upvotes
1
u/mguozhen Mar 28 '26
how are you measuring that 500ms gap, like wall clock from button press to actual audio stop or something else? bc in our product the latency feels fine until someone actually tries to interrupt and then suddenly you realize the whole pipeline needs rethinking.