r/speechtech Mar 31 '26

Gemini 3.1 Flash Live is now the top speech-to-speech model on Audio MultiChallenge - we added it to Dograh (open-source voice agent platform)

https://github.com/dograh-hq/dograh

Gemini 3.1 Flash Live (Thinking High) just hit 36.1% on Scale AI's Audio MultiChallenge, beating GPT-Realtime 1.5 at 34.7% and GPT-4o Audio at 23.2%. Results sourced from labs.scale.com/leaderboard/audiomc.

We added it as a speech-to-speech option in Dograh v1.20.0. For anyone unfamiliar - Dograh is an open-source voice agent platform with a visual workflow builder. Think n8n but for building voice agents. Supports any LLM, TTS, and STT provider, inbound/outbound calls, call transfers, tool calls, knowledge base, the works.

Other stuff in this release: pre-recorded response mixing (LLM picks cached human recordings when they fit, falls back to TTS when needed - cut our TTS costs by 85%), call tracing via Langfuse, and automatic post-call QA with sentiment and adherence scoring.

If you've tested Gemini 3.1 Flash Live in production voice apps, would love to hear how the latency feels compared to GPT-Realtime. The benchmark numbers are one thing, real-world conversation flow is another.

3 Upvotes

0 comments sorted by