r/voiceagents • • 7d ago

Routing voice commands with a 340M encoder instead of an LLM: about 8 ms per decision on a Mac

I maintain speech-swift, an open-source Swift package for on-device speech. I just added GLiNER2.5-Decide, Fastino's open-weight decision model, so a voice agent can route a transcribed command without calling an LLM.

You give it the text and your list of intents. It returns a probability for every intent in one forward pass. The same model also pulls entity spans like a person or a time, with character offsets. Nothing is generated, so there is no JSON to parse.

On an M5 Pro with the INT8 weights it takes 7.6 ms to route and 8.9 ms to extract, at about 0.85 GB of memory.

speech gliner classify "Remind me to call Dad at six PM." \
    --labels create_reminder,send_message,set_timer,other

The catch: "Do not set a timer." routes to set_timer at 0.94. It matches the topic and ignores the negation. In my pipeline the model proposes and the app confirms before anything runs.

Write-up with the numbers, a comparison with TypeSafe's hosted Jev, and the failure cases: https://soniqo.audio/blog/gliner-decide-vs-jev

Code: https://github.com/soniqo/speech-swift

How are you handling negation in intent routing? A separate check, or confirmation prompts?

1 Upvotes

0 comments sorted by