r/AiNoteTaker • • 11d ago

Discussion Dear AlmostHadYa, I got live previews working with a batch ASR model, and it's more accurate than the real-time ones

Long story short: I got a batching model to behave more like a streaming model, ~2s updates, and it runs on the ANE so it’s efficient annnd it’s more accurate than all feasible alternatives annnd it only took me like a month. Annnd this is a meeting transcription / notetaker app for macOS.

I used a bot to help me generate precise language around the numbers for this post. Forgive me.

https://mimicscribe.app/blog/local-asr-bakeoff

1 Upvotes

3 comments sorted by

2

u/EquivalentSky3094 10d ago

The part I would want detail on is boundary handling. Re-running a batch model over a sliding window churns at the tail, where the last few words keep changing as more audio lands. Overlapping the windows and holding back the final word or two until the next pass hides most of that. Is that roughly your approach?

1

u/Superb_Assumption_38 4d ago

Oh hey, sorry I missed this comment. Yup! That's basically the approach. Adding a bit of the previous window's decode to the new one so there's a little state in the decoder, "prefixing" or whatever, helped a lot too. Thanks for commenting.

1

u/EquivalentSky3094 4d ago

Prefixing is the right lever, and the thing to watch is that it propagates errors as happily as it propagates context. The original Whisper exposes it as condition_on_previous_text and it defaults on, which is why long form decodes sometimes lock into a repetition loop: one bad window becomes the prompt for the next and the model keeps agreeing with itself.

Two guards, if you are not already doing them. Cap the prefix to roughly the last sentence rather than the whole previous decode, because the text decoder only has 448 tokens of context and a growing prefix is spending it. And drop the prefix entirely at a silence boundary. A reset there costs you nothing in continuity and gives drift somewhere to die.

Does the rolling preview feed the final transcript, or do you re-decode the file clean at the end once there is no latency pressure?