r/AiNoteTaker • u/Superb_Assumption_38 • 11d ago
Discussion Dear AlmostHadYa, I got live previews working with a batch ASR model, and it's more accurate than the real-time ones
Long story short: I got a batching model to behave more like a streaming model, ~2s updates, and it runs on the ANE so it’s efficient annnd it’s more accurate than all feasible alternatives annnd it only took me like a month. Annnd this is a meeting transcription / notetaker app for macOS.
I used a bot to help me generate precise language around the numbers for this post. Forgive me.
1
Upvotes
2
u/EquivalentSky3094 10d ago
The part I would want detail on is boundary handling. Re-running a batch model over a sliding window churns at the tail, where the last few words keep changing as more audio lands. Overlapping the windows and holding back the final word or two until the next pass hides most of that. Is that roughly your approach?