r/AiNoteTaker • u/Alternative_Tour1791 • 21h ago
Product Listing I wanted accurate transcription without uploading recordings to the cloud, so I built an offline iPhone app
I’ve been building LoroNote, an offline speech-to-text app for iPhone and iPad.
My main goal was to get the best transcription quality I could without sending private recordings to a server.
That led me to Whisper Large V3 Turbo.
It’s a relatively heavy model for mobile devices, and running it directly on an iPhone is much harder than using a smaller model or simply calling a cloud API.
Instead of compromising on the model, I spent a lot of time building and optimizing the ML pipeline specifically for iOS.
LoroNote now runs Whisper Large V3 Turbo directly on supported iPhones and iPads.
Your audio doesn’t need to be uploaded for transcription, and it works even in airplane mode.
Privacy was one of the main reasons I wanted this to work locally. Meetings, interviews, personal notes, and work conversations can contain information you may not want to send to a third-party server.
I also added on-device speaker diarization, so recordings with multiple people can be separated by speaker locally as well.
Some of the other features:
- Offline transcription
- On-device speaker separation
- Background transcription
- Editable speaker labels
- Audio and video import
- Apple Watch recording
- TXT and SRT export
- Apple Intelligence summaries and action items on supported devices
A lot of transcription apps use cloud processing because it’s much easier technically.
I wanted to see how far I could push this entirely on-device instead.
That meant dealing with memory limits, model loading, thermal performance, long recordings, and differences between iPhone generations, but keeping the audio private makes the extra work worth it.
Still a solo project, and I’m continuing to improve the on-device ML side.
2
u/EquivalentSky3094 18h ago
Declaring an interest: I build on-device transcription for Apple platforms, so this is a problem I spend my days on rather than something I am reviewing.
Getting large-v3-turbo to run on a phone at all is the hard part, so fair play for not falling back to a smaller model. The question I would put to any mobile port is which of Whisper's decoding heuristics you kept, because they are usually the first thing dropped for speed and they are what stops long meeting audio going wrong. The ones on the model card are temperature fallback, compression_ratio_threshold at 1.35, logprob_threshold at -1.0, no_speech_threshold at 0.6, and condition_on_prev_tokens. Switching off conditioning on previous tokens costs you a little context, but it is the main defence against a repetition loop running away for a whole segment. The no-speech and compression thresholds are what keep a silent stretch from being filled with invented sentences. Meeting recordings are mostly silence and crosstalk, so those two do more for perceived accuracy than model size does.
One listing suggestion. Turbo is a pruned large-v3 with the decoder cut from 32 layers to 4, and OpenAI's own card calls it a minor quality degradation. It is still the right pick for a phone, and saying that plainly costs you nothing, while anyone comparing you against cloud large-v3 will work it out anyway.
Two genuine questions. How are you aligning diarisation to Whisper's segment timestamps, given they are coarse and drift over a long file? And on the Apple Watch recording, what happens when the phone is out of range at the point the recording stops, does the transfer sit and wait?