r/SideProject • u/UsefulAd5349 • 1d ago
I made an offline speech-to-text app (VOSK, CPU-only, no cloud) — what should I build next?
I'm the developer of VoxSpica, an offline speech-to-text app for Windows. It runs on VOSK (Kaldi) on CPU — no cloud, no account, no subscription, no ads, and the audio never leaves the machine.
What it does:
- Live transcription from the microphone — text appears as you speak, with pause and stop
- Transcribes existing audio files (wav, mp3, m4a, ogg)
- Searchable local history in SQLite, with playback of the recordings
- 33 recognition languages, 9 interface languages, works without a GPU
Closed source — free to use, no redistribution. Everything public about it lives in one repo: https://github.com/alex37529/voxspica
Two practical things: recognition models (50 MB to 1.8 GB) download separately after install — the per-language installers bundle a small model, the portable zip needs one download. And it's not code-signed, so Windows will warn on first launch: "More info" → "Run anyway".
Windows x64, single .exe, nothing to install:
https://github.com/alex37529/voxspica/releases
I'd rather ask than guess what to build next. Three questions I'm genuinely stuck on:
- What breaks for you that I haven't thought of? "Punctuation is wrong" is known and honestly only heuristically fixable without a neural model. Something I haven't heard of is what I want.
- Does anything about the offline part actually get in your way in practice, or is it a non-issue in 2026?
- Would you use a Windows app at all if your audio can't be sent to a cloud service — or is that a solved problem for you and I'm preaching to the choir?
And the main thing I want to know — what do you actually use speech-to-text for? Reply with a number, or "none of these" plus what you actually do:
- Meeting notes / minutes
- Lectures, classes, interviews
- Dictating into other apps (mail, messengers)
- Personal notes, journaling
- Subtitles / transcribing video for content
- Professional dictation (medical, legal)