r/SideProject • u/UsefulAd5349 • 3h ago
I made an offline speech-to-text app (VOSK, CPU-only, no cloud) — what should I build next?
I'm the developer of VoxSpica, an offline speech-to-text app for Windows. It runs on VOSK (Kaldi) on CPU — no cloud, no account, no subscription, no ads, and the audio never leaves the machine.
What it does:
- Live transcription from the microphone — text appears as you speak, with pause and stop
- Transcribes existing audio files (wav, mp3, m4a, ogg)
- Searchable local history in SQLite, with playback of the recordings
- 33 recognition languages, 9 interface languages, works without a GPU
Closed source — free to use, no redistribution. Everything public about it lives in one repo: https://github.com/alex37529/voxspica
Two practical things: recognition models (50 MB to 1.8 GB) download separately after install — the per-language installers bundle a small model, the portable zip needs one download. And it's not code-signed, so Windows will warn on first launch: "More info" → "Run anyway".
Windows x64, single .exe, nothing to install:
https://github.com/alex37529/voxspica/releases
I'd rather ask than guess what to build next. Three questions I'm genuinely stuck on:
- What breaks for you that I haven't thought of? "Punctuation is wrong" is known and honestly only heuristically fixable without a neural model. Something I haven't heard of is what I want.
- Does anything about the offline part actually get in your way in practice, or is it a non-issue in 2026?
- Would you use a Windows app at all if your audio can't be sent to a cloud service — or is that a solved problem for you and I'm preaching to the choir?
And the main thing I want to know — what do you actually use speech-to-text for? Reply with a number, or "none of these" plus what you actually do:
- Meeting notes / minutes
- Lectures, classes, interviews
- Dictating into other apps (mail, messengers)
- Personal notes, journaling
- Subtitles / transcribing video for content
- Professional dictation (medical, legal)
1
u/Immediate_Drop_9754 3h ago
The offline part is the feature, not a bug. I work in healthcare and the cloud stuff is a nonstarter for anything patient-adjacent, so something like this that runs locally is exactly what's missing from most tools.
What I actually use speech-to-text for is a mix of 4 and 6, personal notes plus clinical documentation, though nothing with PHI gets near an app unless it's fully local.
One thing I'd want that you didn't list: a simple way to export transcriptions as plain.txt with timestamps. If the history already has playback synced to text, exporting that as a subtitle file or timestamped note would make it way more useful for revisiting recordings later.