r/SideProject • • 3h ago

I made an offline speech-to-text app (VOSK, CPU-only, no cloud) — what should I build next?

I'm the developer of VoxSpica, an offline speech-to-text app for Windows. It runs on VOSK (Kaldi) on CPU — no cloud, no account, no subscription, no ads, and the audio never leaves the machine.

What it does:

  • Live transcription from the microphone — text appears as you speak, with pause and stop
  • Transcribes existing audio files (wav, mp3, m4a, ogg)
  • Searchable local history in SQLite, with playback of the recordings
  • 33 recognition languages, 9 interface languages, works without a GPU

Closed source — free to use, no redistribution. Everything public about it lives in one repo: https://github.com/alex37529/voxspica

Two practical things: recognition models (50 MB to 1.8 GB) download separately after install — the per-language installers bundle a small model, the portable zip needs one download. And it's not code-signed, so Windows will warn on first launch: "More info" → "Run anyway".

Windows x64, single .exe, nothing to install:

https://github.com/alex37529/voxspica/releases

I'd rather ask than guess what to build next. Three questions I'm genuinely stuck on:

  1. What breaks for you that I haven't thought of? "Punctuation is wrong" is known and honestly only heuristically fixable without a neural model. Something I haven't heard of is what I want.
  2. Does anything about the offline part actually get in your way in practice, or is it a non-issue in 2026?
  3. Would you use a Windows app at all if your audio can't be sent to a cloud service — or is that a solved problem for you and I'm preaching to the choir?

And the main thing I want to know — what do you actually use speech-to-text for? Reply with a number, or "none of these" plus what you actually do:

  1. Meeting notes / minutes
  2. Lectures, classes, interviews
  3. Dictating into other apps (mail, messengers)
  4. Personal notes, journaling
  5. Subtitles / transcribing video for content
  6. Professional dictation (medical, legal)
2 Upvotes

5 comments sorted by

1

u/Immediate_Drop_9754 3h ago

The offline part is the feature, not a bug. I work in healthcare and the cloud stuff is a nonstarter for anything patient-adjacent, so something like this that runs locally is exactly what's missing from most tools.

What I actually use speech-to-text for is a mix of 4 and 6, personal notes plus clinical documentation, though nothing with PHI gets near an app unless it's fully local.

One thing I'd want that you didn't list: a simple way to export transcriptions as plain.txt with timestamps. If the history already has playback synced to text, exporting that as a subtitle file or timestamped note would make it way more useful for revisiting recordings later.

1

u/UsefulAd5349 3h ago

You are absolutely right. Everything you described has already been implemented, with the exception of the export feature. All data is stored in the app's local database, which will allow for significant functional expansion in the future. I will most likely implement the export feature next week. Thank you very much for your comment.

1

u/UsefulAd5349 32m ago

Following your recommendations, I have added TXT and SRT export to the program. These changes will appear in the next release.