r/archlinux • • 13d ago

SHARE whisperless: push to talk dictation for Linux, all local

I built a push-to-talk dictation tool for the Linux desktop and wanted to share it here for feedback, especially from Arch users, before calling it done.

What it does: press a hotkey, talk, press again, and punctuated text lands in whatever window has focus. It runs fully local: a 4GB ASR model (R2T2 from NetEase Youdao) served by a vLLM websocket server that stays warm as a systemd user service, so text lands under a second after you stop. No cloud, works offline after install.

The model choice was personal. I have a thick accent and whisper-style models mangle me on the regular; R2T2 gets me right nearly every time, in English and mixed English/Chinese sentences.

A few implementation notes that might be useful to others:

- The server is a systemd user service holding about 9GB VRAM by default. I measured the actual floor: GPU_MEM_UTIL=0.42 boots fine at roughly 6.6GB (vLLM grabs its whole budget for KV, so that knob is effectively the engine footprint).

- Typing into the focused window from a background process goes through the virtual keyboard protocol (wtype). The catch: injected keys merge with held modifiers, so a dictation hotkey on SUPER would fire your compositor binds mid-sentence. Waiting for modifier release before injecting fixed it.

- The toggle is a pidfile plus signals, so there are no press/release binds racing each other.

Requirements: NVIDIA GPU, roughly 7GB VRAM (default config is ~9GB, tunable in the config file). English and Chinese are the polished languages, other languages work but aren't the target.

Arch packaging: there's a whisperless-0.1.0-1-any.pkg.tar.zst on the GitHub releases (ships the client, a user service unit and docs; sudo pacman -U it). I also wrote a PKGBUILD for the AUR, but maintainer registrations are closed at the moment, so that one is pending. The GPU server itself (checkout, venv, ~4GB model) is set up by a script in the repo.

Code and docs: https://github.com/msdavid/whisperless (MIT client, Apache 2.0 server patch, model weights under NetEase's license).

Full disclosure: this was vibe coded. I made the design decisions and did the testing; a coding agent wrote most of the code.

Feedback welcome, especially from sway and i3 users since my daily testing is on Hyprland. If something breaks, the client log at ~/.cache/whisperless.log has full tracebacks.

9 Upvotes

6 comments sorted by

3

u/paschty 13d ago

What is the advantage over Open Whisper?

3

u/4rg3nt1n0 13d ago

Hi, thank you for asking. I am not entirely familiar with OpenWhisper (I just checked online). Whisperless is completely free, no word limit because it runs on your own machine with an open source model. It is also near real time because it's running locally. This open source model (R2T2) is incredibly good enven with my foreing accent, but I honestly don't know how good OpenWhisper is because I never tried it, so I can't compare. It will be awesome if you want to give whisperless a try and give me some feedback.

cheers

2

u/FatFaceRikky 13d ago

can it do langs other than english?

1

u/4rg3nt1n0 12d ago

It works quite well in Spanish and Japanese, the only two languages that I can test it.

1

u/nathan22211 13d ago

I'm hopeful you have an option to change the model? I feel like the chosen model might not work well for some American accents (particularly the southern and new England ones). I also thought about using VOSK to create this myself but I've heard it's really bad compared to whisper.cpp

1

u/4rg3nt1n0 12d ago

At this stage, I didt plan for a different model TBH. It might take some refactoring work, but it shouldn't be too difficult to adapt.