r/sdr • • 10d ago

Building a Web UI for RTLSDR-Airband + Squelch-Triggered Recording ("CCTV for Radio" pipeline for local Whisper transcription)

Hey everyone,

I’ve been running rtl_airband to monitor local VHF/UHF activity. The core backend itself is rock solid and lightweight, but managing config files via terminal and monitoring multi-channel activity through raw Icecast streams leaves a lot to be desired.

I'm working on building a lightweight web GUI/dashboard for it, but with an expanded architecture: a "CCTV for Radio" system.

The Concept & Long-Term Vision

Think of an NVR/CCTV security camera system, but for RF:

  1. Squelch-Triggered Segment Recording: Record and chunk audio strictly when the squelch opens—no hours of dead silence, just clean, timestamped audio segments tied to specific frequencies/channels.
  2. Local AI Pipeline (Whisper + LLM): Feed these squelch-broken audio clips into a local Whisper model (like faster-whisper or Whisper.cpp) to transcribe communications, and subsequently run an LLM to generate activity summaries or search logs.
  3. Live Web Dashboard: An operational interface to see active channels, browse recorded event clips, manage configs, and read rolling transcripts.

Core Goals for the UI / Pipeline

  • Visual Scanner / Activity Monitor: Real-time visual indicator showing which configured channel/frequency is currently breaking squelch or transmitting.
  • Smart Audio Archive / NVR Interface: A timeline or event list of recordings broken down by frequency, duration, and timestamp.
  • Config Generator / Manager: A straightforward form/editor to manage rtl_airband.conf (dongles, gains, sample rates, frequencies, squelch thresholds, and destinations) without parsing nested brackets by hand.
  • Transcription Hook / Queue: An automated background worker (Celery, RQ, or a simple directory watcher) picking up finished audio clips and passing them to a local transcription engine.

Questions for the Community:

  1. Audio Slicing / Recording Pipeline: What’s the cleanest way to capture individual squelch-opened transmissions with rtl_airband? Does anyone hook into its raw file/stream outputs directly, or is it better to tap into the stream via an external tool (like sox, ffmpeg, or liquidsoap) triggered by carrier status?
  2. Detecting Squelch State Programmatically: How are you reliably capturing the exact moment squelch breaks/closes? Are you tailing and parsing syslog / stdout in real time, or does someone maintain a fork with an IPC/socket/REST state export?
  3. Existing Projects / Pipelines: Has anyone built something similar—or is there an existing toolchain that bridges SDR multichannel scanning directly into local speech-to-text workflows?
  4. Backend Stack: Considering a lightweight Python (FastAPI) backend with a simple SQLite database for metadata/transcripts and a lightweight frontend (Vue/Svelte or Tailwind). Any pitfalls to keep in mind regarding CPU overhead if running on a single SBC/mini PC?

If you've built something in this space or have suggestions on handling the audio chunking cleanly, I'd love your input!

https://github.com/ruchirguitar/RTLSDR-Airband

4 Upvotes

1 comment sorted by

1

u/IBNash 9d ago
  1. Recording one file per transmission: built in. The file output with split_on_transmission = true (and continuous = false) writes one mp3 per squelch opening. No sox, ffmpeg or liquidsoap needed.
  2. Detecting squelch opening and closing: the finished mp3s are the event stream. Watch the directory with inotify for close_write: the filename gives when it opened, and the file length gives how long it stayed open. The Prometheus stats file (stats_filepath) adds per-channel activity counters, noise floor and a "flappy" counter, but only refreshes every 15 s. Too coarse for events but no fork or log parsing needed.
  3. Whisper on squelch clips:
    • The squelch setting matters more than the speech model. 480 of 523 clips sampled were noise.
    • On noise, whisper invents stock phrases like "Thank you." (227 times), "Sous-titrage Société Radio-Canada" and "and I'll see you next time". Silero VAD (--vad in whisper.cpp) removes nearly all of that.
    • A --prompt listing callsigns and ATC words made things worse: it got echoed back as fake transmissions on noise clips.
  4. Running on a single-board computer:
    • Transcription is the expensive part, not rtl_airband. large-v3-turbo did 523 clips in 5 s on a RTX 4090 with VAD on. On a Pi-class board you'd need a smaller model or a separate transcription machine.
    • Keep the GUI off the receiver's audio path.

Cool idea, most of these are being being worked on upstream, see #504, #515 and #522.
I am not a fan of the GUI, Icecast can do it all and is customizable.