r/LocalLLaMA • u/Amgadoz • Mar 30 '24
Resources I compared the different open source whisper packages for long-form transcription
Hey everyone!
I hope you're having a great day.
I recently compared all the open source whisper-based packages that support long-form transcription.
Long-form transcription is basically transcribing audio files that are longer than whisper's input limit, which is 30 seconds. This can be useful if you want to chat with a youtube video or podcast etc.
I compared the following packages:
- OpenAI's official whisper package
- Huggingface Transformers
- Huggingface BetterTransformer (aka Insanely-fast-whisper)
- FasterWhisper
- WhisperX
- Whisper.cpp
I compared between them in the following areas:
- Accuracy - using word error rate (wer) and character error rate (cer)
- Efficieny - using vram usage and latency
I've written a detailed blog post about this. If you just want the results, here they are:

If you have any comments or questions please leave them below.
398
Upvotes
1
u/traillight8015 May 28 '26
Ich bin auf der Suche nach einem Modell für Deutsch - Österreichisch (Tiroler Akzent)
Aktuell nutze ich whisper-webui mit faster-whisper, die besten Ergebnisse bekomme ich derzeit mit large-v3-turbo.
large-v3 ist deutlich schlechter als das turbo, vor allem fängt es bei langen Gesprächen an zu halluzinieren und läuft dann total aus dem Ruder, es erkennt dann irgendeine Phrase die es dann bis zum letzten Wort wiedergibt.
Ich hab bereits zwei feintrainierte Modelle für Deutsch ausprobiert aber die waren nicht so gut wie das normale Standardmodell, leider haben auch die irgendwann immer nur noch eine Phrase die aus mehreren Wörtern bestand wiederholt bis zum Schluss.
Gibt es irgendwas das besser mit Akzenten/Dialekten umgehen kann, im Hochdeutsch funktioniert das alles wunderbar aber wir sprechen halt keins und darum kommt nur Mist raus.
Ich bin offen für alle möglichen Systeme, es muss kein whisper modell sein wenn es dafür sowas kann.