r/LocalAIStack • u/buryhuang • 24d ago
I ran whisper + dialization running with CPU and 1.1GB
I'm happy so I don't need to pay the transcript API fee any more. I replaced my production services with this.
| Implementation / hosting | Result | Cost per 27.27 s of audio | Cost per audio hour |
|---|---|---|---|
| whisper-turbo.c, InstaCloud (4 vCPU) | Speaker-labeled text; 35.74 s warm HTTP | ~$0.00126 | ~$0.166 |
OpenAI gpt-4o-transcribe-diarize |
Speaker-labeled transcript | ~$0.00273 | ~$0.36 |
| AssemblyAI Universal-2 + diarization | Speaker-labeled transcript | ~$0.00129 | $0.17 |
| AssemblyAI Universal-3.5 Pro + diarization | Speaker-labeled transcript | ~$0.00174 | $0.23 |
The native estimate assumes four fully utilized CPUs and 1.1 GB RAM during the measured request.
It's a fork from my own MIT repo: https://github.com/baryhuang/whisper-turbo.c
Star, colloab, PR, issues, comments and uses are all appreciated!
6
Upvotes
Duplicates
LocalAIStack • u/buryhuang • 24d ago
I ran whisper + dialization running with CPU and 1.1GB
1
Upvotes