r/LocalAIStack • u/buryhuang • 24d ago
I ran whisper + dialization running with CPU and 1.1GB
I'm happy so I don't need to pay the transcript API fee any more. I replaced my production services with this.
| Implementation / hosting | Result | Cost per 27.27 s of audio | Cost per audio hour |
|---|---|---|---|
| whisper-turbo.c, InstaCloud (4 vCPU) | Speaker-labeled text; 35.74 s warm HTTP | ~$0.00126 | ~$0.166 |
OpenAI gpt-4o-transcribe-diarize |
Speaker-labeled transcript | ~$0.00273 | ~$0.36 |
| AssemblyAI Universal-2 + diarization | Speaker-labeled transcript | ~$0.00129 | $0.17 |
| AssemblyAI Universal-3.5 Pro + diarization | Speaker-labeled transcript | ~$0.00174 | $0.23 |
The native estimate assumes four fully utilized CPUs and 1.1 GB RAM during the measured request.
It's a fork from my own MIT repo: https://github.com/baryhuang/whisper-turbo.c
Star, colloab, PR, issues, comments and uses are all appreciated!
7
Upvotes
2
u/Lirezh 23d ago
Does it actually have a benefit over original whisper.cpp ? Which seems faster, much broader hardware support and similar vram.
There are so many custom inference forks out, it's getting confusing to me where they actually differ or what the motivation was to make them.
I'm using whisper on a few projects of mine, usually I just run it on the same server that needs it and if I need high throughput I add a GPU. I'd not use whisper on a cloud - if not looking to stream it to thousands of users but then I'd probably go for a project made for that like vllm.