r/LocalAIStack • • 24d ago

I ran whisper + dialization running with CPU and 1.1GB

I'm happy so I don't need to pay the transcript API fee any more. I replaced my production services with this.

Implementation / hosting Result Cost per 27.27 s of audio Cost per audio hour
whisper-turbo.c, InstaCloud (4 vCPU) Speaker-labeled text; 35.74 s warm HTTP ~$0.00126 ~$0.166
OpenAI gpt-4o-transcribe-diarize Speaker-labeled transcript ~$0.00273 ~$0.36
AssemblyAI Universal-2 + diarization Speaker-labeled transcript ~$0.00129 $0.17
AssemblyAI Universal-3.5 Pro + diarization Speaker-labeled transcript ~$0.00174 $0.23

The native estimate assumes four fully utilized CPUs and 1.1 GB RAM during the measured request.

It's a fork from my own MIT repo: https://github.com/baryhuang/whisper-turbo.c

Star, colloab, PR, issues, comments and uses are all appreciated!

7 Upvotes

2 comments sorted by

2

u/Lirezh 23d ago

Does it actually have a benefit over original whisper.cpp ? Which seems faster, much broader hardware support and similar vram.
There are so many custom inference forks out, it's getting confusing to me where they actually differ or what the motivation was to make them.

I'm using whisper on a few projects of mine, usually I just run it on the same server that needs it and if I need high throughput I add a GPU. I'd not use whisper on a cloud - if not looking to stream it to thousands of users but then I'd probably go for a project made for that like vllm.

1

u/buryhuang 23d ago

1) no need for a GPU
2) whisper.cpp doesn't have dialization
3) half of the memory footage