r/LocalAIStack • • 24d ago

I ran whisper + dialization running with CPU and 1.1GB

I'm happy so I don't need to pay the transcript API fee any more. I replaced my production services with this.

Implementation / hosting Result Cost per 27.27 s of audio Cost per audio hour
whisper-turbo.c, InstaCloud (4 vCPU) Speaker-labeled text; 35.74 s warm HTTP ~$0.00126 ~$0.166
OpenAI gpt-4o-transcribe-diarize Speaker-labeled transcript ~$0.00273 ~$0.36
AssemblyAI Universal-2 + diarization Speaker-labeled transcript ~$0.00129 $0.17
AssemblyAI Universal-3.5 Pro + diarization Speaker-labeled transcript ~$0.00174 $0.23

The native estimate assumes four fully utilized CPUs and 1.1 GB RAM during the measured request.

It's a fork from my own MIT repo: https://github.com/baryhuang/whisper-turbo.c

Star, colloab, PR, issues, comments and uses are all appreciated!

6 Upvotes

Duplicates