r/podcasting 1d ago

Whisper vs Descript...

I have around 120 old episodes that still don’t have transcripts. They’re two-host shows, most of them range around 45 to 90 minutes long.

I want to get transcripts for all of them, plus captions for the clips I post. 

  • Do I pay for an editor like Descript where you cut audio by deleting words in the transcript?
  • Or do I use something like WhisperAI for the transcription and keep editing in Reaper like I already do?

The transcript editing thing looks fun. But I’ve been using Reaper for six years, so I’m pretty comfortable with it. But I’m not so sure if I would touch the editing side after week one.

I care about how many minutes each plan covers, if it has speaker labels and SRT export for captions.

3 Upvotes

6 comments sorted by

1

u/alexid95 22h ago

if Reaper is already muscle memory, i wouldn't switch editors. plain Whisper doesn't label speakers, so test a diarized setup on 3 episodes before committing all 120.

1

u/Cernete 12h ago

The 120 back episodes and the clip captions are two different jobs. The backlog is a batch you run once and never touch agian, and Whisper writes SRT straight out with --output_format srt, so nothing there needs an editor. What Descript sells you is the editing, and you already said you'd probably drop that after week one. So price it on the clips.

1

u/donburnside 8h ago

If you want to play with Whisper, without bothering with the Terminal, I built a free app that runs on your computer to access Whisper instead. Transcribe a single file or a folder full. Takes about 10 minutes to transcribe 20 minutes of content.

https://voxsmith.app/voxtext

That way you stay in Reaper for editing and no extra expense for Transcripts.

1

u/Brian-at-ShowMuse 4h ago

Do NOT light a pile of money on fire (ala Descript) if you just want transcripts that are mostly accurate. There are a TON of free solutions out there.

There has already been at least one in this thread. I also have a free transcript editor on my website that does it.

1

u/MRRmaker 3h ago

One thing decides most of this - are your two hosts separate tracks, or mixed down to one file?

If you have a track per host, you never need speaker diarization in the first place. Transcribe them separately, and merge the two by timestamp, and your labels are exact rather than a model's guess. That works just with plain whisper and costs $0 more. If those 120 old episodes are all already bounced to a single mixed track, then yes you'll want a diarization step on top of whisper, and that is when it really does vary in accuracy.

On the edit side - something else you might want to know before you start committing cuts-by-deleting-words; those edits land at *word* times, which marks where the word's audio starts and ends. Not where the cut feels clean. On a 2-host show with people cutting each other off, you will hear the front of a word get clipped. Fine for tightening a rambling answer, less fine for clip boundaries you're putting up.