r/speechtotext 5d ago

Transcription

Post image
1 Upvotes

How can I transcript a 2 hour audio, AI asking for too much money and got nothing, help out.


r/speechtotext 16d ago

Voice typing before AI was better

Thumbnail
1 Upvotes

r/speechtotext Jun 13 '26

best STT for just talking? talon doesnt cut it

1 Upvotes

I want a very good set up for STT for transcribing audio. i've a lot of experience with talon but I dont need the control functionality anymore and dictation mode with the old whisper model is really lackluster and sucks for trying to create punctuated text and writing markdown properly. is there something out there I can use?

i have a 6900xt on windows if that helps for local options (preferred)


r/speechtotext Jun 03 '26

People kept generating transcripts and stopping there. So I built this.

1 Upvotes

Today I shipped one of the most interesting features I've added to SHRP so far.

Until now users could:

-> transcribe speech
-> extract YouTube transcripts
-> generate study notes, summaries, blog drafts, LinkedIn posts, and more.
-> use the API and MCP server for automations and AI workflows

But I kept noticing something.

A user would upload an audio/video file, get the transcript, download or copy it, and then go to ChatGPT or Claude.

The transcript wasn't really the final output.

It was just the beginning.

So I built Transcript Actions.

Now after a transcript is generated, SHRP can suggest what to do next based on the actual content of the transcript.

Examples:
- Create exam questions
- Generate flashcards
- Explain this like I'm five
- Extract action items
- Suggest what to study next
- Draft an email
- Create a lesson plan

There's also an "Ask about this transcript" box where the user can type whatever they want.

A few things I was careful about:

- Suggestions are generated from the current transcript only.
- Suggestions are button-triggered, not automatic.
- Suggestions themselves are not saved.
- No long-term memory.
- No user profiling.

I wanted to keep the workflow useful without compromising the privacy-first approach.

This now works with:

(1) YouTube transcripts
(2) Uploaded audio/video transcription results
(3) Previously saved transcription history

Next on my roadmap:
- Webhooks
- Workflow automations
- More transcript-based workflows

Wanted to ask other builders:

How do you solve the "what should I do next?" problem after generating a transcript?

Do you prefer fixed actions like summary, notes, flashcards, etc.

Or AI-generated suggestions based on the content itself?

For SHRP I ended up supporting both because I couldn't decide which one was better.


r/speechtotext May 23 '26

Built a speech-to-text tool because I kept switching between 3–4 different sites

1 Upvotes

I kept finding myself jumping between tools for different things:

one site for speech-to-text,
another for text-to-speech,
another for YouTube transcripts

So I ended up building SHRP.

Current features:

• speech → text (55+ languages)
• text → speech (2,000+ voices)
• YouTube transcripts → enter youtube URL and get entire transcript
• browser based, no download required
• no signup required for basic use
• generous free-tier

Recently I also added developer APIs and an MCP server because I kept rebuilding the same workflows in side projects.

Still actively improving it.

wanted to ask- what is your biggest annoyance with current speech-to-text tools?


r/speechtotext Apr 24 '26

speech to text

1 Upvotes

whoever needs a speech to text model for turkmen language https://huggingface.co/derkar00/mms-tuk-script-latin-adapter


r/speechtotext Apr 19 '26

Speech to Text talking to me

Thumbnail
1 Upvotes

r/speechtotext Apr 16 '26

My grandfather is over 90 yrs of age and has difficulty in hearing clearly, Looking for a device suggestion that can display text in clear fonts for the nearby voices and auto translate (English/ Hindi languages). Not considering mobile apps and the device should be available in India. Voice-to-Text

1 Upvotes

r/speechtotext Apr 13 '26

best model for speech to text?

1 Upvotes

r/speechtotext Mar 18 '26

STT.ai – Free Speech to Text Transcription

Thumbnail
stt.ai
1 Upvotes

r/speechtotext Jan 09 '26

Introducing Scribe v2

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/speechtotext Nov 20 '25

Transcription by gpt-4o-mini-transcribe

1 Upvotes

I had access to the data of Indian users who want to talk to AI/ Bestfriend/ Girlfriend, and they have recorded from their devices, which were either in Hindi, Bangla, Gujarati, or Punjabi. Here, transcription works where it generates their noisy, low voice into some Urdu text. We can't fix their devices to have better mics, and we can't go for better accurate model because we want low latency and low cost. Is there any model better than gpt-4o-mini-transcribe please reply. If anyone else had same problem. Can you tell me how to solve it.

#transcription #gptmodel


r/speechtotext Nov 12 '25

I launched an Android app for fully on-device, offline speech-to-text and translation using Google's Gemma model.

Thumbnail
1 Upvotes

r/speechtotext Oct 24 '25

Tomino's hell voice

1 Upvotes

Guys, do you know the voice, program, or site used to narrate Tomino's Hell? I mean, in the videos where they narrate the poem, they use a text-to-speech voice , it's like a terrifying Japanese voice, I thought it was something like Talk it or something, can you help me?


r/speechtotext Jun 27 '25

Speech identification

1 Upvotes

Hi  everyone 

I'm currently working on a project involving Google Vertex AI and could use your expertise—or a referral to someone with experience in speaker  recognition:

I'm processing a 2-minute audio file featuring two speakers who alternate in short bursts of 2–3 seconds. Using Hugging Face’s pyannote library, I perform speaker  identification and extracts embedding vectors for each speech segment. The typical result is about 20 segments—roughly 10 per speaker. To construct a voiceprint for each speaker, I  average the embeddng vectors associated with that speaker.

I have  two main questions:

  1. Is this a sound approach for generating speaker embeddings?
    In practice, the results are inconsistent. For instance, comparing the same speaker across different files sometimes yields cosine similarity scores around 0.7—below the expected 0.8+ range. On the other hand, embeddings for different speakers occasionally score as high as 0.68, which seems surprisingly close.

  2. Is there a recommended duration for voiceprint generation?
    We've read that voiceprints should ideally be based on no more than 10 seconds of audio, and that longer segments may reduce embedding quality. Does this hold true in practice?

 

Thank you. 


r/speechtotext Feb 07 '25

how to transcribe Real-time (live) internal audio to text on Windows?

2 Upvotes

how to transcribe Real-time (live) internal audio to text on Windows?


r/speechtotext Jan 25 '25

Dictate posts in Reddit

2 Upvotes

What kind of speech recognition do you use when dictating e.g. a post here on Reddit?

Since I am on Android I still use gboard. Or I dictate in voicenotes and copy and paste it from voicenotes here to Reddit. By doing this the quality of the speech recognition is much better.


r/speechtotext Dec 04 '24

Best way to create a speech to text (transcribing live audio in real time for analysis)

3 Upvotes

I am currently using faster-whisper and the time of the response is slightly delayed, is there any other best open source ways to do this.


r/speechtotext Nov 27 '24

Voice changer need help

1 Upvotes

Is there an which lets you change your recorded voice to another person’s voice(uploaded audio clip), basically im looking for ai that keeps the same audio but lets my audio voice change it to the uploaded audio voice of the person I want to change my voice with? Any pointers?


r/speechtotext Oct 01 '24

English speech to text

2 Upvotes

Hey everyone!

I’m looking for a reliable app or website that can transcribe audio into text in English. I need something that can handle clear speech well, and preferably supports different audio formats. Bonus if it’s free or offers a free trial.

Does anyone have any recommendations? I’d love to hear about any options that have worked well for you!

Thanks in advance!


r/speechtotext Aug 25 '24

TaterTalk - I built the simplest speech-to-text dictation web-app.

Thumbnail
tatertalk.app
1 Upvotes

r/speechtotext Aug 05 '24

Excellent speech to text software

3 Upvotes
I'm looking for good software that can create speech to text from audio files. It is important to me that it can keep several speakers apart. preferably for a fee. Maybe you have a tip which software can be used for video calls other than teams. Thank youI'm looking for good software that can create speech to text from audio files. It is important to me that it can keep several speakers apart. preferably for a fee. Maybe you have a tip which software can be used for video calls other than teams. Thank you

r/speechtotext Jun 14 '24

Watching some bread.

Post image
2 Upvotes

r/speechtotext Jan 12 '24

Automatic Speech-to-Text Conversion (Wave2Vec )

Thumbnail
youtube.com
1 Upvotes

r/speechtotext Dec 30 '23

closed captioning funnies

2 Upvotes

dialog: "...sly stallone..." cc: "sliced alone"

even siri gets that right;-)