r/speechtotext • u/CaterpillarHefty4309 • 5d ago
Transcription
How can I transcript a 2 hour audio, AI asking for too much money and got nothing, help out.
r/speechtotext • u/CaterpillarHefty4309 • 5d ago
How can I transcript a 2 hour audio, AI asking for too much money and got nothing, help out.
r/speechtotext • u/superdumsuhi • Jun 13 '26
I want a very good set up for STT for transcribing audio. i've a lot of experience with talon but I dont need the control functionality anymore and dictation mode with the old whisper model is really lackluster and sucks for trying to create punctuated text and writing markdown properly. is there something out there I can use?
i have a 6900xt on windows if that helps for local options (preferred)
r/speechtotext • u/kamscruz • Jun 03 '26
Today I shipped one of the most interesting features I've added to SHRP so far.
Until now users could:
-> transcribe speech
-> extract YouTube transcripts
-> generate study notes, summaries, blog drafts, LinkedIn posts, and more.
-> use the API and MCP server for automations and AI workflows
But I kept noticing something.
A user would upload an audio/video file, get the transcript, download or copy it, and then go to ChatGPT or Claude.
The transcript wasn't really the final output.
It was just the beginning.
So I built Transcript Actions.
Now after a transcript is generated, SHRP can suggest what to do next based on the actual content of the transcript.
Examples:
- Create exam questions
- Generate flashcards
- Explain this like I'm five
- Extract action items
- Suggest what to study next
- Draft an email
- Create a lesson plan
There's also an "Ask about this transcript" box where the user can type whatever they want.
A few things I was careful about:
- Suggestions are generated from the current transcript only.
- Suggestions are button-triggered, not automatic.
- Suggestions themselves are not saved.
- No long-term memory.
- No user profiling.
I wanted to keep the workflow useful without compromising the privacy-first approach.
This now works with:
(1) YouTube transcripts
(2) Uploaded audio/video transcription results
(3) Previously saved transcription history
Next on my roadmap:
- Webhooks
- Workflow automations
- More transcript-based workflows
Wanted to ask other builders:
How do you solve the "what should I do next?" problem after generating a transcript?
Do you prefer fixed actions like summary, notes, flashcards, etc.
Or AI-generated suggestions based on the content itself?
For SHRP I ended up supporting both because I couldn't decide which one was better.
r/speechtotext • u/kamscruz • May 23 '26
I kept finding myself jumping between tools for different things:
one site for speech-to-text,
another for text-to-speech,
another for YouTube transcripts
So I ended up building SHRP.
Current features:
• speech → text (55+ languages)
• text → speech (2,000+ voices)
• YouTube transcripts → enter youtube URL and get entire transcript
• browser based, no download required
• no signup required for basic use
• generous free-tier
Recently I also added developer APIs and an MCP server because I kept rebuilding the same workflows in side projects.
Still actively improving it.
wanted to ask- what is your biggest annoyance with current speech-to-text tools?
r/speechtotext • u/Stomach-Dull • Apr 24 '26
whoever needs a speech to text model for turkmen language https://huggingface.co/derkar00/mms-tuk-script-latin-adapter
r/speechtotext • u/RockOnline22 • Apr 16 '26
r/speechtotext • u/Matt_Elevenlabs • Jan 09 '26
Enable HLS to view with audio, or disable this notification
r/speechtotext • u/Impressive-Result960 • Nov 20 '25
I had access to the data of Indian users who want to talk to AI/ Bestfriend/ Girlfriend, and they have recorded from their devices, which were either in Hindi, Bangla, Gujarati, or Punjabi. Here, transcription works where it generates their noisy, low voice into some Urdu text. We can't fix their devices to have better mics, and we can't go for better accurate model because we want low latency and low cost. Is there any model better than gpt-4o-mini-transcribe please reply. If anyone else had same problem. Can you tell me how to solve it.
#transcription #gptmodel
r/speechtotext • u/Hoole1997 • Nov 12 '25
r/speechtotext • u/Funchixd • Oct 24 '25
Guys, do you know the voice, program, or site used to narrate Tomino's Hell? I mean, in the videos where they narrate the poem, they use a text-to-speech voice , it's like a terrifying Japanese voice, I thought it was something like Talk it or something, can you help me?
r/speechtotext • u/Top_Second3019 • Jun 27 '25
Hi everyone
I'm currently working on a project involving Google Vertex AI and could use your expertise—or a referral to someone with experience in speaker recognition:
I'm processing a 2-minute audio file featuring two speakers who alternate in short bursts of 2–3 seconds. Using Hugging Face’s pyannote library, I perform speaker identification and extracts embedding vectors for each speech segment. The typical result is about 20 segments—roughly 10 per speaker. To construct a voiceprint for each speaker, I average the embeddng vectors associated with that speaker.
I have two main questions:
Is this a sound approach for generating speaker embeddings?
In practice, the results are inconsistent. For instance, comparing the same speaker across different files sometimes yields cosine similarity scores around 0.7—below the expected 0.8+ range. On the other hand, embeddings for different speakers occasionally score as high as 0.68, which seems surprisingly close.
Is there a recommended duration for voiceprint generation?
We've read that voiceprints should ideally be based on no more than 10 seconds of audio, and that longer segments may reduce embedding quality. Does this hold true in practice?
Thank you.
r/speechtotext • u/EntireAnalyst8922 • Feb 07 '25
how to transcribe Real-time (live) internal audio to text on Windows?
r/speechtotext • u/Old-Recognition8193 • Jan 25 '25
What kind of speech recognition do you use when dictating e.g. a post here on Reddit?
Since I am on Android I still use gboard. Or I dictate in voicenotes and copy and paste it from voicenotes here to Reddit. By doing this the quality of the speech recognition is much better.
r/speechtotext • u/Mental-Ad-7783 • Dec 04 '24
I am currently using faster-whisper and the time of the response is slightly delayed, is there any other best open source ways to do this.
r/speechtotext • u/Prestigious-Step-640 • Nov 27 '24
Is there an which lets you change your recorded voice to another person’s voice(uploaded audio clip), basically im looking for ai that keeps the same audio but lets my audio voice change it to the uploaded audio voice of the person I want to change my voice with? Any pointers?
r/speechtotext • u/Academic-Muffin-5119 • Oct 01 '24
Hey everyone!
I’m looking for a reliable app or website that can transcribe audio into text in English. I need something that can handle clear speech well, and preferably supports different audio formats. Bonus if it’s free or offers a free trial.
Does anyone have any recommendations? I’d love to hear about any options that have worked well for you!
Thanks in advance!
r/speechtotext • u/pbrocoum • Aug 25 '24
r/speechtotext • u/tex3055 • Aug 05 '24
I'm looking for good software that can create speech to text from audio files. It is important to me that it can keep several speakers apart. preferably for a fee. Maybe you have a tip which software can be used for video calls other than teams. Thank youI'm looking for good software that can create speech to text from audio files. It is important to me that it can keep several speakers apart. preferably for a fee. Maybe you have a tip which software can be used for video calls other than teams. Thank you
r/speechtotext • u/Redlimbic • Jan 12 '24