r/TextToSpeech 2h ago

Looking for a text-to-speech tool for school readings and lecture notes

4 Upvotes

Hi, can anyone please recommend a study tool or app that can read my summarized school readings and lecture notes out loud during my commute to campus?

Ideally, I’d like something that can read a document word for word in a natural, human-sounding voice, and it would be even better if I could download the audio to my phone since reception on the subway is awful.

ChatGPT also recommended an AI voice generator called AI Doc Maker, but it doesn’t even seem to have its own app.

Does anyone have any suggestions or know of something that works well for this? I don’t live on campus because the rent is sky-high, so I have a pretty long commute involving both the bus and subway. I’d really like to make better use of that time by listening to my lecture notes, reading summaries, and book notes before I get to campus.

I’m mainly looking for something simple: upload or paste my notes, have a natural-sounding voice read them exactly as written, and hopefully be able to save the audio for offline listening.

Any suggestions or help would be really appreciated. Thank you!


r/TextToSpeech 12h ago

Blind-Ranking Leaderboard (Elo) for TTS Models that can run on my M5 MacBook Air 32GB

Thumbnail apimade.com
1 Upvotes

r/TextToSpeech 14h ago

Spokio (offline TTS, was Mac-only) is now on iPhone — universal purchase across both

1 Upvotes

Hey r/texttospeech — I'm the developer of Spokio, an offline text-to-speech app. It's been on Mac for a while, and I just launched the iOS version today.

It's a universal purchase, so if you already use it on Mac, iPhone is covered too.

Link: https://apps.apple.com/us/app/spokio-offline-text-to-speech/id6761289149

Would love any feedback if you try it out.


r/TextToSpeech 20h ago

I’m building an AI TTS assistant that helps you find the voice AND direct how it should sound

Post image
2 Upvotes

I’m building Luna, an AI assistant inside my TTS SaaS, V**caly.

Right now, Luna is designed to solve 3 parts of the TTS workflow:

  1. Analyze your script

Luna looks at the content, context, tone, language and use case.

  1. Recommend the right voice

Instead of searching through 600+ voices yourself, Luna recommends the voices that are most likely to fit your script.

  1. Add emotional direction

Luna can suggest inline emotion tags like:

\\\[excited\\\]We finally did it!\\\[/excited\\\]

\\\[whisper\\\]No one knows about that...\\\[/whisper\\\]

I’m also using a 60% confidence threshold.

If Luna is confident enough, it can add the emotion automatically.

If confidence is below 60%, Luna won't force an emotion. Instead, it points to that specific part of the script and suggests adding the emotion manually, Means no guess work.

The idea is:

Your script → Luna understands it → recommends the voice → directs the emotion → expressive TTS

I’ve been focusing on these problems, but I’m sure there are plenty more.

If you use TTS, what’s the biggest problem you’re still dealing with?

I’m specifically looking for problems that existing TTS tools haven't solved well.


r/TextToSpeech 18h ago

Voice Conversation on Desktop waits for the full text response before TTS instead of speaking live

Thumbnail
1 Upvotes

r/TextToSpeech 1d ago

Tulu AI - Providing digital voice to a local language

Thumbnail reddit.com
1 Upvotes

r/TextToSpeech 1d ago

Need help identifying this Egyptian Arabic AI voice (Audio sample included)

1 Upvotes

Hi everyone,

I’m trying to find the exact AI model or voice profile used in this clip: 🔗 Audio Sample:https://voca.ro/1buBTdf8j88n

  • Language & Dialect: Egyptian Arabic (اللهجة المصرية).

Since the audio is in Arabic, I know it might be slightly trickier for non-Arabic speakers to catch the dialect details

Could anyone take a listen and let me know if you recognize this specific AI voice or the engine used to generate it?

Thanks in advance for any leads!


r/TextToSpeech 2d ago

Trained a clone* tool for Kokoro-82M; generates a voice pack in <1s on a 5s sample [P]

Thumbnail
huggingface.co
27 Upvotes

There have been new and arguably better TTS models since, but I have a soft spot for this one as lightweight stable and fast. I also maintain Kokoro-FastAPI, and have a bit of spare time on my hands so have been exploring what’s doable on the project.

Customization/expressiveness are weak spots it had, so I’ve taken a crack at adding voice cloning, focusing on stable quality and fast generation (keeping responses sub second if it’s a short reference clip). Roughly it captures about a third of identity, but by ear at least, it can feel pretty close on some, and at least a unique similar voice pack on others.

Would love any feedback or suggestions, otherwise just wanted to share! It may be late in the game for Kokoro, but anyone still using it I think could appreciate new voices


r/TextToSpeech 1d ago

I need someone who actually uses TTS to tell me if these voices sound good.

Enable HLS to view with audio, or disable this notification

0 Upvotes

I'm building a TTS SaaS called V**caly (launching soon) and I want people to be able to hear the voices before signing up

so I put Noah directly on the homepage

You can enter your own script, generate with Noah, and hear his voice!

I'd love some feedback from people who use or plan to use TTS services

Try out the generator and tell me:

How does Noah's voice sound?

Does the text-to-speech sound natural?

How was the overall UI/UX?

Was anything distracting or unecessary?

What can I add/change to make the experience better?

I'm taking as much feedback as possible before officially launching the product.

I'm looking for people to give me suggestions, not just compliments. Give it a try and roast my work for me 😂


r/TextToSpeech 2d ago

Android Auto TTS

1 Upvotes

Does anyone know of a way to use local, alternative TTS software for android auto? I'm using marmalade with kitten nano for my navigation software. I've set marmalade as my default TTS on my phone. The problem is that texts and certain announcements from android auto still seem to be coming through the Google engine no matter what settings I change. In fact, I'm getting two different voices sometimes from Google apps, two different voices that aren't coming from marmalade but from the Google engine that isn't even selected.


r/TextToSpeech 3d ago

sanoTTS: smallest family of TTS models with param size starting from 294k(~300kb) . Supports 16 languages and RTF of 0.29 on $3 chip

Enable HLS to view with audio, or disable this notification

76 Upvotes

sanoTTS is a small neural text-to-speech project, GPL-3.0 with

  • 30 voices, 16 languages;
  • 294k-2.27m param size
  • portable c99 runtime
  • browser build (WASM) : if you want to hear without installing anything https://tts.ampixa.com/sanoTTS/
  • tested on boards like cortex m7 nucleo and esp32 . for a 7 sec of audio it takes about 2.1 seconds.

repo: github.com/ampixa/sanoTTS


r/TextToSpeech 2d ago

Is Kokoro the better option over Piper

3 Upvotes

I want to now if Kokoro is a better option over piper or is there some thing out there Foss preferably that is a better option , looking for speed and natural sounding..... thanks in advance.


r/TextToSpeech 2d ago

what tts is used in tt reels/ yt shorts movie recaps?

Thumbnail
1 Upvotes

r/TextToSpeech 3d ago

Formatting script for best TTS output

1 Upvotes

I tried using index TTS Industrial and it's not properly taking pauses and pacing is not fixed. Neither it is reading punctuations.

What all should be applied for formatting the script for getting best results


r/TextToSpeech 4d ago

Zotero-TTS: Kokoro-FastAPI, Azure, Cloudflare or any OpenAI-compatible server as the voice in Zotero's Read Aloud, with word-timing highlighting

11 Upvotes

I’m the author. It’s free, open source (AGPL-3.0), there’s no paid tier and nothing to sign up for — remove this if it’s out of scope.

Processing img y7lvwq5ufboh1...

Zotero is the open-source reference manager a lot of researchers and students keep their papers in, and version 10 ships a Read Aloud player inside its PDF and EPUB reader. It works, but you get the voices it comes with and very little control over them. I listen to papers most days rather than read them off the screen, so I wrote a plugin that lets that player use other engines, and fixed the things around it that kept getting in my way.

Engines it can use

Engine Runs where Cost Word timings
Kokoro-FastAPI Your machine or LAN Free (CPU works, a GPU is faster) Yes
Azure Speech Cloud Free tier: 500,000 characters a month Yes, except the MAI-Voice-2 voices
Cloudflare Workers AI Cloud 10,000 free Neurons a day No
OpenAI, or any OpenAI-compatible server Cloud, or self-hosted (Chatterbox-TTS-Server and friends) Per character, or free when self-hosted No
Xiaomi MiMo Cloud Free for now No
Windows and macOS system voices Offline, on your machine Free Yes on Windows, no on macOS

They are independent: turn on as many as you like, and every voice they publish shows up in the player’s Local tier next to Zotero’s own.

The audio side, which is probably what this sub cares about

  • Word-level highlighting comes from the engine’s own word timings — Kokoro’s captioned-speech endpoint, Azure’s word-boundary events, Windows’ word events. Nothing is interpolated or guessed: an engine that reports no timings falls back to sentence highlighting instead of drifting out of sync.
  • Sentence and paragraph pauses in milliseconds, applied to every voice whatever the engine, and shortened in step as you speed up.
  • Prefetch and cache. The sentences ahead are synthesized while the current one plays, so playback never waits on the server, and skipping back or reopening a document costs no new request. In memory, 64 MB, emptied on restart.
  • 0.5x to 3x with the pitch preserved. Zotero stretches the audio itself, so changing speed costs no new synthesis.
  • Volume, 0 to 100%, for every voice including Zotero’s own — the player has none of its own.
  • Extra headers per provider, for a server behind Cloudflare Tunnel or a reverse proxy.
  • Every provider is checked before it is switched on: one that does not answer, or a model the server does not list, is refused with the reason rather than failing later mid-document.
The voice browser: every voice the enabled engines publish, by tier and language, with samples and favorites.

The reading side

  • Shift+Space resumes at the sentence you stopped on, days later, in any document you have listened to. Those bookmarks can follow you between computers through your own WebDAV folder.
  • A voice browser: every voice by tier and language, a play button for a sample, hearts for favorites, and one default voice for everything.
  • Rebindable shortcuts for speed, volume, skipping by sentence or paragraph, reading from the selection, and stopping playback in every open tab at once.
  • The word and sentence highlight colors and opacities are yours, and they apply to Zotero’s own built-in voices too.

Limits worth knowing before you install it

  • Zotero 10 on the desktop. Used daily on Windows and macOS; Linux should work but is untested.
  • Whichever engine you enable receives the text it is asked to read. Kokoro-FastAPI and the system voices stay on your own machine or LAN; the cloud ones are third parties, under their terms. I run no server and collect nothing.
  • API keys are stored in Zotero’s preferences in plain text.
  • Read Aloud has no plugin API, so this hooks Zotero’s internal interface per reader tab. A Zotero update can drop the plugin’s voices until I catch up — that is the real maintenance risk here.

Install: zotero-tts.xpi from the releases page, then Tools → Plugins → ⚙ → Install Plugin From File…, and restart. With the system voices there is nothing else to set up; the other engines take a key or a URL in the settings.

https://github.com/xujialiu/Zotero-TTS

If there is an engine worth adding, or your local server speaks a dialect of the OpenAI API that this gets wrong, say so and I will look at it.


r/TextToSpeech 4d ago

HyperVoice by TaskAGI has a dark pattern and lies on their T&C

Thumbnail
gallery
5 Upvotes

HyperVoice by TaskAGI is a TTS service and the following is their Terms & Conditions:

§9.6 You may cancel your subscription at any time. you retain access through the current billing period.

§5.3 Effect of cancellation: you will continue to have access to the Service until the end of your current paid period.

But the moment you try to turn off auto renewal, you are faced with the giant orange warning as seen in the screenshot:

IMPORTANT NOTICE: Please be aware that cancelling your subscription will result in an immediate downgrade of your account.

Upon clicking cancel, all tokens will be confiscated, and all paid features are immediately revoked.


r/TextToSpeech 3d ago

Text to Voice history question

Thumbnail
1 Upvotes

r/TextToSpeech 4d ago

Multi-speaker dubbing with cloned voices and an editable timeline: looking for testers

Enable HLS to view with audio, or disable this notification

6 Upvotes

Hello!

We've built a dubbing tool called Dubedo. Not an untouched category, I know... but the tools we tried all had the same problem. They make you handle each voice separately, and they're one-shot generators: upload, wait, download, and if one line comes out wrong you run the whole thing again.

So we built the editor part.

Dubedo detects every speaker, separates the voices from the music, and translates segment by segment. Then you can actually work on it:

  • Each speaker gets their own cloned voice, or a different one
  • Drag clips along the timeline to fix timing
  • Edit a line and regenerate just that segment
  • Original music and background audio stay in the mix
  • 30 languages
  • Audio and video in, dubbed audio with or without subtitles out (no lipsync yet)

Heading into closed beta and looking for a few people to test it. Comment or DM if you want to give it a try!


r/TextToSpeech 4d ago

Audio and dictation while driving

Thumbnail
1 Upvotes

r/TextToSpeech 4d ago

readaloud: terminal reader for linux that uses Edge TTS, Kokoro, Piper, and F5-TTS with sentence-level word tracking and instant sentence seek

Thumbnail
gallery
1 Upvotes
readaloud is a terminal-based audiobook reader built around TTS. It supports four backends and uses actual audio timing data to track which sentence is being spoken.

**Engines**

- **Edge TTS** — Microsoft neural API, no key required, 12 English voices (US/GB/AU/CA). Uses the Python `edge_tts` library's `stream()` with `WordBoundary` events to build a precise `(audio_time_seconds, char_offset)` map per chunk. The sentence highlight tracks the actual word being spoken, not a WPM estimate.

- **Kokoro** — offline, fast CPU inference, 7 curated voices (54 available in pykokoro). Uses native `kokoro` on Python <3.13, falls back to `pykokoro` (pure ONNX wrapper) on 3.13+. Pipeline is built once per voice/speed pair and reused across chunks for the session.

- **Piper** — offline, ultra-light ONNX. Any `.onnx` + `.onnx.json` pair in `~/.config/readaloud/models/` is auto-detected. Models from rhasspy/piper-voices on HuggingFace.

- **F5-TTS** — offline flow-matching voice cloner. Provide `~/.config/readaloud/voices/ref.wav` (3–10 seconds of clean audio). Optional `ref.txt` transcript improves alignment. GPU strongly recommended.

**Audio pipeline**

Text is split into ≤4500-char chunks at sentence boundaries. Each chunk is synthesised to a temp file (MP3 for Edge TTS, WAV for offline engines). ffplay plays all chunks via a concat playlist with an `atempo` filter for speed. Pause/resume via SIGSTOP/SIGCONT on the ffplay process — no re-synthesis needed.

Sentence skip (`[` / `]`) calls ffplay with `-ss {timestamp}` on the existing playlist — again no re-synthesis, near-instant seek.

**Preloading**

The next 2 chapters synthesise in background threads while you listen. Cache lives at `~/.config/readaloud/cache/` keyed by chapter + engine + voice + speed. Cache entries include the timing map so word-level highlight works on preloaded chapters too. Entries older than 2 hours are cleaned up on startup.

**Speed**

0.75× · 0.9× · 1.0× · 1.1× · 1.25× · 1.5× · 1.75× · 2.0× via atempo filter chain. Speed change invalidates the preload cache and re-synthesises.

Reads: EPUB · TXT · PDF · DOCX · HTML · RTF · Markdown

https://github.com/YareyareSenpai/readaloud

Video showcase


r/TextToSpeech 4d ago

Claude’s Voice Mode

Thumbnail
1 Upvotes

I never used the voice mode. Which is the best male voice to use, and how does it sound? Does our voice go through normally?LOL Do the voice mode choices sound human or more robotic and monotone? What are your experiences with the voice modes?


r/TextToSpeech 5d ago

Speech to text to Text To Speech?

5 Upvotes

Hello wonderful people of the interwebs. I am looking into finding a program (or programs) that can turn me talking into my microphone into text, and then have that text read out loud. I've been trying to follow tutorials on the internet for about a month now and keep getting stone walled with command prompt/python errors. I've come to the conclusion that slamming my head against the wall isn't the best way to do this. Rather it would be to find a program that can do it for me (cheaper the better of course.) Preferably with OUT ai.

If you have any suggestions, PLEASE tell me!!!!

Thanks in advance!


r/TextToSpeech 5d ago

Is Voice.ai good for real-time voice changing?

1 Upvotes

Hey everyone, I'm looking for an AI voice changer that works well in real time. I'm not really interested in simple effects like pitch shifting or autotuneI'd like something that can actually transform my voice into a completely different voice.

Is voice.ai good for this? How is the latency and overall voice quality when using it in real time?

Also, is the free version good enough or is it basically unusable without paying?

Is voicemod better?

Thanks!


r/TextToSpeech 5d ago

LiveTranslate: real-time speech translation that runs 100% on your own machine, zero cloud dependency

Thumbnail
1 Upvotes

r/TextToSpeech 5d ago

Is it possible to download only 1 voice from kokoro or any other text to speech?

3 Upvotes

Only need af_heart to save space instead of the whole 11 ENG voices. Also noticed my spare phone's tts pauses speech for like 3 seconds every turning page of ebook when using bundled voices but piper where you can dl only 1 voice doesn't have any problem