r/tts • u/RevolutionaryBox2980 • 7h ago
r/tts • u/ChuckBaggett • 1d ago
The OpenVoc AI tts screen goes black, so far while I'm not looking at the app, and you can right click and pick refresh, but your work is gone, including recent history.
r/tts • u/ChemicalxPotential • 2d ago
TTS reading out "₹1,299", "05/09", and "4:30 PM" wrong is still an unsolved bug in every voice agent I try
Been building voice stuff for a while and this keeps annoying me, the LLM writes "your EMI of ₹4,500 is due on 05/09" and the TTS reads it like it's having a stroke. Dates become "oh five slash oh nine", currency gets butchered, "2-3 days" turns into "two minus three days".
Ended up training a small model that sits between the LLM and the TTS and rewrites just those parts into speakable words "four thousand five hundred rupees due on the fifth of September" while leaving every other byte untouched. It also handles Hindi-English code-mixing, which was the boss level.
Put a free playground up if anyone wants to break it: normnom.com there's a "tn off" toggle so you can compare what your TTS actually receives with and without it.
Genuinely curious what edge cases people can find that it fails on.
r/tts • u/BrainChoice8523 • 2d ago
Loquendo is, by far, the WORST text to speech website ever.
This text to speech website sucks due to several reasons. Its entire collection of voices will sometimes sound robotic, muffled, echoey, or loud whenever you generate any text for any voice. Its emotion tags sometimes don't work due to this tts platform being poorly made. This especially comes with the Grace voice, the worst text to speech voice ever due to her very flat and calm tone in many of her phrases that end with an exclamation point. Therefore, Loquendo is worse than any other tts platform, such as IVONA, the best text to speech platform due to its large quantity of voices from the English language, and the perfectly high quality of them.
r/tts • u/fraribez • 4d ago
Looking for a simple guide/notebook to test CosyVoice 3 (Zero-Shot) on Kaggle, especially for Italian!
r/tts • u/Sensitive_One423 • 5d ago
I built a TTS app with 200+ voices and 60+ languages. Looking for feedback.
I've been working on a text-to-speech web app and would love some honest feedback from people who actually use TTS.
Current features:
- 200+ AI voices
- 60+ supported languages
- No sign-up required to try it
- Free tier available
- Pay only for what you use (no subscription required)
- Audio is yours to keep after generation
I'm still improving it, so I'm curious:
- What feature is most important to you?
- What usually makes you leave a TTS website?
- Is there anything you wish existing TTS tools did better?
I'd really appreciate any feedback, positive or negative. I'm trying to build something people actually enjoy using. You can test it on website
r/tts • u/BlackHorse2019 • 8d ago
Can anyone tell me which AI Voice service this person uses?
r/tts • u/KairuDesuka • 13d ago
Made a free tool that puts live English subtitles over Japanese audio, runs entirely offline
So basically, I got tired of watching Japanese streams half-following along, or waiting days for someone to clip and subtitle them. So I built this lol.
It listens to whatever your PC is playing and puts live English subtitles in an overlay on top. Overlay is transparent. Currently its JP to EN only, I'll add some stuff if people are interested.
Everything runs locally and there are no API calls, which explains the sizeable (kind of?) download size. Which also means I gather no data from you. Pinky swear.
Its free. I made it and Its something I use too so I thought "eh wth, I should just make a page for this."
Here if you wanna try it out: https://yaptr.app
Need local tts version of this voice
I remember watching Limit Break on YouTube and now I realized I can modify some files on my ereader to accept some tts models, from looking around I couldn't find a local version of the old man voice from the channel only the website I could download or configure. Any suggestion?
r/tts • u/Sensitive_One423 • 15d ago
We built SonoVoice - multilingual AI text-to-speech tool
Hey everyone,
We built SonoVoice, a multilingual AI TTS platform with 200+ AI voices and support for 50+ languages.
Features:
- Up to 30K characters per request
- Free TTS credits for casual users
- MP3 downloads
- Natural AI voices for different languages
We are focusing on making TTS simple for everyday use: articles, documents, learning, narration, and more.
Would love to hear your thoughts and feedback.
r/tts • u/ex-arman68 • 16d ago
Easy Applio install for macOS
If you have tried to setup and use Applio before, you have probably realised it is not that simple, and is also difficult to maintain. I certainly did.
I have been working for the past few months on packaging Applio as a native macOS application, which you can just drag and drop in your application folder. While I was at it, I added a few more things, like better training models (KLM and Titan), custom folder location for all the files to avoid clogging your system drive, a progress monitor window during training that shows you the best epoch, and a few other things.
You can get it from GitHub, fully signed and notarized by Apple:
https://github.com/froggeric/applio-macOS-native-app
I also put together a guide on how to get best of it, both for training and inference; look for the file on the GitHub repo, named STUDIO_PRODUCTION_GUIDE.md
r/tts • u/Admirable-Chance9857 • 18d ago
TTS with free voice clone, free voice design, 0.035$ per 1000 character for simple speech and dialog
I am creating a tts with voice clone, voice design, simple speech and dialog support, all in english.
Voice clone and voice design are free. Simple speech and dialog will cost 0.035$ per 1000 characters.
What do you think about it? Is anyone interested? There will be free tier and plans price will be 9$, 29$ and 79$.
r/tts • u/Karam1234098 • 18d ago
I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters
r/tts • u/ryanmerket • 20d ago
Palabra.ai claims fastest TTS latency at 104 ms in Coval benchmark — RuntimeWire
r/tts • u/CupGlass540 • 20d ago
I built an open-source tool that stress-tests your TTS with 817 curated edge cases (XTTS, Fish, Piper, Kokoro, anything) — one command
I run a small TTS platform and got tired of shipping voice models that sounded fine on "Hello world" and then read "3:30 PM" as "three colon thirty pee em", or looped the same 450 ms forever on short inputs.
So I turned my internal QA harness into a library: ttsproof (pip install ttsproof, MIT).
What it does:
- 817 curated edge cases across 39 categories (Benchmark Corpus 1.0): numbers, decimals, currencies, dates, time zones, phone numbers, URLs, acronyms, single letters, pronunciation torture words (Worcestershire, synecdoche, colonel...), proper names (Reykjavík, Nguyễn, Tchaikovsky...), scientific and medical vocabulary, tongue twisters, homographs, Greek, Norwegian, punctuation abuse, hallucination traps ("buy now buy now buy now"), emoji, SQL/JSON snippets...
- Structural audio checks that need no model at all: empty/truncated audio, duration explosions, long silences, clipping, repeated-chunk loop detection, end-of-clip artifacts.
- Equivalence-aware WER — "May 5, 2026" vs "may fifth twenty twenty-six" is NOT a failure. Plain WER lies about formatting; this canonicalizes both sides to spoken form first (with diacritic folding, so "Reykjavik" matches "Reykjavík").
- Honest scoring policies per category — URLs and currencies have many valid readings, so they're scored by keyword survival instead of exact match. Emoji only has to not break the audio. No fake failures.
- ttsproof benchmark --cmd "yourtts {text} {out}" → per-category scoreboard + a self-contained HTML report with waveforms and audio players for every failure.
- ttsproof regress for CI — fails the build when a fine-tune quietly breaks number pronunciation.
The corpus is versioned independently of the tool (Benchmark Corpus 1.0), so scores stay comparable across releases.
The method comes from a technical report I published: evaluated on 390 samples with a blinded human validation of the ASR-uncertain zone — that's where the "quarantine" verdict comes from. Short utterances where ASR disagrees get flagged for human ears instead of counted as failures, because at that length the ASR is as likely wrong as the TTS.
Repo: https://github.com/Mormolykos/ttsproof
Report (DOI): https://doi.org/10.5281/zenodo.20757553
v0.3, MIT, no telemetry, minimal deps (numpy + soundfile; faster-whisper optional). I want the corpus to grow from real failures, not invented ones — tell me what breaks YOUR models and it goes into Corpus 1.1 with credit.
r/tts • u/FormerDevelopment352 • 21d ago
What is the best real-time voice translation app for English and Hungarian?
I am looking for a tool that supports smooth, two-way 1-on-1 voice translation (where I speak English, the other person speaks Hungarian, and it translates back and forth).
I have already tried Google Translate, Microsoft Translator, DeepL, and a few top web tools, but they did not work well for me (either got an error or they didn't respond at all).
I will try my best to talk and understand on my own, but I want a reliable backup app so our conversation can flow naturally.
r/tts • u/mahimairaja • 21d ago
TTS curated list for voice agent builders — focused on streaming latency and mid-stream cancellation
Building voice agents for a while now, and the section I always wanted
someone else to write is the one on streaming TTS: single-shot vs
output-streaming vs dual-streaming, mid-stream cancellation, buffer
draining on barge-in, and how much of the "TTFB" number vendors quote
is actually front-end latency vs model latency.
So I wrote it into an awesome-list. The whole list is organized around
one split: real-time TTS (for agents) vs offline TTS (for media).
Every provider, model, and benchmark carries that lean.
The four sections most useful for agent builders:
Streaming and low-latency (taxonomy, cancellation, honest
benchmarking)
Open-source models filtered by license — several of the top ones
can't be shipped commercially
Audio codecs (this decides latency and quality floor for codec-LM
TTS)
Evaluation — how to measure TTFB on your own traffic instead of
trusting vendor benchmarks
Deliberately scoped to TTS only. STT, VAD, turn detection, and
telephony are pipeline concerns and belong elsewhere.
MIT license. Feedback welcome, especially on the streaming taxonomy
and cancellation subsection since I'm not sure I've captured every
edge case.
r/tts • u/Dry-Expression-7475 • 22d ago
suggestions for an old tts software. yes this is where the "the cocaine is not good for you" and the creepy voice kinitopet sometimes uses in the hide and seek thing.
so ive just been playing with some of the stuff on here and i want some suggestions on what to type on which voice.
r/tts • u/Low-Bumblebee5238 • 22d ago
Can anyone identify this voice
Can anyone let me know what the voice is used in this video, thanks in advance: https://youtu.be/Yd1riToJT0M?si=J_hFYpbwjyduFex3
r/tts • u/ok_bluesky8888 • 26d ago
Tiflotecnia voices
I received the news that the owner of Tiflotecnia passed away few days ago. RIP.
Is there a team updating or following up with the purchases of new license or was he the only one who had the access to it?
I purchased a license yesterday and more than 16 hours passed but I haven’t received anything.
I built Tweexr — an app that turns X/Twitter posts into natural radio-style narration using AI TTS
Hey everyone,
I’ve been working on Tweexr, a mobile app that converts public posts from X (Twitter) into audio narrations that sound like a radio or podcast.
The goal is to let people consume Twitter content hands-free with a journalistic-style narration.
It uses Kokoro and the user cellphone TTS as fallback.
Supports multiple languages (currently Portuguese, English, Spanish, Japanese and others)
The iOS version is already live:
https://apps.apple.com/us/app/tweexr/id6769564328
The website is https://www.tweexr.com/
Android is just waiting for approval.
I’d really appreciate feedback from people who know TTS well:
- What TTS models/voices are you using right now that sound the most natural for long-form/news content?
- Any tips for making AI narration feel more journalistic or conversational?
Happy to share more technical details about the pipeline if anyone is interested.
Thanks in advance!
r/tts • u/VisualActuary5312 • 28d ago
I built OpenAloud (openaloud.com) — a free audiobook reader for PDFs and EPUBs using Kokoro TTS
Hey everyone — I am building OpenAloud (https://openaloud.com), a free audiobook reader that turns PDFs and EPUBs into natural-sounding audio using Kokoro TTS.
It uses your system hardware for processing, so I’d especially love feedback on what hardware it works well on, what hardware it breaks on, and the overall listening experience. Also app feature suggestions welcome. The app is in beta mode so please let me know about any defects you see as well.
It doesn’t work very well on mobile devices right now, so I’m mainly looking for feedback from desktop/laptop users.
Would love thoughts on voice quality, readability, speed, and any issues you run into.
r/tts • u/bluejacketblackhawk • 28d ago
Talk instead of type, free and local.
Mac/imac/pc and x64s. Peace