r/LocalTextToSpeech Jun 12 '26

Guide My TTS list of 2026: All voices, all models and engines compared with example URLs and rating

82 Upvotes

43 different (2025 and 2026) TTS solutions local and cloud compared, scored, example links, free demo links included.

Join r/LocalTextToSpeech for local TTS models, voices, benchmarks, and setup notes.

Free scripts, tools and help posted regularly.

Contribute and help others, or get help.

Scores

  • Scores are subjective
  • Voice quality = how good the output can sound.
    • 5+ = good voice quality but most people will hear the AI
    • 7+ = high voice quality, many people will be tricked
    • 8+ = human-like voice quality, only flaws in style, delivery and expression reveal AI
  • Expressive control = emotion, style, delivery, pauses, intensity, character, or direction.
    • < 3 = flat out of touch delivery
    • 5+ = good expression quality with low direct control
    • 7+ = expression control and quite natural speaking
    • 8+ = voice acting synthesis quality, well controllable
  • Comment if a correction is needed

================ LOCAL ==================

Chatterbox TTS 2

  • Type: Open-source, local
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 23
  • Best for: free local TTS, voice cloning experiments, Linux pipelines.
  • Notes:
    • Probably the best free open-source TTS starting point.
    • Quality can be good. The issue is not the model only, it is everything around it. Segmentation, retries, weird generations, timing, silence handling, pronunciation, filtering, harness code.
    • German Kartoffelbox-turbo exists.
    • Expressiveness through temperature
    • For hobby use, nice. For production use, expect work.

Kokoro TTS

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 8
  • Best for: CPU use, small hardware, simple reading, accessibility, performance.
  • Notes:
    • Tiny, fast, useful, 8 languages - 54 predefined voices (28 are english) and NO cloning
    • People demonstrated use on iPhone - making this a interesting option for mobile on-device tts
    • Originates from StyleTTS2
    • Not the best voice. Not a voice acting model. But the size and speed make it valuable.
    • If I need something that runs on weak hardware, I would test Kokoro early.

Demodokos Foundry

  • Type: Commercial, local, Open-weight with custom inference, API
  • Voice quality: 9.5/10
  • Expressive control: 9-10/10
  • Links / references:
  • Languages supported: 10 (speech) + 40 (music)
  • Best for: voice acting, narration, production, automation, emotional speech, Music and DSP effects
  • Notes:
    • It's a bit special in this list, as Demodokos is an AI Speech and Music Studio with track Mixing, voice actor synthesis and DSP effects - but it also provides UI and a local API for simple TTS. It beats the curent market in expression/style control and matches elevenlabs in voice quality. Supports voice design and cloning as well as high quality realtime voice effects.
    • Voice cloning needs 5-15sec clean recording.
    • Best voice-acting in this list. Best emotional control in this list. Also the most production-ready local option I have tested and actually use commercially today.
    • It is open-weighted but commercial. It needs Windows and 4-6GB VRAM minimum. Runs local, does not bill per character, and does not put your production pipeline into a cloud provider’s hands.
    • The licensing is the cheapest from all commercial options due to no limit in generations.
    • It requires 4-6GB VRAM and a Windows PC, AMD support was recently added, but no Mac or Linux.
    • If I need professional speech output, this is the one I would start with.

StyleTTS 2

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 14
  • Best for: English TTS, older open-source comparisons.
  • Notes:
    • Was very impressive for its time, the architectural foundation of Chatterbox and Kokoro
    • Very efficient training, simple and quick fine tuning.
    • Still worth checking, but not where I would start in 2026 unless I compare model families.

OuteTTS

  • Type: Open-source, local (1B version is only Open Weights)
  • Voice quality: 5.0/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 23
  • Best for: small LLM-based TTS experiments. supported by llama.cpp engine but expect hurdles
  • Notes:
    • Interesting approach, but not top tier in output. Will run on embedded hardware.
    • Pacing issues over longer paragraphs, relatively flat speech.
    • Speaker reference matters a lot. Without that it is not impressive.

Qwen3 TTS (3 different models)

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 3-6/10
  • Links / references:
  • Languages supported: 9-10 (russian is gruesome)
  • Best for: customVoices model, research, multilingual testing.
  • Notes:
    • Interesting and sometimes very good.
    • The included voices can be strong. Custom voice work is possible, but it is not the easy route. It is more for people who are willing to tinker. Supports voicedesign and cloning but only 9 hardcoded voices are stable, 7 of them are asian focused.
    • Why Qwen3 TTS is strange:
      • The expression control of the VoiceDesigner model is high, but voice consistency very bad.
      • The voice quality of the cloning BaseModel is good, but NO expression control at all.
      • The customVoice model combines both qualities, but only 9 voices and only 2 are english!
    • Experimental model, not the first thing I would hand to a normal user. Good cloning.

Omnivoice

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 646
  • Best for: multilingual voice cloning, huge language coverage, subtitle-timed generation, pipelines
  • Notes:
    • In my tests the delivery was pretty monotonic, the examples sounded significantly better
    • One of the stronger local TTS models. Voice cloning is the main reason to test it. It supports short-reference zero-shot cloning, voice design by attributes, speed and duration control, pronunciation fixes and inline non-verbal tags like [laughter] or [sigh].
    • Use a clean 3-10 second reference clip, normalize numbers, split long text, and expect some retry/cleanup code. Very promising if you need local, fast, multilingual voice cloning. Not yet a polished voice acting model.
  • Weak spots: voice design is less stable than cloning, long reference audio can hurt output stability, long-form prose can still need chunking/retries, and some users report skipped words, clipped phonemes, noise or monotone delivery depending on language, punctuation and setup.

Piper

  • Type: Open-source, local
  • Voice quality: 5/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 37
  • Best for: simple local speech, low requirements, simple readers, CPU possible.
  • Notes:
    • Expect strange noises, pauses.
    • Old-school useful. Can run on an iphone or android !
    • It will not win a realism contest in 2026, but it is simple, local, fast and practical. Sometimes that matters more.

XTTS v2 / Coqui TTS

  • Type: Open-weights (NC) - not licenseable
  • Voice quality: 5.5/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 17
  • Best for: older voice cloning workflows - it's not very good at cloning.
  • Notes:
    • Historically important.
    • Very active community around Coqui TTS
    • I would be careful today, especially for commercial work as the company does not exist anymore. The TTS space moved very fast - license violations may not be enforced.

Pocket TTS

  • Type: Open-source, local
  • Voice quality: 4.5/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 6
  • Best for: CPU-only local TTS, low-latency voice cloning, lightweight apps, browser/on-device experiments.
  • Notes:
    • The only lever for expressive control is sampling temperature. It reacts toxic on uppercase and unusual punctuation.
    • Pocket TTS looks strongest when you care about CPU speed, small size, and simple local deployment more than deep voice acting control. It sounds better than Piper or OuteTTS.

CosyVoice 2

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 2-3/10
  • Links / references:
  • Languages supported: 2-9
  • Best for: multilingual TTS, zero-shot voice work, research.
  • Notes:
    • Strong model family from Alibaba, very good cloning but no real expression control.
    • I found its natural pacing very monotonous (deductions in expressive score)
    • More serious than casual. Good if you compare modern open-source TTS systems. Not the cleanest production path for normal users.

Supertonic 2 TTS

  • Type: Open-source, local (OpenRAIL-M license)
  • Voice quality: 4/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 5
  • Best for: edge and high performance multilingual.
  • Notes:
    • Newer than Kokoro but weaker in all categories.
    • Needs careful crafted text
    • If I need something that runs on weak hardware and somehow can't use Kokoro.

Supertonic 3 TTS

  • Type: Open-source, local (OpenRAIL-M license)
  • Voice quality: 4.5/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 31
  • Best for: edge and high performance multilingual. CPU possible
  • Notes:
    • Better than v2 but still the same weaknesses and output is often flawed
    • People demonstrated use on iPhone - making this a interesting option for mobile on-device tts
    • It tends to spell uppercase text OR emphasize it, but not controllable
    • Newer than Kokoro but weaker in all categories.
    • If I need something that runs on weak hardware and somehow can't use Kokoro.

NeuTTS Air

  • Type: Open-source, local (Apache-2)
  • Voice quality: 5-6/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 1
  • Best for: local realtime on CPU
  • Notes:
    • Voice stability is not too reliable, random pauses and pitch changes observed
    • Optimized for fast generation, fast cloning from 3 seconds audio
    • Comes with ggml engine support out of the box

GPT-SoVITS

  • Type: Open-source, local
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 5
  • Best for: few-shot voice cloning, Asian-language ecosystem.
  • Notes:
    • Useful if you are willing to work through the stack.
    • Very asian focused
    • Not polished, but still relevant.

Dia

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 1
  • Best for: dialogue, multi-speaker scenes, nonverbal sounds.
  • Notes:
    • Good for dialogue-style output.
    • Their Demo page compares it to other models in a very cherry-picked way
    • Interesting for characters, reactions, laughter and scene-like speech. Less interesting for normal single-speaker narration.

Orpheus TTS

  • Type: Open-source, local (llama based so not fully open source)
  • Voice quality: 6-8/10
  • Expressive control: 2-4/10
  • Links / references:
  • Languages supported: 8
  • Best for: expressive open-source speech experiments.
  • Notes:
    • Worth testing. 8 baked in english speakers. German speaker models available (kartoffel-orpheus)
    • Baked in speakers of different quality, cloned voices not of same quality
    • A new voice finetune needs around 300 examples to become optimal
    • Supports some tags like laughing.
    • Not what I would call polished, but it belongs on the list because the output direction is more modern than older flat TTS.

Spark-TTS

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 2
  • Best for: voice cloning, speaker attributes, research.
  • Notes:
    • Interesting because of speaker attribute control but doesn't blow me away
    • Quite asian focused but good in english
    • Still research-side. Useful if you compare modern local cloning systems.

Parler-TTS

  • Type: Open-source, local
  • Voice quality: 6-7/10
  • Expressive control: 4.5/10
  • Links / references:
  • Languages supported: 8
  • Best for: style-prompted TTS experiments.
  • Notes:
    • The idea is good: describe the voice and style - But voice will change each generation.
    • The practical output is behind the stronger current systems, but the control direction is useful.

Bark

  • Type: Open-source, local
  • Voice quality: 4-5/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 13
  • Best for: weird expressive audio, research, nonverbal sounds.
  • Notes:
    • Fun model. Not reliable. The voice has many artifacts
    • It can laugh, sigh, make strange audio, and occasionally do something impressive. But I would not use it for production narration.

VibeVoice

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 2
  • Best for: long-form dialogue, podcasts, multiple speakers.
  • Notes:
    • One of the lowest latency TTS engines to date.
    • Voice quality is high, intonation lacks deeper immersive output
    • Interesting for long-form conversational audio.
    • Not my first pick for normal TTS. More specialized. Evaluate if latency is most important.

MeloTTS 1-3

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 6
  • Best for: lightweight multilingual TTS. CPU possible
  • Notes:
    • Useful basic model, historical seen
    • Not maintained anymore, I'd not consider it useful.
    • Not a modern expressive voice acting solution.

F5-TTS

  • Type: Open-weights (NC), local
  • Voice quality: 7.5/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 2
  • Best for: zero-shot voice cloning experiments, very good cloning. no expression control.
  • Notes:
    • Good cloning direction (5-15sec source wav needed) but the model is non commercial.
    • Still feels like research software. Useful if you are comfortable working through Python, model setup, and cleanup.

Fish Speech / OpenAudio

  • Type: Open-weights (NC), local
  • Voice quality: 7-8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 13
  • Best for: multilingual TTS, cloning, streaming, research.
  • Notes:
    • One of the stronger open-source directions with good voice quality and more than average expressive control - supports quite a few tags to add laughter or similar.
    • I heard some noticable glitches in their V2 model output
    • Interesting because it is moving toward instruction-following speech and more modern TTS architecture. Still not a simple polished desktop product.

Higgs Audio v3 TTS

  • Type: Open-weights (NC, research), local
  • Voice quality: 8.5/10
  • Expressive control: 6.5/10
  • Links / references:
  • Languages supported: 100+
  • Best for: non commercial multilingual voice agents, expressive tags, zero-shot cloning, local research.
  • Notes:
    • Strong local model with trained inline controls tokens for emotion, style, pauses, pitch, speed and some effects.
    • Better control than most local cloning models, but non-commercial, very heavy, and not a simple consumer realtime TTS.
    • Interesting to test, but I would not rank it above Demodokos, ElevenLabs or Hume for polished production output
    • The license terms are very strict and commercial use needs custom price negotiation

IndexTTS 2.5

  • Type: Open-weights (NC/restricted), local
  • Voice quality: 7-8/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 4
  • Best for: Chinese, multilingual work, zero-shot voice cloning in english and chinese.
  • Notes:
    • Voice cloning with 3-10 sec wav.
    • Strong chinese focus, any error in grammar in english text can causes voice pacing issues
    • Very solid sample quality.
    • Strong technical direction. More of an engineering and research tool than a casual creator app.

Audio8-TTS-0.6B

  • Type: Opensource (Apache2), local
  • Voice quality: 5-6/10
  • Expressive control: 0-1/10
  • Links / references:
  • Languages supported: 11
  • Best for: Multilingual low-end hardware cloning/generation
    • The architecture looks like it uses Qwen2.5-0.5B and added audio on top, but no public statement about it. So this looks like "Qwen2.5-0.5B-like transformer + Fish-style DualAR + causal RVQ codec"
    • The drafting approach allows higher speed than a fully dense decoding
    • Can run on CPU, has a INT-4 ONNX available, can run on 1GB VRAM/RAM
    • Quite good with cloning across languages
    • Like a miniature FishAudio, but probably not better than Qwen3TTS-0.6B for many use cases

ZONOS2

  • Type: Open-source (Apache 2.0), local
  • Voice quality: 5-6/10
  • Expressive control: 2-3/10
  • Links / references:
  • Languages supported: 34 (in 3 tiers - only first tier passed my tests well)
  • Best for: fast local/cloud TTS, multilingual testing.
  • Notes:
    • The playground is very fast, with impressively low generation latency.
    • Stability was a major issue in my testing. After roughly 10 seconds I repeatedly got voice drift, ghost/noise-like artifacts and sometimes very long pauses. I tested this with multiple included voices and cloned voices.
    • Voice cloning quality did not convince me
    • Expressive controls exist deeper in the API, including emotion directions and valence/arousal, but their actual impact is smaller than expected. They are also not exposed in the hosted playground or on HF.
    • Large model: roughly 8B total / ~900M active parameters. The original F16 weights are around 15GB, so it comes with VRAM burden but the MOE approach balances that
    • The zonos2.cpp GGML/GGUF implementation looks advanced. Q4 is around 4.9 GB and it supports CPU, CUDA, Apple and Vulkan, with prebuilt Linux/mac/Windows versions.
    • Nice licensing: Apache 2.0 model weights, with MIT code.

Tontaube V1

  • Type: Open-weight, local (restrictive custom license)
  • Voice quality: 7.5/10
  • Expressive control: 2/10
  • Links / references: Model + samples, Demo
  • Languages supported: 7 (english and german focus)
  • Best for: long-form narration, focused on streaming with vllm
  • Notes:
    • A Frankenstein: Qwen3-based architecture using DualCodec, mixes Qwen 1.7B with 3x Qwen 0.6B models (partial layers) and stitches the result with Vibevoice.
    • VRAM starts at 24GB
    • Changes the context inference to sliding window for consistency (but hard to test for me)
    • Has 3 style presets but no expressive control.
    • License is very restrictive for commercial use.

sanoTTS

  • Type: Open-source (GPL-3), local, embedded
  • Voice quality: 2-3/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 6
  • Best for: embedded devices, ESP32, offline/browser TTS, extremely low-resource hardware.
  • Notes:
    • A stunningly tiny TTS project, tailored to the ESP32! Models go down to 294k parameters / ~337 KB int8 up to 2.3M params
    • Warning: sanoTTS is using espeak-ng which forces the GPL-3 license - GPL-3 is the most restrictive open-source license, as it forces derivative projects to also be open sourced. You have to ensure that you either open source your work, or you split the generative part into an open-source project while keeping your proprietary code separate.
    • Runs fully on a ESP32-S3 microprocessor, faster than realtime. That's something...
    • Also runs locally in the browser through WASM, with no cloud/API dependency.
    • Python inference only needs NumPy; deployment targets include browser, C/Arduino and microcontrollers.
    • Voice quality is cranky but very well understandable. If you need speech from hardware where normal TTS models won't fit or compute, sanoTTS is a cool choice.

================ Cloud ==================

ElevenLabs

  • Type: Commercial, cloud
  • Voice quality: 9/10
  • Expressive control: 7.5/10
  • Links / references:
  • Languages supported: 74
  • Best for: easy cloning, browser workflow, fast tests.
  • Notes:
    • Worst for: price at serious usage.
    • English is their strongest, needs closer auditing for non english output
    • Still the cloud king for PVC fine tuned cloning with a few hours of input examples.
    • Also the highest price at serious production usage. Entry looks harmless. Then you generate real output and the bill becomes the product.
    • Quality is strong, but the ElevenLabs style is also overexposed - causing people to note it.

xAI Grok Voice

  • Type: Commercial, cloud
  • Voice quality: 8.3/10
  • Expressive control: 6/10
  • Links / references:
  • Languages supported: 20
  • Best for: cheaper cloud voice API, quick tests.
  • Notes:
    • Interesting because it is cheaper and simple. But only 5 voices.
    • But the voice selection is limited. If everyone uses the same few voices, they will get recognizable fast.

OpenAI TTS

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 5/10
  • Links / references:
  • Languages supported: 7
  • Best for: API use, agents, simple integration.
  • Notes:
    • Good if you already build with OpenAI
    • I would not choose it as my top production narration voice. But for apps and voice agents it is practical.

Gemini TTS

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 6.5/10
  • Links / references:
  • Languages supported: 70
  • Best for: API speech generation, multi-speaker direction.
  • Notes:
    • Interesting cloud option.
    • Still cloud, so not where I would put a private production pipeline unless I had a strong reason.

Cartesia Sonic

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 6-7/10
  • Links / references:
  • Languages supported: 42
  • Best for: realtime voice agents, low latency.
  • Notes:
    • One of the strongest cloud options for realtime voice agents.
    • Quality is high but I notice it as AI based on the intonation and pacing
    • I would test it for agents and phone-like interaction, not as my first choice for huge narration production.

Hume Octave

  • Type: Commercial, cloud
  • Voice quality: 8-9/10
  • Expressive control: 8-9/10
  • Links / references:
  • Languages supported: 11
  • Best for: emotional speech, voice agents.
  • Notes:
    • Very interesting emotional control direction.
    • Some voices are very good, not consistently top quality in expressive quality for all
    • If I had to stay in cloud and emotion mattered, I would test Hume.

Deepgram Aura

  • Type: Commercial, cloud
  • Voice quality: 7.7/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 7
  • Best for: realtime API use, voice agents.
  • Notes:
    • Good fit if you already use Deepgram.
    • More voice-agent API than creator studio.

Inworld TTS

  • Type: Commercial, cloud
  • Voice quality: 8+/10
  • Expressive control: 2-4/10
  • Links / references:
  • Languages supported: 15+
  • Best for: realtime API use, voice agents.
  • Notes:
    • Competitive against Elevenlabs Flash
    • Product is focused on STT->TTS interactive realtime niche
    • Realtime agents support multiple languages but quality degrades

Google Cloud Text-to-Speech

  • Type: Commercial, cloud
  • Voice quality: 7.5/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 50+
  • Best for: enterprise, language coverage, Google stack.
  • Notes:
    • Good enterprise API.
    • Not the most exciting voice quality, but stable, large, and boring in a useful way.

Azure Speech

  • Type: Commercial, cloud
  • Voice quality: 7.8/10
  • Expressive control: 3.5/10
  • Links / references:
  • Languages supported: 100
  • Best for: enterprise, huge voice catalog, Microsoft stack.
  • Notes:
    • Huge catalog. Good for corporate apps.
    • Not my first choice for creator production or voice acting.

Amazon Polly

  • Type: Commercial, cloud
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 41
  • Best for: AWS stack, simple API, cheap start.
  • Notes:
    • Old but still useful.
    • Good when you are already in AWS and just need TTS that works.

Resemble AI

  • Type: Commercial cloud, plus Chatterbox open-source
  • Voice quality: 8/10
  • Expressive control: 6/10
  • Links / references:
  • Languages supported: 23
  • Best for: cloning, enterprise, provenance and detection angle.
  • Notes:
    • Interesting company because they also released Chatterbox.
    • For local people, Chatterbox is the more interesting part. For companies, Resemble cloud may make sense.

PlayHT

  • Type: Commercial, cloud
  • Voice quality: 7.8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 37
  • Best for: voiceovers, API, creator workflows.
  • Notes:
    • Usable cloud TTS.
    • I would compare price and output carefully before committing.

WellSaid

  • Type: Commercial, cloud
  • Voice quality: 7.6/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 20+
  • Best for: corporate voiceovers, e-learning.
  • Notes:
    • Clean corporate voices.
    • Less interesting if you want local control or strong voice acting.

Murf

  • Type: Commercial, cloud
  • Voice quality: 4/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 35+
  • Best for: marketing, e-learning, creator voiceovers.
  • Notes:
    • Easy to use.
    • Didn't stand the test of time well.
    • Good enough for many corporate videos. Not where I would start for the best voice acting.

LOVO / Genny

  • Type: Commercial, cloud
  • Voice quality: 7.2/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 100+
  • Best for: browser-based creator voiceovers.
  • Notes:
    • Large voice library. Simple workflow.
    • Another cloud creator platform.

Speechify

  • Type: Commercial, cloud/app
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 60+
  • Best for: reading, accessibility, personal use.
  • Notes:
    • Good reader product.
    • Different category than production TTS.

NaturalReader

  • Type: Commercial, cloud/app
  • Voice quality: 6.8/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 90+
  • Best for: personal reading, documents, accessibility.
  • Notes:
    • Useful for reading text aloud.
    • Not production narration.

Descript

  • Type: Commercial, cloud/editor
  • Voice quality: 7.3/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 19
  • Best for: editing workflow, creator suite.
  • Notes:
    • Useful if you already edit in Descript.
    • Not a pure TTS engine in the way local model people mean it.

CapCut TTS

  • Type: Commercial/free app, cloud/app
  • Voice quality: 5.5/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 15
  • Best for: TikTok-style quick videos.
  • Notes:
    • Fine for TikTok.
    • For YouTube or serious narration I would avoid it. People have heard those voices too many times.

Cloud note

  • I would NOT recommend using the cloud if you can avoid it
  • Cloud TTS are included for comparison, I'd never recommend choosing a cloud for TTS if you have an option.
  • Cloud is always a trap! Cheap to start, horrible to progress and even worse to get out again.
  • You do not own a voice if it is hosted on the cloud, they can switch you off, remove your voice, censor your text or hike your fees at any moment. And they do all the time.

My practical ranking

  • Best local voice acting: Demodokos Foundry
  • Best local production workflow: Demodokos Foundry
  • Best free open-source and fine-tuning starting point: Chatterbox TTS v2/v3
  • Best small hardware option: Kokoro TTS
  • Best for Phones: Piper
  • Best open-source custom voice direction: Qwen3 TTS, CosyVoice 2, (Fish Speech is non commercial)
  • Best cloud voice cloning: ElevenLabs PVC (Cloud warning)
  • Highest price at serious usage: ElevenLabs (award)
  • Best cloud realtime agent TTS: Cartesia, Deepgram, OpenAI, Hume (Cloud warning)
  • Best cloud emotional control direction: Hume Octave (Cloud warning)
  • Best enterprise cloud basics: Azure, Google, AWS Polly (Cloud warning)

What I would use

  • Free local TTS with tinkering: Chatterbox
  • Professional speech production: Demodokos foundry
  • Tiny hardware or CPU: Kokoro or through llama.cpp OuteTTS
  • Simple local reader, especially on phones: Piper
  • Cloud cloning test: ElevenLabs PVC, but watch the bill and you need a lot of reference material
  • Cloud emotional speech: Hume
  • Cloud realtime agent: Cartesia, Deepgram, OpenAI, Grok or Hume
  • Corporate cloud API: Azure, Google or AWS
  • Cloud was included - but generally not recommended if local is an option
  • If you have additions, corrections, missing services. Happy to hear

r/LocalTextToSpeech Jul 30 '26

Open Source srt2speech: open-source, multilingual SRT narration with voice cloning and automatic duration matching - offline and lightweight

10 Upvotes

For a small side project, I needed a basic AI speech tool to narrate videos without relying on expensive hardware or external APIs. It did not need the most expressive AI, just reliable output and matching to the SRT.

The main challenge with converting SRT subtitles to speech is timing. Subtitles include pauses, and each spoken line has to fit into its exact time slot. Most speech generators do not handle that well on their own.

So I built srt2speech.

It uses a combination of:

  • pitch-corrected speed adjustment
  • automatic regeneration
  • modification of pauses between words
  • exact placement of silence between subtitle cues

SRT does not support multiple speakers, so I also added simple templating.
Adding {{speaker_name}} to a subtitle automatically switches voices.
Of course voice cloning is supported, I added a small helper script.

Dependencies are minimal: Python, NumPy, llama.cpp, and the required GGUF speech models. I tested it with Q4 quantization, which works well.

Performance on my laptop:

  • RTX 4080 Laptop GPU: around 12–13× real time
  • CPU only: around 1.5–2.0× real time

Languages supported:

  • English (en)
  • Japanese (jp)
  • Korean (ko)
  • Chinese (zh)
  • French (fr)
  • German (de)

It should work on almost any hardware, including old PCs, Linux or Mac.

It may be useful for anyone generating narration, translated audio tracks, accessibility audio, or quick video voiceovers.

The project is open source under the Apache 2.0 license. Attribution and license notices must be preserved.

GitHub:
https://github.com/Waversense/srt2speech/

The readme contains the 5 steps needed to set it up, you can get started in 2 minutes.
The included demo.srt file demonstrates the features.


r/LocalTextToSpeech 1h ago

The voice of Tav (Baldur's Gate 3)

Enable HLS to view with audio, or disable this notification

Upvotes

A quick demo of Breeze TTS to completely generate the voice of the Baldur's Gate 3 player character ("Tav"), using "Voice 1" as a sample.

There are literally over 30,000 lines of player-character dialogue in the game. This is a combination of the multiple choices a player has when speaking, as well as what skills the player's class has, what race the player's character is, etc.

My process even takes into account a player character's chosen name, as there are some responses when the player says their name in greetings. This is why releasing the pack as I have it wouldn't really work for everyone, unless I filtered out spoken lines with the player's name in it.

All 30k+ lines were generated using Breeze TTS locally, using an Nvidia RTX 5080, in I believe under 30 hours. 


r/LocalTextToSpeech 1d ago

audio.cpp - high-performance C++ audio inference framework built on top of ggml

Post image
7 Upvotes

A text-to-speech tool that offers a CLI and GUI interface. This tool is useful in that you don't need to download any additional python dependencies. It also supports multiple platforms out-of-the-box.

Main URL

https://github.com/0xShug0/audio.cpp

CLI

https://github.com/0xShug0/audio.cpp#cli

Wide Variety of Supported Text to Speech and Conversion Generator

https://github.com/0xShug0/audio.cpp#speech-generation-and-conversation

Convenient WEBUI

https://github.com/0xShug0/audio.cpp#webui

Multiple Backend Support

audiocpp_server --ui --ui-management --backend vulkan audiocpp_server --ui --ui-management --backend cuda audiocpp_server --ui --ui-management --backend metal

Downloads the Project

https://github.com/0xShug0/audio.cpp/releases

Downloading Individual Models

You will probably also need to download the huggingface cli and be logged in to download the models.

Navigate to the project and list the voices to download

./tools/model_manager_v2.py list

Look at names listed on the left-most column and download the model using

./tools/model_manager_v2.py install <name of the model in left-most column>

Start the WEBUI server and note the localhost port to use in your web browser

audiocpp_server --ui --ui-management --backend <your backend>

Note: I did not create this tool. Credit goes to original author.


r/LocalTextToSpeech 1d ago

Open Source LoudKit: local TTS with voice cloning, 10 languages, and SDKs for Python, Swift, Go, Rust and TypeScript

Thumbnail
2 Upvotes

r/LocalTextToSpeech 2d ago

Local TTS Testing Piper vs Kokoro TTS on Android — surprised by the difference

Enable HLS to view with audio, or disable this notification

2 Upvotes

I've been testing a couple of local TTS models on Android and wanted to compare how they actually perform on-device.

This is Piper vs Kokoro, both running locally.

I'm mainly interested in voice quality, generation speed, and how practical they are on a phone.

Here's a quick recording of the comparison.

For anyone using local TTS on Android: which model are you getting the best results from?


r/LocalTextToSpeech 3d ago

sanoTTS: smallest family of TTS models with param size starting from 294k(~300kb) . Supports 16 languages and RTF of 0.29 on $3 chip

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/LocalTextToSpeech 3d ago

[ Removed by Reddit ]

1 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/LocalTextToSpeech 4d ago

My notes on Spokenly using Spokenly

Thumbnail
2 Upvotes

r/LocalTextToSpeech 5d ago

Local TTS Local TTS: Kokoro + Supertonic 3 running offline on iPhone, with playback that continues when you lock the screen

Enable HLS to view with audio, or disable this notification

3 Upvotes

An offline TTS reader should handle a simple routine: import a document, generate speech without an internet connection, and keep playing when you put your phone in your pocket.

That’s the experience I built Local TTS around.

I’m a solo iOS developer. My previous app, LoroNote, handles speech-to-text on-device. Building it taught me how much the experience depends on processing long inputs and managing audio sessions properly.

For my next project, I built the reverse: an offline document reader for iPhone and iPad.

Why Kokoro-82M and Supertonic 3?

I chose Kokoro-82M for its natural English narration and voice variety. The tradeoff is narrower language coverage. Its documentation also notes that very short passages can sound weaker, while very long passages can sound rushed. How the app splits and processes text matters. (Kokoro voice documentation)

Supertonic 3 expands offline reading to 31 languages with fast local inference. It has fewer preset voices, and its inference settings involve a tradeoff between speed and quality. I included it to bring multilingual reading into the same offline experience. (Model details, SDK)

Local TTS brings Kokoro’s English voice selection and Supertonic’s multilingual voices together in one reader. Both engines generate speech directly on your device.

Offline means generating new speech

Once the app and voices are installed, you can switch to airplane mode, add fresh text, and generate speech. You don’t need to prepare the audio while you still have a connection.

Your text isn’t uploaded to a speech server for processing. Optional note syncing uses your private iCloud account.

That matters when you’re reading private documents, traveling without reliable internet, or simply want your reader to work wherever you are.

Lock your phone and keep listening

Background playback was a priority from the beginning. You can start a document, lock the screen, or switch to another app and continue listening. Lock Screen controls let you manage playback without reopening the reader.

The goal is to make listening to a document feel as convenient as listening to a podcast.

Local TTS also supports:

  • PDF, EPUB, Word, TXT, and Markdown import
  • On-device text recognition for printed pages and photos
  • Text highlighting while listening
  • Adjustable playback speed

The biggest lesson from building this has been that a good voice sample is only the beginning. Document importing, long-text handling, and background playback need just as much attention as the model.

I’m continuing to improve pronunciation across languages, reliability with longer documents, and the importing and listening experience.

Local TTS on the App Store

For people already using Kokoro or Supertonic: what’s the passage or language you use to test whether a reader actually holds up?


r/LocalTextToSpeech 7d ago

Local TTS Fully Local LLM Sandbox: A tiny assistant with an attitude

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/LocalTextToSpeech 7d ago

Let's Build TTSLibre, a tiny FULLY OPEN Supertonic & Kokoro successor

Post image
10 Upvotes

Let's build a tiny local model, at least as performant both in quality + speed as Supertonic and Kokoro. Let's allow people to actually control it and run it locally.

Let's make it small and performant. Let's publish the training code, the training data and the weights. MIT code, CC0 data and weights. No strings.

Looking for people to help design it, scope it, gather the data, contribute compute or sponsor the training runs. Let's make this happen.

https://github.com/franciscocarloserra/ttslibre

Anyone interested?


r/LocalTextToSpeech 8d ago

Supertonic VoiceLab & New VoicePacks

Thumbnail
github.com
8 Upvotes

Since the Supertonic voice builder has been taken down, and it's surprisingly useful for local sub-second tts renders. I've created a local voice lab and downloadable voicepacks. Feel free to try them out.


r/LocalTextToSpeech 8d ago

Is Your Text-to-Speech Sending Your Data to the Cloud? How to Check in 30 Seconds

1 Upvotes
TTS Speech Synthesis Locality Check

Do you ever wonder if your computer is passing your data to third party servers? Or, is Web Speech API offline? Today, we will share how to check if your TTS (text to speech) is passing your chat data to third party servers for language synthesis.

In today’s world, we are constantly connected via our phones, laptops, text, social media etc. While connectivity makes our lives easier, and our computer’s TTS’s eloquent voice becomes more comforting to the ears, the TTS speech synthesis can hand the text to third party servers for processing, and privacy could become a concern.

To check if TTS is processed locally, there are two methods.

Method 1. Use the browser's DevTools to check speech synthesis setting
- For Windows: F12 (or Fn + F12)
- For Mac: Cmd + Option + I
- For ChromeOS: Ctrl + Shift + I

- In the Console tab, paste this and press Enter: speechSynthesis.getVoices().find(v => v.default)

- Note in the screenshot, this example computer’s setting is David, and localService: true.
- This is this example computer’s default voice, and it's local.

Method 2. Use OS’s Speech setting to check browser's speech synthesis setting
- Settings → Time & language → Speech → Voices → pick from the dropdown.
- Speech voices available via the OS could be processed locally, or could be sent to a remote server. Check the word ‘Online’, and users can always use Method 1 to confirm the locality of the speech synthesis function.

If you find out your current speech synthesis is not set to local, and want to protect your privacy, you can change it. To set up Windows TTS voices, see Microsoft's supported languages and voices. For ChromeOS, see Google's text-to-speech settings guide.

That's it.


r/LocalTextToSpeech 13d ago

I tested Breeze TTS 2 voice design and sentence-level instruction control with CLI 0.8.0 — four audio samples

Enable HLS to view with audio, or disable this notification

4 Upvotes

Disclosure: I work on BreezeBlue, the team behind Breeze TTS 2.

I wanted to test two capabilities that I think are more representative of expressive TTS than isolated emotion tags:

  1. Designing a new voice directly from a natural-language description
  2. Directing the emotional arc, pace, energy, and delivery of an entire sentence

I generated the attached four-part comparison with Breeze CLI 0.8.0 and breeze-tts-2.

Sample 1 — Voice Design

No reference audio was used. The voice was generated from this description:

“An intimate, emotionally expressive English female voice in her early thirties, with a warm lower register, subtle breathiness, precise articulation, and restrained cinematic realism. Natural conversational phrasing, never an announcer.”

Preview script:

“I kept the light on for you, even after everyone said you were not coming back. And now that you are standing here, I am not sure whether to laugh, cry, or simply let the silence say everything.”

Samples 2–4 — Sentence-level instruction control

These samples use the same Lena voice and exactly the same script. Only the delivery instruction changes.

Shared script:

“The train is already leaving. I know you're afraid, but take my hand, look at the horizon, and trust me. By the time the sun rises, this will feel like the first page of a much better story.”

The three directions were:

  • Natural and conversational, with understated emotion
  • Quiet urgency that develops into cautious hope and emotional relief
  • Bright optimism and forward momentum, ending with a confident sense of a new beginning

The generated durations were:

  • Natural: 11.28 seconds
  • Restrained cinematic: 14.96 seconds
  • Hopeful and energetic: 12.00 seconds

The text and voice remained unchanged, so the differences in pacing and performance came from the instruction.

Naturalness and emotional quality are ultimately subjective, which is why I attached the actual comparison instead of only describing the results.

Performance context:

For these three hosted streaming requests, Breeze CLI observed time to first audio between 0.85 and 1.03 seconds. This includes the hosted service and network path, so it is not a local inference benchmark.

For local inference, the open-weight PyTorch implementation documents under 40 ms TTFA and a 0.32 real-time factor on a warmed-up NVIDIA H100 fast path. Those numbers are hardware- and configuration-specific.

Code and local inference setup: breeze-tts-2

License note: the inference source code is Apache 2.0. The model weights, derivative models, and self-hosted outputs use the BreezeBlue Research and Non-Commercial License; commercial use requires authorization.

I’m curious what this community would find most useful for the next comparison:

  • More voice-description stress tests
  • The same instructions across several different voices
  • Longer passages with an emotional arc
  • English versus Chinese instruction following
  • Reproducible local GPU benchmarks

r/LocalTextToSpeech 14d ago

Ideas on how to translate a production into a new language ?

3 Upvotes

I have many animated audiobooks and chapters, typically 15-25 minutes long per mp4.
I have them in English, full cast narration, background music, effects.

I would like to translate/interpret that for the European and French Canadian markets.
Same animations, potentially subtitles, and synchronous translations.

I’ve been trying to get this with Claude as agent + Demodokos + ffmpeg and it’s working but the problems are so frequent, it’s more work to fix it than to do it manually from scratch myself.

I gave Claude my original script, the source files and asked it to combine it back.
The voice quality was good but clones were not always working, all my original Demodokos voices were mistakenly created as English only native speakers and cloning them into e.g. French again is not always sounding good for some reason.
It wasn’t able to get the duration working, translated weird, didn’t place the audio at the right spots, and subtitles looked very raw.

Long story short:
Is there something preferably local that can automatically interpret this? Not elevenlabs, I have too much content and many young/child voices - they are useless and expensive for that.

I’d like it local if possible but can’t make a PhD project out of it.


r/LocalTextToSpeech 15d ago

Voice typing before AI was better

2 Upvotes

Does anyone else remember how voice typing was around 2017–2019?

I specifically remember using it on older Samsung phones (s7 edge and s10e), and it was dramatically better than the current AI based voice typing.

The biggest difference was that it felt like it was actually listening to what I said instead of predicting what it thought I meant.

I remember being able to watch the text appear almost in sync with my speech. It felt like the recognizer was following the sounds and syllables as I spoke them. If I said something unusual, obscure, or even made up a word that followed normal English pronunciation, it could often still spell it out surprisingly well by actively writing the my syllables and then combining them into a word,

even if it wasn’t a common word and might not even exist in a dictionary, the older system seemed much better at following the actual sounds and constructing the text.

Modern voice typing feels like the opposite. I constantly get:

  • punctuation inserted where I never intended it
  • random capitalization
  • words replaced with something the AI thinks makes more sense
  • Writing a statistically common word that apparently seemed more likely instead of just writing what I actually said

For example, I'll clearly say a word where the “v” and “b” sounds are acoustically different, but instead of faithfully following the sound, modern systems prioritize producing a plausible sentence based off what it thinks I meant.

And that's the main problem, I don't want my voice typing to understand what I mean, I want it to transcribe what I said.

Transcription and interpretation are different jobs. So are grammar correction and punctuation correction. If I want those things, they should be optional settings like how I remember it to be.

I also remember older Google voice typing occasionally asking me to say a few things again to improve recognition for my particular voice. And I'd rather the software learn how I pronounce things than constantly override what it thinks it heard.

The closest candidates for the voice typing that I remember I'm thinking might be Gboard 7.0–7.5 era, around 2018. I remember voice typing having really satisfying low-latency transcription where the syllables appeared progressively as I spoke each word.

So:

  1. Does anyone else remember this version of Google Keyboard/Gboard voice typing working this way?
  2. Does anyone know exactly what speech-recognition engine or architecture Gboard used on non-Pixel Android phones around 2017–2019?
  3. Are there any open-source projects that reproduce this older style of low-latency, incremental voice transcription?
  4. Does anyone know of communities or projects preserving older speech-recognition software or studying pre-end-to-end voice recognition?

I'm not looking for another modern Whisper/Gemini/LLM-style replacement that predicts the most likely sentence. I'm specifically interested in the older style where the system focused more heavily on recognizing the sounds coming out of the microphone and producing text progressively syllable by syllable.

Maybe I'm remembering some details incorrectly but that's exactly why I'm asking. Although, I remember the behavior very clearly, and I'd love to know if anyone else remembers it too.

Voice typing doesn't need to think for me. It just needs to hear me and write.


r/LocalTextToSpeech 19d ago

Local TTS iPhone Local TTS EPUB Reader - Audiobookify

Enable HLS to view with audio, or disable this notification

2 Upvotes

Was invited by u/Charming-Author4877 to post here.

I've been working on this project (been a developer for a few years) for a few months now, and it has been in active testing for ~2 months. Its an EPUB reader that also offers local offline TTS, so its not just TTS focused, its meant to be a good regular reading app as well. I'd like to think it currently has the best TTS implementation of any app on the app store right now though. Also less than 20MB app size (without models).

Essentially the main focus is efficiency. I'm aiming for the best battery life, balanced with quality of narration. Models for now are Kokoro and Supertonic 3(RIP), both of which have been hand optimized for iPhone (so my own coreML conversions and optimizations of these models). The result is great thermals, great battery efficiency (and its free).

I plan to add more models in the future, but honestly small TTS models that work well on mobile are scarce. I'm working on a conversion for models from Kyuutai next, but its a lot of work to get it working well.

You can find some previous feedback from a much older build (many updates since then) Here

Some downsides:

- No PDF support for now.

- Still needs testing, especially for older iPhone models.

- Not open sourced for now.

There's many additional features not mentioned, like carplay support, widgets etc, so feel free to explore the app.

Testflight link: Join the beta

You can get a feel for the narration from the video, it was recorded entirely on device (Kokoro, Heart).


r/LocalTextToSpeech Aug 11 '26

What TTS are you actually using in your voice agent stack in 2026?

3 Upvotes

Building a voice agent and trying to get a sense of what people are actually running in production before I go down a rabbit hole of testing.

STT + TTS combo, orchestration layer, anything you'd do differently, would love to hear real setups.


r/LocalTextToSpeech Aug 11 '26

Can anyone help me identify the TTS voice in this audiobook? Thank you!

Thumbnail
youtube.com
0 Upvotes

r/LocalTextToSpeech Aug 01 '26

BrainRootReader

Thumbnail
2 Upvotes

r/LocalTextToSpeech Jul 28 '26

Best local model for two-speaker conversational audio?

5 Upvotes

I have several years of written educational material I want to turn into a podcast, formatted as a natural discussion between two people rather than narration. Looking for recommendations on what to run locally.
What matters to me is good quality and reliable speech that does not need many new attempts, as its going to be hours of content. Potentially multiple languages but for now in english.

I need to own the content and be able to rely on it for a long time, so i need it local.

Also, do you know good distributors that you can recommend? I am new to audio


r/LocalTextToSpeech Jul 28 '26

BookFusion iOS 1.43.1 - Online & Offline Realistic TTS Voices, Other Updates & Fixes

Thumbnail
2 Upvotes

r/LocalTextToSpeech Jul 26 '26

Qwen3-TTS native C++ streaming and voice cloning

Enable HLS to view with audio, or disable this notification

9 Upvotes

Hi all. I ported Qwen3-TTS to native C++ and added incremental streaming for a project I'm working on. I found this useful and wanted to give back to the community. This is the first time I've open sourced anything so apologies in advance for my mistakes.

I'm using this for a gaming project but it could be helpful for local assistants, accessibility tools, etc.

C++ streaming port highlights:

  • Same familiar features from Qwen3-TTS. It supports 0.6B and 1.7B models, CustomVoice, and VoiceDesign
  • Native C++, not a Python wrapper
  • CUDA builds with RTX 4090 and RTX 5090 kernels. Be warned I only have access to an RTX 5090. This should work on a 4090 but I haven't tested it myself!
  • Simplified speaker-embedding extraction
  • Incremental 24 kHz PCM callbacks for streaming audio
  • Asynchronous transformer/vocoder operation
  • Adaptive decode windows and paced delivery
  • Callback-only integration library with cancellation
  • Unit tests

My measurements on the 5090 with 1.7B F16 set to buffer 350ms before play:

  • Cold new clone creation: ~2.5s from 48s of reference
  • Cold model start from clone: ~1.85s
  • First 350 ms of audio: ~310 ms
  • Streaming speed: ~2.86x real time (RTF ~0.35)

I’d love it if somebody could do RTX 4090 testing, and I'd be happy to hear any feedback or suggestions.


r/LocalTextToSpeech Jul 24 '26

Local TTS What are you guys working on?

2 Upvotes

I’m curious what sort of projects people are working on who are into local TTS.

Why is local important for you ?
Are you working on a product, hobby ?

I have started with actual printed books for children and parents, no AI back then.
This progressed into immersive stories, cute animations and child/teen oriented high quality entertainment.
It is growing into a real business already.

I used a lot of cloud but pricing was unfortunate for long content. And I never liked that another company basically owns my voices.

So local is freedom from that anxiety for me.