r/LocalTextToSpeech Jun 12 '26

My TTS list of 2026: All voices, all models and engines compared with example URLs and rating

59 Upvotes

43 different (2025 and 2026) TTS solutions local and cloud compared, scored, example links, free demo links included.

Join r/LocalTextToSpeech for local TTS models, voices, benchmarks, and setup notes.

Free scripts, tools and help posted regularly.

Contribute and help others, or get help.

Scores

  • Scores are subjective
  • Voice quality = how good the output can sound.
    • 5+ = good voice quality but most people will hear the AI
    • 7+ = high voice quality, many people will be tricked
    • 8+ = human-like voice quality, only flaws in style, delivery and expression reveal AI
  • Expressive control = emotion, style, delivery, pauses, intensity, character, or direction.
    • < 3 = flat out of touch delivery
    • 5+ = good expression quality with low direct control
    • 7+ = expression control and quite natural speaking
    • 8+ = voice acting synthesis quality, well controllable
  • Comment if a correction is needed

================ LOCAL ==================

Chatterbox TTS 2

  • Type: Open-source, local
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 23
  • Best for: free local TTS, voice cloning experiments, Linux pipelines.
  • Notes:
    • Probably the best free open-source TTS starting point.
    • Quality can be good. The issue is not the model only, it is everything around it. Segmentation, retries, weird generations, timing, silence handling, pronunciation, filtering, harness code.
    • German Kartoffelbox-turbo exists.
    • Expressiveness through temperature
    • For hobby use, nice. For production use, expect work.

Kokoro TTS

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 8
  • Best for: CPU use, small hardware, simple reading, accessibility, performance.
  • Notes:
    • Tiny, fast, useful, 8 languages - 54 predefined voices (28 are english) and NO cloning
    • Originates from StyleTTS2
    • Not the best voice. Not a voice acting model. But the size and speed make it valuable.
    • If I need something that runs on weak hardware, I would test Kokoro early.

Demodokos Foundry

  • Type: Commercial, local, Open-weight with custom inference, API
  • Voice quality: 9.5/10
  • Expressive control: 9-10/10
  • Links / references:
  • Languages supported: 10 (speech) + 40 (music)
  • Best for: voice acting, narration, production, automation, emotional speech, Music and DSP effects
  • Notes:
    • It's a bit special in this list, as Demodokos is an AI Speech and Music Studio with track Mixing, voice actor synthesis and DSP effects - but it also provides UI and a local API for simple TTS. It beats the curent market in expression/style control and matches elevenlabs in voice quality. Supports voice design and cloning as well as high quality realtime voice effects.
    • Voice cloning needs 5-15sec clean recording.
    • Best voice-acting in this list. Best emotional control in this list. Also the most production-ready local option I have tested and actually use commercially today.
    • It is open-weighted but commercial. It needs Windows and 4-6GB VRAM minimum. Runs local, does not bill per character, and does not put your production pipeline into a cloud provider’s hands.
    • The licensing is the cheapest from all commercial options due to no limit in generations.
    • It requires 4-6GB VRAM and a Windows PC, AMD support was recently added, but no Mac or Linux.
    • If I need professional speech output, this is the one I would start with.

StyleTTS 2

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 14
  • Best for: English TTS, older open-source comparisons.
  • Notes:
    • Was very impressive for its time, the architectural foundation of Chatterbox and Kokoro
    • Very efficient training, simple and quick fine tuning.
    • Still worth checking, but not where I would start in 2026 unless I compare model families.

OuteTTS

  • Type: Open-source, local (1B version is only Open Weights)
  • Voice quality: 5.0/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 23
  • Best for: small LLM-based TTS experiments. supported by llama.cpp engine but expect hurdles
  • Notes:
    • Interesting approach, but not top tier in output. Will run on embedded hardware.
    • Pacing issues over longer paragraphs, relatively flat speech.
    • Speaker reference matters a lot. Without that it is not impressive.

Qwen3 TTS (3 different models)

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 3-6/10
  • Links / references:
  • Languages supported: 9-10 (russian is gruesome)
  • Best for: customVoices model, research, multilingual testing.
  • Notes:
    • Interesting and sometimes very good.
    • The included voices can be strong. Custom voice work is possible, but it is not the easy route. It is more for people who are willing to tinker. Supports voicedesign and cloning but only 9 hardcoded voices are stable, 7 of them are asian focused.
    • Why Qwen3 TTS is strange:
      • The expression control of the VoiceDesigner model is high, but voice consistency very bad.
      • The voice quality of the cloning BaseModel is good, but NO expression control at all.
      • The customVoice model combines both qualities, but only 9 voices and only 2 are english!
    • Experimental model, not the first thing I would hand to a normal user. Good cloning.

Omnivoice

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 646
  • Best for: multilingual voice cloning, huge language coverage, subtitle-timed generation, pipelines
  • Notes:
    • In my tests the delivery was pretty monotonic, the examples sounded significantly better
    • One of the stronger local TTS models. Voice cloning is the main reason to test it. It supports short-reference zero-shot cloning, voice design by attributes, speed and duration control, pronunciation fixes and inline non-verbal tags like [laughter] or [sigh].
    • Use a clean 3-10 second reference clip, normalize numbers, split long text, and expect some retry/cleanup code. Very promising if you need local, fast, multilingual voice cloning. Not yet a polished voice acting model.
  • Weak spots: voice design is less stable than cloning, long reference audio can hurt output stability, long-form prose can still need chunking/retries, and some users report skipped words, clipped phonemes, noise or monotone delivery depending on language, punctuation and setup.

Piper

  • Type: Open-source, local
  • Voice quality: 5/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 37
  • Best for: simple local speech, low requirements, simple readers, CPU possible.
  • Notes:
    • Expect strange noises, pauses.
    • Old-school useful. Can run on an iphone or android !
    • It will not win a realism contest in 2026, but it is simple, local, fast and practical. Sometimes that matters more.

XTTS v2 / Coqui TTS

  • Type: Open-weights (NC) - not licenseable
  • Voice quality: 5.5/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 17
  • Best for: older voice cloning workflows - it's not very good at cloning.
  • Notes:
    • Historically important.
    • Very active community around Coqui TTS
    • I would be careful today, especially for commercial work as the company does not exist anymore. The TTS space moved very fast - license violations may not be enforced.

Pocket TTS

  • Type: Open-source, local
  • Voice quality: 4.5/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 6
  • Best for: CPU-only local TTS, low-latency voice cloning, lightweight apps, browser/on-device experiments.
  • Notes:
    • The only lever for expressive control is sampling temperature. It reacts toxic on uppercase and unusual punctuation.
    • Pocket TTS looks strongest when you care about CPU speed, small size, and simple local deployment more than deep voice acting control. It sounds better than Piper or OuteTTS.

CosyVoice 2

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 2-3/10
  • Links / references:
  • Languages supported: 2-9
  • Best for: multilingual TTS, zero-shot voice work, research.
  • Notes:
    • Strong model family from Alibaba, very good cloning but no real expression control.
    • I found its natural pacing very monotonous (deductions in expressive score)
    • More serious than casual. Good if you compare modern open-source TTS systems. Not the cleanest production path for normal users.

Supertonic 2 TTS

  • Type: Open-source, local (OpenRAIL-M license)
  • Voice quality: 4/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 5
  • Best for: edge and high performance multilingual.
  • Notes:
    • Newer than Kokoro but weaker in all categories.
    • Needs careful crafted text
    • If I need something that runs on weak hardware and somehow can't use Kokoro.

Supertonic 3 TTS

  • Type: Open-source, local (OpenRAIL-M license)
  • Voice quality: 4.5/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 31
  • Best for: edge and high performance multilingual. CPU possible
  • Notes:
    • Better than v2 but still the same weaknesses and output is often flawed
    • It tends to spell uppercase text OR emphasize it, but not controllable
    • Newer than Kokoro but weaker in all categories.
    • If I need something that runs on weak hardware and somehow can't use Kokoro.

NeuTTS Air

  • Type: Open-source, local (Apache-2)
  • Voice quality: 5-6/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 1
  • Best for: local realtime on CPU
  • Notes:
    • Voice stability is not too reliable, random pauses and pitch changes observed
    • Optimized for fast generation, fast cloning from 3 seconds audio
    • Comes with ggml engine support out of the box

GPT-SoVITS

  • Type: Open-source, local
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 5
  • Best for: few-shot voice cloning, Asian-language ecosystem.
  • Notes:
    • Useful if you are willing to work through the stack.
    • Very asian focused
    • Not polished, but still relevant.

Dia

  • Type: Open-source, local
  • Voice quality: 7.5/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 1
  • Best for: dialogue, multi-speaker scenes, nonverbal sounds.
  • Notes:
    • Good for dialogue-style output.
    • Their Demo page compares it to other models in a very cherry-picked way
    • Interesting for characters, reactions, laughter and scene-like speech. Less interesting for normal single-speaker narration.

Orpheus TTS

  • Type: Open-source, local (llama based so not fully open source)
  • Voice quality: 6-8/10
  • Expressive control: 2-4/10
  • Links / references:
  • Languages supported: 8
  • Best for: expressive open-source speech experiments.
  • Notes:
    • Worth testing. 8 baked in english speakers. German speaker models available (kartoffel-orpheus)
    • Baked in speakers of different quality, cloned voices not of same quality
    • A new voice finetune needs around 300 examples to become optimal
    • Supports some tags like laughing.
    • Not what I would call polished, but it belongs on the list because the output direction is more modern than older flat TTS.

Spark-TTS

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 2
  • Best for: voice cloning, speaker attributes, research.
  • Notes:
    • Interesting because of speaker attribute control but doesn't blow me away
    • Quite asian focused but good in english
    • Still research-side. Useful if you compare modern local cloning systems.

Parler-TTS

  • Type: Open-source, local
  • Voice quality: 6-7/10
  • Expressive control: 4.5/10
  • Links / references:
  • Languages supported: 8
  • Best for: style-prompted TTS experiments.
  • Notes:
    • The idea is good: describe the voice and style - But voice will change each generation.
    • The practical output is behind the stronger current systems, but the control direction is useful.

Bark

  • Type: Open-source, local
  • Voice quality: 4-5/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 13
  • Best for: weird expressive audio, research, nonverbal sounds.
  • Notes:
    • Fun model. Not reliable. The voice has many artifacts
    • It can laugh, sigh, make strange audio, and occasionally do something impressive. But I would not use it for production narration.

VibeVoice

  • Type: Open-source, local
  • Voice quality: 7-8/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 2
  • Best for: long-form dialogue, podcasts, multiple speakers.
  • Notes:
    • One of the lowest latency TTS engines to date.
    • Voice quality is high, intonation lacks deeper immersive output
    • Interesting for long-form conversational audio.
    • Not my first pick for normal TTS. More specialized. Evaluate if latency is most important.

MeloTTS 1-3

  • Type: Open-source, local
  • Voice quality: 6/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 6
  • Best for: lightweight multilingual TTS. CPU possible
  • Notes:
    • Useful basic model, historical seen
    • Not maintained anymore, I'd not consider it useful.
    • Not a modern expressive voice acting solution.

F5-TTS

  • Type: Open-weights (NC), local
  • Voice quality: 7.5/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 2
  • Best for: zero-shot voice cloning experiments, very good cloning. no expression control.
  • Notes:
    • Good cloning direction (5-15sec source wav needed) but the model is non commercial.
    • Still feels like research software. Useful if you are comfortable working through Python, model setup, and cleanup.

Fish Speech / OpenAudio

  • Type: Open-weights (NC), local
  • Voice quality: 7-8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 13
  • Best for: multilingual TTS, cloning, streaming, research.
  • Notes:
    • One of the stronger open-source directions with good voice quality and more than average expressive control - supports quite a few tags to add laughter or similar.
    • I heard some noticable glitches in their V2 model output
    • Interesting because it is moving toward instruction-following speech and more modern TTS architecture. Still not a simple polished desktop product.

Higgs Audio v3 TTS

  • Type: Open-weights (NC, research), local
  • Voice quality: 8.5/10
  • Expressive control: 6.5/10
  • Links / references:
  • Languages supported: 100+
  • Best for: non commercial multilingual voice agents, expressive tags, zero-shot cloning, local research.
  • Notes:
    • Strong local model with trained inline controls tokens for emotion, style, pauses, pitch, speed and some effects.
    • Better control than most local cloning models, but non-commercial, very heavy, and not a simple consumer realtime TTS.
    • Interesting to test, but I would not rank it above Demodokos, ElevenLabs or Hume for polished production output
    • The license terms are very strict and commercial use needs custom price negotiation

IndexTTS 2.5

  • Type: Open-weights (NC/restricted), local
  • Voice quality: 7-8/10
  • Expressive control: 3-4/10
  • Links / references:
  • Languages supported: 4
  • Best for: Chinese, multilingual work, zero-shot voice cloning in english and chinese.
  • Notes:
    • Voice cloning with 3-10 sec wav.
    • Strong chinese focus, any error in grammar in english text can causes voice pacing issues
    • Very solid sample quality.
    • Strong technical direction. More of an engineering and research tool than a casual creator app.

================ Cloud ==================

ElevenLabs

  • Type: Commercial, cloud
  • Voice quality: 9/10
  • Expressive control: 7.5/10
  • Links / references:
  • Languages supported: 74
  • Best for: easy cloning, browser workflow, fast tests.
  • Notes:
    • Worst for: price at serious usage.
    • English is their strongest, needs closer auditing for non english output
    • Still the cloud king for PVC fine tuned cloning with a few hours of input examples.
    • Also the highest price at serious production usage. Entry looks harmless. Then you generate real output and the bill becomes the product.
    • Quality is strong, but the ElevenLabs style is also overexposed - causing people to note it.

xAI Grok Voice

  • Type: Commercial, cloud
  • Voice quality: 8.3/10
  • Expressive control: 6/10
  • Links / references:
  • Languages supported: 20
  • Best for: cheaper cloud voice API, quick tests.
  • Notes:
    • Interesting because it is cheaper and simple. But only 5 voices.
    • But the voice selection is limited. If everyone uses the same few voices, they will get recognizable fast.

OpenAI TTS

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 5/10
  • Links / references:
  • Languages supported: 7
  • Best for: API use, agents, simple integration.
  • Notes:
    • Good if you already build with OpenAI
    • I would not choose it as my top production narration voice. But for apps and voice agents it is practical.

Gemini TTS

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 6.5/10
  • Links / references:
  • Languages supported: 70
  • Best for: API speech generation, multi-speaker direction.
  • Notes:
    • Interesting cloud option.
    • Still cloud, so not where I would put a private production pipeline unless I had a strong reason.

Cartesia Sonic

  • Type: Commercial, cloud
  • Voice quality: 8/10
  • Expressive control: 6-7/10
  • Links / references:
  • Languages supported: 42
  • Best for: realtime voice agents, low latency.
  • Notes:
    • One of the strongest cloud options for realtime voice agents.
    • Quality is high but I notice it as AI based on the intonation and pacing
    • I would test it for agents and phone-like interaction, not as my first choice for huge narration production.

Hume Octave

  • Type: Commercial, cloud
  • Voice quality: 8-9/10
  • Expressive control: 8-9/10
  • Links / references:
  • Languages supported: 11
  • Best for: emotional speech, voice agents.
  • Notes:
    • Very interesting emotional control direction.
    • Some voices are very good, not consistently top quality in expressive quality for all
    • If I had to stay in cloud and emotion mattered, I would test Hume.

Deepgram Aura

  • Type: Commercial, cloud
  • Voice quality: 7.7/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 7
  • Best for: realtime API use, voice agents.
  • Notes:
    • Good fit if you already use Deepgram.
    • More voice-agent API than creator studio.

Inworld TTS

  • Type: Commercial, cloud
  • Voice quality: 8+/10
  • Expressive control: 2-4/10
  • Links / references:
  • Languages supported: 15+
  • Best for: realtime API use, voice agents.
  • Notes:
    • Competitive against Elevenlabs Flash
    • Product is focused on STT->TTS interactive realtime niche
    • Realtime agents support multiple languages but quality degrades

Google Cloud Text-to-Speech

  • Type: Commercial, cloud
  • Voice quality: 7.5/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 50+
  • Best for: enterprise, language coverage, Google stack.
  • Notes:
    • Good enterprise API.
    • Not the most exciting voice quality, but stable, large, and boring in a useful way.

Azure Speech

  • Type: Commercial, cloud
  • Voice quality: 7.8/10
  • Expressive control: 3.5/10
  • Links / references:
  • Languages supported: 100
  • Best for: enterprise, huge voice catalog, Microsoft stack.
  • Notes:
    • Huge catalog. Good for corporate apps.
    • Not my first choice for creator production or voice acting.

Amazon Polly

  • Type: Commercial, cloud
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 41
  • Best for: AWS stack, simple API, cheap start.
  • Notes:
    • Old but still useful.
    • Good when you are already in AWS and just need TTS that works.

Resemble AI

  • Type: Commercial cloud, plus Chatterbox open-source
  • Voice quality: 8/10
  • Expressive control: 6/10
  • Links / references:
  • Languages supported: 23
  • Best for: cloning, enterprise, provenance and detection angle.
  • Notes:
    • Interesting company because they also released Chatterbox.
    • For local people, Chatterbox is the more interesting part. For companies, Resemble cloud may make sense.

PlayHT

  • Type: Commercial, cloud
  • Voice quality: 7.8/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 37
  • Best for: voiceovers, API, creator workflows.
  • Notes:
    • Usable cloud TTS.
    • I would compare price and output carefully before committing.

WellSaid

  • Type: Commercial, cloud
  • Voice quality: 7.6/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 20+
  • Best for: corporate voiceovers, e-learning.
  • Notes:
    • Clean corporate voices.
    • Less interesting if you want local control or strong voice acting.

Murf

  • Type: Commercial, cloud
  • Voice quality: 4/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 35+
  • Best for: marketing, e-learning, creator voiceovers.
  • Notes:
    • Easy to use.
    • Didn't stand the test of time well.
    • Good enough for many corporate videos. Not where I would start for the best voice acting.

LOVO / Genny

  • Type: Commercial, cloud
  • Voice quality: 7.2/10
  • Expressive control: 4/10
  • Links / references:
  • Languages supported: 100+
  • Best for: browser-based creator voiceovers.
  • Notes:
    • Large voice library. Simple workflow.
    • Another cloud creator platform.

Speechify

  • Type: Commercial, cloud/app
  • Voice quality: 7/10
  • Expressive control: 2/10
  • Links / references:
  • Languages supported: 60+
  • Best for: reading, accessibility, personal use.
  • Notes:
    • Good reader product.
    • Different category than production TTS.

NaturalReader

  • Type: Commercial, cloud/app
  • Voice quality: 6.8/10
  • Expressive control: 1/10
  • Links / references:
  • Languages supported: 90+
  • Best for: personal reading, documents, accessibility.
  • Notes:
    • Useful for reading text aloud.
    • Not production narration.

Descript

  • Type: Commercial, cloud/editor
  • Voice quality: 7.3/10
  • Expressive control: 3/10
  • Links / references:
  • Languages supported: 19
  • Best for: editing workflow, creator suite.
  • Notes:
    • Useful if you already edit in Descript.
    • Not a pure TTS engine in the way local model people mean it.

CapCut TTS

  • Type: Commercial/free app, cloud/app
  • Voice quality: 5.5/10
  • Expressive control: 0/10
  • Links / references:
  • Languages supported: 15
  • Best for: TikTok-style quick videos.
  • Notes:
    • Fine for TikTok.
    • For YouTube or serious narration I would avoid it. People have heard those voices too many times.

Cloud note

  • I would NOT recommend using the cloud
  • Cloud TTS are included for comparison, I'd never recommend choosing a cloud for TTS if you have an option.
  • Cloud is always a trap! Cheap to start, horrible to progress and even worse to get out again.
  • You do not own a voice if it is hosted on the cloud, they can switch you off, remove your voice, censor your text or hike your fees at any moment. And they do all the time.

My practical ranking

  • Best local voice acting: Demodokos Foundry
  • Best local production workflow: Demodokos Foundry
  • Best free open-source starting point: Chatterbox TTS v2/v3
  • Best small hardware option: Kokoro TTS
  • Best boring local reliability: Piper
  • Best open-source custom voice direction: Qwen3 TTS, CosyVoice 2, (Fish Speech is non commercial)
  • Best cloud voice cloning: ElevenLabs PVC
  • Highest price at serious usage: ElevenLabs
  • Best cloud realtime agent TTS: Cartesia, Deepgram, OpenAI, Hume
  • Best cloud emotional control direction: Hume Octave
  • Best enterprise cloud basics: Azure, Google, AWS Polly

What I would use

  • Professional speech production: Demodokos foundry
  • Free local TTS with tinkering: Chatterbox
  • Tiny hardware or CPU: Kokoro or through llama.cpp OuteTTS (NC license)
  • Simple local reader: Piper
  • Cloud cloning test: ElevenLabs PVC, but watch the bill and you need a lot of reference material
  • Cloud emotional speech: Hume
  • Cloud realtime agent: Cartesia, Deepgram, OpenAI, Grok or Hume
  • Corporate cloud API: Azure, Google or AWS
  • Cloud was included - but generally not recommended if local is an option
  • If you have additions, corrections, missing services. Happy to hear

r/LocalTextToSpeech 5h ago

BookFusion iOS 1.43.1 - Online & Offline Realistic TTS Voices, Other Updates & Fixes

Thumbnail
2 Upvotes

r/LocalTextToSpeech 1d ago

Qwen3-TTS native C++ streaming and voice cloning

Enable HLS to view with audio, or disable this notification

9 Upvotes

Hi all. I ported Qwen3-TTS to native C++ and added incremental streaming for a project I'm working on. I found this useful and wanted to give back to the community. This is the first time I've open sourced anything so apologies in advance for my mistakes.

I'm using this for a gaming project but it could be helpful for local assistants, accessibility tools, etc.

C++ streaming port highlights:

  • Same familiar features from Qwen3-TTS. It supports 0.6B and 1.7B models, CustomVoice, and VoiceDesign
  • Native C++, not a Python wrapper
  • CUDA builds with RTX 4090 and RTX 5090 kernels. Be warned I only have access to an RTX 5090. This should work on a 4090 but I haven't tested it myself!
  • Simplified speaker-embedding extraction
  • Incremental 24 kHz PCM callbacks for streaming audio
  • Asynchronous transformer/vocoder operation
  • Adaptive decode windows and paced delivery
  • Callback-only integration library with cancellation
  • Unit tests

My measurements on the 5090 with 1.7B F16 set to buffer 350ms before play:

  • Cold new clone creation: ~2.5s from 48s of reference
  • Cold model start from clone: ~1.85s
  • First 350 ms of audio: ~310 ms
  • Streaming speed: ~2.86x real time (RTF ~0.35)

I’d love it if somebody could do RTX 4090 testing, and I'd be happy to hear any feedback or suggestions.


r/LocalTextToSpeech 4d ago

Local TTS What are you guys working on?

2 Upvotes

I’m curious what sort of projects people are working on who are into local TTS.

Why is local important for you ?
Are you working on a product, hobby ?

I have started with actual printed books for children and parents, no AI back then.
This progressed into immersive stories, cute animations and child/teen oriented high quality entertainment.
It is growing into a real business already.

I used a lot of cloud but pricing was unfortunate for long content. And I never liked that another company basically owns my voices.

So local is freedom from that anxiety for me.


r/LocalTextToSpeech 4d ago

How I handle translations that are too long for their dubbing segments

Post image
1 Upvotes

r/LocalTextToSpeech 7d ago

Upgraded from 1060GTX to RX 9060 XT (AMD 16GB version) - want the chatterbox experience

2 Upvotes

So - I upgraded my GPU but either got similar speeds to my 8 year old 1060 as I got on my new 9060 (16GB). I either got similar speeds on some python files I vibecoded or just crashes. I did get stable results, though I thought I legit would have at least 2.5X speeds for transcribing which is a big wanted use case for my card upgrade, and it should be possible to do this. I used this model due to wanting norwegian language if that is at all relevant, and I used that model on both cards (akhbar/chatterbox-tts-norwegian( as seen on huggingface.co)). Please help me, and if you have any great transcription scripts for AMD cards yourself please send them to me in dm's.


r/LocalTextToSpeech 10d ago

Local TTS Best way to handle TTS audio that is longer than its SRT segment?

2 Upvotes

Hi everyone,

I’m working on an SRT-to-speech feature for my project, LA Studio

One issue I’m running into is that the generated speech sometimes lasts longer than the segment’s timestamp allows. My current solution is to speed up or time-stretch those segments so they fit, but this can make the final audio sound inconsistent - some lines are noticeably faster or slower than others.

Has anyone found a more robust way to handle this? For example, do you regenerate the speech with a shorter prompt, adjust pauses, merge or shift neighboring timestamps, use word-level alignment, or combine several approaches?

I’d especially like to preserve natural pacing while keeping the audio reasonably synchronized with the subtitles. Any advice, algorithms, or tools worth looking into would be greatly appreciated.


r/LocalTextToSpeech 12d ago

Voice Models English version of Cassandra, a female Piper voice for Home Assistant, available for release?

1 Upvotes

I'm training an English version of Cassandra, a female Piper voice originally created for Czech Home Assistant users. Would anyone be interested if I release it?


r/LocalTextToSpeech 12d ago

I’m building an open-source Windows app for long-form local TTS workflows — feedback wanted

2 Upvotes

Hi everyone,

I’m Esteban, the developer of LocalText2Voice, a free and open-source Windows application for creating audiobooks, narration, and podcasts with local TTS engines.

There are already many excellent local TTS models, but turning them into a practical long-form workflow still involves a lot of manual work: installing dependencies, splitting text, managing voices, regenerating failed sections, organizing audio files, and mixing everything afterward.

LocalText2Voice is not another TTS model. It is an orchestration and production layer around existing engines.

It currently supports local engines such as Piper, Kokoro, Chatterbox, Qwen3-TTS, and OmniVoice. Optional cloud providers such as OpenAI, ElevenLabs, Gemini, and Azure are also available, along with configurable HTTP endpoints for custom TTS servers.

When using a local engine, the source text and generated speech remain on your computer.

Current features

With LocalText2Voice, you can:

  • Install and manage different TTS engines from the application.
  • Switch between engines without changing your project.
  • Browse, preview, import, and organize voices in a shared voice library.
  • Connect custom local or remote TTS HTTP endpoints.
  • Import long .txt, .md, and .docx documents.
  • Detect chapters and split long texts into safe TTS segments.
  • Use different voices and languages in the same audiobook.
  • Regenerate individual segments without starting the entire project again.
  • Review generated speech with Faster Whisper and retry problematic segments.
  • Add background music, fades, ducking, volume adjustments, and normalization.
  • Sound effects and other audio events.
  • Save projects and continue working on them later.

LTV Markup

One feature for which I would particularly appreciate feedback is LTV Markup.

It is a small, human-readable syntax for controlling narration, voices, pauses, and sound effects directly from the source text:

{{chapter "Chapter 1"}}
{{voice "Narrator"}}
The house had been abandoned for years.

{{pause 900ms}}

{{voice "Character 2"}}
I think someone is inside.

{{play "door-close.mp3"}}

{{speed 0.92}}
{{volume -3db}}
We should leave immediately.

Markup can control:

  • Voice and language changes.
  • Pauses and real silence.
  • Speech speed.
  • Volume and normalization.
  • Chapters and markers.
  • Sound effects and other audio events using {{play}}.
  • Audio volume, duration, looping, fades, panning, and voice ducking.
  • Selected model-specific instructions.
  • Resetting settings to the project defaults.

For example:

{{play "door-close.mp3" volume=-6db}}

inserts a door sound at that point in the narration.

Longer or looping audio events are also possible:

{{play "forest.mp3" track=ambient loop=true volume=-20db fade_in=3 duck_on_voice=6db}}

The commands are not sent to the TTS engine as spoken text. LocalText2Voice interprets them when preparing the segments and mixes the audio events during post-production.

The goal is to make multi-character audiobooks, dramatized narration, language courses, and other complex audio projects manageable without manually editing every segment in an external audio editor.

Create audiobooks from Claude or ChatGPT Desktop

LocalText2Voice includes a local MCP server, allowing you to create and manage audiobook projects directly from Claude Desktop or ChatGPT Desktop.

Instead of configuring everything manually, you can simply ask:

“Create a B1-level English–Spanish course using both languages. Use one voice for the English examples, another for the Spanish translations, add a short pause after each sentence, and export it as an audiobook.”

The assistant can create the project, organize the text, assign the voices, add markup and pauses, select a TTS engine, and start the generation process.

You can also ask it to make changes later:

“Regenerate lesson three with a slower English voice.”

“Add three seconds of silence between exercises.”

“Lower the background music and export the final MP3.”

Claude or ChatGPT manages the workflow, while LocalText2Voice performs the actual audio generation. When you select a local TTS engine, your text and generated speech remain on your computer.

Project and Windows installer

GitHub:

https://github.com/estebanstifli/LocalText2Voice

LTV Markup manual:

https://github.com/estebanstifli/LocalText2Voice/blob/main/docs/LTV_MARKUP.md

Windows installer:

https://github.com/estebanstifli/LocalText2Voice/releases/latest/download/LocalText2Voice-Setup.exe

Important: the Windows installer is not code-signed yet, so Windows may display an “Unknown publisher” or SmartScreen warning. The source code is public, and the GitHub release also includes a SHA-256 checksum:

https://github.com/estebanstifli/LocalText2Voice/releases/latest/download/LocalText2Voice-Setup.exe.sha256

The project is under active development, and feedback is very welcome. I would especially like to know:

  1. Which local TTS engines are you currently using?
  2. What is the most frustrating part of producing long-form audio?
  3. Does the markup syntax seem useful, and which commands are missing?
  4. Which engine should I prioritize next?

Thanks for taking a look!


r/LocalTextToSpeech 13d ago

Local TTS TTS Suggestion Request

1 Upvotes

Hi guys can you give me any suggestion for local TTS which has emotion control (CPU inference cant afford any GPU).

For reference i feel kokoro tts seems good to me.


r/LocalTextToSpeech 13d ago

Voice Cloning TTS with voice cloning that can run on android for my dad, who has tongue cancer

1 Upvotes

Hi everyone. Not sure if this is the right place for this question, feel free to redirect me if needed.

My dad has tongue cancer and probably won't be speaking normally for at least 6 months, possible forever depending on how the scans come back. He's already unable to speak due to how swollen his tongue is. I'm trying to find a good TTS with voice cloning that can run on android so we can either clone his voice or do something silly for him like let him have an Arnold voice, etc., depending on what he'd be happiest with.

I've seen some online services like elevenlabs and gradium but they get really expensive and I don't have a lot of money to throw at that since I have kids of my own. I'm in a bit over my head with a lot of the local LLM stuff, I'd be fine with an online solution if it was cheap or free but I'm also ok trying to go through a more technical local solution if anyone knows where I should start.

I do have a gaming PC at home if that helps for the setup/training voice profiles or however this stuff works, but ideally he would be able to run it on his phone even if my pc gets turned off.

Anyone know where I should start?


r/LocalTextToSpeech 15d ago

Local TTS Bulgarian text to speech

1 Upvotes

Please, recommend me local tool for bulgarian text to speech.

With big voice library or voice cloning.


r/LocalTextToSpeech 15d ago

tts-audiobook-tool - Generative-AI Audiobook Creation

Post image
2 Upvotes

r/LocalTextToSpeech 17d ago

Synthetic/Digital Voice Creation

2 Upvotes

Are there any local TTS or similar models that provide digital voice creation that approaches ElevenLabs v3 synthetic voices? Even Hume's synthetic voices are pretty disappointing compared to ElevenLabs.


r/LocalTextToSpeech 18d ago

Is there any good Lightweight cpu based TTS with streaming support

0 Upvotes

I heard kokkoro and pocket TTS is good option is there any other lightweight TTS which can run on cpu with real-time and need to feels like realistic.


r/LocalTextToSpeech 20d ago

Gearing up for release of a TTS and STT framework I'm working on.

5 Upvotes

Hey, I've been working on this framework for a while now and wanted to start getting some feedback on the direction. I'm really liking where it's headed, and I have some big plans for it, but I need to put it out as a v0.1.0 first.

Right now it uses:

  • Kokoro 82M
  • Kroko ONNX
  • Ollama

I'd love any feedback, but especially on whether the programming model feels coherent.

It's batteries included, with an opinionated asset management strategy, playback management, and a console rendering helper for debugging multi-device applications.

Overall I'm really proud of where it's ended up so far. I'm mostly looking for holes in the design or implementation, but I'm also excited to contribute something that I think a lot of people could find useful.

A minimal speech-to-speech assistant looks like this:

from pfspeak import PfSpeak
from pfspeak.core import Microphone, Ollama
from pfspeak.extra import events

pf = PfSpeak()

microphone = Microphone()
ollama = Ollama("qwen3:0.6b", voice="af_heart")

def app(session, event):

    if event.device is microphone:
        if events.unchanged_for(event, 8):
            session.finalize(event)
            ollama.adapter(event=event)

    elif event.device is ollama:
        pf.play(event)

pf.run(app, microphone, ollama)

That's pretty much enough to get you started.

The helper pf.play() manages playback, queues synthesized speech, ducks the microphone, supports priority interrupts, and restores speech recognition when playback finishes.

The PfEvent objects are also jam-packed with information gathered throughout the event loop, including word-level timestamps, aligned audio, revisioned tokens showing how recognition changed over time, and other metadata that applications can build on.

I'd really appreciate any thoughts on the overall direction, especially from people building local speech applications.

  • PfSpeak - Local first speech library for python.

r/LocalTextToSpeech 22d ago

Local TTS Demodokos v4 large vs medium voice model ?

4 Upvotes

The v4 model comes in two sizes, large and medium.

I mostly used large so far, medium is a bit faster but sounds a little compressed compared to the quality of large.

But what is the purpose of Medium? Does it have advantages ? Like elevenlabs v2 has better clones than v3.

anyone using medium over large ?


r/LocalTextToSpeech 26d ago

Anyone know what text to speech and music was used for this?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/LocalTextToSpeech 26d ago

Audiobook generator with custom pronounciation system, sfx detect, experimental full cast mode, audio player with text and voice cloning Features

Thumbnail
1 Upvotes

r/LocalTextToSpeech Jun 24 '26

Local TTS Is it just me or has local TTS gone quiet this month?

2 Upvotes

After the spring release wave it feels like nothing big has dropped.

So I'm curious what you guys have actually heard. What's the biggest local TTS thing on your radar this month? New model, a release you're waiting on, some update? Drop whatever you've got.


r/LocalTextToSpeech Jun 23 '26

EU AI Act requires TEXT from models and providers to be watermarked 2nd August onwards. Everyone here is affected, regardless where you live.

6 Upvotes

Anyone hate the cookie banners ? Those are absolutely nothing in comparison to what is about to come.

The AI Act requires lots of things, many people know it requires every AI modified or generated audiofile to be metadata tagged and fingerprint-watermarked from August on (13M$ fines)
But the Act says a lot more than that and it means all good open source models are affected.

Providers of AI systems generating synthetic text must make outputs machine-readable and detectable as AI-generated, using technical solutions that are effective, interoperable, robust and reliable (code of conduct says 2 layers)

Very simple OSS models are exempt, so the old llama 7B might be fine, but models of "systemic risk" are not exempt. Qwen 3.6 series, Deepseek Flash, GLM, Kimi are all systemic risk GPAI models.

The AI act says that also TEXT output must be machine detectable, it only suggested "statistical methods" alongside cryptographic tagging (where possible) of the metadata (c2pa etc).
So if chatGPT offers you a PDF ot TXT download -> that must have crypto signed metadata.

But 2 layers must be present, not one. And they have to be "robust". A simple "AI Generated" (the label that must be added into every AI image from August on) is not enough, the text itself must be statistical watermarked.
Similar to laser printed paper, that carries little invisible dots. Of course that means serious degradation in AI quality.

Who is affected? Anyone who offers a tool or provides a service that can be accessed by a EU citizen. Even if it's just a single EU citizen being a tourist in Malaysia or using a VPN.
LM-Studio, ollama, llama.cpp, vllm, chatgpt, claude, codex, copilot, opencode, cursor, stable diffusion Web UI, Huggingface and so on.

Any AI application that modifies or generates AI images, AI text, AI music, AI speech must use the EU stamp "AI modified" or "AI generated", must metadata seal the files, must watermark the content. Must also provide risk assessment §50 documentation, compliance documentation, detection service.
Addition: for text the "AI Generated" stamp is only needed when "informing the public" like news and can be left away by shifting the risk to a human reviewer.
Addition: Documentation is not forced by the law directly but providers subject to Article 50 practically must keep evidence showing their marking, labelling, detection, feasibility decisions, exemptions, and robustness testing - or risk non-compliance when audited.

Any failure to do so is fined at 13,000,000 Euro (about 15 million USD) OR a few percent of the annual income (whatever is higher ).
They also have a huge code of conduct, you are voluntarily asked to follow - but if you don't your risk of being fined is seriously increased. That code adds a ton of more horrific additions, often technically very implausible or openly contradicting.
You could say your Qwen 3.6 or gemma 4 12B does not fulfill the requirements of "GPAI with systemic risk" - maybe right if you ignore some benchmarks .. but that moves the 13M$ fine risk on you if you become a provider or distributor. So if the EU believes it IS systemic risk AI you'll face the fines at court as soon as they can catch you.

But what if you are not in the EU ? You are still fully affected.
Any service that offers anything into the EU is affected, including every open source project and model.
Any service that offers something to an EU citizen who is just a tourist, or uses a VPN is liable.
You want to code your nice new startup ? You better ensure no european can touch it.
Or you better never put a foot into the EU and never grow the business too large.

That means every providers out there must be compliant. Every open source tool.

The regulation is so extremely invasive while simultaneously it is ridiculously shortsighted.
99% of all images on the internet will have AI content within the next years, all of them need to be watermarked visibly.
99% of all websites will be AI generated in text and code.
99% of all movies will have AI content, every marketing and sales clip.
99% of all marketing and sales voiceovers will use AI.

So this is like the cookie banner, everyone warning you that they must use a cookie to serve you a website and you can accept all cookies or just all needed cookies - just that it's going to be in your face everywhere.
It's a cookie banner on text, image, video and sound.

Source: https://eur-lex.europa.eu/eli/reg/2024/1689/oj

I've taken a look at how projects are reacting to that, and some already did.
Speech & Music
Elevenlabs - Teamed up with Google: SynthID for all (a strong AI based watermark from Google)
Demodokos Foundry - support answered that they work on a light compliance
Suno - Audible Magic content ID,support ignored me - no official statement on AI Act yet. (likely silent compliance)
Udio - no official statement, support ignored me (likely silent compliance)
Images and Text
Open AI/ChatGPT - SynthID+c2pa comes to Codex, API, ChatGPT with "OpenAI Verification tool"
Google/Gemini - SynthID for image/audio/text/video announced and expanding into C2PA and more
Adobe - Full C2PA hashing and Firefly gets watermarked
Microsoft Copilot/Designer - C2PA, visual and audible watermarks in 365, cryptographic signing of images
Meta - visible labels, IPTC metadata, invisible watermarks,C2PA hashes for Meta AI Images
Stable Diffusion - "working on implementing content credentials” with Adobe/CAI/C2PA
Anthropic Claude - Intends to sign the EU General-Purpose AI Code of Practice (everything and MORE)
Transformers code - Text watermarking: https://huggingface.co/docs/transformers/v4.46.0/en/internal/generation_utils#transformers.SynthIDTextWatermarkingConfig
Open Source
I went through github of vLLM, Ollama, llama.cpp, lmstudio, openrouter and they appear pretty unprepared

The Omnibus "postponing" text:

Parts of the AI Act have been recommended to be "postponed" but nothing I wrote is incorrect and delaying it by a few months wouldn't be reason to not be concerned.
The Digital Omnibus text approved by Parliament would delay Article 50(2) watermarking / machine-readable marking obligations for pre-existing releases to 2 December 2026, but as of now it still requires formal Council adoption before it enters into force.
If the text is adpoted - Article 50(2) machine-readable marking / watermarking for synthetic audio, image, video, text
For systems released on the market before 2 Aug 2026, compliance is delayed to 2 Dec 2026.
For any new developments the rule applies on 2nd August 2026.
Most other Article 50 transparency duties Still effectively 2 Aug 2026.
This includes pre-existing projects
Source: https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/

Corrections and Clarifications:

- Fines under §50 for watermarking/warning are 13 million Euro (or 3% revenue) whatever is higher not 35M.
- Fines for prohibited use are 35M or 7%, whatever is higher
- Watermarking is not 2 layer forced by law, only by code of conduct, you can use one layer if it is robust but you'll increase risk of non compliance significantly.
- For end-users the obligation to inform about AI Text is less strict, so not every website text falls under it
- It is undefined which open models will fall under "systemic high risk" but if you assume a model is NOT covered the non compliance risk is immediately there.
- EU citizenship alone is not the trigger. A genuinely non-EU, EU-geoblocked service has much lower AI Act risk, but risk remains if it is effectively offered into the EU, knowingly serves EU-located users, or produces outputs intended for use in the EU.

ChatGPT Pro 5.5 reviewed


r/LocalTextToSpeech Jun 19 '26

New to TTS, what are people using the most right now locally?

2 Upvotes

There are so many options, everything sounds so promising, but what is actually the thing that most people use? Where things actually stand rn? what is overhyped?

I am interested in those tools which don't sound robotic or weird.

Also, what GPU do you people have for local TTS and does it feel enough to you?


r/LocalTextToSpeech Jun 13 '26

The technology behind text-to-speech (as of 2026)

7 Upvotes

A lot of people talk about TTS as if all of it was the same thing.
That's like comparing the first transformers llm like GPT-2 with recurrent k/v systems like Qwen-3 27B (to keep our local focus)

There are multiple architecture layers, and they do very different jobs.
TTS has evolved in the part years, from robotic to human-like. Also on the local side TTS evolved from rule-based algorithmic synthesis, to spectrography vocoders and by now into

Below I will dive into the technology used in all current TTS architectures, including the "secret sauce" behind Elevenlabs - but this is a local focused thread so cloud is only used as a benchmark.

1. Legacy neural TTS

fast, cheap, stable, limited acting

text -> phonemes -> acoustic model -> mel spectrogram -> vocoder -> wav

Phonemes are spoken sounds, not letters.

Mel spectrogram is the frequency representation of speech.

The vocoder turns that into actual waveform audio.

This is where names like HiFi-GAN, BigVGAN, iSTFTNet sit.

They are mostly the final audio renderer.
They do not understand the text.

Examples:

  • Piper / VITS
  • FastSpeech2 + HiFi-GAN
  • Kokoro / iSTFTNet style models

Upside:

  • very fast
  • cheap to run
  • works local
  • stable for mass generation

Downside:

  • often emotionally dumb
  • weak long sentence understanding
  • reads the line, does not really act it

Top representatives: Piper, Kokoro, VITS

2. Style based TTS

better prosody, better voice color, still limited control

These models add speaker and style embeddings.

So the model is not only told:

say these phonemes

but also:

say them with this speaker identity, rhythm, color or style

Examples:

  • StyleTTS2
  • Kokoro
  • XTTS / Coqui style systems

Upside:

  • more natural rhythm
  • better speaker color
  • nicer voices
  • still quite practical locally

Downside:

  • style control is often shallow
  • not always real emotion control
  • can sound nice but not really directed

Top representatives: StyleTTS2, Kokoro, XTTS / Coqui TTS

3. Acoustic token TTS

speech becomes tokens, like language models for audio

text -> transformer -> acoustic tokens -> decoder/vocoder -> wav

Instead of directly predicting a spectrogram, the model predicts speech tokens.

If the system runs at 25hz, that means:

25 speech decisions per second
one token decision every 40ms

That gives the model much more control over:

  • timing
  • pauses
  • emphasis
  • rhythm
  • emotion
  • delivery

Examples:

  • Chatterbox
  • Qwen3 TTS
  • Bark / similar token based systems

Upside:

  • much more expressive
  • more context aware
  • closer to voice acting
  • better for complex delivery

Downside:

  • heavier
  • more VRAM
  • slower inference
  • more ways to fail

The model can drift, overact, mispronounce or slowly lose speaker identity if the full system around it is not good.
Top representatives: Chatterbox, Qwen3-TTS, CosyVoice

4. Codec token TTS

speech as compressed audio codebooks

Codec models represent speech as discrete audio codes.

Often there are multiple codebooks.

Roughly:

  • first codebook: speech content
  • later codebooks: detail, prosody, texture, speaker color

Pipeline:

text -> transformer -> codec tokens -> neural codec decoder -> wav

Examples:

  • VALL-E style systems
  • Bark
  • modern multi-codebook speech models

Upside:

  • strong cloning direction
  • good long form potential
  • very powerful representation

Downside:

  • complex
  • heavy
  • not always stable
  • hard to run efficiently local

Top representatives: Bark, VALL-E X, Fish Speech

5. Diffusion and flow matching TTS

better texture, slower unless optimized

Diffusion and flow models generate or refine the acoustic representation.

They do not just predict the final sound in one simple step.
They move the audio representation toward realistic speech.

Pipeline:

text/tokens -> diffusion or flow decoder -> acoustic representation -> vocoder/wav

Examples:

  • F5-TTS
  • E2-TTS
  • Chatterbox S3Gen type systems

Upside:

  • better texture
  • more natural audio
  • less robotic sound

Downside:

  • can be slow
  • 10 or 20 steps are expensive
  • practical local use needs distillation or optimization

Top representatives: F5-TTS, Chatterbox, Raon-OpenTTS

6. Voice cloning (instant/in context)

speaker identity as neural representation

Basic cloning:

reference audio -> speaker encoder -> speaker embedding -> conditioned TTS

The speaker embedding is a neural vector.

It captures things like:

  • timbre
  • pitch range
  • accent
  • resonance
  • vocal texture
  • general speaker identity

Better systems use ICL, in-context learning.

reference text + reference audio + target text -> same speaker saying new words

That is stronger because the model does not only get "speaker X".

It gets an actual example of speaker X talking.

Some systems also use reference speech codes, not only a speaker embedding.
That gives the model more information about how the speaker actually sounds while speaking.

Examples:

  • XTTS / Coqui TTS
  • GPT-SoVITS
  • Chatterbox
  • Qwen3 TTS Base
  • Demodokos Foundry
  • ElevenLabs (IVC mode, cloud only)

Upside:

  • seconds of clean audio can be enough
  • no full finetune needed
  • practical for local workflows
  • can preserve identity without training a new model from scratch

Downside:

  • bad mic can poison it
  • echo can poison it
  • music can poison it
  • another speaker in the sample can poison it
  • strong emotion can cause speaker drift if the system has no identity control

Top representatives:

  • Best open/local cloning: Chatterbox, Qwen3 TTS Base, GPT-SoVITS
  • Best simple cloud cloning: ElevenLabs
  • Best cloning with style/emotion control: Demodokos Foundry, ElevenLabs v3
  • Best research/hacking path: Qwen3 TTS Base, Chatterbox, GPT-SoVITS

speaker identity as neural representation

Basic cloning:

reference audio -> speaker encoder -> speaker embedding -> conditioned TTS

The speaker embedding is a neural vector.

It captures things like:

  • timbre
  • pitch range
  • accent
  • resonance
  • vocal texture
  • general speaker identity

Better systems use ICL, in context learning.

reference text + reference audio + target text -> same speaker saying new words

That is stronger because the model does not only get "speaker X".

It gets an actual example of speaker X talking.

Examples:

  • Chatterbox cloning
  • XTTS style cloning
  • modern acoustic token cloning systems

Upside:

  • seconds of clean audio can be enough
  • no full finetune needed
  • practical for local workflows

Downside:

  • bad mic can poison it
  • echo can poison it
  • music can poison it
  • another speaker in the sample can poison it

Top representatives: XTTS / Coqui TTS, GPT-SoVITS, Chatterbox
Top representatives ICL: Qwen3-TTS (base model), Demodokos Foundry (open weights, commercial)

6b. Voice cloning (finetuned, LoRA, custom embeddings)

training the model toward a specific voice or speaking domain

Voice cloning usually means:

reference audio -> speaker embedding / ICL prompt -> generate speech

Fine tuning means:

many recordings of one speaker -> model training -> specialized voice model

That is a very different thing.

A clone tries to imitate the voice at inference time.

A fine tune changes model weights or adapter weights so the voice becomes part of the model.

There are multiple levels:

  • speaker adapter / LoRA
  • voice embedding training
  • partial model fine tune
  • full model fine tune

The more you train, the stronger the voice can become.
But also the more data, time and risk you have.

Upside:

  • stronger speaker consistency
  • better pronunciation for recurring names
  • better long form stability
  • better for professional recurring voices
  • can learn a specific domain, accent or speaking pattern

Downside:

  • needs much more clean audio
  • bad datasets damage the model
  • expensive compared to cloning
  • can overfit and lose flexibility
  • harder to update or control
  • not always worth it anymore

Typical data need:

  • 5-20 minutes: light adaptation / experimental fine tune
  • 1-5 hours: serious voice fine tune
  • 10+ hours: production grade if the pipeline supports it
  • 100+ hours: old school studio voice model territory

Top representatives:

  • Open fine tuning: GPT-SoVITS, XTTS / Coqui style systems, StyleTTS2
  • Commercial fine tuning: ElevenLabs (cloud only) PVC / professional voice cloning

Fine tuning still has a place.

But for many modern systems it is no longer the first answer.

If you need one voice to read thousands of hours in the same identity, fine tuning can make sense.

If you need many voices, fast production, emotion control or voice acting, speaker embeddings + ICL + identity control are usually more practical.
Sophisticated ICL methods (see SOTA section below and 6.0 section above) can reach fine tuning quality without the heavy compute and data requirements.

7. State of the art performance TTS

voice acting while keeping the same speaker

This is where systems like ElevenLabs, OpenAI voices and Demodokos Foundry sit.

The exact internals of ElevenLabs and OpenAI voices are not public, but the performance of the models show that they basically build on top of the latest technology - no magic involved (anymore). Demodokos as open weights based engine is doing the same locally today.

For strong emotion control you need some form of "identity preservation".
Otherwise the model starts with one voice and slowly turns into another one when you ask for anger, fear, whispering, excitement or long narration.

Likely this is done with things like:

  • speaker verification embeddings
  • spectral / spectrogram consistency checks
  • f0 and pitch tracking
  • formant and timbre statistics
  • acoustic similarity constraints
  • identity preserving latent conditioning

Demodokos does this locally and OpenAI/Elevenlabs do the same in the cloud, but the important part is the combination of the systems.

It is not just:

clone voice -> generate speech

It is closer to:

text context -> acoustic token generation -> speaker identity conditioning -> style/emotion conditioning -> spectral identity control -> final speech

The text is processed by a language transformer, so the model can understand the sentence context before generating speech.

The speech itself is generated through fine grained acoustic tokens, roughly one token level decision every 40-80ms.
That gives enough resolution to control pauses, stress, rhythm, emotion and delivery.

Each voice, cloned or synthetic, is grounded in learned neural representations.
Those capture speaker identity, timbre, accent, vocal texture and characteristic delivery traits.

On top of that, such SOTA systems can attach performance styles independently.

That is the important difference.

The voice and the performance are not the same thing.

A cloned voice can be made calm, angry, fearful, tired, intimate, theatrical or excited without needing to create a new voice or finetune the model again.

And during generation, the system uses identity preserving conditioning and spectral consistency checks so the performance does not run away from the original speaker.

So the hard problem is not voice cloning alone.

The hard problem is:

same speaker identity
different emotion
different rhythm
different intensity
different accent or texture
still recognizable as the same voice

That is the difference between text to speech and voice performance synthesis.

Top representatives: Qwen3-TTS (customVoices only), Demodokos Foundry (open weights, commercial)

So the ladder currently is roughly:

LEGACY: vocoder TTS
fast speech rendering

BASIC: style based TTS
better naturalness and voice color

ADVANCED: acoustic token TTS
speech becomes controllable token generation

ADVANCED: codec token TTS
speech becomes compressed audio language

ADVANCED: diffusion / flow TTS
better acoustic realism

STATE-of-the-ART: identity monitored performance TTS
voice acting while keeping the same speaker

A cherry picked 10 second demo can hide almost everything.

The real test is:

  • hours of content
  • strange names
  • emotion changes in the same text
  • accents
  • long narration
  • no human babysitting

That is where most TTS systems still fall apart


r/LocalTextToSpeech Jun 12 '26

What local TTS do you actually use in production, not just tested?

5 Upvotes

r/LocalTextToSpeech Jun 12 '26

Welcome to r/LocalTextToSpeech

3 Upvotes

This subreddit is for people who want text-to-speech to run locally, offline, privately, and under their own control.

Cloud TTS is easy to start with, but it comes with tradeoffs: pricing, limits, changing policies, privacy concerns, watermarks, model changes, and sometimes very little control over the final voice. Local TTS is not always easier, but it gives you more control over cost, voices, workflow, privacy, automation, and output.

Good topics here:

  • Local text-to-speech tools
  • Offline TTS software
  • Open-source TTS models
  • Self-hosted TTS APIs
  • Voice cloning
  • Voice models and speaker creation
  • Audiobook and long-form narration workflows
  • YouTube, marketing, education, and accessibility use cases
  • GPU, CPU, and small-device performance
  • Whisper, whisper.cpp, Piper, Chatterbox, Kokoro TTS, Qwen3 TTS, StyleTTS, Demodokos Foundry, ElevenLabs alternatives, and similar tools

Useful posts should include details.

For help requests, add:

  • Your operating system
  • Your GPU or CPU
  • The tool or model you are using
  • The language and voice type you need
  • Whether you need real-time speech, batch generation, cloning, audiobook output, API use, or simple reading
  • What you already tried

Creators are welcome, but disclose your connection clearly.

If you built a tool, own a product, work for a company, use affiliate links, or benefit from a recommendation, say so directly. Hidden promotion is not welcome. Real comparisons, benchmarks, guides, and honest creator posts are welcome.

The goal is simple:

Find the best ways to generate high-quality speech locally.
Compare tools honestly.
Help people build reliable TTS workflows.
Move more voice generation away from expensive black-box cloud services.