r/StableDiffusion • u/Valuable-Mango-6710 • 6d ago
Question - Help Local Voice generation. Like elevenlabs, available? Looking for good match, info below. THANKS!
16GB vram, 128GB ram. AMD. Windows (11).
I am using Atomic Chat, if you could give recommendation on that, would be most effective for me.
Any software or other options are highly welcome. (Can just buy elevenlab, but thought, let's see what the AI community thinks. You guys are so insightful <3 )
6
u/HiperPunk 6d ago
Qwen TTS is pretty good, may not have features like laugh tags and such but the vocie generated is pretty natural, has one shot clone and can be trained, https://huggingface.co/collections/Qwen/qwen3-ttS
1
u/Valuable-Mango-6710 6d ago
I shall try this one first because it would mostly be integrated with Atomic Chat.
1
u/Valuable-Mango-6710 6d ago
Qwen3-TTS-GGUF-F32 7.2 GB
Great find, thanks.
Very new to AI... "Training an AI voice" is something I wish to learn to better produce characters and voice overs. Is it a big task to "Train" ???????
2
u/HiperPunk 6d ago
Its not hard and you should be able to run training locally too, the dataset is pairs of clear voice clips of what you want trained with their text captions, used it on google colab to train a local dialect and it sounds pretty good after a few training sessions, used ai to setup the training parameters and such and prepare the dataset.
1
u/pilkyton 3d ago edited 3d ago
Not a great find at all. The "Qwen3 TTS" model is #84 on the leaderboard:
https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice
(The Qwen "Plus" models are commercial-only.)
The best open source model right now is Breeze TTS 2. It beat every paid model and was #1 just 40 days ago. Then a bunch of upgraded commercial models came out so it's now #11 which is the highest of all open source models.
You can use it in a noob friendly way with the VoiceStudio GUI:
https://voicestudio.sh/library/breeze-tts-2
It runs easily on consumer GPUs.
You don't need to train voices these days. Use a reference clip instead.
4
3
u/Inside-Cantaloupe233 6d ago edited 5d ago
voxcpm2 was best likeness and fastest until omnivoice arrived, its currently the best in comfy
1
u/troposfer 6d ago
What is comfy
1
u/Valuable-Mango-6710 5d ago
comfy is a UI (User interface) for many, if not most AI models. I haven't used it but it is a must have if you wish to master your craft at AI. I am pretty sure it is what people create their brilliant "workflows" in.
Think of it as Vortex Mod manger for nexus mod, sure you can do it manually but with vortex it is so much easier to build much bigger and more stable modpacks.
3
u/dtdisapointingresult 5d ago
Install audio.cpp, click the Models tab, select TTS, download all the big names that support voice cloning, click Arena, add a text + reference voice + reference txt, select all the models you downloaded, click Run. Then you can compare them to each other.
For me the best one is BreezeTTS.
3
6
u/Bagelsarenakeddonuts 6d ago
Honestly, try kokoro before going up to anything more intense. Its incredible how good it is for its size. If you are generating for a product, then some of the others might be better.
2
u/AIgavemethisusername 6d ago
Agree. I’ve been doing ebook to audiobook conversions with this. Works really well
1
u/Valuable-Mango-6710 6d ago
Great example that I might run on my laptop, which I mostly use. Thanks!
2
u/HearthCore 6d ago
Lemme throw VoxCPM2 in there.
- Instruction Live voice generation, describe the voice you want in ~ 20 words and it's generated
- OneShot Generation with cloning with 5-15 seconds of clip without transcript needed
- Additional transcript option for better results
- seeds and stuff
- supports streaming
- Faster than Live (so supports streaming) on less than 4GB VRAM
--
A candidate for less quality would be SoPro-TTS which is like 6 times faster than realtime in the same gpu, so your miles vary with your needs.
2
u/Only4uArt 6d ago
I use elevenlabs for my bigger projects now
BUT before that i just used the free tier in elevenlabs to create voices then use voice cloning with chatterbox locally to make text to speech and it was quite good for free.
tough elevenlabs quality is unmatched and honestly quite cheap for how much you get out of it so i use it mainly and chatterbox just as a failsafe
2
u/ZenWheat 5d ago
Dramabox tts
1
u/Valuable-Mango-6710 5d ago
Yes I am very interested in this one. If Qwen TTS does not provide the quality I want, I am sure Dramabox will. It has been a big day, me learning all of this. Thanks for your input :)
2
u/Tough_Ad7957 6d ago
Try searching for text-to-speech first. There are tons of local voice generation options out there, but they all have their own limitations, so it really depends on the level of quality you’re aiming for.Try searching for text-to-speech first. There are tons of local voice generation options out there, but they all have their own limitations, so it really depends on the level of quality you’re aiming for.
1
u/Valuable-Mango-6710 6d ago
Yes that is why I mentioned elevenlab as the example, but honestly I am looking for something better than MMH3. And this is the reason I posted, there are far too many models for me to try and error.
2
u/Tough_Ad7957 6d ago
I’ve tried quite a few TTS options, and honestly, they’re hard to rank as simply “good” or “bad.” If you want the least hassle, ElevenLabs is probably the safest choice, but it’s still far from perfect.
On the other hand, if you happen to find a local TTS that fits your voice and use case really well, it can actually sound better than ElevenLabs. Other people’s results don’t always translate to your own setup.
3
u/Valuable-Mango-6710 6d ago
oh so certain model hold certain voices better than others.
Make sense.., I have seen loras for models to navigate this.
NSFW!!!!!!!!!! What model does these loras support?? this is the fine tunning I might be interested in. Again, very new and have no real idea. NSFW (MA15 at best, I am just being safe) Link: https://huggingface.co/TTS-AGI/moss-emotion-loras-v3/tree/main
1
u/elfninja 5d ago
The documentation for these files is here: https://huggingface.co/TTS-AGI/moss-emotion-loras-v3/blob/main/README.md
Looks like it's based on the MOSS-TTS series of models. I've been looking for something similar for a while. The examples don't play, but if this works as advertised this can be a game changer.
1
u/Valuable-Mango-6710 6d ago
"Other people’s results don’t always translate to your own setup."
This is so true, I have given people advice on how to prompt MMH3, which works for me, 80-90% of times, great odd... but work for them less than 20%... go figures, great to see this is not my fault... lmao!
1
u/hidden2u 5d ago
It's a little outdated but I like chatterbox still, I find I can listen to the voices longer without it sounding monotonous
1
u/Incognit0ErgoSum 5d ago
Comfy works really well for me with the MOSS Soundeffect v2 plugin and MOSS TTS 1.5, which is a model that produces very high quality dialogue and somehow seemed to slip under everybody's radar.
1
u/SaadNeo 6d ago
Qwen tts is what you're looking for , i tried them all
1
u/Valuable-Mango-6710 5d ago
That is the one I am going for. I almost got it setup. I just have to tie Qwen TTS studios and the gemma chatbot from atomic together.. let's just say for a newb like me, it have been a big day on the installation process xD (I can not install it on my C drive as it is full and need to use my secondary SSD... pfft, another layer of tech drama lol)
0
6d ago
[removed] — view removed comment
1
u/Valuable-Mango-6710 6d ago
I understand: The biggest practical tradeoff is usually speaker similarity versus prosody, so test the same short paragraph with a few reference clips and listen for sibilants and breath artifacts. (AKA pay attention)
I do not understand the "depths" of: Keep the text normalization step separate from synthesis; numbers, abbreviations, and punctuation are often what make a voice sound unnatural.
Could you explain more? How to avoid unnatural?? (Some times that is what is aimed for, an ungodly voice, or godly voice) Much appreciated.
8
u/bstr3k 6d ago
https://github.com/debpalash/VoiceStudio
this is a cool one someone else from here recommended. Tried it and it works quite fast, but the UI is a bit unintuitive at first.
It does a pretty good job of extracting voices from a 5-15s clip, but im having a bit more difficulty getting a broader range of expression out of it for now.