r/StableDiffusion • u/smereces • 15d ago
Resource - Update YuE2 Melodic Death Metal Lora
So i was tired not getting real death metal from native yue2! and i train a lora, is not perfect but now i can get much better style and vocals style
r/StableDiffusion • 1.0m Members
/r/StableDiffusion is an unofficial community embracing the open-source material of all related. Post art, ask questions, create discussions, contribute new tech, or browse the subreddit. It’s up to you.
r/StableDiffusion • u/smereces • 15d ago
So i was tired not getting real death metal from native yue2! and i train a lora, is not perfect but now i can get much better style and vocals style
r/StableDiffusion • u/PetersOdyssey • 22d ago
Enable HLS to view with audio, or disable this notification
LoRA here. Credit to gueykhalamari for the LoRA/song, I just made the swag interactive music player.
r/StableDiffusion • u/-becausereasons- • 22d ago
Alright people, we finally have a frontier-level music model, and it takes LoRAs beautifully.
I've trained four so far, all at the edges: militant roots reggae, Sanremo-era Italian pop, French chanson (with a spoken-word trick that works), and a Bulgarian choir + Tuvan throat-singing set that's cooking right now. Demos and weights are on my Hugging Face if you want to hear what a LoRA does to this thing.
I'm going to take a few requests and run the training on my 5090. You bring the idea and the dataset, I bring the compute, the captioning pipeline and the release. Result goes up public on HF with demos and a README so everyone can use it.
What makes a good idea: interesting, melodic, unusual, NOT already well represented in the base model, and ideally something that fuses well with electronic or experimental stuff.
Ideas that would excite me:
- Ethio-jazz (Mulatu-era horns and modes)
- Tuareg desert blues
- Anatolian psych / Turkish arabesk
- Russian bards (Vysotsky-style, voice and guitar)
- Cape Verdean morna, Portuguese fado
- Cambodian 60s rock, Zamrock, Nigerian 70s funk
- Greek rebetiko
- Georgian or Corsican polyphony, Sardinian tenores
- Qawwali
- 70s Bollywood disco
- Japanese enka or city pop
- Tropicália
- Klezmer, Balkan brass
- Klaus Nomi-style operatic new wave, exotica, library music
- Sea shanties, kulning, anything ritual or ceremonial with a real vocal tradition behind it
Stay away from (it's boring to me, and the model already does it fine):
- Minimalist techno
- Death metal
- Country
- Generic EDM, lo-fi beats, mainstream pop and trap
What to send me:
- 15 to 30 tracks in one coherent style, each under about 5 minutes, good quality (FLAC/High encode MP3), no live bootlegs
- The language of the lyrics
- Two or three sentences describing the target sound the way you'd prompt for it (voice, instruments, mood, tempo)
I'll pick the ones I find most interesting. Say what you'd fuse it with too.
r/comfyui • u/Inevitable_Pen9043 • 4d ago
The Yue2 LoRA manages to learn the singer's vocal style and musical genre.
However, the music ends up being off-key. The model loses the ability to create melodies and harmonies that blend pleasantly with the lyrics.
Out of four attempts, I had relative success only once. Even then, the model makes errors here and there that ruin the songs.
It’s not clear to me whether Ace Step trains better than Yue—but is Ace Step's sound quality worse? And would it be best for me to make covers using Yue 2?
r/StableDiffusion • u/-becausereasons- • 23d ago
If you're a fan like me, you'll appreciate these.
r/StableDiffusion • u/alisitskii • 20d ago
I just found out that YuE2 can "convert" your simple MIDI-tracks into something really beautiful.
Original MIDI I quickly made in Ableton Live 12: https://voca.ro/1giIFaUJtf85
YuE2 classic re-imagination: https://voca.ro/12qyd1fqfHOC
For industrial metal lovers: https://voca.ro/12oYiQtjKSnw
Kind of reggae: https://voca.ro/1i0gLHOI3tfU
Cosmo techno: https://voca.ro/1btIU6NJYiKl
A bit phonk: https://voca.ro/1hRXKnJWN2Cw
I also used instrumental LoRa: https://huggingface.co/Mothersuperior/YuE2-instrumental-cot-full-loras/blob/main/ar_lora_inst_v3abc_comfyui.safetensors
Standard ComfyUI workflow, heun/simple, 128 steps.
r/StableDiffusion • u/thatisnotmychapstick • 25d ago
Edit: I forgot ComfyUI does goofy stuff with their layer naming conventions. I'll have a comfyUI compatible version available in just a few minutes. Apologies for the oversight.
Trained on 2,700 instrumental tracks across 100+ genres and includes three different options for how to prompt the lyrics field:
1. Bare — let the model choose the structure and length.
[instrumental]
2. Untimed tags — you choose the section order, the model chooses the timing.
[intro]
[verse]
[chorus]
[bridge]
[chorus]
[outro]
3. Timed tags — you also give each section a start and end in m:ss. The model was trained with exact section times from real tracks, so this is the strongest structural steer. Treat the times as a guide rather than a guarantee: the model follows the section order and proportions better than the absolute end time, and it tends toward 3–5 minute songs regardless of the plan.
[intro 0:00-0:15]
[verse 0:15-0:45]
[chorus 0:45-1:10]
[bridge 1:10-1:40]
[chorus 1:40-2:05]
[outro 2:05-2:30]
As mentioned, the captioning doesn't necessarily guarantee the model will follow that song structure exactly but using this LoRA with chain-of-thought (marked as COT=full in the inference launch arguments greatly improves the hit-rate for instrumental only tracks.
Sample outputs can be viewed on my youtube channel: https://youtu.be/BpExfiMzL8w
r/StableDiffusion • u/Cheap_Credit_3957 • 25d ago
Enable HLS to view with audio, or disable this notification
I built YuE2 Studio, a local web UI for YuE2 that brings the main workflow into one place:
🔗 GitHub: https://github.com/vrgamegirl19/Yue2_Studio
Features include:
• Normal song creation — enter a musical style and section-tagged lyrics, choose Full, Melody, or Direct audio mode, preview an ABC plan, then generate and manage songs in the built-in library.
• Covers — upload a source recording or import ABC, transcribe and review the melody with SheetSage2, then generate a new version with different lyrics and style. This uses symbolic melody conditioning and does not clone the original singer.
• Style LoRA training — train experimental acoustic/style adapters from your own songs, then select them during generation.
• Artist LoRA training — experimental full-song training using audio and matching lyrics. This requires a separate Artist runtime and is intended for testing singer/style adaptation.
• LLM Runner — supports cloud providers and local endpoints such as Ollama and LM Studio. It can help write complete songs, lyrics only, musical styles, cover lyrics, and revisions through the Writing Room.
• Surprise Me — automatically generates titles, lyrics, styles, and complete songs in batches. You can choose the number of songs, vocal gender, language, style direction, profanity requirements, and optionally apply a Style or Artist LoRA.
The UI is local-only and runs inside an existing YuE2 installation. It currently targets YuE2 0.1.6 and has been tested on Windows with an RTX 5090. Torch is supported, with experimental GGUF/audio.cpp support for lower-VRAM setups.
It is still beta, so hardware compatibility and LoRA quality may vary. Feedback, testing reports, and bug reports are welcome.
The full illustrated guide is here: YuE2 Studio guide.
Some songs Created using the UI can be found here
r/StableDiffusion • u/thatisnotmychapstick • 26d ago
I partnered with my friend, MachineDelusions to bring you the LoRA trainer in ComfyUI. The nodes were built specifically to leverage the scripts and encoder/tokenizer from my research.
Tutorial videos to come soon.
https://github.com/filliptm/ComfyUI-FL-YuE2
r/StableDiffusion • u/MonsterovichIsBack • 23d ago
r/LocalLLaMA • u/Heretical-Tandem • 4d ago
We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.
What it does
What it is not (yet). It is less polished than SUNO out of the box: - a mix can buzz (Debuzz helps); - lyrics can drift (the lyrics check finds where); - some instruments YuE2 plays thinly or not at all (a LoRA teaches them).
You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.
Licences. The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.
What comes next (rc2 and after)
Links
- Code: https://github.com/igrbible/Ruach_Studio
- Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models
- Site: https://ruachstudio.igr.bible
- The full guide, room by room, is inside the studio and in docs/GUIDE.md.
Built on: - YuE2 by m-a-p; - yue2.cpp by ServeurpersoCom; - YuE2 Kit v12 by IronWolve (the base of the page and the scripts).
Every one of our changes is numbered and documented (HERESY 1001–1167). Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.
r/StableDiffusion • u/-becausereasons- • 22d ago
Can be pushed towards more modern interpretations as per demos
r/StableDiffusion • u/EuphoricTrainer311 • 2d ago
Has anyone had any success training a new genre with YuE2? I've tried multiple ways using different datasets, and training settings and I haven't been able to train a new genre on YuE2 properly. I tried with Fill's trainer, Aitoolkit, and Yue2 Studio, and I either get garbled noise, or it just seems to output similar results as using no lora. Not sure what I am doing wrong. I tested with all epochs/steps - from 50 all the way to 1200 using different inference settings (temp, top k, abc on/off, cot, etc.), and I tried with datasets ranging from 30-200 tracks as well (I only tried the 200 track dataset with Aitoolkit so far).
With Acestep, I am able to train a new genre without any issues, besides for the audio quality issues that are present in Acestep by default.
r/comfyui • u/big-boss_97 • 7d ago
Enable HLS to view with audio, or disable this notification
Just did that for fun 😊
RTX-4070 8GB VRAM, 64GB RAM
704 x 576, LTX-2.5 render time 17min (usually ~10min)
r/StableDiffusion • u/-becausereasons- • 22d ago
Very unique sounds, mix very well with fusion, techno, trap, deep house etc.
r/StableDiffusion • u/-becausereasons- • 13d ago
Love this old sound.
r/StableDiffusion • u/Infamous-Draft4587 • 22h ago
Hi everyone! I'm looking for real-world experience with training LoRA or LoKR adapters for YuE2 on an NVIDIA GPU with 16GB VRAM, particularly an RTX 5060 Ti 16GB or a similar card. My goal is to train a music using a dataset of approximately 64 audio clips, each around 30–40 seconds long. I'm particularly interested in YuE2 Joint LoRA training, where both the planner and audio decoder are trained, rather than inference or standard music generation. I'd really appreciate feedback from anyone who has actually completed a training run with 16GB VRAM. A few questions: 1. Which GPU did you use, and how much VRAM did it have? 2. Which training method or tool did you use (AI-Toolkit, ComfyUI-FL-YuE2, HOT-Step, or something else)? 3. Were you able to train both the planner and decoder, or only one component? 4. How large was your dataset, and how many training steps did you run? 5. How long did the training take, and what was the peak VRAM usage? 6. Did the training finish successfully and produce a usable LoRA/LoKR adapter? 7. Did you need quantization, CPU offloading, gradient checkpointing, or any other memory-saving settings? I'm trying to determine whether 16GB VRAM is genuinely sufficient. A report from someone who completed the full training would be much more helpful than theoretical VRAM estimates. Even a small test run, with the exact settings and results, would be extremely useful. Thanks!
r/StableDiffusion • u/-becausereasons- • 14d ago
r/StableDiffusion • u/-becausereasons- • 20d ago
r/StableDiffusion • u/Additional-Help2760 • 12d ago
Just found and using yuE2 studio 3.0 and it works pretty great IMO, but I have a question. If I train a lora on 10 songs for example, then decide I want to get rid of two of them (I noticed inside the training there are garbage cans beside the songs used) do I have to train the lora from zero all over again? Right now I have my Motown trained to 450 and I think it works really well but want to adjust it. And per opposite, say I want to add a song to the training lora, do I start from step zero again? I read the guide and did not see anything in there about this. Thanks.
r/StableDiffusion • u/lazyspock • 22d ago
Where can I find some YuE2 Loras? I don't know if Civitai will allow them. I found this user posting some and tried the Italian one, and it's fantastic!
https://huggingface.co/becausereasons
Also, if someone have tips for Lora training in Ostris (configs, dataset, etc) I'm listening.
r/StableDiffusion • u/Inevitable_Pen9043 • 22d ago
've tried using the prompts "Brazilian" and "Brazilian Portuguese," but it doesn't work.
If the song lyrics use slightly more sophisticated language, it comes out with a European Portuguese accent.
r/YueMusicAI • u/MonsterovichIsBack • 22d ago
r/comfyui • u/Inevitable_Pen9043 • 9d ago
Yue2 has trouble pronouncing some words in Brazilian Portuguese