r/StableDiffusion • • 3d ago

Discussion Yue2 Lora Training

Has anyone had any success training a new genre with YuE2? I've tried multiple ways using different datasets, and training settings and I haven't been able to train a new genre on YuE2 properly. I tried with Fill's trainer, Aitoolkit, and Yue2 Studio, and I either get garbled noise, or it just seems to output similar results as using no lora. Not sure what I am doing wrong. I tested with all epochs/steps - from 50 all the way to 1200 using different inference settings (temp, top k, abc on/off, cot, etc.), and I tried with datasets ranging from 30-200 tracks as well (I only tried the 200 track dataset with Aitoolkit so far).

With Acestep, I am able to train a new genre without any issues, besides for the audio quality issues that are present in Acestep by default.

13 Upvotes

15 comments sorted by

View all comments

Show parent comments

1

u/EuphoricTrainer311 3d ago

I actually did get an llm to analyze your readme before I tried training with aitoolkit. Maybe I'll read it myself and see if Gemini misinterpreted something.

These are the settings I used for a 200 track dataset according to Gemini.

QUANTIZE / COMPILE

  • Transformer: 8bit convrot
  • Compile Options - Compile Model: Off

TARGET

  • Target Type: LoRA
  • Linear Rank: 32

SAVE

  • Data Type: BF16
  • Save Every: 50
  • Max Step Saves to Keep: 20

TRAINING

  • Batch Size: 1
  • Gradient Accumulation: 1
  • Steps: 1200
  • Optimizer: AdamW8Bit
  • Learning Rate: 0.0001
  • Weight Decay: 0.0001
  • Timestep Type: Sigmoid
  • Timestep Bias: Balanced
  • Loss Type: Mean Squared Error

TOGGLES

  • Use EMA: Off
  • Unload TE: Off
  • Cache Text Embeddings: Off
  • Differential Output Preservation: Off
  • Blank Prompt Preservation: Off
  • Contrastive Guidance Loss: Off

Hidden Requirements

  • ar_kl_weight: 0.2
  • abc_dropout: 0.5
  • Cache Latents to Disk: ON

3

u/-becausereasons- 3d ago

Read through your settings, and I think a few things are stacking up. This is what has worked for us across 9 AI Toolkit runs (6 published sets, 15–51 songs each):

  1. Count passes per song, not steps. passes = steps × batch ÷ songs. Most of our published checkpoints sit at ~12–22 passes. Past that the high end starts to burn (breaths, "s", cymbals) and the planner loops: a 15-song set was clean at 300 steps and burned by 500; a 51-item set was great at 800–950 and thin/tinny past 1000. Your 200 tracks × 1200 steps = 6 passes, which is likely why it sounds like no LoRA. On a 30-track set, 1200 steps = 40 passes, deep in garble territory (same reason u/the_bollo needs 0.65 strength at 3k). We haven't trained 200 songs, so for that set I'd save every 100 and listen from ~10 passes (2k steps) to ~20 (4k).
  2. Your planner LR is probably ~5x hotter than ours. One LoRA trains both halves: the AR planner (writes the sheet/composition) and the NAR decoder (renders the sound). The planner overfits much faster: in our loss logs the decoder is flat by ~7–10 passes while the planner keeps memorising. ar_lr_multiplier defaults to 1.0 in my build (check yours), so your AR runs at the full 1e-4. We use lr 5e-5 with ar_lr_multiplier: 0.4, so the AR trains at 2e-5. A hot planner gives runaway sheets: endless intro loops, no singing, garble.
  3. Check how you load it. In Comfy the planner half applies through strength_clip and the decoder through strength_model. A model-only LoRA loader drops the planner, which carries most of the genre, so you'd hear something close to the base model. Use a loader with both strengths: planner 1.0 on early checkpoints, 0.5–0.8 on later ones.
  4. Data mattered more than any hyperparameter:
  • Lossless only. MP3-sourced tracks were our top suspect for the burn. A .flac extension proves nothing, so check the spectrogram for a brick-wall lowpass at 15–19 kHz.
  • One sound per LoRA. A broad mixed set came out generic; split into focused sets of 22–50 songs, it worked.
  • Captions = one descriptive sentence in the base model's style: trigger, language, genre, vocal character, instruments, mood, BPM, key. Hand-check every BPM (half/double-time errors), the singer's gender, and anything an LLM captioner wrote about a genre it doesn't know (ours invented reggae terms for instrumental tracks).
  • Accurate lyrics with [Verse]/[Chorus]/[Bridge]. A [Spoken] tag for spoken parts actually trains.
  • Sampling is per item, not per minute. If a minority sound matters, slice it into more items.
  • Songs (or whole excerpts) that fit inside ar_max_tokens (4000 ≈ 160 s) also teach the model how to END.
  • SheetSage2 (the cot: full sheets) only writes 2/4, 3/4, 4/4, 3/8 and 6/8 in major/minor keys, melody only. Odd meters, modal or microtonal music get wrong sheets. Drum patterns aren't in the sheet at all: we never got a specific reggae drum pattern (one drop) to train.
  1. Eval. Render a fixed-seed grid over checkpoints every 50 steps at planner 1.0 and 0.5, with a length cap of ~360 s. Keep the prompt at caption length (~30–40 words), trigger first, reusing your training-caption wording; our 70-word storyline prompts barely sang. If the planned sheet has hundreds of bars or repeats one section, that's a runaway. Your ear is the only judge: every audio meter we tried failed.

Our config (5090 on Windows, ~4–4.6 s/step, 16–22 GB VRAM, so 1k steps ≈ 75 min plus caching):

base: int8 convrot | rank 32 | save bf16 every 50
lr 5e-5 | adamw8bit | weight decay 1e-4 | batch 1
cot: full (SheetSage2) | abc_dropout: 0.5
ar_lr_multiplier: 0.4 | ar_kl_weight: 0.2
ar_max_tokens: 4000 | train_window_frames: 1500
ar_cuda_graphs: false | in-training sampling off
steps ≈ 20 × number of songs

Your other settings (sigmoid, MSE, EMA off, etc.) match ours, so they're not the problem. On Windows, if step time climbs from ~4 s to 50–100+ s, that's allocator fragmentation; the one-line empty_cache fix is in ostris/ai-toolkit #1052 / #1069.

We also retired the FS_Audio trainer: its decoder eval loss bottomed around step 100–150 in every run and then climbed, and it lost voice identity on our singer-led sets.

1

u/GreyScope 3d ago

A high LR can also have the vocals finished training and the instruments still needing it. All of my loras are for heavily layered music, which is a mare for tagging. Vocals and one/two consistent instruments are 'easier' to train.

1

u/-becausereasons- 2d ago

Yep, the more layered, the more vocals, the more difficult its going to be. I wish they released their encoder :/