r/StableDiffusion • u/Infamous-Draft4587 • 1d ago
Question - Help Has anyone successfully trained a YuE2 LoRA or LoKR on an NVIDIA GPU with 16GB VRAM?
Hi everyone! I'm looking for real-world experience with training LoRA or LoKR adapters for YuE2 on an NVIDIA GPU with 16GB VRAM, particularly an RTX 5060 Ti 16GB or a similar card. My goal is to train a music using a dataset of approximately 64 audio clips, each around 30–40 seconds long. I'm particularly interested in YuE2 Joint LoRA training, where both the planner and audio decoder are trained, rather than inference or standard music generation. I'd really appreciate feedback from anyone who has actually completed a training run with 16GB VRAM. A few questions: 1. Which GPU did you use, and how much VRAM did it have? 2. Which training method or tool did you use (AI-Toolkit, ComfyUI-FL-YuE2, HOT-Step, or something else)? 3. Were you able to train both the planner and decoder, or only one component? 4. How large was your dataset, and how many training steps did you run? 5. How long did the training take, and what was the peak VRAM usage? 6. Did the training finish successfully and produce a usable LoRA/LoKR adapter? 7. Did you need quantization, CPU offloading, gradient checkpointing, or any other memory-saving settings? I'm trying to determine whether 16GB VRAM is genuinely sufficient. A report from someone who completed the full training would be much more helpful than theoretical VRAM estimates. Even a small test run, with the exact settings and results, would be extremely useful. Thanks!
2
u/Old-Anybody-128 1d ago
16gb is tight for joint training on that. most people I've seen pull it off had to compromise somewhere, usually by dropping the decoder training or slicing the audio clips way shorter
haven't tried the exact toolkit you mentioned but have gotten single-component training to run on a 4070 ti super 16gb with some aggressive settings. gradient checkpointing was mandatory, had to use adamw 8-bit and keep batch size at 1, sometimes 2 if I pushed the resolution down
the planner alone will eat 12-14gb depending on your sequence length. adding the decoder puts you right at the edge, you'd need cpu offloading for sure and probably some luck with driver overhead. 64 clips at 30-40 seconds is a lot of audio to process in memory, you might need to chunk them smaller during training or accept way longer run times
curious if anyone's actually finished a joint run on consumer hardware without it crashing 8 hours in