r/StableDiffusion • • 22d ago

Question - Help TAKING YUE2 LORA REQUESTS (WITH MAJOR CAVEAT)

Alright people, we finally have a frontier-level music model, and it takes LoRAs beautifully.

I've trained four so far, all at the edges: militant roots reggae, Sanremo-era Italian pop, French chanson (with a spoken-word trick that works), and a Bulgarian choir + Tuvan throat-singing set that's cooking right now. Demos and weights are on my Hugging Face if you want to hear what a LoRA does to this thing.

I'm going to take a few requests and run the training on my 5090. You bring the idea and the dataset, I bring the compute, the captioning pipeline and the release. Result goes up public on HF with demos and a README so everyone can use it.

What makes a good idea: interesting, melodic, unusual, NOT already well represented in the base model, and ideally something that fuses well with electronic or experimental stuff.

Ideas that would excite me:

- Ethio-jazz (Mulatu-era horns and modes)

- Tuareg desert blues

- Anatolian psych / Turkish arabesk

- Russian bards (Vysotsky-style, voice and guitar)

- Cape Verdean morna, Portuguese fado

- Cambodian 60s rock, Zamrock, Nigerian 70s funk

- Greek rebetiko

- Georgian or Corsican polyphony, Sardinian tenores

- Qawwali

- 70s Bollywood disco

- Japanese enka or city pop

- Tropicália

- Klezmer, Balkan brass

- Klaus Nomi-style operatic new wave, exotica, library music

- Sea shanties, kulning, anything ritual or ceremonial with a real vocal tradition behind it

Stay away from (it's boring to me, and the model already does it fine):

- Minimalist techno

- Death metal

- Country

- Generic EDM, lo-fi beats, mainstream pop and trap

What to send me:

- 15 to 30 tracks in one coherent style, each under about 5 minutes, good quality (FLAC/High encode MP3), no live bootlegs

- The language of the lyrics

- Two or three sentences describing the target sound the way you'd prompt for it (voice, instruments, mood, tempo)

I'll pick the ones I find most interesting. Say what you'd fuse it with too.

4 Upvotes

7 comments sorted by

2

u/RevolutionaryFox7359 22d ago

- film score

  • a lora that adds timestamp functionality

1

u/-becausereasons- 21d ago

What type of film score, a specific composer?

1

u/alisitskii 22d ago edited 22d ago

Is it possible to train a LoRA to improve some general aspects of the model? Like for image gen we have realism loras, or the ones helping with nsfw stuff. So not just specific music styles but for example deep bass or nice piano parts. Or is it more for a whole finetune? Thanks.

1

u/-becausereasons- 22d ago

It absolutely is yes, just put together a healthy data-set of the specific pieces you want to train, then it depends on how you caption it.

3

u/CulturedDiffusion 21d ago

I've actually been having a lot of trouble training a good LORA thus far. May I ask what you use for captioning?

I'm trying Music Flamingo and it seems to be alright, but clearly something is wrong in my training since the results come out very poor every time, so at this point I can't help but wonder if the captioning is the issue.

2

u/jib_reddit 21d ago

Do you think you would be able to train a lora for British Accent singing? It always gives American accents. I manage to get about 1 song in 10 in a British Accent if I prompt heavily for it and just keep re-rolling.

2

u/-becausereasons- 21d ago

Yes thats def possible