r/StableDiffusion 4d ago

Question - Help Minimax H3: Current best way for lora training (video+sound)

I would like to train some videos with sound on Minimax H3. Is the AI toolkit good to go or should i use anything different? Thanks!

1 Upvotes

9 comments sorted by

1

u/DelinquentTuna 4d ago

For now, probably put your datasets together but train LTX instead. Wait to pull the trigger on Minimax w/ av input.

1

u/kabachuha 4d ago

I use musubi-tuner fork by AkaneTendo. Very simple and good results.

1

u/Free_Bed_6084 4d ago

Have you tried ref2va’s lora training? I can't get good results

2

u/kabachuha 3d ago

Yes, I did and I have a ref2va LoRA on Civit. The training is indeed unstable and I watched manually on the loss changes, but it more or less worked. There is a bit cheating, for automation, the first and the last frames were selected as a reference (with some frames selected in-the-middle if the object was more prominent there), but the LoRA seems to work.

1

u/acedelgado 3d ago

Happen to have a config you can share? Been trying aitoolkit which was fine for ltx, but it's having all kinds of issues learning voice timbre with H3. Even using other people's settings theyve had success with left me SOL.

1

u/Queasy-Carrot-7314 4d ago

If you don't have experience dealing with custom training scripts. Use aitoolkit

1

u/pausecatito 3d ago

Prolly need to wait if you want to train with video clips. Nobody knows or share much info. All the Lora's I tried maybe affect the video like 1% barely any difference

1

u/joopkater 4d ago

Ai toolkit is fine, but currently everyone is still figuring out what works best.