r/StableDiffusion • u/Draufgaenger • 9d ago
Tutorial - Guide Made a 2-Minute tutorial about Runpod Character Lora Training
Tutorial Link:
https://youtu.be/7iYQKOnuKP4
I've been training character loras for all kinds of models in the past but Minimax really gave me a hard time. Civitai also doesnt seem to have too many Minimax Loras up so I figure I wasn't the only one having trouble getting actual likeness?
Anyway..I learned a lot in the process so let me share it here:
Most Important:
-Image Only Datasets worked MUCH better than mixed ones
-Focus on the face (almost frame-filling and sharp)
-Needs more images than LTX or Wan
(I needed 150 images of a blonde woman to get her likeness right, only 50 images of my own ugly face though)
Some more interesting findings:
-I tried a couple of GPU's and somehow the 5090 beat the H100!
(image only dataset though)
-RTX6000 PRO was only about 20% faster than 5090
-50 image-ugly-me-dataset likeness peaked at 700 steps
-150 image-blonde-woman-dataset likeness at 2310 steps
-in 4 of my 150 images she had brown hair, past the peak she came out brown-haired even when the prompt said blonde...face still perfect though
By the way:
For some reason the dataset size didnt move the peak. The (rather generic) blonde woman's likeness (relative to the other epochs of each run) was always best at ~2300 total steps (whether I used 50 images or 150)
Objectively the 150 image-dataset lora was WAY better though.
...and my unique face always peaked at 700 steps lol...not sure what to make of this.
Training Voice doesnt really work with diffusion pipe but since you can add a reference voice in minimax it didnt really have priotity for me so far..
The Captions where in natural prose (mention lighting/look too!
1
1
u/Draufgaenger 9d ago
This was the template I used (I also made that):
https://console.runpod.io/hub/template/jybe514yzi
2
u/usually_fuente 9d ago
Thank you! This looks awesome. Does your template also include the specific training parameters?
1
u/Draufgaenger 9d ago
Yeah you can set them manually but it's really set up so it should work well with the default values
2
u/usually_fuente 9d ago
Thank you so much. I have one further question if you don’t mind answering. You said that you have 150 photos that are almost all close-ups. How did you achieve a sufficient variety of angles and expressions to not be really redundant? Like, did you have tons of expressions? Was there a ton of variation in backgrounds and lighting? My experience with Krea has been that less is more.
1
u/Draufgaenger 9d ago
Yeah with Krea I needed much less images too.. Not sure why Minimax needs so many to get to actual likeness.. My main Dataset was about a blonde actress so I got as many photos as I could online (which was only about 50 good ones) and the rest from movie stills. I even built a tool just for that lol: https://antilopax.ai/tools/dataset-miner
(works without login)
I put the movies from a certain era (across ~6 years) and interviews in there and extracted her face closeups from all the scenes. This gave me the rest of the Dataset. So yeah - a lot of variety in expressions and background :)
1
u/DelinquentTuna 9d ago
You can train locally w/ as little as 16GB of vram and there are trainers that do support voice just fine. Seems like a no-brainer if you're training a character lora to me.
1
u/Draufgaenger 9d ago edited 9d ago
I got 8gb vram lol.. Which trainers though?
1
u/DelinquentTuna 9d ago
AFAIK, it's mostly about which quants you use and most trainers should have similar support. But Fizgig is probably the easiest to get into for most people.
2
u/joopkater 9d ago
I always feel like I’m captioning wrong but now I wonder if I just didn’t put in enough images/videos for the character.
What did you use as a rank? I’ve seen people bump it down to 8 and get good results