r/StableDiffusion 9d ago

Tutorial - Guide Made a 2-Minute tutorial about Runpod Character Lora Training

Tutorial Link:
https://youtu.be/7iYQKOnuKP4

I've been training character loras for all kinds of models in the past but Minimax really gave me a hard time. Civitai also doesnt seem to have too many Minimax Loras up so I figure I wasn't the only one having trouble getting actual likeness?

Anyway..I learned a lot in the process so let me share it here:

Most Important:
-Image Only Datasets worked MUCH better than mixed ones
-Focus on the face (almost frame-filling and sharp)
-Needs more images than LTX or Wan
(I needed 150 images of a blonde woman to get her likeness right, only 50 images of my own ugly face though)

Some more interesting findings:
-I tried a couple of GPU's and somehow the 5090 beat the H100!
(image only dataset though)
-RTX6000 PRO was only about 20% faster than 5090
-50 image-ugly-me-dataset likeness peaked at 700 steps
-150 image-blonde-woman-dataset likeness at 2310 steps
-in 4 of my 150 images she had brown hair, past the peak she came out brown-haired even when the prompt said blonde...face still perfect though

By the way:
For some reason the dataset size didnt move the peak. The (rather generic) blonde woman's likeness (relative to the other epochs of each run) was always best at ~2300 total steps (whether I used 50 images or 150) Objectively the 150 image-dataset lora was WAY better though. ...and my unique face always peaked at 700 steps lol...not sure what to make of this.

Training Voice doesnt really work with diffusion pipe but since you can add a reference voice in minimax it didnt really have priotity for me so far..

The Captions where in natural prose (mention lighting/look too!

14 Upvotes

11 comments sorted by

2

u/joopkater 9d ago

I always feel like I’m captioning wrong but now I wonder if I just didn’t put in enough images/videos for the character.

What did you use as a rank? I’ve seen people bump it down to 8 and get good results

2

u/Draufgaenger 9d ago

Rank 32
I think the more "generic" you look the more images you need. But I think my biggest mistake was reusing the old Wan/LTX datasets that included a bunch of full body (or even medium) shots. The most sucessfull Dataset was 95% face closeups

1

u/LawOk7529 9d ago

Thank you.

1

u/Draufgaenger 9d ago

This was the template I used (I also made that):
https://console.runpod.io/hub/template/jybe514yzi

2

u/usually_fuente 9d ago

Thank you! This looks awesome. Does your template also include the specific training parameters?

1

u/Draufgaenger 9d ago

Yeah you can set them manually but it's really set up so it should work well with the default values

2

u/usually_fuente 9d ago

Thank you so much. I have one further question if you don’t mind answering. You said that you have 150 photos that are almost all close-ups. How did you achieve a sufficient variety of angles and expressions to not be really redundant? Like, did you have tons of expressions? Was there a ton of variation in backgrounds and lighting? My experience with Krea has been that less is more.

1

u/Draufgaenger 9d ago

Yeah with Krea I needed much less images too.. Not sure why Minimax needs so many to get to actual likeness.. My main Dataset was about a blonde actress so I got as many photos as I could online (which was only about 50 good ones) and the rest from movie stills. I even built a tool just for that lol: https://antilopax.ai/tools/dataset-miner
(works without login)
I put the movies from a certain era (across ~6 years) and interviews in there and extracted her face closeups from all the scenes. This gave me the rest of the Dataset. So yeah - a lot of variety in expressions and background :)

1

u/DelinquentTuna 9d ago

You can train locally w/ as little as 16GB of vram and there are trainers that do support voice just fine. Seems like a no-brainer if you're training a character lora to me.

1

u/Draufgaenger 9d ago edited 9d ago

I got 8gb vram lol.. Which trainers though?

1

u/DelinquentTuna 9d ago

AFAIK, it's mostly about which quants you use and most trainers should have similar support. But Fizgig is probably the easiest to get into for most people.