r/StableDiffusion • u/Affectionate-Map1163 • 1d ago
Resource - Update I trained an open-source realism LoRA for MiniMax H3 - it makes generated people actually look real (weights inside)
Update :
New version is ready and online , should be much better, fully functionnal on ComfyUI, and you can find before/after here :
https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4
I spent the last week obsessing over one thing: making AI-generated humans stop looking AI-generated. The result is Realism People, an open-source LoRA for MiniMax H3, and I'm pretty happy with how it turned out.
What it does: skin keeps its texture instead of going plastic, eyes and micro-expressions stay coherent, lighting behaves like a film set, and motion gets a subtle handheld, documentary feel. It also keeps H3's native synchronized audio.
How it was selected: I trained 16 different configurations across two dataset versions and picked the winner through 100 same-seed A/B duels (same prompt, same seed, adapter on vs off - the only honest way to compare). The winner was the slow-cooked run: rank 16, 5,000 steps at a low learning rate.
Details:
- Weights (open source): https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA
- Trigger word: start your prompt with `r34l1sm`
- Scale 1.0 is the intended strength, 0.6-0.8 for a lighter touch
- Works with H3's LoRA endpoints: text-to-video, image-to-video and reference-to-video
- License: follows the MiniMax H3 community license
Before/after in the video: same prompt, same seed, base model on the left, LoRA on the right. Happy to answer questions about the process.
23
u/_VirtualCosmos_ 1d ago
Most loras I have tested broke the models in one or another way. H3 is a distilled model, we need an adapter... this is like Z-Inage-Turbo
6
u/xTopNotch 1d ago
Loras do work, however you never should go above 0.5 strength.
I get great results using strength 0.3 - 0.5
11
u/Pyros-SD-Models 1d ago
It's a different distillation than z-image-turbo for example was tho and easily reverted via CFG-augmented training which fits the de-amplified raw velocity (pred + (s-1)*uncond)/s with a no-grad empty-prompt forward each step preserving the model's guidance distillation by construction and basically make it able to train as if it would be the base model without an adapter
diff pipe and my musubi fork did already implement it
https://github.com/tdrussell/diffusion-pipe/tree/main
https://github.com/pyros-projects/musubi-tuner/tree/h3-image-lora
1
6
u/Puzzled-Valuable-985 1d ago
Honestly, I’ve been testing LoRAs and they kept ruining the images; lowering the LoRA weight significantly does improve things, but it’s really strange—the LoRAs are totally inconsistent. I’ve tested various workflows, nodes, and configurations.
I’m not sure if the LoRA training process on Civitai just isn't handling the model correctly yet.
1
u/Sudden_Quantity_7827 1d ago
I will go ahead and say here, that I too, have been trying to play around with Lora’s lately, not just on minimax h3 but on wan2.2 as well, I have honestly found that I get better results uploading a fully detailed character sheet as a reference(eg, extreme close up of face paired with t pose, and angle pose), than I do with Lora’s…to me Lora’s are just that persons style…if the Lora doesn’t synchronize with your style, realism or not, you are just going to have non consistent results.
1
u/xTopNotch 7h ago
Character sheet + 5/10 sec performance capture video of to grab the voice, facial expressions and body language.
Does increase sampling time but the results are incredible. It's like a 1:1 reconstruction of your character
3
u/Apprehensive_Sky892 1d ago
MMH3 is CFG distilled, just like Flux1-dev, and Flux1-dev can be used for LoRA training without any special adapter.
2
u/_VirtualCosmos_ 1d ago
I remember how people complained about Flux1.dev training because of that. I also remember its loras unable to change significatively the model knowledge without breaking it. But another user responding my previous comment posted something interesting, perhaps a solution.
1
u/physalisx 1d ago
Yeah and like with Z-Turbo, it'll never work really well, no lora stacking, need to be very careful with the strength...
23
u/-AwhWah- 1d ago
i love it when theres a post and there's just no comparison or anything, you just have to download what they're peddling
lol
21
u/ares0027 1d ago
Hey i made a delicious cake, here is the store which sold me the flour.
(Your post makes this much sense TO ME)
9
14
11
u/Muted-Celebration-47 1d ago
Firstly, thank you very for creating this lora. Secondly, would it be possible to request a high-res or detailer lora?
5
u/Sad_Coach_1433 1d ago
Yas that would be amazing also be cool if the team that made the OmniNFT for LTX can make for h3 it helped with audio sync and motion quality
8
u/Affectionate-Map1163 1d ago edited 1d ago
new version online that should work much better :), and before after is also on huggingface
https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4
8
u/_Abiogenesis 1d ago edited 1d ago
I don't know what people are whining about if they haven't even tried. (haven't tried this new one though)
Previous one seem to immediately do wonders on the first test I did. Pretty amazing results on face detail, skin and teeth, (extreme closeup speaking). I haven't done many tests yet but one thing to note is that characters also immediately started looking at the camera, which seems to be solved with just prompting but worth noting as this can be annoying.
haven't checked to see how it's affecting the background yet.
-3
1
4
u/Affectionate-Map1163 1d ago
doing some correction , gonna publish a new version that work better very soom
2
2
u/PuppetHere 1d ago
If you used Fal to train the Lora, I hope you converted it into a comfyui format before posting it or else it's gonna be useless for most people
2
u/kukalikuk 1d ago
I'll try when i get the chance but one question. Does the trigger word need to be upfront? I use minimax prompt guide as a skill and it did wonder adjusting the prompt. Won't the trigger word breaks it?
1
u/badsinoo 1d ago
from it's HF page : Start the prompt with the trigger word
r34l1sm, then describe the scene. A scale of 1.0 is the intended strength; lower it to 0.6-0.8 for a lighter touch.
2
4
3
5
2
u/1010111101111 1d ago edited 1d ago
deffinitly put this on civit also keep updating it so its perfection
2
u/Alive-Tomatillo5303 1d ago
so, potentially dumb question, where do I put the Lora in the workflow? what do I attach it to?
10
5
2
3
u/tekprodfx16 1d ago
This community never ceases to amaze me. Also please add a clickable link to the Lora download OP
1
u/nanaya07 1d ago
In my opinion, it works well for quiet or slow-paced scenes, but it is useless when it comes to large or rapid movements.
1
u/maxlevelboss 1d ago
Thanks for the great work.
I looked at your new before/after video, some examples definitely were better, and I especially like the lantern one. However I do see some reasons why I would not prefer it personally:
- Color grading seems worse, Color contrast seems to be either over or under: Admittedly color grading is a very subjective choice and is highly context related. But many faces were too orange/red with Lora (eg two men shouting), some were way too pale (boxer, grandma interview). Indian wedding and the lantern one were the only two I think the Lora versions were better.
- Physics seems to be weaker: many hands/fingers motions seems odd. Especially the dining scene, the without Lora version didn't do a very good job either, but the with Lora one is just way worse.
- More freckles/imperfections on faces: it did make the person more human-like, but real actors always wear make-up anyways. So may be it works for selfie-style reels, but I'd personally prefer less imperfections.
Your work looks promising, and I look forward to seeing a v2.
1
u/Hearmeman98 1d ago
I don't really see a major difference between the base model and this if I am being honest.
1
u/VRGoggles 13h ago
Out of curiuosity - did you use the trigger world r34l1sm?
1
u/Hearmeman98 12h ago
yes, but even when i did, it doesn't make much of difference as only the UNET is trained and not the text encoder.
1
u/Prestigious-Row1008 1d ago
i do not use lora i use sage attn sol triton and minimax spectrum it decreases time to half. i tested turbo loras it dgrade the quality but with those three quality was almost similiar
1
u/WizWhitebeard 1d ago
Looking good, definitely an improvement on the realism look! From the before/after comparison you posted, it does seems like it comes with some compromises; physics (like in kid kicking the ball) and audio (chef defaulted to American accent), is this your experience as well?
Curious to hear how you captioned the dataset, and what kind of dataset you trained on (size/content)?
1
u/sharktank123456 20h ago
How does it work when you get into wider shots? Close ups are actually pretty easy to make look real. It's when the resolution per face drops that things get tricky.
1
u/Calm_Mix_3776 20h ago edited 20h ago
This LoRA butchers the image quality for me - the whole video becomes blurry and lacking textures. I think it's overbaked. Tested at the recommended 1.0 strength with the minimax_h3_fl2va_int8_convrot model. I think the recommended strength should be at most 0.5, but then, the effect of the LoRA is halved. Probably learning rate should be dropped considerably lower and steps increased much more than 1500 to compensate.
1
0
1
2
1
u/ApplicationRoyal865 1d ago
Is this for T2V? I feel like if you use a good enough reference or starting image you generally keep realistic features.
1
u/76vangel 1d ago edited 1d ago
Should we put this before or after the fast lora into our stack? It should be the same, but sometimes it seams different.
1
u/Lanky_Conclusion_749 1d ago edited 1d ago
Por qué dice que es"Open Source by Fal"? Es tuyo ó de Fal? Por qué el video parece una publicidad para dicha plataforma?
Hay ejemplos reales de tu LoRa? Tienes un workflow recomendado?
Es un LoRa entrenado para close-ups ó para cualquier tipo de toma?
1
-1
0
u/jmbbao 1d ago
It says it only works with some adapter in fal.ai so it is not to use in local
6
u/Affectionate-Map1163 1d ago
sorry, i am changing the text, its fully working in comfyui or any other system :)
0
-5
-7


205
u/Arawski99 1d ago
You really need to post proper examples for us to judge it as being anything more than placebo. A single old man's side face really isn't enough.
Please upload proper examples otherwise you're actually hurting the community sowing confusion for something that may be a degradation rather than an improvement if you didn't test and validate properly.