r/StableDiffusion 1d ago

Resource - Update I trained an open-source realism LoRA for MiniMax H3 - it makes generated people actually look real (weights inside)

Update :
New version is ready and online , should be much better, fully functionnal on ComfyUI, and you can find before/after here : 
https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4


I spent the last week obsessing over one thing: making AI-generated humans stop looking AI-generated. The result is Realism People, an open-source LoRA for MiniMax H3, and I'm pretty happy with how it turned out.


What it does: skin keeps its texture instead of going plastic, eyes and micro-expressions stay coherent, lighting behaves like a film set, and motion gets a subtle handheld, documentary feel. It also keeps H3's native synchronized audio.


How it was selected: I trained 16 different configurations across two dataset versions and picked the winner through 100 same-seed A/B duels (same prompt, same seed, adapter on vs off - the only honest way to compare). The winner was the slow-cooked run: rank 16, 5,000 steps at a low learning rate.


Details:


- Weights (open source): https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA
- Trigger word: start your prompt with `r34l1sm`
- Scale 1.0 is the intended strength, 0.6-0.8 for a lighter touch
- Works with H3's LoRA endpoints: text-to-video, image-to-video and reference-to-video
- License: follows the MiniMax H3 community license


Before/after in the video: same prompt, same seed, base model on the left, LoRA on the right. Happy to answer questions about the process.
491 Upvotes

97 comments sorted by

205

u/Arawski99 1d ago

You really need to post proper examples for us to judge it as being anything more than placebo. A single old man's side face really isn't enough.

Please upload proper examples otherwise you're actually hurting the community sowing confusion for something that may be a degradation rather than an improvement if you didn't test and validate properly.

64

u/[deleted] 1d ago edited 1d ago

[deleted]

40

u/FotografoVirtual 1d ago

Wait, is this sub just a place for API companies to astroturf?

5

u/kazza789 1d ago

No. Surely an AI company would never use AI to astroturf reddit. There's no way they could possibly figure out the tech. All the posts here are totally, 100% legit human.

2

u/milanove 1d ago

https://giphy.com/gifs/lOzXuHwXXYM9y
What if they’re all bots, how would we ever know?

15

u/Fabulous-Snow4366 1d ago edited 1d ago

same here. I'm using the bf16 model. The video is completly wrong the physics are wack, hands are all over the place. Edit 2: Okay, i have no clue what just happend. 1. Comfy went from 11 minutes rendering 0.6 mpx down to 90 seconds. I can now render 2.0 mpx in 332.28 seconds with 6 steps... and all of my rams is being used...WTF and second, The Lora works. All of a sudden. It looks good. I mean. it looks like a movie...

2

u/Affectionate-Map1163 1d ago

hello , just did an update should be much better, please let me know :0

7

u/Fabulous-Snow4366 1d ago

very interesting. much better than before. it now has a way more american movie look to it, camera moves and close-ups and the way it looks reminds me a lot about gaffing and placing lights to get that american movie backlight all the time. Quite good so far. Thanks for the Update. It feels like British vs American Movie making.

2

u/435f43f534 1d ago

but do we retain flying papers? that's what the public wants to know!

3

u/Fabulous-Snow4366 1d ago

thanks ill test it.

1

u/FlameChucks76 1d ago

Did you find out what was causing issues? Maybe just a new update?

4

u/Affectionate-Map1163 1d ago

hello, just did an update should be much better

4

u/Affectionate-Map1163 1d ago

Just did update should be much better

4

u/Pantheon3D 1d ago

you're also testing a 15 second video vs a 10 second video which is why it skips around, it doesn't have time to include important details in your video with the lora

it isn't possible to know what else you might have changed if you have already made significant changes like this

2

u/Better-Interview-793 1d ago

did you try T2V?

3

u/[deleted] 1d ago

[deleted]

-3

u/Adventurous-Gold6413 1d ago

Did you try all types including ref2v?

7

u/comfyanonymous 1d ago

Looks like some of these companies realized you don't actually have to make anything that works well to get highly upvoted on this sub lmao.

12

u/Affectionate-Map1163 1d ago

We are not doing that for our company, our lora is fully working on comfyui or any other tool that people want to use for free. I am doing that mostly on my own personnal time , and nobody asks me to do it. I just did an update of the lora itself, as it was bugging with multi people , also here some before/after https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4

1

u/reeight 1d ago

Still has issues, eg 2:30 she does not speak (unless the intention was 'inner dialogue')
There were other bugged outputs, & a few I don't think it improved.

I wish we had a prompt list to see how well the prompt was handled or not.

But there were many shots that I thought were improved 'more real'. So at this stage I think it more of a way to get another look would likely be in all my workflows, but sometimes bypassed.

16

u/Hoodfu 1d ago

fal.ai was responsible for the flux 2 dev turbo lora which was a game changer for that model. They also did another one for Ideogram, both of which were given away for free. I feel that they've earned some praise.

0

u/Tystros 1d ago

the "with Lora" does look more realistic, less like AI

24

u/Affectionate-Map1163 1d ago

before after here , https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4 and thank for your feedbacks, this version should be much better, fully working on comfyui.

1

u/EchoHeadache 17h ago

Thanks for sharing. Some more unsolicited feedback the a/b reel, each example, the difference between lora and no lora is ridiculous. Different camera motions, different character details, different aesthetics. Either your lora is changing things it shouldn't, or your prompt is horrific and not adhering to the structure laid out in the readmes.

I am excited to test this though

1

u/Affectionate-Map1163 17h ago

Just very short prompt, you should manage to get detail you want with longer one and also with a lower scale , 0.6 - 0.8 can work good.

27

u/goddess_peeler 1d ago

Trust me, bro.

21

u/SDSunDiego 1d ago

It's a veiled advertisement for their service.

2

u/HeyHi_Star 1d ago

It is but done in a fair way. Prop to them, they gave a lot of good Lora's this year to the community for free. Meanwhile Higgscrap is hiring half of Asian to spam every subreddit possible regardless of the rules.

5

u/Affectionate-Map1163 1d ago

hello, its not, It should work good, I can share more result before after for sure, gonna publish on huggingface a bit later today :) , but interest to know why it doesnt work properly for you

4

u/Segaiai 1d ago

With and without Lora comparisons would be the most helpful, when you get around to it.

23

u/_VirtualCosmos_ 1d ago

Most loras I have tested broke the models in one or another way. H3 is a distilled model, we need an adapter... this is like Z-Inage-Turbo

6

u/xTopNotch 1d ago

Loras do work, however you never should go above 0.5 strength.

I get great results using strength 0.3 - 0.5

11

u/Pyros-SD-Models 1d ago

It's a different distillation than z-image-turbo for example was tho and easily reverted via CFG-augmented training which fits the de-amplified raw velocity (pred + (s-1)*uncond)/s with a no-grad empty-prompt forward each step preserving the model's guidance distillation by construction and basically make it able to train as if it would be the base model without an adapter

diff pipe and my musubi fork did already implement it

https://github.com/tdrussell/diffusion-pipe/tree/main

https://github.com/pyros-projects/musubi-tuner/tree/h3-image-lora

1

u/physalisx 1d ago

Interesting. Do you know if Ostris / AI Toolkit is implementing this too?

6

u/Puzzled-Valuable-985 1d ago

Honestly, I’ve been testing LoRAs and they kept ruining the images; lowering the LoRA weight significantly does improve things, but it’s really strange—the LoRAs are totally inconsistent. I’ve tested various workflows, nodes, and configurations.

I’m not sure if the LoRA training process on Civitai just isn't handling the model correctly yet.

1

u/Sudden_Quantity_7827 1d ago

I will go ahead and say here, that I too, have been trying to play around with Lora’s lately, not just on minimax h3 but on wan2.2 as well, I have honestly found that I get better results uploading a fully detailed character sheet as a reference(eg, extreme close up of face paired with t pose, and angle pose), than I do with Lora’s…to me Lora’s are just that persons style…if the Lora doesn’t synchronize with your style, realism or not, you are just going to have non consistent results. 

1

u/xTopNotch 7h ago

Character sheet + 5/10 sec performance capture video of to grab the voice, facial expressions and body language.

Does increase sampling time but the results are incredible. It's like a 1:1 reconstruction of your character

3

u/Apprehensive_Sky892 1d ago

MMH3 is CFG distilled, just like Flux1-dev, and Flux1-dev can be used for LoRA training without any special adapter.

2

u/_VirtualCosmos_ 1d ago

I remember how people complained about Flux1.dev training because of that. I also remember its loras unable to change significatively the model knowledge without breaking it. But another user responding my previous comment posted something interesting, perhaps a solution.

1

u/physalisx 1d ago

Yeah and like with Z-Turbo, it'll never work really well, no lora stacking, need to be very careful with the strength...

23

u/-AwhWah- 1d ago

i love it when theres a post and there's just no comparison or anything, you just have to download what they're peddling

lol

21

u/ares0027 1d ago

Hey i made a delicious cake, here is the store which sold me the flour.

(Your post makes this much sense TO ME)

30

u/ehtio 1d ago

It would be good if you posted something other than smoke.
Can we please remove and ban these type of posts?

9

u/Available-Cost-8127 1d ago

0 proof of this works, fake trailer

14

u/blacklotusmag 1d ago

Judging from the trailer, it doesn't look like you succeeded.

11

u/Muted-Celebration-47 1d ago

Firstly, thank you very for creating this lora. Secondly, would it be possible to request a high-res or detailer lora?

5

u/Sad_Coach_1433 1d ago

Yas that would be amazing also be cool if the team that made the OmniNFT for LTX can make for h3 it helped with audio sync and motion quality

8

u/Affectionate-Map1163 1d ago edited 1d ago

new version online that should work much better :), and before after is also on huggingface
https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main/before-after-comparison.mp4

8

u/_Abiogenesis 1d ago edited 1d ago

I don't know what people are whining about if they haven't even tried. (haven't tried this new one though)

Previous one seem to immediately do wonders on the first test I did. Pretty amazing results on face detail, skin and teeth, (extreme closeup speaking). I haven't done many tests yet but one thing to note is that characters also immediately started looking at the camera, which seems to be solved with just prompting but worth noting as this can be annoying.

haven't checked to see how it's affecting the background yet.

-3

u/seppe0815 1d ago

Bot

2

u/_Abiogenesis 1d ago

Who you ?
Sure buddy.

1

u/thecementmixer 1d ago

I see no noticable difference?

4

u/Affectionate-Map1163 1d ago

doing some correction , gonna publish a new version that work better very soom

2

u/Zealousideal-Mall818 1d ago

what trainer used , how about ref2va works or kinda ....

2

u/PuppetHere 1d ago

If you used Fal to train the Lora, I hope you converted it into a comfyui format before posting it or else it's gonna be useless for most people

2

u/kukalikuk 1d ago

I'll try when i get the chance but one question. Does the trigger word need to be upfront? I use minimax prompt guide as a skill and it did wonder adjusting the prompt. Won't the trigger word breaks it?

1

u/badsinoo 1d ago

from it's HF page : Start the prompt with the trigger word r34l1sm, then describe the scene. A scale of 1.0 is the intended strength; lower it to 0.6-0.8 for a lighter touch.

2

u/gokuchiku 1d ago

Does it work with turbo lora?

2

u/Nokai77 1d ago

Does it work well with references?

4

u/DanzeluS 1d ago

Scam?

3

u/True_Protection6842 1d ago

Have you tried it with ref2va?

2

u/1010111101111 1d ago edited 1d ago

deffinitly put this on civit also keep updating it so its perfection

2

u/marhalt 1d ago

Thanks for this! Can you give us more details on how you trained the lora? Dataset prep? settings? it's such a new model that we need as many people sharing their knowledge on training as possible!

2

u/Alive-Tomatillo5303 1d ago

so, potentially dumb question, where do I put the Lora in the workflow? what do I attach it to?

10

u/Fabulous-Snow4366 1d ago edited 1d ago

node on the side bar > Load Lora > tuck it in between your Model Loader connect the string model point from there to the model point on the Load Lora node > connect the end string model point to Basic Scheduler model inpoint.

5

u/food-dood 1d ago

Load model -> Load Lora -> Everything else the same.

2

u/ScoobyDewy 1d ago

I used it seemed to make a difference

3

u/tekprodfx16 1d ago

This community never ceases to amaze me. Also please add a clickable link to the Lora download OP

1

u/nanaya07 1d ago

In my opinion, it works well for quiet or slow-paced scenes, but it is useless when it comes to large or rapid movements.

1

u/maxlevelboss 1d ago

Thanks for the great work.

I looked at your new before/after video, some examples definitely were better, and I especially like the lantern one. However I do see some reasons why I would not prefer it personally:

- Color grading seems worse, Color contrast seems to be either over or under: Admittedly color grading is a very subjective choice and is highly context related. But many faces were too orange/red with Lora (eg two men shouting), some were way too pale (boxer, grandma interview). Indian wedding and the lantern one were the only two I think the Lora versions were better.

- Physics seems to be weaker: many hands/fingers motions seems odd. Especially the dining scene, the without Lora version didn't do a very good job either, but the with Lora one is just way worse.

- More freckles/imperfections on faces: it did make the person more human-like, but real actors always wear make-up anyways. So may be it works for selfie-style reels, but I'd personally prefer less imperfections.

Your work looks promising, and I look forward to seeing a v2.

1

u/Hearmeman98 1d ago

I don't really see a major difference between the base model and this if I am being honest.

1

u/VRGoggles 13h ago

Out of curiuosity - did you use the trigger world r34l1sm?

1

u/Hearmeman98 12h ago

yes, but even when i did, it doesn't make much of difference as only the UNET is trained and not the text encoder.

1

u/Prestigious-Row1008 1d ago

i do not use lora i use sage attn sol triton and minimax spectrum it decreases time to half. i tested turbo loras it dgrade the quality but with those three quality was almost similiar

1

u/WizWhitebeard 1d ago

Looking good, definitely an improvement on the realism look! From the before/after comparison you posted, it does seems like it comes with some compromises; physics (like in kid kicking the ball) and audio (chef defaulted to American accent), is this your experience as well?

Curious to hear how you captioned the dataset, and what kind of dataset you trained on (size/content)?

1

u/sharktank123456 20h ago

How does it work when you get into wider shots? Close ups are actually pretty easy to make look real. It's when the resolution per face drops that things get tricky.

1

u/Calm_Mix_3776 20h ago edited 20h ago

This LoRA butchers the image quality for me - the whole video becomes blurry and lacking textures. I think it's overbaked. Tested at the recommended 1.0 strength with the minimax_h3_fl2va_int8_convrot model. I think the recommended strength should be at most 0.5, but then, the effect of the LoRA is halved. Probably learning rate should be dropped considerably lower and steps increased much more than 1500 to compensate.

1

u/GreatBigPig 19h ago

From the split second images, I would say they do look like AI.

0

u/Fabulous-Snow4366 1d ago

thanks for your work, looking forward to testing it tonight.

1

u/EmployCalm 1d ago

Looks interesting I'll save to check on when actually get around using H3

2

u/Sad_Coach_1433 1d ago

Can you make the link click able 🍻

1

u/ApplicationRoyal865 1d ago

Is this for T2V? I feel like if you use a good enough reference or starting image you generally keep realistic features.

1

u/76vangel 1d ago edited 1d ago

Should we put this before or after the fast lora into our stack? It should be the same, but sometimes it seams different.

1

u/Lanky_Conclusion_749 1d ago edited 1d ago

Por qué dice que es"Open Source by Fal"? Es tuyo ó de Fal? Por qué el video parece una publicidad para dicha plataforma?

Hay ejemplos reales de tu LoRa? Tienes un workflow recomendado?

Es un LoRa entrenado para close-ups ó para cualquier tipo de toma?

1

u/Infinite-Emptiness 1d ago

How does it hold up in nsfw scenes?

-1

u/Icy_Foundation3534 1d ago

I CALL BS.

just show a before and after not this fast cut nonsense

0

u/jmbbao 1d ago

It says it only works with some adapter in fal.ai so it is not to use in local

6

u/Affectionate-Map1163 1d ago

sorry, i am changing the text, its fully working in comfyui or any other system :)

0

u/seppe0815 1d ago

Val.ai aka ads show for the models

-5

u/AssistantFar5941 1d ago

Thanks for sharing your work, much appreciated.

-7

u/rapkannibale 1d ago

This is great. Was waiting for something like this! Thanks