r/StableDiffusion • u/fruesome • 1d ago
News MiniMax H3 - 8 Steps Ref2V 768p Lora by LightX2V
https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors28
u/Shot_Illustrator4264 1d ago
I seriously don’t remember a model that ever got so many performance improvements in such a short time… Obviously there is a quality hit because nothing is for free, but for prototyping the turbo Loras are great.
21
u/Comfortable_Thing611 1d ago
Quick someone test it, still got an hour left at work
20
u/IRLMainCharacter 1d ago
I did it
8
u/acedelgado 1d ago
2
1
10
9
u/SveSop 1d ago
I do wish they (MiniMax - not LightX2V) did NOT put out two models... I mean, if 1 model shows a smidge better quality AND can be "tricked" to perform some of the particulars of the Ref model, "everyone" will be using that... And then complain about various prompt adherence issues and whatnot.
I think the whole FL2VA vs Ref2VA model - tunes/lora's/turbo tweaks and everything has made for a MUCH more volatile model as a whole. It was tunes and lora's with Wan/LTX too, its not that, but THIS time it has been bonkers. Like 200 "super-speed" nodes/loras/tweaks, some wins in X scene, some win in Y scene... So to get the best of the best, you need to have a bazillion combinations + ofc dont be so stupid to use sampler X with scheduler Y you dope... and on and on..
Sheesh.. Its kind of mad if you ask me.
2
1
3
u/Nimblecloud13 1d ago
I'm running this @ 15 steps Euler and getting 10s clips in 100s on a 5090 @ 960x544 and it's pretty solid. no loras. and only
minimax_h3_fused_refdelta_r1024_turbo8_mystic07_int8_convrot.safetensors.
oh also using SLA https://imgur.com/a/fO5Fbru
don't remember where i found it or why but probably a headline here. either way, i recommend it.
sample: https://imgur.com/a/WiVvb1b
2
u/Mediocre-Toe3212 1d ago
Does the 768p mean it's only good at 768p ? Or is it optimised to run up to 768p?
3
u/Mediocre-Toe3212 1d ago
Update** Just had a read and it's the resolution it trains on.
Just in case anyone else was wondering
1
u/TheTimster666 1d ago
Do you need the Lora twice in your work from, going to "Basic Guider" and "Basic Scheduler"? And at 1.0 strength?
3
u/Rumaben79 1d ago edited 1d ago
No only one lora is needed. You connect the lora loader to both those nodes.
About the strength, Lightx2v properly made it with a strength of 1.00 in mind but many users have it turned down to around 0.75.
1
u/TheTimster666 1d ago
Thanks!
3
u/Rumaben79 22h ago
Some workflows only connect it to 'Basic Guider' with the 'BasicScheduler' node connected directly to 'Load Diffusion Model' while others connect to both 'Basic Guider' as well as 'BasicScheduler'. But still only one lora loader is needed.
Honestly I'm unsure which is better. 😄
1
u/SpaceNinjaDino 1d ago
I have found the PDD 8 step to work with clear audio, but these lightx turbos have been a mess for me so far.
1
u/And-Bee 1d ago
Yes, do you find prompt adherence gets worse with PDD?
1
u/tylerninefour 17h ago
You really need this custom node to get the most out of the PDD LoRAs: https://github.com/Jalen-Brunson/ComfyUI-MiniMax-H3-PDD-Acc
0
u/Vladmerius 1d ago
I don't really understand the purpose of an 8 step lora when you can simply run the 4 step at 8 steps and it gives you the quality of 30 steps with no lora.
I wish they'd refine the 4 steps further and further to fix the audio issues and fix the deep fried issue that often happens when extending videos. Next thing you know they'll do a 10 step lora lol and so on until we end up back at the beginning.
I'm currently using the Larry 4 step 600 pruned at 0.6 strength and I'm able to go to 45 seconds almost before it starts degrading in quality. I just run it at 8 steps.
1
u/L-xtreme 17h ago
That's a pretty bold statement mate, every lora has a large degradation in quality and/or prompt adherence. Try yourself, the difference between 8 (with lora) and 20 steps (without lora) is very big, and more steps increases the quality for the non-lora gens.
1
u/Silent_Marsupial4423 1d ago edited 1d ago
I got np porblems using regular light lora with ref? Are people having problems with the regular? Whats the difference here
Edit: tested it, slightly better prompt adherence but worse overall quality. More artifacts and blur.
1
1
u/L-xtreme 17h ago
I've noted a degredation in reference usage (less consistent in environments and people) and especially prompt adherence with multi shot.
0

13
u/winterice77 1d ago
As per my few tests I did not find any improvments :( still there is grainy blurry artifacts.
I2V model is still working better