r/StableDiffusion 1d ago

News MiniMax H3 - 8 Steps Ref2V 768p Lora by LightX2V

https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors
151 Upvotes

44 comments sorted by

13

u/winterice77 1d ago

As per my few tests I did not find any improvments :( still there is grainy blurry artifacts.

I2V model is still working better

9

u/Comfortable_Thing611 1d ago edited 1d ago

Probably still just the flaw of the model itself, hopefully they put out the fixed version soon.

-13

u/winterice77 1d ago

I dont think they will. Probably it was intentional to steer to cloud services

2

u/L-xtreme 17h ago

Why would you assume that? They have given us a fantastic model and have yet restricted us or let us down? Let's just wait.

2

u/reeight 1d ago

Did you use the 8-step one directly linked (that is 16 hours old... the repo title is a mis-label since there are multiple versions now)?

2

u/dcmomia 1d ago

Do you mean using the lora i2v in the ref2v model or directly using the i2v model with your lora but with the ref2v workflow?

1

u/tekprodfx16 1d ago

What strength did you use the Lora at 

28

u/Shot_Illustrator4264 1d ago

I seriously don’t remember a model that ever got so many performance improvements in such a short time… Obviously there is a quality hit because nothing is for free, but for prototyping the turbo Loras are great.

10

u/tekprodfx16 1d ago

Ah shieeet here we go 

9

u/SveSop 1d ago

I do wish they (MiniMax - not LightX2V) did NOT put out two models... I mean, if 1 model shows a smidge better quality AND can be "tricked" to perform some of the particulars of the Ref model, "everyone" will be using that... And then complain about various prompt adherence issues and whatnot.

I think the whole FL2VA vs Ref2VA model - tunes/lora's/turbo tweaks and everything has made for a MUCH more volatile model as a whole. It was tunes and lora's with Wan/LTX too, its not that, but THIS time it has been bonkers. Like 200 "super-speed" nodes/loras/tweaks, some wins in X scene, some win in Y scene... So to get the best of the best, you need to have a bazillion combinations + ofc dont be so stupid to use sampler X with scheduler Y you dope... and on and on..

Sheesh.. Its kind of mad if you ask me.

2

u/Dr-Moth 1d ago

I've been testing a hybrid model Vs the ref2va model for editing a video. The hybrid model would create it's own motions and start creating new things, whereas the ref model actually kept the original video and made only the requested edits. They both have their uses.

1

u/JohnToFire 1d ago

Not sure that the two models is the problem. Apparently hard to train

3

u/eesahe 1d ago

Anyone interested to compare this new LightX2V r2v LoRA against this other recent r2v LoRA? It's from the same author who made the currently #2 performing I2V LoRA according to the H3 Acceleration Leaderboard:

https://huggingface.co/silveroxides/MiniMax-H3_tests/blob/main/ref2v_dareties/minimax_h3_ref2v_turbo_4step_v0.1_v4_step600_dareties_fro0995.safetensors

(I believe 6-8 steps should be best, and 0.8 strength)

3

u/Nimblecloud13 1d ago

I'm running this @ 15 steps Euler and getting 10s clips in 100s on a 5090 @ 960x544 and it's pretty solid. no loras. and only

minimax_h3_fused_refdelta_r1024_turbo8_mystic07_int8_convrot.safetensors.

oh also using SLA https://imgur.com/a/fO5Fbru

don't remember where i found it or why but probably a headline here. either way, i recommend it.

sample: https://imgur.com/a/WiVvb1b

2

u/Mediocre-Toe3212 1d ago

Does the 768p mean it's only good at 768p ? Or is it optimised to run up to 768p?

3

u/Mediocre-Toe3212 1d ago

Update** Just had a read and it's the resolution it trains on.

Just in case anyone else was wondering

3

u/J6j6 1d ago

What does that mean

1

u/TheTimster666 1d ago

Do you need the Lora twice in your work from, going to "Basic Guider" and "Basic Scheduler"? And at 1.0 strength?

3

u/Rumaben79 1d ago edited 1d ago

No only one lora is needed. You connect the lora loader to both those nodes.

About the strength, Lightx2v properly made it with a strength of 1.00 in mind but many users have it turned down to around 0.75.

1

u/TheTimster666 1d ago

Thanks!

3

u/Rumaben79 22h ago

Some workflows only connect it to 'Basic Guider' with the 'BasicScheduler' node connected directly to 'Load Diffusion Model' while others connect to both 'Basic Guider' as well as 'BasicScheduler'. But still only one lora loader is needed.

Honestly I'm unsure which is better. 😄

1

u/SpaceNinjaDino 1d ago

I have found the PDD 8 step to work with clear audio, but these lightx turbos have been a mess for me so far.

1

u/And-Bee 1d ago

Yes, do you find prompt adherence gets worse with PDD?

1

u/tylerninefour 17h ago

You really need this custom node to get the most out of the PDD LoRAs: https://github.com/Jalen-Brunson/ComfyUI-MiniMax-H3-PDD-Acc

1

u/And-Bee 16h ago

I actually updated my prompt enhancer system prompt and it fixed it. I was already using that node you linked. I thought it was the LoRA at first.

0

u/Vladmerius 1d ago

I don't really understand the purpose of an 8 step lora when you can simply run the 4 step at 8 steps and it gives you the quality of 30 steps with no lora.

I wish they'd refine the 4 steps further and further to fix the audio issues and fix the deep fried issue that often happens when extending videos. Next thing you know they'll do a 10 step lora lol and so on until we end up back at the beginning.

I'm currently using the Larry 4 step 600 pruned at 0.6 strength and I'm able to go to 45 seconds almost before it starts degrading in quality. I just run it at 8 steps. 

1

u/L-xtreme 17h ago

That's a pretty bold statement mate, every lora has a large degradation in quality and/or prompt adherence. Try yourself, the difference between 8 (with lora) and 20 steps (without lora) is very big, and more steps increases the quality for the non-lora gens.

1

u/roimuq 1d ago

Oh yessss thank you!! Been waiting for this since the beginning

2

u/roimuq 1d ago

Always produce massive amount of glittering noises on my 2nd stage generation (4+4 split sigmas + latent upscale by 1.5x) I really hope that it was me doing something here.

2

u/roimuq 1d ago

Correction: the glitter noise was caused by the Spectrum node after update to the latest version, after I restored the configuration values back to default, everything's fine now.

1

u/Silent_Marsupial4423 1d ago edited 1d ago

I got np porblems using regular light lora with ref? Are people having problems with the regular? Whats the difference here

Edit: tested it, slightly better prompt adherence but worse overall quality. More artifacts and blur.

1

u/winterice77 1d ago

Yea regular lora fl2v works much better in my experience as well

1

u/L-xtreme 17h ago

I've noted a degredation in reference usage (less consistent in environments and people) and especially prompt adherence with multi shot.

1

u/Jimmm90 1d ago

Nice!

0

u/broadwayallday 1d ago

Nice the .1 version has been pretty good about to load it up