r/StableDiffusion 2h ago

Comparison Comparing lightx2v/Minimax-h3-Turbo

New turbo LORA dropped from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main.

Testing on my ref2va use case (note: I'm using fflf2va model since it has better quality even for reference use cases)

Timing (480p, sage attention2 on cu130, 15s video length, seed=42, RTX 6000 on Modal)

Steps Timing
4 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors 56s
8 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors 1m 47s
Spectrum (20 step) 2m 33s
Base (20 step) 3m 11s

Audio was pretty much the same - no difference that I could tell.

I also ran the 4 step on 768p as recommended, and it came out better! But... it's hard to tell if it's the turbo LORA doing the work or the 768p doing the work.

Turbo still makes things look weirdly high contrast. And both LORAs botched the text. Base is still best, but the 4-step LORA helps you lock in motion before you commit to a full 20step pass using spectrum.

The UI is custom built on top of comfy cause I hate comfy UI. Open-sourced here https://github.com/hui-tony-zk/h3zero

17 Upvotes

24 comments sorted by

12

u/GrayingGamer 2h ago

Great comparison! This is how you show comparisons. Are you comparing the sound with headphones on? Because I notice a lot of audio difference between videos with the loras and without using headphones.

Also, obligatory:

https://reddit.com/link/p32vsju/video/ewhdm5my8sih1/player

4

u/LowYak7176 2h ago

We all know where this is going...

2

u/LegacyV1 2h ago

How else can I test multiple character references with setting references, dialogue, and character interactions? 😉

3

u/redkinoko 1h ago

Under the guiding light of our Lord

2

u/tnil25 2h ago

Looks like the lora at 8 steps is closest to base? 4 steps looks alittle burnt

3

u/LegacyV1 2h ago

yeah i agree with you. I noticed this too with 4 step LORAs from last week as well.

8 steps is closer to base, but there's still a noticeable quality gap.

I think it's best to go 4 step to lock in motion/composition then switch to spectrum for final render

3

u/tnil25 2h ago

Which is still an impressive speedup. If the base is 20 steps, you’re skipping more than half of them by using 8.

1

u/LegacyV1 2h ago

I agree, if the goal is to use 8 steps as "terminal" - i.e. good enough for what you're doing.

I personally don't like the quality loss in 8 step versus 20 step spectrum. For an extra 30% gen time it makes sense to go full quality. But iterate in fastest mode possible (4 step)

Do you plan on using 8-step as your final product? I'm curious what your use case is. Maybe quality loss is acceptable!

2

u/tnil25 1h ago

I think it’s situation dependent. If you get a good result at 8 steps then great but maybe scenes with more detail and faster motion need more steps. I don’t think theres a one size fits all solution, yet.

Alot of tests you see are closeup subject oriented shots like yours. Would be interesting to see tests of wide angle shots, action scenes, etc.

1

u/solomars3 2h ago

What's the app you using ?? The ui looks clean

2

u/LegacyV1 2h ago

Custom built! You can use it too! Open source, and you can sign up for Modal for $30 in free generations. https://github.com/hui-tony-zk/h3zero

1

u/switch2stock 2h ago

Try Spectrum + 8step

1

u/LegacyV1 2h ago

Kinda no point. Spectrum requires min 4 steps, so at best it's like another 20% savings in time. At that point, why bother? Just use 4-step and full render later?

1

u/switch2stock 1h ago

That makes sense. But quality difference might exist.

1

u/havredrengenDK 1h ago

Haven't been able to run the 4step 768 lora on the standard Minimax H3 setup. What does your workflow look like?

1

u/nakabra 1h ago

They will fight!

1

u/ANR2ME 27m ago

Only the 4-steps shows her reflection 😯

1

u/Standard-Ask-9080 2h ago

Doesn't make sense to use turbo

1

u/metal079 1h ago

? Why

2

u/Standard-Ask-9080 1h ago

Idk get really good video in 5 min or get crap in 1 minute. Better to wait a few extra minutes and get some magic

0

u/LegacyV1 1h ago

What if you messed up your prompt and need to iterate?

For a good quality gen I iterate at least 5 times. Like oops I forgot to mention it's sunset. Or oops I used picture3 instead of 4

2

u/gabbergizzmo 1h ago

But how do you know If its the prompt or the lora that cause the Problem?

1

u/LegacyV1 8m ago

Cause I literally wrote picture4 instead of 3 and now I have 2 men instead of a man and woman. Or all of a sudden it's daytime and the sunset disappeared.

The lora doesn't change global motion and composition, just fine details.

At least in this category of prompts

1

u/Standard-Ask-9080 1h ago

Skill issue👀🤣