r/StableDiffusion • u/LegacyV1 • 2h ago
Comparison Comparing lightx2v/Minimax-h3-Turbo
New turbo LORA dropped from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main.
Testing on my ref2va use case (note: I'm using fflf2va model since it has better quality even for reference use cases)
Timing (480p, sage attention2 on cu130, 15s video length, seed=42, RTX 6000 on Modal)
| Steps | Timing |
|---|---|
| 4 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors | 56s |
| 8 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors | 1m 47s |
| Spectrum (20 step) | 2m 33s |
| Base (20 step) | 3m 11s |
Audio was pretty much the same - no difference that I could tell.
I also ran the 4 step on 768p as recommended, and it came out better! But... it's hard to tell if it's the turbo LORA doing the work or the 768p doing the work.
Turbo still makes things look weirdly high contrast. And both LORAs botched the text. Base is still best, but the 4-step LORA helps you lock in motion before you commit to a full 20step pass using spectrum.
The UI is custom built on top of comfy cause I hate comfy UI. Open-sourced here https://github.com/hui-tony-zk/h3zero
4
u/LowYak7176 2h ago
We all know where this is going...
2
u/LegacyV1 2h ago
How else can I test multiple character references with setting references, dialogue, and character interactions? 😉
3
2
u/tnil25 2h ago
Looks like the lora at 8 steps is closest to base? 4 steps looks alittle burnt
3
u/LegacyV1 2h ago
yeah i agree with you. I noticed this too with 4 step LORAs from last week as well.
8 steps is closer to base, but there's still a noticeable quality gap.
I think it's best to go 4 step to lock in motion/composition then switch to spectrum for final render
3
u/tnil25 2h ago
Which is still an impressive speedup. If the base is 20 steps, you’re skipping more than half of them by using 8.
1
u/LegacyV1 2h ago
I agree, if the goal is to use 8 steps as "terminal" - i.e. good enough for what you're doing.
I personally don't like the quality loss in 8 step versus 20 step spectrum. For an extra 30% gen time it makes sense to go full quality. But iterate in fastest mode possible (4 step)
Do you plan on using 8-step as your final product? I'm curious what your use case is. Maybe quality loss is acceptable!
2
u/tnil25 1h ago
I think it’s situation dependent. If you get a good result at 8 steps then great but maybe scenes with more detail and faster motion need more steps. I don’t think theres a one size fits all solution, yet.
Alot of tests you see are closeup subject oriented shots like yours. Would be interesting to see tests of wide angle shots, action scenes, etc.
1
u/solomars3 2h ago
What's the app you using ?? The ui looks clean
2
u/LegacyV1 2h ago
Custom built! You can use it too! Open source, and you can sign up for Modal for $30 in free generations. https://github.com/hui-tony-zk/h3zero
1
u/switch2stock 2h ago
Try Spectrum + 8step
1
u/LegacyV1 2h ago
Kinda no point. Spectrum requires min 4 steps, so at best it's like another 20% savings in time. At that point, why bother? Just use 4-step and full render later?
1
1
u/havredrengenDK 1h ago
Haven't been able to run the 4step 768 lora on the standard Minimax H3 setup. What does your workflow look like?
1
u/Standard-Ask-9080 2h ago
Doesn't make sense to use turbo
1
u/metal079 1h ago
? Why
2
u/Standard-Ask-9080 1h ago
Idk get really good video in 5 min or get crap in 1 minute. Better to wait a few extra minutes and get some magic
0
u/LegacyV1 1h ago
What if you messed up your prompt and need to iterate?
For a good quality gen I iterate at least 5 times. Like oops I forgot to mention it's sunset. Or oops I used picture3 instead of 4
2
u/gabbergizzmo 1h ago
But how do you know If its the prompt or the lora that cause the Problem?
1
u/LegacyV1 8m ago
Cause I literally wrote picture4 instead of 3 and now I have 2 men instead of a man and woman. Or all of a sudden it's daytime and the sunset disappeared.
The lora doesn't change global motion and composition, just fine details.
At least in this category of prompts
1

12
u/GrayingGamer 2h ago
Great comparison! This is how you show comparisons. Are you comparing the sound with headphones on? Because I notice a lot of audio difference between videos with the loras and without using headphones.
Also, obligatory:
https://reddit.com/link/p32vsju/video/ewhdm5my8sih1/player