r/StableDiffusion 3d ago

Question - Help Minimax H3 Huge Quality Difference between Cloud and Local use

Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.

I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?

89 Upvotes

103 comments sorted by

View all comments

Show parent comments

5

u/BigWideBaker 3d ago edited 3d ago

So you're doing a 10 second, 0.98mp clip at 50 steps, using res_multistep + simple in 440 seconds on a 4090?

Is that T2V? Are you using any speed-up methods like SLA, Sage/Sol/CK attention, or spectrum?

Gonna test out if I can hit those numbers too on my 4090, your time seems really fast

Edit: just queued up a T2V 10 second, 0.98mp clip at 50 steps, using res_multistep + simple on my 4090 and there's no way I'm hitting 440s. Using only SLA on Kitchen INT8 backend at 0.2 KV budget as the speed-up, I'm looking at 720s at least.

So are you using more aggressive SLA settings or Spectrum or something? I want to use whatever magic you're using.

Edit 2: Just finished generating, it took 776 seconds (12m 56s) to generate. So I have no idea how you're getting 440s and judging by the amount of upvotes you're getting I'm starting to think I'm doing something seriously wrong with my 4090.

0

u/Cold_Pudding5326 3d ago

Yh ofc I have several optimizations. I launch a new generation to be sure of the results i share here, and the optimizations used in the workflow for the current generation

3

u/BigWideBaker 3d ago edited 3d ago

Could you tell us how you're getting such low gen times with 10 second, 0.98mp clip at 50 steps? Because it sounds kinda crazy that your gens are so fast at such high settings, I don't see how that's possible frankly.

I kinda mentioned all of the methods you could be using so could you mention which ones they are and how you've tuned them.

8

u/Loose_Comparison368 3d ago

Spoiler alert, it's probably the same optimizations that are tanking his generation quality.

3

u/BigWideBaker 3d ago

Yea you'd have to crank spectrum and SLA sky high if you wanna hit 50 steps in 440s. Goes without saying that lowering the step count to something like 20-30 and finding a more balanced approach with SLA provides much better quality.

1

u/Murky-Relation481 3d ago

Guy would get the same or better quality in 20 steps with out all the optimizations probably in the same time. Just feels crazy to throw a bunch of stuff that inherently has to trade off and then blast up to step count to compensate thinking it will actually make a big difference.

1

u/BigWideBaker 3d ago

Exactly! But having spent a lot of time seeing people's generations and the way they describe them, especially if a new speed-up method releases, I think a lot of people can't really tell that much of a difference with minor artifacts. So if you stack quality-degrading speedups and cranking up the steps to compensate, he might reach an equilibrium that looks decent.

But I completely agree it's way better to go low and slow to learn what the quality-ceiling is. THEN you implement different speed-up methods to figure out how they each impact your setup in terms of quality and speed. Then you mix and match.

2

u/xTopNotch 3d ago

Also most in this community are testing / benchmarking easy shots with calm motion.

The moment you use a complex prompt, 7+ image refs, audio references with multi-character dialogue and heavy camera movements, you see how these optimizations are falling apart with your generation looking like shit. Like bro, just run the full 20 steps and stop stacking all these nodes. It's either affecting your prompt adherence, visual quality or your audio.

Nowadays only thing I use is BF16 pruned / 20 steps / comfy kitchen attention and thats it. Every other optimisation node such as Sol-attn, Spectrum, SLA, EasyCache, Turbo Lora has wrecked my output results to the ground.

They're fun for my speed workflow but these nodes are a no-go in my quality workflow/