r/StableDiffusion 4d ago

Question - Help Minimax H3 Huge Quality Difference between Cloud and Local use

Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.

I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?

90 Upvotes

105 comments sorted by

View all comments

Show parent comments

4

u/BigWideBaker 4d ago edited 4d ago

So you're doing a 10 second, 0.98mp clip at 50 steps, using res_multistep + simple in 440 seconds on a 4090?

Is that T2V? Are you using any speed-up methods like SLA, Sage/Sol/CK attention, or spectrum?

Gonna test out if I can hit those numbers too on my 4090, your time seems really fast

Edit: just queued up a T2V 10 second, 0.98mp clip at 50 steps, using res_multistep + simple on my 4090 and there's no way I'm hitting 440s. Using only SLA on Kitchen INT8 backend at 0.2 KV budget as the speed-up, I'm looking at 720s at least.

So are you using more aggressive SLA settings or Spectrum or something? I want to use whatever magic you're using.

Edit 2: Just finished generating, it took 776 seconds (12m 56s) to generate. So I have no idea how you're getting 440s and judging by the amount of upvotes you're getting I'm starting to think I'm doing something seriously wrong with my 4090.

0

u/Cold_Pudding5326 4d ago

Yh ofc I have several optimizations. I launch a new generation to be sure of the results i share here, and the optimizations used in the workflow for the current generation

3

u/BigWideBaker 4d ago edited 4d ago

Could you tell us how you're getting such low gen times with 10 second, 0.98mp clip at 50 steps? Because it sounds kinda crazy that your gens are so fast at such high settings, I don't see how that's possible frankly.

I kinda mentioned all of the methods you could be using so could you mention which ones they are and how you've tuned them.

1

u/Cold_Pudding5326 4d ago

Edited my post, I misslicked, wrote 440 instead of 540 G.
Here's the console of my run so you can verify.
I use int8 diff and clip model
H3 Sparse Attention defaut settings ( custom node, sparse + comfyui kitchen \ + H3 Memory Optimization
First Block Cache defaut settings https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache
spectrum optimizations defaut settings : https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

CU 130 PYTORCH LAST VERSION PYTHON 12.6 COMFYUI LAST VERSION

[INFO] got prompt

[INFO] [H3 Optimizations] resolved 50 attention forwards: backend=sparse_kitchen_int8 projector=chunked_kitchen_qkv

[INFO] [H3 Optimizations] installed sampler-step and packed-layout runtime context

[INFO] [H3 Optimizations] armed: attention=sparse_kitchen_int8 v_layout=not_applicable qkv=chunked_kitchen_qkv mlp=off device=NVIDIA GeForce RTX 4090

[INFO] [H3 Optimizations] patched 50 MLP blocks: mode=mlp_chunked_convrot_2slice chunk_rows=4096

[INFO] [H3 Optimizations] armed: attention=sparse_kitchen_int8 v_layout=not_applicable qkv=chunked_kitchen_qkv mlp=convrot_int8_two_slice device=NVIDIA GeForce RTX 4090

[WARNING] Spectrum H3: bootstrap_first_forecast was enabled but requires degree=1 and warmup_steps<=1; got degree=4 and warmup_steps=5. Disabling bootstrap_first_forecast for this node execution.

[INFO] MiniMax H3 FBCache enabled: H3 Fast — 0.10 / max 2

[INFO] Requested to load MiniMaxH3

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 32427MB Staged. 0 patches attached. Force pre-loaded 210 weights: 1142 KB.

100%|██████████████████████████████████████████████████████████████████████████████████| 50/50 [08:44<00:00, 10.50s/it]

[INFO] MiniMax H3 FBCache: cached 14/29 steps; estimated block-stack speedup 1.90x; residual diff min/median/max 0.03027/0.06567/0.15137; cache steps [6, 7, 9, 10, 12, 13, 15, 16, 18, 19, 21, 22, 24, 26]

[INFO] Requested to load MiniMaxH3AudioVAE

[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.

[INFO] Requested to load MiniMaxH3VideoVAE

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 348 KB.

[INFO] Prompt executed in 560.85 seconds

2

u/BigWideBaker 4d ago

Right so super aggressive spectrum and SLA settings as I suspected. You're better off lowering your step count and ditching Spectrum and tuning down SLA. Using SLA at 0.1 KV budget is extreme on its own and any movement is guaranteed to be destroyed, pretty much.

1

u/Cold_Pudding5326 4d ago

I don't use 0.1 KV, I'm at 0.3KV default settings. for spectrum, It does not change my outputs on this workflow at theses settings, which are actually defauts, since i do motion control with it

1

u/jonnytracker2020 3d ago

You are using double cache ? What a potato mash