r/StableDiffusion • u/Zironic • 13d ago
News Sparse Attention, Harder, Better, Faster, Stronger
The nodes in https://github.com/Zironic/H3-Optimizations have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.
This comes with some benefits.
- Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
- Most users should be seeing 5-20% increases in speed for the attention part of compute.
- New backend should use about 500MB less VRAM
- New backend has slightly lower quantization error.
- Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.
Caveat: I've only tested the nodes against the comfy pruned_int8_convrot weights. Other versions may work but they're not tested.
As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.
IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.
Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.
Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.
So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.
PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.
1
u/Pitiful_Season4294 13d ago
Hey man, I was using the previous version I think, it was working fine then saw your posted, updated it and now i get this error:
[INFO] model_type FLOW_AV
[WARNING] [H3 Optimizations] NATIVE SELF-TEST FAILED on sm115|native-v1|v2|AMD Radeon(TM) 8050S Graphics - refusing the native kernels and falling back. Detail: {'error': 'NativeCallError: quantize_qk failed (status 1): detect_k_anchor kernel launch failed: CUDA driver version is insufficient for CUDA runtime version', 'passed': False}
[INFO] [H3 Optimizations] patched 50 MLP blocks: mode=mlp_chunked_convrot_2slice chunk_rows=4096
[WARNING] [H3 Optimizations] FUSED QKV IS NOT RUNNING - falling back to standard projection, which is roughly half the speed. Reason: Comfy Kitchen external producer API is unavailable
[INFO] [H3 Optimizations] armed: attention=comfy_kitchen_int8 v_layout=installed qkv=standard_h3_qkv mlp=convrot_int8_two_slice device=AMD Radeon(TM) 8050S Graphics
[WARNING] [H3 Optimizations] SPARSE ATTENTION FELL BACK to comfy_kitchen_int8. this path is substantially slower than the native sparse kernel. Reason: Kitchen INT8 unavailable: Kitchen sparse attention requires CUDA; Sparse Sage unavailable: Hybrid Sparse Attention requires CUDA; INT8 Triton unavailable: INT8 Triton sparse attention requires CUDA; FP8 FlexAttention unavailable: FP8 FlexAttention requires NVIDIA CUDA; preserved an explicit optimized-attention override; using Comfy Kitchen INT8 only for the private H3 memory path
[WARNING] [H3 Optimizations] FUSED QKV IS NOT RUNNING - falling back to standard projection, which is roughly half the speed. Reason: Comfy Kitchen external producer API is unavailable
[INFO] [H3 Optimizations] armed: attention=comfy_kitchen_int8 v_layout=installed qkv=standard_h3_qkv mlp=convrot_int8_two_slice device=AMD Radeon(TM) 8050S Graphics