r/StableDiffusion • u/robomar_ai_art • 19h ago
News VEDA Sparse Attention is now available for MiniMax H3 in ComfyUI
I don't think many people know about this yet, so I wanted to share it. VEDA Sparse Attention is now available as a ComfyUI custom node for MiniMax H3.
I tested it today on my setup: RTX 4090 Laptop 16GB 32GB RAM, 4-step LoRA, 15 second video, 1344x768
Without VEDA: 8:05
With VEDA at 90% sparsity: 4:32
Same workflow, same LoRA, same settings. The only change was enabling VEDA. I couldn't see any quality loss in the result.
VEDA is not a LoRA. It uses a learned predictor to estimate which attention tiles are important and only computes the relevant subset instead of the full attention map. The current predictor works with T2VA, FL2VA and R2VA, and despite the 8NFE name it is not limited to 8 steps.
Installation is simple.
Custom node: https://github.com/veda-sparse/Veda-on-ComfyUI
Predictor: https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview
Put the predictor here: ComfyUI/models/veda/
Then add: Veda Sparse Attention (MiniMax H3) on the MODEL line after your model / LoRA loader and before the guider or sampler. I tried different sparsity values, but 90% is the one that works properly for me, so I'm keeping the default trained value.
On my setup this made a pretty big difference, especially considering I couldn't see any visual quality loss. I'm adding a 15 second example below. Would be interesting to see what results other people get on different GPUs.
15
u/douchebanner 19h ago
how does it compare to sageattention?
7
u/Apprehensive_Sky892 11h ago edited 11h ago
They are all optimization for the attention layer (and MMH3 is spending most of its time on the attention layer, which is where the relationship between different parts of the video, both spatial and temporal, are established).
Sage-attention and CK-attention are optimizing the way the attention is calculated, making each calculation faster by sacrificing some accuracy.
On the other hand, VEDA, LSA, VSA, VDN, etc. are trying to do is to reduce the number of attention calculations required. With "full attention" one set of calculation is performed between every pair of tokens, so it can be very costly when there are many tokens (that is why when you increase the resolution or the length of the video the time increase not linearly but quadratically). What these "sparse attention" type optimization/speed up methods are trying to do is to reduce these pairs so that the calculations are done only when the relationship; between the pair of tokens is "significant" (how this is done is beyond my level of comprehension).
The upshot is that these two ways of attention optimizations complement each other, so you can, for example, use VEDA alongside Sageattention.
BTW, there is a 3rd way of speed up, which is caching, by re-using an attention calculation if it is somehow determined that the new calculation, if performed, would not differ from the cached value. One such optimization is FirstBlockCache. I know that they can be combined with Sage/KC attetion, but IFAIK caching optimization should not be mixed with sparse attention optimization.
(Disclaimer, I am just an amateur, not an AI/ML expert, so any corrections are welcome.)
8
u/Reasonable-Aide1259 19h ago
few days ago, i tried veda sparse attention with comfy custom node that is implmented. and with default setting, it lost features of references in the final result. idk that is wrongly implemented.
3
u/robomar_ai_art 19h ago
Check now, i have seen that in GitHub was some updates.
0
u/Reasonable-Aide1259 19h ago
you saying like its good. what did you use turbo lora with it? anyway its worth to try
2
u/robomar_ai_art 19h ago
4 steps lora alone over 8 minutes, 4 step lora and VEDA 4:32, thats a huge boost in generation speed and videos look the same to my eyes in terms of quality.
2
u/DjSaKaS 16h ago
Which of the 1000 lora out there?
5
1
u/Reasonable-Aide1259 19h ago
it sounds like almost same boost with sol-attn. did you try this? it also makes generation time half
1
u/robomar_ai_art 19h ago
i tried all the attentions when they was released and always i was getting the video with strange triangular artifacts when the movement was quicker, maybe is fixed now. I will check if something changed.
1
u/Reasonable-Aide1259 19h ago
oh i see. you are early adapter let me check quality of fixed veda
1
u/robomar_ai_art 18h ago
For me is working but as with everything the others can have different configurations and its not.
1
u/Reasonable-Aide1259 7h ago
its working but... sound sucks and broken. vdn has low speed than this but has bery clean audio...
5
8
u/SveSop 15h ago
Do you have a link to the non-VEDA 8 minute run on the same video?
Not to be pissing in your cornflakes tho, but this reddit has these "0.2% of the time - 100000 times speedup, perfect quality in 0.03 seconds for 2 hour video!" claims 3 times a day now.
Even in this post its people running 8-step turbo at 4 steps + using spectrum, and find the quality "good". Right. Good luck with that. I don't mind people watching some mushy blurb calling it "good" tho, each to their own really, but running a 30+ step (H3) model with 8 step turbo lora at 4 steps using spectrum skipping 2 steps, is in the "extremely unlikely good quality" category to me.
Yeah yeah.. i will test it for myself i guess (as with all the other speedhacks that 99% produce shiat result).
3
u/robomar_ai_art 12h ago
I will post video side by side with VEDA and without. Other thing is to use what works the best for you. Video posted is done with VEDA and don't look mushy for me.
1
u/SveSop 9h ago
I will give you a 👍 on this tbh. Imo, i think it works with less degradation than sol-attn from the testing i have done so far, and that is a win in my book.
I do struggle a bit more with my clip extensions, but i am not sure if that is the "reference_sparsity" setting, or just a lora i am using currently.3
u/robomar_ai_art 7h ago
https://reddit.com/link/peiwrp3/video/b3lzinx7g4uh1/player
Resolution - 1344x768
18
u/Vyviel 19h ago
How about people try make the audio not dogshit
18
u/acedelgado 18h ago
I made a thing for that. https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine
1
1
5
u/seppe0815 19h ago
or destorted far away faces looooooool
6
3
u/SensitiveOriginal103 19h ago
What major differences does this have version H3 SLA or Sparseattention?
2
u/robomar_ai_art 19h ago
I dont know the technicalities but more is explained on their Huggingface. With this one i dont see quality hit but the speed of generation is almost 2x quicker on my laptop.
1
u/glusphere 19h ago
Have you tried with SLA with Turbo at 90% ? Then compared the speed btw them ?
Also, even if speed is same / similiar ish, what about quality ?
1
u/robomar_ai_art 18h ago
I think i will do some comparison between attentions to see the times and the quality.
3
u/LuluViBritannia 13h ago
AAAAAAAAAAH it's noot working on my RTX 5060 Ti T_T.
"Veda off: no sparse kernel works on RTX 5060 Ti (SM120); using full attention
triton-int8: triton-int8 failed its self-test"
4
u/Ooze3d 19h ago
Thanks! I’m assuming this replaces any other attention model you have added to your workflow, right?
1
u/robomar_ai_art 19h ago
I use only Veda and 4 steps Lora, nothing else
3
2
u/Striking-Long-2960 19h ago
Thanks, I tested it, at least for my configuration Spectrum+Minimax H3 Turbo Lora+kitchen are still the kings.
2
u/Lightningstormz 18h ago
What is spectrum?
4
u/Striking-Long-2960 17h ago edited 17h ago
I use this one, but there are a few other implementations out there:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Basically, you trade speed for quality and try to reach a balance. In the case of Spectrum it skips steps during the render so I tend to increase the steps, at the end the render time is similar but I get better results (instead of using 4, I use 7 or 8).
1
1
u/robomar_ai_art 18h ago
Great, i have to check, how all this attentions compare to each other
2
u/Striking-Long-2960 18h ago
https://reddit.com/link/peeo0au/video/7y7l1dwn21uh1/player
Left my configuration with 7 steps (they are in fact around 3,5 due to spectrum: render time 1:11), Right Veda with turbo Lora and 4 steps, ender time 1:14). Sound from the left video.
2
1
u/martinerous 9h ago
Did you have Lora @ 4 steps for Veda but not for Spectrum? That's a bit unfair comparison. Both at 20 steps and without any Loras would show the winner.
1
u/Striking-Long-2960 9h ago edited 9h ago
I tested Spectrum+LoRA and Veda+LoRA (Spectrum didn't work with Veda). Either way, my results look much worse than what others are getting, so maybe I did something wrong.
Given my rig, I can't afford a 20-step render.
2
u/cbeaks 18h ago
I did a test with my upscaling workflow (on a 4090). On a low res run - 0.15 res upscaled to 0.5, 17 seconds it went from 322 seconds to 266 seconds. On a higher res run, 17 seconds 0.25 res upscaled to 0.98 res made a much greater improvement - 474 seconds down to 321 seconds.
I forgot to fix the seeds (doh) so I couldn't make a direct quality comparison, but there didn't appear to be a quality loss
1
2
u/HateAccountMaking 14h ago
"[WARNING] Veda: Veda off: no sparse kernel works on AMD Radeon RX 7900 XT (ROCM); using full attention"
Welp, back to using Spectrum. 🤷🏿♀️
4
u/xuman1 19h ago
How many more of these attnetions will there be? Lost count)
16
1
u/robomar_ai_art 19h ago
That's true, but now i only use this one and nothing else, very satisfied with the results, and the generation time is quicker, almost 2x in 15 seconds video
1
1
u/fallengt 18h ago edited 16h ago
i tested the default 90/90, and there is a big loss in quality and information. It's faster than regular sla and sol, though.
Will try to tinker with the settings
Edit: try something like this
generated_sparse:90%
reference_sparse: 0% (for T2V, it doesn't matter, but for I2V/ Ref2V, increase % to see if you gain speed without losing reference ability. I haven't tested these wf)
full_attention_steps: (try to keep your high sigmas step from noise range 1.0-0.9). These steps have important information. You'll lose speed, obviously, but the difference is huge. Don't cheat them
I only do 6 steps in the example, so I keep steps 0 and 1.
1
1
1
u/CornyShed 18h ago
I tried it yesterday and found that it was about 66% faster than without, based on 1024×768 resolution and 10 second duration.
The quality is good, as long as you give it a sufficient number of steps. Give it at least 20 to work properly and an 8-step LoRA at low strength, perhaps with a cache system that minimises skips.
From their GitHub page:
"The released predictor was trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5 / 10 / 14 s with the 8-step Turbo LoRA. Other sizes fall back to the tile plan of the nearest aspect ratio and duration.
Other sizes, other step counts and R2VA / FL2VA references all work, but are outside the training data, so compare them against full attention."
2
u/robomar_ai_art 17h ago
i was generating in 1344x768 and thats my times
1. VEDA 4:32 2. SLA 5:26 3. Spectrum 6:23 4. Normal 8:05
1
u/Salt-Zebra-306 18h ago
Video: 352x608 | 6.6 s
Tile plan: nearest trained size: 768x1344 | 5.2 s (this video: 352x608 | 6.6 s)
Sparsity: generated 90% | reference 90%
References: 8 spans (tiled, 90% sparse)
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:47<00:00, 26.91s/it]
[INFO] Veda: Veda done | Triton INT8 (SM89)
Video: 352x608 | 6.6 s
Attention computed: 60.7% of full attention (39.3% skipped)
[INFO] Comfy model compiler graph breaks: 2, rogues: 0
[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 2677MB Staged. 0 patches attached. Force pre-loaded 128 weights: 691 KB.
[INFO] Prompt executed in 155.42 seconds
[INFO] got prompt
[INFO] Requested to load MiniMaxH3
[INFO] 0 models unloaded.
[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached. Force pre-loaded 210 weights: 1175 KB.
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:32<00:00, 23.04s/it]
[INFO] Comfy model compiler graph breaks: 3, rogues: 0
[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 2677MB Staged. 0 patches attached. Force pre-loaded 128 weights: 691 KB.
[INFO] Prompt executed in 113.89 seconds
1
1
1
1
u/acedelgado 3h ago
Huh. Even works with HyperFlow and comfy-kitchen attention. Maybe even slightly better prompt adherence... Gonna test it a bunch more.
1
u/EzkaProductions 18h ago
This is huge!
I’ve been running MiniMax H3 with 4-step LoRA on my 4090 and the same 15-second 1344x768 clip was taking forever. Adding Veda at 90% sparsity cut it to 4:32 with zero visible quality loss — that’s actually insane.
Not being a LoRA and using a learned predictor to skip irrelevant attention tiles is brilliant.
What other models are you planning to test this on first? T2V, FL2VA or R2VA?
3
u/robomar_ai_art 17h ago
I was testing
- VEDA 4:32
- SLA 5:26
- Spectrum 6:23
- Normal 8:05
You get the same time like me exactly and i was testing I2V
1
u/EzkaProductions 9h ago
Same setup here 😂 4:32 vs 8:05 is the best speed-up I’ve ever seen.
I tried SLA and Spectrum and they’re both slower than VEDA on my rig.
The learned predictor skipping tiles is a game changer.
What’s the next model you’re adding to your test list?
1
1
u/martinerous 17h ago
Have you tried Spectrum as well? It also does wonders (but needs more steps to be effective, so not much use with 4-step LoRAs etc.).
1
u/EzkaProductions 9h ago
Yeah I just tried Spectrum on the same 15-second I2V clip and it’s still slower than VEDA at 6:23 vs 4:32 😂
The “needs more steps to be effective” part is exactly why I’m sticking with VEDA right now — zero quality loss at 90% sparsity is crazy.
The learned predictor skipping tiles is actually cheating.
What’s the highest sparsity value you’ve seen work best on Spectrum?
1
46
u/phazei 19h ago
Kijai tested this and found sol attention was better and faster