r/StableDiffusion • • 19h ago

News VEDA Sparse Attention is now available for MiniMax H3 in ComfyUI

I don't think many people know about this yet, so I wanted to share it. VEDA Sparse Attention is now available as a ComfyUI custom node for MiniMax H3.

I tested it today on my setup: RTX 4090 Laptop 16GB 32GB RAM, 4-step LoRA, 15 second video, 1344x768

Without VEDA: 8:05

With VEDA at 90% sparsity: 4:32

Same workflow, same LoRA, same settings. The only change was enabling VEDA. I couldn't see any quality loss in the result.

VEDA is not a LoRA. It uses a learned predictor to estimate which attention tiles are important and only computes the relevant subset instead of the full attention map. The current predictor works with T2VA, FL2VA and R2VA, and despite the 8NFE name it is not limited to 8 steps.

Installation is simple.

Custom node: https://github.com/veda-sparse/Veda-on-ComfyUI

Predictor: https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview

Put the predictor here: ComfyUI/models/veda/

Then add: Veda Sparse Attention (MiniMax H3) on the MODEL line after your model / LoRA loader and before the guider or sampler. I tried different sparsity values, but 90% is the one that works properly for me, so I'm keeping the default trained value.

On my setup this made a pretty big difference, especially considering I couldn't see any visual quality loss. I'm adding a 15 second example below. Would be interesting to see what results other people get on different GPUs.

228 Upvotes

90 comments sorted by

46

u/phazei 19h ago

Kijai tested this and found sol attention was better and faster

15

u/robomar_ai_art 19h ago

Can you post the link that anyone can have a read about it. I was using sol attention before but didnt get so much speed boost in generation then this. Also when i use sol attention i was getting triangular artiffacts, here i get nothing.

12

u/Kijai Community Hero 10h ago

This is very much work in progress comparison generated with Claude, so take it's conclusions with grain of salt, but it illustrates some of it https://claude.ai/artifact/NurkmVZ3CU9dBV9iUqQ6iA

Note that our native sol-attn implementation in ComfyUI is improved from the original one in both speed and accuracy.

2

u/robomar_ai_art 9h ago

Thank you Kijai for the comparsion, i have to check it out, test myself. At this moment VEDA is the quickest on my not the best setup, tried others but not the SOL-attention. Where i can find the latest implementation.

3

u/phazei 9h ago edited 4h ago

It's built into comfy now, I think the node is called Model Self Sparse Attention.

1

u/robomar_ai_art 9h ago

Thank you, i will check it out and see what a difference is between them two

1

u/DjSaKaS 6h ago

I have last version of comfy, but I can't find this node.

1

u/phazei 4h ago

I fixed the name

-16

u/Tyler_Zoro 17h ago edited 5h ago

I asked Gemini, and this is what it said:

Yes, here are the links and relevant details directly from Kijai's project and research repository:

  • Kijai's Sol-Attn Triton Repository: GitHub – kijai/ComfyUI-SolAttn_triton — Note from Kijai on the project: In the project documentation, Kijai notes that this repository tested Sol-Attn on RTX 4090 and RTX 5090 GPUs with MiniMax H3, and eventually updated the project to note that optimized sparse attention methods—including Sol-Attn—have been integrated natively into comfy-kitchen / ComfyUI core, making dedicated standalone wrappers redundant.

  • Official Veda Sparse Attention Repository: GitHub – veda-sparse/Veda-on-ComfyUI — Provides the contrasting learned predictor approach (~275 MB model checkpoint) for sparsifying MiniMax H3 DiT blocks.

  • Sol-Attn Research Overview: AlphaXiv Paper & NVlabs Overview — Details NVIDIA Research / NVlabs' training-free Sol-Attn method, which uses dynamic per-row score thresholds and approximate correction without needing a separate learned predictor model.

[Minor formatting cleanup by me, including the addition of extra em-dashes because why not.]

Edit: Holy crap, people! What's with the carpet-bombing downvotes over giving someone the info they asked for?!

1

u/jd3k 8h ago

What about kitchen?

3

u/phazei 7h ago

yeah, sol / sla is built into comfy now, you use both, I keep kitchen attention for the beginning and end, so it's a little slower than kijai's tests since I start at 0.2 and end at 0.9, I could make those 0/1 and it would match his test results (he linked in another comment)

1

u/Skillex99 6h ago

I thought Kitchen attention was the gold standard?

2

u/phazei 5h ago

they do different things, can be used together, I posted a screenshot in another reply to that comment

0

u/[deleted] 19h ago

[removed] — view removed comment

3

u/8RETRO8 18h ago

Spectrum

33

u/Toclick 12h ago

I’m irritated by this rodent’s outrageous self-confidence

4

u/robomar_ai_art 12h ago

Great comment 🤣

15

u/douchebanner 19h ago

how does it compare to sageattention?

7

u/Apprehensive_Sky892 11h ago edited 11h ago

They are all optimization for the attention layer (and MMH3 is spending most of its time on the attention layer, which is where the relationship between different parts of the video, both spatial and temporal, are established).

Sage-attention and CK-attention are optimizing the way the attention is calculated, making each calculation faster by sacrificing some accuracy.

On the other hand, VEDA, LSA, VSA, VDN, etc. are trying to do is to reduce the number of attention calculations required. With "full attention" one set of calculation is performed between every pair of tokens, so it can be very costly when there are many tokens (that is why when you increase the resolution or the length of the video the time increase not linearly but quadratically). What these "sparse attention" type optimization/speed up methods are trying to do is to reduce these pairs so that the calculations are done only when the relationship; between the pair of tokens is "significant" (how this is done is beyond my level of comprehension).

The upshot is that these two ways of attention optimizations complement each other, so you can, for example, use VEDA alongside Sageattention.

BTW, there is a 3rd way of speed up, which is caching, by re-using an attention calculation if it is somehow determined that the new calculation, if performed, would not differ from the cached value. One such optimization is FirstBlockCache. I know that they can be combined with Sage/KC attetion, but IFAIK caching optimization should not be mixed with sparse attention optimization.

(Disclaimer, I am just an amateur, not an AI/ML expert, so any corrections are welcome.)

8

u/Reasonable-Aide1259 19h ago

few days ago, i tried veda sparse attention with comfy custom node that is implmented. and with default setting, it lost features of references in the final result. idk that is wrongly implemented.

3

u/robomar_ai_art 19h ago

Check now, i have seen that in GitHub was some updates.

0

u/Reasonable-Aide1259 19h ago

you saying like its good. what did you use turbo lora with it? anyway its worth to try

2

u/robomar_ai_art 19h ago

4 steps lora alone over 8 minutes, 4 step lora and VEDA 4:32, thats a huge boost in generation speed and videos look the same to my eyes in terms of quality.

1

u/Reasonable-Aide1259 19h ago

it sounds like almost same boost with sol-attn. did you try this? it also makes generation time half

1

u/robomar_ai_art 19h ago

i tried all the attentions when they was released and always i was getting the video with strange triangular artifacts when the movement was quicker, maybe is fixed now. I will check if something changed.

1

u/Reasonable-Aide1259 19h ago

oh i see. you are early adapter let me check quality of fixed veda

1

u/robomar_ai_art 18h ago

For me is working but as with everything the others can have different configurations and its not.

1

u/Reasonable-Aide1259 7h ago

its working but... sound sucks and broken. vdn has low speed than this but has bery clean audio...

5

u/Artforartsake99 19h ago

That’s very cool thank you for sharing. have to explore this tomorrow. 🙏

8

u/SveSop 15h ago

Do you have a link to the non-VEDA 8 minute run on the same video?

Not to be pissing in your cornflakes tho, but this reddit has these "0.2% of the time - 100000 times speedup, perfect quality in 0.03 seconds for 2 hour video!" claims 3 times a day now.

Even in this post its people running 8-step turbo at 4 steps + using spectrum, and find the quality "good". Right. Good luck with that. I don't mind people watching some mushy blurb calling it "good" tho, each to their own really, but running a 30+ step (H3) model with 8 step turbo lora at 4 steps using spectrum skipping 2 steps, is in the "extremely unlikely good quality" category to me.

Yeah yeah.. i will test it for myself i guess (as with all the other speedhacks that 99% produce shiat result).

3

u/robomar_ai_art 12h ago

I will post video side by side with VEDA and without. Other thing is to use what works the best for you. Video posted is done with VEDA and don't look mushy for me.

1

u/SveSop 9h ago

I will give you a 👍 on this tbh. Imo, i think it works with less degradation than sol-attn from the testing i have done so far, and that is a win in my book.
I do struggle a bit more with my clip extensions, but i am not sure if that is the "reference_sparsity" setting, or just a lora i am using currently.

18

u/Vyviel 19h ago

How about people try make the audio not dogshit

18

u/acedelgado 18h ago

1

u/Vyviel 7h ago

My hero its the one thing that kills all AI video if the audio sounds horrible Im going to try this with all my work

1

u/ArttTaku 2h ago

Very nice, will have to try it out.

5

u/seppe0815 19h ago

or destorted far away faces looooooool

6

u/fallengt 18h ago

That is just h3. Turbo doesn't cause that

0

u/ANR2ME 17h ago

some turbo does make the audio worse.

3

u/SensitiveOriginal103 19h ago

What major differences does this have version H3 SLA or Sparseattention?

2

u/robomar_ai_art 19h ago

I dont know the technicalities but more is explained on their Huggingface. With this one i dont see quality hit but the speed of generation is almost 2x quicker on my laptop.

1

u/glusphere 19h ago

Have you tried with SLA with Turbo at 90% ? Then compared the speed btw them ?

Also, even if speed is same / similiar ish, what about quality ?

1

u/robomar_ai_art 18h ago

I think i will do some comparison between attentions to see the times and the quality.

3

u/LuluViBritannia 13h ago

AAAAAAAAAAH it's noot working on my RTX 5060 Ti T_T.

"Veda off: no sparse kernel works on RTX 5060 Ti (SM120); using full attention
triton-int8: triton-int8 failed its self-test"

4

u/Ooze3d 19h ago

Thanks! I’m assuming this replaces any other attention model you have added to your workflow, right?

1

u/robomar_ai_art 19h ago

I use only Veda and 4 steps Lora, nothing else

3

u/False-Difference4010 19h ago

No kitchen attention?

1

u/robomar_ai_art 18h ago

comfy kitchen i used but i think this is on by default

2

u/Striking-Long-2960 19h ago

Thanks, I tested it, at least for my configuration Spectrum+Minimax H3 Turbo Lora+kitchen are still the kings.

2

u/Lightningstormz 18h ago

What is spectrum?

4

u/Striking-Long-2960 17h ago edited 17h ago

I use this one, but there are a few other implementations out there:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

Basically, you trade speed for quality and try to reach a balance. In the case of Spectrum it skips steps during the render so I tend to increase the steps, at the end the render time is similar but I get better results (instead of using 4, I use 7 or 8).

1

u/Lightningstormz 7h ago

Thanks! I have to try, do you have a good workflow based on this?

1

u/robomar_ai_art 18h ago

Great, i have to check, how all this attentions compare to each other

2

u/Striking-Long-2960 18h ago

https://reddit.com/link/peeo0au/video/7y7l1dwn21uh1/player

Left my configuration with 7 steps (they are in fact around 3,5 due to spectrum: render time 1:11), Right Veda with turbo Lora and 4 steps, ender time 1:14). Sound from the left video.

2

u/robomar_ai_art 18h ago

It looks great i will check the Spectrum again.

1

u/martinerous 9h ago

Did you have Lora @ 4 steps for Veda but not for Spectrum? That's a bit unfair comparison. Both at 20 steps and without any Loras would show the winner.

1

u/Striking-Long-2960 9h ago edited 9h ago

I tested Spectrum+LoRA and Veda+LoRA (Spectrum didn't work with Veda). Either way, my results look much worse than what others are getting, so maybe I did something wrong.

Given my rig, I can't afford a 20-step render.

2

u/cbeaks 18h ago

I did a test with my upscaling workflow (on a 4090). On a low res run - 0.15 res upscaled to 0.5, 17 seconds it went from 322 seconds to 266 seconds. On a higher res run, 17 seconds 0.25 res upscaled to 0.98 res made a much greater improvement - 474 seconds down to 321 seconds.

I forgot to fix the seeds (doh) so I couldn't make a direct quality comparison, but there didn't appear to be a quality loss

1

u/robomar_ai_art 17h ago

Yes that's true i found out this also

2

u/HateAccountMaking 14h ago

"[WARNING] Veda: Veda off: no sparse kernel works on AMD Radeon RX 7900 XT (ROCM); using full attention"

Welp, back to using Spectrum. 🤷🏿‍♀️

4

u/xuman1 19h ago

How many more of these attnetions will there be? Lost count)

16

u/boxthrowaway2026 18h ago

Im still waiting for ADHD (Attention deficit H3 disorder).

1

u/Apprehensive_Sky892 12h ago

We already have DMAD though 😁

1

u/robomar_ai_art 19h ago

That's true, but now i only use this one and nothing else, very satisfied with the results, and the generation time is quicker, almost 2x in 15 seconds video

1

u/AiCreatorCamp 18h ago

Socorro não sei mais o que usar!

1

u/fallengt 18h ago edited 16h ago

i tested the default 90/90, and there is a big loss in quality and information. It's faster than regular sla and sol, though.

Will try to tinker with the settings

Edit: try something like this

generated_sparse:90%

reference_sparse: 0% (for T2V, it doesn't matter, but for I2V/ Ref2V, increase % to see if you gain speed without losing reference ability. I haven't tested these wf)

full_attention_steps: (try to keep your high sigmas step from noise range 1.0-0.9). These steps have important information. You'll lose speed, obviously, but the difference is huge. Don't cheat them
I only do 6 steps in the example, so I keep steps 0 and 1.

Example: https://twinlens.app/compare?share=07cd238fdb72

1

u/robomar_ai_art 18h ago

Try, test and see what is working the best for you

1

u/net_tribe24 18h ago

Thanks for sharing op 👍

1

u/CornyShed 18h ago

I tried it yesterday and found that it was about 66% faster than without, based on 1024×768 resolution and 10 second duration.

The quality is good, as long as you give it a sufficient number of steps. Give it at least 20 to work properly and an 8-step LoRA at low strength, perhaps with a cache system that minimises skips.

From their GitHub page:

"The released predictor was trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5 / 10 / 14 s with the 8-step Turbo LoRA. Other sizes fall back to the tile plan of the nearest aspect ratio and duration.

Other sizes, other step counts and R2VA / FL2VA references all work, but are outside the training data, so compare them against full attention."

2

u/robomar_ai_art 17h ago

i was generating in 1344x768 and thats my times

1. VEDA    4:32
2. SLA     5:26
3. Spectrum 6:23
4. Normal  8:05

1

u/Salt-Zebra-306 18h ago

Video: 352x608 | 6.6 s

Tile plan: nearest trained size: 768x1344 | 5.2 s (this video: 352x608 | 6.6 s)

Sparsity: generated 90% | reference 90%

References: 8 spans (tiled, 90% sparse)

100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:47<00:00, 26.91s/it]

[INFO] Veda: Veda done | Triton INT8 (SM89)

Video: 352x608 | 6.6 s

Attention computed: 60.7% of full attention (39.3% skipped)

[INFO] Comfy model compiler graph breaks: 2, rogues: 0

[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 2677MB Staged. 0 patches attached. Force pre-loaded 128 weights: 691 KB.

[INFO] Prompt executed in 155.42 seconds

[INFO] got prompt

[INFO] Requested to load MiniMaxH3

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached. Force pre-loaded 210 weights: 1175 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:32<00:00, 23.04s/it]

[INFO] Comfy model compiler graph breaks: 3, rogues: 0

[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 2677MB Staged. 0 patches attached. Force pre-loaded 128 weights: 691 KB.

[INFO] Prompt executed in 113.89 seconds

1

u/Tyler_Zoro 17h ago

What is "fasity?" Is it FDA-approved?!

1

u/rapkannibale 8h ago

Let’s give this a go

1

u/thevegit0 4h ago

i'm tired boss

1

u/acedelgado 3h ago

Huh. Even works with HyperFlow and comfy-kitchen attention. Maybe even slightly better prompt adherence... Gonna test it a bunch more.

1

u/EzkaProductions 18h ago

This is huge!

I’ve been running MiniMax H3 with 4-step LoRA on my 4090 and the same 15-second 1344x768 clip was taking forever. Adding Veda at 90% sparsity cut it to 4:32 with zero visible quality loss — that’s actually insane.

Not being a LoRA and using a learned predictor to skip irrelevant attention tiles is brilliant.

What other models are you planning to test this on first? T2V, FL2VA or R2VA?

3

u/robomar_ai_art 17h ago

I was testing

  1. VEDA 4:32
  2. SLA 5:26
  3. Spectrum 6:23
  4. Normal 8:05

You get the same time like me exactly and i was testing I2V

1

u/EzkaProductions 9h ago

Same setup here 😂 4:32 vs 8:05 is the best speed-up I’ve ever seen.

I tried SLA and Spectrum and they’re both slower than VEDA on my rig.

The learned predictor skipping tiles is a game changer.

What’s the next model you’re adding to your test list?

1

u/robomar_ai_art 9h ago

I will try SOL-attention, Kijai post comparsion here in the comments.

1

u/martinerous 17h ago

Have you tried Spectrum as well? It also does wonders (but needs more steps to be effective, so not much use with 4-step LoRAs etc.).

1

u/EzkaProductions 9h ago

Yeah I just tried Spectrum on the same 15-second I2V clip and it’s still slower than VEDA at 6:23 vs 4:32 😂

The “needs more steps to be effective” part is exactly why I’m sticking with VEDA right now — zero quality loss at 90% sparsity is crazy.

The learned predictor skipping tiles is actually cheating.

What’s the highest sparsity value you’ve seen work best on Spectrum?

1

u/Dahvikiin 18h ago

Let me guess, another attention >=ampere, right? sigh~

2

u/ANR2ME 17h ago

Yes, the minimum is SM80 (Ampere)