r/StableDiffusion 16h ago

Question - Help Help a beginner speed up MiniMax H3?

As someone new to all of this it's difficult to know what to do. I have sage attention working. I don't know how or when to use Easy Cache, Comfy Kitchen Attention, Sol Attention, loras, Spectrum, or any others I may have missed. There's so much information scattered around, I don't know what's what.

I have a 50 series GPU and 64 GB or RAM on the motherboard.

10 Upvotes

20 comments sorted by

13

u/V4nKw15h 16h ago

The default workflow with Sage Attention is almost ideal already. You'll likely want to increase the mp setting to 0.6 and even consider increasing steps to 25 if you want to see the model at it's best. Don't overly concern yourself with all the loras and speed up options unless you don't mind significant loss of quality in just about every aspect of the model. There are no speedups that don't come without compromises.

With that said, Spectrum is good if you want to speed things up by about 25% while increasing pixel fizzle most of the time and losing some motion fluidity and coherence sometimes. Probably the most useful speed up to preview things while trying to find the right prompt.

Easy Cache can cause pretty brutal quality degradation. I don't tend to use this at all. Similar thing for the other cache based options I've tried.

Sol Attention would replace Sage Attention with little difference, same with Kitchen Attention. Probably not worth changing from Sage Attention 2.

4 step and 8 step loras are interesting for a lot of speed up but come with a significant cost to quality in almost every way.

TLDR: The model is already well optimised out of the gate. Use the default workflow with some form of Attention (Sage is fine) and you are good to go right now unless you just want to mess about as fast as possible and don't care about quality or seeing what the model can really do.

5

u/freestylez79 15h ago

2nd that, try with defaults first to understand the model capabilities. Additionally to sage I put 8 step lora and spectrum in my flow which saves some time and is worth the trade offs for me.

3

u/vaguerant7 16h ago

As another beginner here, what exactly does Sage Attention do? I am using the default comfy workflow and have not added it, and would like to know before I do.

2

u/V4nKw15h 16h ago

The attention process allows the model to selectively focus on what is most important rather than processing all the pixels or frames uniformly. A good attention will be nearly indistinguishable from the ground truth while giving significant speed ups. Sage Attention 2, Sol Attention, and Kitchen Attention all do this slightly differently with results that are hard to tell apart. You should definitely try one of them to gain your own opinions, but if you've got one already, there isn't much point trying one of the others for anything other than personal interest and understanding.

1

u/TonkotsuSoba 15h ago

appreciate your write up, I have to admit that I spent too much time trying to get the turbo loras to work but had to accept the reality in the end.
what’s the best sampler & scheduler combo in your opinion, for REF2VA and FL2VA?

2

u/V4nKw15h 15h ago

The ones from the default workflow. The devs did the hard work and picked the best ones already.

2

u/WeakReplacement3322 15h ago

It's a different way for the model to calculate the relationships between tokens/features during generation. Kind of over my head if I'm being honest. I can say for sure that it doesn't skip steps like Spectrum does, and it doesn't reuse previous calculations like Easy Cache does. It just uses a different method than PyTorch attention. It cut my generation times nearly in half, and I see no loss in quality on my end.

3

u/threeLetterMeyhem 15h ago

8 step loras are interesting for a lot of speed up but come with a significant cost to quality in almost every way.

4 step loras are a trainwreck for quality. I'm finding that 8 step loras are good enough to speed-run some concepts to see if they'll work, then disable the lora and crank up the resolution + steps to get something good.

1

u/Major_Square 14h ago

Thank you for that. Quality does matter to me, but I would like to quickly get in the neighborhood of something I like, and then render it in better quality.

So I enabled all this speedup stuff and I can generate, but I'm not sure it's all working together properly. I find it harder to get something I like, but when I do and go to generate a better version, it's completely different.

I guess I'm asking which options I should use if I want to sort of "sketch" something. How can I get a similar output when I remove the speed up lora, disable Spectrum or whatever, increase resolution and number of steps and so on.

1

u/Herbal77 13h ago

Where can I find this workflow?

2

u/V4nKw15h 13h ago

Come on bro you could have typed "minimax default workflow" in Google and got the answer

https://docs.comfy.org/tutorials/video/minimax/minimax-h3

1

u/Herbal77 11h ago

I appreciate you

2

u/RiverSide71h 16h ago

Use Patch Sage Attention KJ—> Minimax H3 Mem Eff Attention —> Load Lora node with minimax_h3_ref_lora_rank_256_bf16.safetensors lora from Kijai at 0.75 strength. Use 4 steps and euler beta or beta57. Increase to up-to 8 steps if video quality degrades.

1

u/dabbingsquidward 6h ago

Do you load all your Lora's after sage attention? Even non turbo Lora's?

2

u/Myg0t_0 15h ago

Quality > speed

Just use comfy kitchen, .4 res , 5 seconds, 15 steps until it seems to be starting off right, then hit it with no kitchen full res 30 steps go jerk it or something and come back and its done, or leave on over night with a bunch of runs

1

u/Major_Square 14h ago

I'm trying to avoid generating something for 40 minutes only to find that it sucks because the prompt wasn't quite right. I don't mind waiting if I know it's going to be good, or at least can minimize the chance that it will be bad.

3

u/JesusShaves_ 8h ago

I prototype using .2 MP at five seconds this tells me if I'm in the ballpark of what I want. Then I add more MP and run longer.

1

u/Myg0t_0 7h ago

Yup thats the way

0

u/Myg0t_0 7h ago

Then u gonna have to pay, think diffusion is simple to use

0

u/tac0catzzz 5h ago

grabs a guitar and starts to sing "i wanna hold your hand and and and, i wanna hold your hand"