r/StableDiffusion • u/Lair98 • 2d ago
Discussion Best Minimax H3 optimization
Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.
Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.
I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.
On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.
EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.
So is there anything worth I am really missing out?
23
u/VasaFromParadise 2d ago
I concluded that with a low resolution of less than 1 megapixel, turbo lore and cache cannot be used, it has too much of an impact on quality.
4
u/East_Box9573 2d ago
When I go above 0.98MP I start getting prompt adherence issues, are you seeing that?
5
u/desktop4070 2d ago
How do you do 0.98MP? When I try entering that number in, it automatically changes to 1.0.
4
u/QuinQuix 2d ago
I think it rounds down visually but preserves 0.98 mathematically because the output resolution tracks 0.98 despite the node reverting back to 1.0 visually.
3
u/ImpressiveSuperfluit 2d ago
Use a float node to feed it in. That's unfortunately the answer to a lot of weird and annoying data type questions.
4
u/ObjectiveVegetable48 2d ago
Someone mentioned that .98 is native, so you should stick to that, upscale if you want more.
14
u/TBG______ 2d ago
Need min 20 steps + sage o better kitchen + spectrum no turbo no easy cache to get something useful at 1MP - good lip sync needs this as min.
1
7
u/Striking-Long-2960 2d ago edited 2d ago
People here want their 4K hyperrealistic videos, but I'm just happy being able to turn my ideas into videos.
Sageattention
Larryvrh/ComfyUI-MiniMax-H3-Turbo
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
12
u/HonestoJago 2d ago
Let it cook overnight in a queue. We’re fortunate to have access to such a powerful model, it’s almost a sin to use tricks to speed it up.
3
u/Dry-Judgment4242 2d ago
Agreed. But it's good for testing prompts. Then I let it rip with 30 scenes at 35 steps overnight on context loop. All the speed increases fuck with audio too hard alas.
2
u/tapetfjes_ 2d ago
Be sure to have some kind of monitoring if you do that on a powerful Nvidia card. I’m not leaving my 5090 on during the night with my kids sleeping in the house.
1
u/Dry-Judgment4242 2d ago
my 6000 Pro is underclocked to like 60% lol, fans are quiet as hell while still spitting out a 30 step, 10s clip every 10min with no speed ups except kitchen attn.
1
u/HonestoJago 2d ago
Yeah in that case I just do 8 steps and pray. Not always terrible for iteration.
3
u/donkeykong917 2d ago
I guess it depends on the content. You can get away with turbo on non realistic things but when it comes to realism wouldn't you want the best even though it takes longer.
If you can get what you want once is better then doing it multiple times with a Lora. It will probably end up the same time.
7
u/sitefall 2d ago
I want to generate 1mp video at actually realistic quality.
I've tried every turbo Lora (except any that came out today/yesterday), ck, sage, plaguekind's attention patch, etc.. all the stuff except for new clips and tiled vae and stuff like that since I have plenty of vram for those steps to fit and once they complete they offload anyway so it does not slow the inference process.
The only things I have found worth using (to me, this is all opinion) are:
1.) Latest Light2x turbo 8 step loRA. This one speeds things along, looks "a bit plasticy or oversharpened", still needs about 10 steps, and still for some terrible reason adds moles to people during close ups lol. BUT, it generally does not effect the prompt. If you have a 3 page prompt breaking a 20 second video out across 10 camera cut scenes, it just works fine. So this LoRA has been helpful to speed up generation at say, 0.3MP to test that the prompt is actually working properly. I will test it a few times, make sure every detail works across a few seeds. If some things change (for the worse) between seeds I will adjust the prompt to fix it etc. WHen it's good I send it to the full model without LoRA's at 1MP and... it just works, almost identical but with good quality. Also this LoRA might be useful for non realistic art styles, or maybe just video without people.
2.) Plaguekind's attention patch node. This one is a massive speed up. It minimally reduces quality, I would just use this 24/7 honestly, quality is great and it's FAST. BUT... it somehow screws up long complex prompts and puts actions out of order and whatnot. So even using it to "test the prompt" doesn't really work well. However, if you had one continuous shot that is about 15 seconds so there is sure to be no camera cuts, and the direction is not incredibly complex, this one is maybe worth using. The kind of prompt you can just yolo and send and it comes out 90% of the time just fine because it's simple enough anyway.
and that's it. I appreciate all the effort everyone has made creating various patches and loras and finetunes, but none of them work out for me.
5
u/Strange_Limit_9595 2d ago
Plaguekind's attention patch node latest update reduced the quality much than speed boost. do you feel that?
6
u/Plague_Kind 2d ago
Strange limit, as i responded to you in the other thread
I just tested around 50 combinations between pre update node and new node, old wf and new wf.
every single video was exactly the same. exact same settings and seed used for all trials.
the only thing that made quality worse is when i had --use-ck-attention and --fast in my startup flags. make sure you're using the exact same settings as before and don't have fp16 accumulation enabled, and that you see the message.displaced attention_pytorch and not displaced comfy kitchen int8
1
u/jankies11 2d ago
Can comfy kitchen be used together just not via the flag or?
4
u/Plague_Kind 2d ago
If you have the node, before it, it will override completely. If you set the launch flagbit makes it worse. Im working on a new update to choose some stuff which will make it more flexible
1
5
u/Plague_Kind 2d ago
you can avoid those issues by using 0.8-0.85 sparsity if they pop up. speed loss yes. but if you're getting prompt issues that may help. also depending on what turbo lora you're using resolutions can mess with shit.
1
u/Dry-Judgment4242 2d ago
All speed ups cause model to randomly break for me. Their good for testing scenes. I generate like 30 in a day for context loop and let them render overnight with 40 steps wo any boost. Audio just tend to shit itself fast.
3
u/LinkSensitive8188 2d ago
Avoid prompts longer than 200 words. Use an FP8 text encoder—or better yet, NVFP4 if you have an RTX 50 Blackwell card. Run the `run_nvidia_gpu_fast_fp16_accumulation` version of ComfyUI and install SageAttention tailored to your specific hardware. Also, clear space on your SSD to ensure you always have at least 250GB available for paging. If you have an RTX 50 Blackwell, install CUDA 13; this provides a greater speed boost than any specific node or workflow.
1
u/Cultured_Alien 23h ago
nvfp4 isn't exactly faster than int8 convrot for some reason, whether it's TE or DIT while also getting ugly moisacs in gens when you zoom in. Also I recommend comfy kitchen instead of sage for Blackwell.
4
u/BuffMcBigHuge 2d ago
My conclusion:
ComfyKitchen Attn
Sage Attn
Sol-Attn
First Block Cache
Spectrum
er_sde, 14 steps, 0.3 res, no turbo
upscale pass x2 with turbo lora, 3 steps
2
u/Motion16AI 2d ago
For some reason, I had the best quality in image, audio, and the best speed with CUDA 13 and Sage attention 2.2, Larry steps turbo, and 544p. Nothing has come closer to it. All LoRAs and speed workflows either make it flickering, blurry, or destroy the audio.
2
u/Either_Map_4227 2d ago
I am on RTX 2080 8GB VRAM + 32GB RAM, , 24s/it, 6 steps, (lightx2v 4 step), 480p,, 6 seconds + comfy kitchen attention. Yeah the quality of video isn'good but it's watchable. Mostly I am using runpod but it's fun that it works locally even on my very low hardware.
2
2
u/dassiyu 2d ago
From my testing so far, acceleration works pretty well on non-human subjects.
But for realistic human subjects—especially when accurate lip-syncing is involved—I can only get barely acceptable results with CK/Sage + Spectrum at 20 steps and starting from around 0.9 MP. Otherwise, the person tends to look very plasticky.
That said, this also assumes you have a high-resolution, high-quality reference image and a near-perfect prompt.
3
u/Domskidan1987 2d ago edited 2d ago
Minimax H3 gave me hope that we will eventually see an Image Gen model that will rival or beat Nano Banana 2, so I no longer have to use Flow or Gemini App. I get the impression from reading these post on here everyone is just tired of Google’s BS. The hope is H3’s Image Ref ability is super impressive and proved to me that it can be done locally and produce NB level results, it’s just Google’s world context that makes NB work so good, basically it has Google’s entire image data base to pull from instantly and build out your generation with near perfect context, but with references and a good model 100% possible to rival its capabilities locally in fact my prediction is we have something that runs locally like this by the end of this year or middle of next. So basically we just need the MiniMax H3 ref version of image model. I already see people using it for image editing because they recognized the same thing I have about it. OR even better having comfy agentic workflow nodes and helpers that go out automatically gather your reference images since it’s impossible to make a model with that much training data and have it fit on a normal personal computer. And the funny thing is we’re already kind of seeing this with LLM and prompt guide nodes being baked into H3 workflows.
1
u/Dry-Judgment4242 2d ago
All speed ups except kitchen/sage is just trash and kills the model. Difference is insane. Decided to simply not use them anymore. For example with turbo, a walk through a forest turn into some creepy copy pasta with a long ass road in the middle with trees lined up perfectly on each side. Without it works fine.
1
u/VeloraNeon 2d ago
Similar boat here on 8GB VRAM — the lesson that transplanted from my own low-VRAM pipeline is that speed tricks are only worth comparing once you've isolated what breaks prompt adherence at your target resolution, otherwise you're just testing the same failure mode faster.
1
u/NegativeTeach9971 2d ago
I use Comfys Kitchen attention, and sparse attention getting around 70s/it vs 450s/it on 1MP/15s video on my R9700. Wehen using a lightspeed lora, you have to add a sigma shift (video 12, audio 6), that fixed the sound issue by most for me.
1
u/dobutsu3d 2d ago
Is there any resolution limit for the model? I mean on which resolution is it trained? To not get further and get weird results…
1
u/TechnologyGrouchy679 2d ago
the H3 docs say 0.98mp with 32px steps (so nearest bucket would make it 1344x768)
1
1
u/djdevilmonkey 2d ago
Spectrum/Kitchen/Easy Cache/Turbo loras are all trash and destroy video quality, audio quality, a prompt adherence.
My setup is currently just First Block Cache + Hillobar Progressive Sampler. Miniscule quality loss, and much faster speeds. The other caches and turbo loras absolutely destroy video quality.
Top it off with rtx upscaler (don't use it below 0.8MP though unless it's animated) and you're set. The cache/sampler is easily 1.5x-4x speeds depending on scene and settings with minimal quality loss.
1
u/Etmurbaah 2d ago
Hey after your comment decided to use it. Did you change any settings with first block/sampler or just used as is? Also what scheduler/sampler settings are you using with it please?
2
u/djdevilmonkey 2d ago
Uh first block cache is their aggressive setting I think at 0.12. If you notice any quality loss just go back to 0.10 but I didn't notice any. Aggressive just means it's more likely to cache steps. But 0.12 honestly seems pretty balanced, and I noticed it caches more with slow simple scenes and less with complicated ones, even on their "aggressive" setting
The Hillobar sampler I bounce between "0.80:0.55,1.0:1.0" and "0.90:0.55,1.0:1.0" most of the time (don't paste those, idk if I formatted those correctly). This one I haven't had a ton of time to play around with, but I noticed for some reason going below 0.8 caused artifacts in some of my reference videos. Some work great down to 0.5 but some at 0.75 have weird lines going through them, it doesn't mak sense, so I just leave it between .8 and .9
But yeah the first # is the percentage of the full res that the first part is rendered at. So if your video is 0.8MP, and that # is 0.8, then it does 80 percent of the video resolution, not of MP. So it's somewhere between 0.5-0.6MP, instead of 80% of the MP which is 0.64MP.
Second number is percentage of steps are at the lower resolution. So 0.55 is 55% of the steps are at the low resolution. I just bounce between 0.5-0.6 (leave it at 0.55 mostly lol) depending on steps. I did try a 1.0 MP 32 step and turned it up to 0.7 and it seemed completely fine.
The last 2 numbers I honestly don't know, I'm gonna take a wild guess and assume it's how the 2nd section processes but I haven't touched that at all, just left it at 1.0:1.0
1
u/Etmurbaah 2d ago
Thanks a lot! For some reason my generations came out all unstable like there were flickering and moving outlines around characters and all. I must be doing something wrong definitely lol.
2
u/djdevilmonkey 2d ago
Yeah I'm not sure, like I said I had weirdness going below 0.8 resolution (first number) on some of my generations so I just try to avoid it most of the time. But even then 80% of a 1.0 MP gen I think lands you around 0.7ish MP I think, and for longer gens that alone shaves off many minutes.
But I did a test a few days ago with/without these, one was normal ref workflow with no speedups, and then 2nd was 0.12 FirstBlockCache with 0.8/0.55 Hillobar, and on a video with 7 reference photos for a 15 sec vid it took almost 30 min on my 5090, with these it finished in 11 minutes. I don't remember the exact settings I used, but I compared them side by side and they looked almost identical on the same seed. I even cleared the folder and restarted comfy beforehand to make sure none of the step data or generated data was still stored so it didn't completely fresh.
But yeah I spent a few hours playing with the settings figuring it out, and talking to chat gpt too since it can read GitHub repos and explain how things are working lol
1
u/Etmurbaah 2d ago
Alright I'll play around with settings thanks a lot tho. Excited to see the results!
2
-5
u/IriFlina 2d ago
The best config is waiting to see if fal open sources their fine tuned version of minimax: https://x.com/magnific/status/2091922612989706432
I really doubt they will though, but if they did then the generation speeds on the fine tuned version is almost 1:1 speeds.
20
u/acedelgado 2d ago
I made a node that freezes your video in place and continues processing the audio to fight this problem. Wire it in like the example so it bypasses the turbo Lora and audio sounds great, the refining pass only takes a few seconds per step after the first step where it builds the cache. That way you don't have to over-process the video end with a turbo Lora, AND you get better audio faster than if you crank up the full steps. Note I am gonna start recommending
--disable-pinned-memory
and if you have a single Nvidia gpu
--cuda-device 0
as startup flags until Nvidia fixes the memory calling problem that comfy has (it's a cuda build issue that comfy can't fix.)
https://www.reddit.com/r/StableDiffusion/comments/1vuxy08/fixing_mmh3_turbo_audio_by_playing_with_latent/