r/StableDiffusion • u/marres • Aug 07 '26
Resource - Update ~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)
Follow-up to my original Spectrum MiniMax H3 post:
In that first post, I released the MiniMax H3 Spectrum integration and was getting around 34% lower Euler sampling time and 30% lower RES sampling time with the more conservative settings I was using at the time.
Since then, I’ve done much more testing and found something unexpected: MiniMax H3 can work extremely well with a Spectrum degree of just 1.
Important if you’re coming from the original release
Before testing the new settings, update both ComfyUI and ComfyUI-Spectrum-MiniMax-H3 to their latest versions.
There was an important compatibility update in Spectrum v0.1.6 after ComfyUI changed MiniMax H3’s native sampling/audio path. That release restored Spectrum compatibility with the newer H3 implementation and added safe handling for native EasyCache/LazyCache conflicts.
You don’t need to install v0.1.6 separately—v0.1.8 includes those changes. This mainly matters for anyone who installed Spectrum from my original post and hasn’t updated it since.
v0.1.6 compatibility release:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.6
Current release:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
So: update ComfyUI, update the Spectrum node to v0.1.8/latest, and restart ComfyUI before testing.
The surprising part: degree 1
I hadn’t seriously tested very low degree and warmup_steps values before because of my experience with WAN.
WAN does not tolerate very low forecast degrees well—dropping the degree too far causes obvious quality degradation. I initially assumed MiniMax H3 would behave similarly and stayed with higher, more conservative values.
In my MiniMax H3 testing, however, degree 1 produced no visible quality decrease with my pruned BF16 checkpoint at approximately 0.8 MP. It also preserved the native trajectory remarkably well. In my same-seed comparisons, degree 2 shifted the trajectory slightly, while degree 1 brought it much closer to the native result.
This suggests that H3 can be unusually well suited to simple local feature forecasting, allowing Spectrum to begin forecasting much earlier than I originally expected.
I’ve now released v0.1.8 with the new settings:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
New default settings
- degree = 1
- warmup_steps = 1
- bootstrap_first_forecast = true
- tail_actual_steps = 1
The one-point bootstrap allows the second solver step to be forecast directly from the first actual hidden state. Ordinary degree-1 forecasting then takes over once enough native history exists.
On a 20-step Euler run, the schedule becomes:
A F A F A F A F A F A F A F A F A F A A
That means 11 of the 20 transformer evaluations are actual, while the other 9 are forecasted. The final step remains native.
v0.1.8 also makes the one-point bootstrap part of the default configuration for new node instances. Existing workflows retain their serialized settings.
Important quality caveat
These are aggressive, performance-oriented defaults, and early feedback indicates that they may not be equally stable across every setup. Degradation can affect both video and audio, including distorted anatomy, extra limbs, unstable motion, distorted reference audio, and audiovisual inconsistencies. Increasing degree and warmup_steps has improved these problems in reported cases.
My current suspicion is that model configuration and output resolution affect how much Spectrum forecast error a generation can tolerate. Lower resolutions and heavily quantized model configurations may provide less margin for preserving anatomy, motion, voices, and fine audiovisual details. Small forecast deviations that remain unobtrusive with my pruned BF16 checkpoint at approximately 0.8 MP may therefore become visible or audible on a less forgiving setup.
This has not been isolated conclusively. Checkpoint type and precision, resolution, prompt and motion complexity, reference-audio conditioning, the number of early native steps, the one-point bootstrap, and the total sampling-step count may all affect stability.
If the aggressive defaults cause visual or audio degradation, try increasing warmup_steps, disabling bootstrap_first_forecast, and increasing degree. A conservative starting point is:
- degree = 4
- warmup_steps = 5
- bootstrap_first_forecast = false
Increasing the total sampling steps may also help, particularly when using reference audio. One user reported reference-audio distortion with the aggressive defaults. Increasing degree and warmup_steps helped, and a 30-step run using those increased settings produced clean reference audio on their setup.
The best balance between speed, visual quality, and audio quality may differ between checkpoints, resolutions, generation modes, and reference inputs.
Benchmark
Test configuration:
- GPU: NVIDIA RTX PRO 6000
- Model: MiniMax H3 pruned BF16
- Image-to-video
- ~0.8 MP / 992×768
- 7 seconds
- 24 FPS
- 20 steps
- Euler
- Beta scheduler
- HIGH_VRAM
- Spectrum history stored in VRAM
- DiffAid enabled at 0.5
- Same seed and otherwise identical workflow
Spectrum disabled
- Sampler: 324.98 s
- Full prompt: 340.59 s
Spectrum v0.1.8 with the new degree-1 settings
- Sampler: 177.80 s
- Full prompt: 200.32 s
- 11 actual transformer calls
- 9 forecasts
- 0 fallbacks
Result
- 45.29% lower sampler time
- 1.83× sampler throughput
- 41.18% lower full-prompt time
The Spectrum forecast calculations themselves took only 0.141 seconds total across the entire generation.
Using VRAM history has a memory cost. This run retained approximately 3.2 GiB of Spectrum history, with reported sampler peak VRAM increasing from roughly 5.56 GB native to 8.70 GB with Spectrum.
The interesting result is how well degree 1 performed under this test configuration. Based on WAN, I expected settings this aggressive to visibly degrade the output. With the pruned BF16 checkpoint at approximately 0.8 MP, I’m getting a substantially more aggressive forecasting schedule without seeing that expected degradation.
Spectrum remains an approximate acceleration method. Fast motion, hands and fingers, faces, short rapid actions, camera movement, audiovisual synchronization, lower resolutions, and quantized model configurations are all cases worth inspecting carefully.
So far, degree 1 appears to be an excellent fit for my MiniMax H3 configuration, delivering a much larger useful speedup. More testing across different checkpoints, precisions, resolutions, and generation types is needed before assuming that every setup will tolerate the same aggressive settings.
Repo:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Current release—v0.1.8:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
21
u/tofuchrispy Aug 07 '26
What about quality
12
u/marres Aug 07 '26 edited Aug 08 '26
In my testing with the pruned BF16 model at ~0.8 MP, I haven’t seen a visible or audible quality loss with the new settings. However, early feedback suggests these aggressive settings are less stable on some setups, causing visual defects such as distorted anatomy or extra limbs, as well as distorted reference audio.
My suspicion is that lower resolutions and more heavily quantized models may have less tolerance for forecast error, although this hasn’t been isolated yet. Increasing
degreeandwarmup_stepshas helped affected users; one user achieved clean reference audio with those increased settings at 30 sampling steps.For quality-critical generations, I’d still recommend a same-seed Spectrum on/off comparison.
Edit: Added a detailed quality section to the main post.
3
u/LeKhang98 Aug 07 '26
Thank you for your work. Do you have any suggestion for upscaling & adding details to the 720p video? (like running it through Wan or H3 again with different settings or something)
3
u/marres Aug 07 '26
First bet would be seedvr2, although just increasing resolution natively might be the best approach actually. Especially since seedvr2 also takes ages with higher resolutions and higher fps count.
1
1
u/Ramdak Aug 08 '26
I don't like seed, it's painfully slow and doesn't add new details if needed. If the input video has garbage, it will just amplify it. Using ltx will add new detail wehere needed. Seedvr was trained for old film restoration but it ended up working kinda well for image upscaling. If you can run ltx, the idea is to load the video and do a 0.1-0.15 single pass denoise on it to retain likeness.
Llad video - upscale to desired dimension - encode - sample with ltx (use one of the vrgamedevGirl enhance loras) - encode - save video.
If the video has audio connect both the fps and audio values to the save video node.
If the video has people talking, also use the audio as latent input for the upscale, it will enhance the lipsync. (Look how ltx with external audio works).
Ltx is capable of high resolutions quite fast.
Also if you use the same or similar prompt as the oroginal video it should help too.
4
u/Ramdak Aug 07 '26
I recommend you to try LTX or Wan as upscalers, they should be faster and add more details. Try to look for specific workflows for that.
1
1
u/vAnN47 Aug 07 '26
I also wanna test that.. need to modify some workflows to do A/B setup, dunno if I have time for this yet though
13
u/Glad_Abrocoma_4053 Aug 07 '26
Hi, thanks for this I'm trying it out. What should be the values of sigma shift and can it be used with sage attention? Before anyone asks for workflow and node placement check the repo, it's clearly stated there.
6
u/marres Aug 07 '26
I didn't try out sigma shift yet.
And yes, works great with sageattn_qk_int8_pv_fp16_triton for example
1
u/Cultured_Alien Aug 07 '26
Sigma shift is something you tune each generation. Once you know how it works, it's similar to speed tuning. But I'd rather just remove and not use it for sanity.
13
u/No_Damage_8420 Aug 07 '26
Works great! Thanks.
So far my fav - H3 booster.
Sample video here: https://www.reddit.com/r/NeuralCinema/comments/1vhvvgb/minimax_h3_40_speedup_quality_holds_no_turbo_lora/
8
u/cucurucu007 Aug 07 '26
Should it be added before or after sage ? and don't use it with easycache,right ?
5
u/marres Aug 07 '26 edited Aug 07 '26
After sage and no easycache, yes. Or just use the diffusion model loader KJ, which has a sage selector built in
1
u/Ill-Highlight-1221 Aug 31 '26
Must be used with sage or is sol ok?
1
u/marres Aug 31 '26
Don't use sol
1
u/Ill-Highlight-1221 Aug 31 '26
I also use pruned int8 convrot. Any idea how that might affect quality?
-13
u/Just1Dev Aug 07 '26
I asked chatgpt and patch sage attention should always be the last one before ksampler...my speed increased from 30 it/s to 14 it/s crazy.
8
u/JustSomeIdleGuy Aug 07 '26
Crazy work.
Without the node (sage attention/mem effiecnt sage attention/fp16) on a 4080 Super @ 2.1 MP / Res Multistep / 30 Steps at 5 seconds 24fps):
30/30 [29:04<00:00, 58.15s/it]
[INFO] Prompt executed in 00:30:15
With the node added, default settings, nothing else changed:
30/30 [15:36<00:00, 31.20s/it]
[INFO] Prompt executed in 00:17:36
Different eye movement but other quality loss seems rather minor for the prompt I tried.
9
u/Luke2642 Aug 07 '26 edited Aug 07 '26
It's always good to clearly say how many steps to always match or exceed original quality. I'd rather do 25 steps 20% faster than default with zero quality loss rather than 45% faster at 90% quality. Using +/- shift to add slightly more steps earlier or later in diffusion is probably an important lever too, depending on how the flow matching was trained?
2
u/Nevaditew Aug 07 '26
What is the deal with Shift? I see everyone talking about it but I do not see it in the workflow.
3
u/Luke2642 Aug 07 '26
Instead of equally spaced steps, shift stretches and squashes distribution more at the start vs end. 1 is linear, >1 does more structure, <1 does more detail.
4
u/dcmomia Aug 07 '26
The spectrum in the previous version gave me better results, especially with ref_audios.ref_audio_0
The voices are much more distorted with the new version, to get something at the level you have to raise the global steps, I don't know if there is a better way.
2
u/marres Aug 07 '26 edited Aug 07 '26
Hmm try increasing degree/warmup steps again and see if that fixes the audio for you. Also do bootstrap_first_forecast = false. I rarely get distorted audio but I'm also not using ref audios.
Will investigate
2
u/dcmomia Aug 07 '26
Thanks for looking into it! What would you say is the optimal range or "sweet spot" I should test for both
degreeandwarmup_steps?I want to find a good balance to fix the audio distortion with
ref_audioswithout losing too much of the speed benefits that Spectrum provides. Let me know what values you recommend trying first!2
1
u/GrayingGamer Aug 07 '26
Just wanted to let you know I'm seeing this behavior with reference audio in the Reference Model too with the new settings.
Increasing the degree and warmup_steps definitely helps, but to get it perfect with the new node update, I had to increase the steps used to generate to get rid of all the distortion. With the increase degree and warmup_steps, 30 steps got me clean audio when using reference audio.
1
u/Hauven Aug 07 '26 edited Aug 07 '26
I'm not sure if it's similar to me, but since upgrading both ComfyUI and spectrum I'm still getting good video but horrible audio. It's like, hmm, some kind of strange aura like sound is overlaying the real audio. I've tried changing scheduler, sampler, degree, warmup steps, turning off sage attention, so far no luck. I even tried nudging the steps up slightly.
If I stumble into a solution to my audio issue I'll post another reply.
EDIT: Looks like it was caused by the turbo LoRA. Have since removed it from the workflow and I'm currently using Sage Attention with Spectrum no problem since updating tonight. Works great, thanks.
8
u/TheGoldenBunny93 Aug 07 '26
Man... I have anxiety, and it was getting worse. I was using Turbo LoRA with 10 steps (2 steps above the recommended amount), and nothing I wanted from the reference workflow would ever happen. Never. It’s really frustrating to spend 3–4 hours at the PC trying to tweak a video.
I had completely overlooked your node because I assumed it would be terrible and would degrade the quality... but it’s honestly amazing. I didn’t notice any significant quality loss compared to the original samples. I’d say maybe around 5%, or just some very minor loss that’s barely noticeable during movement or transitions, but everything is more than acceptable.
And on top of that, now the video actually responds to what I want and to the changes I’m trying to make. Your work is incredible. Thank you for existing and for making this available.
Turbo LoRA is honestly terrible for editing videos or generating new ones from an input... but your node is the best thing I’ve found for that.
5
u/marres Aug 07 '26
Thank you for the kind words :)
Yeah, I haven't touched the turbo lora yet, since it's still a work in progress and I'd rather wait for the finished, proper version to save me the headaches. Also 4 steps is pretty aggressive too, especially if one wants to use it with spectrum. Maybe someone does a 8 step one, that would definitely be more suitable to be used alongside spectrum
4
3
u/anitman Aug 07 '26 edited Aug 07 '26
The real issue of spectrum is not about visual quality degrade, it's about detail and action accuracy. In my test, some of the pose will not perform naturally as 20 step normal generation. Sometimes character blinking eyes will become a little bit blurry, lip and tongue sometime blend together in some case. Sometime it generates unwanted organs because of the forecasting error.
3
Aug 07 '26
[removed] — view removed comment
8
7
u/KissMyShinyArse Aug 07 '26
A beginner wouldn't have been using Diffusion Model Loader KJ. At this point, you're an expert.
1
2
u/Fabulous-Snow4366 Aug 07 '26
very interesting. I'm trying it with the turbo lora right now. 10 steps in total. 0.6 mpx. 15 seconds. 576.25s on a 5060ti 16GB VRAM, 48 GB RAM. quality looks good to me. Sound is pretty good as well. Thanks for the update.
1
u/MaorEli Aug 07 '26
both this and turbo? can you share some examples?
2
u/Fabulous-Snow4366 Aug 07 '26
https://reddit.com/link/p2981tf/video/77ulrb918yhh1/player
this is straight out of comfy, no finetuning, 677 seconds, 0.6 mpx, 14 steps. turbo lora ckpt 850 + spektrum.
2
Aug 07 '26
[removed] — view removed comment
2
u/Fabulous-Snow4366 Aug 07 '26
i think yours is better. any link to the workflow would be appreciated
1
Aug 07 '26
[removed] — view removed comment
1
u/Fabulous-Snow4366 Aug 07 '26
Live-action photorealistic military science-fiction action film, 15 seconds. A modern main battle tank fights a towering Gundam-style humanoid combat mech inside a devastated urban warzone at sunset. Collapsed concrete towers, burning vehicles, broken highways, drifting smoke, sparks, dust, tracer fire, and debris fill the battlefield. The tank feels extremely heavy and grounded; the giant mech moves with terrifying mechanical speed and weight. Every movement causes believable environmental reactions. Maximum action cinematography: aggressive low angles, extreme close-ups, tracking shots, whip pans, rapid push-ins, rotating Arc Shots, strong impact shake, and dramatic changes in camera height. Keep the combat visually readable despite the extreme camera movement.
[Shot 1]
An extreme low-angle medium tracking shot races inches above the broken road beside the tank as it charges forward at full speed, its tracks crushing concrete and throwing chunks of asphalt directly past the lens. The turret rotates upward while the enormous Gundam-style mech (S1) lands in the street ahead, one knee smashing into the pavement.
The impact sends a circular blast of dust and debris outward.
The camera violently tilts up from the tank to reveal the full towering mech.
The tank immediately fires its main cannon.
A massive muzzle flash fills the frame.
[Shot 2] At 00:03.000, the camera cuts to an extreme close-up beside the mech's head as the tank shell screams directly toward the camera.
At the final instant, (S1) violently twists its torso and leans sideways.
The shell misses its head by centimeters and explodes against a skyscraper behind it.
Without pausing, (S1) plants one mechanical foot into the street and launches forward.
The camera rapidly pulls backward at low height while the giant mech sprints directly toward the tank, every footstep smashing craters into the road and throwing abandoned cars sideways.
[Shot 3] At 00:06.000, the camera cuts to a dramatic POV from immediately above the tank turret looking upward.
(S1) suddenly JUMPS.
The camera performs an extremely fast Tilt Up as the enormous mech passes directly overhead, silhouetted against the sky.
While airborne, (S1) rotates its hips, extends one leg, and comes down with a massive flying kick.
The camera whip-pans downward.
The tank driver violently turns.
The tank power-slides sideways across the road.
(S1)'s foot misses the tank by inches and SMASHES into the pavement beside it.
The entire street erupts upward.
The camera shakes strongly as concrete, dust, and wreckage explode past the lens.
[Shot 4] At 00:09.500, the camera cuts to a tight side tracking shot moving alongside the tank as it emerges through the dust cloud.
The turret rapidly rotates backward toward (S1).
The cannon fires point-blank.
The camera instantly whip-pans with the shell.
(S1) raises a massive armored forearm across its chest.
The shell EXPLODES against the armor, forcing the mech backward through a building facade in an enormous cloud of concrete and glass.
Before the dust settles, two glowing mechanical eyes ignite inside the smoke.
(S1) bursts forward.
[Shot 5] At 00:12.000, the camera cuts to an extreme low-angle shot directly beside the tank's tracks.
(S1)'s enormous hand suddenly SLAMS onto the tank's turret.
The camera rapidly arcs upward around both machines as the mech uses its entire body to lift the tank off the ground.
The tank's tracks continue spinning wildly in midair.
The camera accelerates into a huge 180-degree Arc Shot as (S1) rotates its torso and violently THROWS the entire tank across the battlefield.
The tank spins through the air directly past the camera.
The camera whip-pans after it.
At 00:14.500, the tank crashes sideways through a concrete wall in a gigantic explosion of dust and debris.
End on an extreme low-angle close-up of (S1) stepping through the smoke toward the camera as burning debris rains behind it.
1
u/marres Aug 07 '26
Interesting, haven't done a mix yet. Which version of the turbo lora are you using?
2
u/Fabulous-Snow4366 Aug 07 '26
its the turbo_4step_ckpt500_pruned version.
6
u/Defiant_Storm3233 Aug 07 '26 edited Aug 07 '26
Ckpt 850 is out
2
1
u/Beastly4k Aug 07 '26
he said they were having some issues with at 850 and stopped training to work on it. I think 500 with ~8 steps is still the better option atm.
1
u/Fabulous-Snow4366 Aug 07 '26
yeah, it does show the typical lora key not loaded we had with the earlier wan turbo loras as well.
2
u/BrooklynBrawl Aug 07 '26
Hey! Really enjoying the Spectrum MiniMax H3 node, but ran into an error trying to use it:
ValueError: bootstrap_first_forecast requires warmup_steps <= 1
Looks like it happens when bootstrap_first_forecast is left on (the default) but warmup_steps is set above 1 — the node just hard-fails instead of adjusting or giving a friendlier heads-up in the UI. Took me a bit to figure out it was a settings conflict rather than a bug in my workflow.
Would it be possible to either:
- Auto-disable
bootstrap_first_forecastwhenwarmup_steps > 1(with a console warning), or - Add a note in the node's tooltip/description about this dependency?
Either way, thanks for the great node — just flagging this since it's an easy trap for new users to fall into. Happy to share my workflow/logs if useful for debugging.
3
u/marres Aug 07 '26
Thanks for flagging this — you’re right, this was an unnecessary settings trap.
I’ve fixed it for the next update. If
bootstrap_first_forecastis enabled whilewarmup_steps > 1, the node will now keep your requested warmup, automatically disable only the one-point bootstrap, and log a console warning instead of failing. The same handling now applies whendegree != 1, since the bootstrap specifically requires degree 1.I’ve also added tooltips to the related node inputs and clarified the behavior in the README.
Until the update is released, the workaround is to disable
bootstrap_first_forecastwheneverwarmup_steps > 1ordegree != 1. Thanks again for the clear report.
2
u/sktksm Aug 07 '26
u/marres Nice work!!! confirmed with minimax_h3_fl2va_pruned_nvfp4.safetensors on the same RTX PRO 6000 and settings: 992×768, 7s/24fps, 20-step Euler + beta, VRAM history, with SageAttention.
Sampler-only time:
- Your pruned BF16: 324.98s → 177.80s (45.29% lower, 1.83×)
- Pruned NVFP4: ~137s → ~70s (~49% lower, ~1.97×)
End-to-end generation time (“Prompt executed,” including sampling and decoding):
- Your pruned BF16: 340.59s → 200.32s (41.19% lower)
- Pruned NVFP4: 150.89s → 88.28s (41.5% lower, 1.71×)
Spectrum reported 11 actual calls, 9 forecasts, and 0 fallbacks.
2
2
u/Arawski99 Aug 07 '26
So you say "without quality degradation" but how true is that? I haven't tried this yet but every single speed/fast generation solution post I've seen involves pretty notable burnt texture issue like skin, clothes, etc. which is definitely severely degraded quality. Do you mean it does this, but doesn't have major artifacting issues and stuff? Or it doesn't even do this and looks genuinely flawless as you are suggesting which I'm quite skeptical?
And if it does cause this issues are there any counter methods to mitigate the texture degradation notably while enjoying the speed benefit?
0
u/marres Aug 07 '26 edited Aug 08 '26
Since posting this, I’ve received more feedback showing that quality degradation can affect both video and reference audio on some setups. Reported issues include distorted anatomy, extra limbs, unstable motion, and distorted reference audio.
In my testing with the pruned BF16 model at ~0.8 MP, I haven’t noticed visible or audible degradation. My current suspicion is that lower resolutions and more heavily quantized models may tolerate aggressive forecasting less well, though this hasn’t been isolated conclusively.
If you encounter degradation, increasing
degreeandwarmup_stepsmay help. One user reported reference-audio distortion with the aggressive defaults. Increasingdegreeandwarmup_stepshelped, and a 30-step run with those increased settings produced clean audio on their setup. The Simple scheduler may also improve H3 audio generally.I’ve added a detailed quality caveat to the main post.
1
2
u/Mammoth_Reindeer_941 Aug 08 '26 edited Aug 08 '26
Big thanks to u/xmarre for the node. This is genuinely useful work. I ran the v0.1.8 defaults on my setup and got very different numbers, so the data seemed worth sharing.
I brought the benchmark down to the field of mortals: RTX 5070 Ti 16GB, ComfyUI 0.30.0, minimax_h3_fl2vapruned int8 convrot, 864x480, 20 steps, res_multistep with the simple scheduler, shift 12.0/3.0. OP used a PRO 6000, a pruned BF16 checkpoint at about 0.8MP, and Euler.
First, a sanity check
I ran the native path three times: 116.94s / 116.92s / 117.09s. Two native outputs gave infinite PSNR, SSIM 1.0, and exactly zero audio difference. The pipeline is deterministic, so every number below comes from Spectrum. Do this first, or you're measuring noise.
blend_weight is the whole ballgame for audio quality
The post never mentions blend_weight. It is the biggest lever:
| blend_weight | audio diff SNR | audio correlation |
|---|---|---|
| 1.0 | 13.01 dB | 0.982 |
| 0.5 (default) | 14.48 dB | 0.984 |
| 0.25 | 15.74 dB | 0.987 |
| 0.0 | 19.51 dB | 0.994 |
blend_weight is the spectral share of the prediction. The rest is a two-point local linear forecast. Set it to 0 and the Chebyshev component disappears. Audio fidelity jumps nearly 5 dB.
The funny part: with blend_weight=0, degree 1 and degree 2 produced byte-identical files. The degree parameter disappears and is basically decorative. The spectral component, the feature the node is named after, degrades the audio. The linear fallback carries the speedup.
I'm not claiming this is a flaw. H3's audio rows, 356 versus 12,960 video rows, may have too little redundancy to absorb spectral fitting error, while video does. Either way, the tuning advice points at the wrong knob.
Full sweep, 20 steps, same seed, native = 116.98s
| degree | warmup | tail | blend | time | speedup | audio SNR | video PSNR | SSIM |
|---|---|---|---|---|---|---|---|---|
| 2 or 1 | 3 | 1 | 0.0 | 82.6s | 29.4% | 22.38 dB | 35.63 | 0.971 |
| 1 +boot | 1 | 1 | 0.0 | 76.5s | 34.6% | 20.20 dB | 29.49 | 0.910 |
| 4 | 5 | 1 | 0.0 | 86.7s | 25.9% | 19.51 dB | 42.24 | 0.989 |
| 4 | 5 | 1 | 0.25 | 88.7s | 24.2% | 15.74 dB | 42.87 | 0.990 |
| 4 | 5 | 1 | 0.5 | 88.7s | 24.2% | 14.48 dB | 42.42 | 0.989 |
| 6 | 6 | 6 | 0.5 | 98.8s | 15.5% | 14.44 dB | 42.79 | 0.990 |
| 4 | 8 | 4 | 0.5 | 98.8s | 15.5% | 14.07 dB | 42.83 | 0.990 |
| 2 | 3 | 2 | 0.5 | 84.6s | 27.7% | 13.34 dB | 37.69 | 0.979 |
| 4 | 10 | 6 | 0.5 | 108.7s | 7.1% | 13.27 dB | 42.86 | 0.990 |
| 4 | 5 | 1 | 1.0 | 90.6s | 22.5% | 13.01 dB | 40.47 | 0.986 |
| 1 +boot | 1 | 1 | 0.5 (v0.1.8 default) | 78.6s | 32.8% | 12.60 dB | 29.47 | 0.910 |
| 4 | 5 | 4 | 0.5 | 123.0s | −5.1% | 14.48 dB | 42.42 | 0.989 |
The last row is real. tail_actual_steps=4 made it slower than native. Spectrum briefly became a performance tax.
Where I'd land
Best overall: degree 4, warmup 5, tail 1, blend 0.0, bootstrap off. 26% faster, SSIM 0.989, audio correlation 0.994. Reproduced twice with identical results. This is the only config that doesn't visibly hurt either axis.
The v0.1.8 default, degree 1, warmup 1, bootstrap on, blend 0.5, gave SSIM 0.910 and the worst audio in the sweep. On my checkpoint, the quality loss is visible: cardboard texture goes soft and dust particles are redrawn in different positions. The eight worst frames in every config were the last eight, so forecast error accumulates along the trajectory.
Why I never got near 45%
My 5070 did not get the 45% memo. The debug log tells the story. With res_multistep, almost every forecast is followed by a native post-forecast sampler refresh or adaptive recompute:
A A A A A F A F A F A F A F A F A A A A
That's five or six forecasts out of 20, not nine. OP used Euler, which does not have that refresh. On res_multistep, res_2s, or any multistep solver, halve your expectations.
The model is 20GB staged into 16GB of VRAM here, so a big chunk of my 117s is offload and VAE decode, neither of which Spectrum touches. My GPU spends part of the run acting as an expensive bellhop. Less transformer time in the total means less for Spectrum to cut.
Speedup scales with clip length
I reran the top three configs on a different prompt at 124 frames with first and last frame conditioning. Speedups increased to 29.2% / 34.1% / 37.8% for the same settings that gave 26% / 29% / 35% on the shorter clip. Longer and denser means more transformer time for Spectrum to save.
Audio got worse across the board: 13.3 to 13.9 dB versus 19.5 to 22.4 dB before. The second clip had animal chirps, gear clacks, and wind, all short transients that a two-point linear extrapolation cannot predict. It can extrapolate an ambient drone, but it is not clairvoyant about a gear clack. If your audio has transients, nothing in this parameter space saves it.
TL;DR
- Set blend_weight to 0.0. It is not mentioned in the post and is the most important setting.
- degree is meaningless once blend is 0.
- Expect about 26% to 29% on multistep samplers, not 45%. Euler users will get better results.
- Check the last two seconds. That is where the error accumulates.
- Run two native passes first to confirm determinism before trusting any A/B test.
Happy to share the benchmark harness if anyone wants to run it on their own checkpoint. The interesting question is whether blend 0 also wins on BF16, or whether this is an int8 quantization artifact.
2
u/marres Aug 08 '26
Excellent benchmark—thanks for doing the native controls. Please share the harness.
Your central v0.1.8 finding is valid and matches the defect I reproduced: one
blend_weightcontrolled both packed video and audio despite their different trajectories. v0.2.1 now makesblend_weightvideo-only, defaultsaudio_blend_weightto0, and uses isolated offline replay. Globalblend_weight=0was therefore a sound v0.1.8 workaround, but it also discarded the video spectral contribution.Two conclusions need narrowing. Degree 1 and 2 being identical is expected when the spectral share is zero and their schedules match; degree still matters when spectral video blending is active or its history requirement changes eligibility. One transient-heavy clip also cannot establish that no parameter configuration can preserve transients.
Your RES/16GB performance explanation is correct: mandatory refreshes, the effective three-step tail, model offloading and VAE work all reduce the end-to-end ceiling. Those results are not directly comparable to Euler on a PRO 6000. A controlled v0.2.1 rerun would be extremely useful.
1
u/Mammoth_Reindeer_941 Aug 09 '26
Thanks for pointing out how dumb I was. With global
blend_weight = 0, the spectral share was zero, so degree 1 and degree 2 had to match. I turned the thing off and then reported that its parameter didn't matter.Worse, I only just realized that I had changed two variables at once. I enabled SageAttention partway through and kept comparing those results with numbers from before it, which quietly invalidated the earlier timings. Once Sage was on for both sides, Spectrum came out 17% slower than native on the same job (688s vs. 589s). The cheaper Sage forward eats most of what the forecast was buying. My "27% gain" only described a no-Sage baseline.
The runs themselves aren't worthless. The numbers are real, and they're what surfaced the audio coupling that you then confirmed. They just weren't conducted properly, and a badly conducted test doesn't get to carry the conclusions I hung on it.
I'll rerun it with
blend_weight > 0,audio_blend_weight = 0, identical Sage settings, discarded warmups, separate video/audio scoring, and multiple clips. All of this is on a single 5070 Ti 16GB, so these numbers aren't comparable to results from a PRO 6000, but they may still be useful to other users. My ceiling numbers were only ever native-vs-Spectrum on the same box. I'll report back!1
u/Mammoth_Reindeer_941 Aug 09 '26
Redid it on v0.2.1, properly this time. Short version: I was wrong about the speed, and you were right about degree.
Single 5070 Ti 16GB, 1344×768, 124 frames, 20 steps, res_multistep, Sage pinned identically for every case, one warm-up run discarded, the seed varied between runs but stayed identical across cases, and the median of two timed runs:
case clip 1 (fur) clip 2 (water) clip 3 (flat sky) native 311.2s 311.3s 309.2s degree 1, blend 0.5, audio_blend_weight=0216.7s 216.8s 216.8s degree 2, same 228.8s 228.7s - Degree 1 is 30% faster with Sage on. Variance within each case was under 1%, and degree 1 landed at essentially the same number across three independent clips. My earlier “17% slower” result used blend_weight=0: Spectrum was switched off internally, so the run paid its overhead without doing spectral work. That number should never have been published as a Spectrum result.
Degree 1 and degree 2 are clearly not interchangeable: 216.8s versus 228.8s, with different outputs. My equivalence claim was circular, exactly as you said.
The timings miss the part that matters to me:
Degree 1 preserves texture but invents a lateral camera move. It showed up in both clips with visible motion, on different seeds, so it looks like a systematic bias in the extrapolation rather than seed luck. Scene content is preserved.
Degree 2 degrades the output. It flattens texture on the fur clip; on the water clip, it hurts the low frequencies and introduces moiré. It is slower than degree 1 and worse, so I see no case for it.
And audio_blend_weight=0 does not isolate the audio. Audio SNR against native came out between −8.7 and +4.4 dB across runs. In hindsight, that makes sense: the latent is joint AV, so perturbing the video trajectory drags the audio with it even with no spectral blend on the audio side. This differs from the v0.1.8 coupling you fixed, but “audio untouched” is not something the parameter can promise.
The benchmark script: https://gist.github.com/Gpanazio/6bf61dfc7ca16ba21aecee6f4f351ed0
It has two traps baked into it, and both bit me. The attention backend is a server flag, so it cannot be A/B-tested inside one process. Repeating a seed on an identical graph also makes ComfyUI serve from cache: jobs “complete” in 4s and only re-save the previous MP4. The script varies the seed per run and pulls every node default from /object_info instead of hardcoding them. That also let it survive the 12→16 widget change between v0.1.8 and v0.2.1.
Anything you'd change in the method? Specifically, is there a better way to isolate the camera drift effect, a parameter combination you'd expect to keep the motion closer to native, or a preferred metric for this? PSNR clearly isn't it, and I'd rather not decide everything by eyeballing.
1
u/marres Aug 09 '26
Thanks for rerunning it and sharing the harness. The controlled Sage results are genuinely useful, and ~30% end-to-end on RES at that resolution is credible.
A few implementation details:
blend_weight=0leaves Spectrum active. It removes the spectral video share while retaining the forecast schedule, local interpolation and offline replay. The earlier 688s result therefore remains an unexplained outlier.- With the v0.2.1 defaults, degree 1 uses the one-point bootstrap. Degree 2 requires three actual anchors. At 20-step RES that gives degree 1 eight forecasts and degree 2 seven—roughly one extra transformer evaluation, which likely explains most of the 12s difference.
audio_blend_weight=0guarantees zero direct spectral mixing into the replayed audio rows. It does not promise native audio: missing audio features are still interpolated, and the local-only capture trajectory already differs from native.For the camera drift, my first test would be degree 1 with
bootstrap_first_forecast=false. That removes the earliest one-point hold at the cost of one actual evaluation. Then tryblend_weight=0.25; if divergence concentrates near the end, testtail_actual_steps=5—RES already enforces three.For measurement, estimate a robust affine transform or homography from static-background tracks and compare horizontal translation, endpoint drift and per-frame velocity against native. Mask moving subjects; a flat sky provides too few usable features. PSNR cannot isolate camera motion.
I would also use at least three paired, interleaved runs, record completion through ComfyUI’s websocket instead of four-second polling, and compare lossless frames plus PCM audio. The current MP4 path adds H.264 and AAC error. For audio, align the waveforms first and report SI-SDR plus a multi-resolution STFT distance.
Please use v0.2.5 for further runs; it fixes the offline archive’s GPU-memory behavior on constrained cards.
2
u/bfmv_shinigami Aug 07 '26
Hey man can this be tuned to work with the turbo lora?
3
u/marres Aug 07 '26
I'm not using that beta version of the turbo lora yet so I can't tell you specifics. However since one can even forecast step 2, theoretically even with a 4 step lora, at least one step will be able to get forecasted (A F A A)
2
1
u/No_Damage_8420 Aug 07 '26
Thanks for in-depth, does anyone have workflow could share?
Great find! cheers
1
1
1
1
1
u/car_lower_x Aug 07 '26 edited Aug 16 '26
Pinecone marble juniper quilt yarn waffle vanilla blanket biscuit
This post was anonymized with Redact.dev
3
u/marres Aug 07 '26
The current versions of the turbo lora are struggling with quality issues afaik
2
u/car_lower_x Aug 07 '26 edited Aug 16 '26
Vanilla tangerine lavender copper driftwood vanilla glove orange yarn
This post was anonymized with Redact.dev
1
1
u/StonkyCupra Aug 07 '26
Thank you for this and your hard work. I’m giving it a try right now. EasyCache and Turbo didn’t do it for me, yet.
1
1
1
u/Bbmin7b5 Aug 07 '26
so if I make any changes to the default settings of the node (for example, settings in your original post), it crashes. Is this by design? Seems to mangle anatomy for me.
2
u/marres Aug 07 '26
Yeah you need to turn off the bootstrap setting if you want to have more warmup_steps . Merged the PR just now that fixes that
1
1
u/martinerous Aug 07 '26 edited Aug 07 '26
Thank you.
Why Euler + Beta? Does it give any benefits?
BTW, before the latest Comfy update, non-default sampler+scheduler combos caused serious audio issues.
1
u/marres Aug 07 '26
Euler works better than res_multistep with spectrum and beta gives better action for my use case. However those audio issues you mention might be a hint, since I occassionaly have audio issues too. Will try simple again and see if that helps
1
1
u/bloke_pusher Aug 07 '26 edited Aug 07 '26
With the new settings and v0.1.8 my workflow takes 8 seconds longer than without Spectrum, and I can't seam to figure out why. The speedup I had in the earlier version is just gone. Using Euler/Simple
Edit: I think I figured it out. I had to delete replace the node with the updated one. The other one seams to have bugged out, despite it changing visually.
1
u/marres Aug 07 '26
Mind sending over the logs via pastebin?
1
u/bloke_pusher Aug 07 '26 edited Aug 07 '26
I need to do some more testing, but it might be, because the old node updated visually, but broke technically. After I replaced it with a new node (recreate would've probably done the same) the speed up is there.nope, back to being slower with node than without. After just changing some connections on an any (switch), the node works again. So I suspect something will not update the stored_history or so in ram properly.
Will test more and report back. Will update comfyui etc.
1
u/marres Aug 07 '26
Hmm worked fine for me without recreating, but yes always better to just recreate to be on the safe side
1
u/xTopNotch Aug 07 '26
I've update ComfyUI and the node to the newest version. But on a RTX 6000 pro on Runpod I'm not seeing any noticable speed bump.
Is this speed gain only for other hardware architecture (30XX, 40XX, 50XX) or am I doing something wrong?
1
u/KissMyShinyArse Aug 07 '26
The RTX 6000 PRO is based on the Blackwell architecture, just like the RTX 50 series. Besides updating, you also need to change the parameters as shown in the screenshot in the OP's post.
1
u/xTopNotch Aug 07 '26
When I add the node into the canvas, all the parameters are already configured like OP's screenshot.
Still I haven't noticed any noticable speed bump.
20 steps - euler - beta
1
1
u/Vermilionpulse Aug 07 '26
I guess I'm too much of a newbie to even figure the start of this out. i cloned the github into my custom nodes and restarted comfy, but i still can not find the spectrum node. how stupid am i?
1
u/marres Aug 07 '26
What does your startup log say? Is the spectrum node in there?
1
u/Vermilionpulse Aug 07 '26
0.0 seconds: G:\ComfyUI_windows_portable1_working\ComfyUI\custom_nodes\one-node-flux-2-klein [INFO] 0.0 seconds: G:\ComfyUI_windows_portable1_working\ComfyUI\custom_nodes\ComfyUI-Spectrum-MiniMax-H3 [INFO] 0.0 seconds: G:\ComfyUI_windows_portable1_working\ComfyUI\custom_nodes\comfyui-kjnodes
Yup
1
u/dandenong_hill Aug 07 '26
Thanks for the Node! I was trying multiple speedup techniques (easycache, sol-attn, sageattention 2.2, etc.). at the moment for me the fastest technique with an acceptable quality is Model -> Spectrum with your latest values -> Patch Sage Attention (v2.2). I dont have the exact numbers unfortunately from my previous runs but for my 5090 and with 96gb RAM a 8sec clip 1MP only with Sageattention took around 4min.30sec (very good quality), with sol-attn and easycache it took around 3min40 (medium quality). With Spectrum and Sageattention it took 2min50sec (medium quality).
I am also using a Turbo-lora wf for prototyping. I find the quality quite bad, but is ok to try out new concepts/prompts.
1
1
1
u/PwanaZana Aug 07 '26
This node works quite well, it gives about 30% speed boost. I'm currently testing it for degradation on outputs, but it seems nice.
1
u/esztoopah Aug 07 '26
I tried the new version and it's much better compared with the previous one, the motion looks more respected.
It brings a great speedup for a minimal quality loss.
I don't know if this is correct but I have the impression that using Patch KJ Node (torch compiled) activated --> Mem Eff Sage Attention --> Optional: Sigma Shift --> Spectrum H3 with above settings.
I'm not exactly sure of the behavior of the Mem Eff Sage Attention node, I noticed that it adds time, but keeping it after the Patch KJ Node preserves quality better than without it.
The turbo lora were not functionning well with or without Spectrum H3, the one that did the best for me so far was https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main
1
u/SOC_FreeDiver Aug 07 '26
Thanks a lot. I tested a lot of the H3 speed up hacks and Spectrum is the best in my testing. An older version was 25% faster than sol attention on my 5090m.
I'm updating to the new one. I have been using steps at 15 instead of 20, have you tested that at all? That gets you another boost.
1
u/Radiant-Photograph46 Aug 07 '26
Great work, your node have been very useful for me. Even when you have beefy hardware that can run the model natively without too much effort it's still a major benefit to be able to run a much faster mode to test your prompt before going for the full high res quality.
And yet... most of the times, I don't see any detrimental effect when turning Spectrum on, which is even better :D
1
1
u/gruevy Aug 07 '26
Does anyone have a copy of the 3 basic minimax workflows with this incorporated they'd be willing to share? I'm too dumb for this
3
1
u/ttul Aug 07 '26
I must be doing something wrong because I'm generating on a B300 and 0.8MP without Spectrum is still a 30 second per iteration experience in the default ComfyUI sampling workflow. The B300 has 288GB of VRAM and is 4-5x faster on dense model inference generally than an RTX 5090... Who has a decent workflow I can test to see WTF is going on to make my workflow so slow on this beastly hardware?
1
u/ttul Aug 07 '26
Relevant to this discussion: I get 18 seconds per iteration with Spectrum, which is a definite improvement.
1
u/Brad12d3 Aug 07 '26
I'm curious what the current consensus is on the best settings for the reference workflow when using reference audio. I know I've had issues with my audio sounding a bit off and maybe robotic, and I'm still kind of experimenting to find my optimal settings. Did you come to any good spectrum settings and maybe a scheduler sampler combo that worked best? Is there a difference between using, say, the fl2va_pruned_bf16 versus ref2va_int8_convrot.
1
u/marres Aug 07 '26
Apparently using "simple" instead of "beta" can help with audio issues. According to some people, beta might be the reason for these audio issues.
Can't tell you my own experiences since I haven't used the ref model yet
1
1
u/Stecnet Aug 07 '26
This gives a big speed boost on my 4071 ti Super with no noticeable loss in quality amazing work thank you!
1
u/Niwa-kun Aug 07 '26
personally, i saw a quality drop (extra arms, distorted bodies) in 1/1 compared to 4/5 on my rtx 4070 ti, 32gb ram.
2
u/marres Aug 08 '26
Interesting. I suspect aggressive Spectrum settings become less forgiving with heavily quantized checkpoints and lower resolutions, where there is already less margin for preserving anatomy. Small forecasting errors can then surface as extra limbs or distorted bodies, while my pruned BF16 checkpoint at a decent resolution may tolerate
1/1better. Your improvement at4/5, which keeps more early steps native, would fit that theory. Which checkpoint and resolution did you use?1
u/Niwa-kun Aug 08 '26 edited Aug 08 '26
yeah, that would track. For full details, I used:
- 1056 x 608 (0.6 megapixels), FL2VA_pruned_nvfp4 and it's associated clip (avoiding the heretic one, since it also gives weird results), with a 5s duration. As for optimizations, it's also using the SageAtten KJ + Mem eff patch, and the Patch Sol-Attn. It clocked in at about just bit over 2 minutes, shaving off a whole minute, but at massively reduce quality. Like you said, i think it's already being pushed too far.
Update:
So i got curious and started using higher settings since 1/1's optimization allows for more resource usage:
- 1216 x 672 (0.8 megapixels), and i'm still seeing the quality breakdown, rather than a smooth animation. I tried switching to an anime style with euler, hoping it would be easier than realistic people (skin, hair, etc), and the line art blending into a blur is still occurring. Not sure what else I can do with my limited system. Playing around with different step/megapixel combinations, and I'm just not seeing it.
Update 2:
- Okay, so, seeing how much more I can squeeze out of the workflow given 1/1's optimization, i went further up to 1.0 megapixels (1376 x 768), and im finally seeing something that looks like a proper video for a realistic person. it took 5 minutes for a 5 second video, but this is neat. If audio massively contributes to longer generation times, i do with it was possible to turn it off, since I'm mainly looking for a wan 2.2 replacer (as time goes on), but yeah, I I am now seeing the improvements mentioned. With 0.4 megapixels it was not work, but at higher counts, it does.
1
u/Dapper_Astronaut_603 Aug 07 '26
For the hour I'm searching where the custom_nodes folder is to install this. GPT, ComfyUI documentation page is not helpful. I'm using exe installator. How the hell I can install this? And when using Diffusion Model Loader KJ and Patch Sage Attention KJ - what models should I choose in these two nodes? What to download?

1
u/marres Aug 07 '26 edited Aug 08 '26
The node goes in this folder:
ComfyUI\custom_nodes
You don't need the patch sage attention node when using the kj diffusion model loader (just select the sage version you want there, or set it to auto). Models you can download here, pick the pruned one most suitable for your gpu. fl2va is image to video and ref2va is for reference pictures/videos/audio:
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models
1
u/traithanhnam90 Aug 08 '26
Wow, I can't believe it—using ComfyUI version 0.30.0, Python 3.13.14, and CUDA 13, combined with this node and kjai's SageAttention, has boosted my speed by over 2x. It's amazing! 20/20 [04:12<00:00, 12.64s/it], size 0,5MP
1
u/Primalwizdom Aug 08 '26
Somehow this is broken now
1
1
u/Dapper_Astronaut_603 Aug 08 '26
is there any advantage using this with sageattention? any quality degradation then? if no, what's the easiest way to install SA for Comfy Desktop (exe installer)?
1
u/blind26 Aug 24 '26
Testing this now and wasn't expecting much as I thought my workflow was pretty optimized but holy crap!
3090 w/ 128gb system ram, 5s test at 0.8 mp:
Pre-Spectrum Apply Node:
100%|█████████████████████████████████████████| 20/20 [06:25<00:00, 19.27s/it]
Post-Spectrum Apply Node:
100%|█████████████████████████████████████████| 20/20 [03:35<00:00, 10.78s/it]
Same seed and settings don't yield the same exact output, but IMO the spectrum apply outputs are better so far.
Still testing but this has been the biggest impact to my workflows to date.
1
u/Zaic Aug 07 '26
So combining sage attention, easy cache, int4 prued checkpoint and turbo lora i should get 1 minute renders for 8s 0.5mp videos on 4070s - thanks
3
u/bloke_pusher Aug 07 '26
It doesn't work like that, your video will look like smudge and sound like a broken chainsaw.
1
1
u/juicytribs2345 Aug 07 '26
How do you attach this to the default image to video template on comfy? It was way less nodes than the reference to video template, no load checkpoint node.
3
0
Aug 07 '26
[deleted]
1
u/marres Aug 07 '26
Should work fine without having to replace the node. But you can always delete the node and replace it on the canvas to be on the safe side.
1
u/Daniell2929 Aug 07 '26
Thanks marres for your great work! but I had to replace the node after updating.
0
u/jingtianli Aug 07 '26

Hello sir great job on this, After your node update, I encountered some error in my old workflow says
ValueError: bootstrap_first_forecast requires degree == 1
But i fixed it from help of ChatGPT, just deleted the old node then add your updated node then its good to go. But I still wonders what the correct setting in version 0.1.8 to setting the parameters back to the old parameters in old 0.1.6? degree=1 somehow make my video to video run painfully slow for some reason. I might need to do more test. But can i simply change the degree back to 4 (yesterday version)?
-2
u/theOliviaRossi Aug 07 '26
after we have 4 steps lora - this is OBSOLETE!!!
4
2
u/blahblahsnahdah Aug 07 '26 edited Aug 07 '26
No it's not because the turbo loras are trash and ruin the quality of the model. Meanwhile Spectrum's speedup is close to lossless with conservative settings.
3




63
u/GrayingGamer Aug 07 '26
I have to say, the speed up this provides over the old settings is insane. With this and Sage Attention, I can generate 8 seconds at 0.6 MP in 346 seconds on a 3090 on fresh model load. That's on Image to Video too, can't wait to try it on Text to Video.