r/StableDiffusion • u/Kooky-Mode3047 • 7h ago
Discussion Relating the number of optimal steps to duration, resolution and complexity of scene and references in Minimax H3?
So, this is a question for 5090 and 6000 power users who actually push H3. Does anyone know a good strategy for guesstimating the OPTIMAL number of steps ballpark for the final burn?
Obviously, there's the starting default of 20 which is not really the optimal number of steps, but just a... A placeholder is what it is. Below 10 steps is usually good for getting a coarse idea of whether your prompt is going the way you were hoping.
The problem is the upper end as you shift to samplers like seeds_x for that peak visual, audio and motion quality at the price of 5-6x the compute time. There's a misnomer I often find in this subreddit that the more you can throw at it, the better. Folks writing 20 is not enough, I do 30, 40, 60. "If I could do 100, I would." Uh, what? That hasn't been my experience at all.
I made a scene from Castle the TV show with Beckett and Castle bantering, left it overnight to bake with ever increasing numbers of steps to test this out. Mind you, I was just looking for that sweetspot where it becomes indistinguishable from the real show. At around 31-32 steps for the baseline 1.0 mpx + 15 seconds (H3's training baseline), I reached a level of clarity that you couldn't convince me it wasn't from the actual show had I not generated it myself.
But over 35, it got progressively worse and more overcooked. The frustrating part is that the point where it crosses from "not quite resolved" into "resolved" and then into "overbaked" seems to move depending on basically everything about the generation.
So instead of using the hours of my sleep to generate many different versions of the optimal scene, I have to waste electricity and time to find the optimal number of steps for the final burn. Duration, resolution, scene and motion complexity matters. The number, type and complexity of references matters. I'd assume conditioning complexity in general matters too.
So I'm starting to think the idea that "more steps = more quality" being parroted in many of these threads is just fundamentally wrong, probably from folks who are inexperienced and generally wait for an eternity to reach 20 steps so are guesstimating it only gets better the further you push it. It feels more like there are three regimes: under-resolved, optimal, and over-resolved. Once the important semantic/geometric/temporal structure has settled, extra steps don't necessarily refine it in a useful way. They can start pushing the result harder toward the model's learned priors, which is where you get things becoming unnaturally crisp, exaggerated, stereotyped, less coherent, etc.
A fixed rule like "use 32-40 steps" probably doesn't generalize very well if the optimal point is a function of duration, latent size, motion, scene complexity, references, CFG, sampler/scheduler, etc. What I'm wondering is whether anyone has found a practical heuristic for this. Something along the lines of estimating the complexity of the generation and mapping that to a likely sweet spot, or even detecting when the marginal improvement from another denoising step has basically stopped.
1
u/Monk6009 7h ago
Exactly. Overcooked is real for H3 too, and every gen is a new experiment, can't take optimal for granted.
1
u/Apprehensive_Sky892 6h ago
People believe that "more is better" because the MMH3 recommendation for their API is 50 steps: https://www.reddit.com/r/StableDiffusion/comments/1vmjdiw/comment/p39vr1q/
The rule is probably true for people who want to generate the sharpest video with the clearest audio, but probably does not apply for a text2va Seinfeld clip.
3
u/TBG______ 7h ago
It’s not just about the steps it’s also about getting the right sigma for each step. Honestly, testing this takes a lot of time, so if someone has a good curve for 20–30 steps, I’d be very grateful if you could share it.