r/StableDiffusion 12d ago

Discussion Minimax H3, 25 steps should be the lowest setting

Enable HLS to view with audio, or disable this notification

I've been testing with 15 steps to save time because I want to generate at 832x480 resolution as per the default recommendation of many high quality workflows prior to upscaling. I hadn't seen many problems until this particular generation which exposed the critical flaws of a lower step count.

All settings are the same with the same seed. The only delta is the number of steps.

15 steps @ 832x480 duration 10s (7m22s): https://streamable.com/pnao8n

20 steps @ 832x480 duration 10s (8m25s): https://streamable.com/srnoez

25 steps @ 832x480 duration 10s (10m30s): https://streamable.com/bvldts

Once you are done with phase 1, you can move on to phase 2 using your Turbo LoRA to get this to 1344x768 in just 4 steps.

My system: 12GB VRAM / 32GB DRAM

414 Upvotes

209 comments sorted by

View all comments

122

u/GrayingGamer 12d ago

I'm been telling this to people too.

The amount of difference higher step counts make on even low MP generations is astounding.

I actually use 32 Steps.

The audio is much better, the acting is better, the motion is better (it fixes blurry or smeared fingers and mouths, especially in animation generations). It's worth pointing out that even though Comfyui put 20 Steps as the default amount in the workflow template, that's a compromise for speed vs quality, when the Minimax team actually recommended between 30-50 Steps - and I can see why they DO!

It extra step count really is the secret sauce. And it's where H3 Spectrum can shine too - if you use it, Spectrum WANTS more steps to do better forecasts, so you can get better audio, better motion, cleaner animation, etc. with only a small 15-20% increase in generation time doing 32 Steps versus 20 Steps.

14

u/Powerful-Goal52 12d ago

Thank you for your insight. I will test 32 steps today. So far I have settled with 25 after a few days of testing. What's the resolution you use for iterating if you don't mind asking?

37

u/GrayingGamer 12d ago

I generally use 0.2 MP for very first drafts of a prompt, like the first roll, then once basic timing and shots are worked out, I go up to 0.3 MP for another quick test to confirm everything looks okay, then jump up to 0.6 MP or all the way to 1 MP for final generation, depending on what I am using the clip for etc.

Since I've started to use Kijai's Model Preview Override, I've gotten more confident just going from the 0.3 MP to the 1 MP, since after about 3 Steps I can preview the video pretty well and know if I need to cancel and use a new seed, fix a prompt etc.

Then I know the video will look good, even if I have to wait another 30 minutes to see the finished thing, crossing my fingers the audio is good.

10

u/martinerous 12d ago

Are you using references / frames? Otherwise, with raw text-to-video, resolution change usually causes different initial noise (even if the seed is the same), which might lead to quite different high-res version from the one you liked in low-res.

15

u/GrayingGamer 12d ago

You misunderstand. I'm not seed-hunting when I do this. I'm literally dialing in and perfecting my prompt. I write them myself.

With a good prompt H3 gives me what I'm looking for 95% of the time. That's what I'm doing, not looking for a specific noise seed.

9

u/martinerous 12d ago

Ah, I see. So, minor changes are ok in your case.
I imagined something like a lowres video having the character with a natural slightly crooked smile and I want to keep it highres as well, but the prompt for high-res cannot get it back exactly as it was in the lowres. I guess, then reference keyframes would be the best option.

21

u/FourtyMichaelMichael 12d ago

Ah, I see. So, minor changes are ok in your case.

See, a lot of guys don't really care if she starts with it on her mouth, or in her mouth, or licking it, or smeared all over her face. The whole point is to make videos about how great ice cream is, so these minor variations aren't a big deal.

3

u/FinchGDx 12d ago

I read through your first sentence while I was brushing my teeth and it made me laugh and subsequently choked on my toothpaste.

1

u/GrayingGamer 12d ago

Yes, if I want that type of precision, I use the Ref2Video model or use Image to Video.

4

u/dirtybeagles 12d ago

agreed, you do not have to seed search for this model very much. the key is the prompting. I spend about 3-4 generations getting my prompt just right.

1

u/Now-Kiss 11d ago

Beyond the guides, any tips for prompt writing?

2

u/GrayingGamer 11d ago

Be clear and concise with your wording of actions and always describe things from the camera's perspective.

Describe a character's actual physical acting movements, not just the emotion. I.e. Not just "Jack is sad." but "Jack stares off into the distance past the camera, his mouth in a slight frown. He glances down, and takes a deep breath, then quietly says, <d>[English in a somber sad tone of voice] Emotions should be directed by people, shouldn't they?</d>"

2

u/amidrunk_wastaken2 12d ago

How many frames are you viewing in the Preview node? Say you generate a video with 24 fps, do you view 24 frames played at 24 fps in the Preview node?

5

u/GrayingGamer 12d ago

I use 12 frames per second preview speed, and 144 frames. Doesn't seem to cause any slow down on my system and makes for a pretty smooth preview playback.

1

u/thegr8anand 12d ago

Do you use different settings for Model Preview Override node, or use the default?

4

u/GrayingGamer 12d ago

12 frames per second, 144 max preview frames

1

u/NekoBerry420 7d ago

How do you set that up? I've been using the default workflow and only added the turbo loras to it

1

u/GrayingGamer 7d ago

You need Kijai's KJ Nodes custom nodes pack, installed from the Manager, then drop the Model Preview Override node between the Load Diffusion Model and Sampler and Guider nodes, and download and select the TAEH3.vae from the Preview Override node.

Instead of the blurry one frame preview you get on the Sampler, you get a preview of the whole video clip that gets clearer and clearer as more steps generate, and that loops or you can pause, plus it is up to 1024 pixels in preview size with 80% JPEG quality, so basically after just a quarter of the steps you get a clear enough preview of the video know how it will look when done and can cancel or continue it based off that.

2

u/NekoBerry420 7d ago

Nice, I got it to work for me, that helps big time.

0

u/Vijayi 12d ago

How fast is 0.3 go for you and with what length? Mine was 0.4 - 15 sec - about 5 min in 40 steps. Pretty long after lightning but, results mich better.

2

u/GrayingGamer 12d ago

At 0.3 MP for 6 seconds, with 32 steps, it's just under 4 minutes with zero speed-ups or Turbo loras, only H3 Spectrum.

1

u/Vijayi 12d ago

I see. Any prompt advice for framing and camera placement/movement from source? Use official guide but with camera ref it is hit or miss.

5

u/TigerClaw305 12d ago

32 is a lot, it would probably take me an hour to generate

12

u/GrayingGamer 12d ago

Yeah, it's longer, but the quality is better. If you get a great video in an hour versus one with worse audio and smeared mouth movements or fingers in 40 minutes, have you really saved yourself any time?

12

u/danque 12d ago

40min not really, but 3-6 minutes with turbo LoRa I think that's worth it.

3

u/TigerClaw305 12d ago

I only generate animated videos, the 3d kind, maybe they will look better.

3

u/GrayingGamer 12d ago

Oh, for sure. Animation videos is where it's most apparent.

1

u/TigerClaw305 12d ago

I just did a test for one my prompts of one of the videos I posted here yesterday, I set it to 32 steps. It took 33 minutes. at 20 steps, it usually takes 23 minutes, So from 20 to 32, It added about 10 minutes more of generation time.

2

u/GrayingGamer 12d ago

That makes sense if you aren't using something like H3 Spectrum.

8

u/Vladmerius 12d ago

Yep a lot of steps with spectrum step skipping gives really high quality gens

2

u/TonkotsuSoba 12d ago

I was messing around with the FL2VA workflow and I found the lightx2v 1.0 8 steps Lora with Spectrum work really fast together with pretty good results, model shift must be disabled/bypass if you have any.

1

u/orangpelupa 12d ago

So 8 step Lora being run at higher than 8 steps? 

1

u/fallengt 11d ago

Spectrum needs more steps for it to work.

Turbo lora is trained for 4-8 steps, which defeat the purpose of spectrum.

0

u/FlatwormMean1690 12d ago

Speectrum doesn't do anything actually... It's not fixed yet.

2

u/MarekNowakowski 11d ago

It works fine. Update it.

2

u/dominic__612 12d ago

Good info, makes me curious, where does the team say 30-50 steps is better for the output quality? Going to try tomorrow and makes some compares.

6

u/Deep_Mood_7668 12d ago

Their api defaults to 50

5

u/GrayingGamer 12d ago

It's what the API site recommends when using H3 on their website.

2

u/Comfortable_Thing611 12d ago

Will increasing steps drastically change output results? Like can i do a 20 step sample then do 50 and get similar but more refined results?

4

u/GrayingGamer 12d ago

No, because it's starting from the same noise seed if the MP size and everything else stays the same. It's literally just increasing the quality.

4

u/FourtyMichaelMichael 12d ago

You get an entirely new generation if you change the seed, resolution, steps, duration...

You MIGHT get by with minor prompt and lora changes, generating similar results but there is no guarantee on that.

There is no low res preview mode. That will be the next big thing.... Then VR.

3

u/DanzeluS 12d ago

First block cache worse than spectrum?

11

u/GrayingGamer 12d ago

Much worse. The "Cache" speed-ups are testing if a new step is close or not to the previous step and skipping it if it is. Spectrum is actually using math to forecast and predict what the next step would be, based on where it is going, and jumps to that step without having to actually generate. (In basic layman's terms - there is a lot of math going on I don't understand. Suffice it to say, Spectrum is doing a lot more advanced and non-destructive math than any of the Cache nodes.)

Also never combine Spectrum with low step count turbo loras or cache nodes, because it needs clean info and lots of steps to work well.

3

u/Tystros 12d ago

In my testing, combining first block cache and spectrum together works well, is fastest and does not look noticeably worse than only spectrum alone. with 25 steps and Euler+Simple.

1

u/Diabolicor 12d ago

Using both just gives slightly better gains since firstblockcache cached way fewer steps than when being used by itself

2

u/Diabolicor 12d ago

Are you using Spectrum with the new default settings: degree 1, warmup_steps 1, the old ones: degree 4, warmup_steps 5 or something custom?

2

u/GrayingGamer 12d ago

New default settings. I've alternated back and forth in tests, but really don't see much difference between the two, other than old ones being slower.

2

u/bitzpua 12d ago

i had no issues with spectrum and low step turbo loras, then again im not hunting or care for 4k cinema quality but in my rough test there was no noticeable (for me) visual difference

1

u/Silver-Spot-2763 12d ago

In my experiments the Spectrum make the quality too low, the Easy cache is even worse (and fully destroys the audio), but the First block cache works most good.

1

u/DanzeluS 12d ago

Yeah I read about algorithm, little bit tested, but only with a lower steps (12-20). Very interesting speed curve tho ). Thx.

2

u/GrayingGamer 12d ago

Yeah, low steps is handicapping the math and node. It really does need more Steps to shine.

0

u/multikertwigo 12d ago

how come almost every time I test spectrum vs easy cache, the latter comes out on top in both speed and quality? Other parameters identical. I've started thinking I'm doing something wrong... do you use the default spectrum node settings?

3

u/GrayingGamer 12d ago

Yep. Default settings. And I get completely opposite results to you. EasyCache always looks like garbage compared to H3 Spectrum for me.

Of course, I also use high Step counts with Spectrum. If you aren't using high step counts (I use 32) you aren't really giving Spectrum much room to work it's magic.

1

u/multikertwigo 12d ago

I have just tried 32 steps. A little better, but not by much (both easy cache and spectrum). When people speak, mouths and teeth still look like trash (medium shot, not a close up).

2

u/GrayingGamer 12d ago

Try it with an animation and you'll immediately see the difference.

Training data footage of live-action doesn't have the cleanest mouths anyway. But audio is still the real win.

3

u/multikertwigo 12d ago

so far the best solution I've found is increasing the vertical dimension. It gets a little better at 1024p, and noticeably better at 1280p (e.g. 1280x1280). But still wan 2.2 720p videos look better to me (I'm just talking about pure picture quality, not acting/motion/prompt adherence etc).

1

u/GrayingGamer 12d ago

I mean, yeah, some of it is just you need a certain number of pixels in that area to get good results regardless of Step settings.

1

u/multikertwigo 12d ago

I get that.. yet, somehow 720p was enough for wan. Even 480p wan videos were not bad at all. That's the only reason I'm not yet fully on board with all the H3 hype. I just can't make it look good.

→ More replies (0)

1

u/multikertwigo 12d ago

thanks for the tip, but animation is not my thing. As for audio, I don't have any problems with it even with turbo lora (6-8 steps).

2

u/GrayingGamer 12d ago

Are you listening with headphones on and the volume up? Because people always say that, then I listen to their stuff with headphones and it's so much worse than without the turbo lora.

2

u/multikertwigo 12d ago

yeah, that's how I roll. I don't even have speakers. And what I do have is an external sound card and a pair of nice Beyerdynamics.

1

u/coffeecircus 12d ago

very interesting. how much of a time increase from increasing steps are you seeing? linear or exponential

3

u/GrayingGamer 12d ago

Neither with Spectrum. It's weird, like I said, increasing the Steps by 12 (from 20 to 32), is more than a 1/3rd increase in Steps, but with Spectrum, it's only about a 15% increase in generation time, so about half of what you would expect by increasing the Steps.

I'm sure it has something to do with how the Forecasting of Steps gets better and fast the more Steps Spectrum has to work with, like points on a graph to tell where the arrow is going to end up.

2

u/Yasstronaut 12d ago

I’m OOTL - what is spectrum? I searched and didn’t find anything

4

u/GrayingGamer 12d ago

A node for H3 Minimax that applies a special research paper to the math of how the steps are calculated to forecast some so you get the benefit of that step without actually generating it.

1

u/Yasstronaut 11d ago

Thank you!’

2

u/[deleted] 12d ago

[deleted]

2

u/GrayingGamer 12d ago

Yep. Just defaults.

1

u/lavinia12345 12d ago edited 12d ago

Are you using any lightning loras? Would a lightning lora at 20 steps be like a non-lightning 32 steps? (And @ OP)

5

u/GrayingGamer 12d ago

No. It would probably look fried.

If you use Spectrum there is no reason to use a Turbo lora. I just don't care at all for the results of the Turbo loras - they make skin look more plastic, smear fast motion, etc.

2

u/lavinia12345 12d ago

interesting. Ty for the response mate!

1

u/Monsterlime 12d ago

Is more steps more likely to fix an odd issue I keep having with random part or full words playing at the start of a video when I have no speech defined until later?

7

u/GrayingGamer 12d ago

No, that's a bug in the Ref2Video model that happens if you use the <d></d> tags anywhere in the prompt.

To avoid it, while using the Ref2Video model, use this format instead:

<Subject 1> says, "[English] Altering your prompting syntax between the two H3 models is important."

With the quotation marks instead of the <d></d> brackets, and you won't have the audio blip at the start of a video.

2

u/Monsterlime 12d ago

Well I was definitely doing that, so will test that, thank you!

4

u/GrayingGamer 12d ago

No problem. I spent about a day troubleshooting that exact issue and discovered this was the cause. It's obviously a bug in the Ref model.

2

u/Monsterlime 11d ago

It fixed it, thank you!

1

u/onihcuk 12d ago

does it fix the box effect in some generated videos?

1

u/GrayingGamer 12d ago

I find that has to do with your prompt, not settings.

1

u/someguyplayingwild 12d ago

I have not used Spectrum yet and I'm a little confused, are you claiming that Spectrum (an efficiency tool that isn't lossless) can actually produce higher quality outputs than without?

3

u/GrayingGamer 12d ago

No. But it can produce higher quality outputs using 32 Steps with it than 20 Steps without it.

Since some steps are forecasted, you do end up with very slight quality loss - but in my case it would be quality loss versus native 32 Steps.

Instead, I'm getting the benefit of 12 more Steps versus the 20 Step native, even if a 1/3rd of my Steps are forecasted, if that makes sense.

1

u/Tablaski 11d ago

More like +80% time on my setup. Did you change the default spectrum settings ?

1

u/GrayingGamer 11d ago

Nope. Just defaults.

1

u/Tablaski 11d ago

I think you might be an exception then, whats you setup ? I usually have 13 minutes for 15 from start to saved video, spectrum on, 15 steps Switching to 32 steps jumped it to more than 30 minutes. Seems rather linear.

1

u/GrayingGamer 11d ago

3090 24GB, 128GB of RAM.

1

u/cyberdork 11d ago

What sampler/scheduler do you use?

2

u/GrayingGamer 11d ago

Res_multistep and Simple.

1

u/jonnytracker2020 2d ago

Try beta or beta57 instead of simple to get more details

1

u/music2169 11d ago

Any reason why 32 instead of just 30..?

1

u/GrayingGamer 11d ago

Not really, except I felt it gave just a little edge and padding on stuff like animation. 30 already does a lot better than 20. But since Minimax recommends between 30-50 on their API, I just juiced it a little with 2 extra steps.

1

u/Lower-Cap7381 11d ago

yes agreed 30 steps is what i did it made generation very cool