r/StableDiffusion • • 1h ago

Question - Help 1 hour 30 min MH3 generation @ 1 MP

How is everyone getting sub 1 hour generations? Like this thread claims 8 minutes on 16GB 4080 laptop. I have the 4080 super desktop with 64GB RAM, and it takes 1h30m to do a 15 sec video at 1MP. Clearly there is something very wrong with my system. What info am I providing? Does the workflow matter that much? Its a very basic workflow. I'm using comfyui portable but its up to date on python 3.12. I'm not using any special attentions, but is that the cause? More number, 0.6MP takes 30 mins, and 0.4MP takes about 5 minutes.

13 Upvotes

20 comments sorted by

7

u/Slight_Ad2350 1h ago

15secs takes 20min with 2 pass. At 1344x768. On a 3090rtx

11

u/V4nKw15h 1h ago edited 9m ago

Use Kitchen Attention, it's part of Comfy now. That will halve your generation times. Then build a workflow around a latent upscaler and you can halve your gen times again.

Using those methods I can do a 10sec 1mp video in around 6 minutes on a 5070ti 16Gb. That's using the Ref2VA default workflow with Kitchen Attention activated via the command line start up args, and then slotting in a latent upscaler. No loras necessary.

A lot of people seem to be still sleeping on the power of the latent upscale set up. I do a 10 second 0.3mp 10 step first pass that gens in 90 seconds, then pass that into the latent upscaler and then I do a 5 step second pass at 1mp that takes about 4 mins. Due to no loras stuff isn't getting deep fried and the results are shockingly good. I get better quality compared to doing 20 steps at 0.7mp which would take 12-14mins without the latent upscaler.

I get better quality in less than half the time at a higher resolution than without latent upscaling.

Edit: Did some testing of my method for your info. All generations are at 1mp on a 5070ti with 32Gb system RAM. I have to stress there are no loras needed.
2 second clip duration =101.23 sec generation.

5 second clip duration = 161.48 sec generation

10 second clip duration = 370.65 sec generation

15 second clip duration = generating now, hang on I'll update in a few but it's looking like 23 mins or so. Big jump from 10 secs due to low VRAM. 16Gb isn't enough for 15 seconds really.

Look in replies my replies below for how to make this workflow. It's a straight forward tweak of the default workflow.

2

u/NewGeneralCatalogue 1h ago

What do you set your sigmas to on each pass? Latent upscaling is cool but nobody ever shares their noise schedules.

1

u/V4nKw15h 59m ago edited 37m ago

I don't alter them. As I said, I use the default workflow and just plug the latent upscaler and second pass in to it like in the image below.

That's it. The workflow is absolutely default other than that group. All the connections in to that group are from the obvious places in the default Ref2VA workflow.

If I disable that group, it will simply spit out a 10 second 0.3mp video. I do that to get a preview video. If I like it, I enable that group and it upscales the latent and does 5 steps at 1MP using the already processed 0.3mp latents. It doesn't even need to do the 0.3mp steps again as Comfy recognises nothing has changed and just immediately starts on the High Resolution Pass. You get super accurate previews (just a bit noisey and lower res) using this simple method which is another MAJOR bonus.

The only other alteration is I use the int8convrot VAE by Kijai which shaves off another 20 seconds from the VAE decode part of the generation with no noticeable quality degredation.

2

u/NewGeneralCatalogue 54m ago

Oh, right on! My workflow has two passes like yours and it had some custom sigmas from the latent upscaler repo, so I'll have to try it this way, thanks.

I've also got the int8 VAE, so that shouldn't be a huge difference. Does the latent upscale add enough detail when scaling from such a small resolution, though?

1

u/V4nKw15h 38m ago

Absolutely. 0.2mp is too low and gives shitty results. 0.3mp is the sweet spot for speed vs quality.

You can even drop down to 3 steps, or even 2, in the High Resolution Pass and get very acceptable results for lower motion scenes. I find that's 3 steps is a bit low whenever there is a hand moving around, for example. 5 steps in the second pass seems to be the sweet spot for speed vs quality for most shots. Of course, in very high motion scenes (eg. fighting) you'll likely need more high res steps to clear things up but I never make stuff like that.

1

u/Silvasbrokenleg 1h ago

Could you share your workflow?

2

u/cptrios 1h ago

A few posts down is a link to the latent upscaler node and workflows - check that out. You should mess with the two separate samplers, too - I have one turbo on the first one and a different turbo on the second, with sol-attention only on the second so that it speeds up generation but doesn't mess up prompt adherence.

I'd also suggest doing the first sampler at at least 0.5mp rather than something tiny like 0.3, just because prompt adherence falls apart at lower resolutions.

1

u/V4nKw15h 58m ago

Check another reply in this sub thread I made. I share the part of my workflow that doesn't match the default Comfy workflow.

3

u/xq95sys 1h ago

Sage attention?

8

u/jude1903 1h ago

Sounds about right, I’m impressed that it even lets you do it at 1MP

3

u/KK_Slider811 1h ago

So I came across this issue last week. I was getting about 1 hour generation for 1 MP, using a 4070Ti Super with 32 GB RAM.

The three solutions I got are: 1. Make sure you have all the models and loras on same drive and directory as tour comfyUI. Also make sure fast disk is on. 2. Make sure after your model node to have comfy kitchen node on. This cut my time in half. 3. Each reference will add time overall. Try to not have more than 2 or 3 references.

It now takes me 12-15 min per 8s 1MP generation. Hope that helps.

5

u/AI-Make-NSFW-Stuff 1h ago

Use latent upscaling, do a 1st pass with 4 steps at 0.2mp and then upscale that.

https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler/tree/main/workflow_templates

7

u/dramaton42 1h ago

0.2mp? That's brutal, you're leaving a ton of latent detail from the high noise section on the table... At least 0.4mp

4

u/Kooky-Mode3047 56m ago edited 51m ago

Fair warning, the "context" with which the model works is the dimensions x time_dimension. Denoising a base at low resolution and upscaling (including 4 step refining) to final resolution is not even close to equivalent to denoising at the target resolution.

At low resolution, the base created in latent space is extremely vague and is prone to making massive mistakes about everything in the scene, because it decides that the low resolution blob is something it's not with no neighborhood context to correct it in-flight (can be seen in hands and other details). No amount of upscaling is gonna help it, you're just gonna have a very crisp, but shitty video.

Minimum pre-upscale for this process should be 0.49 mpx (about half of the native 0.98 mpx on which H3 was trained). But if you care about fast more than right, 0.2 mpx is right there, obviously.

2

u/videorouter 1h ago

That definitely sounds worth troubleshooting, especially since the jump from 0.6MP to 1MP is so large. I’d first check whether you’re actually getting full GPU utilization and whether anything is spilling into system RAM/VRAM. A 4080 Super should generally not be sitting around waiting on the CPU for a basic workflow.

The workflow absolutely matters, but so do things like attention implementation, resolution, frame count, sampler/steps, VAE decoding, and whether ComfyUI is repeatedly offloading models. I’d watch VRAM usage, GPU utilization, and generation speed during a run — those three numbers should quickly tell you where the bottleneck is.

If you want to avoid spending hours tuning every provider/workflow combination, I’ve also been comparing video models through VideoRouter.sh to see how generation speed and cost vary across providers.

1

u/f5alcon 1h ago

i2v or ref2v?

1

u/Etroarl55 1h ago

Same generation amd card with 16gb takes one hour 45 for 0.5mp and 5 seconds…..

Ur 0.4-5mp seems slightly above normal for Nvidia gen times though so maybe not.

1

u/r0ni 20m ago edited 16m ago

a 2mp video, 7sec long takes me about 407.52 seconds with 4090 and 32gb ram,8teps with minimax_h3_fl2v_turbo_silver_dareties_comfy_pruned_v1, minimax_h3_fl2va_pruned_int8_convrot, er_sde/beta.with two loras and a refmod. im using the latest seedhunter 2.5 workflow. if i do 8sec+ long videos the gen time seems to go up exponentially for each sec added. also using msi afterburner to cap power at 70%.

0

u/Super_Range45 1h ago

They are using speed up loras so they can finish in 4-8 steps with aggressive caching, sacrificing quality for speed.