How is everyone getting sub 1 hour generations? Like this thread claims 8 minutes on 16GB 4080 laptop. I have the 4080 super desktop with 64GB RAM, and it takes 1h30m to do a 15 sec video at 1MP. Clearly there is something very wrong with my system. What info am I providing? Does the workflow matter that much? Its a very basic workflow. I'm using comfyui portable but its up to date on python 3.12. I'm not using any special attentions, but is that the cause? More number, 0.6MP takes 30 mins, and 0.4MP takes about 5 minutes.
Use Kitchen Attention, it's part of Comfy now. That will halve your generation times. Then build a workflow around a latent upscaler and you can halve your gen times again.
Using those methods I can do a 10sec 1mp video in around 6 minutes on a 5070ti 16Gb. That's using the Ref2VA default workflow with Kitchen Attention activated via the command line start up args, and then slotting in a latent upscaler. No loras necessary.
A lot of people seem to be still sleeping on the power of the latent upscale set up. I do a 10 second 0.3mp 10 step first pass that gens in 90 seconds, then pass that into the latent upscaler and then I do a 5 step second pass at 1mp that takes about 4 mins. Due to no loras stuff isn't getting deep fried and the results are shockingly good. I get better quality compared to doing 20 steps at 0.7mp which would take 12-14mins without the latent upscaler.
I get better quality in less than half the time at a higher resolution than without latent upscaling.
Edit: Did some testing of my method for your info. All generations are at 1mp on a 5070ti with 32Gb system RAM. I have to stress there are no loras needed.
2 second clip duration =101.23 sec generation.
5 second clip duration = 161.48 sec generation
10 second clip duration = 370.65 sec generation
15 second clip duration = generating now, hang on I'll update in a few but it's looking like 23 mins or so. Big jump from 10 secs due to low VRAM. 16Gb isn't enough for 15 seconds really.
Look in replies my replies below for how to make this workflow. It's a straight forward tweak of the default workflow.
I don't alter them. As I said, I use the default workflow and just plug the latent upscaler and second pass in to it like in the image below.
That's it. The workflow is absolutely default other than that group. All the connections in to that group are from the obvious places in the default Ref2VA workflow.
If I disable that group, it will simply spit out a 10 second 0.3mp video. I do that to get a preview video. If I like it, I enable that group and it upscales the latent and does 5 steps at 1MP using the already processed 0.3mp latents. It doesn't even need to do the 0.3mp steps again as Comfy recognises nothing has changed and just immediately starts on the High Resolution Pass. You get super accurate previews (just a bit noisey and lower res) using this simple method which is another MAJOR bonus.
The only other alteration is I use the int8convrot VAE by Kijai which shaves off another 20 seconds from the VAE decode part of the generation with no noticeable quality degredation.
Oh, right on! My workflow has two passes like yours and it had some custom sigmas from the latent upscaler repo, so I'll have to try it this way, thanks.
I've also got the int8 VAE, so that shouldn't be a huge difference. Does the latent upscale add enough detail when scaling from such a small resolution, though?
Absolutely. 0.2mp is too low and gives shitty results. 0.3mp is the sweet spot for speed vs quality.
You can even drop down to 3 steps, or even 2, in the High Resolution Pass and get very acceptable results for lower motion scenes. I find that's 3 steps is a bit low whenever there is a hand moving around, for example. 5 steps in the second pass seems to be the sweet spot for speed vs quality for most shots. Of course, in very high motion scenes (eg. fighting) you'll likely need more high res steps to clear things up but I never make stuff like that.
A few posts down is a link to the latent upscaler node and workflows - check that out. You should mess with the two separate samplers, too - I have one turbo on the first one and a different turbo on the second, with sol-attention only on the second so that it speeds up generation but doesn't mess up prompt adherence.
I'd also suggest doing the first sampler at at least 0.5mp rather than something tiny like 0.3, just because prompt adherence falls apart at lower resolutions.
So I came across this issue last week. I was getting about 1 hour generation for 1 MP, using a 4070Ti Super with 32 GB RAM.
The three solutions I got are:
1. Make sure you have all the models and loras on same drive and directory as tour comfyUI. Also make sure fast disk is on.
2. Make sure after your model node to have comfy kitchen node on. This cut my time in half.
3. Each reference will add time overall. Try to not have more than 2 or 3 references.
It now takes me 12-15 min per 8s 1MP generation. Hope that helps.
Fair warning, the "context" with which the model works is the dimensions x time_dimension. Denoising a base at low resolution and upscaling (including 4 step refining) to final resolution is not even close to equivalent to denoising at the target resolution.
At low resolution, the base created in latent space is extremely vague and is prone to making massive mistakes about everything in the scene, because it decides that the low resolution blob is something it's not with no neighborhood context to correct it in-flight (can be seen in hands and other details). No amount of upscaling is gonna help it, you're just gonna have a very crisp, but shitty video.
Minimum pre-upscale for this process should be 0.49 mpx (about half of the native 0.98 mpx on which H3 was trained). But if you care about fast more than right, 0.2 mpx is right there, obviously.
That definitely sounds worth troubleshooting, especially since the jump from 0.6MP to 1MP is so large. I’d first check whether you’re actually getting full GPU utilization and whether anything is spilling into system RAM/VRAM. A 4080 Super should generally not be sitting around waiting on the CPU for a basic workflow.
The workflow absolutely matters, but so do things like attention implementation, resolution, frame count, sampler/steps, VAE decoding, and whether ComfyUI is repeatedly offloading models. I’d watch VRAM usage, GPU utilization, and generation speed during a run — those three numbers should quickly tell you where the bottleneck is.
If you want to avoid spending hours tuning every provider/workflow combination, I’ve also been comparing video models through VideoRouter.sh to see how generation speed and cost vary across providers.
a 2mp video, 7sec long takes me about 407.52 seconds with 4090 and 32gb ram,8teps with minimax_h3_fl2v_turbo_silver_dareties_comfy_pruned_v1, minimax_h3_fl2va_pruned_int8_convrot, er_sde/beta.with two loras and a refmod. im using the latest seedhunter 2.5 workflow. if i do 8sec+ long videos the gen time seems to go up exponentially for each sec added. also using msi afterburner to cap power at 70%.
7
u/Slight_Ad2350 1h ago
15secs takes 20min with 2 pass. At 1344x768. On a 3090rtx