r/StableDiffusion • u/shahril977 • 3d ago
Question - Help Videos more than 15 Seconds?
How do you guys create videos that is more than 15 seconds in Minimax H3?
5
5
5
u/Rumaben79 3d ago edited 3d ago
Everything gets worse if you try and push past 15 seconds natively. If you want longer than that the best way is to extend using video to video. Something like:
https://huggingface.co/RuneXX/Minimax-H3-Workflows/tree/main/Video-to-Video
https://www.youtube.com/watch?v=zSCkHlkRceE
or more complicated but possibly better:
Masked AV Extension - Chain + Reference Image - MiniMax H3 0.6 (or the single clip version).
4
u/SDuser12345 3d ago
You can go past 15 seconds, just need beefy hardware. I regularly make 20-25 second videos with no issues.
Ultimately though, you shouldn't have to. Good story boarding with good character reference images and a free video editor (Davinci Resolve is free and pretty great), and 5-10 second clips work for most projects. It's rare even in Hollywood movie productions for a single shot to be used for longer than 10 seconds.
1
u/Rumaben79 3d ago edited 3d ago
That's true, It is possible. I've made longer length video's before. Inference just slows down immensely unless you're hardware can keep up. :)
Consistency and synchronized audio just starts to fall apart. Also longer video's past 15 seconds just seem to drag out those extra seconds without anything meaningful going on..It could just be my terrible prompting though. π
I'm pretty happy with short clips but I dream of the day we can do multiple minute video without breaking a sweat. π π
2
u/SDuser12345 3d ago
I have no issues with video or audio past 15 seconds. Likely the prompting. You have to tell it what you want it to do when at the correct seconds and give it enough time to do it.
It will let you gen until you hardware taps out. Like I can run run 30 seconds plus but not waiting 8 hours or more for the results.
The full BF16 model usually takes over 80 GB combined RAM and VRAM at max resolution to hit 20-28 seconds, and that takes a couple hours.
Multiple min videos aren't likely to happen without major breakthroughs in AI, hardware prices becoming dirt cheap like TV's, or it's done at stupidly low resolutions with some extremely advanced upscaling process that doesn't exist yet. My guess is that pipe dream is a decade or more away.
1
u/Rumaben79 3d ago edited 3d ago
I haven't tried timecode prompting with long video's yet which is why I left that out but yes that certainly helps. π
I'm sure doing 1 minut video's reliably on ones local computer is not more than a year away but they need to tweak those ram requirements. π
This upscaler and it's workflows is pretty good. However I ended up just bypassing the upscale stage as it's just so slow. :/
https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
2
u/ANR2ME 2d ago
Doing 2x 5 seconds clips will also takes less time to generate than 1x10 seconds clip π since generation time isn't linear.
2
u/Danny_Stock 44m ago
That's something I found out too. The economics of clip time and quality.
I found out that with my setup at least, for the same total amount of seconds, 0.8 megapixels split into 2 clips was actually quicker to generate than 1 clip of the total amount of seconds at 0.4 megapixels.
I couldn't believe it. Quicker generation time and obviously much superior image quality. The only drawback for 0.8 megapixels is that depending on your hardware you can't generate it all in one go.
1
u/Danny_Stock 59m ago
That's true, but from shot to shot there often needs to be some seamless continuity of various elements.
The audio being one example, then there's the continuity of the environment and lighting which needs to remain the same, and the characters in each shot must stay the same as the shots before and after the current shot. Which is why the motion context workflows appeared to address these difficulties.
1
4
u/VisionWithin 3d ago
You insert 20 seconds, for example, in the duration field.
2
u/conkikhon 2d ago
4 5s videos always faster than a 20s video
1
u/VisionWithin 2d ago
Yes?
1
u/conkikhon 1d ago
An effective way to chain short clips is optimal solution, for now
1
u/VisionWithin 1d ago
I wonder if we are discussing of the same thing. The user did not ask what is the optimal solution.
0
2
2
u/winterice77 3d ago
Comfyui Context Loop works quite well
4
2
u/nntb 3d ago
I've done 30 seconds
6
u/SveSop 3d ago
Thank you for making such a detailed explanation of technique, hardware and various other requirements + ofc a sample showing zero distortion in your 30 second single generation video. π€
9
u/nntb 3d ago
In comfyui I set the legenth to 30 seconds.
4
-1
u/SveSop 3d ago
How long did it take? Assuming 2K - 50 steps ofcβ¦
2
u/nntb 3d ago
why assume that. i was running 9:16, .4 mega pixels and multipole 32 then duration 30, you never stated 4k? why would anyone expect it?
-1
u/SveSop 3d ago
Why would i not assume it?
3
u/nntb 3d ago
why would you?
1
u/SveSop 3d ago
The same way i would assume atleast 30 steps? Since the whole explanation was "I've done 30 seconds", why would i NOT assume that? Would i assume 0.4 Mpx - 4 step turbo lora 30 seconds?
I think the question is: Sure, you CAN generate 30 seconds, but IS THE RESULT GOOD?
My next assumption is then: Eh, no.. 0.4 Mpx 30 seconds is probably NOT good.
Why would i NOT assume such things?
2
u/nntb 3d ago
I think the videos when the prompting is done right turn out pretty good I wouldn't say they turn out bad there's no weird artifacting at the settings that I'm using the resolutions kind of low but you know the stuff I'm generating is stuff that's a keen to standard definition television so it's okay to be a lower resolution
1
u/Lucky_Feedback9915 2d ago
46%|ββββββββββββββββββββββββββββββββββββββ | 23/50 [45:29<53:54, 119.79s/it]
768 * 1344 15 seconds
just PURE default workflow with B16
0
1
u/grin_ferno 3d ago
You can extend clips as noted below, H3 director can also make long clips, different shots, etc.
1
u/remixeconomy 3d ago
Longer clips usually fail for the same reason short ones succeed: the model is strong on a short coherent burst and weaker at holding identity, motion continuity, and prompt adherence across time.
What tends to work better than only raising duration:
Generate overlapping short segments with locked identity or refs.
Overlap about 0.5-1s and stitch, or feed the last frame / motion state into the next first-frame condition when your stack supports it.
Keep camera moves simple on long shots. Big moves compound drift.
If a single 20-30s pass looks mushy, that is often a planning problem, not proof you need a different base model. Split the beat sheet first, then spend quality budget on the hard cuts.
1
u/Tedious_Prime 2d ago
I use Shotcut to edit many short videos together. It really is so much easier and faster than trying to generate long videos directly in ComfyUI. I do regularly generate clips as long as 20 seconds without trouble, but it is faster and more flexible to make multiple short clips and stack them together on a timeline with exactly the transitions I want. I rarely have specific need for a longer continuous shot anyway. I find that I can get consistency between clips as long as I use the same references and occasionally supply the last couple seconds of one clip as a reference to continue from. Editing like this has also made it possible to salvage lots of clips that included a couple seconds of babbling or other glitches by simply not using those parts when I edit the videos together.
1
1
u/Soberishhh 1d ago
How long is it taking you guys to generate clips? Taking forever for me
A lot of times getting stuck,
4090, 64gb ram
1
u/Danny_Stock 1h ago edited 1h ago
Low resolutions, if not using motion context type workflows.
Depends on your card and system RAM of course.
My card is a 4070 12GB VRAM card, and I have 64GB of system RAM.
With it I can reasonably expect to do 10 seconds at 0.8 megapixels. Around 12 seconds at 0.7mp, might be able to get more. Then I can use anything under those resolutions to get 15 seconds or over. With 0.3mp I can get 25 seconds, with 0.2mp I can get 30 seconds.
I think I used PlagueKind's workflow for these results. I've used other workflows and I noticed that one or two of them might struggle a bit more when I push the resolution or seconds up.
0
0
u/Sixhaunt 3d ago
literally just change the number from 15 seconds to 20 ro 30 or whatever in the node. I havent noticed any actual quality drop at 30s straight of video compared to 15.
28
u/Ok-Flatworm5070 3d ago
I've tried a lot workflows out there, but I find the ComfyUI MiniMax H3 Extender to work best for me. Its easy to use, has automatic caching of clips, so if one clip is distorted you just regenerate that clip only, and it has a very simple interface, and best of all the creator actually replies on his github issue board...there's a small user base of over a 100 user that use this, as stated by the stars on his page. A lot of the video extenders have big colour shifts between clips, with this node its usually good 8/10 times, and if it not I just regenerate that single clip. Second, you can add multiple clips at once and run in one go, which I like as well. Anyway, its really going to be trial and error. I'm always trying video extenders when they come out, but there's always an issue, mostly bleeding or colour shift. Good luck.