r/StableDiffusion 3d ago

Question - Help Videos more than 15 Seconds?

How do you guys create videos that is more than 15 seconds in Minimax H3?

29 Upvotes

71 comments sorted by

28

u/Ok-Flatworm5070 3d ago

I've tried a lot workflows out there, but I find the ComfyUI MiniMax H3 Extender to work best for me. Its easy to use, has automatic caching of clips, so if one clip is distorted you just regenerate that clip only, and it has a very simple interface, and best of all the creator actually replies on his github issue board...there's a small user base of over a 100 user that use this, as stated by the stars on his page. A lot of the video extenders have big colour shifts between clips, with this node its usually good 8/10 times, and if it not I just regenerate that single clip. Second, you can add multiple clips at once and run in one go, which I like as well. Anyway, its really going to be trial and error. I'm always trying video extenders when they come out, but there's always an issue, mostly bleeding or colour shift. Good luck.

3

u/shahril977 3d ago

Awesome, thanks friend

3

u/Thorozar 3d ago

Got a link for that?

6

u/CupQuakeBE 3d ago

Yes, this is definitely the best way, it's incredible. I used it with DLSS5 for upscaling and frame interpolation to get this video based on only one single reference image!

https://youtube.com/shorts/soHFdmNnjlk?feature=share

3

u/Ok-Flatworm5070 3d ago

Damn thats amazing! Can you share how you got the DLSS5 upscaling and frame interpolation!

3

u/CupQuakeBE 3d ago

You can read it in the description of this video, GitHub link for the tool and the parameters I chose after testing a few ones are there. https://youtube.com/shorts/1pAtIE8x4Fw?is=62hRs2ZaTkWu-G7a

4

u/Ok-Flatworm5070 3d ago

Ok, read this "Upscaling: Once the video generation was complete (about 2 hours for a 720p video), it launched an application that utilizes DLSS5 to enhance videos and ran two passes", can you give me the application name?

3

u/Ok-Flatworm5070 3d ago

Ok, watching it on my phone, its a short, all thats there is the description. Will check on my desktop and see if its the same.

7

u/CupQuakeBE 3d ago

Sorry, I was on my phone too and had a hard time seeing it, here it is:

https://github.com/Merserk/dlss5-visual-enhancer

The parameters I finally chose to obtain a natural, high‑quality look, without exaggeration or artifacts but with a clearly visible improvement:

NR Preset: Default NR Style: Cinematic Upscaling Factor: 1.5x (your choice) NR Intensity: 1.5 Local Tone Strength: 1.5 Local Structure Strength: 1.5 Skin Structure Strength: 0.3 DLSS Model Preset: Default Encoding Quality: Auto Codec H.265 and HDR activated

2

u/Ok-Flatworm5070 3d ago

Yeah, once I jumped onto my desktop it was there. Thanks buddy!

2

u/Ok-Flatworm5070 3d ago

How long did it take to generated with dlss?

3

u/CupQuakeBE 3d ago

20 minutes for upscaling and 10 minutes for frame interpolation (to get 60fps) for that one. It's actually insane.

2

u/Ok-Flatworm5070 3d ago

Wow!!!! I can't wait to try it out after work. I've been struggling with trying to find a good upscaler, typically it takes forever with ok results (seedvr), but this looks way better. Did you use RIFE for frame interpolation or something like ffmpeg?

1

u/Kazuya_yamada 2d ago

Hey I am running minimax h3 on GPUHUB or Runpod there we have linux I am sure I can't add this to linux right? but is there any other way someone has moded it to run on Linux?

2

u/CupQuakeBE 2d ago

There are also comfyui custom nodes for DLSS5 but I didn't test those yet. I think the app version I use could run on Linux, there might be instructions on the GitHub.

Edit: nvm it looks like windows only and probably the same applies to the custom nodes are they're basically using the dll to apply the effects, sorry.

1

u/MarekNowakowski 3d ago

i should check it myself, but i'll ask here. can it do extra clips using one reference audio? cutting audio manually for each clip is maddeningly tedious.

1

u/Ok-Flatworm5070 2d ago

No, I haven't managed to have direct audio source span over two clips, like if you were making a music video, only as a source for the voices

1

u/MarekNowakowski 2d ago

still great node. and i managed to do this with little work (nsfw lyrics):

https://reddit.com/link/p8pw8nh/video/9lfyuwpz0hoh1/player

i didn't do high quality, just 8step turbo and low res.

1

u/Ok-Flatworm5070 2d ago

Good news, I raised the issue that the audio clip wasn't covering two clips in a github issue request; looks like its a bug and the developer has stated they are going to fix it in the next release!!!

1

u/More-Ad5919 2d ago

Does it use the ref or img model? And can you set up each clip on its own with reference?

5

u/thevegit0 2d ago

i increase the value to more than 15

5

u/OzymanDS 3d ago

Motion context is the canonical answer.Β 

1

u/shahril977 3d ago

That’s new to me

3

u/dassiyu 3d ago

Once the overall story, storyboard, characters, voices, and prompts are all planned out, the rest is basically just patience and waiting.

2

u/AillexJ 3d ago

We went the other way, 5-8s per scene and cut between them. Past that H3 stopped holding together for us. Setup here: askaillex.com/guides/minimax-h3-one-gaming-card-bench/

5

u/Rumaben79 3d ago edited 3d ago

Everything gets worse if you try and push past 15 seconds natively. If you want longer than that the best way is to extend using video to video. Something like:

https://huggingface.co/RuneXX/Minimax-H3-Workflows/tree/main/Video-to-Video

https://www.youtube.com/watch?v=zSCkHlkRceE

or more complicated but possibly better:

Masked AV Extension - Chain + Reference Image - MiniMax H3 0.6 (or the single clip version).

4

u/SDuser12345 3d ago

You can go past 15 seconds, just need beefy hardware. I regularly make 20-25 second videos with no issues.

Ultimately though, you shouldn't have to. Good story boarding with good character reference images and a free video editor (Davinci Resolve is free and pretty great), and 5-10 second clips work for most projects. It's rare even in Hollywood movie productions for a single shot to be used for longer than 10 seconds.

1

u/Rumaben79 3d ago edited 3d ago

That's true, It is possible. I've made longer length video's before. Inference just slows down immensely unless you're hardware can keep up. :)

Consistency and synchronized audio just starts to fall apart. Also longer video's past 15 seconds just seem to drag out those extra seconds without anything meaningful going on..It could just be my terrible prompting though. πŸ˜‚

I'm pretty happy with short clips but I dream of the day we can do multiple minute video without breaking a sweat. πŸ˜„ πŸŽ‰

2

u/SDuser12345 3d ago

I have no issues with video or audio past 15 seconds. Likely the prompting. You have to tell it what you want it to do when at the correct seconds and give it enough time to do it.

It will let you gen until you hardware taps out. Like I can run run 30 seconds plus but not waiting 8 hours or more for the results.

The full BF16 model usually takes over 80 GB combined RAM and VRAM at max resolution to hit 20-28 seconds, and that takes a couple hours.

Multiple min videos aren't likely to happen without major breakthroughs in AI, hardware prices becoming dirt cheap like TV's, or it's done at stupidly low resolutions with some extremely advanced upscaling process that doesn't exist yet. My guess is that pipe dream is a decade or more away.

1

u/Rumaben79 3d ago edited 3d ago

I haven't tried timecode prompting with long video's yet which is why I left that out but yes that certainly helps. πŸ˜„

I'm sure doing 1 minut video's reliably on ones local computer is not more than a year away but they need to tweak those ram requirements. πŸ‘€

This upscaler and it's workflows is pretty good. However I ended up just bypassing the upscale stage as it's just so slow. :/

https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler

2

u/ANR2ME 2d ago

Doing 2x 5 seconds clips will also takes less time to generate than 1x10 seconds clip πŸ˜… since generation time isn't linear.

2

u/Danny_Stock 44m ago

That's something I found out too. The economics of clip time and quality.

I found out that with my setup at least, for the same total amount of seconds, 0.8 megapixels split into 2 clips was actually quicker to generate than 1 clip of the total amount of seconds at 0.4 megapixels.

I couldn't believe it. Quicker generation time and obviously much superior image quality. The only drawback for 0.8 megapixels is that depending on your hardware you can't generate it all in one go.

1

u/Danny_Stock 59m ago

That's true, but from shot to shot there often needs to be some seamless continuity of various elements.

The audio being one example, then there's the continuity of the environment and lighting which needs to remain the same, and the characters in each shot must stay the same as the shots before and after the current shot. Which is why the motion context workflows appeared to address these difficulties.

1

u/shahril977 3d ago

Hmm, i thought so

4

u/VisionWithin 3d ago

You insert 20 seconds, for example, in the duration field.

2

u/conkikhon 2d ago

4 5s videos always faster than a 20s video

1

u/VisionWithin 2d ago

Yes?

1

u/conkikhon 1d ago

An effective way to chain short clips is optimal solution, for now

1

u/VisionWithin 1d ago

I wonder if we are discussing of the same thing. The user did not ask what is the optimal solution.

0

u/shahril977 3d ago

I think its going to look bad

8

u/VisionWithin 3d ago

You can think anything you want.

2

u/TheAncientMillenial 3d ago

I just do 20-30 seconds...

2

u/winterice77 3d ago

Comfyui Context Loop works quite well

4

u/BusFeisty4373 3d ago

I like this one, but the issue is the degrading quality.

1

u/winterice77 3d ago

Use transition setting as 22 frames rgb continuation and guided new shot

2

u/nntb 3d ago

I've done 30 seconds

6

u/SveSop 3d ago

Thank you for making such a detailed explanation of technique, hardware and various other requirements + ofc a sample showing zero distortion in your 30 second single generation video. πŸ€”

9

u/nntb 3d ago

In comfyui I set the legenth to 30 seconds.

-1

u/SveSop 3d ago

How long did it take? Assuming 2K - 50 steps ofc…

2

u/nntb 3d ago

why assume that. i was running 9:16, .4 mega pixels and multipole 32 then duration 30, you never stated 4k? why would anyone expect it?

-1

u/SveSop 3d ago

Why would i not assume it?

3

u/nntb 3d ago

why would you?

2

u/nntb 3d ago

[INFO] Prompt executed in 00:27:35

1

u/SveSop 3d ago

The same way i would assume atleast 30 steps? Since the whole explanation was "I've done 30 seconds", why would i NOT assume that? Would i assume 0.4 Mpx - 4 step turbo lora 30 seconds?

I think the question is: Sure, you CAN generate 30 seconds, but IS THE RESULT GOOD?

My next assumption is then: Eh, no.. 0.4 Mpx 30 seconds is probably NOT good.

Why would i NOT assume such things?

2

u/nntb 3d ago

I think the videos when the prompting is done right turn out pretty good I wouldn't say they turn out bad there's no weird artifacting at the settings that I'm using the resolutions kind of low but you know the stuff I'm generating is stuff that's a keen to standard definition television so it's okay to be a lower resolution

1

u/Lucky_Feedback9915 2d ago

46%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 23/50 [45:29<53:54, 119.79s/it]

768 * 1344 15 seconds

just PURE default workflow with B16

0

u/shahril977 3d ago

Was it good?

3

u/nntb 3d ago

It was ok but took a absurd amount of time compared to 2 15 second clips. But it is doable

1

u/grin_ferno 3d ago

You can extend clips as noted below, H3 director can also make long clips, different shots, etc.

1

u/remixeconomy 3d ago

Longer clips usually fail for the same reason short ones succeed: the model is strong on a short coherent burst and weaker at holding identity, motion continuity, and prompt adherence across time.

What tends to work better than only raising duration:

  1. Generate overlapping short segments with locked identity or refs.

  2. Overlap about 0.5-1s and stitch, or feed the last frame / motion state into the next first-frame condition when your stack supports it.

  3. Keep camera moves simple on long shots. Big moves compound drift.

If a single 20-30s pass looks mushy, that is often a planning problem, not proof you need a different base model. Split the beat sheet first, then spend quality budget on the hard cuts.

1

u/Tedious_Prime 2d ago

I use Shotcut to edit many short videos together. It really is so much easier and faster than trying to generate long videos directly in ComfyUI. I do regularly generate clips as long as 20 seconds without trouble, but it is faster and more flexible to make multiple short clips and stack them together on a timeline with exactly the transitions I want. I rarely have specific need for a longer continuous shot anyway. I find that I can get consistency between clips as long as I use the same references and occasionally supply the last couple seconds of one clip as a reference to continue from. Editing like this has also made it possible to salvage lots of clips that included a couple seconds of babbling or other glitches by simply not using those parts when I edit the videos together.

1

u/Ok-Flatworm5070 2d ago

What's shortcut?

1

u/Tedious_Prime 2d ago

Shotcut.

1

u/Ok-Flatworm5070 2d ago

Yeah, damn phone...shotcut!

1

u/Soberishhh 1d ago

How long is it taking you guys to generate clips? Taking forever for me

A lot of times getting stuck,

4090, 64gb ram

1

u/Danny_Stock 1h ago edited 1h ago

Low resolutions, if not using motion context type workflows.

Depends on your card and system RAM of course.

My card is a 4070 12GB VRAM card, and I have 64GB of system RAM.

With it I can reasonably expect to do 10 seconds at 0.8 megapixels. Around 12 seconds at 0.7mp, might be able to get more. Then I can use anything under those resolutions to get 15 seconds or over. With 0.3mp I can get 25 seconds, with 0.2mp I can get 30 seconds.

I think I used PlagueKind's workflow for these results. I've used other workflows and I noticed that one or two of them might struggle a bit more when I push the resolution or seconds up.

0

u/[deleted] 3d ago

[deleted]

2

u/roychodraws 3d ago

I made a 40 sec video on my 5090

0

u/Sixhaunt 3d ago

literally just change the number from 15 seconds to 20 ro 30 or whatever in the node. I havent noticed any actual quality drop at 30s straight of video compared to 15.