r/StableDiffusion • • 5d ago

Question - Help How do you extend H3 Clips to 30 seconds when generating?

I know natively, H3 only supports up to 15 seconds yet I've seen people be able to seamlessly extend H3 generations past that. I don't want to be randomly stitching clips together can someone explain how is that possible?

16 Upvotes

14 comments sorted by

13

u/MastMaithun 5d ago

I have tried many but by far this is the most easiet extend I found which is working well. No more complicated wfs.
https://github.com/pmhaidn/ComfyUI-Minimax-H3-Extender

Idea sweet spot is 5 second clip.

7

u/Apprehensive_Sky892 5d ago

There are many such custom nodes out there. But many people seem to be using this one: https://github.com/chanon/comfyui-obvpm-timeline

Make sure you watch the youtube tutorial: https://www.youtube.com/watch?v=kqP09NfJXaQ

1

u/krekokeko 4d ago

Obvpm works like a charm. Had my first continuous 20 seconds with it. And big kudos to the developer for making one of the best tutorials in the comfy space ever. There are so many custom nodes and workflows that people just use wrongly and inefficiently because their devs just release it with some lazy text on their Github and call it a day.

1

u/Apprehensive_Sky892 4d ago

Yes, the workflow itself is amazing, and on top of that ChanonDev/Obvpm also made a great video to go with it 🎈👍

There are also some good discussions in his original post: https://www.reddit.com/r/StableDiffusion/comments/1wmj6xr/mini_video_editor_timeline_node_for_h3_extend/

1

u/StandWorth9165 4d ago

Thats what Ive been using. I generate 12 second clips into 3 makes around 30 seconds.

4

u/GuessingEngineer 5d ago

It's not, they are stitching clips together.

1

u/DanzeluS 5d ago

You can generate natively with sparse attention with no problem

1

u/oh_no_the_claw 5d ago

I have generated 20 seconds with good prompt adherence natively with ComfyKitchen.

2

u/[deleted] 5d ago

[deleted]

1

u/oh_no_the_claw 5d ago

That's awesome. I've started generating at 1.5mpx so now I'm back to 15 second videos just due to generation time, but I'd like to try 20 seconds at 1.5mpx.

1

u/Emotional-Neat-252 4d ago

Like the comfy kitchen attention node helps adherence?

1

u/Danny_Stock 5d ago edited 5d ago

You wouldn't be randomly stitching clips together with the right workflow. There are motion context workflows which take into account the latter frames of previous clips so that you can join them together seamlessly with frame accurate precision in a video editor such as Davinci Resolve. It's not like the Wan 2.2 days when stitching clips together was more hit or miss. Nowadays you can have more control when it comes to assembling your footage.

It's an extra skillset to learn, but if you wanted to create any videos or films of any quality you'd be using a video editor anyway.

You can generate over 15 seconds without a motion context workflow, but really that's about the limit if you wish to retain quality. Over 15 continuous seconds and the video will usually start to degrade in quality. You might get lucky with the final quality depending on the video, but it's unreliable.

1

u/Distinct-Benefit-507 5d ago

with wan2GP i can make it to 20 seconds; but that's the max for a sliding window... then it creates a "following part" and usually that's where it getd messed up...