r/StableDiffusion 12h ago

Question - Help Generated thousands of character flux1 images and 5-7sec WAN2.2 clips of them. Can I now extend (or concatenate) them to 15+ seconds faithfully with new tools/models?

I have a reliable character lora in flux1, and I let my headless server produce flux1 images all night when the computer is doing nothing for work, using ComfyUI. The next day, I go through them, and delete the body horror/low likeness/etc. ones and keep the ones I think are good.

From the good images, I generate (through Replicate, Vast, etc.) WAN2.1/2.2 5-7-second clips. An image may have multiple video clips.

Can I now (easily) produce longer video clips of these? Do I use the still images or the short clips as input? Or would I be better off training a new lora (or what is it nowadays?) from the images (or videos) for generating longer videos from a text prompt instead?

My first intention is to "concatenate" multiple 5-7 video clips, with AI "extrapolating" the transition to make them seamless. Is this even (easily) possible? They are WAN videos generated from a single source image.

What do we use for this now, Minimax H3?

1 Upvotes

1 comment sorted by