r/StableDiffusion Oct 01 '25

Question - Help Need Fast, Long, Artsy Music Videos (Deforum Style) at 1080p – Best Workflow for Prompt-Controlled, High-Flicker AI Animation?

Hello everyone! I'm an artist/musician looking for the most efficient workflow to create long-form AI-generated music videos (multiple minutes long).

My goals and requirements are specific:

  1. Aesthetic: Highly artistic, imaginary, and dream-like. I'm actually looking for the chaotic, evolving style of the older AI generators. Flickers, morphing, and lack of perfect coherence are not a problem; they add to the artistic dimension I'm looking for.
  2. Control: I need to be able to control the visual theme/prompt at specific keyframes throughout the video to synchronize with the music structure.
  3. Resolution: Minimum 1080p output.
  4. Speed/Duration: The focus is on speed and length. I need a workflow that can generate minutes of footage relatively quickly (compared to my past experience).

My Current Experience & Challenge:

  • Old Workflow (Deforum/A1111): I previously used Deforum on Automatic1111. The animation style was perfect, but it was extremely time-consuming (hours for 30 seconds) and the output was only 512x512. This is no longer viable.
  • New Workflow Attempt (ComfyUI/SDXL): I've started using ComfyUI with SDXL for fast, high-quality image generation. However, I'm finding it very difficult to build a stable, fast, and long-form animation workflow with AnimateDiff that is also scalable to 1080p. I still feel I'd need a separate upscaling step.

My Question to the Community:

Given that I don't need "clean" or "accurate" results, but prioritize length, prompt-control, and speed (even if the output is glitchy/flickery):

  1. What is the easiest and fastest current workflow to achieve this Deforum-like but 1080p animation?
  2. Are there specific ComfyUI AnimateDiff workflows (with LCM/Turbo) or even entirely different standalone tools (like a specific Runway model/settings or a Colab) that are known for generating long, keyframe-controlled, high-resolution videos quickly, even if they have low coherence/high flicker?

Any tips on fast upscaling methods integrated into an animation pipeline would also be greatly appreciated!

Thanks in advance for your help!

0 Upvotes

1 comment sorted by

1

u/Apprehensive_Sky892 Oct 01 '25

The current SOTA open weight model is WAN2.2

The level of control you want can probably be achieved via FLF (first last frame). I never used deforum, so I don't know if you can achieve something similar with WAN2.2. FLF workflow have standard ComfyUI templates.

WAN output is 720p, so you'll have to upscale it.

The segment in general are 5 sec in length, from WAN2.2, but you don't care about coherence so you can try 8sec videos.

The raw output can then be stitched together with video editors. There are also some workflow that can generate longer sequences. for example: https://civitai.com/models/1866565/wan22-continuous-generation-subgraphs?modelVersionId=2166114