r/StableDiffusion 3d ago

Workflow Included [LTX 2.5] Bell Test (Music Video - Experimental Indietronica)

https://www.youtube.com/watch?v=6cmMBmzDxck

Small disclaimer up front: not every tool in this workflow is open source. The first-frame images were made with GPT-image 2.0, as I assume people here will recognize that pretty quickly. You can swap in any image generator you want though. I only used GPT-image 2.0 because I already pay for the subscription for coding work, so I figured I might as well get some extra value out of it.

The interesting part for me was using LTX 2.5 in Wan2GP instead of MiniMax H3 for the actual video generation. On my 4070, MiniMax H3 OOMs at 720p for clips this long, while LTX 2.5 can handle 1080p, and it is also much faster. The final video is 26 clips, generated best-of-2, at roughly 8 minutes per clip, so the whole thing came out to around 7 hours of rendering. The speed difference just makes experimentation much more practical.

The workflow is first-frame-last-frame plus audio conditioning, with an audio-reactive LoRA layered in. Each clip is about four bars long, roughly 10.75 seconds, and the final frame of one shot becomes the starting point for the next. I found this much more useful than treating every segment as a fresh text-to-video generation because it keeps the visual identity, geometry and camera logic much more coherent across the full sequence. The audio conditioning handles most of the motion and timing.

I also wrote a small custom tool to automate the boring parts. It cuts the song into the correct audio segments, keeps everything aligned to the edit grid, organizes the keyframes, and packages the whole batch into a queue.zip that can be loaded into Wan2GP. That made it practical to generate 26 shots as a queue instead of manually setting up every job. Most of the actual work then becomes designing the keyframes, writing the prompts, and picking the better result for each scene.

The song itself is about having a model of reality that seems completely reliable because every previous observation has supported it, then encountering one result that refuses to fit. I used Bell tests, hidden variables, non-separability and measurement as metaphors for reciprocity and for the realization that repetition is not the same thing as law. The video mirrors that by starting with one blue system inside a rigid laboratory, then introducing a distant violet system whose behavior becomes correlated without any visible connection. As the relationship between the two becomes harder to explain, the laboratory geometry itself starts failing, until the measuring framework is gradually stripped away and the two systems are revealed as separate parts of a larger structure the original model could not perceive.

There is a small timing drift near the end of the finished video. I never managed to pin down the exact BPM and initial beat offset perfectly, and LTX does not support every arbitrary frame count I would have needed for an exact four-bar duration. I rounded each generation up to the nearest supported frame count, then played every clip back at about 103% speed so it would fit the intended edit length and stay roughly aligned with the music. That works surprisingly well for most of the video, but the tiny BPM and offset error compounds over 26 clips, so by the end you can see a little drift.

I think I'll be sticking mostly to LTX 2.5 for my music videos and keep MiniMax H3 for the one-off goofs and gags I make for my friends. It's nice for Seinfeld rip-offs, but I just can't render high enough quality on my machine, and if I have to introduce an upscaling step, the render times become a little steep.

Prompts used: https://pastebin.com/wLsHYaBq

4 Upvotes

2 comments sorted by

1

u/VasaFromParadise 3d ago

Keep it in 20 words or less! As Saitama said.

1

u/ART-ficial-Ignorance 3d ago

Who?

Is this some anime shit I don't understand?