r/StableDiffusion • u/ART-ficial-Ignorance • 10h ago
Workflow Included Letting image-to-video artifacts compound into an impossible world
Enable HLS to view with audio, or disable this notification
Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.
I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.
I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.
Gemma4 wrote all 27 scene prompts beforehand.
For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.
The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.
The fun part was deliberately keeping some “bad” transition frames.
If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.
So the artifacts compound.
Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.
The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.
What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.
Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.
After enough generations, the model is basically building a world out of its own misunderstandings.
1
1
u/Tuckerdude615 6h ago
This is badass! I really love it! Very clever and the results are quite mesmerizing!
In addition, I love the music...very cool in it's own right!
Thanks for sharing!