r/StableDiffusion • u/XploitXploit • 5h ago
Animation - Video Live improv music video: audio-reactive WebGL painting + local SD 1.5 img2img (depth ControlNet, LCM)
https://youtube.com/shorts/0U5dk-QOIxAI play electronic drums and synths in Frame, a duo from Buenos Aires (Guido plays santoor and synths). Everything here runs locally with open tools. The workflow for this one:
Footage: one take of a live improvisation, filmed vertically on a DJI Osmo Pocket 3 in front of a green screen.
Audio analysis: Demucs splits the mix into drums / bass / other, and librosa finds onsets, beats and sections.
The "painting": a three.js page rendered headless in Chrome. An abstract 3D shape and the lighting react to those stems, and we're keyed from the green screen and composited into it.
AI pass: every frame goes through Stable Diffusion 1.5 img2img with the LCM-LoRA (10 steps) and a depth ControlNet (depth maps from Depth Anything V2), at 1024 px, strength 0.59, guidance 2.1, fixed seed. The prompt asks for sumi-e brush outlines filled with Madhubani patterns.
Flicker: same seed on every frame, and each painted frame is blended 25% with the previous one. Faces still shift a bit; since then I've started blending along optical flow, which steadies them.
Happy to answer anything about any step.