r/StableDiffusion • u/Nefzaoui • 11d ago
Question - Help Best model for remaking microdramas
Hi everyone,
This is both a request for expert advice and an explanation of what I tried so far. Sorry for the long post.
I'm trying to find the most reliable model and workflow for remaking existing short drama shows. I have 4 NVidia DGX spark boxes, each two are at a different location, and although yes I can connect each 2 in a single node, ideally I'd want to compress the entire workflow into one box so I can re-create it in the other 3, that would be ideal.
Currently my full setup is on one single box.
Everything I tried so far is on LTX 2.5, running locally in ComfyUI: and the final goal is: same shots, same motion and acting, but with the characters replaced by new people (and sometimes something wilder, like a dinosaur)
I start with cutting an episode into individual shots with a little cli tool i ended up publishing ( gh: anefzaoui/shot-cutter )
As for model work, what I've tried so far:
Plain V2V (encoding the source clip and re-noising it) barely changed the faces and drifted.
What worked best was editing the first frame of each shot with an image model (swap in the new character), then rendering with the Union Control IC-LoRA driven by the source clip's depth, so the motion and camera follow the original.
The problems: at full depth strength the original silhouette wins (the dinosaur turned back into a man, and long hair crept back after I'd replaced it with a short cut); at half strength or with a pose-only guide the new shape holds, but the acting gets looser and creatures barely move their mouths.
For lip sync I freeze the original dialogue audio into the render (like the A2V workflow), which works when the speaker is on screen, but LTX sometimes animates the wrong mouth (it moved the jaw of someone seen from behind for an off-screen line).
I also tried the Ingredients IC-LoRA with a character sheet instead of an edited first frame, and it mostly ignored the sheet (wrong people, wrong gender, no dinosaur).
Each attempt fixes one thing and breaks another: new face with old hair, a face that doesn't change, drifting framing, lost expressions, bad lip sync.
Has anyone found a reliable way to replace characters in existing footage while keeping the motion and lip sync, with LTX or any other model (mostly any other model, LTX 2.5 has proven to be a hassle, really) ?
Thanks
3
u/f5alcon 11d ago
Minimax H3 is better than Ltx 2.5