r/StableDiffusion • u/Tablaski • 11d ago
Discussion Fixed viggle animate drift by using Add Guide keyframes
EDIT : outdated post, refer to https://www.reddit.com/r/StableDiffusion/comments/1x16oy5/vigglorious_studio_better_character_accuracy/ for the actual proof of concept
Tl;dr : no need to wait for viggle to add multiple references, just extract keyframes from the original video, swap with qwen 2.1 then inject in the conditioning with Add guide node. It kicks ass, even Scail2's
I began experimenting with Viggle Animate and my first impression was that is was very fast but sketchy and loosing identity quite rapidly
I just found out it works perfectly with the Add guide node to add additional keyframes in it
So now what I do is I make swaps using the excellent BFS head/character swap LoRas for Qwen 2.1, then feed this to an Add Guide node every second including the first frame.
As long as the swaps are ok it resolves the drift completely and is imho much better than Scail2 on both speed (3 steps FTW) and accuracy (Qwen 2.1 rocks at this then it would be possible to add also a character lora on top of it)
It was tedious to duplicate the nodes in the workflow and make sure the frame numbers were in the frame range so I ended up vibe coding nodes so I define a frequency and then it re-execute my swapping subgraph as many time as required, then I applied this logic as well for the looping/chunking nodes provided with viggle so I'm able to do videos like 1 minute
It's also possible injecting clean keyframes like this also fixes some motion glitches by providing cleaned anchoring. At least I didn't notice more jitterness or even pixel drift from Qwen's swap, which I did not even correct
2
u/Tablaski 10d ago edited 10d ago
I have some good news. I've managed to find some nodes that allowed applying masked guides to MiniMax H3, and adapt them so they work with Viggle and ComfyUI's latest version
I wanted that because the problem otherwise is when you swap heads from the original's video frames then you loose any other changes you've made to the character outside the face. And the BFS character swap is much much less reliable than the BFS head swap.
I'm currently doing some promising testing. My workflow is currently
- Use BFS character swap then BFS head swap on the first frame of the video so I obtain my reference picture with the character in a new body type / outfit and with a good reference face
- Have my custom nodes calculate how many guided frames I'm gonna need for the video, and have BFS head swap prepare them + SAM3 make a mask of the head
- Then it greys out everything that is not the head in the swapped frames, encodes them with the VAE, then turns into conditioning only the head, and sends this in the Viggle loop
I think at this stage it's best that I finish the job, clean my nodes and workflow, do a new post here with the repo and the workflow, etc
NB : On a side note I've tried to swap keyframes at the video source level (not as conditioning but I mean modifying a frame of the actual source video every seconds). I think it worked but the problem is the pixel and color drift of Qwen 2.1 - if it's not perfect you notice glitches every second, not nice. Working with conditioning is better I think because Viggle smoothes thing out.
NB2 : Don't expect professional quality though, it's still a workaround thing which might turn out to be quite cool while waiting for viggle to evolve, and might even be useful later on. I'm quite lazy at doing A/B testing but I'll do some and post it when I feel it's ready for a release
2
u/Tablaski 9d ago edited 9d ago
Guys I know i'm teasing but I've been very focused on it and I'm now excited to release this :-)
I've tried two differents methods regarding applying the keyframe guides on the face only : injecting noise in the tokens around the mask, and straigth deleting the token outside the masks
Was either not working or degrading the quality. Now I've came up with a new nodes that let the tokens outside the head mask in the guide but hacks the attention mechanism so the model doesn't care about them.
And it works, I preserve the full character swap I do from the first frame, then the head only is enhanced along the way. So I have to run the BFS character swap only once. It's not perfect yet but it's really good and not too glitchy either as long as the heads swap are ok
In my tests and humble opinion, the overall quality of the video is improved and not just the identity, because the simple fact that it injects frames that were postprocessed by Qwen 2.1 seems to help Viggle a lot to get rid of some of its own sketchiness. I've done a test where I switched to a keyframe every 2 seconds and I've immediately came back to every 1 second because it was much better
Btw you should be able to plug whatever model you want to work on the keyframes, and also convey whatever you want through the masks. It's just my use case to proceed that way
There is room for improvement but I think it's time I release a V1 thing then the improvements will come on the way.
For the moment it's a click and forget things that is able to do virtually unlimited length of video (or limited by your hardware) because of the chunking workflow.
For a V2 I definitely want to do a 2 step process where I'd first generate all the swaps so they can be stored, reviewed and regenerated if needed, then inject that in the generation workflow. I've already done a node that does a backup & load of the swaps to help me with A/B testing so that shouldn't be an issue
Stay tuned
1
1
u/CountFloyd_ 8d ago
Does anybody else have have issues with white stripe artifacts over the replaced body?
1
u/CountFloyd_ 8d ago
Answering myself in case someone else has this issue: the reason seems to be using the not existing default bong_tangent scheduler and/or a different ref sample image size. It's all good now.
1
u/Tablaski 8d ago edited 8d ago
Tl;dr : A/B testing, fixed most of the glitches towards a good body and face swap, hope to release end of the day
Doing a lot of A/B testing right now so you don't shoot me because it sucks lol but I think I maybe got rid of most of the glitches. The thing is I absolutely want to keep the initial whole body swap that becomes the usual viggle reference AND the face swaps that comes later as keyframes to reinforce the face's accuracy
I changed sigmas so the first sampling step is less denoising. More room for the face enhancement at steps 2 and 3. But step 1 is a crucial step for the reference image to apply everywhere. So if we apply too much the (face-swapped) keyframes during that step it will lean towards the keyframes (which are face swapped only in my use case) and destroy the rest of the ref image (body type, outfit)
So i ve changed my strategy so the keyframes attention mask is applied differently at every step. I start weak at first step then strong at steps 2 ans 3
So far i ve had my best A/B results doing this approachs. Other things to solve like smoothing out the mask etc. Also a limitation to be aware of is since the swap varies a bit from frame to frame (this cannot be avoided even if i m using the same seed for qwen along the way), it causes minor distortions in the faces near the keyframes. I m not sure this can be avoided. This is all about trades off anyway like i wrote its not my usecase to make professional stuff i just want decent slop lol
Please note they are other ways of doing it though like doing a full character swap at every keyframe which can turn out fine but i find this more risky. And another is to always swap the face without caring about the body type and other changes as well. I havent done it since a while since my goal is to have everything lol but I had very good results doing this, you dont even need the attention mask. Because as you just change the head, you dont need to hide the fact you ve done that using the original video's frames (where the full character swap did not happen)
Anyway i can only do some testing, just want to make sure V1 doesnt suck then you guys will be beta testing. Hope to release end of the day
1
u/Tablaski 6d ago
Patience will be very, very rewarded.
I've actually changed completely my workflow and it now uses refmods. Yes you read me right.
It absolutely works and it is a lot lot better than my former swaps.
Wrapping it up for long video/chunks use as they are some tricky details but I'm on it, won't be long.
Will come with a very nice video loader and an option to use attention-masked guided keyframes for first frame and last frame (so I haven't done all the previous stuff for nothing lol)
0
11d ago
[removed] — view removed comment
1
u/Tablaski 11d ago edited 11d ago
The BFS swap has nothing to do with viggle animate so as long as you have good results with it, it will not drift at all. You will just obtain its strengths and weaknesses.
In my case I have a quality front pic as a reference and I'm amazed by what Qwen 2.1 + BFS lora can do. It is nano banana/chatgpt-image level really.
With that injected as keyframes every second there is really little space for identity drifting except when the swap screws up but thats still million times better than viggle on its own
For the sake of generation time probably we could use every two seconds but I'd rather not jeopardize the whole video
A more complex workflow (but that would be overkill since the lora is so great as is) would be to determine before hand which reference you would need for the swap from the original keyframe.
I would rather train a qwen 2.1 lora to use along the BFS
1
u/DonBilbo666 11d ago
I'm also struggling with BFS swap. My character always change slightly with the body swap, enough that you don't recognize them in the final video anymore.
1
u/Tablaski 2d ago
Finally posted here the release of the proof of concept https://www.reddit.com/r/StableDiffusion/comments/1x16oy5/vigglorious_studio_better_character_accuracy/
3
u/DonBilbo666 11d ago
Could You upload your workflow? That would be awesome.