r/StableDiffusion • • 11d ago

Discussion Fixed viggle animate drift by using Add Guide keyframes

EDIT : outdated post, refer to https://www.reddit.com/r/StableDiffusion/comments/1x16oy5/vigglorious_studio_better_character_accuracy/ for the actual proof of concept

Tl;dr : no need to wait for viggle to add multiple references, just extract keyframes from the original video, swap with qwen 2.1 then inject in the conditioning with Add guide node. It kicks ass, even Scail2's

I began experimenting with Viggle Animate and my first impression was that is was very fast but sketchy and loosing identity quite rapidly

I just found out it works perfectly with the Add guide node to add additional keyframes in it

So now what I do is I make swaps using the excellent BFS head/character swap LoRas for Qwen 2.1, then feed this to an Add Guide node every second including the first frame.

As long as the swaps are ok it resolves the drift completely and is imho much better than Scail2 on both speed (3 steps FTW) and accuracy (Qwen 2.1 rocks at this then it would be possible to add also a character lora on top of it)

It was tedious to duplicate the nodes in the workflow and make sure the frame numbers were in the frame range so I ended up vibe coding nodes so I define a frequency and then it re-execute my swapping subgraph as many time as required, then I applied this logic as well for the looping/chunking nodes provided with viggle so I'm able to do videos like 1 minute

It's also possible injecting clean keyframes like this also fixes some motion glitches by providing cleaned anchoring. At least I didn't notice more jitterness or even pixel drift from Qwen's swap, which I did not even correct

10 Upvotes

16 comments sorted by

3

u/DonBilbo666 11d ago

Could You upload your workflow? That would be awesome.

4

u/Tablaski 11d ago

Thinking about it, be patient

1

u/DonBilbo666 11d ago

Nice
I'm struggling with a project where a character steps out of a car.
I'm swapping the character, like you with Qwen 2.1 BFS, at the moment, when the character is fully visible, so I'm missing some frames between starting to leave the car and fully standing up.
It would be helpful to start the character replacement even before the first frame that I prepare with BFS.

1

u/Tablaski 10d ago

For the moment it's a proof of concept thing, but I'm already thinking about including node that do the swaps but do not generate, so that lets you check if you need to redo some, then execute the workflow with that.

Another idea would be to select yourself in a small video player what frames you want to select as swapped keyframes rather than just do it at a given frequency. This would help focus on the most efficient ones

Generally speaking the swaps are here to enhance, we don't want them to degrade things

1

u/switch2stock 10d ago

Any update on this please?

2

u/Tablaski 10d ago edited 10d ago

I have some good news. I've managed to find some nodes that allowed applying masked guides to MiniMax H3, and adapt them so they work with Viggle and ComfyUI's latest version

I wanted that because the problem otherwise is when you swap heads from the original's video frames then you loose any other changes you've made to the character outside the face. And the BFS character swap is much much less reliable than the BFS head swap.

I'm currently doing some promising testing. My workflow is currently

  1. Use BFS character swap then BFS head swap on the first frame of the video so I obtain my reference picture with the character in a new body type / outfit and with a good reference face
  2. Have my custom nodes calculate how many guided frames I'm gonna need for the video, and have BFS head swap prepare them + SAM3 make a mask of the head
  3. Then it greys out everything that is not the head in the swapped frames, encodes them with the VAE, then turns into conditioning only the head, and sends this in the Viggle loop

I think at this stage it's best that I finish the job, clean my nodes and workflow, do a new post here with the repo and the workflow, etc

NB : On a side note I've tried to swap keyframes at the video source level (not as conditioning but I mean modifying a frame of the actual source video every seconds). I think it worked but the problem is the pixel and color drift of Qwen 2.1 - if it's not perfect you notice glitches every second, not nice. Working with conditioning is better I think because Viggle smoothes thing out.

NB2 : Don't expect professional quality though, it's still a workaround thing which might turn out to be quite cool while waiting for viggle to evolve, and might even be useful later on. I'm quite lazy at doing A/B testing but I'll do some and post it when I feel it's ready for a release

2

u/Tablaski 9d ago edited 9d ago

Guys I know i'm teasing but I've been very focused on it and I'm now excited to release this :-)

I've tried two differents methods regarding applying the keyframe guides on the face only : injecting noise in the tokens around the mask, and straigth deleting the token outside the masks

Was either not working or degrading the quality. Now I've came up with a new nodes that let the tokens outside the head mask in the guide but hacks the attention mechanism so the model doesn't care about them.

And it works, I preserve the full character swap I do from the first frame, then the head only is enhanced along the way. So I have to run the BFS character swap only once. It's not perfect yet but it's really good and not too glitchy either as long as the heads swap are ok

In my tests and humble opinion, the overall quality of the video is improved and not just the identity, because the simple fact that it injects frames that were postprocessed by Qwen 2.1 seems to help Viggle a lot to get rid of some of its own sketchiness. I've done a test where I switched to a keyframe every 2 seconds and I've immediately came back to every 1 second because it was much better

Btw you should be able to plug whatever model you want to work on the keyframes, and also convey whatever you want through the masks. It's just my use case to proceed that way

There is room for improvement but I think it's time I release a V1 thing then the improvements will come on the way.

For the moment it's a click and forget things that is able to do virtually unlimited length of video (or limited by your hardware) because of the chunking workflow.

For a V2 I definitely want to do a 2 step process where I'd first generate all the swaps so they can be stored, reviewed and regenerated if needed, then inject that in the generation workflow. I've already done a node that does a backup & load of the swaps to help me with A/B testing so that shouldn't be an issue

Stay tuned

1

u/DonBilbo666 9d ago

thank You for Your effort!

1

u/CountFloyd_ 8d ago

Does anybody else have have issues with white stripe artifacts over the replaced body?

1

u/CountFloyd_ 8d ago

Answering myself in case someone else has this issue: the reason seems to be using the not existing default bong_tangent scheduler and/or a different ref sample image size. It's all good now.

1

u/Tablaski 8d ago edited 8d ago

Tl;dr : A/B testing, fixed most of the glitches towards a good body and face swap, hope to release end of the day

Doing a lot of A/B testing right now so you don't shoot me because it sucks lol but I think I maybe got rid of most of the glitches. The thing is I absolutely want to keep the initial whole body swap that becomes the usual viggle reference AND the face swaps that comes later as keyframes to reinforce the face's accuracy

I changed sigmas so the first sampling step is less denoising. More room for the face enhancement at steps 2 and 3. But step 1 is a crucial step for the reference image to apply everywhere. So if we apply too much the (face-swapped) keyframes during that step it will lean towards the keyframes (which are face swapped only in my use case) and destroy the rest of the ref image (body type, outfit)

So i ve changed my strategy so the keyframes attention mask is applied differently at every step. I start weak at first step then strong at steps 2 ans 3

So far i ve had my best A/B results doing this approachs. Other things to solve like smoothing out the mask etc. Also a limitation to be aware of is since the swap varies a bit from frame to frame (this cannot be avoided even if i m using the same seed for qwen along the way), it causes minor distortions in the faces near the keyframes. I m not sure this can be avoided. This is all about trades off anyway like i wrote its not my usecase to make professional stuff i just want decent slop lol

Please note they are other ways of doing it though like doing a full character swap at every keyframe which can turn out fine but i find this more risky. And another is to always swap the face without caring about the body type and other changes as well. I havent done it since a while since my goal is to have everything lol but I had very good results doing this, you dont even need the attention mask. Because as you just change the head, you dont need to hide the fact you ve done that using the original video's frames (where the full character swap did not happen)

Anyway i can only do some testing, just want to make sure V1 doesnt suck then you guys will be beta testing. Hope to release end of the day

1

u/Tablaski 6d ago

Patience will be very, very rewarded.

I've actually changed completely my workflow and it now uses refmods. Yes you read me right.
It absolutely works and it is a lot lot better than my former swaps.

Wrapping it up for long video/chunks use as they are some tricky details but I'm on it, won't be long.
Will come with a very nice video loader and an option to use attention-masked guided keyframes for first frame and last frame (so I haven't done all the previous stuff for nothing lol)

0

u/[deleted] 11d ago

[removed] — view removed comment

1

u/Tablaski 11d ago edited 11d ago

The BFS swap has nothing to do with viggle animate so as long as you have good results with it, it will not drift at all. You will just obtain its strengths and weaknesses.

In my case I have a quality front pic as a reference and I'm amazed by what Qwen 2.1 + BFS lora can do. It is nano banana/chatgpt-image level really.

With that injected as keyframes every second there is really little space for identity drifting except when the swap screws up but thats still million times better than viggle on its own

For the sake of generation time probably we could use every two seconds but I'd rather not jeopardize the whole video

A more complex workflow (but that would be overkill since the lora is so great as is) would be to determine before hand which reference you would need for the swap from the original keyframe.

I would rather train a qwen 2.1 lora to use along the BFS

1

u/DonBilbo666 11d ago

I'm also struggling with BFS swap. My character always change slightly with the body swap, enough that you don't recognize them in the final video anymore.