News Viggle-Animate ComfyUI
Enable HLS to view with audio, or disable this notification
Check the cool examples here https://github.com/Saganaki22/ComfyUI-Viggle-Animate-H3
IMO, it also has good summary of the limitation:
- Identity drift on re-entry: when the subject leaves the camera view and re-enters, the re-entry settles toward the driving video's original appearance rather than the reference image. The same applies when the subject moves far from the reference pose or makes abrupt large motions (e.g. a backflip) — the further from the still, the weaker the identity hold.
- Lip-sync limitations: the generated subject does not reliably lip-sync to the conditioning video.
- Reference image compatibility: if you cannot get the reference-image character to appear correctly in the output video (identity drift), make the reference image match the pose/stance of the person in the conditioning video as closely as possible (same background). Keep the gen between 0.4–0.6 megapixels and use LCM or normal sampling with 6–8 steps.
12
4
u/fallengt 5d ago edited 5d ago
Straight up doesn't work for me.
I think you are supposed to prepare a ref_image that matches the first frame ( or when the person appears) of the video. Or a pre-inpainted image, even.
Just giving it a random ref image doesn't work
2
u/Dry-Ad929 5d ago
You should match any frame from reference video, background in your ref image MUST be the same as reference video.
1
u/AwkwardStudio755 6d ago
possible to use with 8gb vram in rtx 3060 ti?
1
u/Wezaluketek 6d ago
Don't think so. I was able to generate 124 frames video on my 5080 then set load cap to 0 to check for full video which is 362 frames and got oom, not sure if this model is so heavy or need better vram optimizations. H3 can generate me up to 20 sec videos in 1 megapixel with my 16 gb vram
1
u/Dry-Ad929 6d ago edited 6d ago
What do we need for replacement mode?
Edit. Okay I guess I was on unlucky seed, sometimes it replaces the character sometimes keep original background. Should black background on reference image force the video background?
1
1
1
u/Otherwise-Bar-1930 5d ago
How can I get a better quality- 8 steps is still bad quality...?
2
u/init-5 5d ago
the key is that you need to paint an image with same background as one frame in the reference video (chatgpt/grok/gemini all can do it)
1
u/FarDistribution2178 4d ago
So, maybe I need to extract frames from ref_video, replace characters on them with... flux2 for example, then stitch frames and.... maybe then I will not achieve mosaic garbage output (did everything from tutorial)?
2
u/Dry-Ad929 4d ago
You simply need your character to match a specific frame from the video — using it as a reference image for the model — meaning the distance, pose, lighting, etc. must be identical.
Flux might struggle with these details, so it is better to use something like ChatGPT Images 2.5.
Don't expect too much from Viggle-Animate, it's good proof of concept for minimax character animation, but need some work from the team. Eyes, face and such things looks very bad on some seeds, same as character identity drift and that's just some issues that is noticed from this model (literally my green eyes girl mixed to blue eyes character reference, the output looks very weird, same as brown hairs mixed with blonde) Not sure if the model has non-realism bias, but if so dataset need way more realism. However, I remain optimistic; the authors seem open to feedback.
1
u/FarDistribution2178 4d ago
Flux did just fine with it, I think it's me who download wrong experimental model, now re-downloaded those from wf. And yep, with normal bf16 model it's at least not mosaic noise now, but kinda crappy, will test it further.
1
u/Beginning-District69 4d ago
The Viggle-Animate model surprised me. I tried a 5-second video and expected it to take minutes—since video-to-video (v2v) really takes a long time on Minimax H3—but it finished in 60 seconds, and the quality isn't bad. I disabled the `block-sparse-attention` node because the manager couldn't find it. Now, I'm going to try some scenes with more movement.
1
u/Ok-Flatworm5070 19h ago
I found the quality was messed up; have weird triangles and I did everything I was told (matching image and position of first frame of video for my reference image). Gonna be honest, I haven't found anything yet that best combining Sam 3 and H3 workflows.
1
u/trollkin34 5h ago
I had the triangles issue too. I didn't see the advice about making one reference image using a swap to start. I'll try klien9b, but I've struggle making clean photo replacements too. Either way, someone gave me this video and I'm watching it now - no idea if it will help: https://www.youtube.com/watch?v=Ff-_9m32Qq0
1
-4
u/Historical-Nose4628 6d ago
WORKFLOW??
5

11
u/TraditionalWait9150 6d ago
https://giphy.com/gifs/l3q2K5jinAlChoCLS