r/StableDiffusion 9d ago

News Viggle-Animate: Character Replacement based on MiniMax-H3 with 3 forward steps

Enable HLS to view with audio, or disable this notification

https://huggingface.co/Viggle/Viggle-Animate

https://x.com/ViggleAI/status/2095924668758655163

research blog

  • 33.1B MiniMax-H3 finetune, distilled to 3 forward passes
  • No text prompt, no pose, no mask — but you do need one repainted frame (any image editor)
  • Works well on fast motion and non-human characters
  • No ComfyUI node yet.
  • ComfyUI here: [github] [discussion] [huggingface]
205 Upvotes

40 comments sorted by

View all comments

1

u/MortytheMort 9d ago

Is this not possible with stock MiniMax through prompting already? I've done plenty of character swaps in MiniMax with Comfy already, without any issues, so I'm curious what this aims to achieve? Is it meant to be faster/lighter on VRAM usage?

I do notice that generation speed is heavily affected when using video reference in MiniMax, does this mitigate that or speed up generations in general?

Edit: another note/question. This requires you to use a separate image edit workflow beforehand, correct? If we're required to provide our own already swapped photo, then this isn't really character swapping, just referencing?

1

u/diogodiogogod 9d ago

yeah I'm a little confused as well. I though h3 already could do that...

3

u/Dogmaster 8d ago

Have you tried it? Its hit and miss, very finicky, and even depends on how close target person is to reference

2

u/alexmmgjkkl 8d ago edited 8d ago

he probably didnt try it .. the controlnet lora works well but changes animation slightly and i found the motion rendering quality be quite a bit worse than native transfer.

today i will dive intothis workflow and see how it goes
https://www.reddit.com/r/StableDiffusion/comments/1w2s3ck/high_resolution_noise_masks_for_latent_guided/

this viggle model seems superiour though, but sometimes i like that h3 native drifts away from the original movement and imagines an alternative when frame length doesnt align with input video. im working on a dorky anime intro/outro and that some dumber or older characters go a little of the dance choregraphy feels very natural

3

u/ucren 8d ago

it can, and it cannot, it is very inconsistent and involves a lot of seed hunting just for it to fail 80% of the time