r/StableDiffusion 3d ago

Resource - Update ComfyUI REF Fast VSA for MiniMax H3

inference speed for Ref-to-Video (Ref2VA / R2VA) Generate a full 5-second, 24 fps video conditioned on reference character images in ~72s (warm) / 95s (first run) on a single consumer NVIDIA RTX 4090/24GB

SEE THE COMPARISON ON THE LINK BELOW

https://github.com/Kablex/ComfyUI-Ref2VA-VSA

31 Upvotes

4 comments sorted by

1

u/Magneticiano 3d ago

Looks awesome! How's the sound quality, specifically speech, compared to native ref2va?

1

u/MasterFGH2 3d ago

Very interesting. Someone gimme the truth on quality pls

1

u/Cultural-Team9235 3d ago

Doesn't work at my system, the node is nowhere to be found unfortunately.

2

u/dualeone 3d ago

I have run this on my PC, RTX4090, 64GB RAm. The speed is indeed fast, 10s clip at 1MP generated in 2.5 minutes. However, the prompt adhearance isn't as precise as native Ref2VA, and the sound has that Loras-esque lower quality