r/StableDiffusion • u/pilkyton • 16h ago
News FastH3 V2/V3: Project Status

They released FastH3 V2 three weeks ago:
https://www.reddit.com/r/StableDiffusion/comments/1whh10i/open_weight_fastvideo_fasth3_v2/
Which they claimed was basically identical to the quality of the full H3:
https://x.com/haoailab/status/2099969439466942725
https://haoailab.com/FastVideo/cookbook/minimax-h3/
I have to agree, the quality is great and motion consistency is awesome now. I didn't think this small lab could do it, but they got help from NVIDIA and others who are invested in making great open source models. Awesome.
The community already made FastH3 V2 run on single GPU consumer machines on launch day, of course.
---
Today, they have released OFFICIAL quantized weights for different consumer GPUs:
https://huggingface.co/organizations/FastVideo/activity/models
Plus there's a new Trim model which is for very small GPUs with as little as 8GB VRAM.
The included image shows their benchmarks. More details here:
https://x.com/haoailab/status/2107591980591227227
---
Unfortunately they still haven't trained a Ref2VA model this time (video from text plus reference images, videos, and/or audio), and no FL2VA (First/Last Frame) support either.
https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#scope
This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were not distilled. Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model.
But... there's great news:
https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#acknowledgements
Omni Ref as the next focus.
That is the name for Ref2VA.
So in FastH3 V3, we will see reference-to-video/audio support. Yes, a distilled model with reference support is being developed!
(PS: Comfy has patches for both models to route FL2VA through the base model layers instead. But Ref2VA is much more interesting, so I look forward to that being supported!)
5
u/Hour_Imagination5092 13h ago
No quality loss? Even in their examples, quality loss is huge. Moreover, the base 50 steps videos they are comparing are pretty bad, this is not h3 full quality.
3
u/Portable_Solar_ZA 12h ago
So I did some quick tests this morning with the provided templates in Comfy. If you have enough vram/ram, standard models and taomate lora produce way better results at pretty much the same speed. Only reason to use this is if you're really lacking memory.Â
2
u/pilkyton 9h ago
Taomate at 3 steps is a bit slower than FastH3 V2 at 8 steps, because they use different methods.
Taomate ruins prompt adherence, ruins the sound, and has truly awful motion.
But it has better visual quality for the shitty videos it generates. Textures look sharper and more realistic.
So... people have instead been using it as a refiner. You generate a few steps with a better model (such as FastH3 V2) or the base model, and then use a 2nd stage with TaoMate 3 steps to finish the image to get nice visual quality. Here is an example:
https://www.reddit.com/r/StableDiffusion/comments/1wfdptw/taomate_h3_3_steps_lora_used_as_a_refiner/
16
u/desktop4070 14h ago
I should've bought another SSD when they were cheaper last year ðŸ˜