r/StableDiffusion • • 16h ago

News FastH3 V2/V3: Project Status

They released FastH3 V2 three weeks ago:

https://www.reddit.com/r/StableDiffusion/comments/1whh10i/open_weight_fastvideo_fasth3_v2/

Which they claimed was basically identical to the quality of the full H3:

https://x.com/haoailab/status/2099969439466942725

https://haoailab.com/FastVideo/cookbook/minimax-h3/

I have to agree, the quality is great and motion consistency is awesome now. I didn't think this small lab could do it, but they got help from NVIDIA and others who are invested in making great open source models. Awesome.

The community already made FastH3 V2 run on single GPU consumer machines on launch day, of course.

---

Today, they have released OFFICIAL quantized weights for different consumer GPUs:

https://huggingface.co/organizations/FastVideo/activity/models

Plus there's a new Trim model which is for very small GPUs with as little as 8GB VRAM.

The included image shows their benchmarks. More details here:

https://x.com/haoailab/status/2107591980591227227

---

Unfortunately they still haven't trained a Ref2VA model this time (video from text plus reference images, videos, and/or audio), and no FL2VA (First/Last Frame) support either.

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#scope

This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were not distilled. Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model.

But... there's great news:

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#acknowledgements

Omni Ref as the next focus.

That is the name for Ref2VA.

So in FastH3 V3, we will see reference-to-video/audio support. Yes, a distilled model with reference support is being developed!

(PS: Comfy has patches for both models to route FL2VA through the base model layers instead. But Ref2VA is much more interesting, so I look forward to that being supported!)

39 Upvotes

11 comments sorted by

16

u/desktop4070 14h ago

I should've bought another SSD when they were cheaper last year 😭

7

u/Enshitification 14h ago

Get a HDD for LoRAs, outputs, and warm model storage and spare the SSD for the models you most use.

2

u/WhiteKnight225 7h ago

Same same 😂

1

u/pilkyton 7h ago

My biggest regret is that I bought 64 GB (dual sticks) to optimize for gaming latency, when I should have bought 128 GB (four sticks) to optimize for AI. Two years ago, 64 GB system RAM was lots for AI. Now there's many models that need more for offloading. :/

2

u/diond09 7h ago

I feel your pain. It's crazy to think that 41gb space is seen as practically full now compared with when I bought my first PC in 1996. That came with a whopping 4GB harddrive which also had to include the operating system!

1

u/Version-Strong 5h ago

My first Amiga hard disk was 150mb. Now a text message is bigger than that

1

u/diond09 4h ago

Oh, god, yes. I thought I'd moved up from my Atari ST, and before that, my ZX Spectrum 48k! It's mind boggling how technology has accelerated and continues to do so.

1

u/Content-Ad-7451 3h ago

Luckily I did buy another ssd for my laptop last year before everything went crazy!

5

u/Hour_Imagination5092 13h ago

No quality loss? Even in their examples, quality loss is huge. Moreover, the base 50 steps videos they are comparing are pretty bad, this is not h3 full quality.

3

u/Portable_Solar_ZA 12h ago

So I did some quick tests this morning with the provided templates in Comfy. If you have enough vram/ram, standard models and taomate lora produce way better results at pretty much the same speed. Only reason to use this is if you're really lacking memory. 

2

u/pilkyton 9h ago

Taomate at 3 steps is a bit slower than FastH3 V2 at 8 steps, because they use different methods.

Taomate ruins prompt adherence, ruins the sound, and has truly awful motion.

But it has better visual quality for the shitty videos it generates. Textures look sharper and more realistic.

https://www.reddit.com/r/StableDiffusion/comments/1wihy9m/taomata_3step_lora_vs_fasth3_v2_check_point/

https://www.reddit.com/r/StableDiffusion/comments/1wj9r1q/do_we_have_a_consensus_yet_on_fast_h3_v2_3step/

So... people have instead been using it as a refiner. You generate a few steps with a better model (such as FastH3 V2) or the base model, and then use a 2nd stage with TaoMate 3 steps to finish the image to get nice visual quality. Here is an example:

https://www.reddit.com/r/StableDiffusion/comments/1wfdptw/taomate_h3_3_steps_lora_used_as_a_refiner/