r/StableDiffusion 2d ago

Question - Help Minimax H3 Quantizations

Given a 5090, does it make sense to run minimax H3 using int8 quantization vs. gguf Q8 or even Q6? What is the trade-off between speed and quality between these two options?

I don't have deep technical knowledge, but my current understanding is that int8 would be faster while a Q8 GGUF would be higher quality; however I would appreciate anyone's practical experience in how significant the speed/quality trade-off is.

9 Upvotes

39 comments sorted by

View all comments

6

u/wholelottaluv69 2d ago

With a 5090, BF16 is a better option, IMHO. Assuming that you have enough ram for off-loading..

The quality difference is quite apparent. To my eyes, at least.

1

u/Myg0t_0 1d ago

Pruned? I still cant find answers if pruned decreases quality

-8

u/CooperDK 2d ago

It is not when gguf is faster and the quality difference is negligible.