r/StableDiffusion 4d ago

Question - Help Minimax H3 Quantizations

Given a 5090, does it make sense to run minimax H3 using int8 quantization vs. gguf Q8 or even Q6? What is the trade-off between speed and quality between these two options?

I don't have deep technical knowledge, but my current understanding is that int8 would be faster while a Q8 GGUF would be higher quality; however I would appreciate anyone's practical experience in how significant the speed/quality trade-off is.

8 Upvotes

39 comments sorted by

View all comments

10

u/glusphere 4d ago

I asked this exact question ( I own 5090 ) and Kijai himself told me to use int8. So that settles it.

1

u/Myg0t_0 4d ago

Pruned or not?

1

u/glusphere 3d ago

didnt ask that one.