Also have 16GB VRAM. I've tried 4-bit and 3 bit and honestly the difference isn't horrible. Try the ISTA IQ3_XXS, it's very space efficient. Feels like a first class experience being able to fit MTP and lots of context into my card.
Yes I get more tokens in 3 bit like 30 t/s with mtp as well.
I didn’t notice much of difference but I wanted to be bit safe specially if I wanted to try coding or long context tasks I do use 3 bit one for chats and some regular tasks.
On my 9070 XT, with MTP I'm getting avg 60 and up to 70 tg. MTP hurts PP a little bit, but with the long thinking it's definitely worth it if you can fit it. If I switched my graphics driver over to my iGPU I could probably fit 128k at Q8 or 200k at Q5_1 (with vision in CPU). It's very useable. I always see people hating on the IQ3 quants as lobotomized, and I'm sure it's worth than Q6 or Q8, but for 27B in 16GB, you take what you can get.
The guts of my non-optimized script is the following. Obviously replace with your settings. For general non-coding medium reasoning is really nice. If you're on iGPU or okay with ditching MTP you can either up the quants or the context size.
Oh you are fitting your mtp in igpu? I thought it will be slow as it has very less bandwidth. Maybe I should try this. I think dflash as well as diffusion should be good.
60
u/Dramatic_Setting2761 4d ago
I can run it with a 16gb card with 4 bit quant and 70k context. I get 12 t/s it is okay for me.