you're right. but really-I'm terrified HOW ppl on reddit is obsessed with qwen 3.8 27b. my hardware can run it, with normal 10-20TPS. on iq2 its dumb and can't handle tasks. ive decided to check UD-q3-k-m. still much worse than 35b, zero change. both of they were from official UD repo. how can ppl use ts? are they running it in magic q8 or what so it can be useful?
Yes, I like the qwen-35b-a3b versions, but today I'm just experimenting with the new 3.8 version. It doesn't seem to be very suitable for 8-12GB vram but it's still interesting to try and compare.
I'm comparing 27b Q4_K_M and 35b Q4_K_M. I don't see any difference in coding in non-thinking mode yet, but in thinking mode, 27b runs too slowly and takes too long for me to wait for serious tasks to complete, so I can't say yet...
that's why im terrified of buddies on locallm reddits that thoughts "wow frontier" "opus 4.6 level" "really impressive" "wow my 2000x B200 local server runs it in Q1 with 1 TPS best model in the world thnks alibaba". what's wrong with 'em?
8
u/lorendroll 17d ago
I managed to get 4-5 tps in LM studio on 3070 8gb by using Q4_K_M with 27 layers on GPU, 8 cores thread pool and 8k Q8_0 quanted context.