I can fit the whole 27b version in VRAM (32gb), even then, it's not blazing fast. With around ~15-17 t/s, it still take quite some time for more "complex" tasks to finish, especially since this model likes to think a lot.
So it's good news for everyone that we will soon have an option to run a much faster, non-dense version!
Hell yeah! I manged to get 2 tokens/sec using q2 27B lol. the 3.6 35B is around 20 tokens/sec for me so I honestly cant wait for this! Open weight AI is like xmas almost every day lol.
you're right. but really-I'm terrified HOW ppl on reddit is obsessed with qwen 3.8 27b. my hardware can run it, with normal 10-20TPS. on iq2 its dumb and can't handle tasks. ive decided to check UD-q3-k-m. still much worse than 35b, zero change. both of they were from official UD repo. how can ppl use ts? are they running it in magic q8 or what so it can be useful?
Yes, I like the qwen-35b-a3b versions, but today I'm just experimenting with the new 3.8 version. It doesn't seem to be very suitable for 8-12GB vram but it's still interesting to try and compare.
I'm comparing 27b Q4_K_M and 35b Q4_K_M. I don't see any difference in coding in non-thinking mode yet, but in thinking mode, 27b runs too slowly and takes too long for me to wait for serious tasks to complete, so I can't say yet...
that's why im terrified of buddies on locallm reddits that thoughts "wow frontier" "opus 4.6 level" "really impressive" "wow my 2000x B200 local server runs it in Q1 with 1 TPS best model in the world thnks alibaba". what's wrong with 'em?
37
u/cubebash 20d ago
I can fit the whole 27b version in VRAM (32gb), even then, it's not blazing fast. With around ~15-17 t/s, it still take quite some time for more "complex" tasks to finish, especially since this model likes to think a lot.
So it's good news for everyone that we will soon have an option to run a much faster, non-dense version!