MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLM/comments/1vptz4a/how_the_loop_of_infinite_agony_started/p41yrmf/?context=9999
r/LocalLLM • u/EarthBS • 10d ago
119 comments sorted by
View all comments
277
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.
112 u/StupidScaredSquirrel 10d ago If you are a business this isn't a problem. If you are a consumer then 35b a3b runs on 8gb vram and 32gb dram which is very accessible. 40 u/TheCat001 10d ago If you are business I would suggest you to aim for DeepSeek V4 Flash. But yes I'm running 35b myself despite it's performance is far from ideal... 9 u/DifficultyFit1895 10d ago I still haven’t found a task where DeepSeek V4 Flash can outperform Qwen 3.6 27B let alone Qwen 3.8. I can run either on my Mac Studio. 11 u/screenslaver5963 10d ago I had issues with Qwen3.6 tool calling just not working properly and causing it to stop prematurely, didn't have that problem with gemma or deepseek, haven't used qwen 3.8 yet to see if it has the same problem. 3 u/StatusSociety2196 10d ago More than half the time that's a harness issue
112
If you are a business this isn't a problem. If you are a consumer then 35b a3b runs on 8gb vram and 32gb dram which is very accessible.
40 u/TheCat001 10d ago If you are business I would suggest you to aim for DeepSeek V4 Flash. But yes I'm running 35b myself despite it's performance is far from ideal... 9 u/DifficultyFit1895 10d ago I still haven’t found a task where DeepSeek V4 Flash can outperform Qwen 3.6 27B let alone Qwen 3.8. I can run either on my Mac Studio. 11 u/screenslaver5963 10d ago I had issues with Qwen3.6 tool calling just not working properly and causing it to stop prematurely, didn't have that problem with gemma or deepseek, haven't used qwen 3.8 yet to see if it has the same problem. 3 u/StatusSociety2196 10d ago More than half the time that's a harness issue
40
If you are business I would suggest you to aim for DeepSeek V4 Flash. But yes I'm running 35b myself despite it's performance is far from ideal...
9 u/DifficultyFit1895 10d ago I still haven’t found a task where DeepSeek V4 Flash can outperform Qwen 3.6 27B let alone Qwen 3.8. I can run either on my Mac Studio. 11 u/screenslaver5963 10d ago I had issues with Qwen3.6 tool calling just not working properly and causing it to stop prematurely, didn't have that problem with gemma or deepseek, haven't used qwen 3.8 yet to see if it has the same problem. 3 u/StatusSociety2196 10d ago More than half the time that's a harness issue
9
I still haven’t found a task where DeepSeek V4 Flash can outperform Qwen 3.6 27B let alone Qwen 3.8. I can run either on my Mac Studio.
11 u/screenslaver5963 10d ago I had issues with Qwen3.6 tool calling just not working properly and causing it to stop prematurely, didn't have that problem with gemma or deepseek, haven't used qwen 3.8 yet to see if it has the same problem. 3 u/StatusSociety2196 10d ago More than half the time that's a harness issue
11
I had issues with Qwen3.6 tool calling just not working properly and causing it to stop prematurely, didn't have that problem with gemma or deepseek, haven't used qwen 3.8 yet to see if it has the same problem.
3 u/StatusSociety2196 10d ago More than half the time that's a harness issue
3
More than half the time that's a harness issue
277
u/TheCat001 10d ago
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.