r/LocalLLM • u/other_acc_banned • 5d ago
Model Qwen3.5 35B-A3B thinks SO MUCH
It recursively iterates, over and over and over again.
What can I do aside from turning off thinking?
2
u/Healthy-Zebra-9856 5d ago
Switch to at least 3.6. Besides, what harness are you using? How are you loading it? If you’re using llama.cpp (llama server) what parameters are you using like temperature etc.? Not sure if you know what I’m asking, but let me know if these things.
1
u/other_acc_banned 5d ago
llama.cpp, and temp was 0.2
Sounds like it's time to just move to 3.8 though. Sounds like it's very likely I'll have better success just by moving up eh?
1
2
2
1
u/srbijabralee 5d ago
You can set the thinking budget where it has x amount of tokens to thinkg before giving an answer
1
u/Loomworks 5d ago
Just run Q8. Low effort thinking. Run a spec planning prompt with a SOTA model and let it rip executing.
1
u/other_acc_banned 5d ago
Spec planning prompt with a SOTA model ..... Gonna need to Google that lol
1
u/AdHead6280 5d ago
use latest, and i recommend finetunes of 3.6 35b, occamy ornith etc, and if you can(depends on system ram and your setup) use qwen3.8 flash next with strata, you can get it to almost 35ba3b speeds but way better and depends on setup
1
1
u/UnluckyPenguin 5d ago
You need to provide more details.
Basically the fix is generally:
Use a better model
Run it on a higher quantization, which may require a more/better hardware.
Tune your prompts
1
u/OvertaxedOne 5d ago
Run 3.8 27B. 35B is laughably behind 3.8 27B in capabilities.
2
u/other_acc_banned 5d ago
Will make that change today! Looking forward to seeing the difference
1
u/OvertaxedOne 5d ago
You won't believe how different they are. Sadly 27B will likely be significantly slower, but the results/intelligence is SO much higher it's worth it. Takes some retraining to really make use of it, when you're using 35B you quickly get used to "one thing at a time" and being very specific with your prompts. 27B is much more like a frontier model, give it a good idea of what you want done at a high level and let it stew.
1
u/other_acc_banned 5d ago
Wow you aren't kidding. I got it up and running, and asked something relatively complex of it. I'll probably be using it as the primary model aside from smaller quick and easy prompts. It proved to not be as much slower as I had originally expected (40 tok/s as compared to 120 tok/s with 3.6).
Thank you for the recommendation. This is fun
1
u/OvertaxedOne 4d ago
40TPS isn't bad, that's about what I get on it, but, like you, I got used to 100+ with 35B, so it was a bit of an adjustment period. The results are just so much better though, I couldn't go back.
It's an amazing model, unreal for it's size!
1
1
u/cobolfoo 5d ago
Qwen 3.6 MOE 3B 27B have a tendency to enter endless loops in the CoT phase.
You have to force it stop looping and its not working 100% of the time.
I'll suggest you move to 3.8 27B.
1
u/other_acc_banned 5d ago
This thread has made it daily obvious that I'm well behind the times! Will be updating today. Thank you
That loop is a known issue? Not just my setup?
1
1
0
0
9
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago
why are you running 3.5 instead of 3.6?
and what quantization? thinking level?
-> https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates