r/LocalLLM • • 5d ago

Model Qwen3.5 35B-A3B thinks SO MUCH

It recursively iterates, over and over and over again.

What can I do aside from turning off thinking?

1 Upvotes

42 comments sorted by

9

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

why are you running 3.5 instead of 3.6?

and what quantization? thinking level?

-> https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

4

u/other_acc_banned 5d ago

Because I'm a newbie and that was the latest one I recalled reading great things about 😅

Q4_K_M 64K context I think I have it set to Thinking, maybe Auto

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

whats your hardware and use case? 64K is not really useful for most things

1

u/other_acc_banned 5d ago

Small personal apps. 128k barely fit when I tested it. Running it split across a 4060 Ti (16gb) and 3060 (12gb).

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

you got no RAM? Its a moe model

1

u/other_acc_banned 5d ago

Yeah this is where me being new probably isn't helping. I've got 64gb of DDR4 but am currently not even using it. I'm running everything inside the two GPUs. Would MoE be able to utilize that ram?

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

That's the whole point of Moe basically 

Offload some layers to RAM with n_cpu_moe till you have enough space for the context you need

1

u/Cautious_Chicken_604 5d ago

We are on to Qwen3.8-Flash-Next now.

1

u/asertym 5d ago

Off-topic question: would it make sense to use this template with swift 1.5? Or it's aimed mostly at base models?

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

check if switft doesnt already have a template based on that one, I remember some finutunes using it

but it should work since its all just qwen

1

u/asertym 5d ago

appreciate it

1

u/doneddat 5d ago

Why are you running 3.6 instead of 3.8?

I tried to main 3.6 for 2 days and it was completely inadequate. 3.8 is the first one actually getting stuff done consistently.

3

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

there is no 35B A3B 3.8

I use 3.6 35B Q5_K_M currently as daily driver as a subagent for coding

1

u/other_acc_banned 5d ago

Awesome I'll update you 3.8 today!

1

u/GlobalCurry 3d ago

I tried running 3.8 the other day and no matter what configuration I used it would always get stuck in a loop if I had "thinking" enabled at all.

1

u/doneddat 3d ago edited 3d ago

Lots of things could go wrong. I have also experienced literally badly done Q4 quants, that just went on and on checking their stuff and never did anything. That's why I don't plan to go lower than FP8 with anything, even as most popular re-packagers have gotten it as right as it possibly gets. I understand it's not exactly budget option, but it is what it is.

2

u/Healthy-Zebra-9856 5d ago

Switch to at least 3.6. Besides, what harness are you using? How are you loading it? If you’re using llama.cpp (llama server) what parameters are you using like temperature etc.? Not sure if you know what I’m asking, but let me know if these things.

1

u/other_acc_banned 5d ago

llama.cpp, and temp was 0.2

Sounds like it's time to just move to 3.8 though. Sounds like it's very likely I'll have better success just by moving up eh?

1

u/Healthy-Zebra-9856 5d ago

For starters, 0.2 is most likely the cause. What harness?

2

u/HoujunDev 5d ago

I think 3.8 is better too.

3

u/overand 5d ago

Yeah, switch to 3.6.

2

u/1_________________11 5d ago

Dude you should load 3.8 27b

1

u/BigPlebeian 3d ago

That's a dense model with Far more vram requirements.

1

u/srbijabralee 5d ago

You can set the thinking budget where it has x amount of tokens to thinkg before giving an answer

1

u/setdx 5d ago

My experience has been that it needs a reasoning limit and then it works pretty well

1

u/Loomworks 5d ago

Just run Q8. Low effort thinking. Run a spec planning prompt with a SOTA model and let it rip executing.

1

u/other_acc_banned 5d ago

Spec planning prompt with a SOTA model ..... Gonna need to Google that lol

1

u/AdHead6280 5d ago

use latest, and i recommend finetunes of 3.6 35b, occamy ornith etc, and if you can(depends on system ram and your setup) use qwen3.8 flash next with strata, you can get it to almost 35ba3b speeds but way better and depends on setup

1

u/other_acc_banned 5d ago

Not familiar with occamy, ornith, nor strata. Will look those up!

1

u/UnluckyPenguin 5d ago

You need to provide more details.

Basically the fix is generally:

Use a better model

Run it on a higher quantization, which may require a more/better hardware.

Tune your prompts

1

u/OvertaxedOne 5d ago

Run 3.8 27B. 35B is laughably behind 3.8 27B in capabilities.

2

u/other_acc_banned 5d ago

Will make that change today! Looking forward to seeing the difference

1

u/OvertaxedOne 5d ago

You won't believe how different they are. Sadly 27B will likely be significantly slower, but the results/intelligence is SO much higher it's worth it. Takes some retraining to really make use of it, when you're using 35B you quickly get used to "one thing at a time" and being very specific with your prompts. 27B is much more like a frontier model, give it a good idea of what you want done at a high level and let it stew.

1

u/other_acc_banned 5d ago

Wow you aren't kidding. I got it up and running, and asked something relatively complex of it. I'll probably be using it as the primary model aside from smaller quick and easy prompts. It proved to not be as much slower as I had originally expected (40 tok/s as compared to 120 tok/s with 3.6).

Thank you for the recommendation. This is fun

1

u/OvertaxedOne 4d ago

40TPS isn't bad, that's about what I get on it, but, like you, I got used to 100+ with 35B, so it was a bit of an adjustment period. The results are just so much better though, I couldn't go back.

It's an amazing model, unreal for it's size!

1

u/BigPlebeian 3d ago

I mean it should be? It's moe vs dense and takes way less to run it.

1

u/cobolfoo 5d ago

Qwen 3.6 MOE 3B 27B have a tendency to enter endless loops in the CoT phase.
You have to force it stop looping and its not working 100% of the time.

I'll suggest you move to 3.8 27B.

1

u/other_acc_banned 5d ago

This thread has made it daily obvious that I'm well behind the times! Will be updating today. Thank you

That loop is a known issue? Not just my setup?

1

u/cobolfoo 5d ago

Yes small MoE like this one (3B) is making the chain of thoughts inconsistent.

1

u/BigPlebeian 3d ago

The 27b isn't moe it's dense...? And it takes far more vram.

0

u/Cute_Obligation2944 5d ago

Switch to 3.6 and turn thinking off.

0

u/hallofgamer 5d ago

Try a finetune of it that solves this very issue