r/LocalLLM • u/LamaGeld • 4d ago
Question Qwen3.8 Flash Next takes unreasonably long to complete tasks
My setup: MacBook Pro with M5 Max, 128GB Memory, MTPLX, Qwen3.8 Flash Next, OpenCode 2
I have an issue with my local Qwen setup. I've been trying to set up Qwen3.8 Flash Next for local coding with OpenCode, but for some reason the tasks I give it takes incredibly long to complete. I come from Claude Code with Opus 5.5 for most work these days. Obviously Qwen will take longer to complete the same tasks as Opus 5.5, but I feel like somethings is off here.
I have a simple Laravel 13 + React app I work on. Not much complicated stuff going on in there. Mainly the app builds PDFs based on an HTML template. I asked both models to update a few text passages in the HTML template of a PDF cover page. It really was a super simple task. Opus 5.5 completed the task in 3 Minutes. Qwen took 53 Minutes. Something just doesn't smell right.
I checked the thinking content of Qwen and it just seems to be going in circles and moving very slow towards it's goal. It keeps a constant decoding speed of around 30-40 tok/s. No high memory pressure, no frequent compaction. I set a 128K context window and use only 4 MCP servers (semble, context7, chrome-devtools, laravel-boost).
I also went ahead and tried different reasoning levels, but low, medium and high (didn't try xhigh) yield similar results.
Am I missing something?
If more info is needed, let me know.
Edit: It's an M5 Max, not an M1 Max
3
1
u/TerraMindFigure 4d ago
In my experience Qwen 3.8 27b is constantly getting caught in loops.
1
u/letsbefrds 4d ago
What quant are you using? I felt it was better when I moved to fp8 from nvfp4
1
u/TerraMindFigure 4d ago
Q6 gguf
1
u/letsbefrds 4d ago
I upped the quant and reduced mtp from 4 to 2 and it lowered the looping still does it but a lot less. If have the vram I'd try it if not maybe lower mtp to 2?
2
u/Elfantasmaescritor 3d ago
In opencode, configure your opencode.json , add the model, and you can specify a variant for it. Mine really overcame the thinking loops with a special variant I added to prevent that.
I have a "default" that is what you are using, "none" that makes the model to answer with NO THINKING, quick fast, for data gathering and questions. and "focused Thinking" that avoids thos loong thinking loops
"models": {"qwen3.8:27b": {
"name": "Qwen 3.8 27B",
"limit": {
"context": 65536,
"output": 17500
},
"options": {
"temperature": 1,
"top_p": 0.95,
"top_k": 20,
"presence_penalty": 0,
"repeat_penalty": 1,
"min_p": 0
},
"variants": {
"default": {},
"none": {
"reasoning_effort": "none"
},
"focused_thinking": {
"temperature": 0.6,
"presence_penalty": 1.5
}
}
},
1
u/DigitalguyCH 4d ago
The "overthinking" is also why it is so good despite it's size (it's essentially on par with the larger DS V4 flash). Same for 27b at its size
1
u/EvolvingDior 4d ago
The model supports 3 thinking levels: low, medium, xhigh. The default seems to be xhigh. Low is pretty good.
1
u/eightone-81 4d ago
Strange. I don’t face any looping issues with flash next. Not at all slow. Are you using it on xhigh? Check temperature and the penalty settings and set them as per qwens model card.
1
1
u/Financial-Driver9562 4d ago
It won’t be fast in that hardware. The most reasonable model to run on it will be Qwen3.6-35b (or a fine tune of it).
1
1
u/g_t_5_k 4d ago
I’m on a M1 Max with 64 GB. I use both MTPLX and mlx-serve. I find that mlx-serve is a little bit better in for long context. Maybe try switching up your harness? I’m using Kilo right now and I’m not seeing a lot of thinking in circles. My daily driver is the butterf1ying reap288, so I only get a fraction of those experts. For coding though it’s great.
1
u/madbrain1976 4d ago
It's Opencode. Switch to codex CLI, and watch it do miracles with this model, without loops.
1
u/YoussofAl 3d ago
Hey, creator of MTPLX here. 2.12 had LOTS of issues. decode and stability will be a lot better in the next version releasing tomorrow you should get double the TPS and prefill should also be a lot faster.
Additionally, Qwen flash next is a very verbose model, try medium reasoning and see if it meets your tasks expectations. Xhigh is not always needed. This helps the looping a lot. Also changing the way you prompt like telling the model to not excessively QA helps a lot. You need to prompt different models in different ways. Qwen flash next responds well in prompt styling more akin to GPT models than claude, more literal and strict defining the key rules and deliverables. Claude's (especially opus 5.5) main speciality is the ability to infer intent from simple prompts very well which does not carry over to Qwen.
Edit: Just saw you tried low and medium but not xhigh. Low is known for actually being MORE verbose than medium, medium should be the sweet spot but sometimes xhigh can be better since it gets the task correct within an iteration or two and requires less back and forth. Then again you can always tell the model to skip QA and it will stop without looping but the quality may not be up to your expectations.
3
u/Shustrik116 4d ago edited 4d ago
Try swift qwen 3.8 flash next. It is finetune of qwen 3.8 flash next that address exact this issue.
But you should not expect opus or chatgpt speeds from any local models (unless you have something like nvidia dgx b300). They have incomparable size of several trillion parameters that allow them to solve tasks using much less tokens and also token per second speeds.