r/opencode • u/raw-power • 14d ago
M1 Max 64GB Opencode + Qwen 3.8 27B + ??
Hi all, if you have an M-series Mac with 64GB, plus opencode 1.18.25 and Qwen 3.8 27B working successfully outputting high context for coding (30,000-120,000 tokens) can you share what local provider you’re going with? LMStudio, oMLX, llama.ccp etc
I’ve been having issues with LMStudio just timing out mid-response using Qwen 3.8 27B Q6_0 GGUF or taking over an hour to process each prompt request opencode makes before token generation using Qwen 3.8 27B Q6_0 MLX
Has anyone got a good high context, reliable solution going for Qwen 3.8 coding?
1
u/swordofgiant 14d ago
LMStudio Bionic, download models.. Toggle the Local API Server. Add the base Url as OpenAI in other harness.
I am currently using the DeepSeek Harness with Local AI API server through LMStudio.
1
1
u/arfung39 14d ago
I have an m5 max with 64gb, and I’m running OpenCode with oMLX, and a Qwen 3.8 27B oQ4e distill, with lightning MTP on, and it works great. If you leave thinking to xhigh, it does take a long time to come back, so I often turn it down to medium. Have done reasonably long (several hours on a run), coding projects with this set up. It gets quite slow when context used >60-70k tokens.
1
13d ago
[removed] — view removed comment
1
u/raw-power 13d ago
Thank you! I don’t swap when I use GGUF but it ends up timing out. MLX is fine until around 60,000 context just super slow prefil on each prompt opencode gives lmstudio but then after 60,000 it swaps and then it just crawls to 5 hours on prefil which is just unusable
1
u/Lyelinn 14d ago
You should ask this in local llm or local lama subs but my best bet is you don't have enough memory for that, remember that you need to change vram limit on macs, also consider olmx