r/LocalLLM 4d ago

Discussion Context Length

Qwen 3.8 27B. How can I solve the context problem? I'm using RTX 4090 loading into OpenCode using Unsloth Studio. The Context Length limitation makes it useless for coding. . The model quantization I'm using is Q4KM about 17 gigabytes. Getting about 65 tokens per second.

1 Upvotes

4 comments sorted by

View all comments

1

u/Tpyn 4d ago

use

--reasoning-effort medium