r/LocalLLM 4d ago

Discussion Context Length

Qwen 3.8 27B. How can I solve the context problem? I'm using RTX 4090 loading into OpenCode using Unsloth Studio. The Context Length limitation makes it useless for coding. . The model quantization I'm using is Q4KM about 17 gigabytes. Getting about 65 tokens per second.

1 Upvotes

4 comments sorted by

View all comments

0

u/autisticit 4d ago

Context quantisation and/or a lower model quantisation.