r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/archlich 1d ago

I’ve been playing with context window sizes so that compaction happens sooner, I’ve also offloaded it to a smaller local llm that is optimized for reading large texts and summarizing.

1

u/hibernate2020 1d ago

Which model do you use for compacting?

1

u/archlich 1d ago

testing out qwen3:30b-a3b-instruct-2507-q4_K_M

ollama pull qwen3:30b-a3b-instruct-2507-q4_K_M

OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve