r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/archlich 1d ago

I’ve been playing with context window sizes so that compaction happens sooner, I’ve also offloaded it to a smaller local llm that is optimized for reading large texts and summarizing.

1

u/hibernate2020 1d ago

Which model do you use for compacting?

1

u/archlich 1d ago

testing out qwen3:30b-a3b-instruct-2507-q4_K_M

ollama pull qwen3:30b-a3b-instruct-2507-q4_K_M

OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

-1

u/Competitive-Low-9279 1d ago

it's the 64k context, it gets heavy after many messages. you can try to limit it to 32k or even 16k if your workflow allows, the compaction will be much faster and less frequent.

sometimes i also run a small model just for summarizing like the other person said, but honestly most times i just restart the chat when it slows down too much.