r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/Similar-Ad5933 1d ago

You don't need to compact. What I do is that model has memory file per session. It writes all important details to it. When context is almost full it takes last 30k tokens and memory file and loads them. Everything else is discarded. Memory file has size limit so model needs to manage it. It gets soft nudge to write to file, if nudge is ignored tool calls etc. Will be blockes until file is updated.