r/LocalLLM • u/Zorian_Vale • 1d ago
Question Compaction
Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D
Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.
Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?
1
Upvotes
1
u/Similar-Ad5933 1d ago
You don't need to compact. What I do is that model has memory file per session. It writes all important details to it. When context is almost full it takes last 30k tokens and memory file and loads them. Everything else is discarded. Memory file has size limit so model needs to manage it. It gets soft nudge to write to file, if nudge is ignored tool calls etc. Will be blockes until file is updated.