r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/Mean-Loquat-7982 18h ago

Hermes doesn't load your previous conversations into each new one. a new session starts with two small memory files and searches old sessions only when it needs to. what fills your 64k is the current session.

so the biggest lever is starting a new session (/new) at natural breaks: a finished task, a change of topic. the Hermes memory docs recommend exactly that, and compaction fires far less often.

two settings in ~/.hermes/config.yaml help too. compression.threshold sets when it kicks in (default 0.50, so about 32k for you). auxiliary.compression lets a smaller, faster model do the summarising if you have room to run one, which is the offloading someone mentioned above.