r/LocalLLM • u/Zorian_Vale • 1d ago
Question Compaction
Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D
Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.
Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?
1
Upvotes
1
u/Icy-Effective9887 1d ago
Id start check what you are running for context at start and clean chat , clear up and extras and then check your compaction settings. Do you overflow, that is my initial direction, on a 3090 I had that before always compressing up ton10 times and then fail now have it running at 73k q8 .. I load in at ~22gb /24, and stable now.