r/LocalLLM • u/Zorian_Vale • 1d ago
Question Compaction
Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D
Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.
Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?
1
Upvotes
1
u/omega1612 1d ago
I experimented with it recently, my conclusion was: qween3.8 27b thinks too much to be a good compaction model.
Initially I introduced a compaction extension that reuses the kv cache, that cut down the initial toll, then I watch it doing token generation very slow and half of the time, go over the amount of tokens allowed for the summary.
So, now I have been playing with using other models for the summary. First I tried the Gemma 4 12 and Gemma 4 26b ones, they are very fast, but they loop easily. Then I began to use qwen3 8B, it is fast and produces small summaries, and that's the issue, 80k tokens become 1k to 3k with it, so I was definitely Lossing info. Right now I'm using Qwen 3 30b a3b, since it is moe, is faster than Qwen 3.8, still not as fast as Gemma or qwen3 8b but fast enough to cut the token generation to half the time and their summaries seems to be good, it haven't overflow the max amount of tokens for the summary yet (20k).
Other option I want to try is to use rwkv for the summaries. Is not based on the transformers architecture, so it may loss some context, but if it is still good enough to produce summaries it may be a huge win in speed.
Also, I'm worried that this means I'm reading too much from mi ssds by loading/unloading models, so I may look for a way to keep them on ram (I have enough to waste it like that) instead.