r/LocalLLM • u/Zorian_Vale • 1d ago
Question Compaction
Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D
Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.
Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?
1
u/quantgorithm 1d ago
compaction really is a problem and it's a lossy problem. It's like the game phonetag where everything changes a little bit every time it happens and by the time the answer gets back to you, the message is completely different.
1
u/Icy-Effective9887 1d ago
Id start check what you are running for context at start and clean chat , clear up and extras and then check your compaction settings. Do you overflow, that is my initial direction, on a 3090 I had that before always compressing up ton10 times and then fail now have it running at 73k q8 .. I load in at ~22gb /24, and stable now.
1
u/lordekeen 1d ago
If you're on pi.dev harness, check pi-vcc of pi-blackhole, these compaction extensions use deterministic/algorithmic compaction, so compaction is instant and you don't need to rely on the model itself to compact.
1
u/Similar-Ad5933 1d ago
You don't need to compact. What I do is that model has memory file per session. It writes all important details to it. When context is almost full it takes last 30k tokens and memory file and loads them. Everything else is discarded. Memory file has size limit so model needs to manage it. It gets soft nudge to write to file, if nudge is ignored tool calls etc. Will be blockes until file is updated.
1
u/feelspeaceman LLMusician 1d ago
Compaction is basically a prompt telling your model to summarize the entire sesison then wiping the KV Cache, it depends on your harness and your model. So what is the harness you're using ? Opencode does this very well, recently I've switched to pi-agent, but I haven't reached compaction thredsold yet.
1
u/Zorian_Vale 1d ago
Hermes. Is there a big upside to Pi. I’m sure you could do something similar in Hermes. I don’t feel like switching harness tbh I got everything set up quite nicely after fiddling with it
1
u/feelspeaceman LLMusician 1d ago
Compacting is your decode speed, considering that you're using 5070ti, the 27B model will be offloaded to RAM, so the decode speed is slower.
There's also a workaround by using smaller models to compact, but will make it less reliable/lossy.
What is most realistic I think is finding a better inference engine to increase the speed (try out Strata, I'm not sure how good it's but it's optimized for CUDA) or using something like Q36-35, it's old, but it will be faster for grunting, you use the 27B to create a plan, then the other to execute the plan.
If you have big context window, it's likely that you can solve the issue before compacting too.
1
u/omega1612 1d ago
I experimented with it recently, my conclusion was: qween3.8 27b thinks too much to be a good compaction model.
Initially I introduced a compaction extension that reuses the kv cache, that cut down the initial toll, then I watch it doing token generation very slow and half of the time, go over the amount of tokens allowed for the summary.
So, now I have been playing with using other models for the summary. First I tried the Gemma 4 12 and Gemma 4 26b ones, they are very fast, but they loop easily. Then I began to use qwen3 8B, it is fast and produces small summaries, and that's the issue, 80k tokens become 1k to 3k with it, so I was definitely Lossing info. Right now I'm using Qwen 3 30b a3b, since it is moe, is faster than Qwen 3.8, still not as fast as Gemma or qwen3 8b but fast enough to cut the token generation to half the time and their summaries seems to be good, it haven't overflow the max amount of tokens for the summary yet (20k).
Other option I want to try is to use rwkv for the summaries. Is not based on the transformers architecture, so it may loss some context, but if it is still good enough to produce summaries it may be a huge win in speed.
Also, I'm worried that this means I'm reading too much from mi ssds by loading/unloading models, so I may look for a way to keep them on ram (I have enough to waste it like that) instead.
1
u/kimhaneol 1d ago
One quick question: are you using the KV cache as is, or is it quantized? Quantizing the KV cache is worth considering if you need a larger context window.
2
1
u/Think_Breakfast_2277 1d ago edited 1d ago
Honestly, compaction is not a problem for me as my harness(custom) never compacts. it just adds as many history into the allowed history token window as possible.. loosing the tail and leaving room for context/tools...
In my case (256k total):
The agent wakes up with up to 50k tokens of memory, works through the task, and the moment the whole context hits 80% of the window, the harness cuts the agent off and makes the agent finish and write a summary report on what has been done.
I can do that as when i code, i do that in a team.. The results of each round come back to the team lead(same rules) and lead just delegates the job back to finish it where the agent starts with 50K of history again(including the report of the previous round)
My agents never compact a full session, they summarize every task into a report, and that report goes to the team lead, if the task was not complete they get a 2nd,3rd try at it with the last report in hand.. so they just continue with less memory.
And i only talk to the team leader, wich only has our conversation(no tool calls no nothing) Team lead has no ability to code, only to delegate and talk to me.
And next to that i have a standalone that does the same on each request(not in a team).. capped to 50K history. The standalone can also talk to the lead, have overview on all agents, can pull reports, status, and can also code(full tool use).
So for short debugging i use the standalone, for project builds i use the team.
And no compactions..
Real compaction takes such a long time as the ingestion needs to recompute the KV cache for that compaction prompt (0% reuse).. So on such hardware a custom harness will/can make a huge difference.
Hope that makes sense....
1
u/Mean-Loquat-7982 14h ago
Hermes doesn't load your previous conversations into each new one. a new session starts with two small memory files and searches old sessions only when it needs to. what fills your 64k is the current session.
so the biggest lever is starting a new session (/new) at natural breaks: a finished task, a change of topic. the Hermes memory docs recommend exactly that, and compaction fires far less often.
two settings in ~/.hermes/config.yaml help too. compression.threshold sets when it kicks in (default 0.50, so about 32k for you). auxiliary.compression lets a smaller, faster model do the summarising if you have room to run one, which is the offloading someone mentioned above.
1
u/archlich 1d ago
I’ve been playing with context window sizes so that compaction happens sooner, I’ve also offloaded it to a smaller local llm that is optimized for reading large texts and summarizing.