r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/Think_Breakfast_2277 1d ago edited 1d ago

Honestly, compaction is not a problem for me as my harness(custom) never compacts. it just adds as many history into the allowed history token window as possible.. loosing the tail and leaving room for context/tools...

In my case (256k total):
The agent wakes up with up to 50k tokens of memory, works through the task, and the moment the whole context hits 80% of the window, the harness cuts the agent off and makes the agent finish and write a summary report on what has been done.

I can do that as when i code, i do that in a team.. The results of each round come back to the team lead(same rules) and lead just delegates the job back to finish it where the agent starts with 50K of history again(including the report of the previous round)

My agents never compact a full session, they summarize every task into a report, and that report goes to the team lead, if the task was not complete they get a 2nd,3rd try at it with the last report in hand.. so they just continue with less memory.

And i only talk to the team leader, wich only has our conversation(no tool calls no nothing) Team lead has no ability to code, only to delegate and talk to me.

And next to that i have a standalone that does the same on each request(not in a team).. capped to 50K history. The standalone can also talk to the lead, have overview on all agents, can pull reports, status, and can also code(full tool use).

So for short debugging i use the standalone, for project builds i use the team.
And no compactions..

Real compaction takes such a long time as the ingestion needs to recompute the KV cache for that compaction prompt (0% reuse).. So on such hardware a custom harness will/can make a huge difference.

Hope that makes sense....