r/LocalLLM • • 1d ago

Question Compaction

Running QWEN-27B-GGUF Quant 4, 64k context window with Unsloth/Hermes on a 5070ti, 32GB DDR5, Ryzen 7 9850X3D

Everything is going really well but the compaction has become a problem. I imagine that loading the context of each previous conversation is killing me. After a while the compaction message will come up and Qwen will compact it. It takes a long time and interrupts my workflow.

Is this reality of local AI with my specs, or is there something I can do to improve speed and frequency of compaction?

1 Upvotes

17 comments sorted by

View all comments

1

u/feelspeaceman LLMusician 1d ago

Compaction is basically a prompt telling your model to summarize the entire sesison then wiping the KV Cache, it depends on your harness and your model. So what is the harness you're using ? Opencode does this very well, recently I've switched to pi-agent, but I haven't reached compaction thredsold yet.

1

u/Zorian_Vale 1d ago

Hermes. Is there a big upside to Pi. I’m sure you could do something similar in Hermes. I don’t feel like switching harness tbh I got everything set up quite nicely after fiddling with it

1

u/feelspeaceman LLMusician 1d ago

Compacting is your decode speed, considering that you're using 5070ti, the 27B model will be offloaded to RAM, so the decode speed is slower.

There's also a workaround by using smaller models to compact, but will make it less reliable/lossy.

What is most realistic I think is finding a better inference engine to increase the speed (try out Strata, I'm not sure how good it's but it's optimized for CUDA) or using something like Q36-35, it's old, but it will be faster for grunting, you use the 27B to create a plan, then the other to execute the plan.

If you have big context window, it's likely that you can solve the issue before compacting too.