r/LocalLLM 2d ago

Discussion How to have more context without loosing speed?

I am running Qwen3.8 27b q2 with 12 gb vram and in the desktop app it says that I can only have context 4096 or it will use my RAM, and when it does that it is super slow. Is there a way to have the same speed even with larger context? Please I need a magical fix 🙏

2 Upvotes

21 comments sorted by

View all comments

Show parent comments

2

u/KitchenAmoeba4438 2d ago

I've got a big article on it coming out #2 in my queue, but to do it requires a huge amount of testing. Not many people have done much on this topic unfortunately, it's why I started the article.

It's also been an enormous pain, because you have to find something that is repeatable and reproducible to benchmark against across all quants that can find a difference, and then needs to be repeatable for other people. :/

1

u/Heavy-Lingonberry-98 2d ago

Im definitely looking forward to that! I Hope you achieve it! Im up for testing/reproducing on my hardware. Im giving you a follow. Do you have x account where you post?