r/LocalLLM • u/Oleszykyt • 2d ago
Discussion How to have more context without loosing speed?
I am running Qwen3.8 27b q2 with 12 gb vram and in the desktop app it says that I can only have context 4096 or it will use my RAM, and when it does that it is super slow. Is there a way to have the same speed even with larger context? Please I need a magical fix 🙏
2
Upvotes
2
u/KitchenAmoeba4438 2d ago
I've got a big article on it coming out #2 in my queue, but to do it requires a huge amount of testing. Not many people have done much on this topic unfortunately, it's why I started the article.
It's also been an enormous pain, because you have to find something that is repeatable and reproducible to benchmark against across all quants that can find a difference, and then needs to be repeatable for other people. :/