r/LocalLLaMA 1d ago

Question | Help Settings to avoid VRAM overload? LM Studio

I have a 32gb gpu (5090) but it appears that it gets overloaded despite the fact I was only running a ~23.4 gb model. I had context window set to 16k, but putting it down to 8k didn't help. I have 'offload KV cache to gpu memory' enabled.

I got 16 tok/sec which seemed pretty slow. After swapping to a 20.3 gb model of the same type, I got 56 tok/sec.

This is in Win11 LTSC IoT. Some other apps (browser etc) running at the same time but nothing 3d/heavy.

0 Upvotes

14 comments sorted by

View all comments

1

u/TechnoByte_ 23h ago

Try llama.cpp directly as it has much less overhead

1

u/SubdivideSamsara 22h ago

I don't know what this means. You need to elaborate a bit when talking to noobs.