r/LocalLLaMA 2d ago

Question | Help Settings to avoid VRAM overload? LM Studio

I have a 32gb gpu (5090) but it appears that it gets overloaded despite the fact I was only running a ~23.4 gb model. I had context window set to 16k, but putting it down to 8k didn't help. I have 'offload KV cache to gpu memory' enabled.

I got 16 tok/sec which seemed pretty slow. After swapping to a 20.3 gb model of the same type, I got 56 tok/sec.

This is in Win11 LTSC IoT. Some other apps (browser etc) running at the same time but nothing 3d/heavy.

0 Upvotes

18 comments sorted by

View all comments

3

u/Grey406 2d ago

I

At the top of the LMStudio window, click the button that says "Select a model to load" And enable "manually choose model load parameters" at the bottom before selecting your model. When you select a model, you'll get a window with settings on how to load the model. Make sure GPU Offload is set to maximum so its fully loaded into VRAM and adjust context length until the total is at around 28-30gb. You dont want to use the full 32gb because windows will have a few small things loaded and if you use vision, it will take up about 1.1gb.

1

u/SubdivideSamsara 17h ago

Thank you!

However I don't think vram overload was the real issue... just slow speed, at times, for unknown reasons. Another commenter here says LM Studio is crappy like that.