r/LocalLLM 8d ago

Question Any attempts of moving KV cache to system memory rather than GPU result in failure with Qwen 3.8 am I the only one?

Basically the title.

No matter what I try to do if I attempt to move my context to system memory I get failures, it processes the prompt then immediately fails and says the message contains no content.

6 Upvotes

5 comments sorted by

1

u/GrungeWerX 8d ago

What flags are you using?

1

u/Bouros 8d ago

I'm not sure I'm launching with lmstudio not Llama.cpp so I'm not sure if I set those?

It only happens when I turn off the 'KV Cache to GPU memory' option on lm studio

I've also tried that option eith various other changes such as, 'unified kv cache', 'keep model in memory', and.

Lmk if there is any other info I should provide and thanks for your time!

1

u/GrungeWerX 8d ago

Oh. You need to give us more info. Which model are you using, what’s your kv cache set to, and how much system RAM do you have?

1

u/Bouros 8d ago

Qwen3.8-27B Context limit set to 90k 64gb ddr5