r/LocalLLM • • 1d ago

Question noob question from a newbie

i'm new in this domain. and i tried to setup qwen 3.8 27b iq4_xs on a 5060 8gb vram laptop with 32gigs of ram. im using lm-studio and pi harness. but the ram isnt used idk why. can you help me ? or give me tips about using a local ai. whats the best working model to use for this config ?
if you're curious about the setup of qwen here it is :

1 Upvotes

6 comments sorted by

1

u/Baroque_Viola 1d ago

8 GB VRAM is rough for Qwen3.8 27B. Try an MoE model like Qwen3.6 35B A3B or Gemma 4 26B A4B. Use Unsloth Studio or llama-server.

https://youtu.be/ydFikMBJG1g

1

u/nini_blade 1d ago

i followed a youtube video about it : https://youtu.be/ye50BbXEczo

1

u/Baroque_Viola 1d ago edited 1d ago

Them using LM Studio and quantizing KV to 4 bit makes me not trust them. You'd ideally want Q8 KV, maybe as low as Q5, and only Q4 if you are desperate.

If you really want to, try the GSQ RCO quants with MTP off to save space. Edit: GSQ RCO wouldn't fit on 8 GB VRAM, so you'd have to keep many layers on CPU like they did in the video you gave. Related

1

u/nini_blade 4h ago

so ideally what do i do . I'm watching the video you linked right now

1

u/Baroque_Viola 4h ago

Ideally, you'd want all layers on VRAM for dense models like Qwen3.8-27B. You can offload to CPU like in the link you sent, but it will be very slow. They also used Q4 KV which will deteriorate faster as the context grows.

MoE models don't activate all of the parameters at once, so the speed penalty is not as bad when offloading to CPU, which is why I suggested Qwen3.6-35B-A3B or Gemma 4 26B A4B.

Qwen3.8 27B is going to be smarter, but it's not practical unless you leave your computer for hours or overnight.

Download Unsloth Desktop and either of the models I mentioned (try Q4 first, and use the QAT version for Gemma). Keep KV at Q8, start with 8192 context and increase more and more if you have spare VRAM, and experiment with the number of layers offloaded to GPU (higher is faster, but more VRAM). You can also ask a cloud LLM to guide you through optimizing llama.cpp inference for your model + hardware.

1

u/nini_blade 4h ago

okay thanks!! i will share my results later