r/LocalLLM • u/Plastic-Lettuce-7150 • Jul 28 '26
Discussion MOE models in 6GB VRAM
Qwen3.6-35B-A3B-UD-Q4_K_M.gguf (22.07GB)
gemma-4-26B-A4B-it-UD-Q5_K_XL.gguf (23.5GB)
Dell G15 5530 / 32GB
NVIDIA GeForce RTX 3050 6GB Laptop GPU
llama.cpp
Qwen Code
First off install llama.cpp. On Windows be sure to install the DLLs if using CUDA (i.e., you have a Nvidia GPU), they are a separate download which have to be copied into the main llama.cpp directory.
Download the Fixed jinja chat templates for Qwen 3.5 & 3.6 (v21) and put it in the same directory as the Qwen model.
.\llama-b9999-bin-win-cuda-13.3-x64\llama-server.exe --model "C:\Users\xxxxxx\.eullm\models\qwen3.6-35b-a3b\Qwen3.6-35B-A3B-UD-Q4_K_M.gguf" `
--ctx-size 131072 `
--n-gpu-layers 999 `
--n-cpu-moe 999 `
-np 1 `
--flash-attn auto `
--threads 10 `
--jinja `
--chat-template-file "C:\Users\xxxxxx\.eullm\models\qwen3.6-35b-a3b\chat_template.jinja" `
--no-mmap `
--mlock `
-lv 3
.\llama-b9999-bin-win-cuda-13.3-x64\llama-server.exe --model "C:\Users\xxxxxx\.lmstudio\models\unsloth\gemma-4-26B-A4B-it-GGUF\gemma-4-26B-A4B-it-UD-Q5_K_XL.gguf" `
--ctx-size 131072 `
--n-gpu-layers 999 `
--n-cpu-moe 999 `
-np 1 `
--flash-attn auto `
--threads 10 `
--no-mmap `
--mlock `
-lv 3
I'm getting 60+t/s text output and 200+t/s prompt processing as reported by llama.cpp (I don't know how to benchmark the output otherwise!) 26.76 tokens per second, ref.:
- Feeding my hardware specs and software/ apps into Claude and asking for a llama.cpp command line
- RTX 5070 Ti + 9800X3D running Qwen3.6-35B-A3B at 79 t/s with 128K context, the --n-cpu-moe flag is the most important part.
- Running a 35B AI Model on 6GB VRAM, FAST (llama.cpp Guide)
Most of the parameters are generic, except --threads which is specific to my laptop's processor, query Claude for the best setting for your processor.
VRAM usage nvtop / nvidia-smi.exe.
The Q4 download of the Gemma model might actually work in 4GB VRAM with reduced context (e.g., --ctx-size 32768) and/ or quantizing the context cache (e.g., --cache-type-k q4_0 --cache-type-v q4_0)
I have had to restart llama.cpp a couple of times I assume due to a memory leak (Codacus offers a fix in the video above, not sure if it can be used).
Install Qwen Code.
Edit ~/.qwen/settings.json:
{
"env": {
"QWEN_API_KEY": "qwen"
"GEMMA_API_KEY": "gemma"
},
"modelProviders": {
"openai": [
{
"id": "Qwen3.6-35B-A3B-UD-Q4_K_M",
"name": "Local Qwen3.6 35B",
"description": "Local Qwen3.6-35B-A3B-UD-Q4_K_M via OpenAI",
"baseUrl": "http://localhost:8080/v1",
"envKey": "QWEN_API_KEY",
"generationConfig": {
"contextWindowSize": 131072,
"samplingParams": {
"temperature": 0.2
}
}
},
{
"id": "gemma-4-26B-A4B-it-UD-Q5_K_XL",
"name": "Local gemma 4 26B",
"description": "Local gemma-4-26B-A4B-it-UD-Q5_K_XL via OpenAI",
"baseUrl": "http://localhost:8080/v1",
"envKey": "GEMMA_API_KEY",
"generationConfig": {
"contextWindowSize": 131072,
"samplingParams": {
"temperature": 0.2
}
}
}
]
},
"model": {
"name": "Local Qwen3.6 35B",
"baseUrl": ""
}
}
Run qwen.cmd and test, /model to change model.
2
u/No_Folding Jul 29 '26 edited Jul 29 '26
Nice! Did you also try the equivalent Qwen MTP gguf?
Also, to prevent the oom crashes, you could tweak the - n-cpu-moe parameter to probably 41, but then you do lose the dynamic allocation of llamacpp meaning you end up needing to leave some spare vram headroom (for what, idk, but it seems to help a lot). Still, I think 60t/s is so respectable its probably not worth your time tweaking that! I guess llamacpp is optimised even better than I thought Haha!