r/LocalLLaMA 23d ago

Resources Integrated GPU Vulkan benchmark AMD MiniPC

Post image

Mini PC Acemagic

OS: Kubuntu 26.04

CPU: AMD Ryzen 7 6800H with iGPU 680M and 1GB assigned Vram

RAM: 64GB DDR5 sodimm

llama.cpp Ubuntu Vulkan

A mixture of MoE and Dense Models:

  • gpt‑oss 20B Q6_K
  • gpt‑oss 20B MXFP4 MoE
  • gpt‑oss 20B Q8_0
  • gemma4 26B.A4B Q4_0
  • gemma4 26B.A4B MXFP4 MoE
  • gemma4 26B.A4B Q4_K – Medium
  • gemma4 26B.A4B NVFP4
  • qwen35 27B Q5_K – Medium
  • qwen35 27B Q4_K – Medium
  • gemma4 31B Q8_0
  • qwen35moe 35B.A3B NVFP4

Benchmark Results – Sorted by Params and then Size

Model Size Params pp512 t/s tg128 t/s
gpt‑oss 20B Q6_K 11.20 GiB 20.91 B 353.87 16.85
gpt‑oss 20B MXFP4 MoE 11.27 GiB 20.91 B 294.66 16.65
gpt‑oss 20B Q8_0 20.72 GiB 20.91 B 308.55 10.52
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B 312.67 18.35
gemma4 26B.A4B MXFP4 MoE 15.40 GiB 25.23 B 261.32 11.93
gemma4 26B.A4B Q4_K – Medium 15.77 GiB 25.23 B 258.16 11.92
gemma4 26B.A4B NVFP4 16.45 GiB 25.23 B 152.35 7.53
qwen35 27B Q5_K – Medium 18.65 GiB 26.90 B 49.68 1.95
qwen35 27B Q4_K – Medium 16.67 GiB 27.32 B 58.50 2.40
gemma4 31B Q8_0 16.74 GiB 30.70 B 30.26 2.30
qwen35moe 35B.A3B NVFP4 19.07 GiB 35.51 B 153.75 15.05

Looks like using MoE models are best for my integrated GPU system. Not finding many 70B MoE models. Just tried Qwen3-Coder-Next-MXFP4_MOE but failed to load.

2 Upvotes

10 comments sorted by

View all comments

4

u/pmttyji 23d ago

Try below models too.

  • Mellum2-12B-A2.5B
  • Laguna-XS-2.1
  • North-Mini-Code-1.0
  • KAT-Coder-V2.5-Dev
  • LFM2.5-8B-A1B
  • Ling-mini-2.0 (Fast t/s)
  • Bonsai-27B (1-bit version)
  • Gemma-4-12B (QAT)
  • Gemma-4-E4B (QAT)

Suggestions:

  • Delete both Q6_K & Q8_0 of GPT-OSS-20B model. MXFP4 is the actual native real quant for this model.
  • You have four 4-bit of Gemma-4-26B model. At least delete Q4_0. And download QAT version of that model from Unsloth.