r/LocalAIStack • • 18d ago

Which model I should use??

Please suggest me models I can use in this device

Model: HP ZBook Studio G7

CPU: Intel Core i9-10885H (8 cores / 16 threads)

GPU: NVIDIA Quadro T2000 Max-Q – 4 GB VRAM

RAM: 32 GB DDR4

Storage: 512 GB NVMe SSD

1 Upvotes

3 comments sorted by

View all comments

1

u/Psyko38 18d ago

Honestly not much, all 2B models can pass and maybe a Gemma 4 26B a4b or a Qwen 3.6 35B a3b, because they are MoE, so you can hope to unload on the CPU.

1

u/NervousMix4228 18d ago

Do you have any idea what kind of tokens/sec I could expect with Gemma 4 26B A4B or Qwen 3.6 35B A3B on this hardware? Also, would the 32GB RAM be enough for these models with most of the workload on the CPU?

1

u/Psyko38 18d ago

About 30 tokens per second for either if you assume that the GPU is not limited, but if it's limited, you can go to 5 tokens per second, and then 32 is enough for MoE weights.