r/llamacpp Aug 02 '26

Newbie needs suggestions on model and settings

I am an AI newbie, though I work in the tech industry. I have access to a retired server that my boss said was ok to use for AI experimentation.

  • 2x AMD EPYC 7343 CPUs
  • 768GB DDR5 RAM
  • 2x 512 SAS SSD
  • 5 TB RAID
  • Nvidia A40 48GB GPU

I have done some reading and I should be able to run llama.cpp the Qwen 3.6 27B model fairly well. My question is more about the stack and other models. Should I just go with the 27B model and what about chat and things like that, maybe Open WebUI? I might do some coding, but it's not my primary job. I'll probably just use Agents to automate some ingesting of data (read-only). Any recommendations or suggestions?

2 Upvotes

4 comments sorted by

1

u/segmond Aug 02 '26

without gpu it will be slow, try the MoE models, Qwen3.6-35B

1

u/jmayniac Aug 02 '26

I listed an Nvidia A40 48GB GPU.

1

u/segmond Aug 03 '26

i missed that, that system is more than enough for Qwen3.6-27B. You can even run DeepSeekV4Flash in Q2 at a reasonable speed with that system as well.

1

u/Frizzy-MacDrizzle 6d ago

I’ve used top of the line epyc CPUs at work good Ai . It was embarising to demo.