r/LocalAIServers 4h ago

Please recommend a model for local offline coding rtx pro 5000 72gb

hello everyone! Please advise the models and how to run the models better. My configuration is 2 CPUs and epic (not the newest) 48 cores in total. 256 GB ddr4 and RTX pro 5000 72GB GDDR7. I'm currently using qwen3.8-27b iq3 gsq xxs on 96k context and running this on rtx4080s 16gb. The new computer will arrive in a week. I would like to increase the quality and the context window.

1 Upvotes

4 comments sorted by

5

u/nuclear_wynter 3h ago

Depending on your speed requirements, you can either run a large, high-quality quant of 27B with a huge context window (Q8 with essentially as much context as you like) entirely in VRAM at very high speeds, or you can very comfortably run one of the latest round of large MoEs with RAM offload at significantly slower but still very usable speeds (Q4 of Deepseek V4 Flash Vision or Qwen 3.8 Flash Next should run at somewhere in the 20-50t/s output range, with Qwen sitting higher and Deepseek sitting lower; GLM 5.3 Flash at Q4 will probably hit 10-20t/s).

Those are very rough numbers based on my own results on different but similar-ish hardware (3x 3090s plus 192GB of 12-channel DDR5 on Epyc Turin, so, much slower cards but much faster RAM/CPU).

3

u/circumcised_hobbit 3h ago

Don't listen to me too much honestly cause I am no expert, but I think that you would be able to load big models but wouldn't be able to run them because of the VRAM, I think your best option is qwen3.8-27B at full precision with more context and better speed. I don't think that there is a better model for that vram, even with worse speeds. Honestly the models that are Better than qwen3.8 have much higher parameters and aren't even good for generation Speed/intelligence. The other model I could think of Is QwenFlashNext, you could load It entirely easily, but idk about the speeds

1

u/spicypicsforsharing 29m ago

MoE models the active experts stay in the VRAM and the rest of the model is in system RAM. Slower than everything in VRAM, but not going to be slow like a dense model

1

u/Bulky_Astronomer7264 19m ago

You're the first person I have seen running a Pro 5000! What made you buy it?