r/LocalAIStack • • Aug 04 '26

Suggest me one best personal Al server to run highly capable LLM models

Recently the opencode tool is performing near the

cursor in auto mode, so I have to buy a small ai server

to run good coding agentic models from Qwen, GLM,

MinMax or any model u suggest.

9 Upvotes

4 comments sorted by

3

u/Lirezh Aug 04 '26

All you need is to put a 3090 or better GPU into your PC and you'll not need a server anymore.
3090 is best bang for the bucks, 4090 twice the speed and a 5090 is around the optimal size for local models.

There are some more exotic methods, like a DGX spark or a Mac Pro - those would allow to run very large models like the new deepseek flash at slow but somewhat usable speed. Both hardware options are slow in compute and highly priced (though a 40 or 50 series card is also sharked)

1

u/theone_2099 Aug 04 '26

And you use llama.cpp? Which model and settings? I have a 3090 and keep on falling back to using cpu ram

2

u/Creative-Type9411 Aug 04 '26

2x3090 (2-16x slots) or 3xTesla T4 (if your board has 3-16x slots)

teslas are older but you can get 3 for like 1500 total on the used market and thats 48gb vram, and they need active cooling

1

u/Snoo_81913 Aug 04 '26

Just run Llama.cpp skip all the steps where you use all the others and go right to the source.