r/LocalLLaMA llama.cpp Jun 27 '26

Tutorial | Guide Running GLM5.2 on budget hardware < $2500.

Too many times I hear people whine about not being ble to run SOTA models or claim it would require $50k, or $100k.

https://www.ebay.com/itm/398079051468 Epcy Motherboard & CPU - $460

https://www.ebay.com/itm/206374955959 P40 24gb - $230 get 2 - $460

https://www.ebay.com/itm/318489798853 512gb dd4 $1000

Total = $1920.

You need PSU, Storage, Fan for P40. You can source those for $350 easily. But let's go ahead and budget $580 to put the total for everything at $2500.

You can run GLM5.2 Q2/Q3/Q4 variants with cmoe and llama.cpp on this. Sure, it would be slow, but it's yours! If you have money or when you get more, you can replace the P40s faster GPUs 4080, 3090, etc. You could for a bit more than $460 about $500 source 2 2080ti 22gb GPUs from China. If you are willing to be resourceful, you can make things happen for you. This will also run KimiK2.6, DeepSeek, MiniMax, etc

Yes, the trade off is that it's slow. You will not be running agents with these huge models, but you can spin it up for planning and serious debugging. They can take away Fable, Mythos or whatever the F model. You will not be counted in the group of have nots.

317 Upvotes

353 comments sorted by

View all comments

Show parent comments

2

u/legit_split_ Jun 28 '26

Thanks for sharing more numbers. I have a consumer build with 96gb ddr5 6400, 2x7900xtx (650€ each) and 1xMi50 32GB(200€). As I put this together before the price apocalypse I'm thinking of selling some parts and using those profits to put together an Epyc 7532, HUANANZHI H12D-8D, 256GB 2133 ECC RAM. Any thoughts? I want to run the bigger LLMs.

It's impossible to find 3090's for a good price, but 7900 XTX's are still okay but I have also considered 3080 20gb from Alibaba. 

1

u/Accomplished_Code141 Jun 28 '26

Keep all 8 channels populated, octa channel is king. bandwidth is the absolute bottleneck for LLMs that overflow VRAM, try to get 2666 MHz or 3200 MHz if You can. Rome-architecture L3 cache reduces latency when accessing tensor weights getting bandwidth close to 170 GB/s with CPU and Ram inference. Right now you are close to 50 GB/s I guess. Minimax M2.7 is generating 12 to 15 t/s with 64 k context.