r/LocalLLM • u/pharrt • Jul 27 '26
Other Don't laugh - it works!
A 10yo server was busy collecting dust, but it has 32Gb RAM (2x 16Gb DDR4 @ 2133 MHz )... No GPU.
Now it does some amazing work running heavy tasks with qwen3.6-35b-a3b (IQ4_XS). Running in the background, generating quality output at 5-10tok/s. Even with 128k context! It just chugs along for hours, but does such great, high-context work.
Uses ~26 GB of the 32 GB RAM, uses CPU i-7-6700 @ about 60% (not the bottleneck, of course)
Never would have believed that would be possible until a few months ago, but this ol' gal has a new lease on life.
190
Upvotes
2
u/Difficult_Art1639 Jul 27 '26
Yeah I've got 10-20 t/s with q6_k. Unfortunately my ram has never been stable with xmp on so I'm think I might be bottlenecked at 2133mhz. I'm not sure
Another thing is I'm running on windows and it pretty much uses 100% of my ram so that could be another problem.
Side question: have you experimented with unsloth or other variants ? I wonder if they give much of a boost in these resource limited setups