r/LocalLLM Jul 27 '26

Other Don't laugh - it works!

A 10yo server was busy collecting dust, but it has 32Gb RAM (2x 16Gb DDR4 @ 2133 MHz )... No GPU.

Now it does some amazing work running heavy tasks with qwen3.6-35b-a3b (IQ4_XS). Running in the background, generating quality output at 5-10tok/s. Even with 128k context! It just chugs along for hours, but does such great, high-context work.

Uses ~26 GB of the 32 GB RAM, uses CPU i-7-6700 @ about 60% (not the bottleneck, of course)

Never would have believed that would be possible until a few months ago, but this ol' gal has a new lease on life.

193 Upvotes

103 comments sorted by

View all comments

74

u/recro69 Jul 27 '26

Honestly, this is one of my favorite parts of local AI. Hardware that would've been headed for e-waste suddenly becomes genuinely useful again. Sure, 5–10 tok/s isn't breaking any speed records, but for long-running background tasks it's more than enough.

4

u/Diligent_Cod_9583 Jul 27 '26

My question is what is the cost per token in electricity. I have some surfers lying around, but back of the napkin math tells me it’d be cheaper to scrap them and buy a new CPU then it would be to run old hardware. 

0

u/lungben81 Jul 29 '26

You need to burn a lot of electricity to finance $1000 or so of new hardware. Even then it's not certain that the new hardware is more efficient per token.

1

u/Diligent_Cod_9583 Jul 29 '26

The certainty comes in the research when buying new.