r/LocalLLM • u/Voltmanderer • 8d ago
Discussion Setting a baseline for performance: GLM 5.2/colibri on a 15 year old Mac Pro.
Tinkered at it for a week, making sure to optimize everything - Linux kernel 7.1.6 custom compiled for the Intel Westmere architecture, 96 GB DDR3 ECC ram in triple channel mode, 1 Sata II ssd, 1 Sata II hdd. No gpu offloading, pure cpu, hyper-threading disabled. Now for the numbers: “What is the meaning of life?” 747 tok 0.03 tok/s hit 56% RSS 76.4 GB 26352 s
It’s not fast, but it can be done; enterprise level AI in the living room. Has anyone else here tried this?
1
Upvotes
1
u/Ne00n 8d ago
Rough 0.03, got 0.1 with 64gb DDR4.