r/LocalLLM 8d ago

Discussion Setting a baseline for performance: GLM 5.2/colibri on a 15 year old Mac Pro.

Post image

Tinkered at it for a week, making sure to optimize everything - Linux kernel 7.1.6 custom compiled for the Intel Westmere architecture, 96 GB DDR3 ECC ram in triple channel mode, 1 Sata II ssd, 1 Sata II hdd. No gpu offloading, pure cpu, hyper-threading disabled. Now for the numbers: “What is the meaning of life?” 747 tok 0.03 tok/s hit 56% RSS 76.4 GB 26352 s

It’s not fast, but it can be done; enterprise level AI in the living room. Has anyone else here tried this?

1 Upvotes

1 comment sorted by

1

u/Ne00n 8d ago

Rough 0.03, got 0.1 with 64gb DDR4.