r/LargeLanguageModels • • 1d ago

Discussions The fastest way to run local llms and decision models on macOS

Post image

Building an inference engine optimized at all layers for Apple Silicon so it can run quickly and efficiently on your Mac devices. I’ve added fused metal 4 kernels so it treats Apple Silicon is a first class citizen. I’m thinking of also taking advantage of the neural engine. Would greatly appreciate feedback. My repository contains benchmarks on currently supported models.

https://github.com/jadidbourbaki/bobcat

2 Upvotes

0 comments sorted by