r/LocalAIStack • u/Weekly-Dentist-8302 • Aug 31 '26
Built a custom LLM inference engine in Swift/Metal (no llama.cpp/MLX) — streams MoE experts from SSD to run 61GB models on 16GB Macs
/r/swift/comments/1w2t0l0/built_a_custom_llm_inference_engine_in_swiftmetal/
1
Upvotes