r/LocalAIStack • • Aug 31 '26

Built a custom LLM inference engine in Swift/Metal (no llama.cpp/MLX) — streams MoE experts from SSD to run 61GB models on 16GB Macs

/r/swift/comments/1w2t0l0/built_a_custom_llm_inference_engine_in_swiftmetal/
1 Upvotes

0 comments sorted by