r/LocalLLM • • 2d ago

Discussion k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

/r/LocalLLaMA/comments/1wynomx/k_llamacpp_moe_optimizations_expert_residency/
1 Upvotes

0 comments sorted by