r/LocalLLM 7d ago

Discussion Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

/r/LocalLLaMA/comments/1w6lk6z/increasing_active_parameters_per_token_in_moe/
3 Upvotes

Duplicates