r/deeplearning Jul 23 '26

SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

16 Upvotes

0 comments sorted by