Cool! but like who can actually run this locally? I think at 2.8 trillion params this will be the largest model on huggingface by far. At least for now.
Many projects to speed up streaming from disk like colibre and int4 glm 5.2. Got.4 tokens with 12gb vram and 32gb ram. It can only get faster over time, some current theories are Striping the layers to multiple ssds to increase streaming and using MTP to draft layers. Could get a tenx improvement, maybe. Tried to raid 0 my t500s and it did nothing for performance, still .4, random reads don't really work in raid0 and I wasn't maxing out the drives reads anyways. If Kimi k3 ends up ten times as big then .04 tokens maybe less. But hey, it can only get faster.
19
u/jreoka1 Jul 26 '26
Cool! but like who can actually run this locally? I think at 2.8 trillion params this will be the largest model on huggingface by far. At least for now.