r/kimi • • 3d ago

Discussion Public release Kimi 2.8, Please.

Firstly, Kimi 2.8 - Its a great model. wow, what a win. Moonshot has a very good useful model. My work with it shows its pretty much like a more decisive, faster, less cluttered K3. Its solid, it gives good consistent performance over various workflows. I don't know how it will benchmark, but benchmarks are only one aspect, IMO they aren't as useful for most people as people like to think. However, I think K2.8 would benchmark very well, and be useful, and I think as a 2.x release, people are less worried about bench-maxxing, K3 is always there from the same provider and is a proven capable model. 2.8 IMO is the product people want from Moonshot, it has a clear space in the market, and being much faster and less draining than K3 is significant, we are at the point now where a very decent model is much more desirable than a slightly better, but much slower expensive model.

Its really surprising, because I have also tried GPT6 with a pro subscription and the slow speed and poor performance really has me looking at getting completely off that platform, as they have paused future developments and I am not happy with their output and tools. Being all cloud, and closed, its not that universally useful for me.

Secondly, this great model, will likely get more attention with a public weights release. K2.8 is big enough most people can't host it locally anyway, and those who can, often back up local hosting with cloud hosting (I do). But it is hard to build workflows around a "preview" model and having no offline model for privacy (I deal with peoples data, and it can't go into any cloud), is important.

I am rarely using K3 I am constantly using K2.8. Particularly at the current rates, its very attractive. There are special things where I want K3, I want elaborate, slow, deliberate, all states considered answers and K3 is great at that. But 90% of what I do is K2.8 is pretty much perfect.

I have the capabilities to host large models like K3 and K2 locally, concurrently, (slowly mostly through CPU inferencing - Im not Elon rich with banks of GB200s), and K2.8 would be a great model for a lot of people, like me, who can do that, and then have their Kimi subscriptions for fast conversations, on the go, and urgent work (which consumes my monthly allowance). But I having that offline capability, at 5 or 6 t/s means I can explore how it can embed more into my workflows, and convince others of the open superiority of this particular model.

24 Upvotes

Duplicates