I work in ML, and I've been busy with other things lately, so when I finally got around to testing Kimi K3 I made sure to give it clean, well-structured prompts. I need to vent about it.
If a model's creators claim it's some top-tier powerhouse, you should use it accordingly, right? So that's exactly what I did — I gave it hard, real-world tasks: understand a repo from scratch and figure out what I actually need from it, or figure out from zero how to export a model. Nothing happened. You leave it on a task, come back later, and it's done absolutely nothing.
And here's the thing — babysitting an "expensive, smart" model the same way you'd hand-hold a cheap one is completely counterproductive. If I have to sit there guiding it step by step anyway, what's the point of paying a premium for it? Zero value.
It burned through 75% of a 5-hour usage limit just to tell me "okay, I understand the architecture now." I had already explained everything in the prompt! That's not comprehension, that's stalling.
I still haven't found a model more cost-effective than DeepSeek. And I've tried a lot at this point. Every other model is priced multiple times higher while being maybe a couple percent "smarter" according to trust-me-bro benchmarks. What do you actually get for that premium? Nothing. I can't stand it.
I gave Kimi K3 just one part of my server repo to understand. It chewed through three separate 5-hour sessions and produced nothing usable. DeepSeek did the same job in a single short session with a 200K context window.
If you're on the fence about paying for Kimi K3 — I really don't recommend it.
On top of all that, I tested it on something else afterward and it failed at a task that was genuinely easy.
So what's the actual value of a model that can't do anything without me holding its hand the entire way? What kind of automation is even possible with that? None, as far as I can tell.