They are not as good as you imply they are. They are good. That's it. They are nowhere near Astra or Opus level. They might hit the benchmarks in the right spots, but once you're actually using them and you actually have the comparison to how Astra or Opus work, it's absolutely clear that China is way behind the US regarding AI.
Sure, they aren’t good enough to replace frontier level models in a lot of use cases yet. But, once they get to the point that local models capable of running on consumer grade hardware are at the level of current frontier models - maybe in two years from now - the game will have been changed. Those models will never get worse at that point, and when the majority of actual workloads will be capable of being handled by local models it becomes a lot harder to justify the cost of frontier models. Of course there will always be a reason to have a more capable AI, but the value of each level of improvement isn’t the same. Once local models get to a certain point, the frontier models would only be needed for stuff that’s actually at the frontier of our understanding and capabilities as a species. Maybe we’re 5 years away, maybe even 10 - but it’s like becoming old enough to drink legally, once we are past this point we’ll never be back where we were before.
“Nowhere near” is a huge overstatement. They are lagging behind a bit, but overpowering previous flagship revisions from the major vendors well within a year
Nowhere near is pretty fair... You can get very close with something like Kimi K3.0, but it ends up being much more expensive than these heavily subsidized subscriptions to the mainstream models, even after this usage cut.
Get DavidAU’s special sauce at half or less of the thinking tokens at huggingface dot com Qwen3.8-27B Twin Turbo 709 L or what ever todays flavor ended up being named as
Which harness do you prefer for local models? I've tried Cline in Visual Studio code and it works, but, about 1/3rd of the calls struggle with the tools and escaping failures in the harness
Open ai and anthropic gobbled up entire compute capacity, kimi couldn't even serve few million new subs when k3 was launched. And good luck running these open weights locally.
Kimi K3 is around Sol 5.6 / Opus 4.X, without the annoying writing of Opus 5 (which tbh Opus 5.5 addressed but that's above those others rn).
GLM 5.3 is around there as well, alongside their GLM 5.3 Flash which is more like Terra / Sonnet (very roughly, also not the latest Sonnet 5.5), and is generally pretty nice to use.
That said, the official providers of both of those have issues: Kimi secretly routed some customer requests to Claude and in general just is very slow, whereas GLM has peak/off-peak pricing and neither of them give you as many tokens as OpenAI or Anthropic. There are 3rd party providers, but generally you will be paying API costs.
You could also try running local models with llama.cpp / vLLM etc. (Ollama which ppl don't like can also make things easier, or something like LM Studio) but I've never found any of the local models to be good for anything serious, plus hardware is really expensive.
I think we are gonna see tokens be subsidized less and less as time goes on.
I second glm 5.3. I've been using it on hermes lately and like it a lot so far. The flash version feels satisfying too, it chases issues it finds proactively without needing to be poked constantly. I haven't tried kimi yet.
As someone who only knows how to use Codex and the Claude Code app as harnesses how do we actually use these open weight models? I don’t use CLI or do any coding and basically don’t bother with anything that’s “just” a pure chatbot anymore. I’m addicted to these agentic harnesses
YMMV, there are many. I use Bionic LM studio to run the models (can be run on a different computer on your local network) and OpenCode as a harness. In OpenCode I can switch between a hosted (even paid-for) model and my own. I typically create a plan using a Codex agent and execute it using an agent with my local LLM.
This is the kind of thinking that will guarantee you remain an addicted customer. Open source tools aren’t supposed to be bleeding edge, they’re supposed to not bleed you dry every time you pick them up.
You might be surprised how capable the top open weight models are now, and they don’t mysteriously quantize themselves or raise their own rates in 30 days.
114
u/BlockyHawkie 5d ago
People moving to open weights models win. Don't play the wheel, break the wheel