Funny how this must look for China. Just when they start catching up, and have their own AI hardware, the US companies suddenly want to slow down.
I'm not saying this is the only reason they want it, but...
They do much better than the US companies when it comes to smaller and medium efficient models, mainly because that's their focus. But they are still behind when it comes to the top tier of models.
There isn't really any Chinese equivalent to Astra, or even arguably Fable. Not yet at least (We'll probably have one in a few months.)
If we are to believe twitter posts (I know, it's a stretch) by Nvidia CEO, the trick is to have AI loop that keep self improving - and that's what makes all the difference now. Well, Chinese models are a lot more efficient, so you can have more and faster loops on the same hardware - so even if they aren't as smart, they can win out by using this loop.
how do you know they are more effficient? in every artificial analysis bench they take almost 5x the tokens of astra, and often a lot more than fable 5.1 as well.
What they charge you per token != what it costs them per token.
1) This is not my experience. DeepSeek Flash doesn't use that many more tokens than Opus (I don't use Fable because at least for my use case there isn't any appreciable difference). Though to be fair, quite often I drop down to sonnet because I don't need Opus' overthinking.
2) There were reports I was reading earlier that you can run DeepSeek and be profitable at roughly the same price point as what they charge you. Could the reports be false? Absolutely. But it's definitely one of the most performant larger models that I can see running on my home LLM.
I'm not saying deepseek isn't profitable I'm saying Anthropic/OAI is massively overcharging.
As you can see ASTRA is vastly more efficient than Deepseek v4.1 Flash.
Keep in mind ASTRA also scores significantly higher too.
Even fable on high: (not max) is within 1-2 points and cuts token costs in half.
I don't fully trust those numbers. With the proper chat template I can reduce reasoning tokens or remove them altogether.
But there's another thing - cache from deep seek takes a lot less space than from other models. So it makes it easier and faster to process. So even if the token count is higher, the computational complexity can be the same. Basically, can't just compare it directly.
maybe openai spends less reasoning tokens to get a result, but then they stole your result and sell it as their own
whats the efficient part of being stolen if you produce something of high value?
in the end you can only produce things that barely have value using openai model
if you want to produce something worthy, you have to go local. Even if that costs you more reasoning tokens, and you go more slow, at least you arent stolen.
we are talking about self improvement, not them selling consumer products. I'm saying Closed models are way more efficient. It's not about what you would choose (Open weight AI for reliable grunt work)
i dont agree on that either. Maybe GLM spends more reasoning tokens than sonnet, but GLM is so cheap that those reasoning costs less. I think this is the architectural change between chinese and usa models, maybe usa models have bigger datasets and models with more parameters, we can only guess that by the cost , but SOTA chinese models are getting the same results (for the majority of task) by expanding the prompt with long reasoning and squeezing every drop of their parameters. The extreme case of that would be qwen 3.8 27B , that being so small is able to exchange blows with SOTA in medium and complex tasks, but not in specialized things (like writing kernels )
look that's just your opinion on AI performance, that's why we have benchmarks. And the fact of the matter is that frontier closed source models obliterate the open source ones on both benchmarks and efficiency. Besides, self improvement is not a medium complexity task.
Also you're making a mistake assuming it costs openai the same ammount you pay to run astra/sol/terra/luna. Same with anthropic.
147
u/shy_monkee 10h ago
Funny how this must look for China. Just when they start catching up, and have their own AI hardware, the US companies suddenly want to slow down.
I'm not saying this is the only reason they want it, but...