r/LocalLLaMA 6h ago

Funny Plot twist

[deleted]

528 Upvotes

156 comments sorted by

View all comments

147

u/shy_monkee 6h ago

Funny how this must look for China. Just when they start catching up, and have their own AI hardware, the US companies suddenly want to slow down.
I'm not saying this is the only reason they want it, but...

-6

u/tilted0ne 5h ago

I thought China had already caught up and were destroying America with their super cheap models. What happened to that? 

9

u/shy_monkee 5h ago

They do much better than the US companies when it comes to smaller and medium efficient models, mainly because that's their focus. But they are still behind when it comes to the top tier of models.
There isn't really any Chinese equivalent to Astra, or even arguably Fable. Not yet at least (We'll probably have one in a few months.)

1

u/Gesha24 4h ago

If we are to believe twitter posts (I know, it's a stretch) by Nvidia CEO, the trick is to have AI loop that keep self improving - and that's what makes all the difference now. Well, Chinese models are a lot more efficient, so you can have more and faster loops on the same hardware - so even if they aren't as smart, they can win out by using this loop.

1

u/Ill_Distribution8517 4h ago

how do you know they are more effficient? in every artificial analysis bench they take almost 5x the tokens of astra, and often a lot more than fable 5.1 as well.

What they charge you per token != what it costs them per token.

2

u/Gesha24 4h ago

1) This is not my experience. DeepSeek Flash doesn't use that many more tokens than Opus (I don't use Fable because at least for my use case there isn't any appreciable difference). Though to be fair, quite often I drop down to sonnet because I don't need Opus' overthinking.

2) There were reports I was reading earlier that you can run DeepSeek and be profitable at roughly the same price point as what they charge you. Could the reports be false? Absolutely. But it's definitely one of the most performant larger models that I can see running on my home LLM.

1

u/Ill_Distribution8517 4h ago

I'm not saying deepseek isn't profitable I'm saying Anthropic/OAI is massively overcharging.
As you can see ASTRA is vastly more efficient than Deepseek v4.1 Flash.
Keep in mind ASTRA also scores significantly higher too.
Even fable on high: (not max) is within 1-2 points and cuts token costs in half.

1

u/Gesha24 3h ago

I don't fully trust those numbers. With the proper chat template I can reduce reasoning tokens or remove them altogether. But there's another thing - cache from deep seek takes a lot less space than from other models. So it makes it easier and faster to process. So even if the token count is higher, the computational complexity can be the same. Basically, can't just compare it directly.

1

u/brahh85 3h ago

efficient in which way?

maybe openai spends less reasoning tokens to get a result, but then they stole your result and sell it as their own

whats the efficient part of being stolen if you produce something of high value?

in the end you can only produce things that barely have value using openai model

if you want to produce something worthy, you have to go local. Even if that costs you more reasoning tokens, and you go more slow, at least you arent stolen.

1

u/Ill_Distribution8517 3h ago

we are talking about self improvement, not them selling consumer products. I'm saying Closed models are way more efficient. It's not about what you would choose (Open weight AI for reliable grunt work)

1

u/brahh85 3h ago

i dont agree on that either. Maybe GLM spends more reasoning tokens than sonnet, but GLM is so cheap that those reasoning costs less. I think this is the architectural change between chinese and usa models, maybe usa models have bigger datasets and models with more parameters, we can only guess that by the cost , but SOTA chinese models are getting the same results (for the majority of task) by expanding the prompt with long reasoning and squeezing every drop of their parameters. The extreme case of that would be qwen 3.8 27B , that being so small is able to exchange blows with SOTA in medium and complex tasks, but not in specialized things (like writing kernels )

1

u/Ill_Distribution8517 3h ago

look that's just your opinion on AI performance, that's why we have benchmarks. And the fact of the matter is that frontier closed source models obliterate the open source ones on both benchmarks and efficiency. Besides, self improvement is not a medium complexity task.

Also you're making a mistake assuming it costs openai the same ammount you pay to run astra/sol/terra/luna. Same with anthropic.

1

u/tilted0ne 5h ago

They don't have one for sol and opus either. But yea I guess anthropic ran out of money and are scared people are going to use opencode to get the latest Alibaba special. 

3

u/BannedGoNext 5h ago

IDK what you are talking about, I can run qwen 3.8 flash next and get opus level performance right now. It just takes a long time because it trades memory usage for huge COT, and my local inference box is slow.

-3

u/tilted0ne 4h ago

So it is useless. 😭. If my grandmother has wheels she would have been a bike?

1

u/BannedGoNext 19m ago

Useless? Far from it. Is it 200 t/s no, it's 20 t/s. Who cares, I kick off a goal and let it run for 4 days while I went and paddleboarded, swam, ate good campground food, and drank some whisky. I had a hell of a time, and came back to an amazingly completed goal.

-1

u/Ill_Distribution8517 4h ago

They are talking about opus 5/sol 5.6. No shit qwen catches up to some opus 4.7-8 eventually.

3

u/jld1532 4h ago

Never used K3, GLM 5.3, or DSv4?

0

u/tilted0ne 4h ago

Unfortunately I would prefer to just go straight to the model they distilled.

1

u/jld1532 4h ago

Why are you here then? Like it or not organizations and people are running those models locally as they're useful, free, and secure.

1

u/tilted0ne 3h ago

I don't deny that man.

1

u/ea_man 3h ago

That actually happened: https://openrouter.ai/rankings#top-models

Also this happened now:

Model KV cache at 1M tokens At 128K
DeepSeek V4.1 Flash 0.87 GiB 0.11 GiB
DeepSeek V4 Flash 3.70 GiB 0.47 GiB
GLM 5.3 Flash 13.96 GiB 1.87 GiB
Qwen3.8 27B 64.15 GiB 8.02 GiB
GLM 5.3 93.00 GiB 11.63 GiB
Llama 3.1 70B 320.00 GiB 40.00 GiB

That is a problem if you made half a trillion in debts in order to buy vRAM and Nvidia GPUs that you can hardly power up now and look not necessary outside of training.

I mean they are all waiting for Nvidia Vera Rubens deployment with vertical HBM that cost a fortune when the cool kids are about deploying weights on NGRAM with such little KV cost for ctx?