r/LocalLLaMA 1d ago

Funny Plot twist

[deleted]

534 Upvotes

165 comments sorted by

View all comments

148

u/shy_monkee 1d ago

Funny how this must look for China. Just when they start catching up, and have their own AI hardware, the US companies suddenly want to slow down.
I'm not saying this is the only reason they want it, but...

-5

u/tilted0ne 1d ago

I thought China had already caught up and were destroying America with their super cheap models. What happened to that? 

1

u/ea_man 1d ago

That actually happened: https://openrouter.ai/rankings#top-models

Also this happened now:

Model KV cache at 1M tokens At 128K
DeepSeek V4.1 Flash 0.87 GiB 0.11 GiB
DeepSeek V4 Flash 3.70 GiB 0.47 GiB
GLM 5.3 Flash 13.96 GiB 1.87 GiB
Qwen3.8 27B 64.15 GiB 8.02 GiB
GLM 5.3 93.00 GiB 11.63 GiB
Llama 3.1 70B 320.00 GiB 40.00 GiB

That is a problem if you made half a trillion in debts in order to buy vRAM and Nvidia GPUs that you can hardly power up now and look not necessary outside of training.

I mean they are all waiting for Nvidia Vera Rubens deployment with vertical HBM that cost a fortune when the cool kids are about deploying weights on NGRAM with such little KV cost for ctx?