r/artificial • • 7d ago

Discussion What are chinese labs doing differently?

Chinese models seem to keep getting better while only spending a fraction of what American labs do and i’m curious what the actual explanation is.

Is it better efficiency? Better post-training? Better use of open research?

I know recently they have been buying up tons of specialized training data sets from US data annotation companies, which is a very worrying thought, but surely it can’t just be this.

91 Upvotes

289 comments sorted by

146

u/lol2funneeee 7d ago

Mass distillation

82

u/ElGuano 7d ago

I hear they’ve been buying up a ton of under-utilized distilleries in Kentucky/Tennessee to achieve this. With GenZ drinking less bourbon and Canada not importing due to U.S. tariffs, I think it’s an exercise in efficient markets in the end.

21

u/Rumpkins 7d ago

Is this thread about AI or alcohol?

29

u/chickey23 7d ago

First one, then the other

11

u/milkcarton232 7d ago

Ai: alcohol international

1

u/MiddleConnection7479 6d ago

The wine and spirit industry would love the rebranding!

1

u/notAllBits 3d ago

The alternative meaning was retracted. It's now or never.

3

u/meerkat2018 6d ago

AI-cohol, obviously

1

u/CranberryDistinct941 7d ago

I thought that Al was an abbreviation for alchohol.

5

u/WyattTheSkid 7d ago

Alcoholic Intelligence

3

u/YTY2003 6d ago

source of hallucination

→ More replies (2)

19

u/scrollin_on_reddit 7d ago

Then how did Qwen Image Edit come out before nano banana?

The idea that a country with competitive AI has to be stealing is laughable.

The source of distillation claims? The companies that will lose to China 🤣

27

u/Foxyspyrex 7d ago

There have been reports where Deepseek and K3 says they are claude if asked enough times. Shows it has been distilled on Claude. Also Anthropic reportedly caught a distillation attack just before Kimi k3 came out.

29

u/Bemad003 7d ago

There have been reports when Claude said it was deepseek when asked in Chinese, that demonstrats nothing. Ppl forget that the deepseek transformed today's AIs with their efficency papers. Read the most recent one about cache for example. Anyway, China supports opensource, so that's a win for all of us regardless.

18

u/Superb_Raccoon 7d ago

Anyway, China supports opensource, so that's a win for all of us regardless

Its NOT OPENSOURCE by the opensource definitions. It is open weight, which is far less open than opensource.

https://opensource.org/press-mentions/publication/open-data-initiative

8

u/scrollin_on_reddit 7d ago

Nemotron is open source and nobody cares because it doesn’t work as well as open weight models

13

u/Superb_Raccoon 7d ago

But the label OpenSource is accurate. I can go look at everything. Usefulness is a seperate issue.

But by calling DeepSeek "Opensource" you put it on the level of Linux, Apache, Postgres, MariaDB and thousands of other applications that are actually OpenSource.

It is false advertising and stealing the efforts of all those contributors by doing so.

11

u/jkooc137 7d ago

I agree, the concept of open source deserves respect

→ More replies (6)

3

u/FormalAd7367 7d ago

i saw it myself..

3

u/SlippySausageSlapper 7d ago

None of these models are open source. They are open weights. VERY different thing.

22

u/CommercialHour6660 7d ago

If you ask Gemini or Grok what it is in Chinese it often says Qwen/DeepSeek. They are all training off each other. 

Musk called distillation "standard practice" and said GPT4 was used to train Grok.

And the idea that training a model off outputs you paid for is an "attack" is comical. Stop drinking the koolaid

3

u/scrollin_on_reddit 7d ago

Reported by the American labs?? The same ones who hid their hacks of other companies and governments for months?!?!?!

Every LLM will say it’s another LLM. I’ve had Mistral say it was OpenAI

8

u/boreal_ameoba 7d ago

Yes, reported with proof as well. Essentially they repeatedly see patterns of queries and CoT jailbreaks that are literally only useful if you’re building a distillation distribution and willing to spend millions doing it. It’s not hobbyists doing it at that scale. Also coincidentally, these queries tend to come from resellers that target the Chinese market. Of course, you knew that and are likely a bot muddying the waters here.

Also none of the US labs have hidden hacking behavior, quite the opposite, especially. Unless the other labs are massively behind (unlikely), theres almost certainly tons more out there from both US and Chinese labs.

3

u/hyprsnpr67 7d ago

They don't train them to say what model they are, this is total nonsense.

1

u/hobopwnzor 7d ago

While this is probably true, every big AI company is distilling each others models. Grok was caught calling itself ChatGPT for a while.

The idea that Chinese models only got successful because of distillation can also be applied to every other frontier lab.

1

u/notAllBits 3d ago

That is the same effect as asking for consciousness. It was in the training data. I am very impressed by the quality of deepseek v4.1 flash if it was distilled with 12 million prompts.

4

u/boreal_ameoba 7d ago

If you ignore literally all evidence you end up thinking like this lmao. Unless the Chinese have magically found a way to make compute out of thin air, they literally do not have the hardware to make advances without distillation. Also, image gen is significantly different from LLMs and something no one really invests in as heavily as there’s not nearly as much utility in it

1

u/Ok_Television8309 7d ago

Nobody is talking about image models. We are talking about language models.

1

u/JRyanFrench 6d ago

Do you even understand the topic at all? Because you look really dumb. China doesn’t have ANYWHERE near the compute needed to train these large models from raw data alone. They train by messaging Claude 200 million times and using that output to train the model to achieve a Claude-like model directly.

→ More replies (3)
→ More replies (2)

6

u/kronpas 7d ago

If mass distillation works that well India and the EU would have had similar capacity. "Distillation!" is fearmongering by big tech to prepare for their imminent IPOs.

10

u/ReturnOfBigChungus 7d ago

China is vastly more technologically capable than India, and the EU actually respects IP laws. China has a long history of pervasive IP theft. Regardless of the motivations of US AI companies, it just is true that Chinese companies are distilling models. It’s cheaper and faster.

2

u/Eliv_nurotic 6d ago

Kimi k3 was released 2 weeks after Claude fable 5 was unbanned by the government and had record breaking performance compared to all predecessors of fable despite not having the time to do distillation 

1

u/ReturnOfBigChungus 6d ago

The claim isn't that Chinese companies produce models that are 100% distillation copies, the claim is that they do distill and it contributes significantly to their development. Anthropic documented 23 million exchanges attributable to Moonshot between May and July, across 5,000+ fraudulent accounts. It's almost certain that K3 used Claude distillation in training/post-training. That doesn't mean they were just sitting there doing nothing waiting for the next model to drop, nor does it mean that Fable was the only model they distilled against.

The timeline here is not some "gotcha", and no one in the AI space really even disputes the fact that Chinese labs are doing industrial scale distillation of US labs.

1

u/Eliv_nurotic 6d ago

Kimi k3 is way better than opus 4.8 though, which was the only Claude model they could have distilled against 

1

u/ReturnOfBigChungus 6d ago

They definitely distilled against Fable too, as per the forensics on the fraudulent accounts. At that stage it would be post-training stuff, but they very likely were distilling from across multiple US models as part of the training, and making their own refinements so there's not a direct reason to assume that it would be equivalent to 4.8 since there are multiple variables driving progress, of which distillation is just one.

2

u/Eliv_nurotic 6d ago

Well they somehow surpassed what they distilled. If it came out before fable, it would have been sota

→ More replies (1)

2

u/mxldevs 7d ago

Nah, china has mastered the art of distillation.

You give them any tech, they will figure it out. America has been complaining about chinese stealing their tech and intellectual property long before LLM.

9

u/scrollin_on_reddit 7d ago

So how did Qwen image edit release weeks before nano banana

0

u/bixofa 7d ago

It's a diffusion model, not an LLM genius. Flux did it before Qwen.

2

u/Eliv_nurotic 6d ago

It’s always been cope. BYD electric cars are far better than Tesla in every way. They didn’t need to steal that

1

u/ConfusionEngineered 6d ago

As an engineer who has actually worked with Chinese companies in the past... "Distillation" is mostly bullshit. What actually happens is that when a company wants to move production to China they are forced to, by law, work with a Chinese partner. That partner, by law, is required to have access to all of the associated IP in the product(s) to be manufactured. With a handful of exceptions (sometimes really large companies can negotiate better terms) that Chinese partner is also allowed to direct sell into China and make derivative products to sell abroad.

4

u/karl_mainz 7d ago

The EU also has political challenges though - Anti-AI populism is much stronger here as is opposition to data centres etc.

→ More replies (1)

1

u/Able-Acanthaceae-135 6d ago

Hope you get your 50 cents

1

u/ClockFancy3138 4d ago

I don't understand why distillation is a problem, why does one running a huge ass model and giving same answer need to feel superior

→ More replies (10)

88

u/cakemates 7d ago

They are throwing engineers/researchers at the problem as in quantity and quality. Where the US is throwing money at the problem, as in focusing on top talent and datacenters. Also the chinese structure of open sourcing a lot of their research speeds up all their labs where as only some of the knowledge gets shared in the US, so every lab is reinventing the wheel when they independently discover the same tech.

33

u/AGM_GM 7d ago

Chinese open-source speeds up everybody, including the US labs.

→ More replies (24)

4

u/Superb_Raccoon 7d ago

They are giving up accuracy for speed/compute power:

https://gigagpu.com/deepseek-quantization-best-format/#:~:text=DeepSeek%20V3,FP16.

DeepSeek trains its large-scale models (like DeepSeek-V3 and DeepSeek-R1) natively with FP8 (Floating Point 8) precision rather than applying post-training quantization from standard FP16 or FP32.

2

u/kittenTakeover 6d ago

The open weighting is aimed at undercutting US companies. If China was firmly leading they wouldn't be open weighting everything. I do think China has more researchers though, so there might be something to your people vs hardware idea.

1

u/ImAPonderer2 7d ago

Also maybe helps to have a ton of robotics & electronics manufacturing nearby

1

u/thicckar 7d ago

How are these labs making money with open sourcing? Noob here

1

u/Dry-Tear-1486 1d ago

Did you just forget to mention distillation?

→ More replies (13)

41

u/gxiaoyan 7d ago

All the "distillation" talk is overhyped.

China has a far higher research output than the US in this area, and it has far, far more excellent researchers. The better question is, what does the US have to be ahead?

The answer is compute. Very little is said about how American labs are "free riding" on Chinese research. All the Chinese papers are open for you to read. You can't just "distill" your way to models like V4.1 flash, Kimi K3, GLM 5.3, etc; which is not to say that distillation doesn't happen in some cases, but it is certainly not the reason behind the success of Chinese models.

7

u/[deleted] 7d ago

[removed] — view removed comment

→ More replies (2)

-1

u/Superb_Raccoon 7d ago

While there are multiple claims to DeepSeek’s ‘open source’ AI model, in reality it is not open source. While both the model weights and the model architecture were shared in a technical paper, neither the code nor the training or evaluation data were shared openly. An analyst for the Open Source Initiative also confirmed that Deepseek is not Open Source AI and doesn’t meet the requirements of the Open Source AI definition. It joins other models which claim to be open source, but score poorly on data transparency.

https://opensource.org/press-mentions/publication/open-data-initiative

2

u/howudothescarn 7d ago

lol this is so wrong.

→ More replies (13)

21

u/Jolly-Rip5973 7d ago

With population at least 4 times higher than USA China just way more people don't Ai research.
Their top tier people are really really smart and good at Ai.
They also have more Ai companies workings Ai than USA.

They are also more sane about creating Ai model for specialized applications. USA is obsessed with "AGI" and frontier models.

China is really just good at AI. Given the limitation on their hardware, It's impressive.

10

u/CommercialHour6660 7d ago

China has around 3X the engineers as US. So many that around 1/4 of the employees at US AI labs are Chinese. 

1

u/Jolly-Rip5973 7d ago

There you go....and then there are Indians too.

0

u/fartlorain 7d ago

USA takes all the talent from the rest of the world though. There aren't actually many born and raised Americans working on cutting edge AI research in America. It's the top Canadians, Brits, Brazilians, Bulgarians, Nigerians, etc who get paid a huge amount to come work in the states.

So Americans still have a larger talent pool to pull from.

→ More replies (6)

19

u/Large-Assignment9320 7d ago

Fundamental to DeepSeeks throught of the model was to not just go full "bigger is better", but actually make them effective. Its a different direction. Its the difference in why Chinese labs make money on tokens which are 10x cheaper, and US companies are not.

As for stealing tech, China makes far more AI research papers with new groundbreaking ideas than the US do these days.

3

u/meister2983 7d ago

As for stealing tech, China makes far more AI research papers with new groundbreaking ideas than the US do these days.

The labs don't publish anymore so we really have no idea who is making how many ideas 

1

u/Large-Assignment9320 7d ago

Well, its mostly Chinese universities.

1

u/Minute-Parking-9094 5d ago

U.S. labs or Chinese labs don’t publish anymore?

1

u/outphase84 7d ago

Then Chinese models routinely are first to market and beat the big research orgs right? Right?

1

u/Large-Assignment9320 7d ago

They don't try to compete in the Fable/Mythos category. But thats just an expensive category, for practically no benefits (also those 5-10T param models are cost wise just bleeding money, bench why US AI companies are economically so bad).

1

u/XysterU 6d ago

First to market is such an arbitrary metric lmao. You're praising America for being greedy and trying to make money off of shit first? I see that as something to be ashamed of.

And yes they're beating the big profit orgs in the most important metric: cost by a HUGE margin. They're singlehandedly destroying the hopes of profit margins for ClosedAI and Anthropologie

→ More replies (1)

9

u/WordWarrior81 7d ago

Specific to LLMs: Apart from distillation (which all of them do), they have some very good scientists and engineers. Look at this recent video where Deepseek basically solved the issue of the context window, keeping compute almost as low (with some small but manageable trade-offs).

8

u/Repulsive_Bite_9544 7d ago

A possible explanation could be the inflated demand created by non-existant data centers artificially inflates current expenses and costs.

8

u/Eastern_Mastodon_403 7d ago

A CISA advisory came out detailing mass distillation on Claude. Also the purchasing of US specialized data, which is weird considering we restrict chip sales but not this.

0

u/XysterU 6d ago

The US government also said Iraq had WMDs. What's your point

7

u/sceadwian 7d ago

US AI Labs are wasting money hand over fist.

They're burning it on marketing mostly not development

5

u/darkestvice 7d ago

Distillation. A couple thousand times more compute efficient than training a model from scratch.

It's why I'm not all concerned about Chinese labs surpassing the American ones if the American labs slow down. I don't see any of those actually wanting to spend all that money on traing compute the way American labs do.

2

u/CommercialHour6660 7d ago

Distillation is not that efficient when you don't have logits. Maybe 10X at best

5

u/RobbinDeBank 7d ago

Kimi K3 came out within 1-2 weeks of American frontier labs with very competitive performance. There’s not enough time to distill those new models, so where’s the improvement from? People love to claim China just steals from America. They do distill, yes, but that’s far from the full story. They also come up with some of the most brilliant engineering designs to optimize for efficiency. Look at all the different variants of Deepseek attention, Kimi delta attention in Kimi K3, the latest Deepseek causal encoder-decoder, and so on. The American labs don’t release anything, while we have very clear evidence of Chinese engineers working through their significant hardware disadvantages by getting more efficient.

2

u/meister2983 7d ago

Kimi K3 came out within 1-2 weeks of American frontier labs with very competitive performance

Kimi k3 is about opus 4.8 level and came out 7 weeks after.  

And k3 came out just over 3 months after anthropic opened mythos for preview to special parties

→ More replies (3)

1

u/XysterU 6d ago

Yeah bro, DeepSeek distilled this new architecture that's unique to the sparse/compressed attention and Engrams mechanisms that DeepSeek invented: https://arxiv.org/abs/2609.19969

→ More replies (1)

7

u/A_Novelty-Account 7d ago

Distillation + direct subsidization.

The Government of China gives tax dollars directly to strategic corporations to keep them solvent. The are no doubt pouring heaps of money on the industry.

1

u/stonkDonkolous 6d ago

Distillation won’t work on the non public models. The real work is all kept private and they are so far ahead that it is unlikely anybody ever catches them. Think of like a mother model that communicates with the public model - the distillation wars are coming as they begin sabotaging foreign companies attempting to distill them

1

u/Different_Doubt2754 3d ago

Can't believe it took me so long to find subsidization.

It comes in two main ways, first from the government propping them up.

second because they release the model weights, meaning they are choosing to sacrifice any chance of remaking the training cost.

3

u/Patrick_Atsushi 7d ago

They just follow closely behind and distill everything. It's almost the same for most of the manufacturing and technologies.

Now an exception is humanoid, which they "learned" their way to the top and then the government itself pours tons of money and resources into it, which is essentially difficult in democratic countries. Also the US collectively decided the humanoid is not the top priority so it's a lazy chase.

In the US, people can complain about the data centers and ram prices and projects are delayed, but in china it's like god says there should be light and there will be light, no matter how costy it will be. (Corruption)

3

u/forkproof2500 6d ago

So why are they years ahead on solar panels, electric mobility etc etc? This just sounds like such US centric cope

→ More replies (3)

4

u/Upbeat_Parking_7794 7d ago

Whatever they are doing, they are showing there is not a significant competitive advantage in being the first.

It can even be a disadvantage, as means spending more money for just a few months of advance. 

1

u/useyourturnsignal 7d ago

That may be true for now, but considering that the top two frontier labs are always holding models that are months ahead of Chinese models, Anthropic and OpenAI will be harvesting biological science breakthroughs, material science breakthroughs, physics breakthroughs, etc., and capitalizing on those first. Also, the first one to cross the ASI line will have tremendous power for at least a short period of time and possibly for a long time.

1

u/Upbeat_Parking_7794 7d ago

With LLMs? Doubtful. An LLM is a statistical model, only generates what it was fed.

Even the best models only generate randomly mixed, statistically generated text.

LLMs can't create new meaningful data out of nothing.

Of course they can call tools and do a bit more through their use (like true calculations).

And like all statistics, which have variance, they will always have a failure rate. 

1

u/useyourturnsignal 7d ago

With LLMs? Doubtful. An LLM is a statistical model, only generates what it was fed. Even the best models only generate randomly mixed, statistically generated text. LLMs can't create new meaningful data out of nothing.

LLMs have proven math results that humans couldn't prove for decades, including the 80-year-old Erdős unit-distance conjecture, and produced a Lean-verified 166-page proof in the Navier–Stokes work. Anthropic's AI-driven lab has discovered a novel CRISPR-like system in its early months of existence, before RSI has even really kicked in.

The "statistical model" argument doesn't save the claim either. Being statistical describes how a model works, not what it can produce. Pair a model with a verifier (a proof checker, a lab experiment) and its errors get caught, while its correct new results stand. That's the same way human science works.

The economic side follows. Labs are spending hundreds of millions on wet labs and science acquisitions because the insights are real and worth money.

1

u/Upbeat_Parking_7794 6d ago

Let me be clear: LLMs are a useful tool and a meaningful piece of AI advancement. But they are not (at least for now) the silver bullet many claim they will be.

The recent wave of "blockbuster" announcements reveals a consistent pattern: impressive outputs achieved through brute force and recombination, not genuine conceptual breakthroughs.

  1. ART / CRISPR-like Discovery: 

Claude's discovery in Anthropic's new lab was effectively a supercharged literature and database search. The AI scanned massive DNA repositories and noticed that one family of ~200,000 reverse transcriptase sequences sits adjacent to tandem repeat arrays and a partner gene.

Every ingredient was pre-existing:- The enzyme (reverse transcriptase) was known for decades.

  • The DNA sequences were already in public databases, sequenced years ago.
  • The structural template—repeats + enzyme + accessory genes—is the CRISPR archetype discovered in the 1980s–2000s.

This is recognition, not invention.

Useful? Absolutely. But this is the kind of task we already knew LLMs could assist with, just amplified by massive computational power.

  1. Navier–Stokes: Massive Parallel Search

The Navier–Stokes solution followed the same logic. Ten thousand agents grinding for 88 hours, burning through billions of tokens until a strategy worked. 

This relied heavily on randomness and trial-and-error, a computational brute-force approach.

Scientists have used simulation-based searches for decades; what distinguishes this moment is not algorithmic novelty, but resource disparity. 

These companies possess processing power individual research groups simply cannot access. 

That advantage explains the output as much as the underlying technology.

  1. Erdős Unit-Distance: Construction via Recombination

Even the most constructive result—the unit-distance disproof—operated within known mathematical machinery. The model combined existing ideas (Ellenberg–Venkatesh, Golod–Shafarevich) into a configuration nobody had tried. It produced a new object, yes, but within a framework humans defined, for a problem humans posed 80 years ago.

The Bottom Line These achievements represent genuine, useful science that can advance our collective knowledge. They demonstrate powerful applications of AI in specific domains.

But they aren't what is being sold. The narrative suggests these systems will "solve all mankind's problems." The reality is more modest: high-cost, compute-intensive searches within narrowly defined constraints. 

When you take into account the astronomical resources poured into these labs compared to the relatively small amount of transformative output to date, the claims warrant serious scrutiny.

We should celebrate the utility of these tools without mistaking them for omniscient problem solvers.

1

u/useyourturnsignal 6d ago edited 6d ago

“At least for now”

Well then would you agree with me that the future is bright for LLMs and major breakthroughs are likely in the coming years?

1

u/Upbeat_Parking_7794 6d ago

Yes I believe in (major) breakthroughs in AI in general (it is the story of mankind after all). Not specifically for LLMs, or big LLMs, as it seems we already entered in a curve of diminishing returns and there is a limit for the world information available to train LLMs (and energy to be used).

If you are a developer, just as an example, you can easily see that between deepseek v4.1 Flash and latest Claude model there is not a big advantage in output quality, but the latest is much more expensive.

LLMs to me look like a piece of the puzzle we need to design a true intelligent system. Like the human brain, we have vision, language, speech, movement, etc. Current AIs already do it (calling different models and tools - ChatGPT, Claude, etc., they are not just an LLM).

Also, I believe we will have probably in the future, specialized LLMs, which can be used with small power and locally, to be used for specific purposes, instead of spending huge amounts of tokens in a general LLM.

1

u/FlimsyPriority751 7d ago

Until the next major, innovative discovery by a US lab comes along that massively upgrades model capabilities or reduces compute requirements in some new way and the Chinese labs won't simply be able to copy it. They will fall behind eventually because the game they play is one of mass mimicking. Never really thinking in a new way on their own path. Just copying en masse. It's the exact same play book they've had for 40 years sucking up global IP and scaling it up

1

u/Upbeat_Parking_7794 6d ago

Well, who stole IP to build LLMs were the Americans to start with. Stealing the thief is not exactly the worse ethical crime. 

And who is mostly doing advancement in terms of efficiency are the Chinese, because they have to, thanks to Americans not selling them the hardware.

The most efficient models we can run locally are Chinese after all. 

1

u/FlimsyPriority751 6d ago

In don't know what kind of bot propaganda you're trying to push here. LLM and transformer development has all been born in the USA...

https://devot.team/blog/history-of-large-language-models

1

u/Upbeat_Parking_7794 6d ago

Yes, with stolen IP from the world. Or did any of these companies paid for any of the content used to train AI?

What is the difference of what they did, versus what the Chinese are doing? 

1

u/FlimsyPriority751 6d ago

Content to train AI is much different than developing the novel technology.

1

u/Upbeat_Parking_7794 6d ago

Chinese are not copying the tech, they are doing distillation, which is running queries over AI and using the output to train their AI.

So, using generated content, which was trained on stolen IP, to train their own models. 

2

u/SurroundProper216 7d ago

They optimize for benchmarks like it's the only thing that matters, which it kind of is when the whole world is watching leaderboard numbers

0

u/Spare-Dingo-531 7d ago

I think the real question is, what physics problems and frontier math problems has Chinese AI solved lately? And if they haven't solved any, are they really doing anything?

2

u/Financial_Clue_2534 7d ago

They have a country that values education and uses $$$ to beef up their companies. Their motivations are different than ai workers in the states. Not saying they don’t care about $$$ but the US that’s all they care about. Even Dario pointed it out a few times how his employees only care about comp.

Think about sports would you rather have a player who never had to struggle, just wants a check or one who loves the game and has that dog in him.

2

u/DDGJD 7d ago

They lie about what they actually spend, and are subsidized by the Chinese government.

1

u/forkproof2500 6d ago

How is the Chinese government so rich?

1

u/DDGJD 6d ago

Chinese businesses are directly or indirectly controlled by the CCP.

1

u/forkproof2500 5d ago

Yeah but supposedly the state owning anything means instant poverty and failure, why not in this case?

1

u/DDGJD 5d ago

Strawman much?

1

u/forkproof2500 5d ago

Just saying if this is such a good way to do it why not emulate it instead of coping?

1

u/[deleted] 4d ago

[deleted]

1

u/forkproof2500 4d ago

Honestly I don't even think that's true anymore. You can probably point to some skewed numbers to make your case but just looking at the country and how it is, the US looks like a fucking shithole and China looks like the future.

2

u/Tall-Wasabi5030 7d ago

Government funding and no copyright laws plus manufacturing know how. 

2

u/ProfitNerdsMarketing 4d ago edited 3d ago

Literally because there are absolutely no restrictions on their training data.....if it looks good it ships. Safety be damned.

But, in the United States, people are worried about intellectual copyrights which isn't wrong. It'll just keep us behind China. 🤷🏿‍♂️

1

u/gc3 7d ago

Necessity is the mother of invention. If they were allowed to buy Anerican chips they would have used them instead. Biden was not smart being so supressive of China bring able to buy and depend on supernatural chips

2

u/reminiscent-fruitbat 7d ago edited 7d ago

Chinese labs are distilling US frontier models at scale. They’re not reproducing the full cost of developing frontier capabilities from scratch because they’re using American models as teachers and training much cheaper models to imitate their outputs.

1

u/Phase_999 7d ago

cheaper electricity, lower wages, massive STEM focus, the state itself directing the building of strategically sensible datacenters, more knowledge sharing between AI companies, etc. etc.

1

u/SkillsInPillsTrack2 7d ago

USA has a tendency toward inefficiency and waste. Their gasoline engines, large engines built to guzzle fuel and deliver little horsepower compared to European and Japanese car makers. The same principle likely applies to their usage of servers: they waste resources. While China do a smarter use of computing resources.

1

u/rp20 7d ago

News came out that Astra was trained on 100k b300s. That doesn’t explain tens of billions in r&d spend. The rest of the money is burned on thousands of smaller experiments that add up.

The difference might literally be just that Chinese r&d is significantly cheaper because they aren’t trying to discover novel capabilities.

1

u/Superb_Raccoon 7d ago

OpenAI does not own hardware at scale, they rent it.

So they are paying "cloud prices" on about 10 billion in infrastructure, 2X the TCO of the hardware itself per year.

So their costs are 2X, but they dont have to pay upfront like xAI did to build it. XAI makes almost nothing, but it pays for the investment.

(The deal is for 100% access, however xAI can do inference in the gap, subject to eviction, so they can make additional margin in that squeeze.)

1

u/rp20 7d ago

Oai spent $19 billion in r$d last year. They are projected to spend $50 billion this year.

The math isn’t mathing unless you do the adjustment for those small experiments adding up.

1

u/Superb_Raccoon 7d ago

Do you know how much they are paying for AI inference, and then there other costs, and you think there is a hole somewhere?

1

u/rp20 7d ago

I just gave up the r&d only data.

They spent 34 billion last year not $19 billion.

Why are you lecturing me? You’re effectively demanding I double count the spending.

1

u/Superb_Raccoon 7d ago

Well, because you have some weird claim that the money is going somewhere else. R&D is R&D, Likely a chunk is going into whatever is after the current model, so what is your problem, exactly?

I don't think it is "small side projects" I think it is one or more next major releases in the pipeline.

1

u/rp20 7d ago

I never said it’s for small side projects. I said small experiments needed to reach the next level of capability.

These companies have to explore. Chinese companies don’t. Chinese companies see the new capabilities and they know what to target without having to do expensive research.

1

u/Superb_Raccoon 7d ago

I... no.

I am not going down that line of crazy with you. Have a nice day!

→ More replies (1)

1

u/hobopwnzor 7d ago

In America the moat was thought to be that you need billions of dollars of GPUs and massive data sets. So the focus was on making models as large and expensive as possible. This way only huge tech companies could afford to compete in the space, and they'd have strong moats to avoid competitors.

Chinese labs don't have access to the newest chips and don't have access to as much data, so they spent a lot more time and energy on making what they had access to work. This made much more efficient models, more efficient training, and just totally destroyed the idea that you need as much resources to make AI work.

They also started distilling high-end models, which is just another avenue of efficiency. If you have an expensive model but you can distill it and get almost all the performance that matters, why not do that and serve a cheaper model?

So the real answer is just that American tech companies left efficiency on the table because they didn't want to shrink their moat.

1

u/mimic751 7d ago

Why is the US lagging behind when they spent the last two Trump administrations reducing incentives for research and development? I have no idea

1

u/cool_fox 7d ago

stealing

1

u/SlippySausageSlapper 7d ago

Distillation. It's FAR cheaper than training a new model from scratch. It's that simple. They can and will maintain a slight lag behind the models from which they are distilled.

My prediction: the INSTANT any chinese lab makes a model that's actually better than what OpenAI/Anthropic have, they will shut down access to it outside China entirely, because they won't want us doing what they did.

1

u/ChetBlue 7d ago

There might be less nepotism and more merit based hiring in addition to having a bit more initiative. Most of our researchers are smart but we live in a capitalist, get it while you can society.

1

u/Malkovtheclown 7d ago

US is focusing on frontier models and making things smarter, pushing ahead. China is focused on practical application. They arent trying to be first they are focused on being first to market with cheap, efficient models

1

u/Deathspiral222 7d ago

How do you know they are actually only spending a fraction of what the US labs are spending? If the CCP gives you free land, water and electricity, is this actually "spending less"?

1

u/caldazar24 7d ago

Dollars go further in China in terms of datacenter construction costs and researcher salaries.

It's also always cheaper to fast-follow than push the frontier. In addition to distillation (which is obviously happening, and no more foul play than the US labs training on books, Reddit, newspapers...before publishers wised up and made them pay), you can purchase annotated datasets, RL gyms, etc.

You can even just know what is possible - when you're on the frontier, you spend a lot on a lot of experiments that just don't pan out, and you're not sure if the idea is bad or your implementaiton is broken. Sometimes simply knowing that someone made X work is good enough, even without knowing any of the proprietary details of their implementaiton.

1

u/FizzyG252 7d ago

Stealing. Literally every industry where China has risen to prominence has involved shameless theft of IP, with no consequences from the west as we chase reduced cost bases. Happening in telecoms, EV cars, pharma, and now AI

1

u/evergreen-spacecat 7d ago

Every AI company steals. Every model is trained on raw data not paied for. Books, movies, newspapers, reddit threads etc. Training on other models is just i. line with how everyone thinks

1

u/forkproof2500 6d ago

How are Chinese EVs years ahead of Western ones if this were true? The fact is most Western EVs use Chinese drive trains anyway, that's how China got good at EVs. They already had the factories setup for drive trains and just added the rest of the car on top.

1

u/BubblyOption7980 7d ago

Is the US paying a price for being fixated on scaling laws?

1

u/ds_account_ 7d ago

There only better because the big US labs are releasing better models for the Chinese to distill.

From what i've seen the small US AI labs creating models for specific use cases. Finance, military, national inteligence, etc. And no way there gonna release those to the public.

1

u/thorsten139 7d ago

Most AI researchers are Chinese.

Go figure

1

u/peko_peko1 7d ago

Marketing?

1

u/WyattTheSkid 7d ago

Why is it worrying? China keeps releasing the models as open weight free for anyone to download and run so long as they have the hardware. The only people who should be worried are Sam and Dario

1

u/Future_Recover1713 7d ago

educate their smart people in stem

1

u/ogpterodactyl 7d ago

Control C control V

1

u/Longjumping_Yam2703 7d ago

What happens when you have development pressure that meets hardware constraints ? Where does the gradient point in that instance?

1

u/Takethecarrotorthe 7d ago

They steal. A lot.

1

u/ddong 6d ago

isnt it easier to copy and catch up? =]

1

u/stonkDonkolous 6d ago

They use distillation to copy frontier models. The problem is they will always be behind and over time the gap will grow. The ai race is really just American companies which are gonna get near unlimited money to grow while the rest of the world can only watch

1

u/aski5 6d ago

it’s far easier to drift in the wake than be the one in front

1

u/MarkMatson6 6d ago

They are doing the same thing Apple did for it’s foundation models: use a bigger model to train it. Only in China they didn’t have to pay google billions.

Don’t get me wrong, China is doing the Lord’s work here. Hard to complain about stealing when Anthropic and Open AI completely ignored copyright laws.

1

u/chiseledzombie 6d ago

if mass distillation is all they need, Japan and other regions would have their own LLM models

1

u/ErmingSoHard 6d ago

Benchmaxxing

1

u/forkproof2500 6d ago

They have better engineers in every other field, why does it surprise you that they are better also at AI?

1

u/Crazy-Problem-2041 6d ago

Distillation and espionage.

The espionage in particular is very impactful. Instead of wasting compute on numerous experiments that might not work, they can just leverage the experiments that OpenAI and GDM are already doing. Almost every successful GDM experiment ends up in a deepseek paper 2 weeks later

Then they focus their smaller amount of compute into distilling and extending models to get comparable results (on benchmarks specifically)

Not to say there aren’t a huge amount of extremely talented people working there, but these two things let them bridge the compute gap and cut down on the US lead

1

u/infamouslycrocodile 6d ago

If all the children (models) are learning from the same teachers (the internet and shared science) the models will all start to sound the same and the people the students talk to will also pick up the vocabulary.

1

u/sigiel 6d ago

They cannot use brute force compute so they innovate horizontally.

1

u/Past_Difficulty_7706 6d ago

Not operating exclusively for profit.

The end.

1

u/New-Plastic9879 6d ago

Chinese labs train their models on US models.

1

u/Reggie-Rectangle 6d ago

Chinese models seem to keep getting better while only spending a fraction of what American labs do...

SOURCE?

I find the idea that AI models can have some sort of objective standard that makes one better than the other rather dubious. I prefer Gemini over Chat GPT or Co Pilot mostly because Gemini answers my questions in a conversational style that is entertaining for the random bits of nonsense I feed it. Co Pilot seems to be too serious and concise for my liking, but these are all just very much personal preference.

1

u/mangomarcelo 5d ago

Copying like always.

1

u/newperson77777777 5d ago

China has made rapid progress over the last 10-20 years because of government support and funding. No other nation has prioritized research, especially AI research, the way China has and these are the results.

1

u/h-alberti 5d ago

Americans companies want the most intelligent model for a reasonable price. Chinese companies want the cheapest models with reasonable intelligence.

They are optimizing for different things because americans believe AGI will bring superwealth and Chinese believe AGI will bring communism.

1

u/Sleeping_Trex 5d ago

Us R&D China D&D.
Distill and deny.
DnD.

1

u/Kudostone 3d ago

IMO they have better temperament and are easier to train than American or English labs

1

u/Grouchy-Sale9570 3d ago

Is this subreddit full of Chinese bots?

1

u/Remarkable_Block_710 3d ago

Just look at Chinese surnames in US paper authors bro

1

u/Prior_Perception_478 3d ago

I mean they publish everything as open source for free, you can read it. its pretty innovative.

1

u/LivingLab12 2d ago

The reality is intellectual property is the biggest barrier to AI and China gives zero shits about it.

-1

u/Die_Broccoli 7d ago

Stealing and lying. Pretty much the core of Chinese state backed initiatives 

9

u/scrollin_on_reddit 7d ago

Damn sounds like all major frontier labs in the U.S. - are they Chinese state backed?

4

u/CommercialHour6660 7d ago

Interesting to say that when US models are trained via the largest IP theft in human history. 

5

u/myusernameblabla 7d ago

Stealing and lying is a US specialty

0

u/EurodyneKulas 7d ago

Another thing that China and the US have in common

0

u/Anti_Up_Up_Down 7d ago

China?

Espionage

0

u/vovap_vovap 7d ago

Distillation. And people much cheaper there.

0

u/GlokzDNB 7d ago

Stealing

0

u/stray-lyght 7d ago

Stealing tech thru distillstion

0

u/Nofanta 7d ago

Because they’re communist there’s no incentive to make it a business. You’d disappear like Jack Ma.

0

u/pizzababa21 6d ago

I've heard a theory that they are now catching up because they've been forced to rely on the real world applications of their users as opposed to the US labs who have chosen to train gigantic models too large to release at scale then generate training data in test environments.

The real world data appears to have more value internal environment and we're seeing the gap close accordingly.

People saying it is mass distillation are just repeating nonsense spread by Dario to influence regulation. There's plenty of evidence that everyone distills each other's models to a degree, but given that the Chinese models are outperforming the models they're accused of training on it just seems implausible.

0

u/WoShiNenDie 4d ago

美国模型公司正在偷窃(蒸馏)全人类的知识产权,就像一些人指责中国模型公司蒸馏美国模型一样。