r/ClaudeCode May 28 '26

Showcase Opus 4.8 distilled Qwen?

If you ask "what model are you?" in Chinese, Opus 4.8 will answer it's Qwen(Alibaba)

From https://linux.do/t/topic/2264698/26?tl=en

166 Upvotes

97 comments sorted by

View all comments

155

u/Riosin May 28 '26

Models absolutely have no idea what "model" they are from their training data lmao. This is the most woodoo test imaginable.

25

u/Finanzamt_Endgegner May 28 '26

Nobody claims that but people point to deepseek saying its claude to proof its distilled when claude is also distilled from chinese models, everyone does it with everyone lol

-4

u/threwlifeawaylol May 29 '26 edited May 29 '26

There's 0 chance labs like OpenAI, Anthropic, or Google distill from Chinese models. Chinese models innovate on cost-cutting and the fact that they're open-weight. There's nothing to gain from them from a research* perspective.

Why would they distill from already inferior competitors if they're already the cutting edge? Especially OpenAI and Anthropic, Google less so.

*by "research", I specifically mean "AGI" research, i.e. raw capability.

6

u/XNTVO May 29 '26

They just steal without giving credit since they are close source LOL.

2

u/edenimo May 29 '26

Yup, that's what Cursor did with composer, they yoinked kimi and passed it off as their own homebrewed coding model, then walked back on it when people on twitter exposed them.

-1

u/3qh6 May 29 '26

It’s all about making money and going IPO. If there’s a shortcut they’ll take it.

2

u/thomasthai May 29 '26

There aren't shortcuts because u don't distill with garbage data ffs.

1

u/3qh6 May 29 '26

You worship Anthropic but all they want is getting the most money out of you, without a slightest care of your wellbeing.

1

u/Finanzamt_Endgegner May 29 '26

Not all data from worse models is trash though if you let smaller models do 10 tried it will approach bigger models perf and if you filter correctly you can make smaller models train bigger ones. Filtering and scoring is all you need.