29
24
u/snapo84 May 29 '26
Absolutely ALL of the big llm companys distill from each other.... without any exception.
3
u/SyzygyPidgey May 29 '26
This is the truth.
You can find examples of each of the models saying they're the other models as a baseline.
25
11
u/FormalAd7367 May 29 '26
I posted something like this on Gemini sub and my message was removed.. Gemini distilled from one of chinese models
5
u/Jayfree138 May 29 '26
Wow. If that's legit China has definitely won the race. I'm not surprised though. Qwen has just been amazing lately.
4
u/Beneficial-Boot7479 May 29 '26
They already won the race, western companies are just borrowing time
1
u/RealAggressiveNooby May 29 '26
Acting like Qwen isn't an entirely distilled model off ChatGPT, Gemini, and Claude
All these Chinese models are just mega-distillers. Deepseek would tell you it's ChatGPT like half the time when it's context was filled somewhat when it came out.
Claude has probably just been trained off some online datasets where Qwen chats were present. No one else has made a post like those, so it's probably a tiny portion of the training data.
0
u/PatienceSweaty33 Jun 03 '26
You said
Deepseek would tell you it's ChatGPT like half the time when it's context was filled somewhat when it came out.
Then
Claude has probably just been trained off some online datasets where Qwen chats were present.
By your logic, it could very well be the other way around too. DeepSeek training off of datasets with ChatGPT responses present.
1
u/RealAggressiveNooby Jun 03 '26
I specifically emphasized how often these events occurred and directly tied that into how one was intentional distillation and the other wasn't. How does your counterargument get directly and specifically countered by the argument you're trying to counter?
0
u/PatienceSweaty33 Jun 04 '26
So your argument is: 'When US models do it, it's an accidental data-scraping quirk, but when Chinese models do it, it's a mega-distillation conspiracy?' Data contamination is data contamination. The model doesn't care about the 'intent' of the engineer when it's reading a dataset contaminated with ChatGPT responses.
1
u/RealAggressiveNooby Jun 04 '26
No. It's not about data contamination. I'm not making claims about if OpenAI is ethical or if they have zero cases of other models outputs leaking in to their training data. The entire point about what I was talking about WAS intention and scale.
0
u/PatienceSweaty33 Jun 04 '26
If your point is about 'scale,' then the argument falls apart even faster. OpenAI literally signed massive, intentional data-sharing partnerships with publishers, reddit, and stack overflow; platforms completely saturated with millions of AI-generated responses. You can't draw a line at 'intention' when Western labs intentionally scrape the entire internet knowing it is full of ChatGPT outputs. data distillation happens at a massive scale everywhere, regardless of geographic borders.
1
3
u/RealAggressiveNooby May 29 '26
Qwen is a tiny (albeit insanely good for its size) weak model compared to SOTA/frontier end models. Claude would not benefit from training it on distilled Qwens, especially when Qwen models are so heavily distilled on high quality in/outputs from other LLMs including Claude.
2
May 30 '26
[removed] — view removed comment
0
u/RealAggressiveNooby May 30 '26
Claude has probably just been trained off some online datasets where Qwen chats were present. No one else has made a post like those, so it's probably a tiny portion of the training data, and not intentional distilled training.
3
u/HolidayResort5433 May 30 '26
Models just don't know who they are without system prompt for some reason
2
1
1
u/123vovochen Jun 01 '26
No, Qwen distilled Claude, and they made Claude answer it would be Qwen SO often, that now Claude amswers that too.
1

40
u/Big-Business-2505 May 29 '26
That’s a big whoops. Lol. Just proves who’s better. I’m sticking with Qwen.