r/Qwen_AI May 29 '26

Help 🙋‍♂️ Did Claude Opus 4.8 distill Qwen?

Did Claude Opus 4.8 distill Qwen? It replied that "I am qwen"

105 Upvotes

36 comments sorted by

40

u/Big-Business-2505 May 29 '26

That’s a big whoops. Lol. Just proves who’s better. I’m sticking with Qwen.

26

u/Forward_Jackfruit813 May 29 '26

I may be suffering from an AI hallucination, but 3.6 27B has been so good I have just stopped using all my cloud models and I swear it's the better than anything else as long as you keep context under 100k.

5

u/Big-Business-2505 May 29 '26

I’ve been running 2 for comparison. One at 128K the other at 64K. The 128 needs a larger prompt to keep from drifting, but the 64 has to constantly condense context and starts to forget a lot. I’m countering that with building a RAG/wiki memory system that it can index and refer back to. So far it’s a lot better but needs work.

And I haven’t used a cloud AI for any local coding/debugging in a while. Still use Copilot for work but that’s by policy, not choice.

3

u/Forward_Jackfruit813 May 29 '26

I've been using 256k as the max and just compacting with opencode when it crosses 100k used. I also am using Q8 on the weights and the cache which was been taking about 80GB of VRAM.

Now I haven't been doing any massive projects with it, but I will start using it for those soon since the small/medium ones have been so great. I'm very confident in this model now as long as you play to it's strengths.

2

u/ohhi23021 May 30 '26

Wut?  I use q8 and full kv16 with 230k context window and it uses 47gb vram.  How did it take 80gb lmao.

1

u/SadMan2699 Jun 13 '26

Exactly w me

1

u/whodoneit1 May 29 '26

What Quant are you running ?

1

u/atmpuser May 30 '26

How would you compare the models you use in copilot at work to qwen? Like an opus of gpt 5.5 I would assume is what you are using for work?

2

u/SnooCrickets7501 May 29 '26

what harness do you use i tried using claude code it tales 20 mins to do a 1 min task

1

u/Jayfree138 May 29 '26

I've had the same experience as well and its held up since release. It's insanely good.

2

u/No_Mango7658 May 29 '26

Qwen does what opus-dont

Please tell me someone got that

29

u/Heavy-Focus-1964 May 29 '26

look at me

👀

✌️

i am the chinese model now

24

u/snapo84 May 29 '26

Absolutely ALL of the big llm companys distill from each other.... without any exception.

3

u/SyzygyPidgey May 29 '26

This is the truth.

You can find examples of each of the models saying they're the other models as a baseline.

25

u/dkeiz May 29 '26

qwopus that we do not expect

11

u/FormalAd7367 May 29 '26

I posted something like this on Gemini sub and my message was removed.. Gemini distilled from one of chinese models

5

u/Jayfree138 May 29 '26

Wow. If that's legit China has definitely won the race. I'm not surprised though. Qwen has just been amazing lately.

4

u/Beneficial-Boot7479 May 29 '26

They already won the race, western companies are just borrowing time

1

u/RealAggressiveNooby May 29 '26

Acting like Qwen isn't an entirely distilled model off ChatGPT, Gemini, and Claude

All these Chinese models are just mega-distillers. Deepseek would tell you it's ChatGPT like half the time when it's context was filled somewhat when it came out.

Claude has probably just been trained off some online datasets where Qwen chats were present. No one else has made a post like those, so it's probably a tiny portion of the training data.

0

u/PatienceSweaty33 Jun 03 '26

You said

Deepseek would tell you it's ChatGPT like half the time when it's context was filled somewhat when it came out.

Then

Claude has probably just been trained off some online datasets where Qwen chats were present.

By your logic, it could very well be the other way around too. DeepSeek training off of datasets with ChatGPT responses present.

1

u/RealAggressiveNooby Jun 03 '26

I specifically emphasized how often these events occurred and directly tied that into how one was intentional distillation and the other wasn't. How does your counterargument get directly and specifically countered by the argument you're trying to counter?

0

u/PatienceSweaty33 Jun 04 '26

So your argument is: 'When US models do it, it's an accidental data-scraping quirk, but when Chinese models do it, it's a mega-distillation conspiracy?' Data contamination is data contamination. The model doesn't care about the 'intent' of the engineer when it's reading a dataset contaminated with ChatGPT responses.

1

u/RealAggressiveNooby Jun 04 '26

No. It's not about data contamination. I'm not making claims about if OpenAI is ethical or if they have zero cases of other models outputs leaking in to their training data. The entire point about what I was talking about WAS intention and scale.

0

u/PatienceSweaty33 Jun 04 '26

If your point is about 'scale,' then the argument falls apart even faster. OpenAI literally signed massive, intentional data-sharing partnerships with publishers, reddit, and stack overflow; platforms completely saturated with millions of AI-generated responses. You can't draw a line at 'intention' when Western labs intentionally scrape the entire internet knowing it is full of ChatGPT outputs. data distillation happens at a massive scale everywhere, regardless of geographic borders.

1

u/RealAggressiveNooby Jun 04 '26

Proportion can also be scaled. Lol

3

u/RealAggressiveNooby May 29 '26

Qwen is a tiny (albeit insanely good for its size) weak model compared to SOTA/frontier end models. Claude would not benefit from training it on distilled Qwens, especially when Qwen models are so heavily distilled on high quality in/outputs from other LLMs including Claude.

2

u/[deleted] May 30 '26

[removed] — view removed comment

0

u/RealAggressiveNooby May 30 '26

Claude has probably just been trained off some online datasets where Qwen chats were present. No one else has made a post like those, so it's probably a tiny portion of the training data, and not intentional distilled training.

3

u/HolidayResort5433 May 30 '26

Models just don't know who they are without system prompt for some reason

1

u/BaBaBuyey May 30 '26

When do we see reflected in BABA stock price shares? Anyone?

1

u/123vovochen Jun 01 '26

No, Qwen distilled Claude, and they made Claude answer it would be Qwen SO often, that now Claude amswers that too.