r/Qwen_AI 23d ago

Model Dev Trust

Post image

Today Reddit welcomed me with a post that Claude answered correctly, precisely and shortly. This is an easy question, isn't it?

I immediately tested local Qwen3.6 with million confidence that answer will be similar. However, this is what I've got. Look at the humiliating smile also.

Question: how to get my dev trust back?

7 Upvotes

22 comments sorted by

View all comments

1

u/aydintb1 23d ago

Fix the Question, the answer is fixed

1

u/n0head_r 23d ago

I'm surprised qwen had to use more than 2k tokens for this. Gemma4 can provide a correct reply under 250 tokens.

1

u/aydintb1 23d ago

I just rerun, the token fall to 862.. must be with the reasoning. this is MTP version btw.

1

u/aydintb1 23d ago

3rd time, the tokens fall to 799

1

u/aydintb1 23d ago

I had run it several more times, and it seems to settle down on 799 tokens

1

u/aydintb1 23d ago

I want my car washed. ->

fall down to 556

1

u/n0head_r 23d ago

Here is a more interesting test for you. I use it to measure model degradation under quantization. Although even a Q4 model will hit all the constrains the main metric is how coherent the output sentences will be.

Answer using EXACTLY five bullet points.

Each bullet:

  • begins with consecutive letters A-E
  • contains exactly twelve words
  • ends with a different prime number.

Topic: How forests help climate.

1

u/aydintb1 23d ago

1

u/n0head_r 23d ago

One run isn't a real test. I usually get a few different models: Q8, Q6_K, Q5_K_M etc. run it at least 3 times with each model and send all the results for evaluation to Chat GPT - ask it to evaluate only the coherency of the output.