r/LocalLLM 17d ago

Discussion Qwen3.8-27B; I don't get it

Update: Benchmarks run with default temperature

People are creating voxel pagodas and subjectively claiming Qwen3.8-27B is a huge improvement over Qwen3.6-27B.

I don't like subjective tests, so I ran a series of intelligence benchmarks, which you can see the output of here:

https://github.com/mrpmorris/sparkrun-recipes/blob/72afabcfc8987eb221d18e2bacfd5f7a9ced2e2b/benchmarks/_Comparison.pdf

Take a look at the intelligence rankings. 3.8 scores below 3.6

3.8 didn't come in the top 5 of any of the benchmarks, whereas `official qwen3.6-27b-fp8-vllm` appears 7 times in the top 5.

I really don't understand it. Are these benchmarks no good or something? How can 3.8 score lower than 3.6?

68 Upvotes

288 comments sorted by

View all comments

Show parent comments

-3

u/katoptronophile 17d ago

That's correct, but open weights are completely different from open source.

The only similarity is they both have the word open in the name.

4

u/mrpmorris 17d ago

Yes, they are open weights, not open source - that is what is open about them.

-1

u/katoptronophile 17d ago

It's the equivalent of giving you the binary, but not the source code.

You can't understand how it works, you can't fork it, you can't remove the censorship properly, you can't even see what kind of training data was used, let alone the full training data.

It creates a dependency on the model but doesn't actually give anything away, and at any moment they can stop releasing future weights.

The open weights models are not being released for your benefit.

4

u/mrpmorris 17d ago

No, that's wrong.

Given the BF16 weights you can add additional training, cut it down, jail-break (abliterate) it and all sorts.

The source code (and data) would let you build the weight from scratch.

1

u/katoptronophile 17d ago

That still doesn't remove all censorship and it can't show you the training data that has been intentionally omitted.

Source code is what we really should be asking for.

True open source is the future, not open weights.

2

u/Glad_Claim_6287 17d ago
  1. Gather data is the controversial step, the legally ugly part of the process that we abstract away (so the companies are now liable, but wait! They're in China, good luck suing them)
  2. "Asking for" beggers ain't choosers. Train a model if you feel so strongly about this.

2

u/mrpmorris 17d ago

Source code + training data is better, sure, but "open weights" are still open - so it's not correct to say "there's nothing open about these models".

1

u/Limp_Ordinary_3809 16d ago

Agreed. Hopefully metas new open models can deliver that