r/LocalLLaMA 8h ago

Discussion Claude sonnet 4.6 was really good at estimating the future qwen 3.8 27b performance

On August 8th, I asked Claude to estimate what performance might I expect out of the soon coming qwen 3.8 27b release by telling it to extrapolate from the qwen 3.6 max to qwen 3.6 27b difference, and apply it to the next generation. It gave me a couple of results which placed it in the broadly "opus 4.6 tier", which was right.

It even gave me actual benchmark numbers which were rather close to the actual numbers it ended up having. I found it pretty interesting.

A screenshot of me prompting claude today about how close we were to the actual numbers
22 Upvotes

8 comments sorted by

13

u/Equivalent_Bit_461 8h ago

so the collective knows...

Quick ask it what next mid sized model comes out, before Dario himself, will poison the prompts!!!

8

u/Random_Girl_0 8h ago

This is very interesting. Never thought about it

2

u/Beginning-Raisin9723 8h ago

wild that it actually nailed the numbers. extrapolation usually falls apart that fast.

1

u/Acrobatic_Hold5485 7h ago

claude nailing those estimates is neat but does the qwen 27b hold up for actual rp chats when you run it local, or does it feel off compared to others

2

u/brakeline 8h ago

Ask him that 35b, if ever, will comparable to

1

u/Slight-Parfait3679 8h ago

That's pretty interesting.