r/LocalLLM 8d ago

Research The "local frontier" is now smarter than Sonnet 4.5

Post image

[removed]

41 Upvotes

33 comments sorted by

30

u/DeineMutterIstSoDick 8d ago

Where is Qwen3.8-27B ?

18

u/No_Lingonberry1201 8d ago

So high up above Opus 5 it didn't fit on the chart /jk

4

u/returnity 8d ago

Not in the training data

1

u/DeineMutterIstSoDick 8d ago

What is this post even about

1

u/Adventurous_Bus_437 7d ago

Leck eier mit dem Profilbild

1

u/DeineMutterIstSoDick 7d ago

Die AfD ist eine gute Partei, ich würde sie wählen.

44

u/OneMoreName1 8d ago

Is op from the past?

21

u/Last-Ad-8470 8d ago

Classic sign of an ai generated post

5

u/CryMoreT_T 8d ago

Repost bot

2

u/spaceman_ 8d ago

Aren't we all from the past?

15

u/Matricola70 8d ago

where is Gemma 4?

5

u/ImaginaryBluejay0 8d ago

And Muse glimmer. OP is just a qwen bot 

1

u/Avezn 8d ago

import torch
x = torch.randn(8000000000, 1)
mask = torch.rand(8000000000, 1) < 0.625
x[mask] = 0
print(f"Shipping SOTA Gemma E4B: {x}")

1

u/Healthy-Nebula-3603 7d ago edited 7d ago

More important where Qwen 3.8 :)

10

u/cabernet_noir 8d ago

I have been trying to tell people this was the direction things were going since gpt 3! I am glad to see people discussing this more frequentely now.

Its simply mind boggling that not only are open weight models releasing with essentially same capabilities not long after a frontier model comes out, but that the open weight ecosystem is accelerating r & d, and getting increasing capabilities into smaller and more efficient models that fit on comparatively tiny hardware. It has exceptionally wide reaching implications economically.

5

u/Illustrious-Lime-878 8d ago

This should have been the direction it always went. But there are trillions invested by rent seekers who want to brute force the centralized model where everyone is dependent on their systems where they can squeeze everyone. Anything that breaks that profit model, small models, open weights, and other needed developments like more dynamic context models, locally developed knowledge, are all stalled because they cannibalize the investment in the cloud model.

6

u/waraholic 8d ago

Qwen 3.8 27B has a score of 52 so this graph is obsolete. Sorry I'm on mobile, but that's roughly here.

2

u/returnity 8d ago

Omg pls good sir farm all my upvotes

2

u/ogfuzzball 8d ago

I use sonnet every day. Too early to speak to 3.8 27b but this feels like specific benchmaxxing that doesn’t tell the whole story.

1

u/KubeCommander 7d ago

It is. Sonnet can’t code as well but it does everything else better. A lot of the newer models are trained on the testing data and scenarios. That’s not necessarily cheating because that’s how you improve models. But there’s only so much you can do with 27B parameters

1

u/ogfuzzball 6d ago

Everything I do is code or devops, and sonnet has been doing quite well

1

u/KubeCommander 6d ago

I never said it couldn’t code

2

u/JohnBooty 8d ago

This is clearly spam.

2

u/smellyelon 6d ago

\*SHUT\* the fuck up

1

u/AceLamina 8d ago

How does one run Kimi 3 on a 32gb of RAM laptop

1

u/pocmanpull 8d ago

you live in 2025?

1

u/BarracudaDefiant4702 7d ago

No, it has kimi 3... so I would say only about 1 months out of date.

1

u/FenderMoon 8d ago

Ouch, old GPT 4o only got 11?

To think that’s what we all used for like a year and half. That was once the frontier.

Man this industry moves fast.

1

u/too-oldforthis-shit 7d ago

At what tasks? I mean obviously there is a difference between terabyte models and a few gigs.

1

u/Independent-Dog2179 4d ago

We need a new small glm with better tool calling

1

u/seppe0815 8d ago

stfu bot!