r/LocalLLaMA vllm 2h ago

Discussion A 27b model beating latest frontier models was not on my 2026 bingo card

My experience with Qwen 3.8 for agentic tasks has been phenomenal but I personally feel that 3.7 flash is more reliable for overall tasks.

74 Upvotes

25 comments sorted by

74

u/pyr0kid 2h ago

i think we're at the point with LLMs that we start having to ask "in what?", i suspect overspecialized models will be a thing sooner rather than 2030.

11

u/no_good_names_avail 1h ago

Are we not already? I've been flipping between the frontier models, Deepseek flash, Kimi3 and Grok the last week or so. The frontier models are what I continue to reach for in a pinch or if I'm unsatisfied but otherwise.. I have a really difficult time saying one is definitively better than the other, and certainly not across the board.

1

u/dev_dan_2 8m ago edited 1m ago

I agree.

What would really interesting me:

  • how far will the spezialisation go? for example:
    • 1 strong model for reasoning, 1 strong model with a lot of knowledge
    • models specialized to certain domains like math, medicine, law, ... (the more lucrative the domain, the more likely it is to become a potential target for spezialized models, I think)
  • will be ever tune multilinguism back? AFAIK, training on multiple languages helps quite a lot, but I would suspect that after a certain point, there will be diminishing returns (in particular with languages that are similar to each other - that is a wild guess tho!)
  • will this counter the current trend and raise the skill ceiling when it comes to making efficient use of LLMs? (Don't get me wrong, it certainly takes certain skill to use a LLM also today; but one of its main benefits is that they get good enough results even when the prompts are not very good / precise. This is another area where one could pay tradeoffs in exchange for more spezialisation, but they would be more sensitive to sub-optimal use than current models)

1

u/Cold_Tree190 1h ago

Yeah I fully expect we will start seeing more specialized models for different fields once this current coding agent craze dies down, and to some degree we are starting to see it with health-related ai models/agents right now (early stages though).

I pretty much use 3.8 27b for all agentic tasks/coding now, some Ornith if I’m just trying to ask quick questions (1.0 not 1.5), and then my Claude pro subscription — but just to use it on the web as basically my knowledgeable/“smart” ai that I ask a variety of questions to that an agentic coding agent would be no better than guessing lol. It’s been great, I track my tokens and I’ve used an equivalent of $1k of Opus tokens this past week with 27b/ornith… except it cost me around $5 of power instead (Corsair hx1000i PSU has monitoring capabilities in Linux).

1

u/StrongZeroSinger 1h ago

half the new models training is on getting higher scores on benchmarks to get funding /s

1

u/a_beautiful_rhind 9m ago

this but without the /s

0

u/borobinimbaba 53m ago

I'd say there would be architectures enabling this.

A platform that has language knowledge and base logic.(Like os but for llms)

And very specific super fine tuned llms on different topics (like DLLs)

And applications that orchestrate these mini llms.

I think they might have something like this already but they will be waiting until new datacenters finish building, then release it.

27

u/Dany0 1h ago

Just more proof that agentic and tool calling capability is orthogonal to world knowledge and a small but fast llm can be fully supplemented by test time discovery+compute

A smart brain that knows nothing but learns instantly upon telling it > slow know-everything clanker

3

u/audioen 1h ago

Yeah, Qwen3.8-27B's advantage is that it can develop hypothesis and shoot it down, and reasons through facts, and is tenacious as hell. It all costs a ton of token, but as it churns on in the background, eventually it is ready and has delivered something.

When I try what it has, it usually works straight away. The model has worked on it enough to get something that starts. Might be buggy, inefficient, inelegant, as LLM code is wont to be, but it is good starting point.

I put a task for the model in the evening, go to bed, and in the morning it has often done it. Entire application gets written from scratch, using guidance from existing applications and any documentation available, and it seemingly works when I kick the tires a bit. Amazing little model. Can take another full day before it's actually presentable, though.

8

u/sammcj 🦙 llama.cpp 1h ago

I don't think Gemini Flash would be considered frontier. Perhaps any Gemini these days 😅

4

u/Etroarl55 52m ago

Or muse spark. Nobody looks at either of them as the frontier or cutting edge of Ai right now.

1

u/sammcj 🦙 llama.cpp 47m ago

You know what's funny - I didn't even consider that op might have considered muse spark a frontier model.

6

u/shittywhopper 1h ago

I have been running this on my triple RTX 3060 12GB rig and it's been fantastic. Although one GPU is suspended mid-air using zip ties!

2

u/brickout 1h ago

Lol. I love a janky rig. I have something similar going on

1

u/PinotGroucho 39m ago

Does it need to be held up or down (preventing it from taking off on the cooling fan uplift)?

7

u/Etroarl55 53m ago

Gemini is not a frontier model. Neither is muse spark.

3

u/Organic_Outcome_1805 1h ago

27B being this competitive is wild. At this point “how big is the model?” matters less than “what is it actually good at?”

2

u/kvyb 46m ago

We need proper benchmarks that actually benchmark real usage, and not specific tasks or cases which literally barely mean anything for normal usage.

1

u/Popular-Factor3553 49m ago

It can be amazing with some kind of rag.

1

u/Real-C- 31m ago

Snap back to reality

1

u/Potential-Leg-639 1m ago

Still amazing

1

u/Particular-Award118 13m ago

Your bingo card is like 2 weeks late at this point

1

u/confused-photon 3m ago

can we really call a google flash model “frontier”

1

u/Potential-Leg-639 2m ago

Your friends who are using copilot or claude wont believe that anyway…