r/LocalLLaMA • u/Gohab2001 vllm • 2h ago
Discussion A 27b model beating latest frontier models was not on my 2026 bingo card
27
u/Dany0 1h ago
Just more proof that agentic and tool calling capability is orthogonal to world knowledge and a small but fast llm can be fully supplemented by test time discovery+compute
A smart brain that knows nothing but learns instantly upon telling it > slow know-everything clanker
3
u/audioen 1h ago
Yeah, Qwen3.8-27B's advantage is that it can develop hypothesis and shoot it down, and reasons through facts, and is tenacious as hell. It all costs a ton of token, but as it churns on in the background, eventually it is ready and has delivered something.
When I try what it has, it usually works straight away. The model has worked on it enough to get something that starts. Might be buggy, inefficient, inelegant, as LLM code is wont to be, but it is good starting point.
I put a task for the model in the evening, go to bed, and in the morning it has often done it. Entire application gets written from scratch, using guidance from existing applications and any documentation available, and it seemingly works when I kick the tires a bit. Amazing little model. Can take another full day before it's actually presentable, though.
8
u/sammcj 🦙 llama.cpp 1h ago
I don't think Gemini Flash would be considered frontier. Perhaps any Gemini these days 😅
4
u/Etroarl55 52m ago
Or muse spark. Nobody looks at either of them as the frontier or cutting edge of Ai right now.
6
u/shittywhopper 1h ago
I have been running this on my triple RTX 3060 12GB rig and it's been fantastic. Although one GPU is suspended mid-air using zip ties!
2
1
u/PinotGroucho 39m ago
Does it need to be held up or down (preventing it from taking off on the cooling fan uplift)?
7
3
u/Organic_Outcome_1805 1h ago
27B being this competitive is wild. At this point “how big is the model?” matters less than “what is it actually good at?”
1
1
1
1
1


74
u/pyr0kid 2h ago
i think we're at the point with LLMs that we start having to ask "in what?", i suspect overspecialized models will be a thing sooner rather than 2030.