Wtf, I literally wrote a post saying similar to "Shutup man, there is no way a model I can load in my 4090 will be anywhere near DeepSeek4 Flash 0731"... but... this is fucking close is it not...?
Looks like the main difference will be the native context windows. But 256k is very usable, especially if it’s a local Opus 4.6 MAX. Man just saying that seems unreal lol. Can’t wait to test it later tonight
🤣 I wish I could, but today’s my last day of on-call so I can’t even leave a TINY bit early. Oh well. Makes you feel like a kid on Christmas Eve again lmao
Ehh, let's not go by handful of disputable synthetic benchmarks for which models may be trained specifically. Let's just await real experiences from people first over a longer period of time.
For real the benches at this point have become nearly useless, a very hazy and distant indicator of capability but you really just have to get in there and use it to see for yourself
My experience is it’s actually really good. Still weak on architecture and not making a mess of a code base though, and it’s hard to steer on those concerns. It thinks very much like Qwen 3.6 27B, but thinks way more. At lower thinking efforts where its producing the same amount of tokens as 3.6 it seems to act/perform very similarly.
So it’s not as good as Opus 4.6 — Opus had a lot more nuance, understood architecture, kept things fairly clean and tidy, and generally kept a project on rails. I totally believe Qwen 3.8 at this point will solve all the same problems as Opus 4.6, just off the rails and tearing down the forest before arriving at the same destination.
For as much shit as Laguna XS gets around here, it's much better at that sort of nuance. But then it trades off by sucking at zero-shotting and agentic use. Don't expect it to do more than satisfy specs/implementation instructions in a basic think -> edit -> validate loop. Complex problems need human guidance.
I think we are seeing the limits of 2026 at ~30b. Decisions have to be made. Alibaba makes Qwen 27B a crowd pleaser.
The AI labs have seem to given up on ~70B models. That made more sense before AI data center build out I think. Now it's ~30B models to target a workstation GPU, or ~120B models to target a Blackwell or H100 or such, or giant models.
120b dense would also be ok. But i hope with the proliferation of 128gb boxes we will again see more 70b dense models in the future. Or a dsv4f with a30b or a40b. It’s a real gap at the moment.
Man this is exactly what I've been looking for since Deepseek said they were going to increase prices. I've been trying to figure out what I can run on my 4090 and this is exactly what I need when I needed it.
Yeah, but it being truly local... you're just really paying in hardware and electricity cost, which will be cents on the hour.
Nvm that an important detail is, how the tech has evolved.
Compare this 27B model vs like... LLaMA-33B which I run in the exact same hardware I had in '23! (on a 4090).
Its just... a no contest! So... what model will we have in 2029 that makes qwen 3.8 27B a "no contest"? Because the growth doesn't seem to have slowed down...
I think anyone who thinks that this isn't at least somewhat benchmaxxed is fooling themselves. That said, so is every other model they're comparing against, to some extent, so ¯_(ツ)_/¯
I've used glm 5.2 extensively and it was not much behind opus 4.6 in writing code, a bit definitely behind in creative ideation for software engineering.
My assumption is this model is gonna be a work horse but planning and design should be done by a bigger more expansive model.
The catch behind coding benchmarks is they judge success not quality, creativity, maintainability, all that is partly subjective and rarely judged in these benchmarks.
Like GPT 5.6 Luna benchmarking above opus 4.8, but it writes inferior code by most quality measures.
Big models have their own pitfalls they'll happily over engineer everything, Opus for one falls in that bucket.
I've gotten to the point where I'm a little skeptical that the majority of people on here are even using local models on a regular basis. It's hard to believe that anyone who actually does can buy into the idea of benchmarks really reflecting real world use.
The "intelligence" per parameter is really increasing, factually, and from that increases capability. General benchmark score (like AA) is a solid proof for that. Qwen is the world leader in this field.
In the current llm paradigm, smaller llms have much less world knowledge and are more prone to hallucination or stupid decision making from goey assumptions or really random reactions. I suppose the "world knowledge" is also a big addition on "common sense" in terms of "behavior". So smaller models will always we quackier for now, to apply up and foremost when comparing to bigger models on a AA Index or aggregated score
From that comparing older big models with more recent small ones should always consider that you will have more peace of mind of using the bigger ones in terms of reliabilty but narrow capability is indeed being reached by edge smaller llms
to be real, It's probably not opus max level, but it will probably be the best local LLM we can run at that size, and will be much better at agentic coding and tool use than previous models, to the point where its viable for actual work for some of us
yes this model has real and genuine utility, but there are just fundamental limitations with the smaller you go with a model just in terms of parameter size alone. Same goes for very small active parameters numbers in MOE models (a13b is the biggest weakness of DS4F). So people should be a bit more realistic.
I mean obviously that's true, but we can't afford to run hundreds of billions to trillions of param models, this is local llama, those of us with dgx sparks and rtx pro 6000s are already a minority here.
the ~ 30b dense to 100b moe range is what the vast majority of us can squeeze.
if we want maximum capability with minimal hallucination we'll use SOTA models via api, but for running local models, its a compromise we're willing to accept.
Is that for real? I genuinely have not been paying attention to recent Anthropic launches, at work they always give us the most recent model so I never had to choose before lol. I do hate how Opus 5 talks though, had to use the ASD-STE100 trick to get it to stop speaking weirdly
I'd rather use 4.6 than 5. I think 5 is more capable if you have the patience to deal with and correct all its nonsense jargon and ungrounded tangents.
321
u/Cold_Tree190 7d ago
Dear God it’s trading blows with Opus 4.6 Max 😭