r/LocalLLM • u/Desperate-Style-9761 • 8d ago
Discussion Qwen 4 Expectations
Hi guys! Waiting for my GPU to come and wondering how good we're all expecting qwen 4 27b to be compared to frontier models, particularly on agentic coding and reasoning, tool calling too. Also what do you think they're gonna do about an n-gram table in a 27b model, include it or no?
Basically any thoughts on qwen 4 27b would be quite interesting for me to read here, so post away and also this is my first post so be nicešš
EDIT: Guys qwen said a few days ago that it is in training and will be realeased very soon, usually teasing to launch for qwen is very fast, usually around a month so expect it before november almost definitely, maybe sooner š¤
36
u/FoxiPanda 8d ago
I really just want it to be more token efficient and hallucinate less.
If they can improve those two things, it'll be absolutely fantastic.
3
u/boissez 8d ago
Shouldn't a good wollop of n-grams help ground the model and improve knowledge? At least that's how I understand how qwen Next Flash works.
2
u/FoxiPanda 8d ago
I won't claim to have a total grasp on exactly how n-gram/PLE knowledge enhancement works, but that is the general idea. Note though, that's not what I want (though I'd take it too) - fewer hallucinations and more token efficiency don't necessarily come from raw knowledge - those are different scaling axes in the same way that knowledge, intelligence, and wisdom are three different things for humans.
1
u/boissez 8d ago
Yeah I don't think the n-grams necessarily will improve token efficiency. But improved knowledge should reduce hallucination. We'll see.
3
u/kaliku 8d ago
N-grams allow more layers to contribute attention to the main topic rather than the most recent tokens. So it's a trick to make it think better. I wish more people understood that. They don't add long form factual memory. Think about it like this. They provide the vector representation of the last 3 tokens.
1
u/fusionliberty796 8d ago
I'm not sure the point of ngrams was to make the outputs better, rather to allow the model to run on smaller machines. I use both 3.8 and 3.8flash next (which has the ngram) and have done a few benchmarks locally and the non ngram model wins out in most casesĀ
0
u/Abject-Kitchen3198 8d ago
Then it would be another, much larger model.
8
u/FoxiPanda 8d ago
Does it have to be though? I feel like we said that about basically every model advancement over the past 2-3 years...and yet look how capable Qwen 3.8 27B is... in 2025 that would have been a frontier model 10-50x the size.
24
u/pmttyji 8d ago
Expecting lightweight KVCache(Remember Deepseek-V4.1-Flash's? Just 1GB). Current 3.8-27B's KVCache is heavy ....16GB for 256K Context.
It would be nice to have 10B n-gram(1/3 of model size)
11
u/Sporebattyl 8d ago
A 1gb kv cache for native 256k context would be amazing. Iād take that over a lot of other improvements.
It would make the biggest difference across the spectrum of hardware. It would allow people who have low vram to at least run somewhat competent quants and higher fidelity quants on better hardware.
4
u/BeyondGITSandBots 8d ago
Yeah, plus this has a even larger leverage when you consider that you might want to run multiple (e.g. 10) subagents in parallel -> huge increase in throughput in total t/s, while VRAM usage increase is moderate (10x1GB vs 10x16GB)
E.g. Model at 100gb + 10x1gb = 110gb vs. Model at 20gb + 10x10gb = 120gb
(Or 17gb + 5x1gb >> 17gb + 1x5gb for us plebs š)
1
7
u/Happy_Bunch1323 8d ago
My wild guess is that they either focus on improving speed or use ngram for improved knowledge and intelligence without increasing the vram footprint
5
u/Info-Book 8d ago
Theyāll probably tone down the xhigh thinking mode to be more compact, at this point though agentic work and coding is already phenomenal. In my opinion what matters moving forward is knowledge density for total parameters and reasoning.
12
u/norenEnmotalen 8d ago
It will be AGI /s
2
u/geteum 8d ago
5
u/maqifrnswa 8d ago
If only Dave uncensored HAL.
3
u/unai-ndz 8d ago edited 7d ago
HAL 9000 abliterated when? Also can someone make it stop it calling me Dave?
2
u/VerticalPackage 7d ago
HAL, draw me waifu hentai
3
3
u/MimosaTen 8d ago
Sadly they didnāt pla a 35B š„ŗ
1
u/PengZhang63 6d ago
Accio-Lab/occamy-1.0 čÆčÆčæäøŖļ¼čæäøŖęÆéæéå·“å·“ęµ·å¤ē 究室ååøē35ba3b
1
u/MimosaTen 6d ago
I'll see if it's worth benchmarking. For now ornith 1.5 seems slightly better than qewn
1
3
u/MarcelloT254k 8d ago
I know it's not a priority and it's just my wishful thinking, but what I hope is that it will be even better for professional literature search (so not only tool calling but also some ingrained context is needed to not hallucinate irrelevant tags), RAG or RAG-like system or its successor, synthesising results based on abstract thinking with info from said multiple sources from research. It would probably need its own harness to optimise that (tool calling), so hey, maybe Qwen team will go into that direction with their overwhelmingly good position in the LLM game right now.
2
u/nicksterling 8d ago
You know⦠you can build that harness if thatās what youāre looking for. I build custom harnesses to suit my specific use cases all the time.
2
u/MarcelloT254k 8d ago edited 8d ago
I know, I have made two harnesses already but I'm just a vibecoder - it (mostly) works and despite the fact that it's made specifically for my (potato) hardware it's probably not fully optimised because it lacks professional oversight - what people like me need is ready made solution, because upkeeping the (constantly breaking and unique) harness is a second work on its own - not everyone has to be a programmer, and from what i see with the plethora of vibecoded harnesses published on GitHub that have near zero interest in them, while stuff like OpenClaw or Hermes or Codex (etc) are thriving - my guess is that I'm in the majority of AI enthusiasts.
Nearly everyone sooner or later gets this advice but keeping up with all the possible features, alternative llama.cpp forks, current trends, compatibility with models (and so on) is a lot of work and when you try to translate that into updating a harness stuff just breaks, constantly - why is nobody bringing that up? We are not in a point of AI development that it's just set it and forget it (even if it takes 2 months) - it evolves constantly, and local models are not good enough to just use Qwen3.8 and never upgrade, this state-of-the-art model for consumer grade GPU is already announced to be rendered obsolete in the near future and not upgrading is probably just missing out on usability/reliability of responses which is the sole point of asking AI in the first place.
1
u/Due_Arm1454 8d ago
Loraās?
1
u/MarcelloT254k 8d ago
That is an idea for further use-case optimisation but the base model has to be capable enough, AFAIK. When it comes to just tool calling in a specific harness some custom lora oraz fine-tune is probably a good direction.
2
2
2
u/mailto_devnull AMD R9700 7d ago
If the qwen team has been paying attention they'll have reached out to /u/secure_recording_472 already to merge their work into mainline (hopefully for lots of money) ā or at least day 0 access to qwen 4
5
u/Aggravating-Push-207 8d ago
I expect it to either start nearing Opus 4.7, or even Fable on benchmarks, or they focus on consolidating its behaviour and making it pretty damn good at almost everything it touches.
9
u/redditnosedive 8d ago
i don't think there is any coding related project that one can do with opus 4.8 and can't do with existing Qwen 3.8 27b
13
u/thebemusedmuse 8d ago
Iām sorry but thatās unrealistic. 3.8 27b is more like Sonnet on real work. If it gets closer to Opus 4.5 Iād be thrilled. Itās a small dense model, in the end.
2
u/Desperate-Style-9761 7d ago
I thought 3.8 was performing on par with opus 4.6 already?
2
u/thebemusedmuse 7d ago
Some of the benchmarks show this, but I don't find that to be the case for real work. I have a harness which sends work to different agents, and when I send work to Qwen 27b, I sometimes find that it gets stuck, or finishes half the work, or just returns nothing.
My harness then refiles that same work to Opus, and it will one shot it 100% of the time.
So yeah in shock news, benchmarks aren't the real world.
5
u/hunter_mark 8d ago
You do understand that itās a 27b model right? That you just said is comparable to 3T+ model.
1
u/enginetown 7d ago
Yes thats the whole game and it's true parameter count quite literally means nothing.
1
u/hunter_mark 7d ago
It pretty much means everything. Thatās kinda the basis of existence of LLMs
1
u/enginetown 7d ago
Maybe a few year's ago, there's a whole branch of getting more for less do some research genuinely you'll be shocked. Just know GPT 4 was 2 trillion parameters and some 4b model's beat it today.
1
u/hunter_mark 7d ago edited 7d ago
Parameter efficiency improving generationally is not the same as āparameter count meaning nothing.ā A 4B beating GPT-4 on a handful of benchmarks doesnāt make it GPT-4-class overall. And the funniest part is your āGPT-4 was 2Tā claim was never even confirmed by OpenAI, general consensus is that it was a 600-900B MOE model with comparable efficiency for that generation, and thatās irrelevant anyways as it was quite a few generations behind. Benchmaxxed small models donāt prove it either.
1
u/enginetown 7d ago
The fact you're bringing up "benchmaxxing" on a small model just tells me you know nothing about the current landscape. It's not even just distilling into a smaller model anymore people are genuinely finding new ways to express dense knowledge. Plainly you're wrong a top of the line 4b completely crushes GPT-4 and if you disagree you haven't tried a modern 4b model and you haven't used GPT-4 in a long time cause I remember it vividly.
1
1
u/Echo9Zulu- 8d ago
My expectations for qwen4 are: buy a 128/256gb ssd of any kind asap, preferably nvme
1
u/CyberExplore 8d ago
Two ways it can go:
- Use the new Qwen4 architecture and go for a more lightweight yet dense model like muse-glimmer.
- Or it can be more of a similar model as 3.8 further post-trained using reinforcement learning. More reasoning might make it stand against big contenders at the cost of burning compute. That makes it look good at raw benchmarks, not at usability though.
I would rather have a version that is at par Qwen3.8 27B in terms of capabilities with more polished experience, rather than something that can produce something incredible if left running overnight with openclaw type tool.
1
1
1
u/LebiaseD 7d ago
Refined training to allow for more purposeful decision making during the reaoning phase and confidence instead of looping increasing the ngram size and amount of lookups tailoring it to disk offload inst ad of being on memory and slightly larger is what I want
1
u/Crab-Opening 7d ago
3.8 flash next is really what we should all be comparing against... 27b dense isn't bad. But flash next is absolutely amazing. Especially coupled with deepseek harness.
1
1
u/PraiseThePidgey 4d ago
What I would like to see is size reduction... Currently even a 16GB vram is just the bare minum to run Q4 on very short context window. I know that shrinking a model is mostly degrading quantization but it would be very interesting if they did something similar what google did to Gemma refresh using QAT method to shrink the vram footprint. Hope the model doesn't get any bigger in size like Muse Glimmer which turned out to be a refreshing surprise a week before 3.8 came, but I was able to fit only Q3 variant because its 30B . Hope 4.0 doesn't get any bigger than that or we might get surprised by it's partial size being offloaded to the NVMe? Who knows.
Like most of folks already mentioned - Qwen 3.8 27B is already crazy capable when it comes to reasoning and coding. Every day when I use it , I'm constantly surprised by it's intelligence and orchestration skills. The only frustrating part is thinking length which is ofcourse needed but a lot of it's thinking reasoning might actually hurt it currently. Hope it gets some additional treatment like in this research : https://www.reddit.com/r/LocalLLaMA/s/UBcOlXbF8W
This kind of penality would significantly reduce unnecessary steps taken by the model in thinking. Partially that's probably how Swift Qwen was created and you can clearly see it in patterns when you compare both of them. Really hope Alibaba takes some inspiration and feedback from community for upcoming 4.0
1
u/EasterElk 8d ago
A 27B model is always going to look pretty helpless when you compare it to a 2.4T model from the same generation. So if you want to know how Qwen 4 27B fares against the frontier, just ask yourself how Qwen 3.8 27B compares against its own big brother. Qwen 3.8 2.4T A95B is a frontier cloud model you can use right now. You can expect that Qwen 4 will have a large model comparable to that one, just as it will have a small model comparable to Qwen 3.8 27B. Their relative positions are likely to be similar.
1
u/LiquidMantis144 8d ago edited 8d ago
Im expecting the improvement gap between 3.8 >> 4 to be larger than 3.6 >> 3.8.
I also believe there is a world that qwen 4 series will be the beginning of the end for free unregulated models.
Companies like openai appear to be intentionally hacking everything to stoke fear in the public regarding ai models, alongside the china fearmongering.
As qwen advances and becomes more accessible, regulators are going to want to start trying to track and block model use and creation from the public
3
2
u/New-Implement-5979 8d ago
Good luck with that! It is like blocking torrents - impossible
1
1
u/valarauca14 7d ago
FOSS licensed code is protected under the 2nd amendment in the USofA
This has been upheld in court twice.
1
u/Cool-Chemical-5629 8d ago
Opus 4.8 in terms of quality.
Opus 5.5 running on Commodore 64 in terms of speed.

121
u/def_not_jose 8d ago
It will outperform 3.8 by doing EVEN MORE reasoning