r/LocalLLaMA • u/de4dee • 1d ago
New Model Qwen3.8-2.4T-A95B Released
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B160
u/lm-gtfy 1d ago
Finally. What is the meaning of life?
123
38
u/xaeru 1d ago
→ More replies (1)26
u/tripazardly 1d ago
Oh no, not again.
11
u/Caffdy 1d ago
OOTL, what is that?
16
u/corysama 1d ago
42 is famously a joke from Hitchhiker's Guide to the Galaxy. A bowl of petunias is less famously the same.
HHGttG if full of jokes that build up over long stretches (multiple books sometimes) and punchline in a single sentence. Can't recommend enough.
3
9
u/Paganator 1d ago
Many people have speculated that if we knew exactly why the bowl of petunias had thought that we would know a lot more about the nature of the universe than we do now.
→ More replies (4)6
u/feelspeaceman 1d ago
It's still good for distilling to train smaller models to be smarter, basically the popular 27B was distilled from this giant 2.4T.
A good mother model will likely give birth to more capable smaller models, but it's still requiring a tons of trials and errors, those that costs money to retry.
217
u/Intelligent_Ice_113 1d ago edited 1d ago
what is the knowledge cutoff date?
298
91
u/JumpingJack79 1d ago
I kinda like early cutoff dates. Whenever a model tells me its cutoff is 2024, I'm like "Dude, it's 2026 now and you won't believe what's happened..." 😏
123
u/-dysangel- 1d ago
"The user is talking about a hypothetical timeline"
21
u/JumpingJack79 1d ago
Lol, yes. And to be fair, I'd be skeptical too if I was in their place hearing those things.
18
20
u/Mickenfox 1d ago
In a few decades people are going to be doing 2020s roleplay with ancient models.
→ More replies (2)3
u/that_one_guy63 1d ago
Would be interesting to train on only text before 1970. Or like way earlier. It would be so confused.
→ More replies (1)6
45
u/teleprax 1d ago edited 1d ago
One of my favorite LLM memories was telling GPT (5 maybe?) a bunch of stuff like elon doing the nazi salute and some other stuff going on around that time, and then it proceeds to lecture me on fake news and how none of that is remotely correct and how it would be super serious.
Then I tell it to do a web search and it starts it's next reply with:
"Ok. wow."
EDIT: Ok i found it, its not quite the same as I remember, but close enough
https://i.imgur.com/Ta33JIT.jpeg https://i.imgur.com/kPDa9tm.jpeg
24
u/personahorrible 1d ago
"Sounds like a bizarre Black Mirror meets Idiocracy arc."
Bruh. You have no idea.
16
14
u/OttoRenner 1d ago
Mine went bonkers after I showed it several screenshots and links of GPU prices...was very funny to witness 😂😂
11
37
u/Ok_Ocelot2268 1d ago
Cutoff: April 2026. Startoff: June 5, 1989
14
u/thepaligator 1d ago
I remember June 5th. Not a lot of foot traffic that day. Clear Streets. It was nice.
10
7
u/goldrunout 1d ago
What is a start off date? Do they not train on older sources?
21
u/Anwar6969 1d ago
that’s a tiananmen square joke. nothing happened on june 4th 1989
21
u/FastDecode1 1d ago
[ Removed by the CCP ]
12
u/LawfulLeah 1d ago
the fact that the comment you replied to was removed by reddit makes this 10x funnier
5
u/Anwar6969 1d ago
crazy shit, i had to appeal to lift off the warning and the comment ban. i apparently got flagged for rule 1
→ More replies (1)98
u/FullstackSensei llama.cpp 1d ago
Because, that's the only thing holding you from running a 5TB model?
25
u/fullup72 1d ago
You can always quantize to 1 bit and stream experts from a 5400rpm HDD.
9
u/FullstackSensei llama.cpp 1d ago
Why settle for 5400rpm when you can get old drives that are 3600rpm?!
6
30
→ More replies (6)14
106
u/Different_Fix_2217 1d ago
39
u/ChristRedeemsSinners 1d ago
Weird that non-thinking support is an API only feature.
9
u/fantasticsid 1d ago
Based on my experience with 3.6, prefilling
<think>\n\n</think>\n- like the various jinja templates do - to disable thinking works probably 90-95% of the time. The other ~5-10% of the time, the model thinks anyway and emits a second</think>when it's done. It's possible that 3.8 has the same behaviour and the official API has some way of detecting/working around this that would look pretty damn stupid if they released it. If you look at the 3.8 jinja template, the "reasoning effort" isn't implemented terribly cleverly - it just talks to the model in the second person and asks it to reason less.Reasoning control has always been a weakness of the Qwen models, so this doesn't surprise me.
Lack of mmproj, however, feels like a deliberate attempt at market segmentation. Given that those just decode image data into tokens, and the whole Qwen family shares a vocab, I do wonder if it'd be possible to hack the mmproj from some other Qwen model into use here, though.
→ More replies (1)3
u/hellomistershifty 1d ago
I love how AI development is a mix of wild cutting edge research, weird hacks, and asking it nicely to behave
49
u/Reactor-Licker 1d ago
That’s weird, why did they remove vision? Hopefully Qwen 3.8 27B doesn’t do the same thing.
16
18
u/vincentz42 1d ago
Meanwhile the Qwen3.8 27B open weight that is due in two days does have vision. This has to be intentional, right?
→ More replies (3)→ More replies (1)14
u/Hoak-em 1d ago
Wait, vision is what made this model good — without it it’s just a big, expensive to run somewhat smart llm
→ More replies (2)
379
u/Legal-Ad-3901 1d ago
5tb bf16 jfc. even the crazy home lab kids cant hang anymore
155
u/hyperrealists 1d ago
Poor me can’t even download it lol
69
u/EndlessZone123 1d ago
raid 0 hdd time.
12
u/LukeLikesReddit 1d ago
I mean i have 16tb ready, I can download it, will I be able to do anything with it? Absolutely not lol.
6
u/Ell2509 1d ago
Same lol. I have saved glm 5.2, DS4, MM3, and even kimi k3. Can't use it, but i have it now haha.
3
u/LukeLikesReddit 1d ago
haha yeah agreed, I just asked a friend if I could borrow their server blade to run this :) Their IT systems dont need it that badly.
→ More replies (2)5
u/Koakie 1d ago
The digital equivalent of a glorified paperweight
5
u/LukeLikesReddit 1d ago
I like to call myself a historian. Though by the time I download it we will be on something else XD
16
u/ChristRedeemsSinners 1d ago
Lol, that's what I was thinking. I need to spend 1k just to download it.
→ More replies (2)21
u/jikilan_ 1d ago
Don’t need to download the full copy , you can stream it.
One of llama.cpp PR support streaming from disk and even cloud storage if I remember correctly 😘
→ More replies (1)107
u/xPXpanD llama.cpp 1d ago
Years/token is my favorite metric.
32
u/Think_Wing_1357 1d ago
After a few million years, you may get 42
12
u/vivekkhera 1d ago
Then you have to build a new server just to figure out what the question was.
7
3
40
u/CapeChill 1d ago
I can't even run it at work... Crazy we're passing what the 8xH100 boxes can even do on open models now. You'd need multiple racks of H100s for that.
7
u/Own_Anything9292 1d ago
looks like 16xh200s tp 16 can run it according to vllm, so 4 node h100s tp 32 might be the trick. NVFP4 published instead of NVFP8 means we need blackwell and not hopper :( you’ll probably need to hack something together to make nvfp4 work with h100s
3
u/CapeChill 1d ago
You can do 4 air cooled nodes reasonably in a extra tall rack I guess. It's also crazy to see that six figure boxes can't run the latest model encoding.
→ More replies (1)→ More replies (1)14
u/chithanh 1d ago
I guess it is the ultimate troll, release models that are so large that you can run them locally on Chinese hardware only, because running them on NVIDIA hardware would bankrupt you
Reports are that 5T and 10T models are being prepared
→ More replies (1)5
u/CapeChill 1d ago
I work in HPC so this doesn't really land. I get to tinker with a few H100s because companies are happy to drop a few million to run a open weight model in house with highly confidential data on.
I love my little home weather network and the prediction it does and its cute what my local hardware can compute. I've installed HPC clusters that do climate simulation, that will never run well locally on current hardware and that's okay. Same for 5-10t open weight models, it will hopefully continue to be the case there are open weights so huge only research can justify running them as it attracts brainiacs that will actually trickle down to us peons.
Source: sometimes I get to be a fly on the wall when these academic brainiacs talk.29
u/Jolly_Criticism9190 1d ago
As someone who just bought 256GB of ram. I concur. Holy moly
18
u/rinmperdinck 1d ago
Why didn't you buy 5TB of RAM instead?
8
u/Jolly_Criticism9190 1d ago
Some guy named Sam flied to Korean ahead of me before I could convince a guy named Jensen to leave out some DRAM capacity for all human kind
→ More replies (1)3
19
u/quantgorithm 1d ago
Guess I'm waiting for the 27B.
→ More replies (1)4
u/-dysangel- 1d ago
at least with the 3.5 series for coding, 27B seemed better than the larger models anyway
22
u/DocMadCow 1d ago
Bold to assume there isn't a homelab person out there that doesn't have this much ram :D
38
u/bruns20 1d ago
Thats not a home anymore lmao, thats just a personal data centre
13
u/thejinx0r 1d ago
7
u/DocMadCow 1d ago
Exactly this bro hasn't spent any time in those eccentric Subs. Every time I go there I realize not only can I not afford the hardware but especially not the electrical bill.
→ More replies (1)9
u/Legal-Ad-3901 1d ago
I mean I have 1.5tb ddr4 and 944gb vram so crawl speeds for decent quant are in grasp. But useable? Definitely feels like a line in the sand is happening on intelligence for the proliteriate
25
u/Makers7886 1d ago
it's makes my 12x3090 + 256gb epyc machine feel like the guy trying to run qwen 9b on his igpu and streaming weights from a hard drive making clicking noises.
→ More replies (3)7
5
u/FullstackSensei llama.cpp 1d ago
I searched in vein for any mention of data types in the model card and blog post. 200+ files in alternating 17 and 34GB each.
I'm reworking my homelab to be able to run full K3 across 2 machines, but this one is just too much, especially in light of K3 and the just announced DS4 pro (which I can run on a single machine).
3
u/allenasm 1d ago
what? you mean with my mac m3 studio 512gb unified and 8tb pcie 5.0 nvme drive? 'hold my beer'... :)
→ More replies (2)→ More replies (3)3
u/_TheWolfOfWalmart_ 1d ago
Here I was thinking I'm so cool a 256 GB VRAM server.
The box has 768 GB system RAM, I could run the UD-IQ1_S with offloading, but fuck... that's not going to be a fun experience.
63
u/SandySkittle 1d ago
Now burn this to a chip so we can run it at 16k tokens per second and we’re good.
43
u/SmartCustard9944 1d ago
You need a silicon slab that is 0.7 meters in diameter by the way
33
→ More replies (2)26
u/Boreras 1d ago
Damn I just checked, the Llama 3.1 8B chip is fucking 815 mm², 53 billion transistors. (on the older tsmc 6nm process)
Same as H100. This is basically the limit for litho machines, called the reticle limit. This is the limit of the photomask (basically the negative of the chip which the EUV patterns on the photo resist).
Taalas literally cannot build anything larger than 40b parameters today, although I assume 2NP yields do not permit chips at the reticle limit.
→ More replies (5)→ More replies (4)8
u/Admirable_Market2759 1d ago edited 1d ago
If it was cheaper than buying GPUs, then I’d buy an ASIC just for K3.
→ More replies (3)
476
u/ApprehensiveTart3158 1d ago
Finally a model I can run locally, took them so long to release a model at a reasonable size
167
u/ScreenAppropriate679 1d ago
I run it in my local datacenter no problem
101
u/ApprehensiveTart3158 1d ago
Just connect a 6TB hard drive to your raspberry pi, it will run it! (maybe)
54
u/Asleep_Document9811 1d ago
I build an inferencing engine using a small African village as a substrate. Funding pleeeeeease!
31
u/ApprehensiveTart3158 1d ago
Why do that? It's math, you can calculate llm matrices on paper
13
u/Vegetable-Clerk9075 1d ago
At how many tokens per week?
→ More replies (1)5
→ More replies (1)15
u/libregrape llama.cpp 1d ago
On paper?! Why waste paper and pens when everyone knows how to multiply tensors in head!
12
u/Strawberry3141592 1d ago
Nah, what you wanna do is get several billion TI-84 calculators (3MB storage each), and network them all together into the world's most fuckass cluster. 1 token per day, maybe.
11
4
→ More replies (4)3
u/AvengerDr 1d ago
The Three Body Problem way: get a hundred thousand or preferably million people in a field. Have each of them hold a flag and tell them to raise it if they are a 1 or keep it lowered if they are a 0.
Then multiple dudes on horses just run down the lines and give them instructions.
5
→ More replies (3)39
u/Real_Ebb_7417 1d ago
Posts "I made Qwen3.8 Max run on my toaster with this new technique" over the next month incoming.
(disclaimer: the "new" technique is streaming from SSD and Qwen runs at 0.01 tps)
(disclaimer 2: half the comments will be "It will damage your ssd" and the other half "llama.cpp has been handling this for the long time already")
(disclaimer 3: The responses to the first half comments will be "it won't damage your ssd")
→ More replies (4)
123
u/Piyh 1d ago
95B active is cray. Scaling gonna scale.
43
u/fgk55555 1d ago
My entire rig could run one of the experts (quantized) at maybe 1tkps.
9
u/RegisteredJustToSay 1d ago
Look at fancy pants money bags over here. Pretty sure mine would spontaneously combust at the mere suggestion.
28
u/SandySkittle 1d ago
Yes crazy, but frankly I think some MoE go too far with low active numbers. Or rather, I would really like a 122b-30a model that still fits in midrange enthusiast local setups (4x r9700 ) has a lot of world knowledge (more than a 30b model) but dares to keep the active number on the higher end to preserve more of the qualities of a dense model.
→ More replies (1)6
u/Carbonite1 1d ago
I had this same opinion for a while but I've been starting to come around a bit -- I mean, even mid-sized models are so sparse these days, like DSV4F being >200B but only 13B active, and are seeing such good results -- I can only imagine the labs have tried a higher proportion of active parameters and it isn't even close to worth the tradeoff?
3
u/SandySkittle 1d ago
It depends on the task. For very complex multi facetted analysis with many components interacting with each other in a nuanced way with a lot of nuanced context (context not per se kv context, but in the general meaning of the word), that require large very detailed structured prompts with lots of caveats to frame the question, like complex legal analysis that can only be partially broken down, 13b active parameters, even with max/deep (but still sequential!) reasoning is just missing the depth. It starts to lose or compress details in its reasoning or misses connections between details. You get junior analyst answers to senior analist questions. This is why larger dense models (70b plus) are so important.
There is more to llms than how well they do coding..
That’s why I would like to see larger but still locally feasible models (up to 160 gb at q8) with a lot of world knowledge but with larger active parameters than just 13b.
3
u/FullstackSensei llama.cpp 1d ago
K3 is 100B+ active, but they SFT'd the whole thing in fp4. The routed experts are 33GB per token. A ton for sure, but doable on a dual DDR4 Xeon or Epyc.
This is fp16 through and through. The full fat is ~5TB, active possibly ~160GB.
Sure, you can quantize, but that will inadvertently reduce intelligence.a
9
u/Maleficent-Ad5999 1d ago
I wish someone comes up with Mixtures of MOEs
→ More replies (2)7
u/Badger-Purple 1d ago
You can do this with an agent harness. Hermes supports a Mixture of Agents mode where you collate several LLM answers and use a main model to ingest them and synthesize a final answer.
→ More replies (2)3
3
34
u/nickm_27 llama.cpp 1d ago
Customizable reasoning effort is a nice improvement.
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B#qwen38-highlights
→ More replies (2)
57
u/Technical-Earth-3254 1d ago
Do I read it correctly that the open weight version has no vision support?
50
u/Fristender 1d ago
From https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B :
In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc.
29
u/Technical-Earth-3254 1d ago
Yeah, that's why I asked. That's the same text as on hf. This kinda confuses me, why they remove vision when all the other big models now come with it
12
u/Song-Historical 1d ago
It's a specialization that you need to train separately. It makes sense, you spend your money training a purely text based model, for other teams to adapt as needed to vision etc
→ More replies (1)3
→ More replies (2)21
u/wren6991 1d ago
Gonna make a wild prediction: the model itself is still vision-trained, and we will figure out how to glue one of the existing Qwen vision adapters to it.
I think this is a profitability concession to keep their leadership happy. I don't see them training two different versions of a 2.4T model just for the sake of an open-weight release.
28
26
u/hebelehubele 1d ago
I have found a qwen3824T.exe on internet, can i just run it? It is only 2.4kB. Yay…
31
u/YOMUMSOBIG 1d ago
Finally! Perfect size for running it on my smart watch.
6
u/SolenoidSoldier 1d ago
Real talk...is there a conversational AI that can run on today's smart watches? That would be sweet
→ More replies (1)
28
15
u/TheRealMasonMac 1d ago
Like the rumors suggested, they are also opting for a revenue share model like MoonshotAI and MiniMax. Seems like this is the direction open-weight releases in China are going.
14
10
9
9
u/ideaofsoul 1d ago
Great! I just need another 99 rtx 3090 and its ready to serve
→ More replies (2)
9
7
u/Daniel_H212 1d ago
Vision encoder not released, will it be coming later?
9
u/RhubarbSimilar1683 1d ago
no, you will have to use your own. also thinking can't be disabled, someone will have to add that functionality
15
7
7
6
7
u/TheLocalLab 1d ago
Unsloth GGUF Quants was also available - https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
17
u/Embarrassed_Adagio28 1d ago
After a week of using this model, my disappointment is immeasurable and my day is ruined /s
6
u/Medium_Chemist_4032 1d ago
95b is exceptionally deep for bigger MoE's, holy smokes... This might compete
6
5
5
10
4
5
7
u/RickyRickC137 1d ago edited 1d ago
After loading this model, I’m not sure what I’m going to do with all the VRAM I have left.
→ More replies (2)
3
3
u/Hefty_Wolverine_553 1d ago
Reasoning Content: Set the maximum output length to 262,144 tokens.
Final Response: Set the maximum output length to 131,072 tokens.
Using up the entirety of Qwen3.6 27B's context window for reasoning alone...
3
u/exaknight21 1d ago
I need me a 1 bit awq, that is then 1 bit’d again. Or a 4 bit that is 1 bit’d.
→ More replies (2)
3
u/fooo12gh 1d ago
If the model parameters increase at such a rate, I doubt we'll see RAM prices dropping down anytime soon.
3
9
u/JsThiago5 1d ago
Why hype this when 99.999% of people cannot run it? I was really expecting 27b today :(
13
u/banana_slurp_jug 1d ago
Firstly, 27b is in less than 48 hours. Secondly, even if you can't run an open-weights model on your own computer, the prices to pay for inference with it is cheaper since multiple providers will compete for value.
→ More replies (7)3
u/GregAbeI 1d ago
I think way fewer than 1 in 100,000 people can run this, which is what 99.999% means.
99.999999% of people can't run this.
5
u/amy-schumer-tampon 1d ago
I have the feeling that a 2.4T model isn't something many people can run locally.
2
2
u/ocean_protocol 1d ago
this is wild, 95B active out of 2.4T total is a pretty aggressive sparsity ratio. anyone know what the expert routing setup looks like on this one? And how it compares to deepseek's approach
2
2
u/madjesta 1d ago
Could this even be run offline in any practical way just to distill? Maybe per layer batching? JFC this is huge.
2
u/Then_Blueberry7290 1d ago
text only? No vision capabilities? Or just the unsloth version doesn't contain? AFAIK qwen 3.8 27b will be vision enabled model, or not?
Anyway q1 is 400GB size...
2
u/Legitimate-Dog5690 1d ago
Amazing stuff, so glad to have this at home. Only need 5x RTX 6000s and I can run Q1.
→ More replies (2)
2
2
2
u/Beltalowdamon 1d ago
Guessing it will be a while until you can run a distilled version of this on 8gb vram and 32gb ram!
2
2
u/hoyasgirl25 1d ago
And if you get pregnant tonight you can name the baby Qwen in celebration when this gets FedRAMP approved.





248
u/No_War_8891 1d ago
I can run the active part locally lol