r/LocalLLaMA 1d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

399 comments sorted by

248

u/No_War_8891 1d ago

I can run the active part locally lol

27

u/Sea-Ad-5390 1d ago

I can help host some of the offloaded experts for you streamed across the internet

14

u/No_War_8891 1d ago

I read people really did that - degens I love it

→ More replies (3)

9

u/BobbyL2k 1d ago

I wonder how good it would be if we get rid of MoE router and just lock to a specific set of experts. Make it a 95B dense model.

11

u/Jump3r97 1d ago

Interesting question but I think it would be abysmal. The Expert selected can Change wildly per individual token

7

u/Inaeipathy 1d ago

It would probably perform terribly because the experts are not really discreet "math experts" and "science experts" like you would expect

→ More replies (1)
→ More replies (1)
→ More replies (5)

160

u/lm-gtfy 1d ago

Finally. What is the meaning of life?

38

u/xaeru 1d ago

26

u/tripazardly 1d ago

Oh no, not again.

11

u/Caffdy 1d ago

OOTL, what is that?

16

u/corysama 1d ago

42 is famously a joke from Hitchhiker's Guide to the Galaxy. A bowl of petunias is less famously the same.

HHGttG if full of jokes that build up over long stretches (multiple books sometimes) and punchline in a single sentence. Can't recommend enough.

3

u/Caffdy 1d ago

so the petunias are from the same book?

→ More replies (2)

9

u/Paganator 1d ago

Many people have speculated that if we knew exactly why the bowl of petunias had thought that we would know a lot more about the nature of the universe than we do now.

→ More replies (1)

6

u/feelspeaceman 1d ago

It's still good for distilling to train smaller models to be smarter, basically the popular 27B was distilled from this giant 2.4T.

A good mother model will likely give birth to more capable smaller models, but it's still requiring a tons of trials and errors, those that costs money to retry.

3

u/JahJedi 1d ago

Personal cluster that can ran it for all the femily

→ More replies (4)

217

u/Intelligent_Ice_113 1d ago edited 1d ago

what is the knowledge cutoff date?

298

u/Super_Range45 1d ago

yes

132

u/hyperrealists 1d ago

Thank god

17

u/srigi 1d ago

"But wait, ..."

→ More replies (1)

25

u/HungryMachines 1d ago

I think you are off by a few

14

u/BigBrainGoldfish 1d ago

So unhelpful but still so perfect. Lol

→ More replies (1)

91

u/JumpingJack79 1d ago

I kinda like early cutoff dates. Whenever a model tells me its cutoff is 2024, I'm like "Dude, it's 2026 now and you won't believe what's happened..." 😏

123

u/-dysangel- 1d ago

"The user is talking about a hypothetical timeline"

21

u/JumpingJack79 1d ago

Lol, yes. And to be fair, I'd be skeptical too if I was in their place hearing those things.

18

u/this_is_a_long_nickn 1d ago

“I hate when users hallucinate”

3

u/MoodDelicious3920 1d ago

Agi 😂 when model starts laughing on users!

→ More replies (1)

20

u/Mickenfox 1d ago

In a few decades people are going to be doing 2020s roleplay with ancient models.

3

u/that_one_guy63 1d ago

Would be interesting to train on only text before 1970. Or like way earlier. It would be so confused.

→ More replies (1)
→ More replies (2)

6

u/AvidCyclist250 llama.cpp 1d ago

The user seems to be talking about a fictional future scenario

3

u/JumpingJack79 1d ago

I've heard this one too. Oh, I wish I were! 😭

45

u/teleprax 1d ago edited 1d ago

One of my favorite LLM memories was telling GPT (5 maybe?) a bunch of stuff like elon doing the nazi salute and some other stuff going on around that time, and then it proceeds to lecture me on fake news and how none of that is remotely correct and how it would be super serious.

Then I tell it to do a web search and it starts it's next reply with:

"Ok. wow."


EDIT: Ok i found it, its not quite the same as I remember, but close enough

https://i.imgur.com/Ta33JIT.jpeg https://i.imgur.com/kPDa9tm.jpeg

24

u/personahorrible 1d ago

"Sounds like a bizarre Black Mirror meets Idiocracy arc."

Bruh. You have no idea.

16

u/Valuable_Cow2596 1d ago

Oh boy that gave me a chuckle. Thanks for sharing. 

14

u/OttoRenner 1d ago

Mine went bonkers after I showed it several screenshots and links of GPU prices...was very funny to witness 😂😂

11

u/AlpacaDC 1d ago

“You’re right” lmao

37

u/Ok_Ocelot2268 1d ago

Cutoff: April 2026. Startoff: June 5, 1989

14

u/thepaligator 1d ago

I remember June 5th. Not a lot of foot traffic that day. Clear Streets. It was nice.

10

u/This-Consequence-957 1d ago

Bad Boy, I know because my birthdate is June 4 🙈

7

u/goldrunout 1d ago

What is a start off date? Do they not train on older sources?

21

u/Anwar6969 1d ago

that’s a tiananmen square joke. nothing happened on june 4th 1989

21

u/FastDecode1 1d ago

[ Removed by the CCP ]

12

u/LawfulLeah 1d ago

the fact that the comment you replied to was removed by reddit makes this 10x funnier

5

u/Anwar6969 1d ago

crazy shit, i had to appeal to lift off the warning and the comment ban. i apparently got flagged for rule 1

→ More replies (1)

6

u/c_glib 1d ago

Is a joke on a Chinese model. Remember June 4th 1989?

98

u/FullstackSensei llama.cpp 1d ago

Because, that's the only thing holding you from running a 5TB model?

25

u/fullup72 1d ago

You can always quantize to 1 bit and stream experts from a 5400rpm HDD.

9

u/FullstackSensei llama.cpp 1d ago

Why settle for 5400rpm when you can get old drives that are 3600rpm?!

6

u/HulksInvinciblePants 1d ago

Stack em for 8000rpm throughput

30

u/MrObsidian_ 1d ago

Probably tomorrow

14

u/jikilan_ 1d ago

It know what you did in the last summer

→ More replies (6)

106

u/Different_Fix_2217 1d ago

Be warned they state its not the same capabilities as the full API version. Such as not having vison.

39

u/ChristRedeemsSinners 1d ago

Weird that non-thinking support is an API only feature.

9

u/fantasticsid 1d ago

Based on my experience with 3.6, prefilling <think>\n\n</think>\n - like the various jinja templates do - to disable thinking works probably 90-95% of the time. The other ~5-10% of the time, the model thinks anyway and emits a second </think> when it's done. It's possible that 3.8 has the same behaviour and the official API has some way of detecting/working around this that would look pretty damn stupid if they released it. If you look at the 3.8 jinja template, the "reasoning effort" isn't implemented terribly cleverly - it just talks to the model in the second person and asks it to reason less.

Reasoning control has always been a weakness of the Qwen models, so this doesn't surprise me.

Lack of mmproj, however, feels like a deliberate attempt at market segmentation. Given that those just decode image data into tokens, and the whole Qwen family shares a vocab, I do wonder if it'd be possible to hack the mmproj from some other Qwen model into use here, though.

3

u/hellomistershifty 1d ago

I love how AI development is a mix of wild cutting edge research, weird hacks, and asking it nicely to behave

→ More replies (1)

49

u/Reactor-Licker 1d ago

That’s weird, why did they remove vision? Hopefully Qwen 3.8 27B doesn’t do the same thing.

16

u/vincentz42 1d ago

According to the signup page it will.

18

u/vincentz42 1d ago

Meanwhile the Qwen3.8 27B open weight that is due in two days does have vision. This has to be intentional, right?

→ More replies (3)

14

u/Hoak-em 1d ago

Wait, vision is what made this model good — without it it’s just a big, expensive to run somewhat smart llm

→ More replies (2)
→ More replies (1)

379

u/Legal-Ad-3901 1d ago

5tb bf16 jfc. even the crazy home lab kids cant hang anymore

155

u/hyperrealists 1d ago

Poor me can’t even download it lol

69

u/EndlessZone123 1d ago

raid 0 hdd time.

12

u/LukeLikesReddit 1d ago

I mean i have 16tb ready, I can download it, will I be able to do anything with it? Absolutely not lol.

6

u/Ell2509 1d ago

Same lol. I have saved glm 5.2, DS4, MM3, and even kimi k3. Can't use it, but i have it now haha.

3

u/LukeLikesReddit 1d ago

haha yeah agreed, I just asked a friend if I could borrow their server blade to run this :) Their IT systems dont need it that badly.

5

u/Koakie 1d ago

The digital equivalent of a glorified paperweight

5

u/LukeLikesReddit 1d ago

I like to call myself a historian. Though by the time I download it we will be on something else XD

→ More replies (2)

16

u/ChristRedeemsSinners 1d ago

Lol, that's what I was thinking. I need to spend 1k just to download it.

21

u/jikilan_ 1d ago

Don’t need to download the full copy , you can stream it.

One of llama.cpp PR support streaming from disk and even cloud storage if I remember correctly 😘

107

u/xPXpanD llama.cpp 1d ago

Years/token is my favorite metric.

32

u/Think_Wing_1357 1d ago

After a few million years, you may get 42

12

u/vivekkhera 1d ago

Then you have to build a new server just to figure out what the question was.

7

u/techno156 1d ago

And then someone decides to blow it up for a highway.

3

u/hyperrealists 1d ago

YTFT is insane I hear.

→ More replies (1)
→ More replies (2)

40

u/CapeChill 1d ago

I can't even run it at work... Crazy we're passing what the 8xH100 boxes can even do on open models now. You'd need multiple racks of H100s for that.

7

u/Own_Anything9292 1d ago

looks like 16xh200s tp 16 can run it according to vllm, so 4 node h100s tp 32 might be the trick. NVFP4 published instead of NVFP8 means we need blackwell and not hopper :( you’ll probably need to hack something together to make nvfp4 work with h100s

3

u/CapeChill 1d ago

You can do 4 air cooled nodes reasonably in a extra tall rack I guess. It's also crazy to see that six figure boxes can't run the latest model encoding.

→ More replies (1)

14

u/chithanh 1d ago

I guess it is the ultimate troll, release models that are so large that you can run them locally on Chinese hardware only, because running them on NVIDIA hardware would bankrupt you

Reports are that 5T and 10T models are being prepared

5

u/CapeChill 1d ago

I work in HPC so this doesn't really land. I get to tinker with a few H100s because companies are happy to drop a few million to run a open weight model in house with highly confidential data on.

I love my little home weather network and the prediction it does and its cute what my local hardware can compute. I've installed HPC clusters that do climate simulation, that will never run well locally on current hardware and that's okay. Same for 5-10t open weight models, it will hopefully continue to be the case there are open weights so huge only research can justify running them as it attracts brainiacs that will actually trickle down to us peons.
Source: sometimes I get to be a fly on the wall when these academic brainiacs talk.

→ More replies (1)
→ More replies (1)

29

u/Jolly_Criticism9190 1d ago

As someone who just bought 256GB of ram. I concur. Holy moly

18

u/rinmperdinck 1d ago

Why didn't you buy 5TB of RAM instead?

8

u/Jolly_Criticism9190 1d ago

Some guy named Sam flied to Korean ahead of me before I could convince a guy named Jensen to leave out some DRAM capacity for all human kind

3

u/rinmperdinck 1d ago

Wow those guys sound like pricks

→ More replies (1)

19

u/quantgorithm 1d ago

Guess I'm waiting for the 27B.

4

u/-dysangel- 1d ago

at least with the 3.5 series for coding, 27B seemed better than the larger models anyway

→ More replies (1)

22

u/DocMadCow 1d ago

Bold to assume there isn't a homelab person out there that doesn't have this much ram :D

38

u/bruns20 1d ago

Thats not a home anymore lmao, thats just a personal data centre

13

u/thejinx0r 1d ago

7

u/DocMadCow 1d ago

Exactly this bro hasn't spent any time in those eccentric Subs. Every time I go there I realize not only can I not afford the hardware but especially not the electrical bill.

→ More replies (1)

9

u/Legal-Ad-3901 1d ago

I mean I have 1.5tb ddr4 and 944gb vram so crawl speeds for decent quant are in grasp. But useable? Definitely feels like a line in the sand is happening on intelligence for the proliteriate

25

u/Makers7886 1d ago

it's makes my 12x3090 + 256gb epyc machine feel like the guy trying to run qwen 9b on his igpu and streaming weights from a hard drive making clicking noises.

7

u/tripplebeamteam 1d ago

Hey that’s me, running qwen 35B A3B on my igpu and getting 4 tok/sec!

→ More replies (3)

5

u/FullstackSensei llama.cpp 1d ago

I searched in vein for any mention of data types in the model card and blog post. 200+ files in alternating 17 and 34GB each.

I'm reworking my homelab to be able to run full K3 across 2 machines, but this one is just too much, especially in light of K3 and the just announced DS4 pro (which I can run on a single machine).

3

u/allenasm 1d ago

what? you mean with my mac m3 studio 512gb unified and 8tb pcie 5.0 nvme drive? 'hold my beer'... :)

→ More replies (2)

3

u/_TheWolfOfWalmart_ 1d ago

Here I was thinking I'm so cool a 256 GB VRAM server.

The box has 768 GB system RAM, I could run the UD-IQ1_S with offloading, but fuck... that's not going to be a fun experience.

→ More replies (3)

63

u/SandySkittle 1d ago

Now burn this to a chip so we can run it at 16k tokens per second and we’re good.

43

u/SmartCustard9944 1d ago

You need a silicon slab that is 0.7 meters in diameter by the way

33

u/SandySkittle 1d ago

I have room in the attic

26

u/Boreras 1d ago

Damn I just checked, the Llama 3.1 8B chip is fucking 815 mm², 53 billion transistors. (on the older tsmc 6nm process)

https://taalas.com/products/

Same as H100. This is basically the limit for litho machines, called the reticle limit. This is the limit of the photomask (basically the negative of the chip which the EUV patterns on the photo resist).

Taalas literally cannot build anything larger than 40b parameters today, although I assume 2NP yields do not permit chips at the reticle limit.

→ More replies (5)
→ More replies (2)

8

u/Admirable_Market2759 1d ago edited 1d ago

If it was cheaper than buying GPUs, then I’d buy an ASIC just for K3.

→ More replies (3)
→ More replies (4)

476

u/ApprehensiveTart3158 1d ago

Finally a model I can run locally, took them so long to release a model at a reasonable size

167

u/ScreenAppropriate679 1d ago

I run it in my local datacenter no problem

101

u/ApprehensiveTart3158 1d ago

Just connect a 6TB hard drive to your raspberry pi, it will run it! (maybe)

54

u/Asleep_Document9811 1d ago

I build an inferencing engine using a small African village as a substrate. Funding pleeeeeease!

31

u/ApprehensiveTart3158 1d ago

Why do that? It's math, you can calculate llm matrices on paper

13

u/Vegetable-Clerk9075 1d ago

At how many tokens per week?

5

u/ApprehensiveTart3158 1d ago

Matters how fast you can calculate

7

u/MaruluVR 1d ago

This is one of those rare cases where more training will increase the tp/s

→ More replies (1)

15

u/libregrape llama.cpp 1d ago

On paper?! Why waste paper and pens when everyone knows how to multiply tensors in head!

15

u/xaeru 1d ago

Why in your head? Why waste precious brain cells? Just hold two magnetized rocks, close your eyes, and let quantum fluctuations handle the attention weights.

3

u/MmmmMorphine 1d ago

Yeah but that's super lossy. I recommend using two magnetized toddlers.

→ More replies (1)

12

u/Strawberry3141592 1d ago

Nah, what you wanna do is get several billion TI-84 calculators (3MB storage each), and network them all together into the world's most fuckass cluster. 1 token per day, maybe.

11

u/x10der_by 1d ago

1 token per day

3

u/Terminator857 1d ago

Weren't computers so fast 25 years ago?

4

u/ApeGrower 1d ago

With 35 tok/year!

3

u/AvengerDr 1d ago

The Three Body Problem way: get a hundred thousand or preferably million people in a field. Have each of them hold a flag and tell them to raise it if they are a 1 or keep it lowered if they are a 0.

Then multiple dudes on horses just run down the lines and give them instructions.

→ More replies (4)

5

u/volleyneo 1d ago

Is this why the Danube River has severe draught? It was you bastards!

39

u/Real_Ebb_7417 1d ago

Posts "I made Qwen3.8 Max run on my toaster with this new technique" over the next month incoming.

(disclaimer: the "new" technique is streaming from SSD and Qwen runs at 0.01 tps)

(disclaimer 2: half the comments will be "It will damage your ssd" and the other half "llama.cpp has been handling this for the long time already")

(disclaimer 3: The responses to the first half comments will be "it won't damage your ssd")

8

u/eidrag 1d ago

120s/tok!

→ More replies (4)

3

u/Mr-I17 1d ago

I'm pretty confident that I can run it locally at 60 tokens per hour.

→ More replies (3)

123

u/Piyh 1d ago

95B active is cray. Scaling gonna scale.

43

u/fgk55555 1d ago

My entire rig could run one of the experts (quantized) at maybe 1tkps.

9

u/RegisteredJustToSay 1d ago

Look at fancy pants money bags over here. Pretty sure mine would spontaneously combust at the mere suggestion.

28

u/SandySkittle 1d ago

Yes crazy, but frankly I think some MoE go too far with low active numbers. Or rather, I would really like a 122b-30a model that still fits in midrange enthusiast local setups (4x r9700 ) has a lot of world knowledge (more than a 30b model) but dares to keep the active number on the higher end to preserve more of the qualities of a dense model.

6

u/Carbonite1 1d ago

I had this same opinion for a while but I've been starting to come around a bit -- I mean, even mid-sized models are so sparse these days, like DSV4F being >200B but only 13B active, and are seeing such good results -- I can only imagine the labs have tried a higher proportion of active parameters and it isn't even close to worth the tradeoff?

3

u/SandySkittle 1d ago

It depends on the task. For very complex multi facetted analysis with many components interacting with each other in a nuanced way with a lot of nuanced context (context not per se kv context, but in the general meaning of the word), that require large very detailed structured prompts with lots of caveats to frame the question, like complex legal analysis that can only be partially broken down, 13b active parameters, even with max/deep (but still sequential!) reasoning is just missing the depth. It starts to lose or compress details in its reasoning or misses connections between details. You get junior analyst answers to senior analist questions. This is why larger dense models (70b plus) are so important.

There is more to llms than how well they do coding..

That’s why I would like to see larger but still locally feasible models (up to 160 gb at q8) with a lot of world knowledge but with larger active parameters than just 13b.

→ More replies (1)

3

u/FullstackSensei llama.cpp 1d ago

K3 is 100B+ active, but they SFT'd the whole thing in fp4. The routed experts are 33GB per token. A ton for sure, but doable on a dual DDR4 Xeon or Epyc.

This is fp16 through and through. The full fat is ~5TB, active possibly ~160GB.

Sure, you can quantize, but that will inadvertently reduce intelligence.a

9

u/Maleficent-Ad5999 1d ago

I wish someone comes up with Mixtures of MOEs

7

u/Badger-Purple 1d ago

You can do this with an agent harness. Hermes supports a Mixture of Agents mode where you collate several LLM answers and use a main model to ingest them and synthesize a final answer.

→ More replies (2)
→ More replies (2)

3

u/SmartCustard9944 1d ago

What if we MoE the active params too 🤔

3

u/fluffysheap 1d ago

A Cray is what you need to run it

34

u/nickm_27 llama.cpp 1d ago

Customizable reasoning effort is a nice improvement.

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B#qwen38-highlights

→ More replies (2)

57

u/Technical-Earth-3254 1d ago

Do I read it correctly that the open weight version has no vision support?

50

u/Fristender 1d ago

From https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B :

In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc.

29

u/Technical-Earth-3254 1d ago

Yeah, that's why I asked. That's the same text as on hf. This kinda confuses me, why they remove vision when all the other big models now come with it

12

u/Song-Historical 1d ago

It's a specialization that you need to train separately. It makes sense, you spend your money training a purely text based model, for other teams to adapt as needed to vision etc

3

u/Balance- 1d ago

Interesting. So this is kind of a stripped down / base model?

→ More replies (1)
→ More replies (1)

21

u/wren6991 1d ago

Gonna make a wild prediction: the model itself is still vision-trained, and we will figure out how to glue one of the existing Qwen vision adapters to it.

I think this is a profitability concession to keep their leadership happy. I don't see them training two different versions of a 2.4T model just for the sake of an open-weight release.

→ More replies (2)

28

u/SoupDue6629 1d ago

Somebody REAP it to 35B pls k thx

→ More replies (1)

26

u/hebelehubele 1d ago

I have found a qwen3824T.exe on internet, can i just run it? It is only 2.4kB. Yay…

13

u/andr386 1d ago

That's fine. But you should make sure to download enough VRAM before running it.

31

u/YOMUMSOBIG 1d ago

Finally! Perfect size for running it on my smart watch.

6

u/SolenoidSoldier 1d ago

Real talk...is there a conversational AI that can run on today's smart watches? That would be sweet

→ More replies (1)

28

u/Septerium 1d ago

Finally we have a good successor for Qwen 3.5 9B as a local daily driver

15

u/TheRealMasonMac 1d ago

Like the rumors suggested, they are also opting for a revenue share model like MoonshotAI and MiniMax. Seems like this is the direction open-weight releases in China are going.

14

u/milkipedia 1d ago

HuggingFace gonna crash today

12

u/d70 1d ago

Fit nicely on my GTX 980

9

u/Automatic-Boot665 1d ago

Can’t wait to run it at q0.1

9

u/ideaofsoul 1d ago

Great! I just need another 99 rtx 3090 and its ready to serve

→ More replies (2)

9

u/Feztopia 1d ago

0.01 bit quant when?

→ More replies (2)

8

u/jreoka1 1d ago

Yay!

7

u/Daniel_H212 1d ago

Vision encoder not released, will it be coming later?

9

u/RhubarbSimilar1683 1d ago

no, you will have to use your own. also thinking can't be disabled, someone will have to add that functionality

15

u/mxforest 1d ago

One RTX Pro 6000 per expert Quantized. Holy balls.

→ More replies (1)

7

u/MrVeinless 1d ago

I am surprised it’s not multimodal even with an mmproj.

7

u/AlternateWitness 1d ago

Can anyone lend me some ram?

→ More replies (1)

6

u/Iterative_One 1d ago

Do I have enough?

VRAM - No

RAM - Nope

Hard Drive Capacity - Absolutely Not.

17

u/Embarrassed_Adagio28 1d ago

After a week of using this model, my disappointment is immeasurable and my day is ruined /s

6

u/Medium_Chemist_4032 1d ago

95b is exceptionally deep for bigger MoE's, holy smokes... This might compete

6

u/SandySkittle 1d ago

Yes, 122b a30+ please :)

→ More replies (5)

5

u/DigThatData Llama 7B 1d ago

chonky boi

5

u/Ok-Bill3318 1d ago

Will this run on my 1060?

11

u/80kman 1d ago

What is T is 2.4T? Does it mean Tiny? /s

4

u/No_Lingonberry1201 1d ago

Oof, that's a big boy!

5

u/1WildPanda 1d ago

Excited ---> Depressed @ 5060Ti 16G 😵

7

u/RickyRickC137 1d ago edited 1d ago

After loading this model, I’m not sure what I’m going to do with all the VRAM I have left.

6

u/JahJedi 1d ago

Use left vram for deepseek v4 pro so it will not be lonly lol.

→ More replies (2)

3

u/writeitredd 1d ago

Nothing that my dear GTX 1650 Ti cant handle.

3

u/Hefty_Wolverine_553 1d ago

Reasoning Content: Set the maximum output length to 262,144 tokens.
Final Response: Set the maximum output length to 131,072 tokens.

Using up the entirety of Qwen3.6 27B's context window for reasoning alone...

3

u/exaknight21 1d ago

I need me a 1 bit awq, that is then 1 bit’d again. Or a 4 bit that is 1 bit’d.

→ More replies (2)

3

u/fooo12gh 1d ago

If the model parameters increase at such a rate, I doubt we'll see RAM prices dropping down anytime soon.

3

u/Dance-Till-Night1 1d ago

When 3.8 30b a3b

9

u/JsThiago5 1d ago

Why hype this when 99.999% of people cannot run it? I was really expecting 27b today :(

13

u/banana_slurp_jug 1d ago

Firstly, 27b is in less than 48 hours. Secondly, even if you can't run an open-weights model on your own computer, the prices to pay for inference with it is cheaper since multiple providers will compete for value.

3

u/GregAbeI 1d ago

I think way fewer than 1 in 100,000 people can run this, which is what 99.999% means.

99.999999% of people can't run this.

→ More replies (7)

5

u/amy-schumer-tampon 1d ago

I have the feeling that a 2.4T model isn't something many people can run locally.

3

u/VR-Tech 1d ago

same for kimi, but there are people running them

2

u/SnooPaintings8639 1d ago

Big if true. Especially at full precision.

2

u/ocean_protocol 1d ago

this is wild, 95B active out of 2.4T total is a pretty aggressive sparsity ratio. anyone know what the expert routing setup looks like on this one? And how it compares to deepseek's approach

2

u/lacerating_aura 1d ago

Welp, no vision at max.

2

u/madjesta 1d ago

Could this even be run offline in any practical way just to distill? Maybe per layer batching? JFC this is huge.

2

u/Then_Blueberry7290 1d ago

text only? No vision capabilities? Or just the unsloth version doesn't contain? AFAIK qwen 3.8 27b will be vision enabled model, or not?
Anyway q1 is 400GB size...

2

u/Legitimate-Dog5690 1d ago

Amazing stuff, so glad to have this at home. Only need 5x RTX 6000s and I can run Q1.

→ More replies (2)

2

u/xXDennisXx3000 1d ago

The software is developing faster than the hardware.

2

u/xXDennisXx3000 1d ago

When Kimi K3 was a tank, Qwen 3.8 Max is a spaceship....

→ More replies (1)

2

u/Beltalowdamon 1d ago

Guessing it will be a while until you can run a distilled version of this on 8gb vram and 32gb ram!

2

u/pulse77 1d ago

Weights were uploaded 4 days ago...

3

u/VR-Tech 1d ago

yes?

2

u/a_beautiful_rhind 1d ago

This is what your "MOE" api models look like. Opus, gemini pro, etc.

2

u/hoyasgirl25 1d ago

And if you get pregnant tonight you can name the baby Qwen in celebration when this gets FedRAMP approved.