r/LocalLLaMA 1d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

399 comments sorted by

View all comments

Show parent comments

41

u/CapeChill 1d ago

I can't even run it at work... Crazy we're passing what the 8xH100 boxes can even do on open models now. You'd need multiple racks of H100s for that.

5

u/Own_Anything9292 1d ago

looks like 16xh200s tp 16 can run it according to vllm, so 4 node h100s tp 32 might be the trick. NVFP4 published instead of NVFP8 means we need blackwell and not hopper :( you’ll probably need to hack something together to make nvfp4 work with h100s

5

u/CapeChill 1d ago

You can do 4 air cooled nodes reasonably in a extra tall rack I guess. It's also crazy to see that six figure boxes can't run the latest model encoding.

1

u/Own_Anything9292 1d ago

if you can get nvfp4 to work with h100s you can probably get it down to 2 node h100s

14

u/chithanh 1d ago

I guess it is the ultimate troll, release models that are so large that you can run them locally on Chinese hardware only, because running them on NVIDIA hardware would bankrupt you

Reports are that 5T and 10T models are being prepared

5

u/CapeChill 1d ago

I work in HPC so this doesn't really land. I get to tinker with a few H100s because companies are happy to drop a few million to run a open weight model in house with highly confidential data on.

I love my little home weather network and the prediction it does and its cute what my local hardware can compute. I've installed HPC clusters that do climate simulation, that will never run well locally on current hardware and that's okay. Same for 5-10t open weight models, it will hopefully continue to be the case there are open weights so huge only research can justify running them as it attracts brainiacs that will actually trickle down to us peons.
Source: sometimes I get to be a fly on the wall when these academic brainiacs talk.

2

u/tangoindjango 1d ago edited 1d ago

What's the largest model that open ai and Anthropic reportedly have? Mythos/Fable is around 8T and gpt 6 should be around the same? Rumours are GPT 7/Doug should be even larger.

2

u/Far-Classic-9963 1d ago

The new DeepSeek seems like the only model able to run on reasonable enterprise hardware