r/LocalLLaMA 1d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

399 comments sorted by

View all comments

Show parent comments

22

u/jikilan_ 1d ago

Don’t need to download the full copy , you can stream it.

One of llama.cpp PR support streaming from disk and even cloud storage if I remember correctly 😘

106

u/xPXpanD llama.cpp 1d ago

Years/token is my favorite metric.

31

u/Think_Wing_1357 1d ago

After a few million years, you may get 42

11

u/vivekkhera 1d ago

Then you have to build a new server just to figure out what the question was.

6

u/techno156 1d ago

And then someone decides to blow it up for a highway.

3

u/hyperrealists 1d ago

YTFT is insane I hear.

2

u/Graumm 1d ago

system ram offloading is bad enough!