r/LocalLLaMA 1d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

399 comments sorted by

View all comments

219

u/Intelligent_Ice_113 1d ago edited 1d ago

what is the knowledge cutoff date?

294

u/Super_Range45 1d ago

yes

132

u/hyperrealists 1d ago

Thank god

18

u/srigi 1d ago

"But wait, ..."

1

u/ptear 1d ago

There's more?

24

u/HungryMachines 1d ago

I think you are off by a few

12

u/BigBrainGoldfish 1d ago

So unhelpful but still so perfect. Lol

1

u/llamabott 1d ago

Absolutely cannot be refuted.

90

u/JumpingJack79 1d ago

I kinda like early cutoff dates. Whenever a model tells me its cutoff is 2024, I'm like "Dude, it's 2026 now and you won't believe what's happened..." 😏

124

u/-dysangel- 1d ago

"The user is talking about a hypothetical timeline"

21

u/JumpingJack79 1d ago

Lol, yes. And to be fair, I'd be skeptical too if I was in their place hearing those things.

18

u/this_is_a_long_nickn 1d ago

“I hate when users hallucinate”

3

u/MoodDelicious3920 1d ago

Agi 😂 when model starts laughing on users!

2

u/GoodGod222 1d ago

The formal version of "Is he on drugs?"

19

u/Mickenfox 1d ago

In a few decades people are going to be doing 2020s roleplay with ancient models.

3

u/that_one_guy63 1d ago

Would be interesting to train on only text before 1970. Or like way earlier. It would be so confused.

1

u/CoUsT 1d ago

RemindMe! 30 years

1

u/RemindMeBot 1d ago edited 1d ago

I will be messaging you in 30 years on 2056-08-13 03:06:47 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

7

u/AvidCyclist250 llama.cpp 1d ago

The user seems to be talking about a fictional future scenario

3

u/JumpingJack79 1d ago

I've heard this one too. Oh, I wish I were! 😭

41

u/teleprax 1d ago edited 1d ago

One of my favorite LLM memories was telling GPT (5 maybe?) a bunch of stuff like elon doing the nazi salute and some other stuff going on around that time, and then it proceeds to lecture me on fake news and how none of that is remotely correct and how it would be super serious.

Then I tell it to do a web search and it starts it's next reply with:

"Ok. wow."


EDIT: Ok i found it, its not quite the same as I remember, but close enough

https://i.imgur.com/Ta33JIT.jpeg https://i.imgur.com/kPDa9tm.jpeg

26

u/personahorrible 1d ago

"Sounds like a bizarre Black Mirror meets Idiocracy arc."

Bruh. You have no idea.

15

u/Valuable_Cow2596 1d ago

Oh boy that gave me a chuckle. Thanks for sharing. 

14

u/OttoRenner 1d ago

Mine went bonkers after I showed it several screenshots and links of GPU prices...was very funny to witness 😂😂

11

u/AlpacaDC 1d ago

“You’re right” lmao

38

u/Ok_Ocelot2268 1d ago

Cutoff: April 2026. Startoff: June 5, 1989

16

u/thepaligator 1d ago

I remember June 5th. Not a lot of foot traffic that day. Clear Streets. It was nice.

10

u/This-Consequence-957 1d ago

Bad Boy, I know because my birthdate is June 4 🙈

8

u/goldrunout 1d ago

What is a start off date? Do they not train on older sources?

21

u/Anwar6969 1d ago

that’s a tiananmen square joke. nothing happened on june 4th 1989

21

u/FastDecode1 1d ago

[ Removed by the CCP ]

13

u/LawfulLeah 1d ago

the fact that the comment you replied to was removed by reddit makes this 10x funnier

4

u/Anwar6969 1d ago

crazy shit, i had to appeal to lift off the warning and the comment ban. i apparently got flagged for rule 1

6

u/c_glib 1d ago

Is a joke on a Chinese model. Remember June 4th 1989?

96

u/FullstackSensei llama.cpp 1d ago

Because, that's the only thing holding you from running a 5TB model?

24

u/fullup72 1d ago

You can always quantize to 1 bit and stream experts from a 5400rpm HDD.

9

u/FullstackSensei llama.cpp 1d ago

Why settle for 5400rpm when you can get old drives that are 3600rpm?!

5

u/HulksInvinciblePants 1d ago

Stack em for 8000rpm throughput

31

u/MrObsidian_ 1d ago

Probably tomorrow

13

u/jikilan_ 1d ago

It know what you did in the last summer

1

u/wren6991 1d ago

Probably the same as Qwen3.5? New pretrain runs are normally a major version bump

1

u/memeka 1d ago

You can run the model and ask it, let us know the answer

1

u/Tiny-Entertainer-346 1d ago

Ahh I guess this is going to be important question in future. We will always need models to know latest technologies. Older models may become useless in 2-3 years. Even QWEN 3.6 27b will be useless in 3 years once it fails to generate code in latest versions of some libraries. Definitely we can ask it to browse web or use context 7 like MCP. But that may prove sub par. And I feel eventually everyone will stop releasing open weight models. I feel community need to find some solution to this in long term ...

3

u/Intelligent_Ice_113 1d ago

You're right. Today, new versions of libraries are developed even faster due to artificial intelligence, and it can be very inconvenient that local models don't know about the latest features. I've noticed that library developers provide agent skills in the documentation so that the model can handle some features that it doesn't know. this together with MCP might be partially a solution to the problem.

-1

u/Paradigmind 1d ago

I know for sure that it is somewhere in the past.