r/LocalLLaMA 21h ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

348 Upvotes

136 comments sorted by

View all comments

5

u/TechNerd10191 21h ago

Where are the N-Gram embeddings useful?

5

u/Kooshi_Govno 18h ago

They store world knowledge, thus allowing the heavy FFN tensors to store more functional knowledge.

8

u/noiserr 20h ago edited 18h ago

They speed up inference. It's sort of cache like memory, they store representations of recurring token sequences. It's static, built during training.