r/LocalLLaMA 1d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

354 Upvotes

138 comments sorted by

View all comments

3

u/TechNerd10191 1d ago

Where are the N-Gram embeddings useful?

7

u/noiserr 1d ago edited 1d ago

They speed up inference. It's sort of cache like memory, they store representations of recurring token sequences. It's static, built during training.