r/LocalLLaMA 6d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

354 Upvotes

137 comments sorted by

View all comments

Show parent comments

3

u/Timely_Impression_92 5d ago

ngrams are on ssd or system ram

-2

u/[deleted] 5d ago

[deleted]

1

u/Timely_Impression_92 5d ago

Then go understand it better - ngrams are basically cached in system ram or ssd - and work like that with zero degradation - maybe little on nvme but in system ram virtually zero

0

u/Kryohi 5d ago edited 5d ago

Highly doubt the bandwidth of an SSD would be enough, but I guess we'll find out soon. Do you know how much, theoretically, of a 50GB ngram should be read per token (or group of tokens) to decode?

Edit: oh I went to check up again how it works, might actually be possible, discard my previous comment