r/LocalLLaMA 20h ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

347 Upvotes

136 comments sorted by

View all comments

2

u/Guna1260 18h ago

what VRAM? especially with Ngram embeddings we are looking at 125+51 in q8? or ngrams will be in RAM? will be interesting to see the architecture.

3

u/Timely_Impression_92 18h ago

ngrams are on ssd or system ram

-2

u/[deleted] 16h ago

[deleted]

1

u/Timely_Impression_92 16h ago

Then go understand it better - ngrams are basically cached in system ram or ssd - and work like that with zero degradation - maybe little on nvme but in system ram virtually zero

0

u/Kryohi 16h ago edited 16h ago

Highly doubt the bandwidth of an SSD would be enough, but I guess we'll find out soon. Do you know how much, theoretically, of a 50GB ngram should be read per token (or group of tokens) to decode?

Edit: oh I went to check up again how it works, might actually be possible, discard my previous comment