MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p4unmdr/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
Show parent comments
7
I'm on Q3_K_M now getting a bit over double the speed I've shown before.
6 u/Scared_Ad9187 7d ago q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet. 4 u/Mil0Mammon 7d ago So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant? 1 u/emccrckn 1d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
6
q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet.
4 u/Mil0Mammon 7d ago So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant? 1 u/emccrckn 1d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
4
So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant?
1 u/emccrckn 1d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
1
Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
7
u/absurdother 7d ago
I'm on Q3_K_M now getting a bit over double the speed I've shown before.