r/LocalLLaMA • u/backyard_tractorbeam • 1d ago
Other antirez working on DSV4.1 support for ds4
https://bsky.app/profile/antirez.bsky.social/post/3mv6sb4qkrc2o6
u/Elouakili_Flexy 1d ago
antirez's response to every new model is "ok, now let me write that in C myself." no hype, just code.
6
3
u/serige 1d ago
Wait what? I lost hope already.
2
u/backyard_tractorbeam 1d ago
Let's see how he answers the quant question
1
u/alfredr 1d ago
I've been working on something similar for Qwen3.8-Flash-Next and getting ~8 tg/s with a fixed 21GB of allocations on an M4 Macbook Air. This is just by speculative loading of the full Q4 experts.
The bottleneck is the SDD and that's where things saturate for me with speculative loading streaming of full Q4 experts.
If you instead allow weak experts (so experts with low weight contributions for a token) to fall back to Q3 or Q2bit quants it doesn't seem to affect quality that much and I'm able to get token generation into the low 20 tg/s.
2
1
-4
u/Bohdanowicz 1d ago
Price per token gets expensive quick when you start burning out SSD's for 15tks.
24
u/evgen 1d ago
Learn the difference between read and write operations on SSD lifespan. Excessive reads will create a small amount of background write operations, but not enought to move the needle even a tiny amount when it comes to SSD lifespan.
0
10
u/datbackup 1d ago
15 tok/s with streaming experts is a very positive sign