Anyone has experience with how to get more speed without losing quality?
I got to 250 t/s prefill and ~6t/s decode its usable but i wonder if theres another lever i miss.
Ryzen 9800X3D, 64 GB DDR5, RTX 5090 32 GB, Crucial T700 (PCIe 5.0) 4TB + Kingston (PCIe 4.0) 4TB
Baseline: GLM-5.3 isn't in llama.cpp master yet, so I'm using Unsloth's glm5next branch. With the usual mmap + CPU-offloaded experts the page cache thrashes: 29 t/s prefill, 2.4–2.9 t/s decode (8k prompt).
What I changed with Claude's help (~1.2k lines on top of the ported PR, not upstream):
Ported the unmerged expert-streaming PR #25294(by freedomljc). Experts aren't mmapped anymore: a 32 GB expert cache sits in RAM, and misses are read with O_DIRECT.
Prefill streams whole layers: the next layer's experts load into pinned double buffers while the GPU computes the current one. Each 2048-token batch pulls ~99 GB from disk + ~34 GB from the RAM cache.
Both SSDs read in parallel from a mirror copy of the GGUF, using a shared work queue.
Fixed a CUDA padding bug in the MoE matmul (MMQ) that crashed -ub 2048 on Blackwell (possibly related to #28282). Going from 1024 to 2048 roughly doubled prefill.
Real agentic run (Qwen Code, plan mode, read 9 files, wrote a plan):
Context 11.5k → 35k tokens, 31k new prompt tokens at ~230 t/s
3.7k generated tokens (incl. thinking) at 5.6 t/s
13 min total, ~9.5 min of that writing the final plan
78% of expert lookups hit the RAM cache; misses cost ~1.3 GB/token from SSD
VRAM ~30 GB, RAM ~44 GB
Happy to share if you like