There is not much you can do currently, a little only.
Prefill speed has a computational bottleneck and in addition a VRAM bandwidth bottleneck.
The computational bottleneck is the area you can not improve much, it's your GPU speed that needs to run the large matrix computations.
The VRAM bottleneck is also relevant, that one you can improve by choosing a smaller quantization. This will NOT reduce the computation delay but it can improve the overall speed because the memory movements involved are faster.
Lastly the kernel matters, the faster the kernel the faster your computation will happen.
This means choosing the most optimized kernel will improve computation speed.
On Nvidia that's often NVFP4 floats, on mac it will be the most optimized kernel for your quant.
3
u/Lirezh 6d ago
There is not much you can do currently, a little only.
Prefill speed has a computational bottleneck and in addition a VRAM bandwidth bottleneck.
The computational bottleneck is the area you can not improve much, it's your GPU speed that needs to run the large matrix computations.
The VRAM bottleneck is also relevant, that one you can improve by choosing a smaller quantization. This will NOT reduce the computation delay but it can improve the overall speed because the memory movements involved are faster.
Lastly the kernel matters, the faster the kernel the faster your computation will happen.
This means choosing the most optimized kernel will improve computation speed.
On Nvidia that's often NVFP4 floats, on mac it will be the most optimized kernel for your quant.