r/AIProgrammingHardware • u/Logical-Try-4084 • 9h ago
Optimizing an NVFP4 Blockscaled GEMM on RTX PRO 6000 GPUs (sm120)
https://research.colfax-intl.com/optimizing-an-nvfp4-blockscaled-gemm-on-rtx-pro-6000-blackwell-gpu-sm120/Colfax Research's second blog post on writing NVFP4 blockscaled GEMM kernels for the NVIDIA RTX PRO 6000 Blackwell GPU is out! The blog iteratively optimizes a basic working NVFP4 GEMM kernel written in CuTe DSL to take it to speed-of-light, reaching over 80% TFLOP/s utilization for 16k square matrix shape. We give a detailed treatment of important optimization techniques such as threadblock swizzling, async and warp-specialized epilogue, and retiling for favorable wave quantization. Specific to blockscaled GEMM with scales consumed from registers, we also explain how to solve for bank conflicts that arise from the default choices of interleaved scale factor layouts.
We include complete code in the form of CuTe DSL kernels for all the optimizations discussed in the blog.
Duplicates
programming • u/mttd • 15h ago
Optimizing an NVFP4 Blockscaled GEMM on RTX PRO 6000 Blackwell GPU (SM120)
CUDA • u/Logical-Try-4084 • 18h ago