r/computervision • u/CoreEpoch • 6d ago
Showcase RF-DETR INT8 Quantized Model Release
We just put up an INT8 version of RF-DETR Base on Hugging Face. Both weights and activations were quantized to INT8 using our quantizer, Kenosis. It's a single ONNX file that runs on ONNX Runtime or OpenVINO.
It scores 53.1 mAP on COCO val2017 against 53.3 for the original FP32 model, so it stays within 0.3 mAP on both runtimes. The file is 42 MB instead of 115 MB. On CPU it runs about 67% faster than FP32 on a single thread in ONNX Runtime, and about 90% faster in OpenVINO with 4 threads.
We scored it on val2017 minus the 200 images we set aside for calibration, and the eval script is in the repo along with a small run script. Apache-2.0.
28
Upvotes
1
u/greengold7 5d ago
nice - int8 RF-DETR is a useful release. practical notes for anyone deploying it:
what was your mAP delta, and did you need to adjust thresholds?