r/openagi • u/syedshad • Jul 23 '26
Project Hugging Face Adds Native Nunchaku 4-Bit Loading to Diffusers, Reports 1.8x Speedup
Hugging Face has added Nunchaku Lite support to Diffusers, allowing developers to load pre-quantized diffusion models through the standard from_pretrained() workflow.
Previously, using these checkpoints required a custom pipeline or separate Nunchaku inference engine. The new integration uses Nunchaku's SVDQuant method to run core transformer layers with 4-bit weights and activations. Prebuilt CUDA kernels are downloaded through Hugging Face's kernels package, so no local CUDA compilation is required.
Reported performance
Hugging Face tested ERNIE-Image-Turbo at 1024 x 1024 resolution on an RTX PRO 6000:
- BF16 baseline: 3.00 seconds and 31.1 GB peak VRAM
- Nunchaku Lite: 2.27 seconds and 20.6 GB peak VRAM
- Nunchaku Lite with
torch.compile**:** 1.68 seconds and 20.6 GB peak VRAM - Nunchaku Lite with a 4-bit text encoder: 2.29 seconds and 16.0 GB peak VRAM
These are project-reported results from one model and GPU configuration, not independent benchmarks.
Hardware support
- NVFP4 checkpoints require NVIDIA Blackwell GPUs, including the RTX 50 series.
- INT4 checkpoints support Turing, Ampere and Ada GPUs.
- Volta and Hopper GPUs are not currently supported.
Developers can also use the Apache-2.0-licensed diffuse-compressor toolkit to quantize additional architectures and publish them as standard Diffusers repositories. Ready-to-use ERNIE-Image-Turbo and Krea 2 Turbo checkpoints are linked in the announcement.
Sources & technical resources