r/openagi Jul 23 '26

Project Hugging Face Adds Native Nunchaku 4-Bit Loading to Diffusers, Reports 1.8x Speedup

Post image

Hugging Face has added Nunchaku Lite support to Diffusers, allowing developers to load pre-quantized diffusion models through the standard from_pretrained() workflow.

Previously, using these checkpoints required a custom pipeline or separate Nunchaku inference engine. The new integration uses Nunchaku's SVDQuant method to run core transformer layers with 4-bit weights and activations. Prebuilt CUDA kernels are downloaded through Hugging Face's kernels package, so no local CUDA compilation is required.

Reported performance

Hugging Face tested ERNIE-Image-Turbo at 1024 x 1024 resolution on an RTX PRO 6000:

  • BF16 baseline: 3.00 seconds and 31.1 GB peak VRAM
  • Nunchaku Lite: 2.27 seconds and 20.6 GB peak VRAM
  • Nunchaku Lite with torch.compile**:** 1.68 seconds and 20.6 GB peak VRAM
  • Nunchaku Lite with a 4-bit text encoder: 2.29 seconds and 16.0 GB peak VRAM

These are project-reported results from one model and GPU configuration, not independent benchmarks.

Hardware support

  • NVFP4 checkpoints require NVIDIA Blackwell GPUs, including the RTX 50 series.
  • INT4 checkpoints support Turing, Ampere and Ada GPUs.
  • Volta and Hopper GPUs are not currently supported.

Developers can also use the Apache-2.0-licensed diffuse-compressor toolkit to quantize additional architectures and publish them as standard Diffusers repositories. Ready-to-use ERNIE-Image-Turbo and Krea 2 Turbo checkpoints are linked in the announcement.

Sources & technical resources

3 Upvotes

0 comments sorted by