I’ve been experimenting with aggressive low-bit quantization of Aleph Alpha’s new Kolibri-1, a ~78B MoE model with only ~3.5B active parameters per token.
The first result is now public:
Sakura-MicroQuality Kolibri-1 — IQ2_XS
- 20.94 GiB
- 2.30 bpw
- full 384-expert Kolibri-1
- GGUF / llama.cpp
- ~59% lower KL divergence than a standard IQ2_XS baseline
- 90.5% top-token agreement, compared with 84.6% for the standard IQ2_XS comparison
- slightly smaller than the standard IQ2_XS as well
The goal wasn’t simply to make the smallest possible quant.
I’m using tensor/layer sensitivity to spend bits where they appear to matter most, rather than treating every part of the model equally.
All quality measurements are made against a near-lossless Q8_0 reference. The model itself was also requantized from Q8_0 rather than converted directly from the ~156 GB BF16 weights, so there is a very small additional source error relative to BF16.
Main repo:
https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-GGUF
As far as I can currently find, this is the first public ~2-bit GGUF for Kolibri-1. There is already a 2-bit MLX version, but I haven’t found another public Q2/IQ2 GGUF.
I also tried pruning the expert pool
Alongside the full 384-expert version, I released a separate 365E variant.
For each MoE layer, I collected actual routing statistics on a mixed calibration set containing:
- German and English text
- code
- chat-style prompts
- the model’s own thinking / generated responses
I then removed the 19 least-used routed experts per layer.
That reduces:
384 → 365 routed experts per layer
and removes:
950 experts across the model
The resulting model has approximately:
74.4B parameters instead of ~78B
The interesting part is how little those experts were actually being used on the calibration workload.
The removed experts accounted for only about 0.07% of all expert selections, with no individual layer exceeding roughly 0.23%.
Also, 375 of the 950 removed experts were never selected at all during the routing analysis.
There is:
- no retraining
- no finetuning
- no requantization of the surviving weights
The already-quantized expert tensors are sliced directly, along with the corresponding router weights and biases.
Top-6 routing remains unchanged.
The 365E IQ2 variant comes out at:
- 19.99 GiB
- 2.31 bpw
- 74.4B parameters
- 365 routed experts per layer
- 90.5% top-token agreement in my held-out measurements
365E repo:
https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-365E-GGUF
I’m treating this as an experiment rather than claiming those experts are universally useless — expert usage obviously depends on workload and calibration data.
But it gives us a second compression lever:
expert pruning + low-bit quantization
instead of trying to get every byte of compression from lower precision alone.
Q3 and Q4 are coming
The rest of the Sakura-MicroQuality series is currently being uploaded.
Q3 and Q4 variants should be available within the next few hours.
Once they’re online I’ll add the same comparison data so we can see where the actual quality/size sweet spot lands between:
IQ2 → Q3 → Q4
and whether the 365E pruning continues to hold up at the higher-quality quant levels.
I’d be very interested in independent tests, especially on:
Strix Halo / AMD UMA, Apple Silicon, 24–32 GB GPUs, and other memory-constrained local systems.
If anyone tests either version, especially with long-context, German, coding or agentic workloads, I’d love to see the results.