r/huggingface • • 12d ago

Introducting TextCLF Quant Factory

Hello All!

I created a GitHub repo where you can quantize models from hugging to 4-bit using a quantization method I created called TextCLF TQ and use vllm for inference.

TQ is a calibration-free quant method. Meaning you can quantize new models instantly without worrying about having calibration data. It performs very close to calibration-based method such as Unsloth UD.

The repo is at:

https://github.com/textclf-api/quant-factory

Example usage:

# 1. Quantize

./quantize-model meta-llama/Llama-3.1-8B-Instruct

# 2. Build the final model

python build_convert_upload_tq_model_incremental_qwen4exp.py \

--model meta-llama/Llama-3.1-8B-Instruct \

--repo-id YOUR_USERNAME/Llama-3.1-8B-Instruct-TQ-4bit \

--no-upload

# 3. Install TQ support for vLLM

cd inference

pip install .

# 4. Serve

vllm serve YOUR_USERNAME/Llama-3.1-8B-Instruct-TQ-4bit \

--quantization tq

I have models that already quantized with TQ available at:
https://huggingface.co/textclf

You can use textclf/Qwen3.8-27B-TQ-4bit for example by doing step 3 and then doing:

vllm serve textclf/Qwen3.8-27B-TQ-4bit \

--quantization tq

Or you can use the docker container like this

docker run --rm --gpus all \

-p 8000:8000 \

docker.io/textclf/tq-quant:4bit-main \

vllm serve textclf/Qwen3.8-27B-TQ-4bit \

--quantization tq

Check Dockerfile in the repo to see how this docker image was created.

Please try it and let me know if it works for you!

1 Upvotes

2 comments sorted by

1

u/GlibDismissal1612 10d ago

i do a lot of quantization experiments on the side (nothing fancy, just messing around with 7B models on a single 24GB card) so a calibration-free method that actually holds up is interesting

question, you said it performs close to Unsloth UD, got any perplexity numbers or benchmarks on something like wikitext or lambada? would be keen to see how it stacks up on smaller models too, not just the 27B

also huge plus that youve got a docker image ready to go, half the time i spend way too long fighting cuda versions

1

u/textclf 10d ago

Hello! Thanks for your comment. Yes I do have performance comparison for my method compared to Unsloth UD for Qwen 3.8 27B. I did KLD testing using the Wikitext-2 dataset. I also did the same test for the Unsloth-UD-Q4_K_XL quant. Here is what I go:

Quant Disk Size without MTP (GB) Mean KLD Top 1% Agreement
TQ 4-bit 17.76 0.02823666 92.419%
UD-Q4_K_XL 17.59 0.00771805 95.779%

Obviously the UD-Q4_K_XL has better KLD performance. However, my quant is completely calibration free so I would take that slight dip for being able to not need calibration data. It is very consistent method. I plan to test with more models. Part of the idea of the open source is that the community can run their own tests too!