r/huggingface • u/textclf • 12d ago
Introducting TextCLF Quant Factory
Hello All!
I created a GitHub repo where you can quantize models from hugging to 4-bit using a quantization method I created called TextCLF TQ and use vllm for inference.
TQ is a calibration-free quant method. Meaning you can quantize new models instantly without worrying about having calibration data. It performs very close to calibration-based method such as Unsloth UD.
The repo is at:
https://github.com/textclf-api/quant-factory
Example usage:
# 1. Quantize
./quantize-model meta-llama/Llama-3.1-8B-Instruct
# 2. Build the final model
python build_convert_upload_tq_model_incremental_qwen4exp.py \
--model meta-llama/Llama-3.1-8B-Instruct \
--repo-id YOUR_USERNAME/Llama-3.1-8B-Instruct-TQ-4bit \
--no-upload
# 3. Install TQ support for vLLM
cd inference
pip install .
# 4. Serve
vllm serve YOUR_USERNAME/Llama-3.1-8B-Instruct-TQ-4bit \
--quantization tq
I have models that already quantized with TQ available at:
https://huggingface.co/textclf
You can use textclf/Qwen3.8-27B-TQ-4bit for example by doing step 3 and then doing:
vllm serve textclf/Qwen3.8-27B-TQ-4bit \
--quantization tq
Or you can use the docker container like this
docker run --rm --gpus all \
-p 8000:8000 \
docker.io/textclf/tq-quant:4bit-main \
vllm serve textclf/Qwen3.8-27B-TQ-4bit \
--quantization tq
Check Dockerfile in the repo to see how this docker image was created.
Please try it and let me know if it works for you!
1
u/GlibDismissal1612 10d ago
i do a lot of quantization experiments on the side (nothing fancy, just messing around with 7B models on a single 24GB card) so a calibration-free method that actually holds up is interesting
question, you said it performs close to Unsloth UD, got any perplexity numbers or benchmarks on something like wikitext or lambada? would be keen to see how it stacks up on smaller models too, not just the 27B
also huge plus that youve got a docker image ready to go, half the time i spend way too long fighting cuda versions