CoolFace
Apppublic

AsadIsmail/ternary-quant-demo

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
2likes
App README

ternary-quant — post-training quantization to {−1, 0, +1} for VLMs, seq2seq, and audio models, not just LLMs.

Three pre-quantized checkpoints are available:

ModelCompressionQuality
Qwen3-1.7B2.7×97.5% FP16 retain
Qwen2-VL-2B1.8×≥96% (text backbone)
Gemma 4 E2B1.8×text + vision quantized

Select a model, enter a prompt, and click Generate. The first call downloads and loads the checkpoint (may take ~30s); subsequent calls are fast.

Source: github.com/Asad-Ismail/ternary-quant