CoolFace
Modelpublic

huawei-csl/Qwen3-4B-PreSINQ-GGUF

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
6likes56downloads
Model Card

<p align="center"> <img src="SINQGGUFHF.png" alt="Logo" style="max-width: 80%; height: auto;"> </p>

<p align="center">๐Ÿ™ <a href="https://github.com/huawei-csl/SINQ">Github</a>&nbsp;&nbsp; | &nbsp;&nbsp;๐Ÿ“„ <a href="http://arxiv.org/abs/2509.22944">Paper</a></p>

PreSINQ GGUF Quantized Qwen3-4B Model

This repository contains the official PreSINQ GGUF-quantized versions of the `Qwen3-4B` model. For a detailed explanation of PreSINQ strategy please refer to the the official SINQ repository. SINQ is a fast and high-quality quantization technique designed to significantly reduce Large Language Model size while preserving accuracy.

If you find this project useful, please consider giving a โญ to the official [SINQ](https://github.com/huawei-csl/SINQ) repository.


Model Details

  • โ€”Model Name: Qwen3-4B-PreSINQ-GGUF
  • โ€”Base Model: `Qwen/Qwen3-4B`
  • โ€”Task: Text Generation
  • โ€”Framework: PyTorch / Transformers
  • โ€”License: Apache-2.0
  • โ€”Quantized By: Huawei โ€“ Computing Systems Lab

How to Obtain the PreSINQ Model

The PreSINQ Qwen3-4B models are produced using the PreSINQ GGUF script available in the official SINQ repository.

The models provided here correspond to the best-performing configurations for each quantization type.

๐Ÿ“Š Best PreSINQ Quantization Results (Qwen3-4B)

Results below are measured on the WikiText-2 test set.

MethodBitsSize (GB)Perplexity โ†“
Baseline (FP16)FP167.5014.3128
Baseline + Q4KS4-bit2.2214.9756
PreSINQ + Q4_K_S4-bit2.2214.7121
Baseline + Q3KS3-bit1.7619.0347
PreSINQ + Q3_K_S3-bit1.7616.0734

However, you can generate good PreSINQ models (not the best one) faster by reducing the number of configurations explored during the PreSINQ script execution. The table below shows perplexity for different PreSINQ parameter configurations using Q3_K_S quantization. Evaluation is performed on a 5k-line subset of the **Pile validation dataset**.

Group SizeIterationsRepetitionsPerplexity
322111.2359
324111.1062
328110.7951
3216110.8192
3232110.7918
642111.2507
644111.0779
648110.9395
6416110.9450
1282111.2318
1284110.9561
1288110.8965
12816110.8904

๐Ÿš€ Usage

Usage Example

You can load and run the PreSINQ GGUF models using:

  • โ€”๐Ÿค— Transformers
  • โ€”llama.cpp
  • โ€”Any GGUF-compatible inference framework

๐Ÿงพ How to Cite This Work

If you find SINQ useful in your research or applications:

  • โ€”Please give a โญ to the official SINQ repository
  • โ€”Cite our <a href="http://arxiv.org/abs/2509.22944" target="_blank"><strong>paper</strong></a>:
bibtex
@misc{muller2025sinq,
      title={SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights}, 
      author={Lorenz K. Muller and Philippe Bich and Jiawei Zhuang and Ahmet Celik and Luca Benfenati and Lukas Cavigelli},
      year={2025},
      eprint={2509.22944},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={http://arxiv.org/abs/2509.22944}
}