CoolFace
Modelpublic

LeaderboardModel1/Qwen3.6-35B-A3B-uncensored-heretic-AutoRound-W4A16-RTN

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes96downloads
Model Card

Qwen3.6-35B-A3B-uncensored-heretic-AutoRound-W4A16-RTN

Model Details

This model is a int4 weight-only quantization with group_size 128 and symmetric quantization of llmfan46/Qwen3.6-35B-A3B-uncensored-heretic generated by AutoRound. Please follow the license of the original model.

Quantization Details

AttributeValue
Base Modelllmfan46/Qwen3.6-35B-A3B-uncensored-heretic
Quantization ToolAutoRound
Quantization SchemeW4A16
Quantized Size19777 MB

Evaluation Results

TaskAccuracy
hellaswag0.6316
mmlu0.8230
mmluabstractalgebra0.5800
mmlu_anatomy0.8370
mmlu_astronomy0.9342
mmlubusinessethics0.8300
mmluclinicalknowledge0.8943
mmlucollegebiology0.9444
mmlucollegechemistry0.6100
mmlucollegecomputer_science0.7000
mmlucollegemathematics0.6800
mmlucollegemedicine0.8439
mmlucollegephysics0.6667
mmlucomputersecurity0.8700
mmluconceptualphysics0.9277
mmlu_econometrics0.8070
mmluelectricalengineering0.8483
mmluelementarymathematics0.8069
mmluformallogic0.6905
mmluglobalfacts0.5000
mmluhighschool_biology0.9516
mmluhighschool_chemistry0.7980
mmluhighschoolcomputerscience0.8800
mmluhighschooleuropeanhistory0.8606
mmluhighschool_geography0.9343
mmluhighschoolgovernmentand_politics0.9845
mmluhighschool_macroeconomics0.8846
mmluhighschool_mathematics0.5963
mmluhighschool_microeconomics0.9580
mmluhighschool_physics0.8212
mmluhighschool_psychology0.9578
mmluhighschool_statistics0.8148
mmluhighschoolushistory0.9069
mmluhighschoolworldhistory0.9114
mmluhumanaging0.8251
mmluhumansexuality0.8702
mmlu_humanities0.7583
mmluinternationallaw0.9256
mmlu_jurisprudence0.8704
mmlulogicalfallacies0.9264
mmlumachinelearning0.7946
mmlu_management0.9029
mmlu_marketing0.9444
mmlumedicalgenetics0.9100
mmlu_miscellaneous0.9387
mmlumoraldisputes0.8295
mmlumoralscenarios0.6078
mmlu_nutrition0.8889
mmlu_other0.8574
mmlu_philosophy0.8650
mmlu_prehistory0.9074
mmluprofessionalaccounting0.7234
mmluprofessionallaw0.6721
mmluprofessionalmedicine0.9338
mmluprofessionalpsychology0.8775
mmlupublicrelations0.7455
mmlusecuritystudies0.8245
mmlusocialsciences0.9035
mmlu_sociology0.9353
mmlu_stem0.8069
mmluusforeign_policy0.9500
mmlu_virology0.5663
mmluworldreligions0.9123
piqa0.8166

How to Use

HF Usage

Step 1: Install [AutoRound](https://github.com/intel/auto-round)

bash
pip install auto-round

Step 2: Load and run the quantized model

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen3.6-35B-A3B-uncensored-heretic-AutoRound-W4A16-RTN"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")

# prepare the model input
prompt = "Write a quick sort algorithm."
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(**model_inputs, max_new_tokens=512)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]) :].tolist()

content = tokenizer.decode(output_ids, skip_special_tokens=True)
print("content:", content)

VLLM Usage

bash
vllm serve Qwen3.6-35B-A3B-uncensored-heretic-AutoRound-W4A16-RTN \
    --trust-remote-code \
    --dtype bfloat16 \
    --tensor_parallel_size 1

If you encounter any issues, feel free to open an issue on the AutoRound GitHub repo or provide feedback on the Low-Bit Open LLM Leaderboard.

Ethical Considerations and Limitations

The model can produce factually incorrect output, and should not be relied on to produce factually accurate information. Because of the limitations of the pretrained model and the finetuning datasets, it is possible that this model could generate lewd, biased or otherwise offensive outputs. Therefore, before deploying any applications of the model, developers should perform safety testing.

Caveats and Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Here are a couple of useful links to learn more about Intel's AI software:

Disclaimer

The license on this model does not constitute legal advice. We are not responsible for the actions of third parties who use this model. Please consult an attorney before using this model for commercial purposes.

Cite

@article{cheng2023optimize,
  title={Optimize weight rounding via signed gradient descent for the quantization of llms},
  author={Cheng, Wenhua and Zhang, Weiwei and Shen, Haihao and Cai, Yiyang and He, Xin and Lv, Kaokao and Liu, Yi},
  journal={arXiv preprint arXiv:2309.05516},
  year={2023}
}

arxiv github


This model is part of the [Intel Low-Bit Open LLM Leaderboard](https://huggingface.co/spaces/Intel/low_bit_open_llm_leaderboard) initiative.