CoolFace
Modelpublic

JC1DA/Qwen3.8-27B-DavidAU-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-INT4-W4A16

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
13likes4.2kdownloads
Model Card

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (W4A16)

W4A16 GPTQ quantization of Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU using AutoRound with GPTQ group size 128.

Quantization Details

ParameterValue
AlgorithmAutoRound + GPTQ
Weight bits4
Activation bits16
Group size128
VersionAutoRound 0.12.3

Model size: ~19GB (vs ~54GB FP16) — 65% reduction

Benchmark Comparison: W4A16 vs FP16

Both versions benchmarked on identical hardware (NVIDIA A100 80GB) using vLLM 0.27.1.

BenchmarkFP16W4A16Delta
ARC-Easy (acc)86.91%87.25%+0.34%
ARC-Easy (acc_norm)85.69%86.49%+0.80%
HumanEval (pass@1)79.27%79.88%+0.61%
BBH (exact_match)89.94%90.34%+0.40%
MMLU (acc)86.89%86.69%-0.20%
WikiText2 PPL8.048.17+0.13

Key findings:

  • —W4A16 matches or exceeds FP16 on all classification/generation benchmarks
  • —Perplexity difference is negligible (0.13%) — well within measurement variance
  • —No perceivable quality loss at 4-bit weights with 16-bit activations

Usage

Works with any framework supporting GPTQ/AutoRound checkpoints:

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16/DavidAU_Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-w4g128"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    torch_dtype="auto"
)

Or via vLLM for high-throughput serving:

bash
vllm serve DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16/DavidAU_Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-w4g128 \
    --dtype auto --tensor-parallel-size 1

Model Card

See the FP16 original model card for full details on training methodology, stages, and capabilities.

This quantized version preserves all characteristics of the original:

  • —Heretic/uncensored output
  • —Strong reasoning and instruction following
  • —Reduced overthinking tokens
  • —Auto-variable thinking sizes