CoolFace
Modelpublic

Vishva007/Qwen3.8-4B-Distill-W4A16-AutoRound-LLM-Compressor

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
0likes640downloads
Model Card

Qwen3.8-4B-Distill (W4A16 Quantized via AutoRound)

This repository contains a W4A16 (4-bit weights, 16-bit activations) quantized version of empero-ai/Qwen3.8-4B-Distill, quantized using Intel's AutoRound algorithm.


⚡ Quantization Details

Calibrated and quantized with fine-grained group sizes and high iteration depth to preserve reasoning traces (<think> blocks) and multimodal capabilities:

  • —Algorithm: Intel AutoRound
  • —Precision / Scheme: W4A16 (4-bit weights, 16-bit activations)
  • —Group Size: 32 (fine-grained reconstruction fidelity)
  • —Symmetric (`sym`): True
  • —Calibration Samples (`nsamples`): 512
  • —Sequence Length (`seqlen`): 4096
  • —Tuning Iterations (`iters`): 1000 (Production-grade accuracy)
  • —Vision Tower (`quant_nontext_module`): False (Kept in BF16 to preserve visual reasoning and OCR precision)
  • —Special Modules (`layer_config`): Multi-Token Prediction (mtp, mtp.fc) kept in native bfloat16

📦 Available Formats

Depending on your inference engine, choose the appropriate repository:


🚀 Usage & Quickstart

1. High-Throughput Serving via vLLM

bash
# Using the LLM-Compressor / Compressed-Tensors build
vllm serve Vishva007/Qwen3.8-4B-Distill-W4A16-AutoRound-LLM-Compressor \
    --dtype bfloat16 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.90

📊 VRAM & Performance Benefits

  • —Original Model (BF16): ~8–10 GB VRAM required for full context inference
  • —Quantized Model (W4A16 Group 32): ~2.5–3.5 GB VRAM (runs comfortably on 4GB/6GB consumer GPUs, laptops, and edge devices)
  • —Throughput: Lowers memory bandwidth pressure, accelerating token generation speeds during extended chain-of-thought (<think>) reasoning.

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
TemplateCUDA VersionDocker ImageTemplate IDDeploy
PyTorch 2.14 (CUDA 12.6)12.6vishva123/cuda-12.6-pytorch-2.14-runpodd7lxsa4w9m![Deploy to RunPod](https://runpod.io/console/deploy?template=d7lxsa4w9m&ref=iabrlp7z)
PyTorch 2.14 (CUDA 13.0)13.0vishva123/cuda-13.0-pytorch-2.14-runpodyk0y6j6rpg![Deploy to RunPod](https://runpod.io/console/deploy?template=yk0y6j6rpg&ref=iabrlp7z)
PyTorch 2.14 (CUDA 13.2)13.2vishva123/cuda-13.2-pytorch-2.14-runpodgsp4gwx0nw![Deploy to RunPod](https://runpod.io/console/deploy?template=gsp4gwx0nw&ref=iabrlp7z)
PyTorch 2.13
TemplateCUDA VersionDocker ImageTemplate IDDeploy
PyTorch 2.13 (CUDA 12.6)12.6vishva123/cuda-12.6-pytorch-2.13-runpodgmlupxnxfk![Deploy to RunPod](https://runpod.io/console/deploy?template=gmlupxnxfk&ref=iabrlp7z)
PyTorch 2.13 (CUDA 13.0)13.0vishva123/cuda-13.0-pytorch-2.13-runpody3j8xvk4f4![Deploy to RunPod](https://runpod.io/console/deploy?template=y3j8xvk4f4&ref=iabrlp7z)
PyTorch 2.13 (CUDA 13.2)13.2vishva123/cuda-13.2-pytorch-2.13-runpodvigpissn5w![Deploy to RunPod](https://runpod.io/console/deploy?template=vigpissn5w&ref=iabrlp7z)
PyTorch 2.12
TemplateCUDA VersionDocker ImageTemplate IDDeploy
PyTorch 2.12 (CUDA 12.6)12.6vishva123/cuda-12.6-pytorch-2.12-runpodctmz86zmf0![Deploy to RunPod](https://runpod.io/console/deploy?template=ctmz86zmf0&ref=iabrlp7z)
PyTorch 2.12 (CUDA 13.0)13.0vishva123/cuda-13.0-pytorch-2.12-runpodqjko5yiwzi![Deploy to RunPod](https://runpod.io/console/deploy?template=qjko5yiwzi&ref=iabrlp7z)
PyTorch 2.12 (CUDA 13.2)13.2vishva123/cuda-13.2-pytorch-2.12-runpodifg6xmye0f![Deploy to RunPod](https://runpod.io/console/deploy?template=ifg6xmye0f&ref=iabrlp7z)

📚 Acknowledgments