CoolFace
Modelpublic

amanwalksdownthestreet/Qwen3-Coder-30B-A3B-Instruct-exl3

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes37downloads
Model Card

ExLlamaV3 quantizations of Qwen3-Coder-30B-A3B-Instruct with tensor-level (L3) optimization and boosted attention layers (5-6 bit). Maximum effort applied towards the goal of achieving the best possible quantizations at the expense of time and compute.

Using this measurement.json file and the base quants provided, additional highly-optimized quantizations can be made in seconds at any reasonable bpw by anyone. All work done with ExLlamaV3 v0.0.18.

Optimized

VRAM-targeted quants using exl3's measure.py → optimize.py → recompile.py pipeline.

SizebpwTarget
3.34bpw-h5-opt13 GB3.3416GB @ 128k
4.90bpw-h6-opt18 GB4.9024GB @ 262k

The 4.90bpw quant hit the optimization ceiling - requesting higher bpw (5.30, 6.95) produced identical 4.83bpw pre-boost output, indicating no further beneficial tensor swaps available.

Base

Sizebpw
2.0bpw-h68 GB2.0
3.0bpw-h612 GB3.0
4.0bpw-h615 GB4.0
5.0bpw-h619 GB5.0
6.0bpw-h622 GB6.0
7.0bpw-h626 GB7.0