CoolFace
Modelpublic

amanwalksdownthestreet/Qwen3-235B-A22B-Instruct-2507-exl3

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes33downloads
Model Card

ExLlamaV3 quantizations of Qwen3-235B-A22B-Instruct-2507 with tensor-level (L3) optimization and boosted attention layers (5 bit). Maximum effort applied towards the goal of achieving the best possible quantizations at the expense of time and compute.

Using this measurement.json file and the base quants provided, additional highly-optimized quantizations can be made in seconds at any reasonable bpw by anyone. All work done with ExLlamaV3 v0.0.18.

Optimized

SizebpwTargetPPLvs 2.0vs 3.0
2.26bpw-h6-opt66 GB2.2672GB @ 64k4.54−18%+12%
2.44bpw-h6-opt69 GB2.4496GB @ 256k4.33−22%+7%
2.89bpw-h6-opt81 GB2.8996GB @ 128k4.02−28%−0.5%
3.06bpw-h6-opt86 GB3.0696GB @ 64k3.94−29%−2.5%
3.93bpw-h6-opt109 GB3.93————
4.68bpw-h6-opt129 GB4.68————

Base

SizebpwPPL
2.0bpw-h657 GB2.05.57
3.0bpw-h684 GB3.04.04
4.0bpw-h6112 GB4.0—
5.0bpw-h6139 GB5.0—
6.0bpw-h6166 GB6.0—

Methodology

Optimized quants use exl3's measure.py → optimize.py → recompile.py pipeline. Attention layers replaced with 5bpw precision post-optimization.

Perplexity by optimization stage: | Target | Pre-attn bpw | PPL | +5bpw attn | Final bpw | Size | |:--:|:--:|:--:|:--:|:--:|:--:| | 72GB @ 64k | 2.20 | 4.72 | 4.54 | 2.26 | 66 GB | | 96GB @ 256k | 2.41 | 4.48 | 4.33 | 2.44 | 69 GB | | 96GB @ 128k | 2.86 | 4.20 | 4.02 | 2.89 | 81 GB | | 96GB @ 64k | 3.00 | 4.14 | 3.94 | 3.06 | 86 GB |

Cost vs gain (relative to 2.0 base @ 57 GB, PPL 5.57): | Target | Final bpw | Size | +GB | PPL | Δ PPL | PPL/GB | |:--:|:--:|:--:|:--:|:--:|:--:|:--:| | 72GB @ 64k | 2.26 | 66 GB | +9 | 4.54 | −1.03 | −0.11 | | 96GB @ 256k | 2.44 | 69 GB | +12 | 4.33 | −1.24 | −0.10 | | 96GB @ 128k | 2.89 | 81 GB | +24 | 4.02 | −1.55 | −0.06 | | 96GB @ 64k | 3.06 | 86 GB | +29 | 3.94 | −1.63 | −0.06 |

Attention boost adds 0.03-0.06 bpw (2-3 GB) for meaningful PPL gains. Diminishing returns between 4bpw and 5bpw attention at higher base bitrates.

Lower PPL is better. vs columns show % change from base 2.0 (PPL 5.57) and base 3.0 (PPL 4.04). Measured at 2k context.