wcamon/Agents-A1-4B-Wringer-Q2.6
Wringer p3b_w2 — Agents-A1-4B at 2.655 bits/weight
Draft v0.2 (2026-09-12). HumanEval is reported as the mean ± sd over 4 sampling seeds (temperature 1.0); IFEval and GSM8K are single official runs.
A 2.655 bits-per-weight (body) quantization of InternScience/Agents-A1-4B (Qwen3.5 hybrid: 25 GatedDeltaNet linear-attention layers + 7 full-attention layers, 3.57 B quantized body weights), produced with Wringer: fixed-grid GPTQ codes → one round of low-rank "water" (r128 LoRA, KD to the bf16 parent on the model's own long-trajectory corpus) → closed-form "wring" that re-solves scales and codes so no adapter is shipped. Two fill-and-wring rounds were run.
Scores (official A1 evaluation framework, thinking on, 16k max tokens)
\*comp = mean of the three retention ratios against the bf16 anchor (HumanEval ratio uses the 4-seed means). The three Wringer rows are within sampling noise of each other (HumanEval single runs swing by up to 5.5 pp at temperature 1.0); the second wring round did not hurt and is what we publish. Body b/w counts the quantized linear layers only; embeddings, norms and lmhead stay bf16 (whole-LM figure 4.69 b/w with bf16 embeddings; 2.95 b/w if the tied embedding were stored at Q4K-class 4.5 b/w — not evaluated).
HumanEval with the 10 tasks flagged by our contamination check removed (see below), seed 20260806: 89.61 (research state), 94.16 (anchor).
What is in this repo
wringer_p3b_w2.safetensors— the container (1.10 GiB): packed codes + fp16 / int8 block scales, metadatawringer_meta.wringer_unpack.py— dependency-free materializer: container + bf16 source export → bf16 HF checkpoint.bf16/— the materialized bf16 checkpoint (fake-quant weights; loads withtransformers/ vLLM like the parent).reports/— packing verification (pack_p3b_w2.md), contamination check (contamination_e69.md), verdict JSON.
No custom kernel is provided. The container is a storage format; inference uses the materialized bf16 weights.
Recipe (per module)
Ledger: code 2.588 + scales 0.067 = 2.655 b/w. Codes solved by GPTQ per layer on 128 × 16k self-generated calibration rows; scales by a joint least-squares closed form with a prior (λ = 0.01).
Honest caveats
- Scale precision. The research solver stored scales in fp32 while the ledger charged fp16. Materializing at ledger precision moves 2.95 % of the bf16 weights by one ulp. We re-scored the container-materialized weights: HumanEval 85.37 ± 1.32 vs 86.74 ± 2.74 for the research state (difference 1.37 pp, SE 1.52 — not distinguishable), IFEval and GSM8K slightly higher. What you download is exactly what the first row measures. The solver now projects scales to fp16 at solve time.
- Sampling noise. HumanEval at temperature 1.0 swings by several points between seeds, even for the bf16 parent (91.46–94.51). Single-run HumanEval headlines from this project's earlier write-ups (e.g. 89.02) were lucky draws; all HumanEval numbers here are 4-seed means.
- Contamination. Our calibration corpus is model-generated. An 8-gram check against the three test sets found GSM8K and IFEval clean; 10 HumanEval tasks have ≥25-token overlaps with corpus solutions (list in
reports/). Scores are reported with and without them. - Three benchmarks only. IFEval / HumanEval / GSM8K in thinking mode. No MMLU-class or agentic suites yet.
- One model, one architecture. Generality beyond this hybrid Qwen3.5 model is not yet shown.
How it was made
Article: on the Hugging Face blog (link added after publication). Code, pre-registrations and evidence: https://github.com/wcAmon/wringer.
License
The parent model's license applies to the weights. Wringer code: Apache-2.0.
