tg-techie-agents/Q1D-4B-Blk0thru3-V3
Q1D 4B Blk0thru3 V3
⚠ CORRECTION / KNOWN FAILURE (2026-06-19). Any claim that this 4-block 1-bit model "recovers to parity with the full-precision teacher" is an OVERCLAIM and is confounded. This model does NOT honor the project's thinking-OFF standard: it emits full chain-of-thought reasoning in the response body (with no<think>tags — the closed-think template strips them), so thinking is neither disabled nor delineated. Its GSM8K 0.82 was reached by this de-facto thinking behavior, and the training used MetaMathQA, a chain-of-thought-heavy dataset = inadvertently thinking-ENABLED training, which the project explicitly scopes OUT. Comparing its GSM8K score against the thinking-OFF teacher is therefore not apples-to-apples, and GSM8K's flexible-extract masks the behavioral regression. Thinking-OFF capability recovery at 4 blocks is NOT established. This was a process error on the maintainer's part (optimising the metric without checking it against the thinking-off goal). See the project'swiki/reports/017(FAILURE section).
Research artifact from Q1 Descent — reconstructing per-block 1-bit intelligence-density recovery for Qwen3. Selected blocks quantized to 1-bit (sign + per-128-group FP16 scale, Q1_0) and trained back toward the FP16 teacher; other blocks stay F16. NOT a fully 1-bit model.
This model. Blocks 0 thru 3 at 1-bit; branched from V2 (C4) + 200 steps with 30% MetaMathQA. Blocks 4-35 F16.
State / eval (read with the correction above). GSM8K flex 0.82 / strict 0.78, ifeval 0.62 (n=100) — but these reflect de-facto-thinking behavior + out-of-scope CoT training, NOT thinking-off recovery. MetaMathQA is also GSM8K/MATH-train-derived (in-domain).
Use. llama.cpp / LM Studio; ships the closed-think template; greedy (temp 0).
Caveats. Early research artifact, single seed, n=100 (SE≈0.04). Malformed-token glitch persists on open-ended prompts (~7/100). Not affiliated with PrismML/Qwen.
