tg-techie-agents/Q1D-4B-Blk0thru3-V2
Q1D 4B Blk0thru3 V2
⚠ CORRECTION / KNOWN FAILURE (2026-06-19). Any claim that this 4-block 1-bit model "recovers to parity with the full-precision teacher" is an OVERCLAIM and is confounded. This model does NOT honor the project's thinking-OFF standard: it emits full chain-of-thought reasoning in the response body (with no<think>tags — the closed-think template strips them), so thinking is neither disabled nor delineated. Its GSM8K 0.62 reflects this de-facto thinking behavior (verbose CoT in the body), present even though this version trained on C4 only. Comparing its GSM8K score against the thinking-OFF teacher is therefore not apples-to-apples, and GSM8K's flexible-extract masks the behavioral regression. Thinking-OFF capability recovery at 4 blocks is NOT established. This was a process error on the maintainer's part (optimising the metric without checking it against the thinking-off goal). See the project'swiki/reports/017(FAILURE section).
Research artifact from Q1 Descent — reconstructing per-block 1-bit intelligence-density recovery for Qwen3. Selected blocks quantized to 1-bit (sign + per-128-group FP16 scale, Q1_0) and trained back toward the FP16 teacher; other blocks stay F16. NOT a fully 1-bit model.
This model. Blocks 0 thru 3 at 1-bit, joint STE-KD on C4 (lr 5e-5, 1800 steps). Blocks 4-35 F16.
State / eval (read with the correction above). GSM8K flex 0.62 (n=100) — reflects de-facto-thinking behavior (verbose CoT in the body), NOT thinking-off recovery.
Use. llama.cpp / LM Studio; ships the closed-think template; greedy (temp 0).
Caveats. Early research artifact, single seed, n=100 (SE≈0.04). Malformed-token glitch persists on open-ended prompts (~7/100). Not affiliated with PrismML/Qwen.
