jakeatx/slimder-qwen38-reap384-depth32-mb2-3-4-5
SLIMDER Qwen3.8 REAP-384 depth-32 (macroblocks 2, 3, 4, 5 removed)
This is the promoted BF16 depth-32 structural checkpoint from the S3 beam search. It derives from the public depth-40 REAP-384 checkpoint at revision 2835be64b0429eddef35c001b818929223476c5e.
The parent had already removed original macroblocks 4 and 5. S3 removed local depth-40 macroblocks 2 and 3, which are original macroblocks 2 and 3 (source layers 8-15). The resulting model therefore removes original macroblocks 2, 3, 4, and 5.
Validation evidence is public in sjakek/slimder-qwen38-s3-results-20260831:
- 17-candidate runtime beam screen: residual cosine 0.919, top-1 agreement 0.775, loss increase 0.154;
- disjoint holdout: residual cosine 0.913, top-1 agreement 0.705, loss increase 0.153, ranking first among the three held-out candidates;
- repaired executable functional gate: 7/8 tasks, versus 6/8 for the unchanged depth-40 parent, with no category regression. Executed Python and exact JSON tool calls both passed; the one failure was the same schedule-arithmetic item missed by the parent.
The holdout category slices are noisier than the aggregate: retrieval residual cosine was 0.876 and code/tool top-1 agreement was 0.638/0.631, while the corresponding executable tasks passed. Treat this checkpoint as an experimental Pareto-frontier artifact rather than a drop-in replacement without downstream evaluation.
The checkpoint contains 32 transformer layers and 115,323,137,280 parameters. The routed-expert width remains REAP-384.
<!-- qwen38-perian-lineage:start -->
Qwen3.8 Perian project lineage
This repository is retained in the Qwen3.8 Perian checkpoints collection. Its exact position in the lineage is: Depth-pruning precursor at 32 layers. It retains 384 experts per layer and the full PLE table and predates the Perian QLoRA.
The final Qwen3.8 Perian GGUF release combines three reductions and one post-training stage:
- depth: 48 to 32 transformer layers;
- routed-expert width: 384 to 288 experts per layer;
- PLE n-gram capacity: 320,001,446 to 160,000,768 rows (50%, about 25.60B parameters removed), using activation-aware bigram and frequency-ranked trigram selections validated on a document-disjoint 5M-token holdout;
- rank-32 QLoRA on 12,558 normalized traces spanning math/STEM reasoning, coding/debugging, agentic tool use, retrieval, and general multi-step reasoning. The trace mixture draws from several frontier-model families, including Fable 5, GLM 5.2, Kimi K3, Claude Opus 4.7, Qwen3.8-Max, and GPT-5.6-Sol. The final merged milestone was trained through 9,336,692 supervised assistant tokens.
Earlier checkpoints in this collection do not inherit later stages merely by being listed beside them; the stage statement above is authoritative for this artifact. <!-- qwen38-perian-lineage:end -->
