philipjohnbasile/hy3-family-mini-qwen35b-reap-v1
Hy3-Family Mini — Qwen35B REAP v1
Explore the model guide · All public work
Release at a glance
A REAP-pruned variant of `philipjohnbasile/hy3-family-mini-qwen35b-v1`: 25% of experts removed (256→192) via saliency-guided REAP pruning (arXiv:2510.13999), with soul/domain-protected experts kept (coding, math, science, security, design, fullstack, gamedev, legacy, music, art, perfumery). No LoRA heal applied — measured to beat the healed variant and match-or-beat the un-pruned base on the project's 46-case eval suite.
18 GB → 14 GB. 58.5 → 82.2 tok/s (+40%).
Why no heal
The standard "prune then heal" recipe (used successfully on this project's own Hy3-architecture models) was tested here with a real prune-only control for the first time — and healing measurably hurt this checkpoint (71.7% vs the prune-only 82.6% pass rate), root-caused to heal-induced verbosity/stopping-calibration drift. Full writeup: data-model-brain what-doesnt-work.md #36.
Verified (measured, receipts in the source repo)
- REAP prune only, ratio 0.25, saliency mean = gate × ‖expert output‖ over routed calibration tokens (460 prompts, protected-facet aware).
- Eval suite (coding/tool-calls/agent-repair/json-schema/planning/souls/hard/brutal, 46 cases): 82.6% pass (38/46), beating the un-pruned base's 80.4% (37/46).
- Speed: 82.2 tok/s decode vs base's 58.5 tok/s.
- No structural corruption: loads and generates cleanly, smoke-tested and cross-checked against the base model's behavior.
Usage
mlx_lm.chat --model <this-directory>
mlx_lm.server --model <this-directory> --port 8080 # OpenAI /v1Qwen3.6 quirk: to turn OFF thinking, pass enable_thinking=False as a DIRECT kwarg to apply_chat_template (the nested form and /no_think are ignored) — or --chat-template-args '{"enable_thinking":false}' if serving via this project's own scripts/08_serve_mlx.py.
⚠️ For strict-JSON/structured output, pin `repeat_penalty: 1.0` explicitly in the request — see the base model card for the full incident this warning is based on.
Limitations
- Same base-model caveats as
hy3-family-mini-qwen35b-v1(different base than Hy3-295B, English/agent focus, no built-in tool execution). - Quantization was pushed further (2-bit experts) and found to break the model outright (6/46 pass, genuine content corruption) — 4-bit is this checkpoint's real quantization floor. Don't requantize below 4-bit.
- A follow-up capability-injection attempt (77 teacher-distilled examples targeting the model's weak 'agentic' facet, healed via LoRA) came back a wash — not shipped as a separate variant. See data-model-brain
what-doesnt-work.md#38.
Source, recipe, receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx
