CoolFace
Modelpublic

philipjohnbasile/hy3-family-mini-qwen35b-reap-v1

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
0likes193downloads
Model Card

Hy3-Family Mini — Qwen35B REAP v1

Explore the model guide · All public work

Release at a glance

This artifact
PurposeA smaller Qwen Mini variant with 25% of experts removed and no additional post-prune LoRA heal.
RuntimeAutoregressive MLX tooling; see the source project for the pruning and generation recipe.
StatusRecorded runtime; see the evidence and limits below.
Tensor download14.97 GB (13.95 GiB) of root .safetensors files, including any root sidecars. This is a file-size total, not peak RAM.
Read firstThe project recorded 38/46 passes versus 37/46 for its unpruned comparison. A one-case difference is not broad model superiority.

A REAP-pruned variant of `philipjohnbasile/hy3-family-mini-qwen35b-v1`: 25% of experts removed (256→192) via saliency-guided REAP pruning (arXiv:2510.13999), with soul/domain-protected experts kept (coding, math, science, security, design, fullstack, gamedev, legacy, music, art, perfumery). No LoRA heal applied — measured to beat the healed variant and match-or-beat the un-pruned base on the project's 46-case eval suite.

18 GB → 14 GB. 58.5 → 82.2 tok/s (+40%).

Why no heal

The standard "prune then heal" recipe (used successfully on this project's own Hy3-architecture models) was tested here with a real prune-only control for the first time — and healing measurably hurt this checkpoint (71.7% vs the prune-only 82.6% pass rate), root-caused to heal-induced verbosity/stopping-calibration drift. Full writeup: data-model-brain what-doesnt-work.md #36.

Verified (measured, receipts in the source repo)

  • —REAP prune only, ratio 0.25, saliency mean = gate × ‖expert output‖ over routed calibration tokens (460 prompts, protected-facet aware).
  • —Eval suite (coding/tool-calls/agent-repair/json-schema/planning/souls/hard/brutal, 46 cases): 82.6% pass (38/46), beating the un-pruned base's 80.4% (37/46).
  • —Speed: 82.2 tok/s decode vs base's 58.5 tok/s.
  • —No structural corruption: loads and generates cleanly, smoke-tested and cross-checked against the base model's behavior.

Usage

bash
mlx_lm.chat --model <this-directory>
mlx_lm.server --model <this-directory> --port 8080   # OpenAI /v1

Qwen3.6 quirk: to turn OFF thinking, pass enable_thinking=False as a DIRECT kwarg to apply_chat_template (the nested form and /no_think are ignored) — or --chat-template-args '{"enable_thinking":false}' if serving via this project's own scripts/08_serve_mlx.py.

⚠️ For strict-JSON/structured output, pin `repeat_penalty: 1.0` explicitly in the request — see the base model card for the full incident this warning is based on.

Limitations

  • —Same base-model caveats as hy3-family-mini-qwen35b-v1 (different base than Hy3-295B, English/agent focus, no built-in tool execution).
  • —Quantization was pushed further (2-bit experts) and found to break the model outright (6/46 pass, genuine content corruption) — 4-bit is this checkpoint's real quantization floor. Don't requantize below 4-bit.
  • —A follow-up capability-injection attempt (77 teacher-distilled examples targeting the model's weak 'agentic' facet, healed via LoRA) came back a wash — not shipped as a separate variant. See data-model-brain what-doesnt-work.md #38.

Source, recipe, receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx