philipjohnbasile/hy3-family-mini-qwen35b-mtp-v1
Hy3-Family Mini — Qwen35B MTP v1
Explore the model guide · All public work
Release at a glance
⚠️ MTP variant is staged/recognized, not yet verified-runnable — see Status below. For inference today use the AR sibling.
The MTP-equipped variant of hy3-family-mini-qwen35b-v1: our verifier-healed Qwen3.6-35B-A3B trunk with the Qwen NextN / MTP head (mtp.*, from the clean mlx-community/Qwen3.6-35B-A3B-MTP-4bit) grafted on for self-speculative decoding.
Runtime status — reviewed September 10, 2026
The original Forge probe recognized MTP weights, but its build failed with Model type qwen3_5_mtp not supported. That failure describes the tested release at the time; it is not a claim about every current release.
The maintainer later confirmed Qwen and Hy3 backend work shipped in MTPLX 2.1.0. Upstream release record. This exact grafted checkpoint has not been freshly qualified end to end here. Recognition, loading, accepted-token agreement, and speed are separate checks. Use the AR sibling for the documented inference path.
What it is
- Trunk:
mlx-community/Qwen3.6-35B-A3B-4bit+ our verifier-filtered LoRA heal (val 0.639, 8/10 executed stress — see the AR sibling card). - MTP head: the clean Qwen NextN head (a draft component whose output-equivalence claim depends on a correctly implemented and validated verification path).
- Clean base — no private domain-specific adapters are fused into this build.
Notes
- The fast AR daily-driver path is the AR sibling; this variant carries MTP weights for runtime qualification.
- Graft script + receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx
