Youssofal/Qwen3.6-35B-A3B-MTPLX-Optimized-Balance
Qwen3.6-35B-A3B MTPLX Optimized Balance
Balanced local 35B-A3B inference for Apple Silicon, packaged for MTPLX native Multi-Token-Prediction speculative decoding.
This checkpoint is the balanced 35B release: a 6-bit MLX body paired with calibrated INT4 MTP heads. It is tuned for strong reasoning-on generation speed while keeping prompt processing and memory use practical across normal coding contexts.
Run It
brew install youssofal/mtplx/mtplx
mtplx start
mtplx run "hello" --model Youssofal/Qwen3.6-35B-A3B-MTPLX-Optimized-BalanceFor an OpenAI-compatible local server:
mtplx serve --model Youssofal/Qwen3.6-35B-A3B-MTPLX-Optimized-Balance --profile sustained --max --port 8000 --no-stats-footerWhy This Exists
MTPLX uses the model's own MTP heads to generate draft tokens, then verifies them with the main model. When the draft heads are well-matched, you get higher throughput without running a separate drafter model.
Optimized Balance is built for that path. MTPLX reads mtplx_runtime.json and selects the measured D2 defaults automatically.
Recommended Runtime Defaults
Performance
Measured in MTPLX Sustained Max on Apple Silicon with reasoning enabled.
Generation
D2 is the promoted default because it gives the best balance of throughput, acceptance, and verify cost. A three-run D2 repeat averaged 123.44 tok/s, with every run above 122 tok/s.
Prompt Processing
Average prompt processing across the 512-to-64k ladder was 3,110.5 tok/s.
Model Build
This is not a full-precision checkpoint. It is built for fast local use on Apple Silicon through MTPLX.
Files
model-*.safetensors: MLX 6-bit body shardsmtp.safetensors: calibrated INT4 MTP sidecarmtplx_runtime.json: MTPLX runtime contract and measured defaultsMTPLX_PUBLISH_MANIFEST.json: file sizes and benchmark summary- tokenizer and config files for local loading
