dogfoodai/Qwen3.6-35B-A3B-Escha-W2-trellis-mlx
Qwen3.6-35B-A3B-Escha-W2-trellis-mlx
MLX (Apple Silicon) build of [EschaLabs/Qwen3.6-35B-A3B-Escha-W2](https://huggingface.co/EschaLabs/Qwen3.6-35B-A3B-Escha-W2) that keeps the routed experts verbatim in their packed EXL3 trellis form and decodes them on the fly with fused Metal kernels, while the dense (non-expert) weights are kept bit-exact against the source w8a16 int8 contract.
Requires oMLX (or mlx-lm with the bundled patches): loading swaps qwen3_5_moe's SparseMoeBlock for the trellis decoder, auto-detected from quantization_config.quant_method == "eschamoe".
~12.1 GB on disk / ~12.1 GB resident (measured mx.get_active_memory() after load). Fits a stock 24 GB Mac.
Exactness of every tensor (aligned with EschaLabs/escha-mlx)
Hadamard transforms use MLX's native mx.hadamard_transform(scale=1/sqrt(128))
- bit-exact with the reference, no hand-rolled butterfly.
The non-expert halves are not re-quantized (unlike a naive 4-bit repack): the dense linears carry the exact int8 values the eschamoe runtime ships.
Usage
# oMLX: place this repo in a model directory and serve it; the trellis path
# engages automatically from the config marker.Or with mlx-lm + the oMLX patches:
python - <<'EOF'
import sys; sys.path.insert(0, "/path/to/omlx")
from omlx.patches.escha_trellis import apply_escha_trellis_patch
apply_escha_trellis_patch()
from mlx_lm import load, generate
model, tok = load("dogfoodai/Qwen3.6-35B-A3B-Escha-W2-trellis-mlx")
print(generate(model, tok, prompt="Write a haiku about the ocean:", max_tokens=64))
EOFPerformance (Apple Silicon; decode is bandwidth-bound)
Single stream is trellis-decode work + dense bandwidth, so it sits ~50 tok/s (there is a ~41 tok/s bandwidth ceiling on an entry M4, ~60 on an M5 Pro - see EschaLabs/escha-mlx PERFORMANCE.md). Throughput scales with concurrency because the decode kernels amortize across sequences (batched engine decode):
For heavy concurrency, setESCHA_MLX_WIRED_GB(e.g.19) before loading - MLX's wired-limit can otherwise thrash catastrophically once the working set nears Metal's cap (the reference runtime measured a ~23x cliff on a 24 GB Mac).
Notes
- No MTP head: measured no throughput gain for this checkpoint (draft acceptance ~20-25% on real text; trellis verify cost scales with tokens), so the bundle stays lean.
- Aligned with the upstream reference implementation `EschaLabs/escha-mlx`: same decode hash, same Q8 f32-scale repack, same native Hadamard.
- The affine re-quant artifact (2/3-bit experts, ~112 tok/s, less exact) is published separately as
...-2bit-mlxif raw single-stream speed is the goal.
Reproducing
python tools/convert_escha_mlx.py --src <escha-w2-dir> --out <parts> --expert-format trellis
python tools/finalize_escha_mlx.py --parts <parts> --out <out> --src-config <w2>/config.json --tokenizer-dir <tok-dir> --template <w2>/chat_template.jinja --expert-format trellisLicense
Apache-2.0 (weights per the Escha model card; base Qwen3.6-35B-A3B is Apache-2.0).
