CoolFace
Modelpublic

dogfoodai/Qwen3.6-35B-A3B-Escha-W2-trellis-mlx

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes182downloads
Model Card

Qwen3.6-35B-A3B-Escha-W2-trellis-mlx

MLX (Apple Silicon) build of [EschaLabs/Qwen3.6-35B-A3B-Escha-W2](https://huggingface.co/EschaLabs/Qwen3.6-35B-A3B-Escha-W2) that keeps the routed experts verbatim in their packed EXL3 trellis form and decodes them on the fly with fused Metal kernels, while the dense (non-expert) weights are kept bit-exact against the source w8a16 int8 contract.

Requires oMLX (or mlx-lm with the bundled patches): loading swaps qwen3_5_moe's SparseMoeBlock for the trellis decoder, auto-detected from quantization_config.quant_method == "eschamoe".

~12.1 GB on disk / ~12.1 GB resident (measured mx.get_active_memory() after load). Fits a stock 24 GB Mac.

Exactness of every tensor (aligned with EschaLabs/escha-mlx)

PartFormatExactness
Routed experts gate_up_proj / down_projEXL3 trellis codes, 2/3 bpw (verbatim)bit-exact decode (((x*0xCBAC1FED)&0x8FFF8FFF)^0x3B603B60, two fp16 halves summed RNE)
Attention, shared expert, embed, lm_headaffine Q8 group-128, f32 scales/biasesbit-exact vs the source int8 w8a16 contract (dequant = f32 scale*w8, measured ±1e-6)
Router gates (mlp.gate, shared_expert_gate)fp16exact fp16
Norms / conv1d / A_log / dt_biasbf16norm +1 shift baked into weights (Qwen3.6 convention)

Hadamard transforms use MLX's native mx.hadamard_transform(scale=1/sqrt(128))

  • —bit-exact with the reference, no hand-rolled butterfly.

The non-expert halves are not re-quantized (unlike a naive 4-bit repack): the dense linears carry the exact int8 values the eschamoe runtime ships.

Usage

bash
# oMLX: place this repo in a model directory and serve it; the trellis path
# engages automatically from the config marker.

Or with mlx-lm + the oMLX patches:

bash
python - <<'EOF'
import sys; sys.path.insert(0, "/path/to/omlx")
from omlx.patches.escha_trellis import apply_escha_trellis_patch
apply_escha_trellis_patch()
from mlx_lm import load, generate
model, tok = load("dogfoodai/Qwen3.6-35B-A3B-Escha-W2-trellis-mlx")
print(generate(model, tok, prompt="Write a haiku about the ocean:", max_tokens=64))
EOF

Performance (Apple Silicon; decode is bandwidth-bound)

Single stream is trellis-decode work + dense bandwidth, so it sits ~50 tok/s (there is a ~41 tok/s bandwidth ceiling on an entry M4, ~60 on an M5 Pro - see EschaLabs/escha-mlx PERFORMANCE.md). Throughput scales with concurrency because the decode kernels amortize across sequences (batched engine decode):

concurrent requestsaggregate tok/s (measured)
1~32 (server)
4124
8135
For heavy concurrency, set ESCHA_MLX_WIRED_GB (e.g. 19) before loading - MLX's wired-limit can otherwise thrash catastrophically once the working set nears Metal's cap (the reference runtime measured a ~23x cliff on a 24 GB Mac).

Notes

  • —No MTP head: measured no throughput gain for this checkpoint (draft acceptance ~20-25% on real text; trellis verify cost scales with tokens), so the bundle stays lean.
  • —Aligned with the upstream reference implementation `EschaLabs/escha-mlx`: same decode hash, same Q8 f32-scale repack, same native Hadamard.
  • —The affine re-quant artifact (2/3-bit experts, ~112 tok/s, less exact) is published separately as ...-2bit-mlx if raw single-stream speed is the goal.

Reproducing

bash
python tools/convert_escha_mlx.py --src <escha-w2-dir> --out <parts> --expert-format trellis
python tools/finalize_escha_mlx.py --parts <parts> --out <out> --src-config <w2>/config.json   --tokenizer-dir <tok-dir> --template <w2>/chat_template.jinja --expert-format trellis

License

Apache-2.0 (weights per the Escha model card; base Qwen3.6-35B-A3B is Apache-2.0).