CoolFace
Modelpublic

Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q5_K_XL-mlx

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes37downloads
Model Card

Qwen-AgentWorld-35B-A3B — UD-Q5KXL (mlx-node)

5-bit base mixed-precision quantization of Qwen/Qwen-AgentWorld-35B-A3B for Apple Silicon, using the Unsloth Dynamic per-tensor bit allocation with imatrix-AWQ pre-scaling via mlx-node.

Qwen-AgentWorld-35B-A3B is the first native language world model for agentic environment simulation — a Qwen3.5-VL-MoE (hybrid Gated-DeltaNet + full attention, 256 experts, vision-language) that simulates agentic environments via long chain-of-thought reasoning, predicting the next environment state from an agent's action and interaction history. A single model spans seven interaction domains: MCP (tool calling), Search, Terminal, SWE, Android, Web, and OS. Trained CPT → SFT → RL on Qwen3.5-35B-A3B-Base. (technical report)

Original (BF16)This Model
Size~65 GB26 GB
FormatSafeTensors (sharded)SafeTensors (sharded)
PrecisionBF16 uniformMixed 5/6/8/8-bit affine + BF16 (imatrix-AWQ)

All Variants

RepoFormatSizeDecode (tok/s)
Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q3_K_XL-mlxUD-Q3KXL17 GB112.3
Brooooooklyn/Qwen-AgentWorld-35B-A3B-mxfp4-mlxMXFP421 GB102.6
Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q4_K_XL-mlxUD-Q4KXL22 GB99.8
Brooooooklyn/Qwen-AgentWorld-35B-A3B-nvfp4-mlxNVFP423 GB94.1
[Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q5_K_XL-mlx](https://huggingface.co/Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q5_K_XL-mlx) (this model)UD-Q5_K_XL26 GB95.4
Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q6_K_XL-mlxUD-Q6KXL31 GB92.7
Brooooooklyn/Qwen-AgentWorld-35B-A3B-mxfp8-mlxMXFP836 GB91.0
Brooooooklyn/Qwen-AgentWorld-35B-A3B-UD-Q8_K_XL-mlxUD-Q8KXL37 GB91.1

Benchmarked on a cool Apple M5 Max: median decode throughput over three 512-token generations, with a 60-second idle GPU cooldown after every generation. (Sustained decode on Apple Silicon is thermally sensitive — back-to-back benchmarking on a hot chip can understate throughput by 20–30%, so every model here was measured from a comparable cool start.)

Performance

Steady-state decode: 95.4 tok/s (1.6x vs BF16) on Apple M5 Max. Decode is memory-bandwidth bound on Apple Silicon — fewer bytes per token directly translates to higher throughput. The MoE architecture activates only 8 of 256 experts per token (~3B active out of ~34.7B total), so the active-weight footprint streamed per token is what matters.

Output Quality

Decoded-text quality was verified against the BF16 reference with a multi-judge review of the actual generated output (not a heuristic): a multi-turn factual chat plus a structured reasoning/code task. This UD-Q5_K_XL build produced coherent prose, correct facts, and a correct implementation — no runaway generation, repetition loops, or stray tokens — on par with full precision.

Per-Tensor Quantization

WeightBitsRationale
embed_tokens8-bit affineKLD ~0.15 — very low sensitivity
lm_head8-bit affineKLD ~0.05 — safest tensor
self_attn.q/k/v_proj8-bit affineKLD ~1.5–2.9 — attention-sensitive
linear_attn.in_proj_qkv/z8-bit affineKLD ~2.9 — SSM input gates
self_attn.o_proj8-bit affineKLD ~1.5; row-independent qmv for T=0 exactness
linear_attn.out_proj8-bit affineKLD ~6.0 — worst tensor; kept high
linear_attn.in_proj_a/b8-bit affinetiny low-rank GDN projections
switch_mlp.down_proj6-bit affine"slightly more sensitive" than other FFN
switch_mlp.gate_proj/up_proj5-bit affinebulk of the expert budget
Router gates (mlp.gate, shared_expert_gate)8-bit affineMoE routing accuracy
GDN params (A_log, dt_bias)bf16state-space dynamics
visual.* (vision tower)bf16vision encoder kept full precision

Quantization Strategy

Built on Unsloth Dynamic 2.0 per-tensor KLD analysis: sensitive layers (attention/SSM inputs, downproj, embeddings/head) get higher bits, while the bulk of FFN expert weights are quantized to the base width. `selfattn.oproj`, `linearattn.outproj`, the split low-rank GDN projections (`inproja/b`) and the MoE router gates are pinned to 8-bit affine (groupsize 64). GatedDeltaNet state-space parameters and the vision encoder stay bf16.

imatrix-AWQ: unlike a plain affine quant, these builds apply imatrix activation-aware pre-scaling (AWQ-style) using the unsloth imatrix, so the attention/SSM channels that matter most are scaled before rounding — recovering quality at the lowest bit widths.

Architecture

ParameterValue
Total parameters~34.7B (~3B active per token)
Hidden size2,048
Layers40 (30 linear GatedDeltaNet + 10 full attention, interval 4)
Attention heads16 (2 KV heads, GQA 8:1)
Head dimension256
Experts256 per MoE layer, top-8 routing
Vocab size248,320
Visionyes (Qwen3.5-VL vision tower, kept bf16)
Max context262,144 tokens

Usage

typescript
import { loadSession } from '@mlx-node/lm';

const session = await loadSession('./Qwen-AgentWorld-35B-A3B-UD-Q5_K_XL-mlx');

for await (const event of session.sendStream('An agent runs `ls -la` in /home/user. Predict the terminal output and the resulting environment state.', {
  config: { maxNewTokens: 2048, temperature: 0.6, reasoningEffort: 'low' },
})) {
  if (!event.done) process.stdout.write(event.text);
}

How It Was Made

bash
mlx convert \
  -i Qwen-AgentWorld-35B-A3B \
  -o Qwen-AgentWorld-35B-A3B-UD-Q5_K_XL-mlx \
  -q --q-recipe unsloth --q-bits 5\
  --imatrix-path imatrix_unsloth.gguf_file

The Unsloth recipe's per-tensor bit tiers were applied with imatrix-AWQ pre-scaling (imatrix from unsloth/Qwen-AgentWorld-35B-A3B-GGUF), so activation-weighted channels are scaled before quantization. 7-bit tiers are snapped up to 8-bit (MLX affine supports 2/3/4/5/6/8-bit).

Acknowledgments

  • —[Qwen Team](https://huggingface.co/Qwen) — For the Qwen-AgentWorld model and the Qwen3.5 base architecture
  • —[Unsloth](https://unsloth.ai) — Per-layer KLD bit-allocation (Dynamic 2.0) and the imatrix used for AWQ pre-scaling
  • —[Apple MLX](https://github.com/ml-explore/mlx) — For the Metal-accelerated ML framework

License

Apache-2.0 (inherited from base model).