MoringLabs/Qwen3.5-122B-A10B-MLX-3.7bit-VL-v2
Qwen3.5-122B-A10B-MLX-3.7bit-VL-v2
Enhanced mixed-precision MLX quantization of Qwen3.5-122B-A10B — Alibaba's latest MoE model with full vision support preserved in BF16.
- 3.748 BPW | 53.5 GB | Vision preserved (BF16)
What's New in v2
v2 invests an additional ~1.5GB over v1 to surgically upgrade the most sensitive bottleneck points:
Net effect: +1.5 GB model size for meaningfully better coherence, fewer repetitions, and improved instruction following — especially noticeable in long-form generation and complex reasoning tasks.
🚀 Hardware Optimization
This model brings 122B-class multimodal performance to Apple Silicon. By utilizing advanced mixed-precision quantization, we've compressed the model from uniform 4-bit's 65GB down to 53.5GB — an 11.5GB reduction — while preserving the full vision pipeline at BF16 precision for lossless image understanding.
- 64GB Unified Memory (Minimum): The uniform 4-bit quantization weighs 65GB and simply cannot fit in 64GB at all. This quantization breaks that barrier — fitting a full 122B multimodal model into 64GB for the first time, pushing the hardware boundaries to make local 122B vision+language inference possible on edge devices.
- 96GB+ Unified Memory (Recommended): Delivers an uncompromised, buttery-smooth multimodal experience. The efficient footprint frees up massive headroom for the KV cache, completely unlocking ultimate long-context capabilities for both text and vision tasks.
Quantization
5-tier mixed precision by functional sensitivity:
v1 vs v2 Comparison
Benchmark Results
Primary Intelligence Benchmarks
[!NOTE] Tests were conducted using MLX implementation. MMLU, CMMLU, MBPP and LiveCodeBench were sampled as indicated to balance evaluation time and statistical significance.
Coding & Reasoning (10-Question Test Suite)
Tested against Claude Opus 4.6 reference answers on a 10-question benchmark covering debugging, algorithms, security, concurrency, system design, and logic reasoning:
v2 achieved significant improvements over the previous version, demonstrating that the precision upgrades translate into measurable reasoning improvements.
Quality (WikiText-2 Perplexity)
Lower is better. Same harness for every row: WikiText-2 test split, 128 × 2048 tokens, batch size 1.
Usage
from mlx_vlm import load, generate
model, processor = load("MoringLabs/Qwen3.5-122B-A10B-MLX-3.7bit-VL-v2")
# Text conversation
response = generate(model, processor, prompt="Hello!", max_tokens=200)
# Image understanding
response = generate(model, processor, prompt="Describe this image", image="photo.jpg", max_tokens=200)
print(response)License
Apache 2.0
