mlx-community/Qwen3.8-Flash-Next-oQ8e-mtp
21.4k
Qwen3.8-Flash-Next-oQ8e-mtp
This model was quantized using oQ in oMLX v0.6.4. oQe (imatrix). MTP kept.
Quantization details
- Model type: qwen4_exp
- Bits: 8
- Group size: 64
- Format: MLX safetensors
- oMLX: Jundot v0.6.4 (source)
- Parameters: 125B total, 6B activated, 51B n-gram embedding, 4B MTP
- MTP: preserved (76
mtp.*tensors) - Source: Qwen/Qwen3.8-Flash-Next, converted to MLX, then oQe with
preserve_mtp
Benchmark
Apple M3 Ultra, 256 GB. oMLX 0.6.4. Thinking off, temperature 0, Lightning MTP depth 3. oQ4e is Jundot/Qwen3.8-Flash-Next-oQ4e-mtp on the same machine.
Write speed is how fast the reply is generated (generation_tokens_per_second). Read speed is how fast the prompt is ingested (prompt_tokens_per_second).
Coding task: write a Python Fibonacci function, no comments. All three packs stopped on their own.
Longer prompts (256-token generation cap):
Write speed (tok/s)
Read speed (tok/s)
Raw timings: bench.json (this pack) and bench-compare.json (all three).
