CoolFace
Modelpublic

sahilchachra/Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
6likes751downloads
Model Card

Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx

MLX quantization of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for Apple Silicon.

Note — text tower only. The source model is a Qwen3.5-VL multimodal model (Qwen3_5ForConditionalGeneration, with a vision encoder). This MLX conversion contains only the text/language tower — the vision encoder weights are not included, so this is a text-only model and does not accept image or video input. The text reasoning the original is benchmarked for (GSM8K, MMLU) is unaffected. It loads via the standard MLX LLM path (mlx-lm, LM Studio). For LM Studio compatibility the config carries partial_rotary_factor inside rope_parameters (LM Studio's engine hard-indexes that key, unlike mlx-lm which defaults it); the config is also tagged as a causal LM (architectures: ["Qwen3_5ForCausalLM"], vision/image/video token ids removed) to reflect that it is text-only.

Variant: Block float MX FP4 Disk size: 4557 MB Quantized by: sahilchachra

Benchmark results

Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.

Performance

This modelFP16 baseline
Decode tok/s (avg, long traces)60.03N/A
Peak memory (GB)5.245N/A
Disk size (MB)455717969

Quality

BenchmarkThis modelFP16 baselinen
GSM8K (math, accuracy)92.0%N/A50
MMLU (knowledge, accuracy)74.0%N/A50

Context scaling (decode tok/s)

Context lengthDecode tok/s
~128 tokens60.9
~256 tokens60.6
~512 tokens60.4
~1024 tokens60.6

Usage

bash
pip install mlx-lm
python
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)

All variants in this collection

Notes

Original model

See empero-ai/Qwythos-9B-Claude-Mythos-5-1M for full model details and intended use.