sahilchachra/Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx
6751
Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx
MLX quantization of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for Apple Silicon.
Note — text tower only. The source model is a Qwen3.5-VL multimodal model (Qwen3_5ForConditionalGeneration, with a vision encoder). This MLX conversion contains only the text/language tower — the vision encoder weights are not included, so this is a text-only model and does not accept image or video input. The text reasoning the original is benchmarked for (GSM8K, MMLU) is unaffected. It loads via the standard MLX LLM path (mlx-lm, LM Studio). For LM Studio compatibility the config carriespartial_rotary_factorinsiderope_parameters(LM Studio's engine hard-indexes that key, unlike mlx-lm which defaults it); the config is also tagged as a causal LM (architectures: ["Qwen3_5ForCausalLM"], vision/image/video token ids removed) to reflect that it is text-only.
Variant: Block float MX FP4 Disk size: 4557 MB Quantized by: sahilchachra
Benchmark results
Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.
Performance
Quality
Context scaling (decode tok/s)
Usage
pip install mlx-lmfrom mlx_lm import load, generate
model, tokenizer = load("sahilchachra/Qwythos-9B-Claude-Mythos-5-1M-mxfp4-mlx")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)All variants in this collection
Notes
- Requires Apple Silicon (M1 or later) with MLX
- Benchmarks run on Apple M5 Pro, 24 GB unified memory
- License: see empero-ai/Qwythos-9B-Claude-Mythos-5-1M for the original model's license
Original model
See empero-ai/Qwythos-9B-Claude-Mythos-5-1M for full model details and intended use.
