CoolFace
Modelpublic

sahilchachra/fable-traces-mxfp8-mlx

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes15downloads
Model Card

fable-traces — MLX Block float MX FP8

MLX quantization of **AliesTaha/fable-traces**, a fine-tuned Qwen3-4B-Instruct-2507 for short, conversational replies. This variant uses Block float MX FP8 quantization (8.25 effective bits/weight).

Quantized by: sahilchachra Closest to FP16 quality; 8-bit block-float precision.

About the base model

  • Architecture: Qwen3ForCausalLM — 36 layers, hidden 2560, 32 attention heads, 8 KV heads (GQA)
  • Context length: 262 144 tokens
  • Thinking mode: Qwen3 hybrid — supports <think> chain-of-thought with enable_thinking=True
  • Fine-tune domain: Conversational / instruct (see egypt-won tag)
  • License: Apache 2.0

Quick start

bash
pip install mlx-lm
python
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/fable-traces-mxfp8-mlx")

messages = [{"role": "user", "content": "Tell me something interesting."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True)
print(response)

With thinking mode (Qwen3 chain-of-thought)

python
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True,
    enable_thinking=True,   # injects <think> block before answer
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=1024, verbose=True)

CLI

bash
mlx_lm.generate --model sahilchachra/fable-traces-mxfp8-mlx \
    --prompt "What's the fastest animal on Earth?" \
    --max-tokens 256

Quantization details

VariantFormatbpwDiskPeak RAM
FP16 (original)BF16 safetensors16.07688 MB~8 GB
mxfp8 ← thisBlock float MX FP88.253968 MB3.98 GB
sahilchachra/fable-traces-4bit-mlxAffine int4 (group size 64)4.502184 MB2.22 GB
sahilchachra/fable-traces-mxfp4-mlxBlock float MX FP44.252050 MB2.12 GB
Note on bpw: Embedding and norm layers are kept at bf16; the reported bpw is across all linear weights.

All MLX variants

RepoFormatbpwDisk
sahilchachra/fable-traces-mxfp4-mlxMX FP44.252050 MB
sahilchachra/fable-traces-4bit-mlxAffine int44.502184 MB
sahilchachra/fable-traces-mxfp8-mlx ← thisMX FP88.253968 MB

Credits