CoolFace
Modelpublic

SirSahOl/MiniCPM5-2B-chat-mlx-16bit

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes844downloads
Model Card

MiniCPM5-2B-mlx-16bit

16-bit MLX conversion of openbmb/MiniCPM5-2B optimized for Apple Silicon native GPU inference.

Converted by: SirSahOl Source Model: openbmb/MiniCPM5-2B Framework: MLX by Apple Quantization: 16-bit (Average 16.00 (unquantized bfloat16) bits per weight) Format: safetensors License: apache-2.0


Model Details

  • —Architecture: LlamaForCausalLM
  • —Parameters: 2B
  • —Context Length: 131,072 tokens
  • —Format: MLX (Apple Silicon native GPU format)
  • —Quantization: 16-bit (Average 16.00 (unquantized bfloat16) bits per weight)
  • —Active VRAM Footprint: ~4.8 GB (Minimum recommended: 8 GB Unified Memory)

Quick Start

Installation

bash
pip install mlx-lm

Usage

CLI
bash
# Chat interactively
mlx_lm.chat --model SirSahOl/MiniCPM5-2B-chat-mlx-16bit

# Generate text
mlx_lm.generate --model SirSahOl/MiniCPM5-2B-chat-mlx-16bit --prompt "Write a short poem about artificial intelligence."
Python API (with Chat Template)
python
from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/MiniCPM5-2B-chat-mlx-16bit")

messages = [
    {"role": "user", "content": "Explain quantum superposition in simple terms."}
]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

response = generate(model, tokenizer, prompt=prompt, verbose=True)
print(response)

Performance Benchmarks

Measured Benchmarks (Apple M1)

Metric4-bit8-bit16-bit
Tokens/sec35.2520.356.53
TTFT28.38 ms49.23 ms179.1 ms
Peak Memory443.7 MB47.8 MB29.5 MB
Benchmarked on Apple M1 with 8GB unified memory. Average over 5 runs with 256 max tokens.

Multi-Quantization Comparison

Evaluate your hardware budget and choose the optimal precision:

VariantDisk SizeVRAM FootprintTarget Apple Silicon HardwareKey Advantage
[4-bit MLX](https://huggingface.co/SirSahOl/MiniCPM5-2B-chat-mlx-4bit)~1.5 GB~1.5 GBM1 / M2 / M3 / M4 (8GB+)Maximum generation speed and lowest RAM overhead.
[8-bit MLX](https://huggingface.co/SirSahOl/MiniCPM5-2B-chat-mlx-8bit)~2.6 GB~2.6 GBM1 / M2 / M3 / M4 Pro/Max (16GB+)Balanced accuracy and generation speed; near-lossless reasoning.
16-bit MLX (This Repository)~4.8 GB~4.8 GBM2 / M3 / M4 Max/Ultra (32GB+)Full unquantized precision; reference evaluation quality.

Who Should Use This?

Your HardwareRecommended Quantization
M1/M2/M3/M4 (8GB – 16GB)4-bit — Best balance of speed, low memory, and multitasking capability
M1/M2/M3/M4 Pro/Max (18GB – 36GB)8-bit — Higher quality reasoning with comfortable memory headroom
M1/M2/M3/M4 Max/Ultra (36GB – 192GB)16-bit — Unquantized full precision, zero quality degradation

General guidance:

  • —Use 4-bit if you want to run this model alongside IDEs, browsers, and background development tools.
  • —Use 8-bit if you have 16GB+ unified memory and require superior reasoning and code accuracy.
  • —Use 16-bit for research, benchmarking, evaluation, or high-end workstation deployments.

Other Quantization Variants


LM Studio & Local Inference Setup Guide

To prevent runaway loops and ensure correct conversational turn-taking, configure custom stop tokens in your local inference runtime:

  1. 1.<|im_start|>
  2. 2.<|im_end|>
  3. 3.<|endoftext|>

Prompt Template Formatting

  • —System Prefix: <|im_start|>system\n
  • —System Suffix: <|im_end|>\n
  • —User Prefix: <|im_start|>user\n
  • —Assistant Suffix: <|im_end|>\n<|im_start|>assistant\n

Ollama Quickstart

dockerfile
FROM SirSahOl/MiniCPM5-2B-chat-mlx-16bit
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER temperature 0.7
bash
ollama create minicpm5-2b-chat-mlx-16bit -f Modelfile
ollama run minicpm5-2b-chat-mlx-16bit

Conversion Details

PropertyValue
Source Modelopenbmb/MiniCPM5-2B
Quantization16-bit
mlx-lm Version0.31.3
Conversion Time92.22s
Output Size4.7 GB
Date2026-09-22T07:07:26.693110+00:00

Reproduction

To reproduce this conversion:

bash
pip install mlx-lm==0.31.3
python3 -m mlx_lm convert --hf-path /Users/z4/.cache/huggingface/hub/models--openbmb--MiniCPM5-2B/snapshots/12a3808a956f869c767195e9266b59c4d21d92e2 --mlx-path output/MiniCPM5-2B-mlx-16bit

Limitations & Known Issues

  • —4-bit group-wise quantization introduces minor precision loss compared to unquantized weights; for deep mathematical derivations or precision-critical reasoning, test the 8-bit or 16-bit variants.
  • —High context sequences (>32K tokens) require sufficient unified memory headroom; ensure unified memory is not overcommitted.
  • —This is a weight-only MLX conversion designed specifically for Apple Silicon GPUs (M1/M2/M3/M4 series).

License

This model conversion inherits the license of the source model: apache-2.0.

See the original model card for full license details.


Changelog

VersionDateChanges
v1.02026-09-22Initial conversion

Converted with [MLX Foundry](https://github.com/SirSahOl/mlx-foundry) — a professional pipeline for converting models to Apple MLX format.