CoolFace
Modelpublic

88plug/MiniCPM-o-4.5-W4A16

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
3likes124downloads
Model Card

Load path (important)

Hub catalog is a compressed-tensors (pack-quantized) checkpoint. Prefer vLLM ≥ 0.21.

RuntimeSupported
vLLM ≥ 0.21Yes — preferred (auto-detect CT; no --quantization flag)
transformers + compressed-tensorsYes for many text models; multimodal may need custom code
Text Generation Inference (TGI)Not supported for these CT packs
Hugging Face Inference WidgetOften fails — use vLLM locally instead
bash
# Preferred
vllm serve 88plug/<ModelName> --trust-remote-code

Do not deploy via TGI “text-generation-inference” paths — that backend does not load our CT format and produces opaque worker/load errors.

MiniCPM-o-4.5-W4A16

INT4 post-training quantization of openbmb/MiniCPM-o-4.5 — a compact omni model with vision (SigLIP2), audio (Whisper), and speech synthesis (CosyVoice2) built on a Qwen3-8B backbone. ~4–5 GB on disk. Runs on a single 8 GB GPU.


At a Glance

PropertyValue
Base modelopenbmb/MiniCPM-o-4.5
Release tierGold (AutoRound iters=200)
Quant methodAutoRound W4A16-G32 iters=200
FLAC statusNot measured (T+7d milestone)
ArchitectureQwen3-8B LLM + SigLIP2 vision + Whisper audio + CosyVoice2 TTS
Quant formatcompressed-tensors (native vLLM)
SchemeW4A16
Group sizedefault (128)
Quantizedmodel.llm transformer Linear layers (Qwen3-8B backbone)
Kept BF16Vision encoder (SigLIP2), audio encoder (Whisper), TTS (CosyVoice2), embeddings, LM head, norms
Disk sizeHub catalog complete (CT pack)
Min GPU1× RTX 3080 10 GB

Memory Requirements

ConfigurationBF16W8A16W4A16
Weights~18 GB~9 GB~4–5 GB
Min GPU1× A100 40 GB1× RTX 3090 24 GB1× RTX 3080 10 GB

Note: The non-quantized modal encoders (SigLIP2 ~1 GB, Whisper ~390 MB, CosyVoice2 ~100 MB) are included in all footprint estimates above. Only the Qwen3-8B LLM backbone is quantized to 4-bit.


Quick Start

Tested with vLLM v0.21.0 (vllm/vllm-openai:v0.21.0-cu129-ubuntu2404). Weights are in compressed-tensors format — vLLM detects and loads quantization automatically. No --quantization flag needed.

vLLM — text output

bash
docker run --gpus device=0 -p 8080:8080 \
  vllm/vllm-openai:v0.21.0-cu129-ubuntu2404 vllm serve \
  88plug/MiniCPM-o-4.5-W4A16 \
  --kv-cache-dtype fp8 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

Weights are in compressed-tensors format — no --quantization flag needed. Mainline vLLM returns text only; CosyVoice2 TTS output is not supported.

Python client

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1", api_key="token")

response = client.chat.completions.create(
    model="88plug/MiniCPM-o-4.5-W4A16",
    messages=[
        {"role": "user", "content": "Describe the architecture of MiniCPM-o 4.5."}
    ],
    max_tokens=512,
)
print(response.choices[0].message.content)

Quantization Design

What is quantized

Only the Qwen3-8B LLM backbone (model.llm) is quantized. AutoRound (blocked — see KNOWN-FAILURES) applies W4A16 to all Linear layers within model.llm, using a round-to-nearest rotation-based optimization with 200 calibration iterations per block.

What stays BF16

ComponentModule pathPrecisionReason
Vision encodervision_model.*BF16Excluded from recipe
Audio encoderaudio_model.*BF16Excluded from recipe
CosyVoice2 TTS decodertts.*BF16Excluded from recipe
Embedding layersre:.*embed_tokens$BF16Standard practice (ignore list)
Layer normsre:.*norm$BF16Standard practice (ignore list)
LM headlm_headBF16Standard practice (ignore list)

The full MiniCPM-o-4.5 checkpoint is saved via model.save_pretrained() after in-place quantization of model.llm, so the output contains all modalities — vision, audio, and TTS encoders remain at full BF16 fidelity.

Implementation notes

MiniCPM-o-4.5 required four patches to run cleanly through llmcompressor:

  1. 1.`get_imports` patch — filters minicpmo, librosa, and soundfile imports to avoid the librosa→soxr cascade during quantization.
  2. 2.`MiniCPMTTSConfig.__getattr__` patch — backfills top_p, top_k, and related attributes missing from the shipped config.json.
  3. 3.`_move_missing_keys_from_meta_to_device` wrap — handles all_tied_weights_keys not being set by MiniCPMO's remote code under transformers 5.8.1.
  4. 4.`is_mllm_model=False` override — forces AutoRound (blocked — see KNOWN-FAILURES) through the standard LLM path instead of the multimodal MLLM compressor, which would fail trying to load a processor from model.llm directly.

Additionally, torch.nn.Module.apply and torch.nn.Module.train are replaced with iterative equivalents to avoid stack overflow on MiniCPM-o's ~985-deep module tree.


Quality Targets

MetricTarget
KL divergence vs BF16< 0.014
MMLU recovery≥ 99%
RULER@128k≥ 97%

Competitor Comparables

MiniCPM-o-4.5 is an omni model — meaningful comparisons must also support vision + audio input. As of publication, no other compressed-tensors or vLLM-native quantization of this model exists.

ModelSourceFormatCompare angle
openbmb/MiniCPM-o-4.5officialBF16Quality ceiling
88plug/MiniCPM-o-4.5-W8A1688plugcompressed-tensors W8A16Higher-precision sibling
88plug/MiniCPM-o-4.5-W4A1688plugcompressed-tensors W4A16This model

First-to-market claim: No compressed-tensors or vLLM-native W4A16 quant was found for this model at publication time. This is the only production-ready W4A16 quant for direct vLLM serving.


Benchmarks

Results pending.

EngineFormatBatchctxtok/sTTFT p50TTFT p99VRAM
vLLM v0.21.0W4A16 compressed-tensors132k
vLLM v0.21.0W4A16 compressed-tensors832k
vLLM v0.21.0W4A16 compressed-tensors1128k
SGLang v0.5.8BF16 (baseline)132k
llama.cpp b9297Q4KM GGUF132k
llama.cpp b9297IQ4_XS GGUF132k

Hardware: A6000 48 GB, CUDA 12.9, driver 570.


SGLang

SGLang does not natively support compressed-tensors. To use this model with SGLang, serve the BF16 base (openbmb/MiniCPM-o-4.5) or an AWQ variant.

bash
docker run --gpus device=0 -p 30000:30000 \
  lmsysorg/sglang:v0.5.8-cu129 python -m sglang.launch_server \
  --model-path openbmb/MiniCPM-o-4.5 \
  --tp 1 \
  --mem-fraction-static 0.85 \
  --port 30000

SGLang results are BF16 baseline — useful as a throughput ceiling reference, not a direct quality comparison to this quant.


llama.cpp

Mainline llama.cpp supports MiniCPM-V (vision + text). For full CosyVoice2 speech output, use the `tc-mb/llama.cpp-omni` fork. Convert and quantize from the BF16 base — do not convert from compressed-tensors weights.

bash
python convert_hf_to_gguf.py openbmb/MiniCPM-o-4.5 \
  --outfile MiniCPM-o-4.5-BF16.gguf

llama-quantize MiniCPM-o-4.5-BF16.gguf MiniCPM-o-4.5-Q4_K_M.gguf Q4_K_M
llama-quantize --imatrix calibration_datav3.txt \
  MiniCPM-o-4.5-BF16.gguf MiniCPM-o-4.5-IQ4_XS.gguf IQ4_XS

llama-server \
  --model MiniCPM-o-4.5-Q4_K_M.gguf \
  --n-gpu-layers 999 \
  --ctx-size 32768 \
  --port 8081

Technical Details

ParameterValue
QuantizerAutoRound (blocked — see KNOWN-FAILURES) (via llmcompressor AutoRound (blocked — see KNOWN-FAILURES)Modifier)
Targets["Linear"] within model.llm
SchemeW4A16
Pipelinesequential
Max seq length2048
Ignore listlm_head, re:.*embed_tokens$, re:.*norm$
ActivationsFP16 (unquantized — W4A16)
trustremotecoderequired

Citation

bibtex
@misc{minicpmo,
  title  = {MiniCPM-o: A GPT-4o Level Multimodal LLM on Your Phone},
  author = {MiniCPM Team, OpenBMB},
  year   = {2025},
  url    = {https://huggingface.co/openbmb/MiniCPM-o-4.5}
}

About

**88plug AI Lab** ships FLAC-target compressed-tensors quantizations — AutoRound iters=200, native vLLM v0.21.0+, no extra flags.

This release: Gold tier — full AutoRound calibration (1024 samples, UltraChat + WikiText-103). Targets ≥99% MMLU recovery.

All weights use compressed-tensors format. vLLM reads quantization_config automatically.

Browse all releases → huggingface.co/88plug