88plug/MiniCPM-o-4.5-W4A16
Load path (important)
Hub catalog is a compressed-tensors (pack-quantized) checkpoint. Prefer vLLM ≥ 0.21.
# Preferred
vllm serve 88plug/<ModelName> --trust-remote-codeDo not deploy via TGI “text-generation-inference” paths — that backend does not load our CT format and produces opaque worker/load errors.
MiniCPM-o-4.5-W4A16
INT4 post-training quantization of openbmb/MiniCPM-o-4.5 — a compact omni model with vision (SigLIP2), audio (Whisper), and speech synthesis (CosyVoice2) built on a Qwen3-8B backbone. ~4–5 GB on disk. Runs on a single 8 GB GPU.
At a Glance
Memory Requirements
Note: The non-quantized modal encoders (SigLIP2 ~1 GB, Whisper ~390 MB, CosyVoice2 ~100 MB) are included in all footprint estimates above. Only the Qwen3-8B LLM backbone is quantized to 4-bit.
Quick Start
Tested with vLLM v0.21.0 (vllm/vllm-openai:v0.21.0-cu129-ubuntu2404). Weights are in compressed-tensors format — vLLM detects and loads quantization automatically. No --quantization flag needed.
vLLM — text output
docker run --gpus device=0 -p 8080:8080 \
vllm/vllm-openai:v0.21.0-cu129-ubuntu2404 vllm serve \
88plug/MiniCPM-o-4.5-W4A16 \
--kv-cache-dtype fp8 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90Weights are in compressed-tensors format — no --quantization flag needed. Mainline vLLM returns text only; CosyVoice2 TTS output is not supported.
Python client
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="token")
response = client.chat.completions.create(
model="88plug/MiniCPM-o-4.5-W4A16",
messages=[
{"role": "user", "content": "Describe the architecture of MiniCPM-o 4.5."}
],
max_tokens=512,
)
print(response.choices[0].message.content)Quantization Design
What is quantized
Only the Qwen3-8B LLM backbone (model.llm) is quantized. AutoRound (blocked — see KNOWN-FAILURES) applies W4A16 to all Linear layers within model.llm, using a round-to-nearest rotation-based optimization with 200 calibration iterations per block.
What stays BF16
The full MiniCPM-o-4.5 checkpoint is saved via model.save_pretrained() after in-place quantization of model.llm, so the output contains all modalities — vision, audio, and TTS encoders remain at full BF16 fidelity.
Implementation notes
MiniCPM-o-4.5 required four patches to run cleanly through llmcompressor:
- `get_imports` patch — filters
minicpmo,librosa, andsoundfileimports to avoid thelibrosa→soxrcascade during quantization. - `MiniCPMTTSConfig.__getattr__` patch — backfills
top_p,top_k, and related attributes missing from the shippedconfig.json. - `_move_missing_keys_from_meta_to_device` wrap — handles
all_tied_weights_keysnot being set by MiniCPMO's remote code under transformers 5.8.1. - `is_mllm_model=False` override — forces AutoRound (blocked — see KNOWN-FAILURES) through the standard LLM path instead of the multimodal MLLM compressor, which would fail trying to load a processor from
model.llmdirectly.
Additionally, torch.nn.Module.apply and torch.nn.Module.train are replaced with iterative equivalents to avoid stack overflow on MiniCPM-o's ~985-deep module tree.
Quality Targets
Competitor Comparables
MiniCPM-o-4.5 is an omni model — meaningful comparisons must also support vision + audio input. As of publication, no other compressed-tensors or vLLM-native quantization of this model exists.
First-to-market claim: No compressed-tensors or vLLM-native W4A16 quant was found for this model at publication time. This is the only production-ready W4A16 quant for direct vLLM serving.
Benchmarks
Results pending.
Hardware: A6000 48 GB, CUDA 12.9, driver 570.
SGLang
SGLang does not natively support compressed-tensors. To use this model with SGLang, serve the BF16 base (openbmb/MiniCPM-o-4.5) or an AWQ variant.
docker run --gpus device=0 -p 30000:30000 \
lmsysorg/sglang:v0.5.8-cu129 python -m sglang.launch_server \
--model-path openbmb/MiniCPM-o-4.5 \
--tp 1 \
--mem-fraction-static 0.85 \
--port 30000SGLang results are BF16 baseline — useful as a throughput ceiling reference, not a direct quality comparison to this quant.
llama.cpp
Mainline llama.cpp supports MiniCPM-V (vision + text). For full CosyVoice2 speech output, use the `tc-mb/llama.cpp-omni` fork. Convert and quantize from the BF16 base — do not convert from compressed-tensors weights.
python convert_hf_to_gguf.py openbmb/MiniCPM-o-4.5 \
--outfile MiniCPM-o-4.5-BF16.gguf
llama-quantize MiniCPM-o-4.5-BF16.gguf MiniCPM-o-4.5-Q4_K_M.gguf Q4_K_M
llama-quantize --imatrix calibration_datav3.txt \
MiniCPM-o-4.5-BF16.gguf MiniCPM-o-4.5-IQ4_XS.gguf IQ4_XS
llama-server \
--model MiniCPM-o-4.5-Q4_K_M.gguf \
--n-gpu-layers 999 \
--ctx-size 32768 \
--port 8081Technical Details
Citation
@misc{minicpmo,
title = {MiniCPM-o: A GPT-4o Level Multimodal LLM on Your Phone},
author = {MiniCPM Team, OpenBMB},
year = {2025},
url = {https://huggingface.co/openbmb/MiniCPM-o-4.5}
}About
**88plug AI Lab** ships FLAC-target compressed-tensors quantizations — AutoRound iters=200, native vLLM v0.21.0+, no extra flags.
This release: Gold tier — full AutoRound calibration (1024 samples, UltraChat + WikiText-103). Targets ≥99% MMLU recovery.
All weights use compressed-tensors format. vLLM reads quantization_config automatically.
Browse all releases → huggingface.co/88plug
