edgemindroboticslabs/MiniCPM5-1B-GGUF
031
MiniCPM5 1B GGUF
This repository contains a Q4KM GGUF quantization of `openbmb/MiniCPM5-1B`.
MiniCPM5-1B is a small on-device text-generation model with long-context and tool-calling tags. This quantized release is intended for local inference with llama.cpp-compatible runtimes such as LM Studio.
Files
Model Details
- Base model:
openbmb/MiniCPM5-1B - Architecture: Llama-compatible
- Parameters: ~1.08B
- Context length in GGUF metadata: 131072
- Languages: English, Chinese
- License: Apache-2.0
Use With llama.cpp
llama-cli \
-m minicpm5-1b-q4_k_m.gguf \
-p "<|im_start|>user\nWrite a small Python function that validates an email address.<|im_end|>\n<|im_start|>assistant\n" \
-n 200 \
--temp 0.7Use With LM Studio
Download minicpm5-1b-q4_k_m.gguf and import it as a local GGUF model.
Recommended hardware label:
- Apple Silicon: Apple M1 Pro or newer
- Unified memory: 16 GB works for this Q4KM file
Conversion
Converted locally with llama.cpp:
hf download openbmb/MiniCPM5-1B --local-dir work/MiniCPM5-1B
uv run --python /opt/homebrew/bin/python3.11 \
--with-requirements work/llama.cpp/requirements/requirements-convert_hf_to_gguf.txt \
work/llama.cpp/convert_hf_to_gguf.py \
work/MiniCPM5-1B \
--outfile outputs/minicpm5-1b-f16.gguf \
--outtype f16
work/llama.cpp/build-gguf2/bin/llama-quantize \
outputs/minicpm5-1b-f16.gguf \
outputs/minicpm5-1b-q4_k_m.gguf \
Q4_K_MNotes
This is a quantized distribution of the upstream model, not a new fine-tune. Quality and behavior are inherited from openbmb/MiniCPM5-1B.
