AXERA-TECH/MiniCPM5-1B-C256-P12K-CTX16K
MiniCPM5-1B C256 P12K CTX16K on AXERA NPU
Ready-to-run deployment package for openbmb/MiniCPM5-1B on AX650 / NPU3.
- This release packages the AX650
axllmruntime together with the compiled text.axmodelfiles. - The packaged runtime is configured for text-only inference on AX650 / NPU3.
- The packaged context layout is
prefill_len=256,kv_cache_len=16383, andprefill_max_token_num=12544. - Thinking is disabled by default and can be enabled per request through the public OpenAI-compatible API.
- The package includes the tokenizer, runtime config files, and the validated
bin/axllmbinary for board-side deployment.
Supported Platform
- [x] AX650 / NPU3
Validated Devices
This package has been validated on the following AX650-based device:
- AX650 / NPU3 development board
This package was validated with AX650 SDK V3.16.0:
libax_interpreter.so V3.16.0libax_engine.so V3.16.0
For this CTX16K profile, use AX650 SDK V3.16.0 or later.
Performance
All measurements below were taken on AX650 / NPU3 with the packaged axllm runtime. TTFT stands for time to first token. In this table, TTFT is measured end-to-end from request arrival at axllm serve to the first generated token.
The validated text prompt below spans multiple 256-token prefill chunks. To avoid one-time startup effects, the reported TTFT excludes the first request for this prompt pattern.
The packaged runtime uses the following context layout:
prefill_len=256kv_cache_len=16383prefill_max_token_num=12544
The Long text generation reference row is the recommended sustained text-only decode figure for this package.
Startup Runtime Footprint
The runtime CMM figure above is a reference measurement from the validated AX650 board. Actual free or remaining CMM depends on the board memory layout and any other services running on the system.
Package Layout
.
├── README.md
├── config.json
├── post_config.json
├── bin/
│ ├── axllm
│ └── axllm.version.json
├── minicpm5_tokenizer.txt
├── model.embed_tokens.weight.bfloat16.bin
├── llama_p256_l0_together.axmodel
├── ...
├── llama_p256_l23_together.axmodel
└── llama_post.axmodelThis package uses a flat runtime layout. The packaged axllm binary reads the root-level runtime files directly, so serving the repository root is sufficient.
Direct Inference with axllm
Download the Model Package
mkdir -p AXERA-TECH/MiniCPM5-1B-C256-P12K-CTX16K
cd AXERA-TECH/MiniCPM5-1B-C256-P12K-CTX16K
hf download AXERA-TECH/MiniCPM5-1B-C256-P12K-CTX16K --local-dir .Install axllm
Option 1: use the validated binary included in this repository:
chmod +x ./bin/axllmOption 2: install from the public repository:
git clone -b axllm https://github.com/AXERA-TECH/ax-llm.git
cd ax-llm
./install.shOption 3: install with a one-line command:
curl -fsSL https://raw.githubusercontent.com/AXERA-TECH/ax-llm/axllm/install.sh | bashOption 4: download the prebuilt binary from GitHub Actions CI:
If you do not have a local build environment, download the latest CI-generated axllm binary from GitHub Actions: https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm Then run:
chmod +x axllm
sudo mv axllm /usr/bin/axllmRun on the Board
From the package root on the board:
chmod +x ./bin/axllm
./bin/axllm serve . --port 8000Expected model id:
AXERA-TECH/MiniCPM5-1B-AX650-C256-P12K-CTX16KHealth check and model listing:
curl http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/v1/modelsText Request
By default, this package runs in no-thinking mode because config.json sets enable_thinking=false.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "AXERA-TECH/MiniCPM5-1B-AX650-C256-P12K-CTX16K",
"messages": [
{
"role": "user",
"content": "请用一句话回答:AX650 是什么平台?"
}
],
"max_tokens": 64,
"temperature": 0
}'Example output:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "AX650 是一个基于 ARM 架构的嵌入式系统平台。"
},
"finish_reason": "stop"
}
]
}Enable Thinking Per Request
Pass top-level enable_thinking=true to enable explicit reasoning output for a single request.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "AXERA-TECH/MiniCPM5-1B-AX650-C256-P12K-CTX16K",
"messages": [
{
"role": "user",
"content": "中国的首都是哪里?请简短思考后给最终答案。"
}
],
"enable_thinking": true,
"max_tokens": 384,
"temperature": 0
}'Typical output shape:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "<think>\n...\n</think>\n\n中国的首都是北京。"
},
"finish_reason": "stop"
}
]
}Browser UI with lite_webui
If you want a browser UI for the OpenAI-compatible service started by axllm serve, use AXERA-TECH/lite_webui.
Set the OpenAI base URL to http://<board-ip>:8000 and the model name to AXERA-TECH/MiniCPM5-1B-AX650-C256-P12K-CTX16K.
Conversion References
If you need the original model files or want to rebuild the deployment artifacts, start with:
- Original Hugging Face model: openbmb/MiniCPM5-1B
- Public runtime repository: AXERA-TECH/ax-llm
Discussion
- GitHub Issues
- QQ group:
139953715
