CoolFace
Modelpublic

GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
216likes354kdownloads
Model Card

<p align="center"> <img src="assets/banner.png" alt="MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF" width="100%"/> </p>

MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF

GGUF quantizations of [MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking) for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.

中文说明

This repository provides local-deployment builds of a 1B Thinking model fine-tuned on Fable 5 data (V2) atop openbmb/MiniCPM5-1B. Compared with V1, V2 strengthens tool calling / function calling, while keeping MiniCPM5's native chat template embedded in the GGUF files.

Transformers checkpoint: [MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking)

Previous GGUF version: [MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF) (V1)


Files

FileQuantSizeNotes
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.ggufQ8_0~1.1 GBrecommended default
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-F16.ggufF16~2.1 GBfull-precision conversion base

Q8_0 is the recommended default quant for this 1B model.


Quick start

llama.cpp (llama-cli)

bash
llama-cli \
  -m MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf \
  -p "Write a Python function to merge two sorted lists." \
  -n 512 \
  --temp 0.9 --top-p 0.95 \
  -c 8192
The model supports up to 128K tokens (131,072) per config.json. Set -c according to your available VRAM/RAM.

llama.cpp server

bash
llama-server \
  -m MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf \
  -c 8192 --port 8080

LM Studio / jan / KoboldCpp

Load any .gguf file from this repository. The MiniCPM5 chat template is embedded in the GGUF metadata.


Sampling recommendations

Generation defaults are inherited from [MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B):

ModeParams
Think (default)temperature=0.9, top_p=0.95
No Thinktemperature=0.7, top_p=0.95, enable_thinking=False

Capabilities

  • Tool calling (enhanced in V2) — stronger function-calling / tool-use behavior
  • Fable 5 fine-tune (V2) — post-trained on Fable 5 data
  • Coding — code generation, debugging, and software-engineering workflows
  • Instruction following — more reliable adherence to user prompts and task constraints
  • Thinking mode — chain-of-thought reasoning; MiniCPM5 chat template baked into the GGUF
  • Long context — up to 128K tokens (131,072 tokens per upstream config.json)

Benchmark

Scores for the Transformers checkpoint MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking:

BFCL + API-Bank

ModelBFCL non_liveBFCL liveAPI-Bank
MiniCPM5-1B (Base)41.51%60.24%7.30%
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking43.06%63.33%22.10%

Tau-Bench

DomainMiniCPM5-1B (Base)MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking
Airline0.34 (17/50)0.36 (18/50)
Retail0.052 (6/115)0.070 (8/115)

Limitations

  • Thinking outputs — the model may emit reasoning blocks before the final answer
  • 1B scale — lightweight local deployment; not frontier-scale
  • Runtime context — actual usable context depends on your GGUF runtime and hardware limits

Provenance & licensing

Apache-2.0, inherited from MiniCPM5-1B.

Acknowledgements