datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Transmem_ecsd_minicpm5_1b_hotpotqa_n4_n8MiniCPM5-1B-atlas
juiceb0xc0de/MiniCPM5-1B-atlas
A brain atlas for openbmb/MiniCPM5-1B, a 1B on-device model with a 130k bilingual vocabulary. This is not a chat dataset or a benchmark. It is an internal-mechanics map, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know which parts of this model are safe to edit, where its output-vocabulary directions live, or which layers are carrying the most… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/MiniCPM5-1B-atlas.MiniCPM-RobotManip-LIBERO
MiniCPM-RobotManip LIBERO
This dataset contains the four LIBERO suites converted to LeRobot v3 format
for the MiniCPM-RobotManip LIBERO full-parameter fine-tuning example in
starVLA.
Dataset summary
Suite
Episodes
Frames
Videos
LIBERO-10
358
95,740
716
LIBERO-Goal
405
48,131
810
LIBERO-Object
450
66,294
900
LIBERO-Spatial
423
51,707
846
Total
1,636
261,872
3,272
Format: LeRobot v3
Frequency: 20 Hz
Cameras: agent view and wrist view
Video… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/MiniCPM-RobotManip-LIBERO.minicpm5-1b-SAEOne JumpReLU SAE per layer of MiniCPM5-1B. All 24 layers, complete.
MiniCPM5-1B: 24 layers, 1536-dim residual stream, 130,560-token bilingual vocab.
Every SAE in this repo: d_in=1536, 49,152 features (32x expansion), JumpReLU activation, streamed FineWeb-Edu, target sparsity L0=50. Same settings on every layer, no hyperparameter changes were applied in the run.
Each layer_NN_s0/ holds:
sae.pt - the weights
meta.json - config and final metrics
checkpoint_full.pt - full optimizer state… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/minicpm5-1b-SAE.details_indischepartij__MiniCPM-3B-OpenHermes-2.5-v2
Dataset Card for Evaluation run of indischepartij/MiniCPM-3B-OpenHermes-2.5-v2
Dataset automatically created during the evaluation run of model indischepartij/MiniCPM-3B-OpenHermes-2.5-v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_indischepartij__MiniCPM-3B-OpenHermes-2.5-v2.minicpmv_overfit_lora
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/minicpmv_overfit_lora.minicpm-v46-strict-cycle-runpod-serverless
MiniCPM Strict Caption Harness
Strict 9-field cycling harness for MiniCPM-V-4.6 v2 captions. The harness asks the model for one field at a time, cleans each field, and assembles the final caption with deterministic v2 headers.
Smoke test without loading the model:
cd /Users/dustinpainter/datasets/image-datasets
PYTHONPATH=tools/mini-cap-harness python3 tools/mini-cap-harness/run_strict_cycle.py \
--backend mock \
--input… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/minicpm-v46-strict-cycle-runpod-serverless.latent-state-tracking-minicpm5
Tracking and Intervening on Latent State Dynamics in a Small Language Agent (MiniCPM5-2B)
Date: 2026-09-17
Model studied: openbmb/MiniCPM5-2B (2.52B params, 42 layers, hidden dim 2048)
Hardware: single RTX 3070 Ti (8GB) — all experiments run on consumer-grade hardware
Summary
We ask whether a small (2.5B-parameter) language model's hidden-state trajectory during generation contains a stable, low-dimensional structure that (a) is linearly decodable into task type… See the full description on the dataset page: https://huggingface.co/datasets/B2J/latent-state-tracking-minicpm5.minicpm5-stock-v2-forward-return
MiniCPM5 Stock v2 — Forward-Return Labels
Binary BUY/SELL stock-direction dataset where labels come from actual forward
5-day returns (BUY > +2%, SELL < -2%, middle band dropped), not news sentiment.
All features are strictly causal (no look-ahead): last 20 daily returns, RSI(14),
volume ratio vs 20d MA, 20d volatility, 5d/20d momentum, 20d relative strength vs SPY.
train_minicpm5_v2.jsonl — 5,056 rows, 16 tickers, class-balanced
val_minicpm5_v2.jsonl — 1,586 rows, 4 held-out… See the full description on the dataset page: https://huggingface.co/datasets/ewin-reg/minicpm5-stock-v2-forward-return.minicpm5-2b-damage-labels
MiniCPM5-2B Damage Labels (MERNIK teacher)
Per-group measured quantization damage for MiniCPM5-2B (dense 2.6B, 42 layers).
What
damage_minicpm5_2b.jsonl — 169 rows: 1 BASELINE + 168 tied-group units.
Each unit row: the group dropped Q5_K → Q3_K while everything else stays at
Q5_K, scored by wikitext-2 PPL (-c 1024 -n 64 --seed 7).
{"unit": "ffn_down@7", "tensors": ["blk.7.ffn_down.weight"],
"ppl": 13.5364, "damage": 0.1732}
ssim_minicpm.npz — measured structural… See the full description on the dataset page: https://huggingface.co/datasets/wepiqx/minicpm5-2b-damage-labels.FLAME-ReCap-CC3M-MiniCPM-Llama3-V-2_5
Dataset description
Recaptioned CC3M by MiniCPM-Llama3-V-2_5.
Uses
The images are equivalent to https://huggingface.co/datasets/pixparse/cc3m-wds. Use data keys to index the original CC3M.
See https://github.com/MIV-XJTU/FLAME.
Citation
@article{cao2024flame,
title={FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training},
author={Cao, Anjia and Wei, Xing and Ma, Zhiheng},
journal={arXiv preprint arXiv:2411.11927}… See the full description on the dataset page: https://huggingface.co/datasets/caj/FLAME-ReCap-CC3M-MiniCPM-Llama3-V-2_5.details_indischepartij__MiniCPM-3B-Bacchus
Dataset Card for Evaluation run of indischepartij/MiniCPM-3B-Bacchus
Dataset automatically created during the evaluation run of model indischepartij/MiniCPM-3B-Bacchus on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_indischepartij__MiniCPM-3B-Bacchus.minicpm-o45-native-gate-dataOfficial-native gate training/eval bundle for issue #8 (deployed gate_native.json, 8bq recipe, train_n=5228).
Contents (official_native_bundle.zip, 187 files, unpacks to official_native_bundle/data/):
Training tags caliboff, expoff, exp2off, exp3off, exp3zhoff, freshoff: frozen_native_<tag>_feats.shard*.npz (ids + 12288-d float32 X), frozen_native_<tag>_traces.jsonl.shard* (official-native answer_text, no_speak, eot_seen, n_ans_chunks), frozen_native_<tag>_judged.parquet (id, query… See the full description on the dataset page: https://huggingface.co/datasets/dyyfk/minicpm-o45-native-gate-data.FLAME-ReCap-YFCC15M-MiniCPM-Llama3-V-2_5
Dataset description
Recaptioned YFCC15M by MiniCPM-Llama3-V-2_5.
Uses
See https://github.com/MIV-XJTU/FLAME.
Citation
@article{cao2024flame,
title={FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training},
author={Cao, Anjia and Wei, Xing and Ma, Zhiheng},
journal={arXiv preprint arXiv:2411.11927},
year={2024}
}
@article{yao2024minicpmv,
title={MiniCPM-V: A GPT-4V Level MLLM on Your Phone},
author={Yao, Yuan… See the full description on the dataset page: https://huggingface.co/datasets/caj/FLAME-ReCap-YFCC15M-MiniCPM-Llama3-V-2_5.minicpm5-stage1-dataminicpm5-1b-quantization-benchmark
openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート
対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM)
検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core)
実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0
検証日: 2026-09-19 12:12:34
1. エグゼクティブサマリー(全体比較)
NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.details_openbmb__MiniCPM-2B-dpo-bf16-llama-format
Dataset Card for Evaluation run of openbmb/MiniCPM-2B-dpo-bf16-llama-format
Dataset automatically created during the evaluation run of model openbmb/MiniCPM-2B-dpo-bf16-llama-format on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openbmb__MiniCPM-2B-dpo-bf16-llama-format.details_indischepartij__MiniCPM-3B-Hercules-v2.0
Dataset Card for Evaluation run of indischepartij/MiniCPM-3B-Hercules-v2.0
Dataset automatically created during the evaluation run of model indischepartij/MiniCPM-3B-Hercules-v2.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_indischepartij__MiniCPM-3B-Hercules-v2.0.minicpm5-vivamais-text-sft-v1
MiniCPM5 Viva Mais Text SFT v1
This dataset is the exact JSONL training/evaluation package used for the
MiniCPM5 Viva Mais text QA candidate v1 run.
Files
minicpm5_text_sft.jsonl: 12,000 SFT rows.
vivamais_qa_eval.jsonl: 32 fixed Viva Mais dashboard QA eval rows.
Training Mix
The SFT mix was generated by the Viva Mais repository pipeline from the Modal
volume minicpm5-vivamais-text-data:
2,400 rows from Polygl0t/gigaverbo-v2-sft
5,400 Viva Mais… See the full description on the dataset page: https://huggingface.co/datasets/marinarosa/minicpm5-vivamais-text-sft-v1.details_gmonsoon__MiniCPM-2B-Base-v2
Dataset Card for Evaluation run of gmonsoon/MiniCPM-2B-Base-v2
Dataset automatically created during the evaluation run of model gmonsoon/MiniCPM-2B-Base-v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_gmonsoon__MiniCPM-2B-Base-v2.details_gmonsoon__MiniCPM-2B-Base-v3
Dataset Card for Evaluation run of gmonsoon/MiniCPM-2B-Base-v3
Dataset automatically created during the evaluation run of model gmonsoon/MiniCPM-2B-Base-v3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_gmonsoon__MiniCPM-2B-Base-v3.minicpm5-vivamais-text-sft-v4
MiniCPM5 Viva Mais text SFT v4
This dataset contains the redacted training and evaluation artifacts used for
marinarosa/minicpm5-1b-vivamais-v4. It was built for Viva Mais, a local-first Portuguese WhatsApp
travel-agency copilot that answers grounded questions from an extracted CRM
context.
Files
data/train.jsonl: 4000 chat-format SFT rows.
data/eval/vivamais_qa_eval.jsonl: 158 dashboard QA eval
rows.
data/teacher/rio31_teacher_distill.jsonl: 80 accepted
rows… See the full description on the dataset page: https://huggingface.co/datasets/marinarosa/minicpm5-vivamais-text-sft-v4.details_gmonsoon__MiniCPM-2B-Base
Dataset Card for Evaluation run of gmonsoon/MiniCPM-2B-Base
Dataset automatically created during the evaluation run of model gmonsoon/MiniCPM-2B-Base on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_gmonsoon__MiniCPM-2B-Base.minicpm5-gguf-stock-analyst
Stock Analyst Financial Trading Signals Dataset
A specialized dataset for fine-tuning Large Language Models (LLMs) to act as quantitative financial analysts. This dataset contains structured technical indicator data for stocks paired with their resulting directional trading signals (BUY, SELL, HOLD).
It is formatted specifically for Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT), utilizing a hard-negative rejection strategy to force the model to learn… See the full description on the dataset page: https://huggingface.co/datasets/ewin-reg/minicpm5-gguf-stock-analyst.visualize-tagged-minicpm-prompt-allminicpm_sftminicpm5-agent-corpus-canonicalPeng_UNO_Flux_Images_MiniCPM_V4_5_ZH_Captionedminicpm5-tool-calling-xmlwellness_voice_triplets_20250908_Chi_All_MiniCPM_r25to45
