datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dyck-k128-seq_len_2048-1B
dyck-k128-seq_len_2048-1B
Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256.
Grammar parameters
param
value
k (bracket types)
128
max_depth
16
p_open
0.5
seq_length
2048
file
split
tokens
train.bin
train
999,999,488
val.bin
val
10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
sutra-1B
Sutra 1B Pretraining Dataset
A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational patterns
Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.InternVL-SA-1B-Caption
Dataset Card for InternVL-SA-1B-Caption
Overview
The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.magpie-llama-3.2-1b-instructgithub-code-nanochatbpe-1B
github-code-nanochatbpe-1B
GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
1,000,000,000
val.bin
val
10,000,000
train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.llama-3-2-1B-f32ibm-granite__granite-3.0-1b-a400m-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-1b-a400m-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-1b-a400m-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-1b-a400m-instruct-details.ibm-granite__granite-3.0-1b-a400m-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-1b-a400m-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-1b-a400m-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-1b-a400m-base-details.realtreetune__rho-1b-sft-MATH-details
Dataset Card for Evaluation run of realtreetune/rho-1b-sft-MATH
Dataset automatically created during the evaluation run of model realtreetune/rho-1b-sft-MATH
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/realtreetune__rho-1b-sft-MATH-details.pankajmathur__orca_mini_v9_5_1B-Instruct_preview-details
Dataset Card for Evaluation run of pankajmathur/orca_mini_v9_5_1B-Instruct_preview
Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v9_5_1B-Instruct_preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v9_5_1B-Instruct_preview-details.Novaciano__Cultist-3.2-1B-details
Dataset Card for Evaluation run of Novaciano/Cultist-3.2-1B
Dataset automatically created during the evaluation run of model Novaciano/Cultist-3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Novaciano__Cultist-3.2-1B-details.oopere__pruned40-llama-3.2-1B-details
Dataset Card for Evaluation run of oopere/pruned40-llama-3.2-1B
Dataset automatically created during the evaluation run of model oopere/pruned40-llama-3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/oopere__pruned40-llama-3.2-1B-details.minicpm5-1b-quantization-benchmark
openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート
対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM)
検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core)
実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0
検証日: 2026-09-19 12:12:34
1. エグゼクティブサマリー(全体比較)
NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details
Dataset Card for Evaluation run of agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
Dataset automatically created during the evaluation run of model agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details.adaptive-retro-gpt-1b-corpus
Adaptive-RETRO-GPT-1B Pretraining Corpus
Cleaned causal language modeling corpus for the Adaptive-RETRO-GPT-1B run.
Source: HuggingFaceFW/fineweb-edu / sample-10BT
Train rows: 80000
Validation rows: 4000
Format: JSONL with text and source
Lil-R__PRYMMAL-ECE-1B-SLERP-V1-details
Dataset Card for Evaluation run of Lil-R/PRYMMAL-ECE-1B-SLERP-V1
Dataset automatically created during the evaluation run of model Lil-R/PRYMMAL-ECE-1B-SLERP-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__PRYMMAL-ECE-1B-SLERP-V1-details.bdh-cl_phase-1b
BDH Continual-Learning phase-1b checkpoints (RA2, single-copy subset)
17 PyTorch checkpoints from the RA2 route-aware growth ladder — the 20-phase
continual-learning run over 20 European languages (europarl corpora) that
precedes RA2b and whose run names (ladRA2-*) are cited throughout bdh/docs.
This repo contains exactly those files of the RA2 ladder that exist in a single
place (on the gx10 training host) and therefore cannot be reclaimed locally
until replicated here: the bg… See the full description on the dataset page: https://huggingface.co/datasets/Saga-AI-Labs/bdh-cl_phase-1b.cognitivecomputations__Dolphin3.0-Llama3.2-1B-details
Dataset Card for Evaluation run of cognitivecomputations/Dolphin3.0-Llama3.2-1B
Dataset automatically created during the evaluation run of model cognitivecomputations/Dolphin3.0-Llama3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__Dolphin3.0-Llama3.2-1B-details.Slimpajama_downsample_32k_1Bminimind-stage1b-mixKingNish__qwen-1b-continued-v2-details
Dataset Card for Evaluation run of KingNish/qwen-1b-continued-v2
Dataset automatically created during the evaluation run of model KingNish/qwen-1b-continued-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KingNish__qwen-1b-continued-v2-details.ultron-t1-live-feedultrafeedback_olmo1b_refNexesenex__Llama_3.2_1b_AquaSyn_0.1-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.2_1b_AquaSyn_0.1
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.2_1b_AquaSyn_0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.2_1b_AquaSyn_0.1-details.1b-sft-dataset
1B-SFT — High-Quality Instruction Dataset for a ~1B LLM
A curated, deduplicated instruction-tuning dataset in ShareGPT format, built for fine-tuning a 1B-class model (Qwen2.5-1.5B, Llama-3.2-1B, Gemma-2-2B).
5,235 samples — train 4,919 / val 158 / test 158.
Quality guarantees
Math: every answer is computed by the generator (arithmetic, percentages, word problems, multi-step problems, unit conversions, linear equations, sequences, fractions) with step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/andro124543/1b-sft-dataset.1B-question-answerLlama-3.2-1B-Instruct and gemma-3-1b-it responses to MuskumPillerum/General-Knowledge dataset.
sokoban_easy_1box_remap_prev
VisGym sokoban_easy_1box Remap-Prev
Solver-success Sokoban trajectories generated from the VisGym Sokoban environment.
HF repo: https://huggingface.co/datasets/novastar112/sokoban_easy_1box_remap_prev
task: sokoban_easy_1box
env_id: sokoban/easy
env_kwargs: {"action_mode": "move_only", "dim_room": [7, 7], "max_steps": 20, "num_boxes": 1, "num_gen_steps": 23, "success_mode": "stop_then_solved"}
action schema: move-only ('move', 1..4) plus ('stop', 'stop')
success semantics:… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/sokoban_easy_1box_remap_prev.Novaciano__La_Mejor_Mezcla-3.2-1B-details
Dataset Card for Evaluation run of Novaciano/La_Mejor_Mezcla-3.2-1B
Dataset automatically created during the evaluation run of model Novaciano/La_Mejor_Mezcla-3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Novaciano__La_Mejor_Mezcla-3.2-1B-details.KingNish__qwen-1b-continued-details
Dataset Card for Evaluation run of KingNish/qwen-1b-continued
Dataset automatically created during the evaluation run of model KingNish/qwen-1b-continued
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KingNish__qwen-1b-continued-details.
