datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-1.5-proofs-only-strict
NuminaMath-1.5-proofs-only-strict
A strictly filtered version of the NuminaMath-1.5-proofs-only dataset, containing ONLY
validated mathematical proof problems.
📊 Filtering Results
Original dataset: Numina1.5 -> filter for proofs -> 110,998 rows
Filters applied:
✓ Kept rows where answer = "proof" (proof problems only)
✓ Kept rows where solution_is_valid = "Yes"
✓ Kept rows where problem_is_valid = "Yes"
✓ Dropped validation columns after filtering
Filtered dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-proofs-only-strict.finevisionmax-strict-ans-ablation
FineVisionMax — Strict Numerical Ablation
Filtered subset of HuggingFaceM4/FineVisionMax,
for an ablation study on the emergence of approximate-number-system (ANS)
representations in vision-language models.
Filter
Strict ablation: rows where ANY user or assistant turn contains a match from
any of 15 categories spanning the REMOVE class (digits, number words,
counting verbs, comparisons, ordinals, etc.) and the EXPERIMENT class (vague
quantifiers, absence… See the full description on the dataset page: https://huggingface.co/datasets/WenqingCao/finevisionmax-strict-ans-ablation.BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.BabyLM-2026-Strict
Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 strict training set. Total: 100M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.gameplay-benchmark-strict
Strict Gameplay Benchmark
This public release contains 1079 five-second gameplay clips
from 303 canonical game rows. Every clip is 1280x720, 30 FPS,
150 frames, H.264, and passed the strict black-frame, geometry, and dense content
gates. Videos are stored without recompression in benchmark_videos.zip.
benchmark_annotations.json contains one benchmark_clip annotation per archived
video and maps every record to its published_video_path. Source-specific license
review metadata is… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/gameplay-benchmark-strict.BabyLM-2026-Strict-Evalsmovement-strict-164
movement-strict-164: High-Quality Filtered Pose Dataset
164,390 clips from Kinetics-700 that pass both programmatic continuity checks and a 235B-parameter Vision-Language Model judge that evaluated the rendered skeleton overlay against the action label. Roughly 60% of clips have been re-tracked through a dedicated multi-frame YOLO + Qwen oracle + sticky IoU tracker pipeline before judgment, replacing the original tracking with a cleaner result.
This is a filtered subset of… See the full description on the dataset page: https://huggingface.co/datasets/maxsegan/movement-strict-164.metaworld_spatial_cardinal_15cm_strict_50deg_close_high_approach_improved_mwThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "metaworld",
"total_episodes": 800,
"total_frames": 97649,
"total_tasks": 8,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 24,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/metaworld_spatial_cardinal_15cm_strict_50deg_close_high_approach_improved_mw.gaokao-sft-chinese-strict-abcd-v3
Gaokao SFT Chinese Strict ABCD V3
This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis.
Composition
Total samples: 88466
Train samples: 86670
Validation samples: 1796
Subject Counts
{
"biology": 33104,
"chemistry": 35796,
"english": 174,
"general_exam": 7982,
"geography": 888,
"history": 229,
"physics": 9986,
"politics": 307
}
Fields
id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.minicpm-v46-strict-cycle-runpod-serverless
MiniCPM Strict Caption Harness
Strict 9-field cycling harness for MiniCPM-V-4.6 v2 captions. The harness asks the model for one field at a time, cleans each field, and assembles the final caption with deterministic v2 headers.
Smoke test without loading the model:
cd /Users/dustinpainter/datasets/image-datasets
PYTHONPATH=tools/mini-cap-harness python3 tools/mini-cap-harness/run_strict_cycle.py \
--backend mock \
--input… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/minicpm-v46-strict-cycle-runpod-serverless.incompebench-strictgameplay-dataset-strict
Strict Gameplay Training Segments
This public release contains 7304 variable-length clean gameplay
segments from 4363 source videos and 254
canonical games, totaling 54.218 hours.
The videos were partitioned on whole-video boundaries into three size-balanced ZIP
shards. No compressed byte stream was split. dataset_annotations.json contains
one training_dataset_segment annotation for every archived video; raw-source
metadata is nested only as provenance. Source-specific license… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/gameplay-dataset-strict.strict-verification-reasoning
Strict Verification Reasoning Dataset
Description
A dataset for training language models to verify facts, check sources, evaluate arguments, and avoid overthinking.
Content
1,010,000 examples
5 categories: anti-overthink, comparisons, strict facts, strict sources, strict arguments
English language
Categories
Category
Description
%
Anti-Overthink
Simple, direct answers
15%
Comparisons
Hallucination vs correct answer
20%… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/strict-verification-reasoning.so101_grey_cylinder_blue_cup_currentcal_transport_corrections16_strict_v1_20260811This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 16,
"total_frames": 9019,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DylanSh/so101_grey_cylinder_blue_cup_currentcal_transport_corrections16_strict_v1_20260811.mmlu_gneissweb_strict_prunedopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match
N8 Rejection Sampling (Strict Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.atomic-metrics-ncw-strict
Atomic Metrics NCW: Strict English Audit
This version contains 100 training and 200 test preference pairs for practical
nonfiction writing. Examples and existing preference labels are preserved;
no replacement responses or preference labels were generated for this release.
Sources
Source
Train
Test
Community Alignment
72
152
Writing Preference Bench
12
12
OASST1
5
23
OASST2
11
13
Source identifiers are retained in source_dataset and… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-ncw-strict.Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.metaworld_spatial_diagonal_15cm_strict_50deg_close_high_approach_improved_mwThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "metaworld",
"total_episodes": 800,
"total_frames": 99144,
"total_tasks": 8,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 24,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/metaworld_spatial_diagonal_15cm_strict_50deg_close_high_approach_improved_mw.floorplan-aligned-strictso101_grey_cylinder_blue_cup_currentcal_precision_replacements7_strict_v1_20260811This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 7,
"total_frames": 5509,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DylanSh/so101_grey_cylinder_blue_cup_currentcal_precision_replacements7_strict_v1_20260811.strictly-speaking
Strictly Speaking
Does a model's mathematical understanding hold up strictly speaking, at Lean-grade precision,
or is it only right in the ordinary, looser sense that informal writing usually gets away with?
Each row is derived from a real, already-formalized Lean 4 theorem (drawn from
Pradheep1647/lean-verifier-formalizations).
The theorem's informal statement is split into two pieces: the hypotheses/setup (prompt), and
the conclusion that was elided from it (answer) - the… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/strictly-speaking.omni-retract-era5ns-strict-rawCAD-experiment-manifests-seed42-vision-qwen3vl32b-v1-strict-coder-v1StrictPRMBench
StrictPRMBench
A strict counterfactual audit benchmark for stress-testing PRM aggregation rules beyond final-answer accuracy.
Anonymous submission artifact.
Quick Start
pip install -r requirements.txt
python3 smoke_test.py
bash reproduce_all.sh
Hardware
No GPU required. All reproduced outputs read from cached step scores and precomputed result tables. Included strict-trace files cover 500 GSM8K problems and 500 MATH-500 problems; BON candidate directories are… See the full description on the dataset page: https://huggingface.co/datasets/staged-benchmark-2026/StrictPRMBench.qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed
Qwen3 4B Thinking SFT v54 Processed Training View
This dataset is the processed and filtered training view used by the v54
Qwen3-4B-Thinking SFT recipe. It starts from
eewer/swerebench-traces-raw-source-targeted-limitations-compaction-full-20260616-2030 and uses the strict-passed raw2030 mini-swe
aligned view.
Rows are compressed JSONL.zst files under data/. Each row contains a
top-level messages column, optional tools, and scalar source mapping fields
such as source_uuid… See the full description on the dataset page: https://huggingface.co/datasets/eewer/qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed.so101_grey_cylinder_blue_cup_currentcal_precision_topup16_strict_v2_20260811This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 16,
"total_frames": 12876,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DylanSh/so101_grey_cylinder_blue_cup_currentcal_precision_topup16_strict_v2_20260811.babylm2-rewritten-clean_multi-adj-strict-reversedgaokao-sft-chinese-strict-abcd
Gaokao SFT Chinese Balanced
This is the strict balanced Chinese SFT dataset version.
Only multiple-choice samples with explicit A/B/C/D option-level explanations are kept in this balanced release.
Composition
Total samples: 645
Train samples: 632
Validation samples: 13
Subject Counts
{
"biology": 199,
"chemistry": 170,
"english": 13,
"geography": 34,
"history": 118,
"physics": 77,
"politics": 34
}
Fields
id
lang
subject
source… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd.
