strict
babylm_2025_submission_strict-small2babylm_2025_submission_strictchinese-babylm-cog-a197-strict-bestKaren_TheEditor_V2_STRICT_Mistral_7B-i1-GGUFBabyLM-2026-Baseline-GPT2-StrictQwen3.5-13B-Strict-Instruct-i1-GGUFchinesebabylm-a7-overall-strict-20260611ConcreteGPT-124M-128ctx-Strict-Small-mixed_bins-i1-GGUF
Datasets
All datasets matching “strict”NuminaMath-1.5-proofs-only-strict
NuminaMath-1.5-proofs-only-strict
A strictly filtered version of the NuminaMath-1.5-proofs-only dataset, containing ONLY
validated mathematical proof problems.
📊 Filtering Results
Original dataset: Numina1.5 -> filter for proofs -> 110,998 rows
Filters applied:
✓ Kept rows where answer = "proof" (proof problems only)
✓ Kept rows where solution_is_valid = "Yes"
✓ Kept rows where problem_is_valid = "Yes"
✓ Dropped validation columns after filtering
Filtered dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-proofs-only-strict.finevisionmax-strict-ans-ablation
FineVisionMax — Strict Numerical Ablation
Filtered subset of HuggingFaceM4/FineVisionMax,
for an ablation study on the emergence of approximate-number-system (ANS)
representations in vision-language models.
Filter
Strict ablation: rows where ANY user or assistant turn contains a match from
any of 15 categories spanning the REMOVE class (digits, number words,
counting verbs, comparisons, ordinals, etc.) and the EXPERIMENT class (vague
quantifiers, absence… See the full description on the dataset page: https://huggingface.co/datasets/WenqingCao/finevisionmax-strict-ans-ablation.BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.BabyLM-2026-Strict
Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 strict training set. Total: 100M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.gameplay-benchmark-strict
Strict Gameplay Benchmark
This public release contains 1079 five-second gameplay clips
from 303 canonical game rows. Every clip is 1280x720, 30 FPS,
150 frames, H.264, and passed the strict black-frame, geometry, and dense content
gates. Videos are stored without recompression in benchmark_videos.zip.
benchmark_annotations.json contains one benchmark_clip annotation per archived
video and maps every record to its published_video_path. Source-specific license
review metadata is… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/gameplay-benchmark-strict.
