datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.BabyLM-2026-Strict
Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 strict training set. Total: 100M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.BabyLM-2026-Strict-Evalsbabylm-2026-fr-92m-5seed-resultsbabylm-2026-nexus-boundaries
BabyLM Nexus Boundary-Preserving Corpus
Metadata
dataset: babylm.curation.strict_small_nexus_boundaries
repo_id: miguelcsx/babylm-2026-nexus-boundaries
license: other
language: en
tags: babylm, strict-small, document-boundaries
Size
total_words: 9999979
documents: 17081
Files
records: data.jsonl
train: training/train.txt
validation: training/val.txt
shards: training/shards/
hf_disk: hf/
manifest: publish_manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-nexus-boundaries.babylm-2026-fr-92m-v2-eval-resultsbabylm-2026-nexus-episodic
BabyLM Nexus Episodic Corpus
Metadata
dataset: babylm.curation.strict_small_nexus_episodic
repo_id: miguelcsx/babylm-2026-nexus-episodic
license: other
language: en
tags: babylm, strict-small, dialogue, narrative
Size
total_words: 9999955
documents: 16525
Files
records: data.jsonl
train: training/train.txt
validation: training/val.txt
shards: training/shards/
hf_disk: hf/
manifest: publish_manifest.json
Source Words… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-nexus-episodic.bind1-babylm2026-eval-artifacts
bind1 — BabyLM 2026 Strict-Small entry: full evaluation artifacts
Complete, per-item evaluation outputs for the
SecludedCorner/bind1-babylm2026-strict-small
entry and its ablation families, released so that every number we report is downloadable and
recomputable — including the negative results.
All outputs were produced by the official babylm-eval pipeline (zero-shot backend causal;
GLUE via the official fine-tuning protocol). Reports are the pipeline's own… See the full description on the dataset page: https://huggingface.co/datasets/SecludedCorner/bind1-babylm2026-eval-artifacts.BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/poopoobabylm/BabyLM-2026-Strict-Small.babylm-2026-mtl-datababylm-2026-tournament-evidencebabylm-2026-mosaic-bpe-corpus
MOSAIC BBPE16K Encoded Corpus
Reproducible encoded training artifact shared by the MOSAIC D384 model
family. It contains the fixed 10M-word Strict-Small corpus representation used
for a controlled 100M-word exposure schedule.
Contents
packed token and word arrays for training and validation;
document offsets and token frequencies;
document-complexity metadata;
lexical recombination neighbors used by the variation-set arms;
manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.babylm-2026-fr-92m-eval-resultsbabylm-2026-sds-bbpe16-corpusbabylm2026-strict-small-custom-corpora
BabyLM 2026 Strict-Small — custom training corpora
Datasheet for the custom (teacher-generated) corpora used to train our BabyLM 2026
Strict-Small submissions. Each corpus stays within the 10M-word Strict-Small
budget (~9.98M words seen per condition). Prepared following the
Datasheets for Datasets framework (Gebru et al., 2021).
Contents
Path
What
Words
real_base/babylm10m_clean.txt
Cleaned BabyLM Strict-Small English base corpus (see cleaning below)… See the full description on the dataset page: https://huggingface.co/datasets/ksu-help/babylm2026-strict-small-custom-corpora.babylm-2026-fr-corpusbabylm-2026-random-init-eval-resultsbabylm-2026-fr-92m-strict-resultsbabylm-2026-fr-en-bpe-eval-results
