datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stratified_10m_curriculum
Dataset Card for Stratified 10M Curriculum
This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange.
Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5).
Child-directed speech accounts for nearly half of the original dataset by word count.
In preliminary experiments using a training data influence estimation method, this category was by far the most influential.
This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.babylm_2024_10m_curriculum
Dataset Card for BabyLM 2024 10M Curriculum
The documents from the 10M dataset provided by the 2024 BabyLM challange.
We add a validation split we with additional documents from the 100M dataset.
The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt).
Pretraining split (train)
Stage
Words
Documents
C1: Child Directed Speech
2839591
28.53%
580000
49.19%
C2: Unscripted Dialogue
1079286
10.84%
108000
9.16%
C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
zhoblimpBabyLM-2026-Strict
Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 strict training set. Total: 100M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.stratified_equitoken_10m_curriculumformatted-CHILDESBabyLM-BLIMP-Filteredbabylm-2026-fr-92m-5seed-resultsSpanish-BabyLMbabylm-100M
BabyLM 100M
This curated dataset is originally from the BabyLM Challenge.
It consists of ~100M words of mixed domain, consisting of the following sources:
CHILDES (child-directed speech)
Subtitles (speech)
BNC (speech)
TED talks (speech)
children's books (simple written language)
babylm-xho
babylm-xho
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: xho
Script: Latin
Number of Documents: 12352
Total Tokens: 664966
Tokens Per Category
child-books: 98144 tokens
educational: 65208 tokens
padding-mt: 60511 tokens
padding-wikipedia: 387662 tokens
qed: 29099 tokens
simplified-text: 24342 tokens
Data Fields
text: The document text
category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.babylm-german
German BabyLM dataset
This is a pre-training dataset for training developmentally plausible language models in German (also called BabyLMs), compiled by the Computational Linguistics Group (CLAUSE) at Bielefeld University.
If you are looking for ways to evaluate your German BabyLMs, we recommend our own lexical decision dataset, CLAMS for syntactic evaluation and XCOMPS for conceptual semantics/world knowledge.
The composition is inspired by the original, English BabyLM dataset (see… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/babylm-german.hanzi-pinyinhanzi-structurebabylm-surprisal-resultsbabylm1sentscounterfactual_babylm_aann_all_det_removal
Dataset Card for "counterfactual_babylm_aann_all_det_removal"
More Information needed
babylm-opensub-bilingual-50M
babylm-opensub-bilingual-50M
Bilingual OpenSubtitles training corpora (50M EN anchor, byte-premium partners).
Interleaved sentence streams for aligned/offset mixing modes.
Configs
en_nld_aligned
pair: en-nl / mode: aligned
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14607622
en_nld_offset
pair: en-nl / mode: offset
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14570422… See the full description on the dataset page: https://huggingface.co/datasets/zhzh98/babylm-opensub-bilingual-50M.babylm-zho-100M
babylm-zho-100M
A filtered version of BabyLM-community/babylm-zho, a Chinese-language corpus designed for the BabyLM Challenge.
This is the official training data for Chinese BabyLM Challenge.
Size
The filtered dataset contains approximately 101,343,320 tokens (tokenized with jieba).
Modifications
The original babylm-zho dataset was filtered to reduce the proportion of speech-derived text. Specifically, 1/2 of the entries sourced from WenetSpeech… See the full description on the dataset page: https://huggingface.co/datasets/chinese-babylm-org/babylm-zho-100M.region-embeddings
region-embeddings
Per-region visual features for the Augustinian BabyLM project: every word-region
pair from the grounding data, encoded with a frozen vision model.
Grounding comes from Flickr30k Entities, RefCOCO+, RefCOCOg and THINGS, together
563k region annotations. Each region is cropped and encoded; the region feature is
the mean of the encoder's patch embeddings inside the bounding box. Three encoders
are provided (DINOv3, iBOT ViT-B/16, SAM ViT-B), all trained on images… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/region-embeddings.babylmATTENTION
This is preprocessed data for the BabyLM challenge https://babylm.github.io/
If you want the raw unprocessed files, you should download them directly.
babylm-10M
BabyLM 10M
This curated dataset is originally from the BabyLM Challenge.
It consists of ~10M words of mixed domain, consisting of the following sources:
CHILDES (child-directed speech)
Subtitles (speech)
BNC (speech)
TED talks (speech)
children's books (simple written language)
vpswap-checkpoint-scores
VP-Swap checkpoint scores
Per-item correctness on the VP-Swap benchmark for nine models at twenty
points in training. This is the raw material behind Figures 4-6 of
Augustinian BabyLM (paper, code).
Layout
<model>/<revision>.jsonl, one line per benchmark item:
{"property": "color", "line": 6, "which": 1,
"pll_orig": -21.4213, "pll_swap": -23.9077, "correct": true}
property and line identify the source line in
eval/vpswap_bb24/vp_swap_<property>_pairs.txt; which… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/vpswap-checkpoint-scores.babylm2-clean-spacybabylm-deu
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: deu
Script: Latn
Tier: 100M
Byte Premium Factor: 1.053648
Size (MB): 568.98
Expected Size (MB): 572.13
Number of Documents: 36,550
Total Tokens: 107,910,839
Tokenizer: separate by whitespace
Tokens Per Category
child-available-speech: 1,267,991 tokens
child-books: 2,096,048… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-deu.babylm_10Mcounterfactual_babylm_pipps_and_keys_to_it_all_10k
Dataset Card for "counterfactual_babylm_pipps_and_keys_to_it_all_10k"
More Information needed
slightly-cleaner-babylm
