datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BabyLMDataset for the shared baby language modeling task.
The goal is to train a language model from scratch on this data which represents
roughly the amount of text and speech data a young child observes.stratified_10m_curriculum
Dataset Card for Stratified 10M Curriculum
This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange.
Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5).
Child-directed speech accounts for nearly half of the original dataset by word count.
In preliminary experiments using a training data influence estimation method, this category was by far the most influential.
This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.babylm_2024_10m_curriculum
Dataset Card for BabyLM 2024 10M Curriculum
The documents from the 10M dataset provided by the 2024 BabyLM challange.
We add a validation split we with additional documents from the 100M dataset.
The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt).
Pretraining split (train)
Stage
Words
Documents
C1: Child Directed Speech
2839591
28.53%
580000
49.19%
C2: Unscripted Dialogue
1079286
10.84%
108000
9.16%
C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.BabyLM-2026-Strict-Small
Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 Strict-Small training set. Total: 10M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
zhoblimpBabyLM-2026-Strict
Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM)
BabyLM 2026 strict training set. Total: 100M tokens.
Please cite the following:
@misc{choshen2026babylmturns4papers,
title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop},
author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.synthetic-grounding-images
Synthetic grounding images
3,162 images generated to extend visual grounding to concrete words that no
photograph dataset covers, for Augustinian BabyLM
(paper, code).
How they were made
Starting from 1,986 concrete words with no image support, an LLM
(claude-sonnet-4-6, temperature 0.8) wrote short scene descriptions
placing as many target words as fit naturally into one scene. Each of the
1,054 resulting descriptions was rendered three times with SDXL-Turbo
(2… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/synthetic-grounding-images.stratified_equitoken_10m_curriculumformatted-CHILDESBabyLM-2026-Strict-EvalsIPA-BabyLM-evaluation
BabyLM 2024 evaluation data in IPA
A version of the BabyLM 2024 evalution data converted to IPA using G2P+. Scripts for producing this data are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes.
token-embeddings
token-embeddings
Per-token visual embedding tables for the Augustinian BabyLM project: [V, 768]
float32 matrices used to initialize the input embedding matrix of a DeBERTa-v3-base
masked LM before text training.
Organized as <encoder>/<vocab>/, for encoder in dinov3 / sam / ibot and vocab in
50k / 75k / 100k. Each directory holds E_init.safetensors (the table) and a
seeded_mask marking which rows carry visual information, roughly 24-38% of rows
depending on vocabulary size.… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/token-embeddings.BabyLM-BLIMP-FilteredBabyLM-devbabylm-2026-fr-92m-5seed-resultsSpanish-BabyLMbabylm-100M
BabyLM 100M
This curated dataset is originally from the BabyLM Challenge.
It consists of ~100M words of mixed domain, consisting of the following sources:
CHILDES (child-directed speech)
Subtitles (speech)
BNC (speech)
TED talks (speech)
children's books (simple written language)
babylm-xho
babylm-xho
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: xho
Script: Latin
Number of Documents: 12352
Total Tokens: 664966
Tokens Per Category
child-books: 98144 tokens
educational: 65208 tokens
padding-mt: 60511 tokens
padding-wikipedia: 387662 tokens
qed: 29099 tokens
simplified-text: 24342 tokens
Data Fields
text: The document text
category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.babylm-german
German BabyLM dataset
This is a pre-training dataset for training developmentally plausible language models in German (also called BabyLMs), compiled by the Computational Linguistics Group (CLAUSE) at Bielefeld University.
If you are looking for ways to evaluate your German BabyLMs, we recommend our own lexical decision dataset, CLAMS for syntactic evaluation and XCOMPS for conceptual semantics/world knowledge.
The composition is inspired by the original, English BabyLM dataset (see… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/babylm-german.hanzi-pinyinBabylm-processed-2023
Dataset Preprocessing for 10M and 100M Text-Only Tracks
Overview
This document describes the preprocessing steps applied to the datasets used for the 10M and 100M text-only tracks. The datasets are a mixture of 10 different corpora, as shown in Table 1 below.
Table 1: Dataset Contents
Dataset
Domain
# Words STRICT-SMALL
# Words STRICT
Proportion
CHILDES (MacWhinney, 2000)
Child-directed speech
0.44M
4.21M
5%
British National Corpus (BNC), dialogue… See the full description on the dataset page: https://huggingface.co/datasets/SrikrishnaIyer/Babylm-processed-2023.hanzi-structurebabylm-surprisal-resultsBabyLM-Testbabylm1sentscounterfactual_babylm_aann_all_det_removal
Dataset Card for "counterfactual_babylm_aann_all_det_removal"
More Information needed
babylm-opensub-bilingual-50M
babylm-opensub-bilingual-50M
Bilingual OpenSubtitles training corpora (50M EN anchor, byte-premium partners).
Interleaved sentence streams for aligned/offset mixing modes.
Configs
en_nld_aligned
pair: en-nl / mode: aligned
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14607622
en_nld_offset
pair: en-nl / mode: offset
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14570422… See the full description on the dataset page: https://huggingface.co/datasets/zhzh98/babylm-opensub-bilingual-50M.babylm-wordlevel-16k-structured-priors-v2babylm-zho-100M
babylm-zho-100M
A filtered version of BabyLM-community/babylm-zho, a Chinese-language corpus designed for the BabyLM Challenge.
This is the official training data for Chinese BabyLM Challenge.
Size
The filtered dataset contains approximately 101,343,320 tokens (tokenized with jieba).
Modifications
The original babylm-zho dataset was filtered to reduce the proportion of speech-derived text. Specifically, 1/2 of the entries sourced from WenetSpeech… See the full description on the dataset page: https://huggingface.co/datasets/chinese-babylm-org/babylm-zho-100M.
