datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stratified_10m_curriculum
Dataset Card for Stratified 10M Curriculum
This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange.
Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5).
Child-directed speech accounts for nearly half of the original dataset by word count.
In preliminary experiments using a training data influence estimation method, this category was by far the most influential.
This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.babylm_2024_10m_curriculum
Dataset Card for BabyLM 2024 10M Curriculum
The documents from the 10M dataset provided by the 2024 BabyLM challange.
We add a validation split we with additional documents from the 100M dataset.
The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt).
Pretraining split (train)
Stage
Words
Documents
C1: Child Directed Speech
2839591
28.53%
580000
49.19%
C2: Unscripted Dialogue
1079286
10.84%
108000
9.16%
C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.babylm-xho
babylm-xho
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: xho
Script: Latin
Number of Documents: 12352
Total Tokens: 664966
Tokens Per Category
child-books: 98144 tokens
educational: 65208 tokens
padding-mt: 60511 tokens
padding-wikipedia: 387662 tokens
qed: 29099 tokens
simplified-text: 24342 tokens
Data Fields
text: The document text
category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.babylm-deu
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: deu
Script: Latn
Tier: 100M
Byte Premium Factor: 1.053648
Size (MB): 568.98
Expected Size (MB): 572.13
Number of Documents: 36,550
Total Tokens: 107,910,839
Tokenizer: separate by whitespace
Tokens Per Category
child-available-speech: 1,267,991 tokens
child-books: 2,096,048… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-deu.translated-babylm-telugu
Translated BabyLM — Telugu (translated-babylm-telugu)
Dataset Description
This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup.
Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B)
Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.babylm-ar-subtitles
babylm-ara
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: ara
Script: Unknown
Number of Documents: 65951
Total Tokens: 399142332
Tokens Per Category
subtitles: 399142332 tokens
Data Fields
text: The document text
doc_id: Unique identifier for the document
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data
script:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ar-subtitles.translated-babylm-hindi
Translated BabyLM — Hindi (translated-babylm-hindi)
Dataset Description
This dataset is a Hindi translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Hindi, following the BabyLM challenge setup.
Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B)
Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-hindi.babylm-nld
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: nld
Script: Latn
Tier: 100M
Byte Premium Factor: 1.051606
Size (MB): 569.49
Expected Size (MB): 571.02
Number of Documents: 304,611
Total Tokens: 109,885,564
Tokenizer: separate by whitespace
Tokens Per Category
child-books: 4,576,823 tokens
child-directed-speech: 3,304,756… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nld.babylm-bg-subtitles
babylm-bg
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: bg
Script: Cyrillic
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 43179
Total Tokens: 277270105
Tokens Per Category
subtitles: 277270105 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-bg-subtitles.babylm-eng
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: eng
Script: Latn
Tier: 100M
Byte Premium Factor: 1.000000
Size (MB): 539.18
Expected Size (MB): 543.00
Number of Documents: 137,710
Total Tokens: 98,878,321
Tokenizer: separate by whitespace
Tokens Per Category
child-available-speech: 9,102,166 tokens
child-books: 26,748,028… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-eng.babylm-zho
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: zho
Script: Hani, Hans, Latn
Tier: > 100M
Byte Premium Factor: 0.935966
Size (MB): 518.85
Expected Size (MB): 508.23
Number of Documents: 203,891
Total Tokens: 137,835,046
Tokenizer: Qwen/Qwen3-0.6B
Tokens Per Category
child-available-speech: 7,403,441 tokens
child-books: 15… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.babylm-pt-subtitles
babylm-pt
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: pt
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 50869
Total Tokens: 356455068
Tokens Per Category
subtitles: 356455068 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-pt-subtitles.babylm-fa-subtitles
babylm-fa
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: fa
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 41312
Total Tokens: 249245044
Tokens Per Category
subtitles: 249245044 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fa-subtitles.babylm-ell
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: ell
Script: Greek, Grek
Tier: 10M
Byte Premium Factor: 1.967262
Size (MB): 106.81
Expected Size (MB): 106.82
Number of Documents: 11,104
Total Tokens: 10,882,556
Tokenizer: separate by whitespace
Tokens Per Category
child-available-speech: 1,673,255 tokens
child-books: 1,390… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ell.babylm-de-subtitles
babylm-de
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: de
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 32073
Total Tokens: 224733295
Tokens Per Category
subtitles: 224733295 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-de-subtitles.babylm-cy-subtitles
babylm-cy
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: cy
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 61
Total Tokens: 407354
Tokens Per Category
subtitles: 407354 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data
script:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-cy-subtitles.babylm-sv-subtitles
babylm-sv
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: sv
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 24147
Total Tokens: 148074740
Tokens Per Category
subtitles: 148074740 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-sv-subtitles.babylm-id-subtitles
babylm-id
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: id
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 45408
Total Tokens: 264979169
Tokens Per Category
subtitles: 264979169 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-id-subtitles.babylm-ko-subtitles
babylm-ko
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: ko
Script: Korean (Hangul + Han)
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 7909
Total Tokens: 34475400
Tokens Per Category
subtitles: 34475400 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ko-subtitles.babylm-hr-subtitles
babylm-hr
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: hr
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 43519
Total Tokens: 294492455
Tokens Per Category
subtitles: 294492455 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-hr-subtitles.babylm-sr-subtitles
babylm-sr
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: sr
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 70570
Total Tokens: 473404509
Tokens Per Category
subtitles: 473404509 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-sr-subtitles.babylm-fas
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: fas
Script: Arab
Tier: 100M
Byte Premium Factor: 1.597326
Size (MB): 867.30
Expected Size (MB): 867.35
Number of Documents: 217,776
Total Tokens: 98,506,081
Tokenizer: separate by whitespace
Tokens Per Category
child-books: 67,165 tokens
educational: 94,320,928 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fas.babylm-nl-subtitles
babylm-nl
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: nl
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 46973
Total Tokens: 346620742
Tokens Per Category
subtitles: 346620742 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nl-subtitles.babylm-heb
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: heb
Script: Hebr
Tier: 1M
Byte Premium Factor: 1.355477
Size (MB): 7.37
Expected Size (MB): 7.36
Number of Documents: 210
Total Tokens: 818,910
Tokenizer: separate by whitespace
Tokens Per Category
child-directed-speech: 309,854 tokens
padding-wikipedia: 509,056 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-heb.babylm-ro-subtitles
babylm-ro
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: ro
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 114687
Total Tokens: 770426959
Tokens Per Category
subtitles: 770426959 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ro-subtitles.babylm-ja-subtitles
babylm-ja
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: ja
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 1047
Total Tokens: 8742849
Tokens Per Category
subtitles: 8742849 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ja-subtitles.babylm-fra
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: fra
Script: Latn
Tier: 100M
Byte Premium Factor: 1.173979
Size (MB): 634.88
Expected Size (MB): 637.47
Number of Documents: 81,950
Total Tokens: 126,580,785
Tokenizer: separate by whitespace
Tokens Per Category
child-available-speech: 1,989,852 tokens
child-books: 1,244,842… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fra.babylm-is-subtitles
babylm-is
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: is
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 2550
Total Tokens: 19122537
Tokens Per Category
subtitles: 19122537 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-is-subtitles.babylm-pl-subtitles
babylm-pl
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: pl
Script: Latin
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 86519
Total Tokens: 488352413
Tokens Per Category
subtitles: 488352413 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-pl-subtitles.babylm-uk-subtitles
babylm-uk
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: uk
Script: Cyrillic
Category: subtitles
Source: Unknown
Age Estimate: n/a
Number of Documents: 4762
Total Tokens: 30327514
Tokens Per Category
subtitles: 30327514 tokens
Data Fields
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-uk-subtitles.
