babylm-anon/babylm_2024_10m_curriculum
Dataset Card for BabyLM 2024 10M Curriculum The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset. The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt). Pretraining split (train) Stage Words Documents C1: Child Directed Speech 2839591 28.53% 580000 49.19% C2: Unscripted Dialogue 1079286 10.84% 108000 9.16%… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.
Dataset Card for BabyLM 2024 10M Curriculum
The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset.
The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt).
Pretraining split (train)
Validation split (validation)
Dataset Sources
- C1: Child Directed Speech
- CHILDES (MacWhinney 2000)
- C2: Unscripted Dialogue
- Switchboard Dialog Act Corpus (Stolcke et al. 2000)
- British NationalCorpus (BNC), dialogue portion
- C3: Scripted Dialogue
- OpenSubtitles (Lison and Tiedemann 2016)
- C4: Wiki
- Simple Wiki (Warstadt et al. 2023)
- C5: Written English
- Standardized Project Gutenberg Corpus (Gerlach and Font-Clos 2018)
Uses
To be used for pretraining experiments with a custom Trainer that samples in a predetermined order.
