CoolFace
20 results

BabyLM

cambridge-climb /BabyLMDataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.10M<n<100M3 likes2.8k downloads2y agoHugging Facebabylm-anon /stratified_10m_curriculum Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.texttext-generation1M<n<10M0 likes1.3k downloads1y agoHugging Facebabylm-anon /babylm_2024_10m_curriculum Dataset Card for BabyLM 2024 10M Curriculum The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset. The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt). Pretraining split (train) Stage Words Documents C1: Child Directed Speech 2839591 28.53% 580000 49.19% C2: Unscripted Dialogue 1079286 10.84% 108000 9.16% C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.texttext-generation1M<n<10M0 likes1.1k downloads1y agoHugging FaceBabyLM-community /BabyLM-2026-Strict-Small Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 Strict-Small training set. Total: 10M tokens. Please cite the following: @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.text1M<n<10M3 likes863 downloads6mo agoHugging Facephonemetransformers /IPA-BabyLM Phonemized BabyLM Pre-training Data This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available. The scripts used to produce the dataset are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here. text10M<n<100M2 likes558 downloads1y agoHugging Facechinese-babylm-org /zhoblimptext10K<n<100K0 likes551 downloads5mo agoHugging Face