CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cambridge-climb /BabyLMDataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.10M<n<100M3 likes2.8k downloads2y agoHugging Face02babylm-anon /stratified_10m_curriculum Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.texttext-generation1M<n<10M0 likes1.3k downloads1y agoHugging Face03babylm-anon /babylm_2024_10m_curriculum Dataset Card for BabyLM 2024 10M Curriculum The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset. The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt). Pretraining split (train) Stage Words Documents C1: Child Directed Speech 2839591 28.53% 580000 49.19% C2: Unscripted Dialogue 1079286 10.84% 108000 9.16% C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.texttext-generation1M<n<10M0 likes1.1k downloads1y agoHugging Face04BabyLM-community /BabyLM-2026-Strict-Small Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 Strict-Small training set. Total: 10M tokens. Please cite the following: @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.text1M<n<10M3 likes863 downloads6mo agoHugging Face05phonemetransformers /IPA-BabyLM Phonemized BabyLM Pre-training Data This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available. The scripts used to produce the dataset are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here. text10M<n<100M2 likes558 downloads1y agoHugging Face06chinese-babylm-org /zhoblimptext10K<n<100K0 likes551 downloads5mo agoHugging Face07BabyLM-community /BabyLM-2026-Strict Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 strict training set. Total: 100M tokens. Please cite the following: @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.text10M<n<100M7 likes438 downloads6mo agoHugging Face08augustinian-babylm /synthetic-grounding-images Synthetic grounding images 3,162 images generated to extend visual grounding to concrete words that no photograph dataset covers, for Augustinian BabyLM (paper, code). How they were made Starting from 1,986 concrete words with no image support, an LLM (claude-sonnet-4-6, temperature 0.8) wrote short scene descriptions placing as many target words as fit naturally into one scene. Each of the 1,054 resulting descriptions was rendered three times with SDXL-Turbo (2… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/synthetic-grounding-images.image1K<n<10K0 likes319 downloads16d agoHugging Face09babylm-anon /stratified_equitoken_10m_curriculumtext100K<n<1M0 likes273 downloads1y agoHugging Face10BabyLM-community /formatted-CHILDEStext10K<n<100K0 likes237 downloads1y agoHugging Face11BabyLM-community /BabyLM-2026-Strict-Evals0 likes209 downloads5mo agoHugging Face12phonemetransformers /IPA-BabyLM-evaluation BabyLM 2024 evaluation data in IPA A version of the BabyLM 2024 evalution data converted to IPA using G2P+. Scripts for producing this data are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. 0 likes208 downloads1y agoHugging Face13augustinian-babylm /token-embeddings token-embeddings Per-token visual embedding tables for the Augustinian BabyLM project: [V, 768] float32 matrices used to initialize the input embedding matrix of a DeBERTa-v3-base masked LM before text training. Organized as <encoder>/<vocab>/, for encoder in dinov3 / sam / ibot and vocab in 50k / 75k / 100k. Each directory holds E_init.safetensors (the table) and a seeded_mask marking which rows carry visual information, roughly 24-38% of rows depending on vocabulary size.… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/token-embeddings.tabular100K<n<1M0 likes196 downloads16d agoHugging Face14BabyLM-community /BabyLM-BLIMP-Filteredtext10K<n<100K0 likes163 downloads5mo agoHugging Face15BabyLM-community /BabyLM-dev0 likes133 downloads6mo agoHugging Face16openhonest /babylm-2026-fr-92m-5seed-resultstabularn<1K0 likes124 downloads2mo agoHugging Face17climb-mao /Spanish-BabyLMtext10K<n<100K1 likes123 downloads1y agoHugging Face18nilq /babylm-100M BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language) text10M<n<100M0 likes112 downloads3y agoHugging Face19BabyLM-community /babylm-xhogated babylm-xho Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: xho Script: Latin Number of Documents: 12352 Total Tokens: 664966 Tokens Per Category child-books: 98144 tokens educational: 65208 tokens padding-mt: 60511 tokens padding-wikipedia: 387662 tokens qed: 29099 tokens simplified-text: 24342 tokens Data Fields text: The document text category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.texttext-generation10K<n<100K0 likes108 downloads11mo agoHugging Face20bbunzeck /babylm-german German BabyLM dataset This is a pre-training dataset for training developmentally plausible language models in German (also called BabyLMs), compiled by the Computational Linguistics Group (CLAUSE) at Bielefeld University. If you are looking for ways to evaluate your German BabyLMs, we recommend our own lexical decision dataset, CLAMS for syntactic evaluation and XCOMPS for conceptual semantics/world knowledge. The composition is inspired by the original, English BabyLM dataset (see… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/babylm-german.text1M<n<10M2 likes105 downloads10mo agoHugging Face21chinese-babylm-org /hanzi-pinyintext1K<n<10K0 likes88 downloads4mo agoHugging Face22SrikrishnaIyer /Babylm-processed-2023 Dataset Preprocessing for 10M and 100M Text-Only Tracks Overview This document describes the preprocessing steps applied to the datasets used for the 10M and 100M text-only tracks. The datasets are a mixture of 10 different corpora, as shown in Table 1 below. Table 1: Dataset Contents Dataset Domain # Words STRICT-SMALL # Words STRICT Proportion CHILDES (MacWhinney, 2000) Child-directed speech 0.44M 4.21M 5% British National Corpus (BNC), dialogue… See the full description on the dataset page: https://huggingface.co/datasets/SrikrishnaIyer/Babylm-processed-2023.0 likes86 downloads2y agoHugging Face23chinese-babylm-org /hanzi-structuretext1K<n<10K0 likes85 downloads4mo agoHugging Face24Lanni-ni /babylm-surprisal-resultstext1M<n<10M0 likes82 downloads10mo agoHugging Face25BabyLM-community /BabyLM-Test0 likes80 downloads6mo agoHugging Face26kanishka /babylm1sentstext10M<n<100M0 likes79 downloads7mo agoHugging Face27kanishka /counterfactual_babylm_aann_all_det_removal Dataset Card for "counterfactual_babylm_aann_all_det_removal" More Information needed text10M<n<100M0 likes75 downloads3y agoHugging Face28zhzh98 /babylm-opensub-bilingual-50M babylm-opensub-bilingual-50M Bilingual OpenSubtitles training corpora (50M EN anchor, byte-premium partners). Interleaved sentence streams for aligned/offset mixing modes. Configs en_nld_aligned pair: en-nl / mode: aligned pairs selected: 7303811 anchor units: 55796159 mixed sentences: 14607622 en_nld_offset pair: en-nl / mode: offset pairs selected: 7303811 anchor units: 55796159 mixed sentences: 14570422… See the full description on the dataset page: https://huggingface.co/datasets/zhzh98/babylm-opensub-bilingual-50M.tabular10M<n<100M0 likes75 downloads3mo agoHugging Face29miguelcsx /babylm-wordlevel-16k-structured-priors-v20 likes75 downloads2mo agoHugging Face30chinese-babylm-org /babylm-zho-100M babylm-zho-100M A filtered version of BabyLM-community/babylm-zho, a Chinese-language corpus designed for the BabyLM Challenge. This is the official training data for Chinese BabyLM Challenge. Size The filtered dataset contains approximately 101,343,320 tokens (tokenized with jieba). Modifications The original babylm-zho dataset was filtered to reduce the proportion of speech-derived text. Specifically, 1/2 of the entries sourced from WenetSpeech… See the full description on the dataset page: https://huggingface.co/datasets/chinese-babylm-org/babylm-zho-100M.text100K<n<1M4 likes71 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.