CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01babylm-anon /stratified_10m_curriculum Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.texttext-generation1M<n<10M0 likes1.3k downloads1y agoHugging Face02babylm-anon /babylm_2024_10m_curriculum Dataset Card for BabyLM 2024 10M Curriculum The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset. The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt). Pretraining split (train) Stage Words Documents C1: Child Directed Speech 2839591 28.53% 580000 49.19% C2: Unscripted Dialogue 1079286 10.84% 108000 9.16% C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.texttext-generation1M<n<10M0 likes1.1k downloads1y agoHugging Face03BabyLM-community /BabyLM-2026-Strict-Small Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 Strict-Small training set. Total: 10M tokens. Please cite the following: @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict-Small.text1M<n<10M3 likes863 downloads6mo agoHugging Face04phonemetransformers /IPA-BabyLM Phonemized BabyLM Pre-training Data This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available. The scripts used to produce the dataset are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here. text10M<n<100M2 likes558 downloads1y agoHugging Face05chinese-babylm-org /zhoblimptext10K<n<100K0 likes551 downloads5mo agoHugging Face06BabyLM-community /BabyLM-2026-Strict Detoxified 100M BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 strict training set. Total: 100M tokens. Please cite the following: @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/BabyLM-2026-Strict.text10M<n<100M7 likes438 downloads6mo agoHugging Face07babylm-anon /stratified_equitoken_10m_curriculumtext100K<n<1M0 likes273 downloads1y agoHugging Face08BabyLM-community /formatted-CHILDEStext10K<n<100K0 likes237 downloads1y agoHugging Face09BabyLM-community /BabyLM-BLIMP-Filteredtext10K<n<100K0 likes163 downloads5mo agoHugging Face10openhonest /babylm-2026-fr-92m-5seed-resultstabularn<1K0 likes124 downloads2mo agoHugging Face11climb-mao /Spanish-BabyLMtext10K<n<100K1 likes123 downloads1y agoHugging Face12nilq /babylm-100M BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language) text10M<n<100M0 likes112 downloads3y agoHugging Face13BabyLM-community /babylm-xhogated babylm-xho Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: xho Script: Latin Number of Documents: 12352 Total Tokens: 664966 Tokens Per Category child-books: 98144 tokens educational: 65208 tokens padding-mt: 60511 tokens padding-wikipedia: 387662 tokens qed: 29099 tokens simplified-text: 24342 tokens Data Fields text: The document text category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.texttext-generation10K<n<100K0 likes108 downloads11mo agoHugging Face14bbunzeck /babylm-german German BabyLM dataset This is a pre-training dataset for training developmentally plausible language models in German (also called BabyLMs), compiled by the Computational Linguistics Group (CLAUSE) at Bielefeld University. If you are looking for ways to evaluate your German BabyLMs, we recommend our own lexical decision dataset, CLAMS for syntactic evaluation and XCOMPS for conceptual semantics/world knowledge. The composition is inspired by the original, English BabyLM dataset (see… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/babylm-german.text1M<n<10M2 likes105 downloads10mo agoHugging Face15chinese-babylm-org /hanzi-pinyintext1K<n<10K0 likes88 downloads4mo agoHugging Face16chinese-babylm-org /hanzi-structuretext1K<n<10K0 likes85 downloads4mo agoHugging Face17Lanni-ni /babylm-surprisal-resultstext1M<n<10M0 likes82 downloads10mo agoHugging Face18kanishka /babylm1sentstext10M<n<100M0 likes79 downloads7mo agoHugging Face19kanishka /counterfactual_babylm_aann_all_det_removal Dataset Card for "counterfactual_babylm_aann_all_det_removal" More Information needed text10M<n<100M0 likes75 downloads3y agoHugging Face20zhzh98 /babylm-opensub-bilingual-50M babylm-opensub-bilingual-50M Bilingual OpenSubtitles training corpora (50M EN anchor, byte-premium partners). Interleaved sentence streams for aligned/offset mixing modes. Configs en_nld_aligned pair: en-nl / mode: aligned pairs selected: 7303811 anchor units: 55796159 mixed sentences: 14607622 en_nld_offset pair: en-nl / mode: offset pairs selected: 7303811 anchor units: 55796159 mixed sentences: 14570422… See the full description on the dataset page: https://huggingface.co/datasets/zhzh98/babylm-opensub-bilingual-50M.tabular10M<n<100M0 likes75 downloads3mo agoHugging Face21chinese-babylm-org /babylm-zho-100M babylm-zho-100M A filtered version of BabyLM-community/babylm-zho, a Chinese-language corpus designed for the BabyLM Challenge. This is the official training data for Chinese BabyLM Challenge. Size The filtered dataset contains approximately 101,343,320 tokens (tokenized with jieba). Modifications The original babylm-zho dataset was filtered to reduce the proportion of speech-derived text. Specifically, 1/2 of the entries sourced from WenetSpeech… See the full description on the dataset page: https://huggingface.co/datasets/chinese-babylm-org/babylm-zho-100M.text100K<n<1M4 likes71 downloads5mo agoHugging Face22augustinian-babylm /region-embeddings region-embeddings Per-region visual features for the Augustinian BabyLM project: every word-region pair from the grounding data, encoded with a frozen vision model. Grounding comes from Flickr30k Entities, RefCOCO+, RefCOCOg and THINGS, together 563k region annotations. Each region is cropped and encoded; the region feature is the mean of the encoder's patch embeddings inside the bounding box. Three encoders are provided (DINOv3, iBOT ViT-B/16, SAM ViT-B), all trained on images… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/region-embeddings.text1M<n<10M0 likes71 downloads16d agoHugging Face23vesteinn /babylmATTENTION This is preprocessed data for the BabyLM challenge https://babylm.github.io/ If you want the raw unprocessed files, you should download them directly. text10M<n<100M5 likes68 downloads3y agoHugging Face24nilq /babylm-10M BabyLM 10M This curated dataset is originally from the BabyLM Challenge. It consists of ~10M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language) text1M<n<10M1 likes66 downloads3y agoHugging Face25augustinian-babylm /vpswap-checkpoint-scores VP-Swap checkpoint scores Per-item correctness on the VP-Swap benchmark for nine models at twenty points in training. This is the raw material behind Figures 4-6 of Augustinian BabyLM (paper, code). Layout <model>/<revision>.jsonl, one line per benchmark item: {"property": "color", "line": 6, "which": 1, "pll_orig": -21.4213, "pll_swap": -23.9077, "correct": true} property and line identify the source line in eval/vpswap_bb24/vp_swap_<property>_pairs.txt; which… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/vpswap-checkpoint-scores.tabular1M<n<10M0 likes64 downloads16d agoHugging Face26kanishka /babylm2-clean-spacytext10M<n<100M0 likes61 downloads1y agoHugging Face27BabyLM-community /babylm-deugated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: deu Script: Latn Tier: 100M Byte Premium Factor: 1.053648 Size (MB): 568.98 Expected Size (MB): 572.13 Number of Documents: 36,550 Total Tokens: 107,910,839 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 1,267,991 tokens child-books: 2,096,048… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-deu.texttext-generation10K<n<100K3 likes60 downloads1y agoHugging Face28omarmomen /babylm_10Mtext1M<n<10M0 likes58 downloads3y agoHugging Face29kanishka /counterfactual_babylm_pipps_and_keys_to_it_all_10k Dataset Card for "counterfactual_babylm_pipps_and_keys_to_it_all_10k" More Information needed text10M<n<100M0 likes56 downloads3y agoHugging Face30qing-yao /slightly-cleaner-babylmtext10M<n<100M0 likes56 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.