CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01babylm-anon /stratified_10m_curriculum Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.texttext-generation1M<n<10M0 likes1.3k downloads1y agoHugging Face02babylm-anon /babylm_2024_10m_curriculum Dataset Card for BabyLM 2024 10M Curriculum The documents from the 10M dataset provided by the 2024 BabyLM challange. We add a validation split we with additional documents from the 100M dataset. The order in which to load this dataset is provided as .pt files (e.g., curriculum.pt or random.pt). Pretraining split (train) Stage Words Documents C1: Child Directed Speech 2839591 28.53% 580000 49.19% C2: Unscripted Dialogue 1079286 10.84% 108000 9.16% C3:… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/babylm_2024_10m_curriculum.texttext-generation1M<n<10M0 likes1.1k downloads1y agoHugging Face03BabyLM-community /babylm-xhogated babylm-xho Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: xho Script: Latin Number of Documents: 12352 Total Tokens: 664966 Tokens Per Category child-books: 98144 tokens educational: 65208 tokens padding-mt: 60511 tokens padding-wikipedia: 387662 tokens qed: 29099 tokens simplified-text: 24342 tokens Data Fields text: The document text category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-xho.texttext-generation10K<n<100K0 likes108 downloads11mo agoHugging Face04BabyLM-community /babylm-deugated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: deu Script: Latn Tier: 100M Byte Premium Factor: 1.053648 Size (MB): 568.98 Expected Size (MB): 572.13 Number of Documents: 36,550 Total Tokens: 107,910,839 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 1,267,991 tokens child-books: 2,096,048… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-deu.texttext-generation10K<n<100K3 likes60 downloads1y agoHugging Face05pulipakav-1 /translated-babylm-telugu Translated BabyLM — Telugu (translated-babylm-telugu) Dataset Description This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.texttext-generation10M<n<100M0 likes52 downloads5mo agoHugging Face06BabyLM-community /babylm-ar-subtitles babylm-ara Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: ara Script: Unknown Number of Documents: 65951 Total Tokens: 399142332 Tokens Per Category subtitles: 399142332 tokens Data Fields text: The document text doc_id: Unique identifier for the document category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data script:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ar-subtitles.texttext-generation10K<n<100K0 likes42 downloads1y agoHugging Face07pulipakav-1 /translated-babylm-hindi Translated BabyLM — Hindi (translated-babylm-hindi) Dataset Description This dataset is a Hindi translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Hindi, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-hindi.texttext-generation10M<n<100M0 likes42 downloads5mo agoHugging Face08BabyLM-community /babylm-nldgated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: nld Script: Latn Tier: 100M Byte Premium Factor: 1.051606 Size (MB): 569.49 Expected Size (MB): 571.02 Number of Documents: 304,611 Total Tokens: 109,885,564 Tokenizer: separate by whitespace Tokens Per Category child-books: 4,576,823 tokens child-directed-speech: 3,304,756… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nld.texttext-generation100K<n<1M0 likes37 downloads1y agoHugging Face09BabyLM-community /babylm-bg-subtitles babylm-bg Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: bg Script: Cyrillic Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 43179 Total Tokens: 277270105 Tokens Per Category subtitles: 277270105 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-bg-subtitles.texttext-generation10K<n<100K0 likes37 downloads1y agoHugging Face10BabyLM-community /babylm-enggated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: eng Script: Latn Tier: 100M Byte Premium Factor: 1.000000 Size (MB): 539.18 Expected Size (MB): 543.00 Number of Documents: 137,710 Total Tokens: 98,878,321 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 9,102,166 tokens child-books: 26,748,028… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-eng.texttext-generation100K<n<1M3 likes37 downloads1y agoHugging Face11BabyLM-community /babylm-zhogated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: zho Script: Hani, Hans, Latn Tier: > 100M Byte Premium Factor: 0.935966 Size (MB): 518.85 Expected Size (MB): 508.23 Number of Documents: 203,891 Total Tokens: 137,835,046 Tokenizer: Qwen/Qwen3-0.6B Tokens Per Category child-available-speech: 7,403,441 tokens child-books: 15… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.texttext-generation100K<n<1M1 likes36 downloads8mo agoHugging Face12BabyLM-community /babylm-pt-subtitles babylm-pt Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: pt Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 50869 Total Tokens: 356455068 Tokens Per Category subtitles: 356455068 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-pt-subtitles.texttext-generation10K<n<100K0 likes30 downloads1y agoHugging Face13BabyLM-community /babylm-fa-subtitles babylm-fa Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: fa Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 41312 Total Tokens: 249245044 Tokens Per Category subtitles: 249245044 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fa-subtitles.texttext-generation10K<n<100K0 likes28 downloads1y agoHugging Face14BabyLM-community /babylm-ellgated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: ell Script: Greek, Grek Tier: 10M Byte Premium Factor: 1.967262 Size (MB): 106.81 Expected Size (MB): 106.82 Number of Documents: 11,104 Total Tokens: 10,882,556 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 1,673,255 tokens child-books: 1,390… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ell.texttext-generation10K<n<100K0 likes27 downloads1y agoHugging Face15BabyLM-community /babylm-de-subtitles babylm-de Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: de Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 32073 Total Tokens: 224733295 Tokens Per Category subtitles: 224733295 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-de-subtitles.texttext-generation10K<n<100K0 likes26 downloads1y agoHugging Face16BabyLM-community /babylm-cy-subtitles babylm-cy Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: cy Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 61 Total Tokens: 407354 Tokens Per Category subtitles: 407354 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data script:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-cy-subtitles.texttext-generationn<1K0 likes21 downloads1y agoHugging Face17BabyLM-community /babylm-sv-subtitles babylm-sv Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: sv Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 24147 Total Tokens: 148074740 Tokens Per Category subtitles: 148074740 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-sv-subtitles.texttext-generation10K<n<100K0 likes21 downloads1y agoHugging Face18BabyLM-community /babylm-id-subtitles babylm-id Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: id Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 45408 Total Tokens: 264979169 Tokens Per Category subtitles: 264979169 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-id-subtitles.texttext-generation10K<n<100K0 likes20 downloads1y agoHugging Face19BabyLM-community /babylm-ko-subtitles babylm-ko Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: ko Script: Korean (Hangul + Han) Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 7909 Total Tokens: 34475400 Tokens Per Category subtitles: 34475400 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ko-subtitles.texttext-generation1K<n<10K0 likes20 downloads1y agoHugging Face20BabyLM-community /babylm-hr-subtitles babylm-hr Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: hr Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 43519 Total Tokens: 294492455 Tokens Per Category subtitles: 294492455 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-hr-subtitles.texttext-generation10K<n<100K0 likes20 downloads1y agoHugging Face21BabyLM-community /babylm-sr-subtitles babylm-sr Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: sr Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 70570 Total Tokens: 473404509 Tokens Per Category subtitles: 473404509 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-sr-subtitles.texttext-generation10K<n<100K0 likes20 downloads1y agoHugging Face22BabyLM-community /babylm-fasgated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: fas Script: Arab Tier: 100M Byte Premium Factor: 1.597326 Size (MB): 867.30 Expected Size (MB): 867.35 Number of Documents: 217,776 Total Tokens: 98,506,081 Tokenizer: separate by whitespace Tokens Per Category child-books: 67,165 tokens educational: 94,320,928 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fas.texttext-generation100K<n<1M0 likes20 downloads1y agoHugging Face23BabyLM-community /babylm-nl-subtitles babylm-nl Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: nl Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 46973 Total Tokens: 346620742 Tokens Per Category subtitles: 346620742 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nl-subtitles.texttext-generation10K<n<100K0 likes19 downloads1y agoHugging Face24BabyLM-community /babylm-hebgated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: heb Script: Hebr Tier: 1M Byte Premium Factor: 1.355477 Size (MB): 7.37 Expected Size (MB): 7.36 Number of Documents: 210 Total Tokens: 818,910 Tokenizer: separate by whitespace Tokens Per Category child-directed-speech: 309,854 tokens padding-wikipedia: 509,056 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-heb.texttext-generationn<1K0 likes19 downloads1y agoHugging Face25BabyLM-community /babylm-ro-subtitles babylm-ro Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: ro Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 114687 Total Tokens: 770426959 Tokens Per Category subtitles: 770426959 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ro-subtitles.texttext-generation100K<n<1M0 likes19 downloads1y agoHugging Face26BabyLM-community /babylm-ja-subtitles babylm-ja Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: ja Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 1047 Total Tokens: 8742849 Tokens Per Category subtitles: 8742849 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ja-subtitles.texttext-generation1K<n<10K0 likes18 downloads1y agoHugging Face27BabyLM-community /babylm-fragated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: fra Script: Latn Tier: 100M Byte Premium Factor: 1.173979 Size (MB): 634.88 Expected Size (MB): 637.47 Number of Documents: 81,950 Total Tokens: 126,580,785 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 1,989,852 tokens child-books: 1,244,842… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fra.texttext-generation10K<n<100K0 likes15 downloads1y agoHugging Face28BabyLM-community /babylm-is-subtitles babylm-is Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: is Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 2550 Total Tokens: 19122537 Tokens Per Category subtitles: 19122537 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-is-subtitles.texttext-generation1K<n<10K0 likes14 downloads1y agoHugging Face29BabyLM-community /babylm-pl-subtitles babylm-pl Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: pl Script: Latin Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 86519 Total Tokens: 488352413 Tokens Per Category subtitles: 488352413 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-pl-subtitles.texttext-generation10K<n<100K0 likes14 downloads1y agoHugging Face30BabyLM-community /babylm-uk-subtitles babylm-uk Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: uk Script: Cyrillic Category: subtitles Source: Unknown Age Estimate: n/a Number of Documents: 4762 Total Tokens: 30327514 Tokens Per Category subtitles: 30327514 tokens Data Fields text: The document text category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source of the data… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-uk-subtitles.texttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.