CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01phonemetransformers /IPA-BabyLM Phonemized BabyLM Pre-training Data This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available. The scripts used to produce the dataset are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here. text10M<n<100M2 likes558 downloads1y agoHugging Face02chinese-babylm /hanzi-pinyintext1K<n<10K0 likes21 downloads6mo agoHugging Face03asparius /babylm-10mtext1M<n<10M0 likes20 downloads3y agoHugging Face04climb-mao /Bulgarian-BabyLMBulgarian BabyLM Dataset curated by Mila Marcheva (University of Cambridge). 28,467,275 tokens (excluding punctuation) Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count. Data Sourcing We additionally release information about the source of sentence in the Bulgarian BabyLM Dataset, which you can find here:… See the full description on the dataset page: https://huggingface.co/datasets/climb-mao/Bulgarian-BabyLM.text1M<n<10M2 likes15 downloads1y agoHugging Face05chinese-babylm /hanzi-structuretext1K<n<10K0 likes5 downloads6mo agoHugging Face06deru35 /babylm-tagged-by-common-complexity-metricstabular100K<n<1M0 likes3 downloads3mo agoHugging Face07BabyLM-community /baby-wikistext100K<n<1M0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.