datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
hanzi-pinyinbabylm-10mBulgarian-BabyLMBulgarian BabyLM Dataset curated by Mila Marcheva (University of Cambridge).
28,467,275 tokens (excluding punctuation)
Dataset Overview
A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count.
Data Sourcing
We additionally release information about the source of sentence in the Bulgarian BabyLM Dataset, which you can find here:… See the full description on the dataset page: https://huggingface.co/datasets/climb-mao/Bulgarian-BabyLM.hanzi-structurebabylm-tagged-by-common-complexity-metricsbaby-wikis
