babylm_ipa
IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
IPA-BabyLM-evaluation
BabyLM 2024 evaluation data in IPA
A version of the BabyLM 2024 evalution data converted to IPA using G2P+. Scripts for producing this data are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes.
IPA-BabyLM-2026
