datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
IPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.ipa-phonemes-word-pairs
license: cc-by-sa 4.0
size: ~275k pairs, ~7mb (~4mb parquet)
generated using: phonemizer/espeak
check out openphonemizer for more details!
celex-en-phoneme-sample
CELEX English Phoneme Sample (gated)
A small sample derived from the CELEX Lexical Database of English
(LDC96L14),
used by the phoneme-entropy
library for phoneme-level entropy and informativity analysis.
Gated / restricted. CELEX is licensed by the Linguistic Data Consortium.
This mirror is permissions-only: access is granted manually to users who
already hold a valid CELEX/LDC license for LDC96L14.
Contents
A single CSV, celex_en_sample.csv, with columns:… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/celex-en-phoneme-sample.phoneme-pseudolabels-csvchild-asr-phoneme-pseudolabels
