phonemetransformers/IPA-CHILDES
IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.
Update README.md
Update README.md
Update README.md
Update README
Update column order and remove language_code
Update README.md
Fix Mandarin
Fix German
Remove train-valid split
Slight updates
Add child utterances
Fix double space issue
Update README
Update Cantonese, German, Norwegian and Romanian
Update Mandarin
Update all languages with use of folding
Add Polish and Serbian
Fix gloss processing
Update Cantonese and Mandarin to combine tone with vowel
Add UK English
Add Catalan, Irish, Italian, Korean, Portuguese, Quechua, Romanian, Swedish and Welsh
New English as default
Fix README
Add alternative English with BabySLM phoneme set
Update German and Indonesian after analysis
Fix Turkish target_child_sex column loading issue
Fix Dutch target_child_sex column loading issue
Add Basque, Cantonese, Croatian, Danish, Dutch, Estonian, Farsi, Hungarian, Icelandic, Indonesian, Japanese, Mandarin, Turkish
Fix German
Remove test sets and use constant size for validation
Update English and German
Add Spanish
Add English, French, German and update README
Initial commit
initial commit
