CoolFace
Datasetpublic

phonemetransformers/IPA-CHILDES

IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.

sourceHugging Faceupdated 1y agoView on Hugging Face
7likes295downloads
Dataset Card

IPA-CHILDES Dataset

This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.

Description

Key Columns

The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:

ColumnDescription
processed_glossThe pre-processed orthographic utterance. This includes lowercasing, fixing English spelling and adding punctuation marks. This is based on the AOChildes preprocessing.
ipa_transcriptionA phonemic transcription of the utterance, space-separated with word boundaries marked with the WORD_BOUNDARY token.
character_split_utteranceA space-separated transcription of the utterance, produced simply by splitting the processed gloss by character. This is intended to have a very similar format to ipa_transcription for studies comparing phonetic to orthographic transcriptions.
is_childWhether the utterance was spoken by a child or not. Note that this is set to False for all utterances in this dataset, but the processing script has the ability to preserve child utterances.

character_split_utterance and ipa_transcription are designed for training character-based (phoneme-based) language models using a simple tokenizer that splits around whitespace. The processed_gloss column is suitable for word-based (or subword-based) language models with standard tokenizers.

Note that the data has been sorted by the target_child_age column, which stores child age in months. This can be used to limit the training data according to a maximum child age.

Dataset Sections

The following languages are included (ordered by number of phonemes):

LanguageDescriptionPhoible Inventory IDSpeakersUtterancesWordsPhonemes% Child
EnglishNATaken from 49 corpora in the EnglishNA collection of CHILDES and phonemized using phonemizer with language code en-us.21753,6872,564,6149,993,74430,986,21835.83
EnglishUKTaken from 16 corpora in the EnglishUK collection of CHILDES and phonemized using phonemizer with language code en-gb.22528692,043,1157,147,54121,589,84239.00
GermanTaken from 10 corpora in the German collection of CHILDES and phonemized using epitran with language code deu-Latn.23988291,525,5595,825,16621,442,57643.61
JapaneseTaken from 11 corpora in the Japanese collection of CHILDES and phonemized using phonemizer with language code ja.2196489998,6422,970,67411,985,72944.20
IndonesianTaken from 1 corpora in the EastAsian/Indonesian collection of CHILDES and phonemized using epitran with language code ind-Latn.1690438813,7952,347,6429,370,98334.32
FrenchTaken from 15 corpora in the French collection of CHILDES and phonemized using phonemizer with language code fr-fr.22691,277721,1212,973,3188,203,64940.07
SpanishTaken from 18 corpora in the Spanish collection of CHILDES and phonemized using epitran with language code spa-Latn.1641,009533,3082,183,9927,742,55045.93
MandarinTaken from 16 corpora in the Chinese/Mandarin collection of CHILDES and phonemized using pinyin_to_ipa with language code mandarin.24572,118530,0222,264,1986,605,91338.88
DutchTaken from 5 corpora in the DutchAfricaans/Dutch collection of CHILDES and phonemized using phonemizer with language code nl.2405107403,4721,475,1744,786,80335.08
PolishTaken from 2 corpora in the Slavic/Polish collection of CHILDES and phonemized using phonemizer with language code pl.1046511218,8601,042,8414,361,79763.26
SerbianTaken from 1 corpora in the Slavic/Serbian collection of CHILDES and phonemized using epitran with language code srp-Latn.2499208319,3051,052,3373,841,60029.14
EstonianTaken from 9 corpora in the Other/Estonian collection of CHILDES and phonemized using phonemizer with language code et.2181157186,921843,1893,429,22844.71
WelshTaken from 2 corpora in the Celtic/Welsh collection of CHILDES and phonemized using phonemizer with language code cy.2406269181,292666,3501,939,28669.18
CantoneseTaken from 2 corpora in the Chinese/Cantonese collection of CHILDES and phonemized using pingyam with language code cantonese.230995205,729777,9971,864,77133.54
SwedishTaken from 3 corpora in the Scandinavian/Swedish collection of CHILDES and phonemized using phonemizer with language code sv.115041154,064581,4511,782,69244.63
PortuguesePtTaken from 4 corpora in the Romance/Portuguese collection of CHILDES and phonemized using phonemizer with language code pt.220645134,543499,5221,538,40839.47
KoreanTaken from 3 corpora in the EastAsian/Korean collection of CHILDES and phonemized using phonemizer with language code ko.423127105,281263,0301,345,27636.76
ItalianTaken from 5 corpora in the Romance/Italian collection of CHILDES and phonemized using phonemizer with language code it.114510994,361352,8611,309,48939.02
CroatianTaken from 1 corpora in the Slavic/Croatian collection of CHILDES and phonemized using epitran with language code hrv-Latn.11395490,992305,1121,109,69639.24
CatalanTaken from 6 corpora in the Romance/Catalan collection of CHILDES and phonemized using phonemizer with language code ca.255518089,103319,7261,084,59436.49
IcelandicTaken from 2 corpora in the Scandinavian/Icelandic collection of CHILDES and phonemized using phonemizer with language code is.25681778,181279,9391,057,23535.21
BasqueTaken from 2 corpora in the Other/Basque collection of CHILDES and phonemized using phonemizer with language code eu.216128671,537230,500942,72548.82
HungarianTaken from 3 corpora in the Other/Hungarian collection of CHILDES and phonemized using epitran with language code hun-Latn.219111669,690237,062918,00247.95
DanishTaken from 1 corpora in the Scandinavian/Danish collection of CHILDES and phonemized using phonemizer with language code da.22652984,019275,170824,31441.71
NorwegianTaken from 2 corpora in the Scandinavian/Norwegian collection of CHILDES and phonemized using phonemizer with language code nb.4993461,906227,856729,64942.58
PortugueseBrTaken from 2 corpora in the Romance/Portuguese collection of CHILDES and phonemized using phonemizer with language code pt-br.220733122,439174,845577,86544.42
RomanianTaken from 3 corpora in the Romanian collection of CHILDES and phonemized using phonemizer with language code ro.24433354,982152,465537,66942.62
TurkishTaken from 2 corpora in the Other/Turkish collection of CHILDES and phonemized using phonemizer with language code tr.221711829,31779,404421,12950.58
IrishTaken from 2 corpora in the Celtic/Irish collection of CHILDES and phonemized using phonemizer with language code ga.25212927,818105,867338,42534.37
QuechuaTaken from 2 corpora in the Other/Quechua collection of CHILDES and phonemized using phonemizer with language code qu.1041422,39746,848281,47840.06
FarsiTaken from 2 corpora in the Other/Farsi collection of CHILDES and phonemized using phonemizer with language code fa-latn.5162922,61343,432178,52340.45

Papers

This dataset has been used in the following key papers: