CoolFace
14 results

childes

fdemelo /ipa-childes-split IPA-CHILDES split This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular, the following changes have been implemented: column processed_gloss dropped as it duplicates information of gloss up to punctuation column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+) column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.tabular10M<n<100M0 likes714 downloads1y agoHugging Faceclimb-mao /MAO-CHILDESMultilingual CHILDES Dataset for Pretraining Small Multilingual BabyLMs in Salhan et al (2024) arxiv.org/abs/2410.22886 0 likes539 downloads1y agoHugging Facew-nicole /childes_data_no_tagstext1M<n<10M0 likes380 downloads5y agoHugging Facephonemetransformers /IPA-CHILDES IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.tabular10M<n<100M7 likes297 downloads1y agoHugging Facew-nicole /childes_datatext1M<n<10M0 likes261 downloads5y agoHugging Facew-nicole /childes_data_with_tags_text1M<n<10M0 likes248 downloads5y agoHugging Face