CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fdemelo /ipa-childes-split IPA-CHILDES split This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular, the following changes have been implemented: column processed_gloss dropped as it duplicates information of gloss up to punctuation column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+) column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.tabular10M<n<100M0 likes651 downloads1y agoHugging Face02climb-mao /MAO-CHILDESMultilingual CHILDES Dataset for Pretraining Small Multilingual BabyLMs in Salhan et al (2024) arxiv.org/abs/2410.22886 0 likes540 downloads1y agoHugging Face03w-nicole /childes_data_no_tagstext1M<n<10M0 likes381 downloads5y agoHugging Face04phonemetransformers /IPA-CHILDES IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.tabular10M<n<100M7 likes295 downloads1y agoHugging Face05w-nicole /childes_datatext1M<n<10M0 likes256 downloads5y agoHugging Face06w-nicole /childes_data_with_tags_text1M<n<10M0 likes248 downloads5y agoHugging Face07BabyLM-community /formatted-CHILDEStext10K<n<100K0 likes237 downloads1y agoHugging Face08w-nicole /childes_data_no_tags_text1M<n<10M0 likes202 downloads5y agoHugging Face09w-nicole /childes_data_with_tagstext1M<n<10M0 likes194 downloads5y agoHugging Face10YasiruR /ChildesForwSlashaudion<1K0 likes136 downloads4y agoHugging Face11superBigPigeon /IPA-CHILDES IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/superBigPigeon/IPA-CHILDES.tabular10M<n<100M0 likes108 downloads8mo agoHugging Face12nshah-fbcs /childes-engUK-conversational-pairs CHILDES Eng-UK Conversational Pairs Curated naturalistic parent-child conversational pairs extracted from the English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test set of 5 complete child histories that no model in the accompanying paper has seen during training. Dataset Summary 278,458 conversation pairs total across train, validation, and test Train: 250,757 pairs from 2,784 transcripts Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.texttext-generation100K<n<1M1 likes97 downloads5mo agoHugging Face13gianjaeger /Childes-OCSC-curated-speech-corpus Childes OCSC Curated Speech Corpus This repository contains a curated subset of recordings and corresponding transcripts from the CHILDES English OCSC Corpus. Contents The data is organized by age groups: 4y/ - 4-year-old speakers 5y/ - 5-year-old speakers 6y/ - 6-year-old speakers 7y/ - 7-year-old speakers 8y/ - 8-year-old speakers 9y/ - 9-year-old speakers Each folder contains audio recordings paired with their transcripts. NOTE: It is incomplete, there are more… See the full description on the dataset page: https://huggingface.co/datasets/gianjaeger/Childes-OCSC-curated-speech-corpus.audio0 likes83 downloads4mo agoHugging Face14wonderwind271 /CHILDES-raw Dataset Card for "CHILDES-raw" More Information needed text10M<n<100M0 likes73 downloads2y agoHugging Face15YasiruR /Childestaudion<1K0 likes58 downloads4y agoHugging Face16suchirsalhan /MAO-CHILDEStext1M<n<10M0 likes54 downloads2y agoHugging Face17BabyLM-community /formatted-CHILDES-cleantext10K<n<100K0 likes49 downloads1y agoHugging Face18kanishka /childes0 likes39 downloads3mo agoHugging Face19llm-slice /babylm-childes-preprocessedtext1M<n<10M0 likes38 downloads1y agoHugging Face20JudithKalinowski /CHILDES_word2vecAll uploaded word embeddings are trained with Word2Vec on CHILDES data. available languages: nor: Norwegian en: English 0 likes27 downloads3y agoHugging Face21kanishka /childes-do_expected-givetext1M<n<10M0 likes27 downloads3mo agoHugging Face22benoitfavre /childes-fra-picto_nllbAutomatically generated text-pictogram pairs using nllb-200-distilled-600m_text2picto. The original corpus is extracted from BabyLM-community/formatted-CHILDES. License varies per instance, but can be assimilated to cc-by-nc-sa. text100K<n<1M0 likes26 downloads6mo agoHugging Face23kanishka /childes-random-givetext1M<n<10M0 likes23 downloads3mo agoHugging Face24CLAUSE-Bielefeld /childes-tripletsThis dataset contains three text files with 10M lexical tokens of interaction triplets from the English section of CHILDES. The first file was used to train the llamalogue model. If you use this dataset, please cite the following paper: @inproceedings{padovani-etal-2025-dialogue, title = "Dialogue Is Not Enough to Make a Communicative {B}aby{LM} (But Neither Is Developmentally Inspired Reinforcement Learning)", author = "Padovani, Francesca and Bunzeck, Bastian and Ali, Manar and… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/childes-triplets.text1M<n<10M0 likes18 downloads11mo agoHugging Face25bbunzeck /childes-dialogue-tripletstext1M<n<10M0 likes18 downloads1y agoHugging Face26kanishka /childes-equal_do_expected_full-givetext1M<n<10M0 likes15 downloads3mo agoHugging Face27kanishka /childes-equal-givetext1M<n<10M0 likes14 downloads3mo agoHugging Face28YasiruR /ChildesAudioMcAllern0 likes12 downloads4y agoHugging Face29Seed42Lab /childes-pretraintext10K<n<100K0 likes12 downloads1y agoHugging Face30kanishka /childes-equal_do_expected_half-givetext1M<n<10M0 likes11 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.