datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.childes_data_no_tagsIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.childes_datachildes_data_with_tags_formatted-CHILDESchildes_data_no_tags_childes_data_with_tagsChildesForwSlashIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/superBigPigeon/IPA-CHILDES.childes-engUK-conversational-pairs
CHILDES Eng-UK Conversational Pairs
Curated naturalistic parent-child conversational pairs extracted from the
English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test
set of 5 complete child histories that no model in the accompanying paper
has seen during training.
Dataset Summary
278,458 conversation pairs total across train, validation, and test
Train: 250,757 pairs from 2,784 transcripts
Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.CHILDES-raw
Dataset Card for "CHILDES-raw"
More Information needed
MAO-CHILDESformatted-CHILDES-cleanbabylm-childes-preprocessedchildes-do_expected-givechildes-fra-picto_nllbAutomatically generated text-pictogram pairs using nllb-200-distilled-600m_text2picto.
The original corpus is extracted from BabyLM-community/formatted-CHILDES. License varies per instance, but can be assimilated to cc-by-nc-sa.
childes-random-givechildes-tripletsThis dataset contains three text files with 10M lexical tokens of interaction triplets from the English section of CHILDES.
The first file was used to train the llamalogue model.
If you use this dataset, please cite the following paper:
@inproceedings{padovani-etal-2025-dialogue,
title = "Dialogue Is Not Enough to Make a Communicative {B}aby{LM} (But Neither Is Developmentally Inspired Reinforcement Learning)",
author = "Padovani, Francesca and Bunzeck, Bastian and Ali, Manar and… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/childes-triplets.childes-dialogue-tripletschildes-equal_do_expected_full-givechildes-equal-givechildes-pretrainchildes-equal_do_expected_half-giveCHILDES-Aligned
[!IMPORTANT]
How to access this dataset: the official public release is hosted by TalkBank at
https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives +
CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there.
This Hugging Face copy is retained gated, for internal use; access requests are
approved manually and general requests may be declined — use the TalkBank release instead.
CHILDES-Aligned: Curated Child-Speech Dataset (BEACON)
English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.childes_phones
Dataset Card for "childes_phones"
More Information needed
ASD_CHILDESchildes_phones_allophoneschildes-rawCHILDES_Asymmetries
