datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.MAO-CHILDESMultilingual CHILDES Dataset for Pretraining Small Multilingual BabyLMs in Salhan et al (2024)
arxiv.org/abs/2410.22886
childes_data_no_tagsIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.childes_datachildes_data_with_tags_formatted-CHILDESchildes_data_no_tags_childes_data_with_tagsChildesForwSlashIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/superBigPigeon/IPA-CHILDES.childes-engUK-conversational-pairs
CHILDES Eng-UK Conversational Pairs
Curated naturalistic parent-child conversational pairs extracted from the
English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test
set of 5 complete child histories that no model in the accompanying paper
has seen during training.
Dataset Summary
278,458 conversation pairs total across train, validation, and test
Train: 250,757 pairs from 2,784 transcripts
Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.Childes-OCSC-curated-speech-corpus
Childes OCSC Curated Speech Corpus
This repository contains a curated subset of recordings and corresponding transcripts from the CHILDES English OCSC Corpus.
Contents
The data is organized by age groups:
4y/ - 4-year-old speakers
5y/ - 5-year-old speakers
6y/ - 6-year-old speakers
7y/ - 7-year-old speakers
8y/ - 8-year-old speakers
9y/ - 9-year-old speakers
Each folder contains audio recordings paired with their transcripts. NOTE: It is incomplete, there are more… See the full description on the dataset page: https://huggingface.co/datasets/gianjaeger/Childes-OCSC-curated-speech-corpus.CHILDES-raw
Dataset Card for "CHILDES-raw"
More Information needed
ChildestMAO-CHILDESformatted-CHILDES-cleanchildesbabylm-childes-preprocessedCHILDES_word2vecAll uploaded word embeddings are trained with Word2Vec on CHILDES data.
available languages:
nor: Norwegian
en: English
childes-do_expected-givechildes-fra-picto_nllbAutomatically generated text-pictogram pairs using nllb-200-distilled-600m_text2picto.
The original corpus is extracted from BabyLM-community/formatted-CHILDES. License varies per instance, but can be assimilated to cc-by-nc-sa.
childes-random-givechildes-tripletsThis dataset contains three text files with 10M lexical tokens of interaction triplets from the English section of CHILDES.
The first file was used to train the llamalogue model.
If you use this dataset, please cite the following paper:
@inproceedings{padovani-etal-2025-dialogue,
title = "Dialogue Is Not Enough to Make a Communicative {B}aby{LM} (But Neither Is Developmentally Inspired Reinforcement Learning)",
author = "Padovani, Francesca and Bunzeck, Bastian and Ali, Manar and… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/childes-triplets.childes-dialogue-tripletschildes-equal_do_expected_full-givechildes-equal-giveChildesAudioMcAllernchildes-pretrainchildes-equal_do_expected_half-give
