datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChildesForwSlashChildes-OCSC-curated-speech-corpus
Childes OCSC Curated Speech Corpus
This repository contains a curated subset of recordings and corresponding transcripts from the CHILDES English OCSC Corpus.
Contents
The data is organized by age groups:
4y/ - 4-year-old speakers
5y/ - 5-year-old speakers
6y/ - 6-year-old speakers
7y/ - 7-year-old speakers
8y/ - 8-year-old speakers
9y/ - 9-year-old speakers
Each folder contains audio recordings paired with their transcripts. NOTE: It is incomplete, there are more… See the full description on the dataset page: https://huggingface.co/datasets/gianjaeger/Childes-OCSC-curated-speech-corpus.ChildestCHILDES-Aligned
[!IMPORTANT]
How to access this dataset: the official public release is hosted by TalkBank at
https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives +
CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there.
This Hugging Face copy is retained gated, for internal use; access requests are
approved manually and general requests may be declined — use the TalkBank release instead.
CHILDES-Aligned: Curated Child-Speech Dataset (BEACON)
English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.ChildesOneCHILDES_Asymmetrieschildes-ocsc
